Method and device for displaying image-text page, electronic equipment and program product
By obtaining user prompt words and canvas content and using the model to generate posters, the problem of inaccurate understanding of user intentions in the prior art is solved, and the accuracy and user experience of generated results are improved.
Patent Information
- Application Number
- CN202510338586.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-24
AI Technical Summary
The existing generative big models cannot accurately understand the user's intentions when generating posters, resulting in the generation results not meeting user expectations and require multiple rework.
By obtaining the user's prompt words and canvas content in the interface, including the content and layout information of the graphic and text page elements determined by the user, the model is used to generate a poster that meets the user's intentions, and the generation results are displayed in the interface.
It improves the controllability and accuracy of model generation posters, reduces the number of user adjustments, and improves the efficiency of posters generation.
Smart Images

Figure CN120198542A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of computers, and more particularly, to methods, apparatuses, electronic devices, and computer program products for displaying picture - text pages. Background Art
[0002] A picture - text page (such as a poster) is a graphic design work that conveys information, promotional content, or expresses creativity through visual elements. In many fields, posters can be used to attract users' attention, convey a specific theme, promote an event or a product, etc. Through poster elements such as images, colors, texts, and layouts, posters can balance aesthetics and functionality.
[0003] With the development of neural networks, using generative models to generate posters has gradually become a popular application. Through techniques such as text - to - image generation and automated design tools, the layout construction, color matching, and copywriting generation of posters can be achieved, greatly improving the generation efficiency of posters. However, current generative large models also have defects such as being unable to accurately understand the user's intention and requiring multiple reworks. Summary of the Invention
[0004] In an embodiment of the present disclosure, there are provided a method, an apparatus, an electronic device, and a computer program product for displaying a picture - text page.
[0005] In a first aspect of the present disclosure, there is provided a method for displaying a picture - text page. The method includes obtaining a prompt word for guiding a model to generate a picture - text page that conforms to the user's intention. The method further includes obtaining the canvas content in an interface, where the canvas content includes the content and layout information of the picture - text page elements determined by the user, and the picture - text page elements include at least one of text and pictures. In addition, the method further includes displaying the picture - text page in the interface, where the picture - text page is generated by the model based on the prompt word and the canvas content.
[0006] In a second aspect of the present disclosure, there is provided a method for displaying a picture - text page. The method includes receiving a prompt word and canvas content from a client, where the prompt word is used to guide a model to generate a picture - text page that conforms to the user's intention, and the canvas content includes the content and layout information of the picture - text page elements determined by the user in the interface, and the picture - text page elements include at least one of text and pictures. The method further includes generating a picture - text page by the model based on the prompt word and the canvas content. In addition, the method further includes sending the picture - text page to the client so that the client displays the picture - text page in the interface.
[0007] In a third aspect of the present disclosure, a device for displaying a graphic and text page is provided. The device includes a prompt word acquisition module configured to acquire a prompt word for guiding a model to generate a graphic and text page that conforms to the user's intention. The device further includes a canvas content acquisition module configured to acquire the canvas content in the interface, where the canvas content includes the content and layout information of the graphic and text page elements determined by the user, and the graphic and text page elements include at least one of text and pictures. In addition, the device further includes a graphic and text page display module configured to display the graphic and text page in the interface, and the graphic and text page is generated by the model based on the prompt word and the canvas content.
[0008] In a fourth aspect of the present disclosure, a device for displaying a graphic and text page is provided. The device includes a prompt word and canvas content receiving module configured to receive a prompt word and canvas content from a client, where the prompt word is used to guide a model to generate a graphic and text page that conforms to the user's intention, the canvas content includes the content and layout information of the graphic and text page elements determined by the user in the interface, and the graphic and text page elements include at least one of text and pictures. The device further includes a graphic and text page generation module configured to generate a graphic and text page by the model based on the prompt word and the canvas content. In addition, the device further includes a graphic and text page sending module configured to send the graphic and text page to the client so that the client displays the graphic and text page in the interface.
[0009] In a fifth aspect of the present disclosure, an electronic device is provided. The electronic device includes a processor. The electronic device further includes a memory coupled to the processor, and the memory has instructions stored therein, and when the instructions are executed by the processor, the electronic device is caused to execute the method according to the first aspect or the second aspect of the present disclosure.
[0010] In a sixth aspect of the present disclosure, a computer program product is provided, which includes computer-executable instructions, where the computer-executable instructions are executed by a processor to implement the method according to the first aspect or the second aspect of the present disclosure.
[0011] In a seventh aspect of the present disclosure, a computer-readable storage medium is provided. Computer-executable instructions are stored on the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, the method according to the first aspect or the second aspect is implemented.
[0012] The summary of the invention is to introduce the selection of concepts in a simplified form, which will be further described in the following detailed implementation. The summary of the invention is not intended to identify the key features or main features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Brief Description of the Drawings
[0013] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0014] Figure 1 A schematic diagram of an example environment in which some embodiments of the present disclosure can be implemented is shown;
[0015] Figure 2 A flowchart of a method for displaying a graphic and text page according to some embodiments of the present disclosure is shown;
[0016] Figure 3A A schematic diagram of a first interface for prompt enhancement according to some embodiments of the present disclosure is shown;
[0017] Figure 3B A schematic diagram of a second interface for prompt enhancement according to some embodiments of the present disclosure is shown;
[0018] Figure 4A A schematic diagram of a first interface for style enhancement according to some embodiments of the present disclosure is shown;
[0019] Figure 4B A schematic diagram of a second interface for style enhancement according to some embodiments of the present disclosure is shown;
[0020] Figure 5A A schematic diagram of a first interface for style enhancement according to some other embodiments of the present disclosure is shown;
[0021] Figure 5B A schematic diagram of a second interface for style enhancement according to some other embodiments of the present disclosure is shown;
[0022] Figure 6A A schematic diagram of a first interface for adding text to a canvas according to some embodiments of the present disclosure is shown;
[0023] Figure 6B A schematic diagram of a second interface for adding text to a canvas according to some embodiments of the present disclosure is shown;
[0024] Figure 6C A schematic diagram of a third interface for adding text to a canvas according to some embodiments of the present disclosure is shown;
[0025] Figure 7A A schematic diagram of a first interface for adding an image to a canvas according to some embodiments of the present disclosure is shown;
[0026] Figure 7B A schematic diagram of a second interface for adding an image to a canvas according to some embodiments of the present disclosure is shown;
[0027] Figure 7C Schematic diagram of the third interface for adding pictures to the canvas in some embodiments of the present disclosure;
[0028] Figure 8A Schematic diagram of the first interface for adding poster templates to the canvas in some embodiments of the present disclosure;
[0029] Figure 8B Schematic diagram of the second interface for adding poster templates to the canvas in some embodiments of the present disclosure;
[0030] Figure 8C Schematic diagram of the third interface for adding poster templates to the canvas in some embodiments of the present disclosure;
[0031] Figure 9A Schematic diagram of the first interface for uploading posters to the canvas in some embodiments of the present disclosure;
[0032] Figure 9B Schematic diagram of the second interface for uploading posters to the canvas in some embodiments of the present disclosure;
[0033] Figure 9C Schematic diagram of the third interface for uploading posters to the canvas in some embodiments of the present disclosure;
[0034] Figure 10A Schematic diagram of the first interface for uploading posters to the canvas in some other embodiments of the present disclosure;
[0035] Figure 10B Schematic diagram of the second interface for uploading posters to the canvas in some other embodiments of the present disclosure;
[0036] Figure 10C Schematic diagram of the third interface for uploading posters to the canvas in some other embodiments of the present disclosure;
[0037] Figure 11A Schematic diagram of the first interface for displaying poster results in some embodiments of the present disclosure;
[0038] Figure 11B Schematic diagram of the second interface for displaying poster results in some embodiments of the present disclosure;
[0039] Figure 12 Schematic diagram of the process for generating posters in some embodiments of the present disclosure;
[0040] Figure 13A block diagram of an apparatus for displaying a graphic page according to some embodiments of the present disclosure; and
[0041] Figure 14 A schematic block diagram of an electronic device according to some embodiments of the present disclosure.
[0042] In all the drawings, the same or similar reference numerals denote the same or similar elements. Detailed Description of Specific Embodiments
[0043] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0044] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations.
[0045] For example, when receiving the user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require the acquisition and use of the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0046] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0047] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure, and other manners that meet the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0048] Embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0049] In the description of the embodiments of the present disclosure, the term "including" and its like terms shall be understood as open inclusion, that is, "including but not limited to". The term "based on" shall be understood as "at least partially based on". The term "an embodiment" or "the embodiment" shall be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects, unless otherwise specified. There may also be other explicit and implicit definitions hereinafter.
[0050] To evaluate the quality of a poster, it can generally be done from three aspects: visual appeal, information transmission efficiency, and theme fit. Among them, the main factors affecting visual appeal include poster color matching, picture style (including the style of picture elements and background pictures), and font design, etc. The main factors affecting information transmission efficiency mainly include the accuracy of the copywriting, information grading, and the rationality of the layout of the main and subheadings, etc. For theme fit, its requirements are higher than visual appeal and information transmission efficiency, and a design that resonates needs to be generated in combination with the poster theme. For example, for an environmental protection-themed poster, the background can be selected as a water-scarce desert, a polluted ocean, and endangered animals, etc.
[0051] In the related art, users can generate posters through a generative model. Specifically, the user designs a prompt, and based on functions such as text-to-image of the generative model, a poster is generated in one stop. However, since elements such as the color matching, copywriting, and design (including fonts and backgrounds) of the poster often have strong personal aesthetic factors, the model cannot accurately parse the user's intention based on the prompt, and thus cannot generate a poster that meets the user's intention, resulting in a low poster export rate in the related art. For the posters generated by the model, a large amount of adjustment by the user is often required.
[0052] In the embodiments of the present disclosure, first, a prompt for generating a poster is obtained, and this prompt is used to guide the model to generate a poster that meets the user's intention. Then, the canvas content determined by the user in the interface is obtained, and this canvas content includes the content and layout information of the poster elements (including text and pictures) set by the user. Further, the model generates a poster based on the prompt and the canvas content, and displays the poster in the interface. In this way, the richness of intention description can be improved by combining the prompt provided by the user and the canvas content, thereby reducing the pressure for the model to infer the user's intention. Thus, the controllability and accuracy of the model for generating posters can be improved.
[0053] Figure 1 FIG. shows a schematic diagram of an exemplary environment 100 in which some embodiments of the present disclosure can be implemented. Refer to Figure 1, the example environment 100 includes a client 102 and a server 110. In some embodiments, the client 102 can be set as devices such as a personal computer, a mobile phone, a tablet computer, etc. The server 110 can be set as a computing system, a single server, a distributed server, or a cloud-based server, etc. Communication can be carried out between the client 102 and the server 110 in a wired or wireless manner. In some embodiments, a model 112, such as a generative large language model, etc., is deployed in the server 110.
[0054] In some embodiments, the user 114 can interact with the client 102 through input-output devices. First, the user 114 can input a prompt 104 into the client 102 through input devices such as a screen and a keyboard. The prompt 104 can be used to guide the model 112 to generate a poster that meets the user's intention (which can be referred to as a graphic page. It should be understood that the graphic page in the present disclosure can also include a cover image, a software interface, etc., and is not limited to the poster in this embodiment).
[0055] Then, the client 102 can display a canvas in the interface for the user to initially draw a poster, such as the text content, picture content, layout, etc. in the poster. After the user completes the initial drawing in the canvas, the client 102 can obtain the canvas content 106 in the canvas. In some embodiments, the canvas content 106 includes the content and layout information of poster elements (which can be referred to as graphic page elements, including text and pictures). The content of the poster elements includes specific picture objects and text descriptions, etc. The layout information of the poster elements is used to describe the typesetting of the poster elements in the poster, and it can include the positions of the poster elements, etc.
[0056] Furthermore, after obtaining the prompt 104 and the canvas content 106, the client 102 sends the prompt 104 and the canvas content 106 to the server 110. After receiving the prompt 104 and the canvas content 106, the server 110 processes the prompt 104 and the canvas content 106 through the model 112, thereby generating a corresponding poster 108 and sending it to the client 102. After receiving the poster 108, the client 102 displays the poster 108 to the user 114 in the interface.
[0057] In an embodiment of the present disclosure, first, a prompt 104 for generating a poster is obtained. The prompt 104 is used to guide the model 112 to generate a poster 108 that conforms to the user's intention. Then, the canvas content 106 determined by the user 114 in the interface is obtained. The canvas content 106 includes the content and layout information of the poster elements (including text and pictures) set by the user 114. Further, the model 112 generates the poster 108 based on the prompt 104 and the canvas content 106, and the poster 108 is displayed in the interface. In this way, the richness of the intention description can be improved by combining the prompt 104 provided by the user 114 and the canvas content 106, thereby reducing the pressure on the model 112 to infer the user's intention. Thus, the controllability and accuracy of the model 112 in generating the poster 108 can be improved.
[0058] The following will be combined with Figures 2 to 14 The process according to the embodiments of the present disclosure will be described in detail. For ease of understanding, the specific data mentioned in the following description are all exemplary and are not used to limit the protection scope of the present disclosure. It should be understood that the embodiments described below may also include additional actions not shown and / or actions shown may be omitted, and the scope of the present disclosure is not limited in this regard.
[0059] Figure 2 The flowchart of a method 200 for displaying a graphic page according to some embodiments of the present disclosure is shown. In some embodiments, the method 200 for displaying a graphic page may be executed by the Figure 1 client 102 therein. It should be understood that the client 102 may include devices such as a personal computer, a mobile phone, a tablet computer, etc.
[0060] At 202, a prompt is obtained. The prompt is used to guide the model to generate a graphic page that conforms to the user's intention. In some embodiments, the user 114 may input the prompt 104 in the client 102 through input devices such as a screen and a keyboard. The prompt 104 may be used to guide the model 112 to generate a graphic page that conforms to the user's intention, such as a poster. For the sake of illustration, the graphic page will be described by taking a poster as an example below, but the graphic page of the present disclosure is not limited to a poster, and it may also be various visual combination pages of text and pictures such as a cover picture and a software interface.
[0061] At 204, the canvas content in the interface is obtained. The canvas content includes the content and layout information of the graphic and text page elements determined by the user, and the graphic and text page elements include at least one of text and pictures. In some embodiments, the client 102 may display a canvas in the interface, and the user initially draws a poster on the canvas. After the user completes the initial drawing in the canvas, the client 102 may obtain the canvas content 106 in the canvas. In some embodiments, the canvas content 106 includes the content and layout information of poster elements (which may be referred to as graphic and text page elements, including text and pictures). The content of the poster elements includes specific picture objects, text descriptions, etc., and the layout information of the poster elements is used to describe the typesetting of the poster elements in the poster, which may include the positions of the poster elements, etc.
[0062] At 206, a graphic and text page is displayed in the interface. The graphic and text page is generated by the model based on the prompt words and the canvas content. In some embodiments, after obtaining the prompt words 104 and the canvas content 106, the client 102 sends the prompt words 104 and the canvas content 106 to the server 110. After receiving the prompt words 104 and the canvas content 106, the server 110 processes the prompt words 104 and the canvas content 106 through the model 112, thereby generating the corresponding poster 108 and sending it to the client 102. After receiving the poster 108, the client 102 displays the poster 108 to the user 114 in the interface.
[0063] In the embodiments of the present disclosure, first, the prompt words for generating the poster are obtained, and the prompt words are used to guide the model to generate a poster that meets the user's intention. Then, the canvas content determined by the user in the interface is obtained, and the canvas content includes the content and layout information of the poster elements (including text and pictures) set by the user. Further, the model generates a poster based on the prompt words and the canvas content, and the poster is displayed in the interface. In this way, the richness of the intention description can be improved by combining the prompt words and the canvas content provided by the user, thereby reducing the pressure for the model to infer the user's intention. Thus, the controllability and accuracy of the model for generating posters can be improved.
[0064] Figure 3A A schematic diagram of the first interface 300A for prompt word enhancement according to some embodiments of the present disclosure is shown. Figure 3B A schematic diagram of the second interface 300B for prompt word enhancement according to some embodiments of the present disclosure is shown. Both the interface 300A and the interface 300B are interfaces in the client 102. In some embodiments, the interface 300A includes an operation bar 302, a function selection area 304, a prompt word input box 306, a custom layout button 308, a style enhancement button 310, a style selection area 312, a recommendation control 314, an upload poster control 316, a poster selection area 318, a generation control 320, a canvas 322, and a prompt word enhancement button 324.
[0065] In some embodiments, the operation bar 302 includes a plurality of operation controls for implementing a plurality of operations on the interface 300A. From left to right in the operation bar 302 are: a return control, where the client 102 can return to the previous interface after the user clicks; an undo control, where the client 102 can undo the previous operation in the interface 300A after the user clicks; a redo control, where the client 102 can, after the user clicks, redo the undone operation after undoing an operation once; a zoom-in control, where the client 102 can, after the user clicks, zoom in on the interface 300A to the entire screen of the client 102; and a more control, where the client 102 can display other collapsed operation controls after the user clicks.
[0066] In some embodiments, the function selection area 304 includes a plurality of function controls for implementing different functions. From top to bottom in the function selection area 304 are: a background control, where the client 102 can enter the background editing interface after the user clicks; a poster control, where the client 102 can enter the poster editing interface after the user clicks; an image-to-video control, where the client 102 can convert the selected picture of the user into a video after the user clicks; a text addition control, where the client 102 can enter the text addition interface after the user clicks to add text to the canvas 322; an upload control, where the client 102 can enter the picture upload interface after the user clicks to add a picture to the canvas 322; and a history control, where the client 102 can display the previously generated posters after the user clicks.
[0067] In some embodiments, the prompt input box 306 is used to display the text input by the user, and the prompt input box 306 includes a prompt enhancement button 324. The prompt enhancement button 324 includes a closed state (which can be referred to as the first display state, referring to the prompt enhancement button 324 in the interface 300A) and an open state (which can be referred to as the second display state, referring to the prompt enhancement button 324 in the interface 300B). When the prompt enhancement button 324 is in the closed state, the client 102 can switch the prompt enhancement button 324 to the open state after the user clicks (i.e., switch from the interface 300A to the interface 300B). When the prompt enhancement button 324 is in the open state, the client 102 can switch the prompt enhancement button 324 to the closed state after the user clicks (i.e., switch from the interface 300B to the interface 300A).
[0068] In some embodiments, after the user enters text in the prompt input box 306, when the prompt enhancement button 324 is in the closed state, the client 102 may not enhance the text entered by the user, and thus directly use the text entered by the user in the prompt input box 306 as the prompt for the model to generate a poster. When the prompt enhancement button 324 is in the open state, the client 102 may use the result of enhancing the text entered by the user as the prompt for the model to generate a poster. In some embodiments, the language model used to enhance the text entered by the user may be deployed either in the client 102 or in the server 110. In this way, the user can choose to enhance the text based on the prompt enhancement control 306 to improve the comprehensiveness and accuracy of the prompt, or choose not to enhance the text to improve the efficiency of generating a poster.
[0069] In some embodiments, the custom layout button 308 also includes an open state and a closed state (shown as the closed state in interfaces 300A and 300B). When the custom layout button 308 is in the open state, the model in the server 110 will generate a poster based on the layout of the poster elements in the canvas 322 by the user. When the custom layout button 308 is in the closed state, the model in the server 110 will not generate a poster based on the layout of the poster elements in the canvas 322 by the user. In this way, the flexibility of the user in setting the layout during the process of generating a poster can be improved.
[0070] In some embodiments, the style enhancement button 310 includes an open state and a closed state (shown as the closed state in interfaces 300A and 300B), and the user can decide whether to select a poster style based on the style button 310. The style selection area 312 displays a plurality of poster styles preset in the client 102. The recommendation control 314 is used to display a plurality of poster templates preset in the client 102, and the plurality of poster templates are displayed in the poster selection area 318. The upload poster control 316 is used to upload the user's poster. After the user completes the drawing and layout of the poster elements in the canvas 322, the generate control 320 can be clicked, so that the server 110 generates a poster and displays it to the user in the client 102. The above-mentioned style enhancement button 310, style selection area 312, recommendation control 314, upload poster control 316, poster selection area 318, generate control 320, canvas 322, and prompt enhancement button 324 will be further described below. Figure 4A A schematic diagram of a first interface 400A for style enhancement according to some embodiments of the present disclosure is shown. Figure 4BA schematic diagram of a second interface 400B for style enhancement according to some embodiments of the present disclosure is shown. Both interface 400A and interface 400B are interfaces in the client 102. In some embodiments, when the style enhancement button 310 is in the off state (which can be referred to as the fourth display state, referring to the style enhancement button 310 in interface 400A), the model in the server 110 does not generate a poster based on the poster style selected by the user (which can be referred to as the graphic page style).
[0071] If the user needs to select a style set in the client 102, the user can click the style enhancement button 310. After the user clicks the style enhancement button 310, the client 102 switches the style enhancement button 310 to the on state (which can be referred to as the third display state, referring to the style enhancement button 310 in interface 400B). At this time, the server 110 will generate a poster based on the poster style selected by the user. In some embodiments, the poster style can include styles such as retro, black and white, and science fiction. By setting the click of the style enhancement button 310, the user can flexibly select whether to generate a poster based on the poster style, thereby improving the matching degree between the poster and the user's needs.
[0072] In some embodiments, when the style enhancement button 310 is in the on state, the client 102 can display a style selection area 312 in the interface 400B. Styles 1, 2, and 3 (which can be collectively referred to as multiple graphic page styles) are displayed in the style selection area 312. The user can click and select the desired style in the style selection area 312. After the user selects one of the styles, for example, style 1, the client 102 can highlight the area corresponding to style 1. After the client 102 determines the style 1 selected by the user, the model in the server 110 can generate a poster based on the prompt words input by the user, the canvas content, and this style 1, and then the client 102 displays the poster.
[0073] Figure 5A A schematic diagram of a first interface 500A for style enhancement according to some other embodiments of the present disclosure is shown, Figure 5B A schematic diagram of a second interface 500B for style enhancement according to some other embodiments of the present disclosure is shown. Both interface 500A and interface 500B are interfaces in the client 102. In some embodiments, the interface 500A and the interface 500B include a model selection control 326 and a style expansion control 328. When the user clicks the model selection control 326, the user can select the model in the server 110 used to generate the poster, such as the poster generation model 1.
[0074] In some embodiments, when the user does not click the style expansion control 328, the client 102 does not display the poster style in the interface 500A. When the user clicks the style expansion control 328, the client 102 displays a style pop-up window 330 in the interface 500B. Styles 1, 2, 3, 4, 5, 6, 7, 8, and 9 (which can be collectively referred to as multiple graphic page styles) are displayed in the style pop-up window 330. When the user selects a desired poster style, such as Style 4, in the style pop-up window 330, the client 102 can highlight Style 4. During the process of generating a poster by the server 110 using the model, the poster will be generated in combination with the user-selected Style 4. In this way, multiple poster styles can be displayed in a pop-up window only when the user needs to select a poster style, thereby improving the simplicity of the poster editing interface.
[0075] Figure 6A A schematic diagram of a first interface 600A for adding text to a canvas according to some embodiments of the present disclosure is shown. Figure 6B A schematic diagram of a second interface 600B for adding text to a canvas according to some embodiments of the present disclosure is shown. Figure 6C A schematic diagram of a third interface 600C for adding text to a canvas according to some embodiments of the present disclosure is shown, where the interfaces 600A, 600B, and 600C are all interfaces in the client 102.
[0076] In some embodiments, the client 102 displays a canvas 322 in the interface, and the user can draw poster elements, including text and pictures, on the canvas 322. When the user needs to insert text, the user can click the add text control in the function selection area 304. At this time, the client 102 can highlight the add text control. After clicking the add text control, the client 102 displays an interface 600A for adding text. In the interface 600A, there is an add a text control 332, a text style selection area 334, and a text style selection area 336. Among them, the text style selection area 334 shows the text styles that the user has recently used, including Text Style 1, Text Style 2, Text Style 3, and Text Style 4. The text style selection area 336 includes Text Style 5, Text Style 6, Text Style 7, Text Style 8, Text Style 9, Text Style 10, Text Style 11, and Text Style 12. The user can click in the text style selection area 334 and the text style selection area 336 to select the desired text style.
[0077] In some embodiments, when the user clicks to add a text control 332 (which may be referred to as a text insertion control), at this time the client 102 switches from the interface 600A to the interface 600B. The client 102 can display a text bounding box 338 in the canvas 322, and then the user can enter the text to be displayed in the poster, such as "Happy New Year!", in the text bounding box 338. In some embodiments, when the text bounding box 338 is displayed in the interface 600B, the text is in an editable state, so the client 102 can display a text toolbar 340 in the interface 600B. The text toolbar 340 includes multiple text operation controls for implementing multiple operations on the text. For example, in the text toolbar 340, from left to right, there may be a font control (currently displayed as regular script), a font size control (currently displayed as 299), a bold control, an italic control, and an underline control in sequence.
[0078] When the user needs to move the text position to adjust the canvas layout, the text bounding box 338 can be dragged to any position in the canvas 322. In addition, in the editing state, the user can also perform operations such as color adjustment, shadow, stroke, and bending on the text. When the user no longer needs to edit the text in the text bounding box 338, the editing state of the text can be exited (for example, by clicking on any blank area in the canvas 322). At this time, the client 102 no longer displays the text bounding box 338 and the text toolbar 340, thereby switching the display interface 600B to the display interface 600C. In this way, the efficiency of inserting text and editing text in the canvas can be improved, and thus the efficiency of generating posters can be improved.
[0079] Figure 7A A schematic diagram of the first interface 700A for adding a picture to the canvas according to some embodiments of the present disclosure is shown. Figure 7B A schematic diagram of the second interface 700B for adding a picture to the canvas according to some embodiments of the present disclosure is shown. Figure 7C A schematic diagram of the third interface 700C for adding a picture to the canvas according to some embodiments of the present disclosure is shown, where the interfaces 700A, 700B, and 700C are all interfaces in the terminal 102.
[0080] In some embodiments, the client 102 displays a canvas 322 in the interface, where the user can draw poster elements, including text and pictures. When the user needs to upload and insert a picture, the user can click the upload control in the function selection area 304. At this time, the client 102 can highlight the upload control. After clicking the upload control, the client 102 displays an interface 700A for uploading pictures. In the interface 700A, there is an added picture insertion control 342 and a picture selection area 344. Among them, the pictures shown in the picture selection area 344 are the pictures that the user has recently uploaded, including Picture 1, Picture 2, Picture 3, and Picture 4. If there is a picture that the user needs in the picture selection area 344, the user can directly click and select the picture in the picture selection area 344.
[0081] In some embodiments, after the user clicks the picture insertion control 342, the user can select the picture to be inserted into the canvas 322. After the user selects the picture to be inserted, the client 102 can display the picture (such as the car shown in the interface 700B) and a picture bounding box 346 in the canvas 322. At this time, the client 102 switches from the interface 700A to the interface 700B. In some embodiments, when the picture bounding box 346 is displayed in the interface 700B, the picture is in an editable state. Therefore, the client 102 can display a picture toolbar 348 in the interface 700B. The picture toolbar 348 includes multiple picture operation controls for implementing multiple operations on the picture. For example, in the picture toolbar 348, from left to right, there may be a brush control, a horizontal flip control, a vertical flip control, and a high-definition control in sequence.
[0082] When the user needs to move the position of the picture to adjust the canvas layout, the user can drag the picture bounding box 346 to move to any position in the canvas 322. In addition, in the editing state, the user can perform operations such as zooming, resizing, super-resolution, and matting on the picture. When the user no longer needs to edit the picture in the picture bounding box 346, the user can exit the editing state of the picture (such as clicking on any blank area in the canvas 322). At this time, the client 102 no longer displays the picture bounding box 346 and the picture toolbar 348, thereby switching the display interface from 700B to 700C. In this way, the efficiency of inserting pictures and editing pictures in the canvas can be improved, and thus the efficiency of generating posters can be improved.
[0083] Figure 8A A schematic diagram of a first interface 800A for adding a poster template to a canvas according to some embodiments of the present disclosure is shown. Figure 8B A schematic diagram of a second interface 800B for adding a poster template to a canvas according to some embodiments of the present disclosure is shown. Figure 8CSchematic diagram of the third interface 800C for adding a poster template to a canvas, showing some embodiments of the present disclosure, where interfaces 800A, 800B, and 800C are all interfaces in the client 102.
[0084] In some embodiments, the client 102 displays a canvas 322 in the interface, and the user can draw poster elements, including text and pictures, on the canvas 322. When the user needs to edit the canvas using a poster template, the user can click the poster function control in the function selection area 304 to enter the poster editing interface 800A. Some content in the interface 800A has been described in detail above and will not be repeated here. After entering the interface 800A, the user can click the recommendation control 314 (which can be called the graphic page template insertion control), and at this time, the client 102 can display a poster selection area 318 below the recommendation control 314. Posters 1, 2, 3, 4, 5, 6, 7, and 8 (which can be collectively referred to as multiple graphic page templates) are displayed in the poster selection area 318.
[0085] The user can select the required poster template in the poster selection area, such as poster 2. Referring to the interface 800B, when the user clicks poster 2, the client 102 will highlight poster 2. At this time, the client 102 will display the poster elements corresponding to poster 2 on the canvas 322, and the bounding boxes corresponding to the poster elements will also be displayed on the canvas 322, including the text bounding box 350 corresponding to the text and the picture bounding box 352 corresponding to the picture. When the text bounding box 350 and the picture bounding box 352 are displayed, the text and the picture are in an editable state. When the user exits the editing state (such as clicking on any blank position in the canvas 322), the client 102 no longer displays the text bounding box 350 and the picture bounding box 352, and at this time, it switches to the interface 800C. In this way, the user can directly set the content and layout of the canvas based on the poster templates set in the client 102, thereby improving the efficiency of the user in drawing poster elements on the canvas and further improving the efficiency of poster generation.
[0086] In some embodiments, referring to the interfaces 800B and 800C, when the user selects poster 2, the client 102 can determine the prompt words corresponding to poster 2 and display the prompt words in the prompt word input box, such as "The theme is summer style, including two cars side by side, and the text is above the cars" as shown in the interfaces 800B and 800C. In some embodiments, the prompt words can be pre-stored prompt words corresponding to poster 2. Alternatively or additionally, the prompt words can be prompt words generated by real-time parsing based on poster 2. In this way, the efficiency of the user in setting the prompt words is improved, thereby improving the efficiency of generating posters.
[0087] Figure 9A A schematic diagram of the first interface 900A for uploading a poster in a canvas, showing some embodiments of the present disclosure. Figure 9B A schematic diagram of the second interface 900B for uploading a poster in a canvas, showing some embodiments of the present disclosure. Figure 9C A schematic diagram of the third interface 900C for uploading a poster in a canvas, showing some embodiments of the present disclosure. Among them, the interfaces 900A, 900B, and 900C are all interfaces in the client 102.
[0088] In some embodiments, the client 102 displays a canvas 322 in the interface, and the user can draw poster elements, including text and pictures, in the canvas 322. When the user needs to use a self-uploaded poster to edit the canvas 322, the user can click the poster function control in the function selection area 304 to enter the poster editing interface 900A. Some contents in the interface 900A have been elaborated in detail above and will not be repeated here. After entering the interface 900A, the user can click the upload poster control 316 (which can be called the graphic page upload control) to self-upload a poster as a poster template. In some embodiments, the client 102 can set the format display of the uploaded poster, such as language (Chinese or English) restrictions, size restrictions, etc. The server 110 can parse the poster uploaded by the user. If it does not meet the relevant format, the parsing fails. If it meets the requirements, the poster is parsed into multiple poster elements and drawn into the canvas 322 after adjusting the poster size.
[0089] Referring to the interface 900B, when the user self-uploads the poster x, the first position in the poster selection area 318 changes to display the poster x and is highlighted, and the display positions of the remaining poster templates are sequentially postponed. At this time, the client 102 will display the poster elements corresponding to the poster x in the canvas 322, and the bounding boxes corresponding to the poster elements will also be displayed in the canvas 322, including the text bounding box 354 corresponding to the text and the picture bounding box 356 corresponding to the picture. When the text bounding box 354 and the picture bounding box 356 are displayed, the text and the picture are in an editable state. When the user exits the editing state (for example, clicks on any blank position in the canvas 322), the client 102 no longer displays the text bounding box 354 and the picture bounding box 356, and at this time, it switches to the interface 900C. In this way, the user can directly set the content and layout of the canvas based on the uploaded poster, thereby improving the efficiency of the user in drawing poster elements in the canvas and further improving the generation efficiency of the poster.
[0090] In some embodiments, referring to interfaces 900B and 900C, when the user uploads poster x, client 102 can determine the prompt words corresponding to poster x and display the prompt words in the prompt word input box 306, such as "The theme is the opening of the Automobile City, with a stronger color tone" shown in interfaces 900B and 900C. In some embodiments, the prompt words can be prompt words generated by real-time parsing of poster x. In this way, the efficiency of the user setting prompt words is improved, thereby improving the efficiency of generating posters.
[0091] Figure 10A A schematic diagram of the first interface 1000A for uploading a poster in a canvas according to some other embodiments of the present disclosure is shown. Figure 10B A schematic diagram of the second interface 100B for uploading a poster in a canvas according to some other embodiments of the present disclosure is shown. Figure 10C A schematic diagram of the third interface 100C for uploading a poster in a canvas according to some other embodiments of the present disclosure is shown. Among them, interfaces 1000A, 1000B, and 1000C are all interfaces in client 102.
[0092] In some embodiments, client 102 displays a canvas 322 in the interface, and the user can draw poster elements in the canvas 322, including text and pictures. When the user needs to use a self-uploaded poster to edit the canvas 322, the user can click the poster function control to enter the poster editing interface 1000A. Some contents in interface 1000A have been elaborated in detail above and will not be repeated here. After entering interface 1000A, the user can click the upload poster control 358 (which can be called the graphic page upload control) to self-upload a poster as a poster template. Among them, the upload poster control 358 is displayed in the first position in the poster selection area 358, and the remaining positions are used to display the poster templates set in client 102, including Poster 1, Poster 2, Poster 3, Poster 4, Poster 5, Poster 6, Poster 7, and Poster 8.
[0093] Referring to the reference interface 900B, when the user uploads the poster x by themselves, the upload poster control 358 in the poster selection area 318 changes to display the poster x and highlights it. At this time, the client 102 will display the poster elements corresponding to the poster x on the canvas 322, and the bounding boxes corresponding to the poster elements will also be displayed on the canvas 322, including the text bounding box 354 corresponding to the text and the picture bounding box 356 corresponding to the picture. When the text bounding box 354 and the picture bounding box 356 are displayed, the text and the picture are in an editable state. When the user exits the editing state (for example, clicks on any blank position in the canvas 322), the client 102 no longer displays the text bounding box 354 and the picture bounding box 356, and at this time, it switches to the interface 1000C. In this way, the user can directly set the content and layout of the canvas based on the uploaded poster, thereby improving the efficiency of the user in drawing poster elements on the canvas, and further improving the generation efficiency of the poster.
[0094] In some embodiments, for the multiple poster elements arranged on the canvas 322 through the above embodiments, the user can delimit a candidate area in the canvas 322, such as a rectangular candidate area. Then, the client 102 can determine all the poster elements within the candidate area and display the bounding boxes of the determined poster elements. It should be understood that at this time, the poster elements are in an editable state. If all the poster elements within the candidate area are text, the client 102 can display the text toolbar 340 on the canvas 322. If all the poster elements within the candidate area are pictures, the client 102 can display the picture toolbar 348 on the canvas 322. If the candidate area contains both text and pictures, the client 102 can either not display the text toolbar 340 and the picture toolbar 348 on the canvas 322, or display the text toolbar 340 and the picture toolbar 348 on the canvas 322 at the same time. In this way, the text toolbar 340 and the picture toolbar 348 can be quickly displayed when the user has an editing need, thereby improving the efficiency of the user in drawing poster elements on the canvas.
[0095] Figure 11A A schematic diagram of the first interface 1100A for displaying the poster result according to some embodiments of the present disclosure is shown. Figure 11B A schematic diagram of the second interface 1100B for displaying the poster result according to some embodiments of the present disclosure is shown. Among them, both the interface 1100A and the interface 1100B are interfaces in the client 102. In some embodiments, when the user sets the prompt words, poster style, and canvas content, etc. and clicks the generate control 320, the client 102 can display the interface 1100A including the poster result, including the result selection area 360 and the generate more results control 362.
[0096] The result selection area 360 contains multiple results generated by the model in the server 110, including Result 1, Result 2, Result 3, and Result 4. The user can select the required poster result in the result selection area 360, such as Result 1. In addition, the user can also click the Generate More Results control 362 to generate more poster results. When the user selects Result 1, the client 102 can highlight Result 1 and display Result 1 in the canvas 322.
[0097] In some embodiments, if the client 102 detects a defect in Result 1, such as the text in the interface 1100A overlapping with the background pattern, resulting in unclear text, the client 102 will display the text bounding box 364 corresponding to the text in the canvas 322. And the client 102 will also display the Reproduce Generate Poster Pop-up Window 366 (which can be called a pop-up window for regenerating the graphic page) at the bottom of the canvas 322. When the user clicks the Reproduce Generate Poster Pop-up Window 366, the model in the server 110 will regenerate the poster, and the client 102 will also display the regenerated poster in the canvas 322 again. As shown in the interface 1100B, the text and the background pattern in the regenerated poster no longer overlap, and the defect existing in the original Result 1 is overcome. In this way, the user can conveniently adjust the generated poster, thereby improving the accuracy of the poster.
[0098] In some embodiments, after the client 102 obtains the prompt word and the canvas content through the above embodiments, the client 102 can send the prompt word and the canvas content to the server 110. The server 110 receives the prompt word and the canvas content, and processes the prompt word and the canvas content through the model to generate the corresponding poster. Then, the server 110 sends the generated poster to the client 102, and the client 102 displays the poster in the interface after receiving the poster. The following will be combined with Figure 12 to describe in detail the specific process of the model generating the poster.
[0099] Figure 12 A schematic diagram of a process 1200 for generating a poster according to some embodiments of the present disclosure is shown. In some embodiments, the process 1200 for generating a poster can be executed by Figure 1 the server 110 therein. In some embodiments, the server 110 can be set as a computing system, a single server, a distributed server, or a cloud-based server, etc. In some embodiments, Figure 1 the model 112 shown in Figure 12 can include an intent model 1212 (which can be called a first model) and a generation model 1220 (which can be called a second model) in
[0100] In some embodiments, the server 110 first obtains the prompt 1202, which is input by the user through the client 102 and sent by the client 102 to the server 110, such as the text input through the above-mentioned prompt input box 306. Then, the server 110 enhances the prompt 1202 through the deployed language model 1204 to obtain the enhanced prompt 1206, and updates the prompt 1202 based on the prompt 1206, so as to improve the accuracy and comprehensiveness of the prompt 1202.
[0101] In some embodiments, the server 110 obtains the poster elements 1208, that is, the content of the pictures and texts arranged by the user in the canvas 322 in the above manner. And the server 110 also obtains the canvas control information 1210 (which can be called layout information). In some embodiments, the canvas control information 1210 may include the bounding box position or the center point position of the poster elements 1208. For example, if the canvas control information 1210 is the bounding box position or the center point position, the vertex coordinates of the bounding box or the center point coordinates can be used as its specific content.
[0102] In some embodiments, the server 110 inputs the prompt 1202, the poster elements 1208, and the canvas control information 1210 into the intent model 1212 to generate the picture copy layout 1214 and the background description information 1216. Among them, the intent model 1212 can be, for example, a multimodal large language model for processing graphic and text information. The picture copy layout 1214 mainly includes the attribute information of the poster elements 1208 (that is, pictures and texts), such as the position, size, color, etc. of the pictures. The background description information 1216 is used as a prompt to guide the generation model 1220 to generate the corresponding background. In addition, during the process of the intent model 1212 generating the picture copy layout 1214 and the background description information 1216, other information can also be combined, such as the poster style selected by the user, the canvas ratio size, etc., so as to further improve the accuracy of the generation result.
[0103] In some embodiments, after obtaining the picture copy layout 1214, the server 110 can render the foreground rendering diagram 1218 (which can be called the foreground diagram) based on the picture copy layout 1214. In some embodiments, the poster involves three types of layers, namely the layer corresponding to the text, the layer corresponding to the picture, and the layer corresponding to the background. The foreground rendering diagram 1218 targets the layer corresponding to the text and the layer corresponding to the picture. Then, the server 110 inputs the foreground rendering diagram 1218 and the background description information 1216 into the generation model 1220, and generates the background diagram 1222 through the generation model 1220. In some embodiments, the generation model 1220 can be, for example, a diffusion model. Further, the server 110 fuses the foreground rendering diagram 1218 and the background diagram 1222 to generate the poster 1224.
[0104] In this way, the server 110 can understand the user's intention for the poster layout based on the intention model 122 and the canvas control information 1210, thereby improving the accuracy of the picture and text layout 1214. Moreover, in combination with the generation model 1220, the background of the poster 1224 is generated, thus dividing the generation of the poster 1224 into two parts: foreground and background, further improving the accuracy and quality of the poster.
[0105] In some embodiments, for the foreground rendering diagram 1218 and the background description information 1216, the user can make adjustments in the client 102. For example, if the user modifies any one of the prompt word 1202, the poster element 1208, and the canvas control information 1210, or modifies the poster style requirement, etc., the server 110 regenerates the picture and text layout based on the model 1212 and renders it into a new foreground rendering diagram 1228, and regenerates new background description information 1226. Then, the server 110 inputs the regenerated background description information 1226 and the foreground rendering diagram 1228 into the generation model 1220 to generate the background image. In this way, an adjustment link can be introduced into the model algorithm, thereby improving the efficiency of the user changing the poster.
[0106] In some embodiments, for the background image 1222, the user can also make adjustments in the client 102. For example, the user can modify the prompt word 1202, the poster element 1208, or the canvas control information 1210, or modify the foreground rendering diagram 1218 or the background description information 1216 to obtain a new background image 1230. Then, the server 110 can generate the final poster 1224 based on the foreground rendering diagram 1218 and the new background image 1230. In this way, the accuracy of the background can be improved, and further the accuracy of the poster can be improved.
[0107] In some embodiments, for the training of the intention model 1212, the server 110 can obtain training data for the intention model 1212. Among them, the training data can include pairs of input data (which can be called the first input data) and annotation data. The input data is used to be input into the model when training the intention model 1212, including the prompt word, the content of the poster element, and the canvas control information (including the bounding box position and / or the center point position of the poster element). The annotation data includes the attributes of the pre-prepared poster element and the background description information.
[0108] Then, the server 110 inputs the above input data into the intent model 1212, and the intent model 1212 infers the input data to obtain the output data in the training phase. Among them, the output data includes the attributes of the poster elements and the background description information. Further, the server 110 calculates the difference degree between the output data and the labeled data to obtain the loss of the intent model 1212, and adjusts the parameters in the intent model 1212 based on the feedback of the loss until the loss converges (that is, the difference degree between the output data and the labeled data meets the conditions).
[0109] In some embodiments, for the training of the generation model 1220, the server 110 may obtain the training data for the generation model 1220, including multiple pairs of input data (which may be referred to as the second input data) and labeled background images. Among them, the input data is used to be input into the generation model 1220 during training, and includes the foreground rendering of the poster elements and the background description information.
[0110] Then, the server 110 inputs the above input data into the generation model 1220, and the generation model 1220 infers to obtain the background image output in the training phase. Further, the server 110 calculates the difference degree between the labeled background image and the output background image to obtain the loss of the generation model 1220, and adjusts the parameters in the generation model 1220 based on the feedback of the loss until the loss converges (that is, the difference degree between the labeled background image and the generated background image meets the conditions).
[0111] Figure 13 The block diagram of the device 1300 for displaying the graphic and text page according to some embodiments of the present disclosure is shown. The device 1300 includes a prompt word acquisition module 1302 configured to acquire a prompt word, which is used to guide the model to generate a graphic and text page that meets the user's intent. The device 1300 further includes a canvas content acquisition module 1304 configured to acquire the canvas content in the interface, and the canvas content includes the content and layout information of the graphic and text page elements determined by the user, and the graphic and text page elements include at least one of text and pictures. In addition, the device 1300 further includes a graphic and text page display module 1306 configured to display the graphic and text page in the interface, and the graphic and text page is generated by the model based on the prompt word and the canvas content.
[0112] It can be understood that the device 900 of the present disclosure can achieve at least one of the many advantages that the methods or processes described above can achieve. For example, the device 900 first obtains a prompt word for generating a poster, and this prompt word is used to guide the model to generate a poster that meets the user's intention. Then, it obtains the canvas content determined by the user in the interface, and this canvas content includes the content and layout information of the poster elements (including text and pictures) set by the user. Further, the model generates a poster based on the prompt word and the canvas content, and displays the poster in the interface. In this way, it is possible to combine the prompt word provided by the user and the canvas content to enhance the richness of the intention description, thereby reducing the pressure on the model to infer the user's intention. Thus, the controllability and accuracy of the model in generating posters can be improved.
[0113] Figure 14 FIG. shows a schematic block diagram of an electronic device 1400 according to some embodiments of the present disclosure. The device 1400 may be the device or apparatus described in the embodiments of the present disclosure. As Figure 14 shown, the device 1400 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 1402, which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 1404 or computer program instructions loaded from a storage unit 1416 into a random access memory (RAM) 1406. In the RAM 1406, various programs and data required for the operation of the device 1400 can also be stored. The CPU / GPU 1402, the ROM 1404, and the RAM 1406 are connected to each other through a bus 1408. An input / output (I / O) interface 1410 is also connected to the bus 1408. Although not shown in Figure 14 FIG., the device 1400 may further include a coprocessor.
[0114] A plurality of components in the device 1400 are connected to the I / O interface 1410, including: an input unit 1412, such as a keyboard, a mouse, etc.; an output unit 1414, such as various types of displays, speakers, etc.; a storage unit 1416, such as a disk, an optical disc, etc.; and a communication unit 1418, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1418 allows the device 1400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0115] Each of the methods or processes described above may be performed by the CPU / GPU 1402. For example, in some embodiments, the method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1416. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 1400 via the ROM 1404 and / or the communication unit 1418. When the computer program is loaded into the RAM 1406 and executed by the CPU / GPU 1402, one or more steps or actions of the methods or processes described above may be performed.
[0116] In some embodiments, the methods and processes described above may be implemented as a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure.
[0117] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not construed to be a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0118] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to various computing / processing devices, or may be downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0119] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages and conventional procedural programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit may execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0120] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when the instructions are executed by the processing unit of the computer or other programmable data processing apparatus, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause a computer, a programmable data processing apparatus, and / or other devices to operate in a specific manner, so that the computer-readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0121] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operation steps are executed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.
[0123] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
[0124] Although the present disclosure has been described in terms of specific structural features and / or method logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms for implementing the claims.
Claims
1. A method for displaying a picture and text page, comprising: Obtaining prompt words, wherein the prompt words are used to guide the model to generate a picture and text page that meets the user's intention; Acquire the canvas content in the interface, the canvas content including the content and layout information of the graphic page element determined by the user, the graphic page element including at least one of text and picture; as well as The graphic page is displayed in the interface, and the graphic page is generated by the model based on the prompt word and the canvas content.
2. The method according to claim 1, wherein obtaining the prompt word comprises: Obtaining the text input by the user in the prompt word input box of the interface; as well as Based on the display status of the prompt word button, perform the following actions: When the prompt word button is in a first display state, determining that the text input by the user is the prompt word, the first display state indicates that the text input by the user is not enhanced; or When the prompt word button is in the second display state, the enhanced text is determined to be the prompt word, and the second display state indicates that the text input by the user is enhanced through a language model.
3. The method according to claim 2, further comprising: When the prompt word button is in the first display state, in response to detecting that the user clicks on the prompt word button, switching the prompt word button to the second display state; or When the prompt word button is in the second display state, in response to detecting that the user clicks on the prompt word button, the prompt word button is switched to the first display state.
4. The method according to claim 1, wherein displaying the graphic page in the interface comprises: Displaying multiple graphic page styles in the interface; Determine the graphic page style selected by the user from among the multiple graphic page styles; as well as The graphic page is displayed in the interface, and the graphic page is generated by the model based on the prompt word, the graphic page style selected by the user, and the canvas content.
5. The method according to claim 4, wherein displaying a plurality of graphic page styles in the interface comprises: When the style button is in the third display state, the plurality of graphic page styles are displayed, and the third display state indicates that the graphic page is generated based on the graphic page style selected by the user; And the method further comprises: When the style button is in the fourth display state, in response to detecting that the user clicks the style button, the style button is switched to the third display state, and the fourth display state indicates that the graphic page is not generated based on the graphic page style selected by the user.
6. The method according to claim 4, wherein displaying a plurality of graphic page styles in the interface comprises: In response to detecting that the user clicks on a style expansion control, displaying a style pop-up window in the interface; as well as The multiple graphic page styles are displayed in the style pop-up window.
7. The method according to claim 1, wherein obtaining the canvas content in the interface comprises: Displaying a canvas for drawing a graphic page in the interface; Displaying the graphic page element input by the user in the canvas includes: In response to detecting that the user clicks on a text insertion control, displaying a text bounding box and text input by the user in the text bounding box in the canvas; and / or In response to detecting that the user clicks on the picture insertion control, a picture bounding box and the picture inserted by the user in the picture bounding box are displayed in the canvas.
8. The method according to claim 1, wherein acquiring the canvas content in the interface comprises: Displaying a canvas for drawing a graphic page in the interface; In response to detecting that the user clicks on a graphic page template insertion control, displaying a plurality of graphic page templates in the interface; Determine the graphic page template selected by the user from among the multiple graphic page templates; as well as Displaying the graphic page template selected by the user in the canvas includes: Displaying the text in the graphic page template and a text boundary box of the text in the canvas, wherein the text is in an editable state when the text boundary box is displayed; and / or The picture in the graphic page template and the picture boundary box of the picture are displayed in the canvas, and the picture is in an editable state when the picture boundary box is displayed.
9. The method according to claim 1, wherein obtaining the canvas content in the interface comprises: Displaying a canvas for drawing a graphic page in the interface; In response to detecting that the user clicks on the picture and text page upload control, obtaining the picture and text page uploaded by the user; as well as Displaying the image and text page uploaded by the user in the canvas includes: Displaying the text in the graphic page and a text boundary box of the text in the canvas, wherein the text is in an editable state when the text boundary box is displayed; and / or The picture in the picture-text page and the picture boundary frame of the picture are displayed in the canvas, and the picture is in an editable state when the picture boundary frame is displayed.
10. The method according to claim 8 or 9, wherein obtaining the prompt word comprises: Determining the prompt word based on the graphic page template or the graphic page uploaded by the user; as well as The prompt word is displayed in the prompt word input box of the interface.
11. The method according to any one of claims 7 to 9, further comprising: In response to detecting that the user has delimited a candidate area in the canvas, determining graphic page elements in the candidate area; The bounding box of the graphic page element displayed in the candidate area includes a text bounding box and a picture bounding box; Based on the type of bounding box, do the following: In the case where all the bounding boxes in the candidate area are text bounding boxes, displaying a text toolbar in the interface; or In the case that the bounding boxes in the candidate area are all picture bounding boxes, a picture toolbar is displayed in the interface.
12. The method according to claim 1, further comprising: Displaying a pop-up window for regenerating the graphic page in the interface; as well as In response to detecting that the user clicks on the pop-up window, a regenerated graphic page is displayed in the interface, where the regenerated graphic page is regenerated by the model based on the prompt word and the canvas content.
13. A method for displaying a graphic page, comprising: Receive a prompt word and canvas content from a client, wherein the prompt word is used to guide the model to generate a graphic page that meets the user's intention, and the canvas content includes content and layout information of graphic page elements determined by the user in the interface, and the graphic page elements include at least one of text and pictures; The model generates the graphic page based on the prompt word and the canvas content; as well as The graphic page is sent to the client so that the client displays the graphic page in the interface.
14. The method according to claim 13, wherein the model comprises a first model and a second model, the layout information comprises a bounding box position and / or a center point position of the graphic page element, and generating the graphic page by the model based on the prompt word and the canvas content comprises: The first model generates attribute information and background description information of the image-text page element based on the prompt word, the content of the image-text page element, and the bounding box position and / or the center point position of the image-text page element; Based on the attribute information of the graphic page element, generating a foreground image for the graphic page element; The second model generates a background image based on the foreground image and the background description information; as well as The graphic page is generated based on the foreground image and the background image.
15. The method according to claim 14, wherein generating the background image by the second model based on the foreground image and the background description information comprises: In response to detecting that the user adjusts the foreground image and / or the background description information, acquiring the foreground image and / or the background description information adjusted by the user; as well as The background image is generated by the second model based on the adjusted foreground image and / or background description information.
16. The method according to claim 14, wherein generating the image-text page based on the foreground image and the background image comprises: In response to detecting that the user adjusts the background image, acquiring the background image adjusted by the user; as well as The graphic page is generated based on the foreground image and the adjusted background image.
17. The method according to claim 14, wherein the training process of the first model comprises: Acquire first input data and corresponding annotation data for training the first model, wherein the first input data includes prompt words, content of the graphic page element, and the bounding box position and / or center point position of the graphic page element, and the annotation data includes attributes and background description information of the graphic page element; The first model generates output data based on the first input data, wherein the output data includes attributes and background description information of the generated graphic page element; as well as Based on the labeled data and the output data, parameters in the first model are adjusted so that the difference between the labeled data and the output data satisfies a condition.
18. The method according to claim 14, wherein the training process of the second model comprises: Acquire second input data and a corresponding annotated background image for training the second model, wherein the second input data includes a foreground image and background description information for an element of a graphic page; generating a background image by the second model based on the second input data; as well as Based on the annotated background image and the generated background image, the parameters in the second model are adjusted so that the difference between the annotated background image and the generated background image satisfies a condition.
19. A device for displaying a picture and text page, comprising: A prompt word acquisition module is configured to acquire prompt words, wherein the prompt words are used to guide the model to generate a picture and text page that meets the user's intention; A canvas content acquisition module is configured to acquire canvas content in the interface, wherein the canvas content includes content and layout information of the graphic page element determined by the user, and the graphic page element includes at least one of text and picture; as well as The graphic page display module is configured to display the graphic page in the interface, wherein the graphic page is generated by the model based on the prompt word and the canvas content.
20. A device for displaying a picture and text page, comprising: A prompt word and canvas content receiving module is configured to receive prompt words and canvas content from a client, wherein the prompt words are used to guide the model to generate a graphic page that meets the user's intention, and the canvas content includes the content and layout information of the graphic page elements determined by the user in the interface, and the graphic page elements include at least one of text and pictures; A picture and text page generation module, configured to generate the picture and text page based on the prompt word and the canvas content by the model; as well as The image and text page sending module is configured to send the image and text page to the client, so that the client displays the image and text page in the interface.
21. An electronic device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, wherein when the instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 12 or claims 13 to 18.
22. A computer program product comprising computer executable instructions, wherein the computer executable instructions are executed by a processor to implement the method according to any one of claims 1 to 12 or claims 13 to 18.
Citation Information
Cited By
Method, apparatus, device and product for authoring media data
CN121349347A