Visual object generation method and device, electronic equipment and storage medium
Through the visual object generation method based on large models, the problem of low visual object design efficiency in the prior art is solved, and visual objects such as posters and promotional pictures are automatically generated, which improves design efficiency and content accuracy.
Patent Information
- Application Number
- CN202510315889.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-06
AI Technical Summary
The existing technology is less efficient when designing visual objects such as posters, promotional pictures, invitations, etc., and users need to manually select image materials and adjust layouts and edit copywriting, resulting in large workloads and low efficiency.
A visual object generation method is provided, by determining the target template and area description information, generating area content based on the large model based on the demand information, and combining the template with area content to generate the target visual object.
It reduces the user's workload, improves the efficiency of visual object production, ensures that the generated content is consistent with regional needs, and improves the accuracy of regional content.
Smart Images

Figure CN120107394A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the fields of generative models, large models, etc. More specifically, the present disclosure provides a visual object generation method, device, electronic device, storage medium and computer program product. Background Art
[0002] In work and life, users sometimes need to design visual objects such as posters, promotional images, invitations, etc. Summary of the invention
[0003] The present disclosure provides a visual object generation method, device, electronic device, storage medium and computer program product.
[0004] According to one aspect of the present disclosure, a visual object generation method is provided, including: determining a target template, the target template including at least one area, each area of the at least one area corresponding to area description information, the area description information being used to describe content to be displayed for the area; for each area, based on a large model, generating area content of the area according to demand information and area description information of the area; and combining the target template with the respective area content of at least one area to obtain a target visual object, so as to display the target visual object.
[0005] According to another aspect of the present disclosure, a visual object generation device is provided, comprising: a template determination module, a generation module and a combination module. The template determination module is used to determine a target template, the target template includes at least one region, each region of the at least one region corresponds to region description information, and the region description information is used to describe the content to be displayed for the region. The generation module is used to generate the region content of each region based on the large model, according to the demand information and the region description information of the region. The combination module is used to combine the target template with the respective region content of at least one region to obtain a target visual object, so as to display the target visual object.
[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided by the present disclosure.
[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the method provided by the present disclosure.
[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the method provided by the present disclosure is implemented.
[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0011] Figure 1 is a schematic diagram of an application scenario of the visual object generation method and device according to an embodiment of the present disclosure;
[0012] Figure 2 is a schematic flow chart of a visual object generation method according to an embodiment of the present disclosure;
[0013] Figure 3 is a schematic flow chart of determining a target template according to an embodiment of the present disclosure;
[0014] Figure 4A to Figure 4E is a front-end schematic diagram of a visual object generation method according to an embodiment of the present disclosure;
[0015] Figure 5 is a schematic structural block diagram of a visual object generating device according to an embodiment of the present disclosure; and
[0016] Figure 6 It is a structural block diagram of an electronic device used to implement the visual object generation method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0017] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0018] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0019] In the technical solution of the present disclosure, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0020] In work and life, users sometimes need to design pages such as posters, promotional images, invitations, etc. In some technical solutions, users can manually select images and other materials according to their own needs, and then adjust the layout of each material and edit the copy on the page to create the required page. However, this production method is inefficient.
[0021] The present disclosure aims to provide a visual object generation method, which can automatically generate target visual objects such as posters, promotional pictures, invitations, etc. according to the demand information input by the user, thereby reducing the workload of the user and improving the production efficiency. It can also ensure that the generated content is compatible with the content required by the region, thereby improving the accuracy of the regional content.
[0022] The visual object generation method provided by the present disclosure is suitable for generating H5 (HyperText Markup Language 5, 5th edition) pages such as promotional pictures, posters, invitation letters, conference invitations, and training course registration pages. The generated H5 page can be a single page or a multi-page page, and the page can be a long page.
[0023] The technical solution provided by the present disclosure will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0024] Figure 1 It is a schematic diagram of an application scenario of the visual object generation method and device according to an embodiment of the present disclosure.
[0025] It should be noted that Figure 1 What is shown is merely an example of a system architecture to which the embodiments of the present disclosure can be applied, in order to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0026] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0027] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Terminal devices 101, 102, 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptops, desktop computers, etc.
[0028] The server 105 may be a server that provides various services, such as a background management server (only an example) that provides support for websites browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as the target visual object obtained or generated according to the user request, etc.) to the terminal device.
[0029] It should be noted that the visual object generation method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the visual object generation device provided in the embodiment of the present disclosure can generally be set in the server 105. The visual object generation method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the visual object generation device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.
[0030] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0031] Figure 2 It is a schematic flowchart of a visual object generation method according to an embodiment of the present disclosure.
[0032] like Figure 2 As shown, the visual object generating method 200 may include operations S210 to S230.
[0033] In operation S210, a target template is determined, the target template includes at least one area, each of the at least one area corresponds to area description information, and the area description information is used to describe the content to be displayed for the area.
[0034] For example, the target template can be a basic framework for generating a target visual object. The target template contains at least one area, and each area in the target template is an area for displaying content, such as a title area, a picture area, a text area, etc. The area description information is used to describe the content to be displayed in the area, such as describing that the area is used to display time and place, the area is used to display pictures and text, a sub-area in the area is used to display images and another sub-area is used to display text. For example, an area description information is "This area is used to display activity slogans, and no pictures are required."
[0035] For example, multiple templates can be pre-configured, and then a template that matches the required information (such as activity type, style preference, etc.) can be selected as the target template. The degree of matching can be determined by calculating the similarity between the required information and the template. For another example, a template can be randomly selected from multiple templates and used as the target template.
[0036] In operation S220 , for each region, regional content of the region is generated based on the macro model according to the demand information and the regional description information of the region.
[0037] For example, the user's demand information may be acquired in advance, where the demand information includes, for example, the activity name, the type of visual object to be generated, the style of the target visual object, and the like.
[0038] The requirement information and the area description information can be input into the macro model, and the macro model generates the area content. The area content can be the specific information that needs to be filled into the area in the target template. The area content can include at least one of the area text and the area image. For example, the area content includes text, pictures, icons, etc. For example, for the "activity title" area, a title copy can be generated, and for the "activity details" area, a detailed activity introduction can be generated. For the "background image" area, a background image is generated.
[0039] In operation S230 , the target template is combined with the respective region content of at least one region to obtain a target visual object so as to display the target visual object.
[0040] For example, the generated regional content can be filled into the corresponding area of the target template to form a target visual object. In the process of generating the target visual object, the generated regional content can be filled into the corresponding position according to the predefined regional position information in the target template, for example, the generated title copy is filled into the title area, and the generated picture is filled into the picture area. In addition, if the size or format of the regional content does not match the target template, the size, font, color, etc. of the content can also be adjusted to ensure the coordination of the overall layout.
[0041] This embodiment can first obtain user needs, then determine the target template, generate regional content for each area in the target template, and then combine the target template with the regional content of each area to obtain the target visual object. In this way, the user inputs personal demand information, and the target visual objects such as posters, promotional pictures, invitations, etc. can be automatically generated, which reduces the user's workload, improves production efficiency, and ensures the quality of the generated content through the generation capability of the large model.
[0042] In addition, this embodiment processes each region separately, and uses a large model to generate regional content based on the regional description and demand information during the processing. Since the content generation of each region is based on the description information and demand information of the region, it can ensure that the generated content is compatible with the content required by the region, thereby improving the accuracy of the regional content. And when it is necessary to modify the local content of the template, since each region is independent of each other, it is only necessary to modify the regional description information of the corresponding region, and there is no need to modify the template as a whole, which is more flexible. In addition, in the actual processing process, the regional content of each region can be generated in parallel, which can shorten the overall generation time and further improve the processing efficiency.
[0043] Next, the process of determining the demand information is described.
[0044] In one example, the user can describe his or her needs by voice, text or file upload, such as the required visual object category, usage scenario, key information in the visual object, etc. After obtaining the input information input by the user, the electronic device can perform intent recognition to determine the user's intention, which includes, for example, invitations, birthday wishes, promotional pictures, posters, etc. If the user's intention does not match the pre-configured intention, it means that the user's real intention is irrelevant to the pre-configured intention, and the user can be prompted that the function is not supported, thereby preventing the user from asking topics irrelevant to the generation of visual objects.
[0045] In addition, the correspondence between the intention and the questions to be clarified can be pre-configured, and after determining the user's intention, the questions to be clarified corresponding to the intention can be output through the front end, so that the user can answer the questions by selection, input, etc., and the details of the user's needs can be clarified. For example, if the user's intention is to generate a poster, the user can be asked about the theme and style of the poster. For another example, if the user's intention is to generate an invitation letter, the user can be asked about the content of the invitation letter, such as the address, time, event, invited people, etc. After obtaining the user's answer to the question, the question, answer and other information can be summarized, for example, the question and answer can be combined in a predetermined format, or the question and answer can be summarized using a large model to obtain the demand information, which can be text.
[0046] Figure 3 is a schematic flowchart of determining a target template according to an embodiment of the present disclosure.
[0047] like Figure 3 As shown, next, the method 300 for determining the target template is described. The method 300 may include operations S311 to S317.
[0048] In operation S311, a plurality of candidate framework templates and a plurality of candidate scenes are configured.
[0049] For example, a candidate frame template may include one or more areas, and the layout of each area and the content to be displayed in the area may be pre-configured. For example, a candidate frame template is pre-configured, in which the first area is used to display the activity slogan, the second area is used to display the details, and the bottom of the third area includes the participants and contact information waiting to be collected. The overall style of the candidate frame template is business simplicity. For example, candidate scenes include: children's birthday party scenes, car company annual meeting scenes, etc.
[0050] In operation S312, determine whether each candidate framework template and the requirement information satisfy the first matching condition. For example, the first matching condition may include: the similarity between the framework description information of the candidate framework template and the requirement information is greater than or equal to a first threshold. In this embodiment, both the framework description information and the requirement information may be texts, and the similarity between the two texts is calculated to determine whether the first matching condition is satisfied. The two texts may more accurately represent the features of the framework and the requirement, thereby improving the accuracy of the matching result. For another example, the first matching condition may include the matching degree between the regional content in the candidate framework template of the candidate framework template and the key elements of the requirement information, etc. This embodiment does not limit the first matching condition.
[0051] In operation S313, if at least one candidate framework template satisfies the first matching condition with the requirement information, a target template may be determined from the at least one candidate framework template. For example, the candidate framework template with the highest similarity may be determined as the target template, or a candidate framework template may be randomly selected as the target template from at least one candidate framework template that satisfies the first matching condition. In this embodiment, the target template is first screened by the preset first matching condition, so that the target template with high consistency with the requirement information can be quickly screened out, thereby improving the efficiency and accuracy of determining the target template.
[0052] In operation S314, if multiple candidate framework templates and demand information do not meet the first matching condition, it can be determined whether multiple candidate scenes and demand information meet the second matching condition. For example, the second matching condition may include: the similarity between the scene description information of the candidate scene and the demand information is greater than or equal to the second threshold. In this embodiment, the scene description information and the demand information can both be texts, and the similarity between the two texts is calculated to determine whether the second matching condition is met. The two texts can more accurately represent the characteristics of the scene and the demand, thereby improving the accuracy of the matching results. For another example, the second matching condition may include: the candidate scene corresponds to a candidate area set, the matching degree between the area description information of the candidate area set and the key elements of the demand information, etc. This embodiment does not limit the second matching condition.
[0053] It should be noted that the main difference between the candidate frame template and the scene is that the granularity of the two is different. The granularity of the candidate frame template is smaller and has clearer template information; while the granularity of the candidate scene is larger and more generalized. The scene description information of the candidate scene is more generalized, more vague, and less accurate than the frame description information of the candidate frame template. For example, the frame description information of the candidate frame template includes clear visual object categories (such as children's birthday posters, elderly birthday posters, Internet company promotional posters, and exhibition invitations of automobile manufacturers), clear regional content requirements (for example, the first area is the exhibition title, and the second area is the exhibition event details), whether to include fields to be collected, what kind of fields to be collected are included (including participant fields), etc. For example, the scene description information of the scene may include the category of the scene (for example, one category is a poster, and the other category is an invitation, and there is no need to distinguish what kind of poster it is), the included regional information (for example, including the activity details area, and it is not necessary to distinguish what kind of activity it is).
[0054] In operation S315, if the plurality of candidate scenarios and the demand information do not satisfy the second matching condition, a pre-configured general template may be used as a target template, or a target template may be directly generated according to the demand information based on a large model.
[0055] In operation S316, if at least one candidate scene and the requirement information satisfy the target scene of the second matching condition, the target scene may be determined from the at least one candidate scene. For example, the candidate scene with the highest similarity may be determined as the target scene, or a candidate scene may be randomly selected from at least one candidate scene satisfying the second matching condition as the target scene.
[0056] In operation S317, after the target scene is obtained, a target template may be determined according to a candidate region set corresponding to the target scene. The candidate region set includes a plurality of regions, and each region may have a pre-configured order.
[0057] For example, the demand information, the area description information of each candidate area in the candidate area set, and the predetermined prompt information can be input into the large model to obtain the target template, wherein the predetermined prompt information is used to guide the large model to screen areas from the candidate area set, and combine the screened areas into the target template, or use the screened areas to generate the target template. In this embodiment, the target scene that matches the demand information is first determined, and then the candidate areas are screened from the candidate area set corresponding to the target scene to obtain the target template. On the one hand, the target scene can be used to make each area in the target template meet the demand information, thereby improving the accuracy of the target template. On the other hand, the large model screens the candidate areas, so that different target templates can be generated for different demand information matching the same target scene, thereby improving the diversity of the target template.
[0058] For another example, the candidate region set corresponding to the target scene may not be screened, but the candidate region set may be used as a target template.
[0059] This embodiment adopts a multi-level matching mechanism, namely, matching of candidate framework templates and scene matching. By matching the similarity between the candidate framework templates and the demand information, the overall more suitable target template can be quickly determined to reduce the computational complexity. If the matching fails, the target template is further determined by scene matching and screening of the candidate area set. In this way, even if the database lacks a candidate framework template that completely meets the demand information, the candidate area set that is more in line with the demand information can still be determined according to the scene. If the scene cannot be successfully matched, the reasoning ability of the large model can be combined to dynamically generate an adapted target template. The above-mentioned multi-level matching mechanism is applicable to a variety of demand information, so that the target template has a high degree of consistency with various demand information.
[0060] It should be noted that the above embodiment adopts a multi-level matching mechanism, where the first-level matching mechanism is to determine whether the candidate framework template and the requirement information meet the first matching condition, and the other-level matching mechanism is to determine whether the candidate scene and the requirement information meet the second matching condition. In other embodiments, only a single-level matching mechanism may be adopted, for example, only any one of the above matching mechanisms is adopted, while omitting another matching mechanism.
[0061] Next, the process of generating the area content of an area is described.
[0062] In this embodiment, the target template includes multiple regions, each of which can be processed independently to generate regional content. The category of the region can be pre-configured, and the category field can be pre-configured in the description information of the region. The category of the region includes image region, text region, and graphic region.
[0063] In one example, the category of the region is an image region, and this type of region only needs to generate a region image, and does not need to generate a region text. The demand information, the region description information of the region, and the pre-configured prompt information can be combined and then input into the large model, and the large model generates a region image. The prompt information can guide the large model to determine whether to generate an image based on the region description information of the current region, and if necessary, generate a region image. The region image can be a background image, a real scene image, a decorative image, etc.
[0064] In another example, the category of the region is a text region, and this type of region only needs to generate region text, and does not need to generate region images. The demand information, the region description information of the region, and the pre-configured prompt information can be combined and then input into the large model, and the large model generates the region text. The prompt information can guide the large model to determine whether to generate copywriting based on the region description information of the current region, and generate region text if necessary. The large model can generate a summary based on the demand information and use the summary as the region text.
[0065] In another example, the category of the area is a graphic and text area, and this type of area needs to generate a regional text and a regional image. Based on the large model, the regional text in the regional content can be generated according to the demand information and the regional description information of the area, and then based on the large model, the regional image in the regional content can be generated according to the demand information, the regional description information of the area and the regional text. For example, the demand information, the regional description information of the area and the pre-configured prompt information are combined, and then input into the large model, and the regional text is generated by the large model. Then the demand information, the regional description information of the area, the regional text and the pre-configured prompt information are combined, and then input into the large model, and the regional image is generated by the large model. In this example, for graphic and text areas, after the regional text is generated, the regional image is generated based on the regional text, so that the regional text can guide the generation process of the regional image, thereby improving the consistency of the text and image of the regional content.
[0066] According to another embodiment of the present disclosure, sometimes it is necessary to collect information of users who browse the target visual object. For example, when the target visual object is a conference invitation or invitation letter, the name, contact information, etc. of the visitor can be collected. When the target visual object is a promotional page, the contact information of the person of interest can be collected. Therefore, it is necessary to generate the name, contact information and other collection fields in the target visual object. Next, the generation process of the fields to be collected is explained. It should be noted that the users are aware of and agree to the acquisition and use of user information, and they are in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0067] In one example, if the target template is a frame template, and the target template does not include elements to be collected, there is no need to generate fields to be collected. It should be noted that the layout and area of the target template can be pre-configured. If the target template contains multiple areas, and the area description information of each interval does not describe the area as an area where information needs to be submitted, it can be determined that the target template does not include elements to be collected. At this time, it means that there is no need to collect information for users browsing the target visual object, and therefore, there is no need to generate fields to be collected.
[0068] In another example, if the target template is a candidate framework template, and the target template includes elements to be collected, the fields to be collected can be generated based on the large model and the demand information. For example, for target templates of the type of invitation letters, conference invitations, etc., the description information of a certain area in the target template characterizes the area as an area where information needs to be submitted. The demand information can be input into the large model, and the large model extracts the fields to be collected from the demand information. For example, a user inputs the demand information of "generate an invitation letter, and count the names and contact information of the visitors" through the front end, then the fields to be collected extracted from the demand information by the large model may include "name" and "contact information". The fields to be collected can then be displayed at a predetermined position in the area so that users browsing the invitation letter can fill in the reply information of the fields to be collected according to their own needs.
[0069] In this example, when the target template is a candidate framework template, by judging whether the target template contains elements to be collected, it is dynamically determined whether to generate fields to be collected, thus avoiding unnecessary field addition. When fields to be collected need to be added, the large model is used to automatically extract the fields to be collected from the demand information, ensuring that the fields to be collected meet the demand information, thus improving the accuracy and flexibility of the fields to be collected.
[0070] In another example, if the target template is determined based on the target scenario, it can be determined whether information needs to be collected based on at least one of the target scenario and the demand information, and if it is determined that information needs to be collected, based on the large model, the fields to be collected are generated according to the demand information and the reference fields associated with the target scenario. For example, it is pre-configured whether each scenario needs to collect information, such as configuring scenes such as invitations, and scenes such as posters and promotional pages do not need to collect information. Therefore, it can be determined whether information needs to be collected based on the configuration information. In addition, the user can express whether information needs to be collected and what information needs to be collected through demand information. Therefore, the demand information can be processed by using large models, regular matching, keyword extraction, etc. to determine whether information needs to be collected. If information needs to be collected, the association relationship between the scene and the reference field can be pre-configured, and then the reference field associated with the target scene can be retrieved from the database based on the association relationship, and then the reference field, demand information and prompt information are input into the large model. The prompt information can be used to guide the large model to preferentially extract the fields to be collected from the demand information. If the field cannot be extracted from the demand information, the reference field is used to determine the field to be collected.
[0071] In this example, for the case where the target scenario is used to determine the target template, determine whether information needs to be collected, and generate the fields to be collected if necessary. If information does not need to be collected, there is no need to call the large model, which can reduce the number of large model calls and reduce data processing costs. In addition, in the process of generating the fields to be collected, it can be generated based on the demand information first, so that the fields to be collected can better meet the user's requirements. If the demand information cannot be extracted to the fields to be collected, then use the reference field to generate them, which can ensure that the generated fields to be collected are highly relevant to the scenario.
[0072] According to another embodiment of the present disclosure, after obtaining the target template, the regional content of each area and the fields to be collected, these data can be combined to obtain the target visual object. Exemplarily, the regional content of at least one area can be filled into the target template based on the location information of at least one area, and the fields to be collected can also be filled into the predetermined positions in the target template, and then the target visual object can be determined based on the filled target template.
[0073] For example, if the target template is a candidate frame template, the position information of each area in the candidate frame template can be pre-configured, and then the area content can be filled into the corresponding area according to the position information. For another example, if the target template is determined based on a set of candidate areas corresponding to the target scene, the size and arrangement order of each area can be pre-configured, so that after screening the candidate areas, the position information of each area can also be determined, and then the area content can be filled into the corresponding area according to the position information. For another example, the position of the field to be collected can be pre-configured. For example, the field to be collected is usually at the bottom of the target visual object such as posters and invitations, and the field to be collected can be filled into the predetermined position. The target template after filling the area content and the field to be collected can be determined as the target visual object.
[0074] This embodiment pre-configures the target template's regional location information, size, arrangement order, and the predetermined location of the fields to be collected, so that the regional content and the fields to be collected can be efficiently and accurately filled, thereby quickly generating the target visual object. In addition, this data filling method can ensure that the generated target visual object meets expectations in terms of structure, layout, and content, thereby improving generation efficiency and quality.
[0075] Next, the visual object generation method provided in this embodiment is explained from the front end.
[0076] like Figure 4AAs shown, the user can input information through the front-end page Page_1, for example, input "create a birthday-themed electronic card". Next, the back-end can perform intent recognition on the input information, for example, recognize that the user's intent is a birthday card, and then, based on the correspondence between the pre-configured intent and the question to be clarified, the user can be asked to input questions to be clarified through the front-end page Page_1. For example, questions and options about the style of the greeting card are output.
[0077] like Figure 4B As shown, the user can select the desired style through the front-end page Page_2, such as "warm and cute". Next, the back-end receives the user's response and can continue to output the next question to be clarified through the front-end page Page_2, such as the question and option of whether information needs to be collected.
[0078] like Figure 4C As shown, the user can choose whether to collect information through the front-end page Page_3, for example, select "need". Next, the back-end receives the user's response and can continue to output the next question to be clarified through the front-end page Page_3, such as outputting questions about the fields to be collected.
[0079] like Figure 4D As shown, the user can input the fields to be collected through the front-end page Page_4. After the user clicks the "Generate Application" option, the back-end can generate the target visual object based on the user's input information. The process of generating the target visual object can be referred to above and will not be repeated in this embodiment.
[0080] like Figure 4E As shown, after the target visual object is generated, the target visual object can be displayed through the front-end page Page_5.
[0081] Figure 5 It is a schematic structural block diagram of a visual object generating device according to an embodiment of the present disclosure.
[0082] like Figure 5 As shown, the visual object generating apparatus 500 may include a template determining module 510 , a generating module 520 and a combining module 530 .
[0083] The template determination module 510 is used to determine a target template. The target template includes at least one area. Each area of the at least one area corresponds to area description information. The area description information is used to describe the content to be displayed for the area.
[0084] The generation module 520 is used to generate the regional content of each region based on the large model, according to the demand information and the regional description information of the region.
[0085] The combining module 530 is used to combine the target template with the respective regional content of at least one region to obtain a target visual object so as to display the target visual object.
[0086] According to another embodiment of the present disclosure, it further includes: a target scene determination module and a first template determination module. The target scene determination module is used to determine a target scene that meets a second matching condition with the requirement information from a plurality of candidate scenes in response to determining that none of the plurality of candidate framework templates meets a first matching condition with the requirement information. The first template determination module is used to determine a target template based on a candidate region set corresponding to the target scene.
[0087] According to another embodiment of the present disclosure, the target template determination module includes: an input submodule, used to input demand information, area description information of each candidate area in the candidate area set, and predetermined prompt information into the large model to obtain the target template; wherein the predetermined prompt information is used to guide the large model to filter areas from the candidate area set, and combine the filtered areas into the target template.
[0088] According to another embodiment of the present disclosure, the first matching condition includes: the similarity between the frame description information of the candidate frame template and the requirement information is greater than or equal to a first threshold. The second matching condition includes: the similarity between the scene description information of the candidate scene and the requirement information is greater than or equal to a second threshold.
[0089] According to another embodiment of the present disclosure, it also includes: a second template determination module, which is used to determine a target template from at least one candidate framework template in response to determining that at least one candidate framework template among multiple candidate framework templates meets a first matching condition with the requirement information.
[0090] According to another embodiment of the present disclosure, it also includes: a field generation module and a processing module. The field generation module is used to generate fields to be collected based on the big model and the demand information in response to determining that the target template is a candidate framework template and the target template includes elements to be collected. The processing module is used to determine whether information needs to be collected based on at least one of the target scenario and the demand information in response to detecting that the target template is determined based on the target scenario; and if it is determined that information needs to be collected, generate fields to be collected based on the big model, the demand information and the reference fields associated with the target scenario.
[0091] According to another embodiment of the present disclosure, the combination module includes: a first filling submodule, a second filling submodule and a determination submodule. The first filling submodule is used to fill the area content of at least one area into the target template according to the position information of at least one area. The second filling submodule is used to fill the field to be collected into a predetermined position in the target template. The determination submodule is used to determine the target visual object according to the filled target template.
[0092] According to another embodiment of the present disclosure, the generation module includes: a generation sub-module for generating, for an area classified as a graphic and text area, based on a large model, regional text in the regional content according to demand information and regional description information of the area; and based on the large model, generating, according to demand information, regional description information of the area and regional text, a regional image in the regional content.
[0093] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, comprising at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned visual object generation method.
[0094] According to an embodiment of the present disclosure, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the above-mentioned visual object generation method.
[0095] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, including a computer program, which implements the above-mentioned visual object generation method when executed by a processor.
[0096] Figure 6 1 is a block diagram of an electronic device for implementing the visual object generation method of an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0097] like Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0098] A number of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0099] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as the visual object generation method. For example, in some embodiments, the visual object generation method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the visual object generation method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the visual object generation method in any other appropriate manner (e.g., by means of firmware).
[0100] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0101] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0102] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0103] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0104] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0105] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.
[0106] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0107] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for generating a visual object, comprising: Determine a target template, where the target template includes at least one area, each of the at least one area corresponds to area description information, and the area description information is used to describe content to be displayed for the area; For each of the regions, based on the large model, according to the demand information and the region description information of the region, generate the region content of the region; as well as The target template is combined with the region content of each of the at least one region to obtain a target visual object so as to display the target visual object.
2. The method according to claim 1, further comprising: In response to determining that none of the plurality of candidate framework templates and the requirement information satisfy a first matching condition, determining a target scene from the plurality of candidate scenes that satisfies a second matching condition with the requirement information; as well as The target template is determined according to a candidate region set corresponding to the target scene.
3. The method according to claim 2, wherein: The determining the target template according to the candidate area set corresponding to the target scene comprises: Inputting the demand information, the region description information of each candidate region in the candidate region set, and the predetermined prompt information into the large model to obtain the target template; The predetermined prompt information is used to guide the large model to filter regions from the candidate regions, and combine the filtered regions into the target template.
4. The method according to claim 2, wherein: The first matching condition includes: the similarity between the framework description information of the candidate framework template and the requirement information is greater than or equal to a first threshold; The second matching condition includes: the similarity between the scene description information of the candidate scene and the requirement information is greater than or equal to a second threshold.
5. The method according to claim 2, further comprising: In response to determining that at least one candidate framework template among the plurality of candidate framework templates satisfies the first matching condition with the requirement information, the target template is determined from the at least one candidate framework template.
6. The method according to claim 1, further comprising: In response to determining that the target template is a candidate framework template and the target template includes elements to be collected, generating fields to be collected based on the large model and according to the demand information; as well as In response to detecting that the target template is determined based on a target scenario, determining whether information needs to be collected according to at least one of the target scenario and the requirement information; And when it is determined that information needs to be collected, based on the big model, the fields to be collected are generated according to the demand information and the reference fields associated with the target scenario.
7. The method according to claim 6, wherein: The combining the target template with the respective area contents of the at least one area to obtain the target visual object comprises: Filling the region content of each of the at least one region into the target template according to the position information of the at least one region; Filling the to-be-collected fields into predetermined positions in the target template; and The target visual object is determined according to the filled target template.
8. The method according to claim 1, wherein: For each region, based on the large model, according to the demand information and the region description information of the region, generating the region content of the region includes: For areas classified as graphics and text, Based on the large model, generating a region text in the region content according to the demand information and the region description information of the region; and Based on the large model, a region image in the region content is generated according to the demand information, the region description information of the region and the region text.
9. A visual object generating device, comprising: A template determination module, used to determine a target template, wherein the target template includes at least one area, each of the at least one area corresponds to area description information, and the area description information is used to describe the content to be displayed for the area; A generating module, for generating, for each of the regions, regional content of the region based on the macro model, according to the demand information and the regional description information of the region; as well as A combining module is used to combine the target template with the respective area content of the at least one area to obtain a target visual object so as to display the target visual object.
10. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
11. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 8.
12. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.