Object generation method, device, intelligent agent, equipment and medium based on large model

Through the large model generation method, preset templates are used to generate and modify the style of graphic objects, which solves the problem of low efficiency of large models when generating complex styles, and improves the accuracy of generation results and user experience.

CN119228946BActive Publication Date: 2025-10-03BAIDU COM TIMES TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411347225.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2025-10-03
Estimated Expiration
2044-09-25

AI Technical Summary

Technical Problem

Large models need to perform complex and tedious operations when generating objects with complex styles, resulting in low efficiency and lack of accuracy.

Method used

A large model is used to generate a first type of graph object based on a preset template, and a second type of graph object that meets the grammatical constraints of the preset template is generated through style modification, thereby reducing the amount of calculation and improving the accuracy of the generated results.

Benefits of technology

By using the preset template generation and style modification method, the computational complexity of large models is reduced, and the accuracy of generated results and user experience are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119228946B_ABST
    Figure CN119228946B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, intelligent agent, electronic device, and storage medium for object generation based on a large model, relating to the fields of artificial intelligence technology, particularly large models, computer vision, and deep learning. The method comprises: generating a first-type graphical object using the large model based on a preset template in response to a received generation request; and modifying the style of the first-type graphical object using the large model to obtain a second-type graphical object, wherein the description data of the second-type graphical object satisfies the grammatical constraints of the preset template.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as large models, computer vision and deep learning, and specifically to object generation methods, devices, electronic devices, intelligent agents and storage media based on large models. Background Art

[0002] With the development of computer and network technologies, neural network models are widely used in various scenarios. For example, large language models (LLMs) can output multimodal data to meet user browsing needs. However, large models also require complex operations to achieve complex styles. Summary of the Invention

[0003] The present disclosure provides a method, device, intelligent agent, electronic device and storage medium for object generation based on a large model.

[0004] According to one aspect of the present disclosure, a method for generating an object based on a big model is provided, comprising: in response to a received generation request, generating a first type of graphic object based on a preset template using the big model; modifying the style of the first type of graphic object using the big model to obtain a second type of graphic object, wherein the description data of the second type of graphic object satisfies the grammatical constraints of the preset template.

[0005] According to another aspect of the present disclosure, a data display method is provided, comprising: in response to obtaining a second type of graph object generated based on a generation request, displaying the second type of graph object; wherein the second type of graph object is generated using the object method provided in an embodiment of the present disclosure.

[0006] According to another aspect of the present disclosure, a large model-based object generation device is provided, including: a generation module for generating a first type of graphic object based on a preset template using the large model in response to a received generation request; a modification module for modifying the style of the first type of graphic object using the large model to obtain a second type of graphic object, wherein the description data of the second type of graphic object meets the grammatical constraints of the preset template.

[0007] According to another aspect of the present disclosure, a data display device is provided, including: a display module, used to display a second type of graph object in response to obtaining a second type of graph object generated based on a generation request; wherein the second type of graph object is generated using the object generation device provided by an embodiment of the present disclosure.

[0008] According to another aspect of the present disclosure, an artificial intelligence agent is provided, which is configured to execute the method provided by the embodiment of the present disclosure.

[0009] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.

[0010] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method described above.

[0011] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the method described above when executed by a processor.

[0012] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0014] Figure 1 Schematically illustrates an application scenario in which the method and apparatus for generating an object based on a large model according to an embodiment of the present disclosure can be applied;

[0015] Figure 2 Schematically shows a flow chart of a method for generating an object based on a large model according to an embodiment of the present disclosure;

[0016] Figure 3 The following schematically shows the principle of the object generation method according to an embodiment of the present disclosure;

[0017] Figure 4 The following schematically shows a principle diagram of an object generation method according to another embodiment of the present disclosure;

[0018] Figure 5 A schematic diagram schematically illustrates a first type of graph object and a second type of graph object according to an embodiment of the present disclosure;

[0019] Figure 6 The following schematically shows a flow chart of a data display method according to an embodiment of the present disclosure;

[0020] Figure 7 Schematically shows a structural block diagram of a large model-based object generation device according to an embodiment of the present disclosure;

[0021] Figure 8 The following schematically shows a structural block diagram of a data display device according to an embodiment of the present disclosure;

[0022] Figure 9 A block diagram schematically illustrates a structure of an artificial intelligence agent according to an embodiment of the present disclosure; and

[0023] Figure 10 A block diagram of an electronic device according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0024] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0025] Figure 1 The following schematically illustrates an application scenario in which the method and apparatus for generating an object based on a large model according to an embodiment of the present disclosure can be applied.

[0026] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not imply that the embodiments of the present disclosure may not be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary application scenario to which the large-model-based object generation method and apparatus may be applied may include a terminal device, but the terminal device may implement the large-model-based object generation method and apparatus provided by the embodiments of the present disclosure without interacting with a server.

[0027] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a terminal device 101 and a server 102 .

[0028] Terminal device 101 can be connected to server 102 via a network. User 103 can send a generation request through the interactive interface provided by terminal device 101. Terminal device 101 can generate a request and send it to server 102, causing server 102 to call the large model and output the second-type graph object 104. The server sends the second-type graph object 104 to terminal device 101, causing terminal device 101 to display the second-type graph object 104 to user 103.

[0029] Various communication client applications may be installed on the terminal device 101, such as intelligent assistant applications, knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only). A user may input a generation request in the interactive interface of these client applications, and these client applications will display the generated second-type image object 104 to the user 103.

[0030] In one embodiment, the server 102 may generate multimodal data using a large language model and invoke a rendering tool to generate the multimodal data. For example, the rendering tool may generate the second-type graph object 104 so as to display the second-type graph object 104 in the form of an image on the terminal device 101.

[0031] The terminal device 101 can be configured as various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, etc.

[0032] Server 102 can be a server that provides various services, such as a backend management server (for example only) that supports the content viewed by users through the interactive interface of terminal device 101. The backend management server can invoke a large model to generate graph objects in response to received user instructions, and then feed the graph objects back to terminal device 101 for display through the interactive interface. Server 102 can also be a cloud server, also known as a cloud computing server or cloud host. This is a host product within the cloud computing service system that addresses the management difficulties and poor scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or "VPS"). Server 1302 can also be a server for a distributed system or a server integrated with blockchain.

[0033] It should be noted that the object generation method based on the big model provided in the embodiment of the present disclosure can generally be executed by the server 102. Accordingly, the object generation device based on the big model provided in the embodiment of the present disclosure can also be set in the server 102. The object generation method based on the big model provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 102 and can communicate with the terminal device 101 and / or the server 102. Accordingly, the object generation device based on the big model provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 102 and can communicate with the terminal device 101 and / or the server 102.

[0034] Alternatively, the object generation method based on the large model provided in the embodiment of the present disclosure may also be generally executed by the terminal device 101. Accordingly, the object generation apparatus based on the large model provided in the embodiment of the present disclosure may generally be provided in the terminal device 101.

[0035] It should be understood that Figure 1 The number of terminal devices and servers in the embodiment is merely illustrative. Any number of terminal devices and servers may be used as required.

[0036] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.

[0037] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.

[0038] In view of this, the embodiments of the present disclosure provide an object generation method based on a large model, which uses the known types of graph object generation methods of the large model to generate other types of graph objects, thereby providing the output data capability of the large model, expanding the form of the large model output data, and enriching the expression form of the output data, so that the large model can meet the application of more scenarios and improve the user experience.

[0039] The following will be combined Figures 2 to 5 The object generation method based on the large model provided by the present disclosure is described in detail.

[0040] Figure 2 The flowchart of the method for generating an object based on a large model according to an embodiment of the present disclosure is schematically shown.

[0041] like Figure 2 As shown, the method 200 includes operations S210 to S230.

[0042] In operation S210 , in response to the received generation request, a first type graph object is generated based on a preset template using a large model.

[0043] According to an embodiment of the present disclosure, the big model can determine the user's generation intention based on the generation request sent by the user. The generation intention can characterize the content generated by the big model required by the user and the expression method of the generated content. For example, the generated content can be any topic and any event, such as the preparation process of a certain sport, the system structure of a certain software, and the development history of a certain object. The expression method of the generated content can be the way in which the generated content is presented to the user, such as in the form of images and charts. For example, the generated content can be presented in the form of text, audio, video, etc. For example, the generation request can be in the format of text, audio, picture, video, chart, etc.

[0044] According to an embodiment of the present disclosure, a large model can be used to generate first data for describing a first type of graphic object based on a preset template. The graphic object can be an object displayed in the form of a graphic, such as a picture, icon, chart, graph, etc. For example, a renderer can be used to render the first data to generate a chart of the first type. The first type of graphic object generated by the large model can be output in the form of first data, and the first type of graphic object corresponding to the first data can be transmitted only in the server without being displayed to the user through a terminal device.

[0045] According to embodiments of the present disclosure, a large model may be an artificial intelligence (AI) large model, which is a machine learning model with extremely large parameters and complex computational structures. AI large models can process massive amounts of data and complete various complex tasks, such as natural language processing, image recognition, and computer vision. AI large models may include language large models, vision large models, and multimodal large models.

[0046] According to an embodiment of the present disclosure, a preset template is used to generate a first type of graph object. The preset template may be obtained in advance, for example, during the training process of a large model. The preset template corresponds to the class of the graph object. Different preset templates can be trained for different types of graph objects. When the large model is used to generate a graph object of this type, the large model can generate corresponding graph object data based on the corresponding preset template. The graph object data is rendered using a renderer to obtain the graph object of this type.

[0047] For example, the preset template may be pre-stored in a memory, and the server may call the preset template from the memory, so that the large model has the ability to generate the first type of graph object.

[0048] For example, the preset template may record the grammatical rules of the language for describing the graph object of this type. For example, the preset template for the graph object may record the description language used to describe each element in the graph object, the description language used to describe the connection between multiple elements, the description language used to describe the attributes of each element, etc. For example, the preset template for the flowchart may record the description language used to describe the structure of each flow box in the flowchart, the description language used to describe the sequence between multiple flow boxes, the description language used to describe flow boxes that serve as decision points, etc. For example, the description data generated by the large model based on the preset template may include data describing each element in the graph object, data used to describe the connection between multiple elements, and data used to describe the attributes of each element.

[0049] For example, a preset template may include a description language pre-set with various elements for describing a graph object. Based on the generation request, the large model may determine a generation intent, obtain raw data related to the generation intent, and populate the preset template with the raw data to generate first data for the graph object of that type. The first data represents the first type of graph object.

[0050] In operation S220 , the style of the first type of graph object is modified using the large model to obtain a second type of graph object.

[0051] According to an embodiment of the present disclosure, the style of a first-type graph object is described by first data. In response to a generation request indicating the generation of a second-type graph object, the style described by the first data is modified using a large model to obtain second data. The style described by the second data is the style of the second-type graph object.

[0052] According to an embodiment of the present disclosure, the description data of the second type of graphic object satisfies the grammatical constraints of the preset template of the first type of graphic object. The description data of the second type of graphic object describes the second type of graphic object using the grammatical rules for describing the first type of graphic object.

[0053] According to an embodiment of the present disclosure, the graph object obtained by rendering the second data can present the generated result to the user in the second type.

[0054] According to an embodiment of the present disclosure, the pre-set templates pre-stored in the memory may not cover the type of graph object indicated by the user's generation request. If the large model cannot obtain the preset template corresponding to the target type of graph object indicated by the generation request, the large model can first use the existing preset template to generate a first type of graph object, then modify the style of the first type of graph object to obtain the target type of graph object, and then display the generation result to the user using the target type of graph object.

[0055] For example, the first type and the second type are different: the first type may be a flowchart, and the second type may be an architecture diagram. For example, a request is generated for "Generate an architecture diagram for an environmental protection strategy." Because the memory does not have a pre-set template for an architecture diagram, the large model can first retrieve a pre-set template for a flowchart stored in the memory, generate first data describing the flowchart, and then modify the first data to modify the flowchart style described by the first data, thereby obtaining second data describing the architecture diagram style. The second data is then rendered to obtain an architecture diagram for an environmental protection strategy.

[0056] If the pre-set templates built into the large model can't meet the user's generation request, the large model may not be able to directly generate the target type graph object specified in the generation request, or the large model may be unable to accurately generate the target type graph object. If the target type graph object is rebuilt, the large model must perform generative operations based on a large amount of parameter data, which will incur a huge amount of computation and cannot guarantee the accuracy of the generated results.

[0057] According to an embodiment of the present disclosure, a first type of graph object is generated by a large model using a pre-stored preset template, and then the first type of graph object is converted to a second type of graph object based on a generation request. This can directly reuse the descriptive language in the preset template, and make full use of the existing resources and computing resources of the large model to generate graph objects that meet user needs, thereby reducing the amount of calculation of the large model, improving the accuracy of the generation results, and improving the user experience.

[0058] The following will be combined Figure 3 and Figure 4 The principle of the object generation method based on the large model is schematically described.

[0059] Figure 3 The schematic diagram schematically shows the principle of the object generation method according to an embodiment of the present disclosure.

[0060] like Figure 3 As shown, in embodiment 300, a generation request 301 is input into the big model 310, and the big model 310 performs intent recognition on the generation request 301 to obtain a generation intent. The generation intent can indicate the type of graph object described by the output data output by the big model 310 and the content information described by the graph object, so as to indicate that the output data output by the big model 3100 can describe the content information with this type of graph object.

[0061] For example, the generation request 301 may be “generate an architecture diagram for an environmental protection strategy”, and the large model 310 may identify that information related to the environmental protection strategy needs to be displayed in the form of an architecture diagram.

[0062] According to an embodiment of the present disclosure, when the server calls the large model 310 , the server may obtain prompt information 302 related to the generation request 301 , concatenate the prompt information Prompt 302 with the generation request 301 , and input the concatenated information into the large model 310 .

[0063] According to an embodiment of the present disclosure, the prompt information 302 may be pre-set. For example, the prompt information 302 may be obtained during the training process of the large model 310. The prompt information 302 may describe the role, workflow, and precautions of the large model 310.

[0064] For example, the prompt information 302 may be an instruction sent to the large model 310 , or may be a text description, or may be a parameter description in a certain format.

[0065] For example, during the training process of large model 310, the user can adjust the large model by editing prompt information 302 and pre-store the edited prompt information 302 in the server's memory. During the use of large model 310, the server can call prompt information 302 from the memory and combine it with the generation request 301 to input into large model 420.

[0066] For example, during the use of the large model 310 , the user may also adjust the prompt information 302 according to actual operation requirements and operation results to improve the flexibility of the large model.

[0067] For example, the prompt information 302 may describe the role of the large model 320. For example, the role of the large model is a model that can write mermaid code. The brief introduction of the role is that a graph object with a simple structure and clear process can be drawn using the mermaid language according to the user's request.

[0068] For example, if prompt 302 instructs large model 310 to identify the generation intent from generation request 301, prompt 302 describes the workflow of large model 310 in identifying the generation intent. For example, please describe the relevant concepts and knowledge involved in generation request 301 in a detailed, comprehensive, and accurate manner, systematically and in-depth analyzing and describing the relevant concepts and knowledge, and ensure that the content is detailed and informative.

[0069] For example, the precautions described in the prompt information 302 may include requirements on the format of the mermaid code output by the large model 310 and suggestions on writing the mermaid code.

[0070] According to an embodiment of the present disclosure, the prompt information 302 may further instruct the large model 310 to generate a graphic object based on preset templates of other graphic object types if there is no preset template of the graphic object type indicated by the generation intention.

[0071] According to an embodiment of the present disclosure, during the training process of the large model 310, the large model 310 can be trained to learn graph object types of similar types. For example, a first type of graph object is similar to a second type of graph object. During the training process of the large model 310, the large model 310 learns preset templates for the first type of graph object and the third type of graph object, and generates corresponding types of graph objects based on the respective preset templates. The large model 310 also learns that graph objects of similar types to the first type of graph object include the second type of graph object, and graph objects of similar types to the third type of graph object include the fourth type of graph object.

[0072] During use of the large model 310, if a generation request indicates generation of a second-type graph object, the large model 310 may use a preset template of the first-type graph object to generate the second-type graph object. If a generation request indicates generation of a fourth-type graph object, the large model 310 may use a preset template of the third-type graph object to generate the fourth-type graph object.

[0073] According to an embodiment of the present disclosure, the large model 310 generates first data based on the generation request 30 and the prompt information 302, and determines multiple graphic primitives in a first-type graphic object and the logical relationships between the multiple graphic primitives. For example, the first-type graphic object is a graphic object that represents the logical relationships between multiple graphic primitives. A graphic primitive can be a graphic element in a graphic object, and different types of graphic elements can represent different information. For example, the graphic object can be a flowchart, and the graphic elements can include flow boxes and connecting lines. The descriptive data of the flow box can describe the structure of the flow box and the data within the flow box, and the descriptive data of the connecting line can describe the shape of the connecting line and the connection relationship between the flow boxes.

[0074] According to an embodiment of the present disclosure, the large model 310 can convert the logical relationships in the first-type graph object into hierarchical relationships between multiple graph elements based on the generation intent indicated by the generation request 310, and generate a second-type graph object 304 based on the hierarchical relationships and the multiple graph elements. The second-type graph object is a graph object that represents the hierarchical relationship between multiple graph elements. For example, the second-type graph object can be an architectural diagram, in which the graph elements have hierarchical relationships.

[0075] For example, each process block in a flowchart is a process node, and the node level of each process block can be determined based on the logical relationship. For example, the macro model 310 can determine the node level based on the process sequence between multiple process blocks. In the process sequence, the process node executed first has a higher node level, and the process node executed last has the lowest node level. The macro model 310 can determine that process blocks with the same node level belong to the same hierarchy.

[0076] For example, the large model 310 may also determine multiple process branches in the flowchart based on the logical relationship. Multiple process boxes belonging to the same process branch have the same hierarchical relationship.

[0077] In some embodiments, the large model 310 determines layout information of the plurality of graphics primitives in the second-type graphics object based on the hierarchical relationship, and generates the second-type graphics object 304 based on the layout information and the plurality of graphics primitives.

[0078] According to embodiments of the present disclosure, when the relationship between graphic primitives changes, the representation of these graphic primitives in the graphic object also changes accordingly. Based on the determined hierarchical relationships between the multiple graphic primitives, the large model 310 re-determines the position information of the multiple graphic primitives, so that the graphic elements corresponding to the multiple graphic primitives can be laid out in the style of the second-type graphic object 304. The layout information can reflect the hierarchical relationships between the multiple graphic primitives.

[0079] According to an embodiment of the present disclosure, when the large model 310 generates the second-type graphic object 304 , the large model 310 outputs the plurality of graphic elements and layout information between the plurality of graphic elements as the second data 303 to the renderer 320 .

[0080] In some embodiments, the large model 310 outputs the second data 303 describing the second type of graph object 304 in a webpage data format. The webpage data format can be rendered as a webpage and output to the user. The webpage data format has greater compatibility and can be rendered as a webpage displayed in a multimodal style.

[0081] For example, the web page data format may be HTML format. The second data 303 output in HTML format may be rendered in a multimodal style by extending the content in the web page frame.

[0082] According to an embodiment of the present disclosure, the second data 303 is rendered using the renderer 320 so that the second data 303 is displayed to the user on the front end in the style of the corresponding second-type graphic object 304 .

[0083] For example, the renderer 430 may call a display control of the second type of graph object and render the second data 303 in the display control to obtain the second type of graph object 304 .

[0084] For example, renderer 320 inserts a display control for a second-type graphic object into a webpage and populates the display control with second data 303. Renderer 320 renders the second data 303 in the display control, causing the second-type graphic object 304 to be displayed on the front-end webpage. Display controls are used to display graphic objects of a specific type. The style of the graphic object can be edited and modified within the display control.

[0085] For example, the second data 303 is a mermaid code. The renderer 320 can insert the second data 303 described in the mermaid code into a display control for displaying a chart, and render the display control to display the second type graph object 304 in a chart form on the front end.

[0086] Figure 4 The schematic diagram schematically shows the principle of an object generation method according to another embodiment of the present disclosure.

[0087] like Figure 4 As shown, in embodiment 400, prompt information 402 includes first prompt information 421 and second prompt information 422. First prompt information 421 instructs large model 410 to generate first data 412 based on generation request 401. Second prompt information 422 instructs large model 410 to modify the style described in first data 412 to obtain second data 403.

[0088] In some embodiments, the big model 410 is used to determine multiple text data based on a generation request; the big model 410 is used to convert the multiple text data into multiple graphic elements and logical relationships between the multiple graphic elements based on the grammatical constraints of a preset template; and the big model 410 is used to generate a first type of graphic object based on the multiple graphic elements and the logical relationships.

[0089] According to an embodiment of the present disclosure, text data is the data content that needs to be displayed by the first-type graph object. There is a certain logical relationship between multiple text data 411. For example, the first-type graph object can be a flowchart, and the text data 411 can be the text data 411 that needs to be filled in each flow box. The logical relationship is the logical relationship between the flow boxes to which the multiple text data 411 belong.

[0090] According to embodiments of the present disclosure, a graphic primitive includes the data required to represent a graphic element in a graphic object. For example, a graphic primitive includes data on the structure of a flow box and the text within the flow box. This includes information such as the size, position, and filler text of the flow box. A graphic primitive can be described using graphic metadata, which is governed by the grammatical rules of a pre-set template. For example, the graphic metadata describes the structure of each flow box and the text within the flow box using the description rules of a flowchart.

[0091] Since there are logical relationships between the text data 411, the large model 410 also converts the logical relationships between the text data 411 into logical relationships between diagram elements. For example, based on a preset template of a flowchart, the large model 410 also converts the logical relationships between the text data 411 into logical relationships between flow boxes.

[0092] The large model 410 combines a plurality of graphic elements describing graphic elements and logical relationships between graphic elements into first data 412 .

[0093] According to an embodiment of the present disclosure, the first prompt information 421 further describes a workflow for generating the first data based on a preset template. The large model 410 can generate the first data 412 based on the first prompt information 421 .

[0094] In some embodiments, the large model 410 may obtain prompt information 402 related to the generation request 401. The prompt information 402 is preset. The second prompt information 422 in the prompt information 402 may describe the style of the second type of graph object.

[0095] Based on second prompt information 422, large model 410 modifies the style described by first data 412 to obtain second data 412. Second prompt information 422 may describe the expression method and necessary graphic elements of the second type of graphic object. Based on second prompt information 422, large model 410 modifies first data 412 describing the style of the first type of graphic object so that the modified first data includes the necessary graphic elements of the second type of graphic object and the expression method that conforms to the second type of graphic object, thereby obtaining second data 403. Second data 403 describes the style of the second type of graphic object.

[0096] The large model 410 further deletes and modifies the data describing the style in the first data 412 based on the second prompt information 422 to delete the graphic elements that do not belong to the second type of graphic object.

[0097] In some embodiments, the large model 410 can determine the connection symbols of the logical relationship between multiple graphic elements in the first type of graphic object based on the second prompt information 422; and modify the connection symbols based on the generation request 401 to obtain the second type of graphic object.

[0098] According to an embodiment of the present disclosure, a connection symbol is a diagram element in a diagram object. A connection symbol can be a diagram element that represents a logical relationship between multiple diagram elements. For example, if the first type of diagram object is a flow chart, the connection symbol can be a connecting line used to connect the flow boxes.

[0099] For example, the large model 410 determines whether the second type graph object includes the connection symbol based on the generation request 401. For example, if the large model 410 determines that the second type graph object does not include the connection symbol based on the generation request 401, the connection symbol may be deleted.

[0100] For example, based on the generation request 401, the large model 410 determines whether the connection symbol of the second-type graph object is consistent with the connection symbol of the first-type graph object. For example, if the large model 410 determines based on the generation request 401 that the connection symbol of the second-type graph object is inconsistent with the connection symbol of the first-type graph object, the connection symbol may be modified. For example, if the large model 410 determines based on the generation request 401 that the connection symbol of the second-type graph object is consistent with the connection symbol of the first-type graph object, the connection symbol may be retained.

[0101] According to an embodiment of the present disclosure, after large model 410 modifies first data 412 describing the style of a first-type graphic object into second data 403 describing the style of a second-type graphic object, large model 410 deems second data 403 to be descriptive data describing the first-type graphic object. After a renderer renders second data 403, the user visually displays the second-type graphic object.

[0102] The following will be combined Figure 5 The generation of the second type of graph object is schematically described. Figure 5 The following schematic diagram shows a first-type diagram object and a second-type diagram object according to an embodiment of the present disclosure. The first-type diagram object may be a flowchart, and the second-type diagram object may be an architecture diagram.

[0103] like Figure 5 As shown, in embodiment 500, the first data may be data describing a first-type graph object 501, and the first data may include multiple graph metadata in the first-type graph object 501 and the logical relationship between the multiple graph metadata. The second data may be data describing a second-type graph object 502, and the second data may include multiple graph metadata in the second-type graph object 502 and the hierarchical relationship between the multiple graph metadata.

[0104] The first data output by the large model may be rendered as a first-type graph object 501, and the second data output by the large model may be rendered as a second-type graph object 502. The second data may be converted from the first data of a single first-type graph object 501, or may be converted from the first data of multiple first-type graph objects 501.

[0105] According to an embodiment of the present disclosure, in response to a user's generation needs, first data is generated using a pre-stored preset template, and then the first data is converted into second data that meets the user's needs, so that the data is displayed in the type of graphic object required by the user.

[0106] Based on the object generation method of the large model provided by the present disclosure, the present disclosure also provides a data display method. Figure 6 The data display method is described in detail.

[0107] like Figure 6 As shown, the data presentation method 600 of this embodiment may include operation S610.

[0108] In operation S610 , in response to acquiring the second-type graph object generated based on the generation request, the second-type graph object is displayed.

[0109] According to an embodiment of the present disclosure, the generation request may be inputted by the user into the terminal device through an interactive interface provided by the terminal device. After receiving the input data query request, the terminal device may send the generation request to the server.

[0110] For example, the terminal device can convert the generation request into a language recognizable by the large model, or add prompt information to the generation request, and then splice the generation request and the prompt information and input them into the large model.

[0111] For example, the prompt information added by the terminal device can be used to instruct the large model to modify the style of the graphic object.

[0112] According to an embodiment of the present disclosure, the second type of graph object is generated using the large model-based object generation method described above.

[0113] For example, after receiving the generation request, the server can output the second type of graph object using the large model-based object generation method described above, and feed the second type of graph object back to the terminal device. The terminal device can then display the second type of graph object after obtaining it.

[0114] Through the data display method provided by the present invention, the second type of graph object can be displayed to the user only based on the user's generation request in the expression of the generation request indication, thereby enriching the graph object types output by the large model and improving the user experience.

[0115] Based on the object generation method based on the large model provided by the present disclosure, the present disclosure also provides an object generation device based on the large model. Figure 7 The device is described in detail.

[0116] Figure 7 A block diagram of a large model-based object generation device according to an embodiment of the present disclosure is schematically shown.

[0117] like Figure 7 As shown, the large model-based object generation device 700 may include: a generation module 710 , a modification module 720 and a rendering module 730 .

[0118] The generation module 710 is configured to generate a first type of graph object based on a preset template using a large model in response to a received generation request.

[0119] The modification module 720 is used to modify the style of the first type of graphic object using the large model to obtain a second type of graphic object, where the description data of the second type of graphic object meets the grammatical constraints of the preset template.

[0120] According to an embodiment of the present disclosure, the modification module 720 includes a first determination submodule for determining the logical relationship between multiple graphic elements and multiple graphic metadata of a first type of graphic object, where the first type of graphic object represents the logical relationship between multiple graphic elements; a first conversion submodule for converting the logical relationship into a hierarchical relationship between multiple graphic elements, where the second type of graphic object represents the hierarchical relationship between multiple graphic elements; and a first generation submodule for generating a second type of graphic object based on the hierarchical relationship and multiple graphic elements.

[0121] According to an embodiment of the present disclosure, the first generation submodule includes: a determination unit for determining the layout information of multiple graphic elements in a second type of graphic object based on a hierarchical relationship; and a generation unit for generating a second type of graphic object based on the layout information and multiple graphic elements.

[0122] According to an embodiment of the present disclosure, the modification module 720 includes: a second determination submodule, used to determine the connection symbol of the logical relationship between multiple graphic elements in the first type of graphic object; and a first modification submodule, used to modify the connection symbol based on the generation request to obtain the second type of graphic object.

[0123] According to an embodiment of the present disclosure, the modification module 720 includes: an acquisition sub-module, used to obtain prompt information related to the generation request, the prompt information is pre-set, and the prompt information indicates the style of the second type of graphic object; and a second modification sub-module, used to use the large model to modify the style of the first type of graphic object based on the prompt information to obtain the second type of graphic object.

[0124] According to an embodiment of the present disclosure, the large model-based object generation device 700 may further include: an output module, configured to output description data of the second type of graph object in a web page data format using the large model.

[0125] According to an embodiment of the present disclosure, the generation module 710 includes: a third determination submodule, used to use the big model to determine multiple text data based on the generation request; a second conversion submodule, used to use the big model based on the grammatical constraints of a preset template to convert the multiple text data into multiple graphic elements and logical relationships between the multiple graphic elements; a second generation submodule, used to use the big model to generate a first type of graphic object based on multiple graphic elements and logical relationships.

[0126] Based on the data display method provided by the present disclosure, the present disclosure also provides a data display device. Figure 8 The device is described in detail.

[0127] Figure 8 It is a structural block diagram of a data display device according to an embodiment of the present disclosure.

[0128] like Figure 8 As shown, the data display device 800 of this embodiment may include a display module 810 .

[0129] The display module 810 is configured to display the second type of graph object in response to obtaining the second type of graph object generated based on the generation request.

[0130] The second type of graph object is generated by the large model-based object generation device described above. In one embodiment, the display module 810 can be used to perform the operation S610 described above, which will not be described in detail here.

[0131] Figure 9 The structural block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is schematically shown.

[0132] In the embodiments of the present disclosure, inspired by the von Neumann structure in modern computer theory, such as Figure 9 As shown, the AI ​​agent 900 may include five core modules: an input module 910 , a control module 920 , a storage module 930 , a calculation module 940 and an output module 950 .

[0133] Input module 910 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment) and converting it into a format that AI agent 900 can understand and process. Input module 910 is the primary link for AI agent 900 to interact with the outside world. It enables AI agent 900 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.

[0134] In an example, the input module 910 may input the generation request described above.

[0135] In the example, the control module 920 is the core support for the AI ​​agent 900 to handle complex tasks. The control module 920 can execute the large model-based object generation method described above.

[0136] In the example, the control module 920 will continuously interact with the storage module 930, the computing module 940, and / or the output module 950 during operation. However, it should be noted that in the embodiment of the present disclosure, the control module 920 acts as a single initiator to initiate communication with the storage module 930, the computing module 940, and / or the output module 950, and there is no communication coupling between the storage module 930, the computing module 940, and the output module 950.

[0137] In this example, the performance of control module 920 may be closely related to the large model underlying AI agent 900. To fully leverage the capabilities of the large language model, the internal structure of control module 920 may be designed to be highly configurable and extensible to handle a variety of different tasks and requirements in real-world scenarios.

[0138] The storage module 930 may be responsible for memorizing information such as historical conversations, event streams, etc. The aforementioned prompt information and preset templates may be included in the storage module 930 .

[0139] In this example, after receiving the generation request, the AI ​​agent 900 can retrieve relevant prompt information and preset templates from the storage module 930 and feed them back to the control module 920. The control module 920 can then use the fed-back prompt information and preset templates to generate a second-type graph object corresponding to the generation request and pass the second-type graph object to the output module 950.

[0140] The operation module 940 can be regarded as a predefined tool library, and the aforementioned renderer and the like can be included in the operation module 940 .

[0141] In the example, when the AI ​​agent 900 needs to render multiple output data, it can call the relevant renderer from the operation module 940 and feed it back to the control module 920. Then, the control module 920 can use the feedback renderer to render a second type of graph object and pass the second type of graph object to the output module 950. It can be understood that although the large language model has excellent language understanding and generation capabilities, it is the same as a human being. Without the help of any tools, the tasks that can be solved are very limited. When the AI ​​agent 900 is given the ability to call tools, it can achieve tasks such as completing mathematical operations with the help of a calculator, completing data analysis with the help of Python, and completing weather forecasts with the help of a search engine.

[0142] In an example, the output module 950 may output the second type of graph object described above.

[0143] The AI ​​agent 900 according to the embodiment of the present disclosure can simply and effectively improve the level of intelligence, and enhance flexibility and versatility.

[0144] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0145] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.

[0146] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the above method.

[0147] According to an embodiment of the present disclosure, a computer program product includes a computer program, and the computer program implements the above method when executed by a processor.

[0148] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0149] like Figure 10 As shown, electronic device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of electronic device 1000 may also be stored in RAM 1003. Computing unit 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.

[0150] Multiple components in the electronic device 1000 are connected to the I / O interface 1005, including an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the electronic device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0151] The computing unit 1001 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as the large-model-based object generation method. For example, in some embodiments, the large-model-based object generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the large-model-based object generation method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to execute the large model-based object generation method in any other appropriate manner (eg, by means of firmware).

[0152] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0153] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0154] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0155] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0156] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0157] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0158] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0159] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for generating an object based on a large model, comprising: In response to the received generation request, if the large model cannot obtain a first preset template corresponding to the target type graph object indicated by the generation request, generating the first type graph object based on the second preset template using the large model; as well as Using the large model, modifying the style of the first type of graph object to obtain a second type of graph object corresponding to the target type of graph object indicated by the generation request, wherein the description data of the second type of graph object satisfies the grammatical constraints of the second preset template; The step of modifying the style of the first-type graph object to obtain the second-type graph object includes: determining a plurality of graphic primitives in the first-type graphic object and logical relationships between the plurality of graphic primitives, wherein the first-type graphic object represents the logical relationships between the plurality of graphic primitives; Converting the logical relationship into a hierarchical relationship between the plurality of graphic elements, wherein the second-type graphic object represents the hierarchical relationship between the plurality of graphic elements; and Determining layout information of the plurality of graphic elements in the second-type graphic object based on the hierarchical relationship; and The second-type graphic object is generated based on the layout information and the plurality of graphic primitives.

2. The method according to claim 1, wherein The modifying the style of the first-type graph object to obtain the second-type graph object includes: Determining connection symbols for logical relationships between a plurality of graphic elements in the first type of graphic object; and Based on the generation request, the connection symbol is modified to obtain the second type of graph object.

3. The method according to claim 1, wherein The modifying the style of the first-type graph object by using the large model to obtain the second-type graph object includes: Acquire prompt information related to the generation request, where the prompt information is preset and indicates a style of the second-type graph object; and The large model is used to modify the style of the first type of graphic object based on the prompt information to obtain the second type of graphic object.

4. The method according to claim 1, further comprising: The description data of the second type of graph object is output in a web page data format using the large model.

5. The method according to claim 1, wherein The step of generating the first type of graph object based on the second preset template using the large model in response to the received generation request includes: Determining a plurality of text data based on the generation request using a large model; Using the large model based on the grammatical constraints of the second preset template, the plurality of text data are converted into a plurality of graphic elements and logical relationships between the plurality of graphic elements; The first type of graphic object is generated according to the multiple graphic elements and the logical relationship using a large model.

6. A data presentation method, comprising: In response to acquiring a second-type graph object generated based on the generation request, displaying the second-type graph object; The second type of graph object is generated by using the large model-based object generation method according to any one of claims 1 to 5.

7. A large model-based object generation device, comprising: a generation module configured to, in response to a received generation request, generate a first-type graph object using the large model based on a second preset template if the large model cannot obtain a first preset template corresponding to a target-type graph object indicated by the generation request; a modification module, configured to modify the style of the first type of graph object using the large model to obtain a second type of graph object corresponding to the target type of graph object indicated by the generation request, wherein the description data of the second type of graph object satisfies the grammatical constraints of the second preset template; Wherein, the modification module includes: A first determining submodule, configured to determine a plurality of graphic primitives in the first-type graphic object and a logical relationship between the plurality of graphic primitives, wherein the first-type graphic object represents the logical relationship between the plurality of graphic primitives; A first conversion submodule is configured to convert the logical relationship into a hierarchical relationship between the plurality of graphic elements, wherein the second-type graphic object represents the hierarchical relationship between the plurality of graphic elements; and The first generation submodule includes: a determining unit, configured to determine layout information of the plurality of graphic elements in the second-type graphic object based on the hierarchical relationship; and A generating unit is configured to generate the second-type graphic object based on the layout information and the plurality of graphic primitives.

8. The device according to claim 7, wherein The modification module includes: A second determining submodule is configured to determine connection symbols of logical relationships between a plurality of graphic elements in the first type of graphic object; and The first modification submodule is used to modify the connection symbol based on the generation request to obtain the second type of graph object.

9. The device according to claim 7, wherein The modification module includes: an acquisition submodule, configured to acquire prompt information related to the generation request, wherein the prompt information is preset and indicates a style of the second-type graph object; and The second modification submodule is used to modify the style of the first type of graphic object based on the prompt information using the large model to obtain the second type of graphic object.

10. The apparatus according to claim 7, further comprising: An output module is used to output the description data of the second type of graph object in a web page data format using the large model.

11. The device according to claim 7, wherein The generation module includes: A third determining submodule is configured to determine a plurality of text data based on the generation request using the large model; A second conversion submodule, configured to convert the plurality of text data into a plurality of graphic elements and logical relationships between the plurality of graphic elements by using the large model based on the grammatical constraints of the second preset template; The second generating submodule is used to generate the first type of graphic object according to the multiple graphic elements and the logical relationship by using the large model.

12. A data display device, comprising: a display module, configured to display the second-type graph object in response to obtaining the second-type graph object generated based on the generation request; The second type of graph object is generated by the object generation device according to any one of claims 7 to 11.

13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 6.

15. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Chart generation method and device based on large model and electronic equipment

    CN117785165A

  • Document generation method and system and electronic equipment

    CN118551740A