Information generation method and device and related product

By generating multimodal output information through an intelligent system, the problem of insufficient diversity and accuracy in information generation in existing technologies is solved, thereby improving the diversity and accuracy of information generation.

CN122019040APending Publication Date: 2026-05-12BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2026-02-05
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies struggle to meet the requirements of diversity and accuracy when generating information, especially when generating multimodal information, where it is difficult to guarantee both diversity and accuracy.

Method used

The system receives input information and generates output information in at least two modalities, including a first content and a second content. The second content is generated based on the input information and a third content, which is obtained through a search. The output information is then displayed in conjunction with an interactive interface.

Benefits of technology

It improves the diversity and accuracy of generated information, ensuring that the output information can better match user needs in a multimodal context.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019040A_ABST
    Figure CN122019040A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an information generation method and device and a related product, and the method comprises the steps: receiving input information at an interaction interface which is used for interacting with an intelligent system; in response to the fact that the input information indicates that output information at least comprises two modes, the output information is generated through the intelligent system, the output information comprises first content and second content, the first content corresponds to the first mode, and the second content corresponds to the second mode; the second content is generated based on the input information and third content, the third content corresponds to the second mode, and the third content is obtained by searching based on the input information; and displaying the output information on the interactive interface.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to an information generation method, apparatus and related products. Background Technology

[0002] With the rapid development of artificial intelligence (AI) technology, users can generate the information they need. For example, a user can input the information "generate a picture of a dinosaur" into an AI model or agent, and the AI ​​model or agent can generate a picture of a dinosaur and display it to the user. However, with the diversification of information generation needs, higher requirements are being placed on the diversity and accuracy of the generated information. Summary of the Invention

[0003] This disclosure provides an information generation method, apparatus, and related products that can improve the diversity and accuracy of the generated information.

[0004] In a first aspect, embodiments of this disclosure provide an information generation method, including: The system receives input information through an interactive interface, which is used to interact with the intelligent system. In response to the input information indicating that the output information includes at least two modalities, the intelligent system generates the output information, which includes a first content and a second content, wherein the first content corresponds to a first modality and the second content corresponds to a second modality; the second content is generated based on the input information and a third content, wherein the third content corresponds to the second modality and is obtained by searching based on the input information; The output information is displayed on the interactive interface.

[0005] Secondly, embodiments of this disclosure provide an information generation apparatus, including: A receiving unit is used to receive input information at an interactive interface, which is used to interact with the intelligent system. A generation unit is configured to generate output information through the intelligent system in response to the input information indicating that the output information includes at least two modalities. The output information includes first content and second content, the first content corresponding to a first modality and the second content corresponding to a second modality. The second content is generated based on the input information and a third content, the third content corresponding to the second modality, and the third content is obtained by searching based on the input information. A display unit is used to display the output information on the interactive interface.

[0006] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the method described in the first aspect above.

[0007] Fourthly, embodiments of this disclosure provide a computer-readable storage medium for storing computer-executable instructions that, when executed by a processor, implement the method described in the first aspect.

[0008] Fifthly, embodiments of this disclosure provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the method described in the first aspect above.

[0009] In this embodiment, firstly, input information is received at the interactive interface, which is used to interact with the intelligent system. Then, in response to the input information indicating that the output information includes at least two modalities, the intelligent system generates output information, which includes first content and second content. The first content corresponds to the first modality, and the second content corresponds to the second modality. The second content is generated based on the input information and third content, which corresponds to the second modality and is obtained by searching based on the input information. Finally, the output information is displayed at the interactive interface. As can be seen, this embodiment enables the display of output information when the input information indicates that the output information includes at least two modalities. The output information includes the first content corresponding to the first modality and the second content corresponding to the second modality, improving the diversity of the output information. Furthermore, the second content is generated based on the input information and the third content, which corresponds to the second modality and is obtained by searching based on the input information. By first generating the third content and then generating the second content based on the third content and the input information, the accuracy of the second content is improved, thereby effectively improving the diversity and accuracy of the generated information. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in one or more embodiments or related technologies of this disclosure, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Figure 1 This is a schematic diagram illustrating an application scenario of the information generation method provided in an embodiment of the present disclosure; Figure 2 A flowchart illustrating an information generation method provided in an embodiment of this disclosure; Figure 3aThis is a schematic diagram of the process for generating output information provided in an embodiment of the present disclosure; Figure 3b This is a schematic diagram of the process for generating output information provided in another embodiment of this disclosure; Figure 3c A schematic diagram of the process for generating output information is provided for yet another embodiment of this disclosure; Figure 3d This is a schematic diagram of the process for generating output information provided in yet another embodiment of the present disclosure; Figure 3e This is a schematic diagram of the process for generating output information provided in yet another embodiment of the present disclosure; Figure 4a This is a schematic diagram illustrating the output information provided in one embodiment of the present disclosure; Figure 4b This is a schematic diagram illustrating the output information provided in another embodiment of this disclosure; Figure 4c This is a schematic diagram illustrating the output information provided in yet another embodiment of this disclosure; Figure 4d This is a schematic diagram illustrating the output information provided in another embodiment of the present disclosure; Figure 5 This is a schematic diagram of the process for generating an information graph according to an embodiment of the present disclosure; Figure 6 A schematic diagram of an infographic provided for an embodiment of this disclosure; Figure 7 This is a schematic diagram of the structure of an information generation apparatus provided in an embodiment of the present disclosure; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation

[0011] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this disclosure, the technical solutions in one or more embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of the embodiments. Based on one or more embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this disclosure.

[0012] It is understood that before using the technical solutions disclosed in the embodiments of this disclosure, relevant parties should be informed of the type, scope of use, and usage scenarios of the information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and authorization from the relevant parties should be obtained.

[0013] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0014] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0015] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0016] This disclosure provides an information generation method, apparatus, and related products that can improve the diversity and accuracy of the generated information. The information generation method can be applied to and implemented by a terminal device, which includes, but is not limited to, laptops, tablets, desktop computers, set-top boxes, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), smartphones, smart speakers, smartwatches, smart TVs, in-vehicle terminals, and various other types of terminal devices.

[0017] This section will first introduce the concepts that may be involved.

[0018] An intelligent agent is an autonomous entity trained based on one or more AI models, capable of perceiving its environment, making autonomous decisions, and executing actions to achieve specific goals. They are widely used in software, hardware, or hybrid systems. Intelligent agents can possess various capabilities such as drawing, generating videos, writing documents, and creating tables.

[0019] Figure 1 This is a schematic diagram illustrating an application scenario of the information generation method provided in an embodiment of this disclosure, such as... Figure 1 As shown, the scenario includes terminal device 101. Figure 1 In this context, terminal device 101 can be any terminal device of a user with information generation needs.

[0020] Figure 1In this system, terminal device 101 runs an intelligent system, which includes, but is not limited to, an intelligent agent or an AI model. Terminal device 101 also displays an interactive interface 1001 for interacting with the intelligent system. The interactive interface 1001 serves as the human-computer interaction interface for the intelligent system. Users can provide input information to the interactive interface 1001. For example, when a user requests the generation of an image, the input information could be, "Please generate an image combining text and images introducing a trip around the world." Based on the input information, terminal device 101 generates corresponding output information through the intelligent system, such as generating the corresponding image.

[0021] In one scenario, the intelligent system is configured offline in terminal device 101. In this case, terminal device 101 does not need to call the corresponding backend server of the intelligent system; terminal device 101 generates the corresponding output information through the intelligent system configured therein. In another scenario, the intelligent system is configured online in terminal device 101. In this case, terminal device 101 needs to call the corresponding backend server of the intelligent system, sending the input information or instructions generated based on the input information to the backend server, and then calling the backend server to generate the corresponding output information through the intelligent system. Terminal device 101 also obtains the output information returned by the backend server.

[0022] Figure 1 In the process, the terminal device 101 displays output information in the interactive interface 1001, which may include various types of images, text, and videos.

[0023] Figure 2 This is a flowchart illustrating an information generation method provided in an embodiment of the present disclosure, as shown below. Figure 2 As shown, the process includes the following steps: Step S202: Receive input information at the interactive interface, which is used to interact with the intelligent system; Step S204: In response to the input information indicating that the output information includes at least two modalities, the intelligent system generates output information, which includes first content and second content. The first content corresponds to the first modality, and the second content corresponds to the second modality. The second content is generated based on the input information and third content, which corresponds to the second modality and is obtained by searching based on the input information. Step S206: Display the output information on the interactive interface.

[0024] In this embodiment, firstly, input information is received at the interactive interface, which is used to interact with the intelligent system. Then, in response to the input information indicating that the output information includes at least two modalities, the intelligent system generates output information, which includes first content and second content. The first content corresponds to the first modality, and the second content corresponds to the second modality. The second content is generated based on the input information and third content, which corresponds to the second modality and is obtained by searching based on the input information. Finally, the output information is displayed at the interactive interface. As can be seen, this embodiment enables the display of output information when the input information indicates that the output information includes at least two modalities. The output information includes the first content corresponding to the first modality and the second content corresponding to the second modality, improving the diversity of the output information. Furthermore, the second content is generated based on the input information and the third content, which corresponds to the second modality and is obtained by searching based on the input information. By first generating the third content and then generating the second content based on the third content and the input information, the accuracy of the second content is improved, thereby effectively improving the diversity and accuracy of the generated information.

[0025] The terminal device runs an intelligent system, which is a system based on AI technology capable of generating information according to user needs. The intelligent system includes, but is not limited to, intelligent agents or AI models. The terminal device also displays an interactive interface for interacting with the intelligent system. The interactive interface can be the human-computer interaction interface of the intelligent system. Users can provide input information to the intelligent system through the interactive interface. In one example, when a user requests the generation of an image, the input information could be, for example, "Please help me generate an image combining text and images introducing a trip around the world." Based on this, in step S202 above, the user's input information is received at the interactive interface.

[0026] Modalities include, but are not limited to, data modalities such as text, images, and videos. Input information indicates that the output information includes at least two modalities. For example, if the input information is "to represent the theme of traveling around the world through text and images," this input information indicates that the output information includes both text and image modalities. Similarly, if the input information is "to represent the theme of traveling around the world through text and video," this input information indicates that the output information includes both text and video modalities. Again, if the input information is "to represent the theme of traveling around the world through text, images, and videos," this input information indicates that the output information includes text, image, and video modalities. In response to the input information indicating that the output information includes at least two modalities, the intelligent system generates the output information and displays it on the interactive page. For example, it displays the generated images and text.

[0027] In one scenario, a third piece of content is retrieved based on the input information; a second piece of content is generated based on the third piece of content and the input information; and a first piece of content is generated based on the input information. The first piece of content corresponds to a first modality, the second piece of content corresponds to a second modality, and the third piece of content corresponds to a second modality. The first modality and the second modality are different. The first and second pieces of content are displayed as output information in the interactive interface. Of course, the third piece of content can also be displayed.

[0028] In one example, the input information includes text information, or includes voice information; the first modality includes a text modality, the second modality includes an image modality; the first content includes text content, the second content includes image content, and the third content includes image content. In yet another example, the input information includes text information, or includes voice information; the first modality includes a text modality, the second modality includes a video modality; the first content includes text content, the second content includes video content, and the third content includes video content.

[0029] In one scenario, the first content is obtained in the following way: The intelligent system identifies the information elements associated with the input information. The intelligent system generates a first text based on information elements; the first content includes the first text; the first text is used to introduce the information elements.

[0030] First, the intelligent system identifies the information elements associated with the input information. Based on the input information, the intelligent system can expand to obtain various knowledge points represented by the input information; each knowledge point is an information element. For example, if the input information includes "generate an image combining text and images introducing traveling around the world," it can be expanded to obtain multiple knowledge points related to traveling around the world, including but not limited to "preparing documents," "planning routes," "must-see attractions," and "safety precautions," each of which is an information element.

[0031] Then, through an intelligent system, a first text is generated based on the information elements. The first content includes the first text, which is used to introduce the information elements. One or more text search keywords can be generated for each information element. For example, multiple text search keywords can be generated for "route planning," including but not limited to "Asia route planning guide," "North America route planning guide," and "Africa route planning guide."

[0032] Next, a text search is performed based on the text search keywords to obtain the text materials corresponding to each keyword. Each keyword can yield one or more web pages as the text materials for that keyword. A text validator then performs an information sufficiency check on each text material obtained for each keyword, verifying whether the information content of each text material meets the information content requirements of the output information. The verification result is either: the sum of the information content of each text material meets the information content requirements of the output information (i.e., sufficient to generate the output information), or the sum of the information content of each text material does not meet the information content requirements of the output information (i.e., insufficient to generate the output information).

[0033] The system includes a pre-trained text validator for verifying the information sufficiency of text materials. This validator checks whether the information content of each text material meets the information content requirements of the output information. The text validator can be an AI model or an intelligent agent. Through this validator, it ensures that the text materials have sufficient information, thus obtaining the first set of content with sufficient information. The intelligent system can autonomously generate the information content requirements corresponding to the output information based on its AI capabilities, and then call the text validator to determine whether the information content of each text material meets the requirements of the output information.

[0034] If the sum of the information content of each text element meets the information content requirement of the output information, then the content of each text element is summarized, and the first summarized text is used as the first content. Since the text elements are obtained based on information element searches, the first text can be used to introduce the information elements. The intelligent system can pre-set a threshold for the number of information elements; for example, it can summarize the text elements corresponding to at least 90% of the information elements associated with the input information to generate the first text. Alternatively, the intelligent system can pre-set a threshold for the number of text elements; for example, it can summarize at least 90% of the searched text elements to generate the first text.

[0035] If the sum of the information content of all text materials does not meet the information content requirement of the output information, it is determined whether the current text search count has reached the maximum text search count. Multiple text search keywords are generated each time, and a text search is performed based on each generated keyword; this is considered one text search. If the current text search count has not reached the maximum, text search keywords are regenerated, and the text search operation is re-executed. The text validator is repeatedly used to determine whether the sum of the information content of the text materials obtained in this search, along with previous searches, meets the information content requirement of the output information. This progressive text search method, which supplements the generation of text search keywords, ensures that sufficient information content is obtained within a limited number of text searches.

[0036] This process is repeated until the information content of each text source obtained from each joint search meets the information content requirement of the output information, or until the current text search count reaches the maximum text search count. After the current text search count reaches the maximum text search count, the content of each text source obtained is summarized, and the first summarized text is taken as the first content.

[0037] Therefore, through this embodiment, the intelligent system can determine the information elements associated with the input information, generate first text based on the information elements, and the first content includes the first text. The first text is used to introduce the information elements. The first content is obtained based on the information elements associated with the input information, thereby improving the accuracy of obtaining the first content and improving the matching degree between the first content and the input information.

[0038] The process of obtaining the second and third contents is described below.

[0039] In one scenario, the second and third contents are obtained in the following way: The intelligent system identifies the information elements associated with the input information. The first visual object is obtained by searching based on information elements through an intelligent system; the third content includes the first visual object. Through an intelligent system, a second visual object is generated based on the input information and the third content; the second content includes the second visual object; the first content and the second content are combined to introduce information elements.

[0040] First, the intelligent system identifies the information elements associated with the input information. Based on the input information, the intelligent system can expand to obtain various knowledge points represented by the input information; each knowledge point is an information element. For example, if the input information includes "generate an image combining text and images introducing traveling around the world," it can be expanded to obtain multiple knowledge points related to traveling around the world, including but not limited to "preparing documents," "planning routes," "must-see attractions," and "safety precautions," each of which is an information element.

[0041] Then, the intelligent system searches for the first visual object based on information elements. The third content includes the first visual object. Information elements can be represented in text form. In one example, the third content includes an image; therefore, the first visual object can be obtained by searching for the image based on the text, and the third content includes the searched image. In another example, the third content includes a video; therefore, the first visual object can be obtained by searching for the video based on the text, and the third content includes the searched video.

[0042] Finally, a second visual object is generated by the intelligent system based on the input information and the third content. The second content includes the second visual object. The first and second content are combined to introduce information elements. Since the first content is used to introduce information elements associated with the input information, and the second content includes the second visual object, which is derived from the input information and the third content, the second content can also be used to introduce information elements associated with the input information. Therefore, the information elements associated with the input information can be introduced by combining the first and second content.

[0043] In one example, both the second and third content items include images. In this case, a second visual object can be generated by combining text with images to create an image. The second visual object includes the image, and the third content item includes that image. In another example, both the second and third content items include videos. In this case, a second visual object can be generated by combining text with videos to create a video. The second visual object includes the video, and the third content item includes that video.

[0044] As can be seen, by first searching for the third content based on the input information and then generating the second content based on the input information and the third content, the second content can accurately represent the various information elements associated with the input information, thus improving the accuracy of the second content.

[0045] In one scenario, an intelligent system retrieves the first visual object based on information elements, including: The system generates primary search keywords based on information elements. The intelligent system searches for visual objects based on the first search keyword, and obtains the first visual object based on the search results.

[0046] First, the intelligent system generates corresponding primary search keywords for each information element. Each information element can generate one or more primary search keywords. For example, multiple primary search keywords can be generated for "must-visit attractions" in the previous example, including but not limited to "must-visit attractions in Asia," "must-visit attractions in North America," and "must-visit attractions in Africa."

[0047] Then, the intelligent system searches for visual objects based on the first search keyword, and obtains the first visual object based on the search results. When the first visual object includes an image, it can be searched for using the first search keyword, and the first visual object is obtained based on the search results. When the first visual object includes a video, it can be searched for using the first search keyword, and the first visual object is obtained based on the search results. In a specific example, the system can first filter the first search keywords corresponding to each information element to obtain first search keywords that meet preset semantic requirements, and then search for visual objects based on the filtered first search keywords. Accordingly, the search results include multiple visual objects obtained through the search. When the first visual object includes an image, the search results include multiple images obtained through the search; when the first visual object includes a video, the search results include multiple videos obtained through the search.

[0048] Next, based on the search results, the first visual object is obtained. In one case, all the visual objects obtained from the search are used as the first visual object. In another case, the degree of matching between each visual object obtained from the search and its corresponding first search keyword is determined, and the visual object with the highest degree of matching is selected as the first visual object. In a specific embodiment, an intelligent agent or AI model is pre-trained as an image matching degree verifier. The intelligent system calls this image matching degree verifier to determine the degree of matching between each visual object obtained from the search and its corresponding first search keyword, and selects the visual object with the highest degree of matching as the first visual object.

[0049] In one scenario, the image matching verifier is a visual understanding model, which directly verifies the degree of match between the content of the searched visual object and its corresponding first search keyword. In another scenario, the image matching verifier is a text understanding model, which obtains the descriptive text of the searched visual object and verifies the degree of match between the content of the searched visual object and its corresponding first search keyword by verifying the match between the descriptive text and the first search keyword. Here, the "corresponding first search keyword" refers to the first search keyword of the searched visual object.

[0050] In a specific example, firstly, an intelligent system generates multiple first search keywords based on various information elements. Then, the intelligent system filters these first search keywords to select those that meet preset semantic requirements. Visual objects are then searched based on these selected first search keywords. Next, the intelligent system calls an image matching verifier to determine the degree of matching and resolution between each searched visual object and its corresponding first search keyword. Visual objects with a matching degree higher than a matching threshold and a resolution higher than a resolution threshold are selected as retained visual objects. Next, the intelligent system determines whether the information content of each retained visual object is greater than a preset information content, which is the information content corresponding to the output information. If yes, the retained visual object is selected as the first visual object. If not, the process returns to generating the first search keywords and repeats until the information content of each retained visual object after each search is greater than the preset information content, or until the number of iterations reaches a threshold. When the number of iterations reaches the threshold, the retained visual object is selected as the first visual object.

[0051] In the above method, the first search keyword is generated based on the information elements, and then the visual object is searched based on the first search keyword. Based on the search results, the first visual object is obtained. Since the first search keyword is generated based on the information elements, it can ensure that the first visual object obtained by search is strongly correlated with the information elements and with the input information, thereby improving the accuracy of the first visual object.

[0052] In another scenario, the first visual object is obtained through an intelligent system by searching for information elements, including: The system generates primary search keywords based on information elements. The intelligent system supplements the first search keyword based on the first content to obtain the second search keyword; The intelligent system searches for visual objects based on the second search keyword, and obtains the first visual object based on the search results.

[0053] First, an intelligent system generates corresponding primary search keywords for each information element. Each information element can generate one or more primary search keywords. This process is described in the previous section and will not be repeated here.

[0054] Next, unlike the previous situation, the intelligent system supplements the first search keyword based on the first content to obtain the second search keyword. This ensures that the second search keyword is not only related to the input information but also to the first content, enriching the diversity of the third content obtained from the search. Specifically, the first content includes text, and keywords can be extracted from the first content and added to the first search keyword. The supplemented first search keywords then become the second search keywords.

[0055] For example, if the input information is "describe how to make a certain dish using pictures and text," the resulting information elements include "XX dish." The generated first text includes the specific steps in making XX dish, including the specific processes of washing, chopping, and cooking. Then, a first search keyword "XX dish" can be generated based on the information element "XX dish." By supplementing the first search keyword with keywords from the specific steps in the first text, a second search keyword can be obtained, including: "XX dish," keywords from the washing, chopping, and cooking processes. Thus, by supplementing the first search keyword with the first content, the second keyword is not only relevant to the input information but also to the first content, increasing the richness of the first visual object found in the search results. Finally, the intelligent system searches for visual objects based on the second search keywords, and obtains the first visual object based on the search results. When the first visual object includes an image, images can be searched using the second search keywords, and the first visual object is obtained based on the search results. When the first visual object includes a video, videos can be searched using the second search keywords, and the first visual object is obtained based on the search results. In a specific example, second search keywords that meet preset semantic requirements can be filtered from each second search keyword, and visual objects can be searched based on each filtered second search keyword. Accordingly, the search results include multiple visual objects obtained through the search. When the first visual object includes an image, the search results include multiple images obtained through the search; when the first visual object includes a video, the search results include multiple videos obtained through the search.

[0056] Next, based on the search results, the first visual object is obtained. In one case, all the visual objects obtained from the search are used as the first visual object. In another case, the degree of matching between each visual object obtained from the search and its corresponding second search keyword is determined, and the visual object with the highest degree of matching is selected as the first visual object. In a specific embodiment, the above-mentioned image matching degree verifier is pre-trained. The intelligent system calls the image matching degree verifier to determine the degree of matching between each visual object obtained from the search and its corresponding second search keyword, and selects the visual object with the highest degree of matching as the first visual object. The working principle of the image matching degree verifier can be referred to the previous description, and will not be repeated here.

[0057] In a specific example, firstly, the intelligent system generates multiple first search keywords based on various information elements. Then, keywords are extracted from the first content and added to the first search keywords, resulting in second search keywords. Next, the intelligent system filters the second search keywords corresponding to each information element to obtain second search keywords that meet preset semantic requirements. Visual objects are then searched based on these filtered second search keywords. Next, the intelligent system calls an image matching degree validator to determine the matching degree and resolution between each searched visual object and its corresponding second search keyword. Visual objects with a matching degree higher than a matching threshold and a resolution higher than a resolution threshold are selected as retained visual objects. Then, the intelligent system determines whether the information content of each retained visual object is greater than a preset information content, which is the information content corresponding to the output information. If yes, the retained visual object is selected as the first visual object. If not, the process returns to generating the first search keywords and repeats until the information content of each retained visual object is greater than the preset information content after each search, or until the number of iterations reaches a threshold. When the number of iterations reaches the threshold, the retained visual object is selected as the first visual object.

[0058] In the above method, the first search keyword is generated based on information elements. This second keyword is then supplemented with the first content, ensuring it is relevant not only to the input information but also to the first content, thus enriching the search results for the first visual object. Next, the visual object is searched based on the second search keyword. The first visual object is then obtained from the search results. Because the second search keyword is generated from both information elements and the first content, a strong correlation between the obtained first visual object and both the information elements and the input information is guaranteed, improving the accuracy of the first visual object.

[0059] When generating the second content, an intelligent system can generate a second visual object based on the input information and the third content. The second content includes the second visual object. In one scenario, considering that the third content includes the first visual object, the intelligent system can select an object whose content matches the input information from among the various first visual objects as the second visual object, and use this second visual object as the second content. For example, if the first visual object includes an image, the intelligent system can select an image whose content matches the input information from among the various images as the second content. Alternatively, the selected matching image can be edited, such as by resizing or adjusting filters; the processed image becomes the second content.

[0060] In some cases, intelligent systems generate second visual objects based on input information and third-party content, including: The intelligent system generates the first initial text and the first initial visual object based on the input information. The intelligent system uses the third content as the constraint content and adjusts the first initial text and the first initial visual object according to the constraint content to obtain the first content and the second visual object.

[0061] In this scenario, the third content can be retrieved based on the input information. The first and second content can be generated from the input information and the third content. In one example, the input information includes text, the third content includes an image, the first content includes text, and the second content includes an image. This is equivalent to generating text and an image based on the text and the image. In another example, the input information includes text, the third content includes a video, the first content includes text, and the second content includes a video. This is equivalent to generating text and a video based on the text and the video.

[0062] Specifically, firstly, the intelligent system generates a first initial text and a first initial visual object based on the input information. The first initial visual object corresponds to the second modality and includes an image or video. In one example, the intelligent system uses the input information as a prompt word to generate the first initial visual object, which is an image or video created based on the prompt word, and generates explanatory text explaining the content of the image or video as the first initial text. In another example, the intelligent system uses the input information as a prompt word to generate the first initial text, which is text obtained by expanding the content based on the prompt word or answering a question, and generates an image or video matching the text as the first initial visual object.

[0063] Then, through an intelligent system, using third content as constraint content, the first initial text and the first initial visual object are adjusted according to the constraint content to obtain the first content and the second visual object. Using the third content as a constraint means restricting the content described by the first initial text from matching with the third content, and also restricting the content described by the first initial visual object from matching with the third content. Taking the second modality as an image modality as an example, the third content includes multiple images. Based on these multiple images, the content of the first initial text is adjusted so that the content of the first initial text matches the content of the multiple images. Similarly, based on these multiple images, the content of the first initial visual object is adjusted so that the content of the first initial visual object matches the content of the multiple images.

[0064] When adjusting the first initial text and the first initial visual object based on the third content as the constraint content, prompt words can be generated based on the third content, and the first initial text and the first initial visual object can be adjusted based on the prompt words.

[0065] In one example, the input information includes "Please help me generate an image combining text and images to introduce traveling around the world." Based on the input information, the third content obtained through the above method includes images of famous scenic spots in various countries. Then, based on the input information, the first initial text includes introductions to famous scenic spots in various countries, travel routes in various countries, and travel promotions in various countries. Furthermore, based on the input information, the first initial visual object includes images of scenic spots in various countries and street images of different seasons in various countries. Next, based on the third content, text for introducing famous scenic spots in various countries can be selected from the first initial text. This text can be edited or expanded to obtain the first text. Additionally, images related to famous scenic spots in various countries can be selected from the first visual object, and the selected images can be edited, such as by adding filters, to obtain the second visual object.

[0066] Therefore, in this scenario, we can first obtain the first initial text and the first initial visual object based on the input information. Then, using the third content as constraint content, we adjust the first initial text and the first initial visual object according to the constraint content to obtain the first content and the second visual object, i.e., obtain the first content and the second content. This ensures that both the first content and the second content are related to the third content. Since the third content is obtained based on the input information, this not only improves the relevance between the first content and the second content and the input information but also improves the relevance between the first content and the second content themselves, making the output information more closely match the input information and improving the accuracy of the output information.

[0067] In some cases, intelligent systems generate second visual objects based on input information and third-party content, including: The intelligent system uses the third content as a prompt, and generates a second initial visual object based on the prompt. The second visual object is obtained by adjusting the second initial visual object based on the input information through an intelligent system.

[0068] In this scenario, the third content can be obtained by searching based on the input information, or by searching based on both the input information and the first content. First, the intelligent system uses the third content as a prompt, and generates a second initial visual object based on this prompt. If the third content includes an image, that image is used as a reference image, and AI-generated images are used to generate content-matching images as second initial visual objects. If the third content includes a video, that video is used as a reference video, and AI-generated videos are used to generate content-matching videos as second initial visual objects.

[0069] Next, using the input information as a constraint, the second initial visual object is adjusted based on the input information to obtain the second visual object. The content of the second visual object matches the input information. A prompt word can be generated based on the input information, and the second initial visual object is adjusted based on this prompt word to obtain the second visual object, ensuring that the content of the second visual object matches the input information.

[0070] In one example, the input information includes "Please help me generate an image combining text and images to introduce traveling around the world." Based on the input information, the third content obtained through the above method includes images of famous scenic spots in multiple countries during the same season. Then, using the third content as a reference, images of famous scenic spots in multiple countries during different seasons are generated as the second initial visual object. Next, a prompt is generated based on the input information. Based on the prompt, the second initial visual object is adjusted by adding images of scenic spots in countries not included in the second initial visual object, as well as tourist route maps for scenic spots in each country. The resulting images and the second initial visual object constitute the second visual object.

[0071] Therefore, in this scenario, after obtaining the third content, it can be used as a prompt. Based on the prompt, a second initial visual object of the same modality can be generated. Then, the second initial visual object can be adjusted based on the input information to obtain the second visual object. This ensures that the second visual object is not only related to the third content but also to the input information, thereby guaranteeing the accuracy of the second content.

[0072] Of course, in other examples, an intelligent system can be used to treat the input information as a prompt, generate a second initial visual object based on the prompt, and then adjust the second initial visual object based on the third content to obtain the second visual object. Taking the second modality including the image modality as an example, in this process, an image (the second initial visual object) is first generated based on the text input information, and then the generated image is adjusted based on the third content (image modality), and the resulting image is the second content.

[0073] In some cases, intelligent systems generate second visual objects based on input information and third-party content, including: A third initial visual object is generated based on the input information and the first content through an intelligent system; The intelligent system uses the third content as a constraint, and adjusts the initial visual object according to the constraint to obtain the second visual object.

[0074] In this scenario, the third content can be obtained by searching based on the input information, or it can be obtained by searching based on both the input information and the first content. First, the intelligent system generates a third initial visual object based on the input information and the first content. For example, using text-modal input information and the first content as image prompts, an image or video is generated, and this image or video serves as the third initial visual object.

[0075] Then, the third content is used as the constraint content. The third initial visual object is adjusted according to the constraint content to obtain the second visual object. The content of the second visual object matches the third content. The second visual object and the third content are used to represent the same or related content.

[0076] In one example, the input information includes "Please describe how to cook dish XX using pictures and text." The first content includes the cooking process of dish XX in text form. Based on the input information and the first content, images of dish XX and images of eating dish XX are generated as the third initial visual object. Then, the third content, which includes images of each cooking step of dish XX, is obtained. Based on the third content, the third initial visual object is adjusted: the image of dish XX is retained, the image of eating dish XX is deleted, and images related to each cooking step are added to obtain the second visual object.

[0077] Therefore, in this scenario, after obtaining the first content, the input information and the first content can be used as prompts. A third initial visual object can be generated based on the prompts. Then, the third initial visual object can be adjusted based on the third content to obtain a second visual object. This ensures that the second visual object is strongly correlated with the third content, the first content, and the input information, thereby guaranteeing the accuracy of the second content.

[0078] Of course, in other examples, an intelligent system can be used to provide a third piece of content as a prompt, generate a third initial visual object based on the prompt, and then adjust the third initial visual object based on the first content and input information to obtain a second visual object. Taking the second modality, which includes an image modality, as an example, in this process, an image (the third initial visual object) is first generated based on the third content in image form, and then the generated image is adjusted based on the input information in text form and the first content. The resulting image is the second content.

[0079] Figure 3a This is a schematic diagram of the process for generating output information according to an embodiment of the present disclosure, such as... Figure 3a As shown, in this process, firstly, the third content is obtained by searching based on the input information, and then the first content and the second content are generated based on the input information and the third content.

[0080] Figure 3b This is a schematic diagram of the process for generating output information provided in another embodiment of this disclosure, such as... Figure 3b As shown, in this process, firstly, the first content is generated based on the input information; then, the third content is obtained by searching based on the input information and the first content; finally, the second content is generated based on the third content and the input information.

[0081] Figure 3c This is a schematic diagram of the process for generating output information provided in another embodiment of this disclosure, such as... Figure 3c As shown, in this process, firstly, the first content is generated based on the input information; then, the third content is obtained by searching based on the input information and the first content; finally, the second content is generated based on the third content, the input information, and the first content.

[0082] Figure 3d This is a schematic diagram of the process for generating output information provided in another embodiment of the present disclosure, such as... Figure 3d As shown, in this process, firstly, the first content is generated based on the input information; then, the third content is obtained by searching based on the input information; finally, the second content is generated based on the third content and the input information.

[0083] Figure 3e This is a schematic diagram of the process for generating output information provided in another embodiment of the present disclosure, such as... Figure 3e As shown, in this process, firstly, the first content is generated based on the input information; then, the third content is obtained by searching based on the input information; finally, the second content is generated based on the third content, the input information, and the first content.

[0084] The above describes how the first, second, and third content are generated. The output information includes the first and second content.

[0085] In one scenario, the first and second content are combined into an image or card and presented to the user. That is, an image or card containing both the first and second content is displayed to the user. When the first content includes text and the second content includes an image, they can be combined into an image, which can be one or more. When the first content includes text and the second content includes video, they can be combined into a card, which can be one or more. The user can trigger the second content within the card to play the video.

[0086] In another scenario, the first and second content are presented to the user separately. That is, the first and second content are output to the user separately. The number of first content items can be one or more, and the number of second content items can also be one or more. For example, if the first content includes text and the second content includes an image, the user is presented with one image accompanied by a piece of text; or, multiple images are presented, each followed by a piece of text; or, all images are presented with a piece of text. Similarly, if the first content includes text and the second content includes a video, the user is presented with one video accompanied by a piece of text; or, multiple videos are presented, each followed by a piece of text; or, all videos are presented with a piece of text.

[0087] Therefore, in some situations, the output information is displayed on the interactive interface, including: In response to the coupling relationship between the first and second content, output information is displayed on the interactive interface.

[0088] The coupling relationship includes coupling and decoupling. Coupling means that the first and second content are combined into a single image or card, while decoupling means that the first and second content are presented to the user separately. Based on the coupling relationship between the first and second content, the output information is displayed in the interactive interface in a corresponding manner, making the display of output information more precise and flexible.

[0089] In some cases, the above coupling relationship indicates that: the first content and the second content in the output information are decoupled; there are one or more first contents; there are multiple second contents; in response to the coupling relationship between the first and second contents, the output information is displayed on the interactive interface, including: displaying each second content on the interactive interface, and displaying the corresponding first content at the associated position of the second content. The associated position includes after each second content, or after the last second content.

[0090] Figure 4a This is a schematic diagram illustrating the output information provided in one embodiment of the present disclosure, such as... Figure 4aAs shown, the user's input includes "introduce City A using images and text." The interactive interface can display output information based on this input. The output includes multiple images related to City A as the second part of the content, and a piece of text related to City A as the first part. The images and text are decoupled, allowing the display of multiple images related to City A followed by the text. The diagram illustrates this using two images related to City A as an example.

[0091] Figure 4b This is a schematic diagram illustrating the output information provided in another embodiment of this disclosure, such as... Figure 4b As shown, the user's input includes "introduce city A using pictures and text." The interactive interface can display output information based on this input. The output includes multiple images related to city A as the second part of the content, and multiple paragraphs of text related to city A as the first part. The images and text are decoupled, allowing the display of multiple images related to city A followed by a paragraph of text related to city A after each image. The diagram illustrates this using two images related to city A as an example.

[0092] In some cases, the above coupling relationship means that: the first content and the second content are coupled to form output information; there are multiple output information; in response to the coupling relationship between the first content and the second content, the output information is displayed on the interactive interface, including: displaying the current output information on the interactive interface; and in response to a trigger on the current output information, displaying subsequent output information on the interactive interface.

[0093] As described above, the first and second content can be coupled and presented to the user as images or cards. These images or cards constitute the output information. When there are multiple images or cards, the current output information can be displayed on the interactive interface. In response to a trigger on the current output information, subsequent output information can be displayed on the interactive interface. For example, displaying one image or card, and responding to the user's swipe gesture on the currently displayed image or card, displaying the next image or card. Alternatively, all images or cards can be arranged sequentially and displayed together.

[0094] Figure 4c This is a schematic diagram illustrating the output information provided in yet another embodiment of this disclosure, such as... Figure 4c As shown, taking the example of coupling the first and second content into an image, the system can display output information on the interactive interface based on the user's input information "introduce city A in a graphic and textual way." The output information includes multiple images related to city A, each containing both image and text. These images can be arranged and displayed together on the interactive interface. The diagram illustrates this using four images related to city A as an example.

[0095] Figure 4dThis is a schematic diagram illustrating the output information provided in another embodiment of the present disclosure, such as... Figure 4d As shown, taking the example where the first and second content can be coupled into an image, the output information can be displayed on the interactive interface based on the user's input information "introduce city A in a graphic and textual way". The output information includes multiple images related to city A, each containing both image and text. One image can be displayed first, and if the user swipes on that image, the next image will be displayed. The diagram illustrates this using four images related to city A as an example.

[0096] As can be seen, the display methods of the first and second content can be flexibly arranged according to the coupling relationship between the first and second content in the output information to meet the user's browsing needs.

[0097] In a specific scenario, the above process can be used to generate infographics that combine text and images. Infographics, also known as one-image streams, are a form of information presentation that presents text, images, and other materials in a logical order. Users can typically browse a large amount of information within a single infographic without needing additional text or images. Therefore, infographics have the advantage of high information presentation efficiency.

[0098] Figure 5 This is a schematic diagram of a process for generating an infographic according to an embodiment of the present disclosure. This process can be executed by an intelligent system, such as... Figure 5 As shown, the process specifically includes: Input information is obtained through the interactive interface of the intelligent system. Text is searched based on this input information, and an external text validator is invoked to validate the searched text. If the searched text has sufficient information, is authentic, and meets timeliness requirements, or if the maximum number of text searches has been reached, the process continues; otherwise, it returns to the previous step and repeats the process. The number of text searches can be the number of iterations. The text obtained in this step is the first content. Alternatively, the text obtained in this step can be summarized to obtain the first content.

[0099] Next, a first search keyword is generated based on the input information. An external semantic validator is called to perform semantic validation on the first search keyword. An image search is then performed using the semantically validated first search keyword, and an external image matching validator is called to determine whether the searched images match the corresponding first search keyword. If the number of validated images matching the first search keyword is sufficient or the maximum number of image searches is reached, the process continues; otherwise, it returns to the step of generating the first search keyword and repeats the process. The number of image searches can be the loop count. The image obtained in this step is the third content.

[0100] Next, using any of the above methods, generate second content based on the input information and the third content. The second content includes an image. Then, generate an initial prompt word based on the first and second content. Call an external prompt word validator to perform description logic validation on the initial prompt word. If the validation passes, proceed to the next step; if it fails, return to regenerate the initial prompt word, until the generated prompt word passes the description logic validation.

[0101] In this step, initial prompts are generated to describe at least the layout relationships between the first content, the second content, and the content itself. These initial prompts may also describe information such as the font, size, color, and style of the first content, and the resolution, filter, and style of the second content. Furthermore, the initial prompts undergo description logic validation. This validation verifies whether the layout relationships described by the initial prompts meet layout logic requirements, which can be determined autonomously by the intelligent system. Meeting these requirements indicates that the layout relationships described by the initial prompts are clear, unambiguous, and unambiguous. When the initial prompts also describe the font, size, color, and style of the first content, and the resolution, filter, and style of the second content, the intelligent system can further validate whether the information described by the initial prompts conforms to a description specification, which can be generated autonomously by the intelligent system or pre-configured. The intelligent system can also validate whether the initial prompts match the input information and whether they can generate sufficiently informative output information. The sufficiency of information can be determined by the intelligent system's model capabilities.

[0102] Next, an initial infographic is generated based on the verified initial prompts. Then, using the text repair prompts as a reference, the text portion of the initial infographic is repaired to obtain the target infographic.

[0103] In this step, an initial infographic is first generated based on the validated initial prompts. Considering that the text in the initial infographic may contain typos and grammatical errors, the image portion, text style, and text layout of the initial infographic are used as a reference benchmark to repair the text portion, correcting typos and grammatical errors. Finally, the repaired initial infographic is used as the target infographic. Alternatively, the image portion, text style, and text layout of the initial infographic can be kept unchanged, while redundant, erroneous, and grammatically incorrect text is repaired, ensuring that the output text is free of obvious typos, redundant text, and grammatical errors. In practice, prompts can be input into the intelligent system, instructing it to repair redundant, erroneous, and grammatically incorrect text in the initial infographic while maintaining its image portion, text style, and text layout. An example prompt could be: "Maintain 4K ultra-high-definition redraw, only repairing typos, redundant text, and grammatically incorrect parts of the text."

[0104] Figure 6 This is a schematic diagram of an infographic provided according to an embodiment of the present disclosure, with the theme of "Around the World". Figure 6 The document uses a combination of text and images to illustrate various information elements, including "preparing documents," "planning routes," "must-see attractions," and "safety precautions."

[0105] In summary, through the above embodiments, when the input information indicates that the output information includes at least two modalities, the output information can be displayed. The output information includes first content corresponding to the first modality and second content corresponding to the second modality, which improves the diversity of the output information. Furthermore, the second content is generated based on the input information and a third content, which corresponds to the second modality and is obtained by searching based on the input information. By generating the third content first and then generating the second content based on the third content and the input information, the accuracy of the second content is improved, thereby effectively improving the diversity and accuracy of the generated information.

[0106] Figure 7 This is a schematic diagram of the structure of an information generation apparatus provided in an embodiment of the present disclosure, as shown below. Figure 7 As shown, the device includes: The receiving unit 71 is used to receive input information at the interactive interface, which is used to interact with the intelligent system. The generation unit 72 is configured to generate the output information through the intelligent system in response to the input information indicating that the output information includes at least two modalities. The output information includes first content and second content, the first content corresponding to a first modality and the second content corresponding to a second modality. The second content is generated based on the input information and a third content, the third content corresponding to the second modality, and the third content is obtained by searching based on the input information. The display unit 73 is used to display the output information on the interactive interface.

[0107] Optionally, the second content and the third content are obtained by: determining the information elements associated with the input information through the intelligent system; searching for a first visual object based on the information elements through the intelligent system; the third content including the first visual object; generating a second visual object based on the input information and the third content through the intelligent system; the second content including the second visual object; and combining the first content and the second content to introduce the information elements.

[0108] Optionally, obtaining the first visual object by searching for the information elements through the intelligent system includes: generating a first search keyword through the intelligent system based on the information elements; searching for the visual object through the intelligent system based on the first search keyword; and obtaining the first visual object based on the search results.

[0109] Optionally, the step of obtaining the first visual object by searching for the information elements through the intelligent system includes: generating a first search keyword through the intelligent system based on the information elements; supplementing the first search keyword with first content through the intelligent system to obtain a second search keyword; searching for a visual object based on the second search keyword through the intelligent system; and obtaining the first visual object based on the search results.

[0110] Optionally, generating a second visual object based on the input information and the third content through the intelligent system includes: generating a first initial text and a first initial visual object according to the input information through the intelligent system; and adjusting the first initial text and the first initial visual object according to the constraint content using the third content as constraint content through the intelligent system to obtain the first content and the second visual object.

[0111] Optionally, generating a second visual object based on the input information and the third content through the intelligent system includes: using the third content as prompting information through the intelligent system, generating a second initial visual object based on the prompting information; and adjusting the second initial visual object based on the input information through the intelligent system to obtain the second visual object.

[0112] Optionally, generating a second visual object based on the input information and the third content through the intelligent system includes: generating a third initial visual object through the intelligent system according to the input information and the first content; and adjusting the third initial visual object according to the constraint content using the third content as constraint content to obtain the second visual object.

[0113] Optionally, the first content is obtained by: determining the information elements associated with the input information through the intelligent system; generating first text based on the information elements through the intelligent system; the first content includes the first text; the first text is used to introduce the information elements.

[0114] Optionally, the display unit 73 is specifically used to: display the output information on the interactive interface in response to the coupling relationship between the first content and the second content.

[0115] Optionally, the coupling relationship means that the first content and the second content in the output information are decoupled; the number of the first content is one or more; the number of the second content is multiple; the display unit 73 is specifically used to: display each of the second content in the interactive interface, and display the corresponding first content at the associated position of the second content.

[0116] Optionally, the coupling relationship indicates that the first content and the second content are coupled to form the output information; the number of output information is multiple; the display unit 73 is specifically used to: display the current output information on the interactive interface; and in response to a triggering of the current output information, display subsequent output information on the interactive interface.

[0117] The information generation device in this embodiment can implement the various processes of the above-described information generation method embodiment and achieve the same effect and function, which will not be repeated here.

[0118] One embodiment of this disclosure also provides an electronic device. Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure, as shown below. Figure 8As shown, electronic devices can vary considerably due to differences in configuration or performance. They may include one or more processors 801 and memories 802, with the memory 802 storing one or more application programs or data. The memory 802 can be temporary or persistent storage. The application programs stored in the memory 802 may include one or more modules (not shown), each module including a series of computer-executable instructions from the electronic device. Furthermore, the processor 801 may be configured to communicate with the memory 802, executing the series of computer-executable instructions stored in the memory 802 on the electronic device. The electronic device may also include one or more power supplies 803, one or more wired or wireless network interfaces 804, one or more input or output interfaces 805, one or more keyboards 806, etc.

[0119] In one specific embodiment, the electronic device includes a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the following process: The system receives input information through an interactive interface, which is used to interact with the intelligent system. In response to the input information indicating that the output information includes at least two modalities, the intelligent system generates the output information, which includes a first content and a second content, wherein the first content corresponds to a first modality and the second content corresponds to a second modality; the second content is generated based on the input information and a third content, wherein the third content corresponds to the second modality and is obtained by searching based on the input information; The output information is displayed on the interactive interface.

[0120] The electronic device in this embodiment can implement the various processes of the above-described information generation method embodiment and achieve the same effects and functions, which will not be repeated here.

[0121] Another embodiment of this disclosure also provides a computer-readable storage medium for storing computer-executable instructions that, when executed by a processor, implement the following process: The system receives input information through an interactive interface, which is used to interact with the intelligent system. In response to the input information indicating that the output information includes at least two modalities, the intelligent system generates the output information, which includes a first content and a second content, wherein the first content corresponds to a first modality and the second content corresponds to a second modality; the second content is generated based on the input information and a third content, wherein the third content corresponds to the second modality and is obtained by searching based on the input information; The output information is displayed on the interactive interface.

[0122] The computer-readable storage medium in this disclosure embodiment can implement the various processes of the above-described information generation method embodiment and achieve the same effects and functions, which will not be repeated here.

[0123] Another embodiment of this disclosure also provides a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the following process: The system receives input information through an interactive interface, which is used to interact with the intelligent system. In response to the input information indicating that the output information includes at least two modalities, the intelligent system generates the output information, which includes a first content and a second content, wherein the first content corresponds to a first modality and the second content corresponds to a second modality; the second content is generated based on the input information and a third content, wherein the third content corresponds to the second modality and is obtained by searching based on the input information; The output information is displayed on the interactive interface.

[0124] The computer program product in this disclosure embodiment can implement the various processes of the above-described information generation method embodiment and achieve the same effects and functions, which will not be repeated here.

[0125] In various embodiments of this disclosure, the computer-readable storage medium includes read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc.

[0126] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0127] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0128] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0129] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, in implementing the embodiments of this disclosure, the functions of each unit can be implemented in one or more software and / or hardware.

[0130] Those skilled in the art will understand that one or more embodiments of this disclosure can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0131] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0132] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0133] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0134] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0135] One or more embodiments of this disclosure can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices.

[0136] The various embodiments in this disclosure are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0137] The above description is merely an embodiment of this disclosure and is not intended to limit the scope of this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of the claims of this disclosure.

Claims

1. An information generation method, comprising: The system receives input information through an interactive interface, which is used to interact with the intelligent system. In response to the input information indicating that the output information includes at least two modalities, the intelligent system generates the output information, which includes a first content and a second content, wherein the first content corresponds to a first modality and the second content corresponds to a second modality; the second content is generated based on the input information and a third content, wherein the third content corresponds to the second modality and is obtained by searching based on the input information; The output information is displayed on the interactive interface.

2. The method according to claim 1, wherein the second content and the third content are obtained by the following means: The intelligent system determines the information elements associated with the input information. The intelligent system searches for a first visual object based on the information elements; the third content includes the first visual object. The intelligent system generates a second visual object based on the input information and the third content; the second content includes the second visual object; the first content and the second content are combined to introduce the information elements.

3. The method according to claim 2, wherein obtaining the first visual object through the intelligent system based on the information elements includes: The intelligent system generates a first search keyword based on the information elements. The intelligent system searches for visual objects based on the first search keywords and obtains the first visual object based on the search results.

4. The method according to claim 2, wherein obtaining the first visual object through the intelligent system based on the information elements includes: The intelligent system generates a first search keyword based on the information elements. The intelligent system supplements the first search keyword with the first content to obtain the second search keyword; The intelligent system searches for visual objects based on the second search keyword, and obtains the first visual object based on the search results.

5. The method according to claim 2, wherein generating a second visual object through the intelligent system based on the input information and the third content comprises: The intelligent system generates a first initial text and a first initial visual object based on the input information. Using the intelligent system, the third content is used as a constraint, and the first initial text and the first initial visual object are adjusted according to the constraint to obtain the first content and the second visual object.

6. The method according to claim 2, wherein generating a second visual object through the intelligent system based on the input information and the third content comprises: The intelligent system uses the third content as a prompt, and generates a second initial visual object based on the prompt. The intelligent system adjusts the second initial visual object based on the input information to obtain the second visual object.

7. The method according to claim 2, wherein generating a second visual object through the intelligent system based on the input information and the third content comprises: The intelligent system generates a third initial visual object based on the input information and the first content. The intelligent system uses the third content as a constraint and adjusts the third initial visual object according to the constraint to obtain the second visual object.

8. The method according to claim 1, wherein the first content is obtained by means of: The intelligent system determines the information elements associated with the input information. The intelligent system generates first text based on the information elements; the first content includes the first text; the first text is used to introduce the information elements.

9. The method according to claim 1, wherein displaying the output information on the interactive interface includes: In response to the coupling relationship between the first content and the second content, the output information is displayed on the interactive interface.

10. The method according to claim 9, wherein the coupling relationship indicates that: the first content and the second content in the output information are decoupled; the number of the first content is one or more; the number of the second content is multiple; and displaying the output information on the interactive interface in response to the coupling relationship between the first content and the second content includes: The interactive interface displays each of the second contents, and at the associated position of the second contents, the corresponding first contents are displayed.

11. The method according to claim 9, wherein the coupling relationship represents: the first content and the second content are coupled to form the output information; the number of output information is multiple; and displaying the output information on the interactive interface in response to the coupling relationship between the first content and the second content includes: The current output information is displayed on the interactive interface; In response to a triggering of the current output information, subsequent output information is displayed on the interactive interface.

12. An information generation apparatus, comprising: A receiving unit is used to receive input information at an interactive interface, which is used to interact with the intelligent system. A generation unit is configured to generate output information through the intelligent system in response to the input information indicating that the output information includes at least two modalities. The output information includes first content and second content, the first content corresponding to a first modality and the second content corresponding to a second modality. The second content is generated based on the input information and a third content, the third content corresponding to the second modality, and the third content is obtained by searching based on the input information. A display unit is used to display the output information on the interactive interface.

13. An electronic device, comprising: processor; as well as, A memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the method described in any one of claims 1-11.

14. A computer-readable storage medium for storing computer-executable instructions that, when executed by a processor, implement the method described in any one of claims 1-11.

15. A computer program product comprising a computer program that, when executed by a processor, implements the method described in any one of claims 1-11.