Method, device and equipment for generating image-text reply, medium and program product
The generative question-answering system solves the problems of intuitiveness and response time in information transmission of traditional question-answering systems by generating graphic and text responses, realizes the organic combination of images and text, and improves user experience and information acquisition efficiency.
Patent Information
- Application Number
- CN202510908755.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-10-17
AI Technical Summary
When traditional question-and-answer systems provide data presentation, process descriptions, and explanations of complex concepts, pure text cannot convey information intuitively and efficiently. Images and text are difficult to match, and response times are long, resulting in a poor user experience.
A generative question-answering system is used to generate graphic and text responses. The generative model analyzes user query requests, combines visual elements and text content, and generates images and texts with consistent and accurate styles, adapting to the visualization methods of different types of information.
It improves the efficiency of information acquisition, reduces cognitive burden, and enhances user experience. The organic combination of images and text makes complex concepts and data relationships more intuitive and clear.
Smart Images

Figure CN120804350A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates generally to the field of computers, and more particularly to a method, an apparatus, a device, a computer-readable storage medium, and a computer program product for generating an image-text reply. BACKGROUND
[0002] Generative model technology has made significant progress in recent years. A generative model can analyze a query request input by a user and a natural language instruction to generate various forms of content output.
[0003] A generative model can learn from existing data, capture the potential distribution rules in the data, and generate new data. Current question and answer systems based on generative models have been widely applied in various fields, such as text generation creation, data analysis, etc., and are constantly driving the development of various industries. SUMMARY
[0004] According to an example embodiment of the present disclosure, a method, an apparatus, a device, a computer storage medium, and a computer program product for generating an image-text reply are provided.
[0005] In a first aspect of the present disclosure, a method for generating an image-text reply is provided, the method comprising receiving a query request of a user. The method further comprises obtaining an image-text reply generated based on the query request, wherein the image-text reply comprises at least text and an image, and wherein the image is generated by a generative model by analyzing the query request. The method further comprises outputting the image-text reply to the user.
[0006] In a second aspect of the present disclosure, an apparatus for generating an image-text reply is provided, the apparatus comprising a receiving module configured to receive a query request of a user. The apparatus further comprises an obtaining module configured to obtain an image-text reply generated based on the query request, wherein the image-text reply comprises at least text and an image, and the image is generated by a generative model by analyzing the query request. The apparatus further comprises an outputting module configured to output the image-text reply to the user.
[0007] In a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processing unit; at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which when executed by the at least one processing unit, cause the electronic device to perform the method described in the first aspect of the present disclosure.
[0008] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, the computer-readable storage medium having stored thereon machine executable instructions which, when executed by a device, cause the device to perform the method described in the first aspect of the present disclosure.
[0009] In a fifth aspect of the disclosure, there is provided a computer program product comprising computer executable instructions, wherein the computer executable instructions implement the method described according to the first aspect of the disclosure when executed by a processor.
[0010] The summary is provided to introduce a selection of concepts that are further described in the detailed description below. It is not intended to identify key or essential features of the disclosure or to delineate the scope of the disclosure. Other features of the disclosure will be apparent from review of the disclosure herein. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 A schematic diagram illustrating an example environment in which embodiments of the disclosure can be implemented is shown;
[0012] Figure 2 A flow diagram illustrating a method for generating a text-image reply according to embodiments of the disclosure is shown;
[0013] Figure 3 A further process flow diagram illustrating a method for generating a text-image reply according to certain embodiments of the disclosure is shown;
[0014] Figure 4 A schematic diagram illustrating an interface for generating a text-image reply according to embodiments of the disclosure is shown;
[0015] Figure 5 A schematic diagram illustrating a further interface for generating a text-image reply according to embodiments of the disclosure is shown;
[0016] Figure 6 A schematic diagram illustrating a further interface for generating a text-image reply according to embodiments of the disclosure is shown;
[0017] Figure 7 A schematic block diagram illustrating an apparatus for generating a text-image reply according to some embodiments of the disclosure is shown; and
[0018] Figure 8 A block diagram of an example device that can be used to implement embodiments of the disclosure is shown.
[0019] In all of the drawings, the same or similar reference numerals designate the same or similar elements. DETAILED DESCRIPTION
[0020] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information. It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.
[0021] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0022] For example, when receiving the active request of the user, the prompt information is sent to the user to explicitly prompt the user that the operation requested to be executed will require the acquisition and use of the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as the electronic device, the application program, the server or the storage medium, etc. that executes the operation of the technical solutions of the present disclosure according to the prompt information.
[0023] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending the prompt information to the user may, for example, be the manner of a pop-up window, and the prompt information may, for example, be presented in the form of text in the pop-up window. In addition, the pop-up window may, for example, also carry a selection control for the user to select “agree” or “disagree” to provide the personal information to the electronic device.
[0024] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation manners of the present disclosure, and other manners that meet the relevant laws and regulations can also be applied to the implementation manners of the present disclosure.
[0025] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.
[0026] In the description of the embodiments of the present disclosure, the term “comprising” and similar terms are to be understood as open-ended, i.e., “including but not limited to”. The term “based on” is to be understood as “based at least in part on”. The term “one embodiment” or “the embodiment” is to be understood as “at least one embodiment”. The terms “first”, “second”, etc. can refer to different or same objects unless explicitly stated otherwise. Further explicit and implicit definitions can be included below.
[0027] Traditional question and answer systems mainly rely on pure text form to provide information and answers to users. When it comes to data display, process description, spatial relationship or complex concept explanation, pure text form often cannot intuitively and efficiently convey information. Users need to read a large amount of text and build a visual model in their minds, which increases cognitive burden and reduces information acquisition efficiency. In addition, when traditional question and answer systems such as large models or multi-modal models need to introduce images, they often insert an original picture material, which is difficult to match with the generated content, and the text in the image is not accurate, which cannot accurately convey information and layout. In addition, traditional question and answer systems are relatively complex and have a long response time, which makes the user experience poor.
[0028] At least to solve the above and other potential problems, embodiments of the present disclosure provide a method for generating a picture-text reply, which includes receiving a query request of a user. The method further includes obtaining a picture-text reply generated based on the query request, wherein the picture-text reply at least includes text, an image, and the image is generated by a generative model by analyzing the query request. The method further includes outputting the picture-text reply to the user.
[0029] The method implemented by the present disclosure can make the generated image more matched, more accurate, more consistent in style and improve the user experience.
[0030] Embodiments of the present disclosure will be described in detail below with further reference to the accompanying drawings, in which Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. As Figure 1 As shown, a user 102 can input information to be queried to a generative question and answer system 104 implemented according to the present disclosure. As an example, the user 102 can input to the generative question and answer system 104, analysis of summer travel destination trends this year. After receiving the query input, the system can then perform semantic analysis on the received query, thereby identifying key information in the query, such as travel destination, summer this year, trend analysis, and other elements. In some embodiments, the method implemented according to the present disclosure can be applied to various application scenarios such as adaptive education, data analysis, scientific research, business reporting, etc., to achieve multi-scene adaptation.
[0031] According to embodiments of the present disclosure, the generative question and answer system 104 can be pre-trained to determine whether visualization technology needs to be used to enhance the answer based on the nature and content of the user's question. For example, the most suitable visualization method is selected for different types of information, such as data comparison, trend analysis, process display, etc.
[0032] As an example, the generative question answering system 104 can determine that the current query is a query that needs to analyze the travel trends of the current season based on the key information identified in the user input query, and needs to show the ranking changes of the travel destinations.
[0033] In some embodiments, the generative question answering system 104 can determine to first query the relevant data and determine, for example, the top five or top ten travel cities or attractions in the current trends. According to some embodiments of the present disclosure, the generative question answering system 104 can then generate structured text content, including text introductions and analysis of different travel destinations, such as Beijing, Shanghai, New York, etc.
[0034] In some embodiments, the generative question answering system 104 can also generate a picture-text reply 106 for popular scenic photos and images of these different travel destinations. For example, according to embodiments of the present disclosure, the generative model can analyze the intent, theme, background, contextual information, user needs, etc. in the user input query. Based on this information, the generative model can draw image elements in the image, such as sky, ground, plants, sea, etc., and assign corresponding image space and proportion relationship. At the same time, the generative model can also consider the modeling relationship between elements, for example, the sunlight affects the color of the sea and the beach, and the movement of the sea waves can be related to the sea breeze. In some embodiments, the generative model can perform rendering and style processing, for example, selecting brush strokes, textures and color matching of oil painting when rendering details, rather than just generating a real photo. Compared with pure text answers, the picture-text reply 106 based on the picture-text fusion mode can help users quickly understand complex concepts, data relationships and spatial information, reduce cognitive burden, and improve information acquisition efficiency.
[0035] For example, according to some embodiments of the present disclosure, the generative question answering system 104 can also generate visualized outputs for these tourist destinations. For example, the generative question answering system 104 generates trend charts for these tourist destinations to reflect the changing trends of the number of tourists for different tourist destinations. According to some embodiments of the present disclosure, the generative question answering system 104 can also generate ranking charts for these tourist destinations to reflect the ranking of the number of tourists for the current major tourist destinations. Additionally or alternatively, in some embodiments, the generative question answering system 104 can also generate heat maps for these tourist destinations to reflect the current hottest city heat distribution, thereby identifying tourist hotspots. For complex content such as multidimensional data analysis, system architecture, process description, etc., visualized presentation can concretize abstract concepts and make difficult-to-understand content intuitive. Additionally or alternatively, in some embodiments, the generative question answering system 104 can also generate multi-modal content such as videos, audio animations, interactive web page components / modules, specially formatted texts, etc. for each tourist destination, and the present disclosure does not make any limitation in this regard. By using interactive components, users can actively explore data, adjust parameters, and observe changes, thereby changing from passive information reception to active knowledge exploration, improving user engagement and satisfaction.
[0036] As an example, the generated text content can be output in the format of Markdown, while multi-modal content such as images, tables, etc. can be embedded in the text-image reply 106 in the form of HTML or SVG. In some embodiments of the present disclosure, the generative question answering system 104 can also adjust the style of the generated content, such as color, font, icon, etc., so that the overall page design is coordinated and unified. Compared with solutions that need to call external APIs or services, the present disclosure can directly generate complete HTML code in the model answer, reducing system complexity and response delay, and improving system efficiency.
[0037] According to some embodiments of the present disclosure, the client for displaying the text-image reply content 106 can stream these multi-modal content to the user. For example, in some embodiments, the generative client can decompose the generation task into multiple sub-tasks based on the multi-modal content to be generated, such as text generation task, image generation task, chart drawing task, audio synthesis task, etc. In some embodiments, these tasks can be further divided. For example, the generated long text is split into one or more text parts according to paragraphs, or the text is split based on the theme of the content, the information block of each theme is extracted, and these text parts are output in sequence. In some embodiments, for the image content in the generated text-image reply 106, the client can adjust the image size, for example, scale a high-definition image to the standard size required by the website or social media platform, while maintaining the image proportion to avoid stretching or distortion of the image.
[0038] In some embodiments, the client can also crop the image to remove unnecessary parts and keep the key parts. In some embodiments, the client can adjust the resolution of the image based on the network condition to reduce the image file size or to improve the clarity of the image. Additionally or alternatively, in some embodiments, the client can also convert the generated image to a different format, such as from PNG grid to JPEG format, to reduce the file size or to accommodate different application needs. In some embodiments, the client can also compress the generated image, such as lossy compression or lossless compression, to reduce the storage space.
[0039] In some embodiments, when generating the chart, the client can check for missing values, outliers, and duplicated values in the data and perform data cleaning to avoid the impact of these data on the analysis result. In some embodiments, the client can also perform data analysis on the cleaned data, such as data grouping, data aggregation, trend analysis, correlation analysis, etc. to determine the regularity between the data. In some embodiments, the client can also generate the chart based on the analyzed data, such as bar chart, line chart, pie chart, scatter plot, heat map, etc. to show the key content of the data.
[0040] The client can then process these tasks in parallel. According to some embodiments of the present disclosure, the generative question answering system 104 can also determine the generation order of these tasks. For example, in some embodiments, the client can determine to complete the text description task first and then output the image task. In some embodiments, the client can determine the execution time of each subtask and sort the tasks according to the execution time. In some embodiments, the execution time of each subtask can be determined based on historical data, such as based on the experience data of past execution of similar tasks to determine the execution time of each subtask.
[0041] In some embodiments, the execution time of each subtask can be determined based on the complexity of the task, such as the time required to process the text content, the time required to process the image adjustment. In some embodiments, the subtasks that take less time can be placed at the front of the sequence for priority processing, thereby achieving partial output and improving the response experience of the user. Additionally or alternatively, in some embodiments, the subtasks that take longer time can also be placed at the front of the sequence for priority processing, achieving priority start of complex tasks, and other subtasks can be processed in parallel, thereby improving the overall processing efficiency.
[0042] For example, in some embodiments, the client can render the multi-modal content to be generated, such as text, images, audio, etc., step by step. For example, in some embodiments of the present disclosure, the text part can be displayed to the user immediately after being generated, while the background rendering process of the chart, image, or video continues. As an example, the user can be first shown the description text, the reason for recommendation, the relevant data, etc. information about a travel destination.
[0043] When the remaining part is completed, the loading and pushing can continue. For example, after the text description, the generative Q&A system 104 can display to the client the information associated with the travel destination, such as images, charts, and maps, etc. In some embodiments, the image content can be generated by the image processing module of the generative Q&A system 104. In some embodiments, the chart and map content can be generated by the generative Q&A system 104 in real time by calling the visualization library and embedded into the response content. The whole process can maintain asynchronous processing so that the user will not feel any lag.
[0044] As understood by those of ordinary skill in the art, the server where the generative Q&A system 104 is located can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDNs, and basic cloud computing services such as big data and artificial intelligence platforms. The servers can be directly or indirectly connected through wired or wireless communication, which is not limited in the present application.
[0045] The user device used by the user 102 can be any type of mobile computing device, including a mobile computer (e.g., a personal digital assistant, a laptop computer, a notebook computer, a tablet computer, a netbook, etc.), a mobile phone (e.g., a cellular phone, a smartphone, etc.), a wearable computing device (e.g., a smartwatch, a head-mounted device, including smart glasses, etc.), or other types of mobile devices. In some embodiments, the user device can also be a stationary computing device, such as a desktop computer, a game console, a smart television, etc. It should be understood that, where the user device has sufficient computing power, the user device can replace the server to complete the above operations, or the user device and the server can work together to complete the above operations.
[0046] It should be understood that the architecture and functions in the example environment 100 are described for illustrative purposes only, without implying any limitation on the scope of the present disclosure. Embodiments of the present disclosure can also be applied to other environments with different structures and / or functions.
[0047] The process according to embodiments of the present disclosure will be described in detail below in conjunction with other drawings. For ease of understanding, the specific data mentioned in the following description are all exemplary and are not intended to limit the protection scope of the present disclosure. It can be understood that the embodiments described below can also include additional actions not shown and / or can omit the actions shown, and the scope of the present disclosure is not limited in this respect.
[0048] Figure 2 A flow chart of a method 200 for generating a picture-text reply according to certain embodiments of the present disclosure is shown. In this embodiment, the present method can be performed by an application program of a user device of a user 102. At block 202, a query request of a user is received. In some embodiments, the request of the user can be in natural language and additionally can include text, picture description, data requirement or other forms of information.
[0049] At block 204, a picture-text reply is generated based on the query request. According to embodiments of the present disclosure, the picture-text reply content can include various multi-modal content such as visual modules or interactive modules generated by html / JavaScript / CSS or the like code. In some embodiments, the visual modules or interactive modules can include dynamic charts, forms, data display windows, information streams, etc. In some embodiments, the visual modules or interactive modules can also include various interactive components to allow the user to interact in real time, which can be operated with the content through clicking, dragging, sliding, etc., such as picture carousel components, menu switching components, map navigation components, question and answer system components, etc. elements to allow the user to dynamically adjust the content to be displayed.
[0050] According to embodiments of the present disclosure, the picture-text reply can also include various multimedia embedded elements such as video players, audio players, etc. for displaying multimedia content. In some embodiments, the picture-text reply can also include various picture-text cards such as shopping cards, take-out ordering cards, etc. element content.
[0051] Upon receiving the user query request, the generative question answering system implemented according to the present disclosure can analyze and extract the key information in the request. As an example, when the user inputs the query request of “working principle of electric vehicles”, the generative question answering system can extract the key words “electric vehicles” and “working principle”, and then generate the corresponding image based on these key words. For example, the basic elements in the image can be constructed first, such as the vehicle skeleton, tires, doors, etc. In some embodiments, the generative question answering system can process image details such as light and shadow, color, texture, etc. to ensure that the generated content meets the user's query requirements. For example, during the generation process, the surface reflection of the vehicle body, the transparency and reflection effect of the vehicle window, the color and texture processing of the vehicle body, etc. are adjusted to make the generated image consistent with the user's desired image. In some embodiments, the generative question answering system can also generate text associated with the image to ensure accuracy.
[0052] At block 206, the image-text reply is output to the user. The generative question answering system can push the generated image-text reply to the user device. In some embodiments, adjustments can also be made based on user needs and device adaptation to ensure display effects on different screens or devices. Through the method implemented according to the present disclosure, the generated image can be more matched, more accurate, more consistent in style, and the user experience can be improved.
[0053] Figure 3 Another process flow diagram of the method 300 for generating an image-text reply according to some embodiments of the embodiments of the present disclosure is shown. As shown at block 302, the user can ask a question to the generative question answering system implemented according to the present disclosure, such as what is the capital of country A. Figure 3
[0054] At block 304, upon receiving the question asked by the user, the instruction module in the generative question answering system implemented according to the present disclosure can perform question analysis on the received question, for example, by extracting key words such as country A, capital, GDP, area, etc. through natural language processing method; the subsequent decision module in the generative question answering system performs semantic analysis to determine whether the user is asking for facts or requesting trends (e.g., GDP growth trend of country A) and classifies the user's intention to determine the intention, theme, background and context information, etc. associated with the query request. The user's intention can be determined according to these information.
[0055] According to embodiments of the present disclosure, the generative question answering system can be trained based on system cues to regulate its output behavior. In some embodiments, the system cues can set rules for the behavior of the generative question answering system, the generated content, and the format of the output. In some embodiments, the system cues can specify the type of task, such as performing a question answering task, a dialogue task, a summarization task, etc. In some embodiments, the system cues can limit the output range of the generative question answering system, thereby avoiding the generation of redundant or irrelevant content by the generative question answering system. In some embodiments, the system cues can set the language style and tone, output format, etc. for the generative question answering system, thereby helping the generative question answering system to generate content that meets expectations.
[0056] In some embodiments, after determining the user's intention, the generative question answering system can make a visual decision to determine the image-text reply expected by the user. For example, in some embodiments, the rendering module in the generative question answering system can determine that the user expects to generate a picture of a riverside scene, and then determine that the style of the image should be realistic, the structure should be a boat on the river in the foreground, a tower in the middle distance, and a distant mountain in the distance, etc. Then the generative question answering system can draw elements in the image, such as the riverside, the small boat, etc. In some embodiments, light sources and shadows, textures, etc. can also be added to the image to achieve rendering. In some embodiments, tonal and detail modifications can also be added to the image to perform stylization processing.
[0057] In some embodiments, for short factual questions, such as what is the capital of country A, the generative question answering system can return a relatively concise text answer. For data-driven or trend-related questions, such as a comparison of the GDP of country A and country B, the generative question answering system can generate relevant data comparison charts, such as bar charts, line charts, trend analysis charts, etc. In some embodiments, for process questions, such as how to apply for a visa for country A, the generative question answering system can generate a corresponding process flowchart. In some embodiments, for questions that include spatial location or geographical information, such as what are the interesting places in Paris, a corresponding geographical map can be generated to identify the corresponding scenic spots.
[0058] In some embodiments, the interaction control module in the generative question answering system can also generate one or more interactive components for controlling the generated image-text reply. For example, the interactive components can include buttons, sliders, chart interaction components, etc., and the user can drag these components to perform hover tips, click filtering, data exploration, and adjust the content, color, font, size, style, etc. in the image-text reply.
[0059] At block 306, according to some embodiments of the present disclosure, the generative question answering system can combine these generated textual and visual reply contents into a mixed content answer in a mixed format, ensuring that the multi-modal content can be correctly displayed. In some embodiments, the generative question answering system can adjust the visual consistency of the visualized content with the overall interface through style adjustment. Compared to general generative image models, the method implemented according to the present disclosure can organically combine images and text through the processes of creation of visual elements, construction of image space, and rendering of visual details, ensuring that visual and language content can work together to meet the multi-modal needs of users, and the visual elements shown by the image are accurately consistent with the content described by the text. The method implemented according to the present disclosure takes into account the generation of text and images, rather than generating images based on probability distribution.
[0060] At block 308, according to some embodiments of the present disclosure, the generative question answering system can send the generated content to the user client for rendering. According to embodiments of the present disclosure, the generated textual and visual reply can be divided into one or more sub-tasks. The generative question answering system can pre-process and cache the one or more sub-tasks and process the one or more sub-tasks based on a predetermined order to output the textual and visual reply. For example, the generative question answering system can prioritize processing of text tasks and image tasks for immediate display at the client without loading, and subsequently the generative question answering system can generate tables, video content, etc.
[0061] In some embodiments, upon sending, the generative question answering system can adjust the style of images, text, colors, and fonts in the textual and visual reply to adjust the layout of the textual and visual reply based on the size of the display screen of different devices, thereby ensuring adaptability, format specification, and style control under different devices and screen sizes. Additionally or alternatively, in some embodiments, the generative question answering system can check the security and stability of the generated textual and visual reply, for example, providing a secure and reliable CDN library reference mechanism, implementing a secure reference of picture resources, and an external resource reference mechanism, ensuring copyright compliance and link stability, etc.
[0062] At block 310, the user can acquire knowledge by reading the textual content, while deeply understanding complex information or exploring data details through interactive textual and visual content. In some embodiments, the user can provide feedback to the generative question answering system on the evaluation of the generated textual and visual reply content, such as good textual and visual reply content or bad textual and visual reply content. Based on these feedback evaluations, the generative question answering system can be trained again to adjust the generated textual and visual reply content.
[0063] Figure 4 A schematic diagram of an interface 400 for generating a textual and visual reply according to embodiments of the present disclosure is shown. As shown in FIG. 4, the interface 400 can include a textual and visual reply 402, a textual and visual reply content 404, and a textual and visual reply content 406. Figure 4As shown, a user can ask the generative Q&A system implemented according to the present disclosure “explain the tourist attractions in City A”. Upon receiving the user’s inquiry request, the generative Q&A system can extract keywords such as “explain”, “City A”, and “tourist attractions” to determine that the user needs information about the description, location, and features of the attractions.
[0064] In some embodiments, the generative Q&A system can then select one or more representative tourist attractions from the tourist resources of City A and generate corresponding images or photos, such as images 402 and 404. These images can be generated by the generative Q&A system based on the generated textual description, for example, by extracting key visual elements such as the city and the river, the beautiful scenery, etc. from the text, and after understanding the text, the corresponding generation can be performed to ensure that the elements in the image are properly combined together and maintain visual consistency. The rendering of light and shadow, color, and corresponding stylization are then performed to conform to the details in the description.
[0065] In some embodiments, the user can also be provided with close-up photos of the attractions. In some embodiments, the generative Q&A system can also generate corresponding attraction information, such as the name of the attraction, the location, the historical background, the features of the attraction, and the tour suggestions. In some embodiments, the user can also continue to ask more specific questions based on the provided content, such as “What special exhibitions are there in the attraction?” or “Are there any interactive exhibitions in the museum?” etc. In some embodiments, the system can recommend other similar tourist attractions based on the user’s interests, for example, if the user likes history, it can recommend historical museums, etc.
[0066] Through the method implemented according to the present disclosure, the user can not only obtain clear attraction information, but also intuitively experience the charm of the attraction through images, and at the same time, the generative Q&A system can provide further interaction and recommendations to enhance the user’s experience.
[0067] Figure 5 A schematic diagram of another interface 500 for generating a text-image reply according to an embodiment of the present disclosure is shown. As Figure 5 As shown, a user can ask the generative Q&A system implemented according to the present disclosure “explain Company B”. Upon receiving the user’s inquiry request, the generative Q&A system can extract keywords such as “explain” and “Company B” to determine that the user needs information about the basic information, asset status, and business trends of Company B.
[0068] In some embodiments, the generative question answering system can first generate a company profile about B Company, and then generate the asset status of B Company based on the collected data and provide the user with data of a column chart to intuitively reflect the changing trend to the display user. In some embodiments, the generative question answering system can also perform data cleaning and data analysis processing based on the collected data, and generate a histogram 502 and a curve chart 504 based on the processed data to reflect the market demand for B Company's products, changes in the external economic environment, etc. By combining charts and text, the user can more intuitively understand the development status of B Company, thereby helping the user make rational judgments and a more comprehensive understanding.
[0069] Figure 6 A schematic diagram of yet another interface 600 for generating a graphic text reply according to embodiments of the present disclosure is shown. As shown, a user can ask the generative question answering system implemented according to the present disclosure "explain the flower viewing places in C City". When receiving the user's inquiry request, the generative question answering system can extract keywords such as "flower viewing", "C City", etc. to determine that the user needs information about flower viewing places in C City, including the best time for flower viewing, ticket price, and flower viewing types, etc. Figure 6
[0070] The generative question answering system implemented according to the present disclosure can then generate a response based on the flower viewing data of C City, for example, recommended D Park and E Park, and their associated best flower viewing time, ticket, flower viewing features, etc. In some embodiments, the generative question answering system implemented according to the present disclosure can also call map data to mark the locations of D Park and E Park in the map 602. Through the generative question answering system implemented according to the present disclosure, the user can not only quickly understand the flower viewing spots in C City, but also intuitively understand the specific locations, characteristics and recommended times of major scenic spots through the map illustration, improving the overall interactive experience. Additionally or alternatively, in some embodiments, text, images, maps, etc. can also be included in the generated graphic text reply at the same time. These generated data can correspond to each other and have consistent styles, so that the user can fully grasp the information needed to understand and improve the user experience.
[0071] Figure 7 A schematic block diagram of an apparatus 700 for generating a graphic text reply according to some embodiments of the present disclosure is shown. The apparatus 700 can be implemented by software, hardware, or a combination of both. As shown, Figure 7 The apparatus 700 includes a receiving module 710, a generating module 720, and an output module 730.
[0072] In some embodiments, the receiving module 710 is configured to receive a query request of a user. In some embodiments, the obtaining module 720 is configured to obtain a text-image reply generated based on the query request, wherein the text-image reply comprises at least text and an image, and the image is generated by a generative model by analyzing the query request. In some embodiments, the output module 730 is configured to output the text-image reply to the user.
[0073] In some embodiments, the output module is configured to divide the generated text-image reply into one or more sub-tasks, pre-process and cache the one or more sub-tasks, and process the one or more sub-tasks based on a predetermined order to output the text-image reply.
[0074] In some embodiments, the one or more sub-tasks comprise a text task, an image task, and a chart task. The output module 730 is configured to process the text task includes splitting the text into different text parts and processing the different text parts in sequence, process the image task includes resizing, formatting and compressing the image, and process the chart task includes cleaning and data analysis processing the obtained data, and drawing a data chart based on the processed data.
[0075] In some embodiments, the apparatus 700 comprises a processing module configured to process the one or more sub-tasks based on the predetermined order includes: sorting the one or more sub-tasks based on the time required to process each sub-task, and outputting the one or more sub-tasks based on the sorting; or outputting the one or more sub-tasks based on a preference order of the user.
[0076] In some embodiments, the apparatus 700 comprises an adjusting module configured to adjust the text-image reply, adjust the style patterns in the text-image reply associated with the image, text, color and font; and adjust the layout of the text-image reply based on the size of the display screen of different devices.
[0077] In some embodiments, the text-image reply is generated by the generative model based on the query request includes: analyzing the query request by the generative model to identify intent, theme, background and context information associated with the query request; determining the user demand of the user by the generative model based on the identified intent, theme, background and context information; and determining the text-image reply expected by the user by the generative model based on the user demand.
[0078] In some embodiments, generating the image by the generative model comprises: analyzing, by the generative model, one or more of the intent, the topic, the context, the background, the contextual information, the user needs, determining the style and structure of the image to be generated; rendering, by the generative model, the image elements in the image to be generated; rendering, by the generative model, the rendered image elements; and stylizing, by the generative model, the rendered image elements.
[0079] In some embodiments, the image-text reply further comprises one or more of a chart, an animation, an audio, a video, an interactive component; and wherein the chart further comprises one or more of a data comparison chart, a trend analysis chart, a line chart, a flow chart, an interactive chart, a geographical chart.
[0080] In some embodiments, rendering the data chart based on the processed data comprises: in response to the query request including a data comparison, generating a data comparison chart to display to the user; in response to the query request including a trend analysis, generating a trend analysis chart or a line chart to display to the user; in response to the query request including a description of a process or steps, generating a flow chart based on the description of the process to display to the user; in response to the query request including a request for spatial location or geographical information, generating a geographical chart based on map data.
[0081] In some embodiments, further comprising, in response to the query request including a user need to interact with the generated image-text reply, generating, by the generative model, an interactive component; and wherein the interactive component comprises one or more of a button, a slider, a chart interactive component, and is configured to control one or more elements in the image-text reply.
[0082] In some embodiments, further comprising controlling one or more elements in the image-text reply comprises: controlling, by the interactive component, the content, the color, the font, the size, the style of one or more elements in the image-text reply.
[0083] In some embodiments, further comprising checking the security and stability of the generated image-text reply.
[0084] In some embodiments, the generative model has been pre-trained based on multi-modal content to output different image-text replies according to different types of content.
[0085] In some embodiments, receiving, by the generative model, feedback from the user on the image-text reply; and re-training, by the generative model, based on the feedback.
[0086] Figure 8 A block diagram of an example device 800 that can be used to implement embodiments of the present disclosure is shown. It should be understood that Figure 8The device 800 shown is merely an example and should not be construed as limiting the functionality and scope of the implementations described herein. Figure 1 The user equipment described above can be used to perform the above-described Figures 1 to 7 For another example, the device 800 may correspond to the electronic device of the third aspect of the invention summary.
[0087] like Figure 8 As shown, device 800 is in the form of a general-purpose computing device. Components of device 800 may include, but are not limited to, one or more processors or processing units 810, memory 820, storage devices 830, one or more communication units 840, one or more input devices 850, and one or more output devices 880. Processing unit 810 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of device 800.
[0088] Device 800 typically includes multiple computer storage media. Such media can be any available media accessible to device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory) or some combination thereof. Storage device 830 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks or any other media, which can be used to store information and / or data (e.g., training data for training) and can be accessed within device 800.
[0089] The device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 8As shown in FIG. 8, a disk drive 820 or other computer readable media drive can be provided for reading from or writing to a removable, non- volatile, magnetically encoded disk (e.g., a "floppy disk"), and an optical disk drive 822 can be provided for reading from or writing to a removable, non- volatile, magnetically encoded disk (e.g., a "floppy disk"). In such cases, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 820 can include a computer program product 825 having one or more program modules configured to carry out the various methods or actions of the various implementations of the present disclosure.
[0090] The communication unit 840 enables communications with other computing devices over a communication medium. Additionally, the functionality of the components of the device 800 can be implemented in a single computing cluster or multiple computer machines that are capable of communicating over a communication connection. Thus, the device 800 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the networking environment.
[0091] The input device 850 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 860 can be one or more output devices, such as a display, a speaker, a printer, etc. The device 800 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc., through the communication unit 840, as needed, with one or more devices that enable a user to interact with the device 800, or with any devices (e.g., a network card, a modem, etc.) that enable the device 800 to communicate with one or more other computing devices. Such communication can be carried out via an Input / Output (I / O) interface (not shown).
[0092] According to example implementations of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the methods described above. According to example implementations of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the methods described above. According to example implementations of the present disclosure, a computer program product is provided having a computer program stored thereon, which program is executed by a processor to implement the methods described above.
[0093] Various aspects of the disclosure can be described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, systems, and computer program products according to this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0094] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer- readable storage medium having no data storage cycles that change state. The instructions can be executed by one or more processors of a computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions which execute via the one or more processors of the computer or other programmable data processing devices create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0095] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0096] The flow and block diagrams in the drawings show the architectural, functional, and operational views of possible implementations of systems, methods, and computer program products according to the implementations of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, a segment, or a portion of instructions which contain one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession can in fact be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems which perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0097] The implementations of the present disclosure have been described above with the understanding that these implementations are exemplary, and are not exhaustive, and are not limited to the implementations disclosed. Many modifications and variations of the described implementations are possible, without departing from the scope and spirit of the described implementations. The choice of words in this document is intended to best describe the principles of the implementations, practical application, or improvement to the technology in the market, or to enable other ordinary skilled persons in the art to understand the various implementations disclosed herein.
Claims
1. A method for generating a graphic text reply, comprising: Receive user query requests; Obtaining a graphic-text response generated based on the query request, wherein the graphic-text response includes at least text and an image, and the image is generated by a generative model by analyzing the query request; as well as Output the graphic reply to the user.
2. The method according to claim 1, wherein outputting the graphic reply to the user comprises streaming the graphic reply to the user: Divide the generated graphic and text reply into one or more subtasks; Preprocessing and caching the one or more subtasks; as well as The one or more subtasks are processed based on a predetermined order to output the graphic and text reply.
3. The method according to claim 2, wherein the one or more subtasks include a text task, an image task, and a chart task; and Processing the text task includes splitting the text into different text parts, and processing the different text parts in sequence; Processing the image task includes adjusting the size and format of the image and compressing the image; and Processing the chart task includes cleaning and analyzing the acquired data, and drawing a data chart based on the processed data.
4. The method of claim 2, wherein processing the one or more subtasks based on a predetermined order comprises: sorting the one or more subtasks based on a time required to process each subtask, and outputting the one or more subtasks based on the sorting; or The one or more subtasks are output based on the user's preferred order.
5. The method according to claim 1, further comprising adjusting the graphic reply: Adjusting the style associated with images, text, colors, and fonts in the graphic response; and The layout of the graphic reply is adjusted based on the size of the display screen of different devices.
6. The method according to claim 1, wherein generating the graphic-text reply based on the query request by the generative model comprises: analyzing the query request by the generative model to identify intent, subject matter, background, and contextual information associated with the query request; determining, by the generative model, a user need of the user based on the identified intent, the subject, the background, and the contextual information; as well as The generative model determines the graphic and text reply desired by the user based on the user's needs.
7. The method of claim 6, wherein generating the image by the generative model comprises: The generative model analyzes one or more of the intent, the subject, the background, the contextual information, and the user needs to determine the style and structure of the image to be generated; The generative model draws image elements in the image to be generated; Rendering the drawn image elements by the generative model; as well as The rendered image elements are stylized by the generative model.
8. The method according to claim 1, wherein the graphic reply further comprises one or more of a chart, animation, audio, video, and interactive components; and The charts may include one or more of a data comparison chart, a trend analysis chart, a line chart, a flow chart, an interactive chart, and a geographic chart.
9. The method of claim 8, wherein plotting data charts based on the processed data comprises: In response to the query request including data comparison, generating the data comparison chart for display to the user; In response to the query request including trend analysis, generating the trend analysis graph or the line graph for display to the user; In response to the query request including a description process or steps, generating the flowchart based on the description process to display to the user; In response to the query request including a request for spatial location or geographic information, the geographic map is generated based on map data.
10. The method according to claim 8, further comprising: In response to the query request including the user needing to interact with the generated graphic and text reply, the generative model generates the interaction component; and The interactive component includes one or more of a button, a slider, and a chart interactive component, and is configured to control one or more elements in the graphic reply.
11. The method according to claim 10, wherein controlling the one or more elements in the graphic reply comprises: The content, color, font, size, and style of one or more elements in the graphic reply are controlled through interactive components.
12. The method according to claim 1, further comprising: The security and stability of the generated graphic and text responses are checked by the generative model.
13. The method according to claim 11, wherein the generative model has been pre-trained based on multimodal content to output different graphic and text responses according to different types of content.
14. The method according to claim 1, further comprising: The generative model receives feedback from the user on the graphic and text reply; as well as The generative model is retrained based on the feedback.
15. A device for generating a graphic reply, comprising: A receiving module configured to receive a query request from a user; an obtaining module configured to obtain a graphic-text response generated based on the query request, wherein the graphic-text response includes at least text and an image, and the image is generated by a generative model by analyzing the query request; as well as An output module is configured to output the graphic and text reply to the user.
16. An electronic device comprising: at least one processing unit; At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit causing the electronic device to perform the method according to any one of claims 1 to 14.
17. A computer program product having a computer program stored thereon, which, when executed by a processor, implements the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Multi-mode reply generation method and device, electronic equipment and storage medium
CN116401349A
User question and answer method and device, electronic equipment, storage medium and program product
CN119938951A
Visual indicators of generative model response details
US12266065B1