Interactive processing method and device, equipment, medium and program product
By blending AI response text with interactive components in the interactive dialogue interface, the problem of the traditional AI customer service interface being monotonous in layout and cumbersome in operation is solved, achieving complex layouts and instant response, thus improving user experience and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENPAY PAID TECH
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-26
AI Technical Summary
Traditional AI customer service replies are presented in basic Markdown text format, which cannot achieve complex layouts, resulting in low information transmission efficiency and requiring users to switch between different interfaces to complete the operation, thus affecting the user experience.
The interactive dialogue interface blends AI-generated response text with interactive components to achieve complex layouts and multimodal expression, and binds operational behaviors to components, allowing business operations to be completed directly within the interface.
Improve information transmission efficiency, enable immediate response and processing, and significantly enhance user experience and query efficiency.
Smart Images

Figure CN122086274A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device, medium, and program product for interactive processing. Background Technology
[0002] In the current era of booming artificial intelligence (AI), AI responses have become an indispensable form of interaction in many scenarios. For example, from intelligent customer service answering user questions in real time to virtual assistants accurately planning daily schedules, AI, with its powerful natural language processing and deep learning capabilities, quickly understands user intent, extracts key information from massive amounts of data, and generates accurate and relevant responses.
[0003] However, traditional AI customer service responses often use basic Markdown text to present the AI's reply, but due to its limited expressive capabilities, it only supports basic text formatting and cannot achieve complex layouts, resulting in low information transmission efficiency. Furthermore, regarding subsequent operations, the currently presented AI reply only provides simple instructions, requiring the user to switch from the current customer service dialogue interface to the business function interface and rely on their memory of the instructions to complete the relevant operations. This results in numerous steps, poor efficiency, and a negative impact on user experience. Summary of the Invention
[0004] This application provides an interactive processing method, apparatus, device, medium, and program product that not only breaks through the limitations of basic text layout, enabling complex layouts and multimodal expression and improving information transmission efficiency, but also allows for the direct completion of relevant business operations on the interactive dialogue interface, achieving immediate response and ultimate simplification of operations, significantly improving user experience and query efficiency.
[0005] This application provides a method for interactive processing. The method includes: In response to an interactive trigger operation of an artificial intelligence (AI) interactive application, an interactive dialogue interface related to the AI interactive application is displayed, the interactive dialogue interface including input controls; In response to an input operation in the input control, a first query question is displayed; In response to a trigger operation for the first query question, the first query question and M first responses to the first query question are displayed in the interactive dialogue interface. Each first response includes an AI response text and an interactive component associated with the AI response text. The interactive component is used to guide the trigger to perform a target behavior on the corresponding AI response text. The AI response text is used to answer the first query question, and M is an integer greater than or equal to 1.
[0006] This application also provides an interactive processing apparatus. The interactive processing apparatus includes: A response display unit is used to respond to an interactive trigger operation of an artificial intelligence (AI) interactive application and display an interactive dialogue interface related to the AI interactive application, the interactive dialogue interface including input controls; The response display unit is used to display a first query question in response to an input operation in the input control; The response display unit is configured to respond to a trigger operation for the first query question by displaying the first query question and M first responses to the first query question in the interactive dialogue interface. Each first response includes an AI response text and an interactive component associated with the AI response text. The interactive component is configured to guide the triggering of a target behavior on the corresponding AI response text. The AI response text is used to answer the first query question, and M is an integer greater than or equal to 1.
[0007] In one possible design, in another implementation of another aspect of the embodiments of this application, the response display unit is further specifically used for: After displaying the first query question and M first responses to the first query question in the interactive dialog interface in response to a trigger operation for the first interactive component, a second query question is displayed in the interactive dialog interface in response to a trigger operation for the first interactive component. The first interactive component and the second query question have a one-to-one association relationship. The second query question is generated based on the business logic of the first interactive component. The first interactive component is an interactive component included in any one of the M first responses. In response to the second query question, a second answer matching the second query question is displayed in the interactive dialog interface.
[0008] In one possible design, in another implementation of another aspect of the embodiments of this application, the interactive component includes an information container class component; the response display unit is further specifically used for: In response to a trigger operation for the first query question, after displaying the first query question and M first responses to the first query question in the interactive dialogue interface, in response to a trigger operation for the information container class component, aggregated information of the first AI response text is displayed. The aggregated information of the first AI response text is used to present the multi-dimensional features of the first AI response text. Wherein, the first AI response text is the AI response text contained in any one of the M first response contents, and the information container class component is associated with the first AI response text.
[0009] In one possible design, in another implementation of another aspect of the embodiments of this application, the response display unit is further specifically used for: After displaying aggregated information of the first AI response text in response to a trigger operation on an information container component, the target filter conditions are displayed in response to a configuration operation on the information filter slider in the aggregated information of the first AI response text. In response to the target filtering conditions, first content is displayed, which is information from the aggregated information of the first AI response text that meets the target filtering conditions.
[0010] In one possible design, in another implementation of another aspect of the embodiments of this application, the response display unit is further specifically used for: After displaying aggregated information of the first AI response text in response to a trigger operation on an information container component, the current information viewing progress is displayed in response to a drag operation on the progress bar slider in the aggregated information of the first AI response text. In response to the current information viewing progress, second content is loaded, which is an information fragment in the aggregated information of the first AI reply text that matches the current information viewing progress.
[0011] In one possible design, in another implementation of another aspect of the embodiments of this application, the interactive component includes an operation triggering class component; the response display unit is further specifically used for: After displaying the first query question and M first responses to the first query question in the interactive dialogue interface in response to the trigger operation for the operation triggering class component, third content is displayed on the first information display page in response to the trigger operation for the operation triggering class component. The third content is related information related to the second AI response text generated based on the business logic of the operation triggering class component. Wherein, the second AI response text is the AI response text contained in any one of the M first response contents, the operation triggering component is associated with the second AI response text, and the first information display page is the associated hierarchical page of the interactive dialogue interface.
[0012] In one possible design, in another implementation of another aspect of the embodiments of this application, the interactive component includes a content display component; the response display unit is further specifically used for: After displaying the first query question and M first responses to the first query question in the interactive dialogue interface in response to a trigger operation for the content display component, multimedia display information associated with the third AI response text is displayed in response to a trigger operation for the content display component. The multimedia display information associated with the third AI response text is used to present the target flow of the third AI response text. The target flow includes at least one of a business operation flow and a content generation flow. The third AI response text is the AI response text contained in any one of the M first response contents, and the content display component is associated with the third AI response text.
[0013] In one possible design, in another implementation of another aspect of the embodiments of this application, the content display component includes an image preview component; the response display unit is specifically used for: In response to a trigger operation on the image preview component, the first image associated with the third AI response text is enlarged and displayed. The first image is used to present the target flow associated with the third AI response text.
[0014] In one possible design, in another implementation of another aspect of the embodiments of this application, the first image includes multiple sub-images arranged in sequence; the response display unit is specifically used for: In response to a trigger operation on the image preview component, the first sub-image is enlarged and displayed. The first sub-image is the first frame sub-image in the first image and is used to describe the first stage in the target process. In response to the switching operation of the first sub-image, the second sub-image is enlarged and displayed. The second sub-image is used to describe the second stage in the target process. The first stage and the second stage are adjacent. The first sub-image and the second sub-image are two adjacent sub-images among the multiple sequentially arranged sub-images. In response to a switching operation for all said sub-images, the last sub-image is displayed, which is used to describe the end stage in the target process.
[0015] In one possible design, in another implementation of another aspect of the embodiments of this application, the content display component includes a video playback component; the response display unit is specifically used for: In response to a trigger operation on the video playback component, a target video associated with the third AI response text is played, the target video being used to present the target flow associated with the third AI response text.
[0016] In one possible design, in another implementation of another aspect of the embodiments of this application, the interactive processing device further includes an acquisition unit and a processing unit; The acquisition unit is specifically used to acquire the format syntax information and component configuration information of the first text format. The first text format is used to describe the text format of the AI response text. The component configuration information includes the component type and the component display method, or the component configuration information includes the component type, the component display method, and the component triggering behavior. The processing unit is specifically used for: Based on the format syntax information of the first text format and the component configuration information, construct the component protocol description information; Based on a preset large language model, the component protocol description information, candidate components, and the first query question are processed to generate target text content. The target text content is rendered to generate the M first response contents.
[0017] In one possible design, in another implementation of another aspect of the embodiments of this application, the processing unit is specifically used for: Based on the format structure of the first text format, the target text content is parsed to obtain the M AI response texts and each of the interactive components; Each AI response text is treated as a text node, and each interactive component is treated as a component node. Based on the logical relationship between the text nodes and the component nodes, the text nodes and the component nodes are constructed to generate a target tree data structure. The text nodes and component nodes in the target tree data structure are traversed and rendered to generate the M first response contents.
[0018] In one possible design, in another implementation of another aspect of the embodiments of this application, the processing unit is specifically used for: Based on the format structure of the first text format, the target text content is parsed to obtain the M AI response texts and each of the interactive components; Based on the logical relationship between each AI response text and the corresponding interactive component, the corresponding interactive component is inserted into a preset slot in the AI response text to generate the M first response contents.
[0019] In one possible design, in another implementation of another aspect of the embodiments of this application, the processing unit is further specifically used for: After processing the component protocol description information, candidate components and the first query question based on the preset large language model to generate target text content, the message display format of the target application is identified. When the message display format does not match the first text format, the message display format and the target text content are converted based on the preset large language model to obtain the converted target text content. The converted target text content is rendered to obtain M third response contents and an interactive component corresponding to each third response content, wherein each third response content includes AI response text; Display the content of the M third responses and the corresponding interactive components.
[0020] In another aspect, this application provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the methods described above.
[0021] Another aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods described above.
[0022] Another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the methods described above.
[0023] As can be seen from the above technical solutions, the embodiments of this application have the following advantages: In this embodiment, an interactive dialogue interface related to the AI interactive application is displayed in response to an interactive trigger operation by the target object. This interactive dialogue interface includes input controls. Further, in response to the target object's input operation in the input controls, a first query question is displayed. Finally, in response to the target object's trigger operation on the first query question, the first query question and M first responses to the first query question are displayed in the interactive dialogue interface, where M is an integer greater than or equal to 1. It should be noted that each first response in this application includes AI response text and an interactive component associated with the AI response text. Furthermore, the interactive component is used to guide the target object to perform a target action on the corresponding AI response text, and the AI response text is used to answer the first query question.
[0024] Through the above methods, after entering a first query question in the interactive dialogue interface, the relevant first response content can be displayed on the same interactive dialogue interface. From this first response content, one can not only understand the corresponding AI response text, but also directly guide and trigger the target object to perform subsequent operations such as querying and manipulating the AI response text through interactive components. On the one hand, by mixing and rendering the AI response text and associated interactive components in the same interactive dialogue interface, the limitations of basic text layout are broken, enabling complex layouts and multimodal expression, thus improving information transmission efficiency. On the other hand, this application treats the interactive component as an operation entry point, binding the target behavior to the interactive component. Without exiting the interactive dialogue interface, relevant business operations can be completed directly on the interactive dialogue interface, achieving immediate response and extreme simplification of operations, significantly improving user experience and query efficiency. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This illustration shows a schematic diagram of an embodiment of this application applied to a financial services scenario; Figure 2 This illustration shows a schematic diagram of an embodiment of this application applied to an e-commerce scenario; Figure 3 A schematic diagram of an implementation environment for the interactive processing method in an embodiment of this application is shown; Figure 4 This illustration shows another implementation environment diagram of the interactive processing method in the embodiments of this application; Figure 5 A flowchart illustrating an interactive processing method provided in an embodiment of this application is shown. Figure 6A A schematic diagram of an interface for an interactive processing method provided in an embodiment of this application is shown. Figure 6B Another interface diagram of the interactive processing method provided in the embodiments of this application is shown; Figure 7 This illustration shows an optional schematic diagram of an information container class component provided in an embodiment of this application; Figure 8A This illustration shows a schematic diagram of an interactive interface based on an information filtering slider provided in an embodiment of this application; Figure 8BThis illustration shows a schematic diagram of an interactive interface based on a progress bar slider provided in an embodiment of this application; Figure 9 This illustration shows an interface diagram based on an operation-triggered component provided in an embodiment of this application; Figure 10A This illustration shows a schematic diagram of an interface display based on an image preview component provided in an embodiment of this application; Figure 10B This illustration shows another interface display diagram based on the image preview component provided in an embodiment of this application; Figure 11 This illustration shows a schematic diagram of an interface based on a video playback component provided in an embodiment of this application; Figure 12 This application provides a flowchart illustrating a process for determining the content of M first responses. Figure 13 This illustration shows an optional schematic diagram of component protocol description information provided in an embodiment of this application; Figure 14 This illustration shows a schematic diagram of the target text content provided in an embodiment of this application; Figure 15A This illustration shows a schematic diagram of a rendering process provided in an embodiment of this application; Figure 15B This paper illustrates another framework diagram of the rendering process provided in an embodiment of this application; Figure 16 This paper illustrates another framework diagram of the rendering process provided in an embodiment of this application; Figure 17 The image shows the rendering effect of the application on different application platforms; Figure 18A This paper illustrates another framework diagram of the rendering process provided in an embodiment of this application; Figure 18B This paper illustrates another framework diagram of the rendering process provided in an embodiment of this application; Figure 19 This paper presents a schematic diagram of the overall processing framework of the interactive processing method provided in this application. Figure 20 A schematic diagram of the layered rendering and multi-platform adaptation framework provided in this application is shown. Figure 21 This paper presents a schematic diagram of the overall process of the interactive processing method provided in this application; Figure 22 A schematic diagram of one embodiment of the interactive processing device in this application is shown; Figure 23 A schematic diagram of the structure of a computer device provided in an embodiment of this application is shown. Detailed Implementation
[0027] This application provides an interactive processing method, apparatus, device, medium, and program product that not only breaks through the limitations of basic text layout, enabling complex layouts and multimodal expression and improving information transmission efficiency, but also allows for the direct completion of relevant business operations on the interactive dialogue interface, achieving immediate response and ultimate simplification of operations, significantly improving user experience and query efficiency.
[0028] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0029] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “corresponding to,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] The development of AI technology has driven a profound transformation in human-computer interaction, blurring the lines between information acquisition and business processing. When a user intends to inquire about a specific matter, they can use an AI system to initiate a request via natural language. This AI system automatically analyzes the user's intent, connects it to multi-source backend data, and generates a precise response based on the context. For example, when a user asks "How to transfer social security," the AI system can not only identify the keywords in the query but also dynamically generate relevant AI-generated response text, such as a "prepared materials checklist" and "processing instructions," based on contextual information like the user's location and insurance status. This response is then displayed visually on the user's interface.
[0031] However, current solutions for querying and displaying AI-generated response texts often use basic Markdown text format. Due to its limited expressive power, it only supports basic text formatting and cannot handle complex layouts, resulting in low information delivery efficiency. Furthermore, the AI-generated response only provides simple instructions for subsequent operations, such as "Go to the repayment page → Select bank card → Enter amount → Confirm repayment." However, the user needs to switch from the current customer service interface to the business function interface, relying on their memory of the previous instructions to complete the operation. Since the customer service interface and the business function interface are two completely independent interfaces, requiring frequent switching between them to complete the repayment operation results in numerous steps, poor efficiency, and a negative impact on user experience.
[0032] Therefore, while the relevant solutions can use AI to generate basic responses to address users' initial inquiry needs, their problems of simplistic information presentation and fragmented interaction processes remain prominent. Markdown text can only linearly display data and lacks integration with subsequent operation entry points, forcing users to manually filter content and navigate to other business function interfaces to complete subsequent actions. This operation path is lengthy and prone to errors. Especially when dealing with operational logic or complex business processes, users need to repeatedly switch between different interfaces to verify information, which can easily lead to operational errors due to memory bias, reducing operational accuracy and efficiency.
[0033] To address the aforementioned issues, this application provides an interactive processing method. This method achieves complex layouts and multimodal expression by blending and rendering AI-generated response text with associated interactive components within the same interactive dialogue interface, overcoming the limitations of basic text formatting and improving information transmission efficiency. Furthermore, this application treats interactive components as operation entry points, binding operations to these components. This allows for direct completion of relevant business operations within the interactive dialogue interface without exiting the interface, achieving immediate response and simplification of operations, significantly improving user experience and query efficiency. The interactive processing method of this application can be applied in at least one of the following scenarios.
[0034] Scenario 1: Financial and wealth management services scenario; In financial and wealth management service scenarios such as banking, securities, and insurance, AI-powered responses are reshaping customer experience with intelligent and personalized services. When a target inquires about financial matters through mobile banking or online platforms, the AI assistant can recognize the intent and provide accurate responses, whether it's account inquiries, wealth management recommendations, or loan application processes. Integrating the interactive processing methods in this application into financial and wealth management service scenarios allows for the mixed rendering and display of AI-generated response text and associated interactive components within the same interactive dialogue interface. Furthermore, by treating these interactive components as operation entry points, actions are bound to these components, enabling relevant business operations to be completed directly within the interactive dialogue interface without exiting it.
[0035] For example, Figure 1 This illustration shows a schematic diagram of an embodiment of this application applied to a financial services scenario. For example... Figure 1 As shown, when a user inquires about the issue of "payment not received," the interactive dialogue interface of this "Financial Intelligent Assistant" returns two responses to the question. Response 1 includes AI-generated text such as "XX Bank, Card Number 7766, Amount X1 Yuan, Time MM-D1 10:00, Status: Repayment Successful," and embeds a "View Details" button as an interactive component. Clicking this button expands the complete processing flowchart for the transaction. Response 2 includes AI-generated text such as "YY Bank, Card Number 8899, Amount X2 Yuan, Time MM-D2 14:30, Status: Processing," and simultaneously provides two interactive components: "Contact Customer Service" and "Repay." Users can directly click "Contact Customer Service" to immediately connect with human support, or select "Repay" to be directly redirected to the payment page to complete the operation. By deeply integrating information feedback with interactive components as an entry point, not only is the cognitive load on users reduced, but the business loop path is also significantly shortened, achieving a seamless connection from "question and answer" to "completion," truly reflecting the efficiency and convenience of intelligent services.
[0036] Scenario 2: E-commerce and retail scenarios; In e-commerce and retail scenarios such as product recommendations, promotional activities, and after-sales support, when a target user inquires about product-related issues through an AI interactive application, the application can accurately identify the user's intent and respond in real time by combining dynamic data such as product inventory and order status. Integrating the interactive processing method of this application into e-commerce and retail scenarios allows for the mixed rendering and display of AI-generated response text and associated interactive components within the same interactive dialogue interface. Furthermore, by treating the interactive components as operation entry points, the operation is bound to the interactive components, allowing relevant business operations to be completed directly on the interactive dialogue interface without exiting it.
[0037] For example, Figure 2 This illustration shows a schematic diagram of an embodiment of this application applied to an e-commerce scenario. For example... Figure 2 As shown, when a user inquires about the delay in shipping order #123456, the e-commerce intelligent assistant returns two responses in the interactive dialogue interface. Response 1 includes the AI-generated text: "The product is out of stock; the estimated restocking time is MM month 10th," and embeds an "Arrival Alert" button, allowing users to subscribe to inventory notifications with one click. Response 2 displays: "The backup item is in stock," and simultaneously shows a "Backup Item Product Image Display" preview component, which allows users to view the appearance and detailed parameters of the backup item. In this way, users can not only keep track of order status in a timely manner but also take corresponding actions based on real-time information, such as subscribing to arrival alerts or viewing recommended alternative products. By deeply integrating AI responses with interactive components, information acquisition and business processing are seamlessly connected on the same interface, greatly improving the functionality and user experience of the dialogue system.
[0038] Scenario 3: Entertainment and Media Scenarios; In entertainment and media scenarios such as content recommendation, interactive Q&A, and game NPCs, AI systems can recommend personalized content, such as movies, music, or game items, based on users' interests and preferences. Applying the interactive processing method of this application to entertainment and media scenarios allows AI response text and interactive components to be presented together in the same interactive dialogue interface. Furthermore, by treating the interactive components as operation entry points, operations are bound to the interactive components, allowing relevant business operations to be completed directly on the interactive dialogue interface without exiting it.
[0039] For example, when a user asks "What are some popular sci-fi movies to recommend recently?" in a movie-sharing application, the AI assistant returns a list of recommendations on the chat interface. Each recommendation includes a thumbnail of the movie poster, the movie title, rating, and a brief description, along with clickable buttons such as "Watch Now," "Add to Favorites," and "Share with Friends." When the user clicks the "Watch Now" button for a movie, the system automatically redirects to the playback page and loads the corresponding resource, all without leaving the original chat interface. By deeply integrating content display with the action entry point, a seamless transition from interest arousal to viewing is achieved, significantly enhancing user engagement and service immersion. Similarly, in a music recommendation scenario, when a user asks "What light music is suitable to listen to while working?", the AI assistant returns multiple audio cards on the chat interface. Each card includes the track title, performer information, and a 30-second preview, along with function buttons such as "Favorite," "Create Playlist," and "Timer Off." Users can directly listen to tracks, add to playlists, or share with contacts on the current interface, with all interactions continuing within the existing chat context. By deeply integrating multimedia content with functional components, users can operate efficiently without interrupting the communication context, further enhancing the immediacy of service response and the integration of interactive experience.
[0040] It should be noted that the above application scenarios are merely examples, and the interactive processing method provided in this embodiment can also be applied to other scenarios. For example, these include, but are not limited to, online education, medical consultation, travel, smart homes, and legal consultation, etc., and are not limited here.
[0041] It is understood that the AI interactive application involved in this application is an application with intelligent dialogue and real-time interaction capabilities. For example, the AI interactive application may include, but is not limited to, intelligent chatbots, intelligent bot assistants, intelligent agents, etc., which are not limited in this application.
[0042] The method provided in this application can be applied to different implementation environments, which will be described below with reference to the illustrations.
[0043] Implementation Environment 1: The method provided in this application can be applied to... Figure 3 The implementation environment shown includes at least a terminal 110. The terminal 110 involved in this application includes, but is not limited to, mobile phones, tablets, laptops, desktop computers, smart voice interaction devices, virtual reality devices, smart home appliances, vehicle terminals, and aircraft. The client is deployed on the terminal 110 and can run on the terminal 110 via a browser, a standalone application (APP), a mini-program, or a public account.
[0044] In the first implementation environment described above, in step A1, terminal 110, in response to the target object's interactive trigger operation on the AI interactive application, displays an interactive dialogue interface related to the AI interactive application. The interactive dialogue interface includes input controls. The interactive trigger operation can be a command issued by the target object in a text chat box using natural language; or it can be a command issued by the target object via voice, which is not limited in this application. In step A2, terminal 110, in response to the target object's input operation in the input controls, displays a first query question. This first query question can be understood as a query request made by the target object to obtain certain information. The query format can be based on natural language, keywords, or structured query statements, etc., which is not limited in this application. Thus, after displaying the first query question, the target object can perform a trigger operation on the first query question. Thus, in step A3, in response to the target object's triggering operation on the first query question, terminal 110 displays the first query question and M first responses to the first query question in the interactive dialogue interface. Each first response includes AI response text and an interactive component associated with the AI response text. It should be noted that the interactive component is used to guide the triggering target object to perform a target behavior on the corresponding AI response text. The described target behavior may include, but is not limited to, query behavior, operation behavior, etc., where operation behavior can be understood as performing operations such as displaying content or redirecting links on the AI response text, which is not limited in this application. Furthermore, the mentioned AI response text is used to answer the first query question. M is an integer greater than or equal to 1.
[0045] Implementation Environment Two: The method provided in this application can be applied to... Figure 4 The illustrated implementation environment includes a terminal 210 and a server 220, and the terminal 210 and server 220 can communicate via a communication network 230. The communication network 230 uses standard communication technologies and / or protocols, typically the Internet, but can also be any network, including but not limited to Bluetooth, a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), mobile, private networks, or any combination of virtual private networks. In some embodiments, customized or dedicated data communication technologies may be used to replace or supplement the aforementioned data communication technologies.
[0046] For details regarding terminal 210, please refer to the description of terminal 110 in the foregoing embodiments; it will not be repeated here.
[0047] The server 220 involved in this application can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence (AI) platforms.
[0048] In the second implementation environment described above, in step B1, the target object triggers the operation to launch the AI interactive application through terminal 210, for example, by clicking or long-pressing. In step B2, in response to the target object's interactive trigger operation on the AI interactive application, terminal 210 displays an interactive dialogue interface related to the AI interactive application, which includes input controls. In step B3, in response to the target object's input operation in the input controls, terminal 210 displays a first query question. This first query question can be understood as a query request made by the target object to obtain certain information. The query form can be based on natural language, keywords, or structured query statements, etc., and is not limited in this application. Thus, after displaying the first query question, the target object can perform a trigger operation on the first query question. In step B4, in response to the target object's trigger operation on the first query question, terminal 210 sends the first query question to server 220 through communication network 230. In step B5, server 220 determines M first responses to the first query question. Each first response includes AI response text and an interactive component associated with the AI response text. For example, server 220 can construct component protocol description information based on the text format of the AI response text and component configuration information. Then, by combining candidate components and the first query question, and using the prompting engineering of a preset large language model, target text content is generated. Rendering this target text content yields the M first responses. It should be noted that the interactive component is used to guide and trigger the target object to perform the target behavior on the corresponding AI response text. The AI response text is used to answer the first query question, and M is an integer greater than or equal to 1. Thus, in step B6, server 220 sends the M first responses to terminal 210 through communication network 230. In step B7, terminal 210 displays the first query question and the M first responses to the first query question in the interactive dialogue interface.
[0049] Based on the above introduction, the interactive processing method in this application will be described below. Please refer to [link / reference]. Figure 5The interactive processing method in this application embodiment can be completed independently by the terminal or in cooperation with the server. The method of this application includes: 501. In response to an interactive trigger operation of an AI interactive application, display an interactive dialogue interface related to the AI interactive application, the interactive dialogue interface including input controls.
[0050] In one or more embodiments, the interactive dialogue interface refers to a page where a target object interacts with an AI interactive application. Within this interactive dialogue interface, the target object can engage in interactive behaviors such as dialogue and communication with the AI interactive application.
[0051] When a target user intends to consult on certain business-related issues, the AI interactive application can be launched in response to an interaction trigger operation related to that business. The terminal then loads and displays the interactive dialogue interface related to that AI interactive application.
[0052] The AI interactive application provided in this application is an intelligent application capable of interacting, communicating, conversing, and providing information feedback to a target object. This AI interactive application can be a standalone application, such as an AI assistant application like ChatGPT. It can also be an AI interactive function module integrated into other applications, such as an intelligent assistant module embedded in office software, an AI customer service module embedded in instant messaging software, an intelligent customer service module integrated into an e-commerce platform, or a web-based intelligent dialogue plugin. Other chatbots, intelligent agents, etc., are not limited in this application.
[0053] The displayed interactive dialog interface includes at least an input control. The input control in this application is mainly used to receive query questions or interactive instructions input by the target object, and it supports multiple input modalities such as text input, voice input, and image upload, which are not specifically limited in this application.
[0054] 502. In response to an input operation in the input control, display the first query question.
[0055] In one or more embodiments, when the target object enters a first query question in the input control of the interactive dialog interface, the terminal will respond to the input operation in the input control and then display the first query question in the interactive dialog interface.
[0056] The first query question in this application refers to the query question submitted by the target object through an input control. This first query question can be about business inquiries, information retrieval, operational guidance, or other questions requiring AI assistance, and has clear semantic content and query intent. By entering the first query question, the terminal can clearly understand the target object's query needs and initiate subsequent interactive processing accordingly.
[0057] For example, the input of the first query question can be text input, where the target user types the question text into the input control. For instance, the first query question could be "Payment not received" or "How to check refund progress," etc., and this application does not limit this. Alternatively, the first query question can also be input via voice input, where the target user asks the question by voice input, such as saying "Why hasn't my order been shipped yet?" The terminal will then use a voice recognition model to recognize the voice and convert it into a text-based question. Alternatively, an image input method can be used, where the target user uploads an image containing the query question, and the terminal uses optical character recognition technology to extract the text content from the image and convert it into a processable first query question.
[0058] It should be noted that this application does not limit the input method for the first query question.
[0059] 503. In response to the trigger operation for the first query question, display the first query question and M first answers to the first query question in the interactive dialogue interface. Each first answer includes AI answer text and an interactive component associated with the AI answer text. The interactive component is used to guide the trigger to perform the target behavior on the corresponding AI answer text. The AI answer text is used to answer the first query question. M is an integer greater than or equal to 1.
[0060] In one or more embodiments, after a target object enters a first query question in the input control, the target object can further perform a triggering operation on the first query question, such as clicking or long-pressing the first query question, to confirm submission. Thus, the first query question and M first responses to the first query question can be displayed in the interactive dialog interface. Each first response includes corresponding AI response text and an interactive component associated with the AI response text.
[0061] The interactive components associated with the AI response text provided in this application are understood as entry points used to guide and trigger the execution of target behaviors on the corresponding AI response text, and are functional controls capable of enabling further operations on the AI response text. For example, the interactive components may exist in forms including but not limited to at least one of the following: information container components, operation triggering components, and content display components.
[0062] Among them, the information container component is used to hold and structure aggregated information related to AI responses. This aggregated information can reflect the multi-dimensional characteristics of the AI response text. This information container component includes, but is not limited to, cards, information panels, list items, etc., which can intuitively display key information in the AI response text in a block, category, or layered manner.
[0063] Action-triggered components provide interactive entry points to guide the target object to perform specific actions related to the AI's response text. For example, action-triggered components include, but are not limited to, buttons, links, and switches, which can respond to user clicks, swipes, long presses, and other actions to trigger corresponding functional flows, such as navigating to a specified page, launching a specific service, submitting feedback, or performing a task.
[0064] Content display components are used to display multimedia content related to the AI's response text, such as images and videos. These content display components can take various forms, including image preview components and video playback components, and are not limited to any particular form in this application.
[0065] In practical applications, interactive components associated with AI-generated text responses can exist independently or be combined as needed to collaboratively achieve more complex functional guidance and information delivery. For example, information container components can embed operation triggering components, providing immediate operation entry points while presenting structured information, thus improving interaction efficiency. Another example is the linkage between content display components and operation triggering components, which can overlay interactive functions on top of multimedia content, allowing users to complete actions such as sharing or downloading while browsing images or videos.
[0066] For ease of understanding, Figure 6A A schematic diagram of an interface for an interactive processing method provided in an embodiment of this application is shown. For example... Figure 6A As shown in Part (a), when the target object intends to query "finance" related business, it can first launch the AI interactive application of the "finance" related business through the terminal (e.g., intelligent assistant 601).
[0067] Next, as Figure 6A As shown in section (b), after performing an interactive triggering operation such as clicking on the smart assistant 601, the interactive dialogue interface 602 of the smart assistant 601 can be displayed. The interactive dialogue interface 602 includes an input control 603. Further, the target object can enter a first query question 604 in the input control 603, such as ("Payment not received").
[0068] Finally, as Figure 6AAs shown in section (c), after the input of the first query question 604 is completed, the first query question 604 can be displayed in the interactive dialogue interface 602, along with two responses to the first query question 604, such as response content 605 and response content 606. Response content 605 includes AI response text 6051 and an embedded information container component 6052. This information container component 6052 presents structured information related to "payment not received" in card form, including details such as transaction time, amount details, and processing progress. Response content 606 includes AI response text 6061 and the associated information container component 6062. This information container component 6062 presents actions related to "payment not received" in the form of buttons, such as "repay" or "contact customer service."
[0069] In this embodiment, by rendering and displaying the AI-generated response text and associated interactive components in the same interactive dialogue interface, the limitations of basic text layout are overcome, enabling complex layouts and multimodal expression, thus improving information transmission efficiency. Furthermore, this application treats the interactive components as an operation entry point, binding the operation to the interactive components. Without exiting the interactive dialogue interface, relevant business operations can be completed directly on the interactive dialogue interface, achieving immediate response and extreme simplification of operations, significantly improving user experience and query efficiency.
[0070] Optionally, in the above Figure 5 Based on one or more corresponding embodiments, in another optional embodiment provided by the present application, after displaying the first query question and M first answers to the first query question in the interactive dialog interface in response to a triggering operation for the first query question, it may further include: In response to a trigger operation on the first interactive component, a second query question is displayed in the interactive dialogue interface. The first interactive component and the second query question have a one-to-one relationship. The second query question is generated based on the business logic of the first interactive component. The first interactive component is an interactive component contained in any one of the M first response contents. In response to the second query, the second answer that matches the second query is displayed in the interactive dialog interface.
[0071] In one or more embodiments, a method is introduced to trigger deeper queries through interactive components, building upon the display of interactive components. As described above, the interactive components in the M first response contents are all associated with the corresponding AI response text. For each interactive component, a corresponding second query question is pre-configured according to the business logic of each component. This allows the target object to automatically trigger the second query question bound to the interactive component when it triggers the corresponding interactive component, eliminating the need for the target object to manually input query content into the input control. This enables rapid querying of relevant query questions.
[0072] Specifically, for ease of understanding, let's take an interactive component (i.e., a first interactive component) contained in any one of the M first response contents as an example. When the target object triggers this first interactive component, the second query question associated with the first interactive component can be automatically identified and displayed on the interactive dialog interface. Furthermore, it automatically responds to the second query question, thus directly presenting the second response content matching the second query question in the interactive dialog interface, without requiring the user to re-enter or navigate to another page.
[0073] For example, Figure 6B Another interface diagram of the interactive processing method provided in an embodiment of this application is shown. For example... Figure 6B As shown above, in the aforementioned Figure 6A Based on the displayed response content 605 and response content 606, the target user can click on the information container component 6052 in response content 605. Upon triggering, a second query question 607 bound to the information container component 6052 will be directly displayed in the interactive dialog interface 602, such as "I would like to inquire about this repayment order of X1 yuan from XX Bank 7766 dated D1 of MM month". After recognizing and responding to this second query question 607, a second response content 608 corresponding to this second query question will be displayed in the interactive dialog interface 602. For example, the second response content 608 could be "Your repayment order is currently in the status of successful repayment. The bank has processed your repayment application, and the credit limit has been restored in real time. You can perform the following operations..."
[0074] By employing the above method, triggering the interactive component automatically initiates the associated in-depth query question and retrieves the corresponding answer, eliminating the need for manual input or page redirection, thus significantly improving operational and information retrieval efficiency. Furthermore, this process is completed within the same session flow, ensuring operational continuity and contextual consistency, and significantly reducing the length of the user's operation path.
[0075] Optionally, in the above Figure 5Based on one or more corresponding embodiments, in another optional embodiment provided by this application, the interactive component may include an information container class component. Therefore, after displaying the first query question and M first responses to the first query question in the interactive dialog interface in response to a triggering operation for the first query question, the method further includes: In response to a trigger operation on an information container component, aggregated information of the first AI response text is displayed. The aggregated information of the first AI response text is used to present the multi-dimensional features of the first AI response text. Among them, the first AI response text is the AI response text contained in any one of the M first response contents, and the information container component is associated with the first AI response text.
[0076] In one or more embodiments, a triggering method between an information container component and AI response text is described. As can be seen from the foregoing, the interactive component may include an information container component. When a target object triggers an information container component, in response to the triggering operation, a first AI response text associated with the information container component can be extracted, and the first AI response text can be aggregated and displayed with multi-dimensional features to present aggregated information including the first AI response text.
[0077] Aggregated information can cover features across multiple dimensions. For example, it may include, but is not limited to, the confidence level of the AI response, generation time, data source, associated knowledge graph nodes, and recommended operation suggestions, enabling the target audience to comprehensively evaluate the reliability and applicability of the response.
[0078] It should be noted that the aggregated information in this application can be displayed through overlays, pop-ups, or interactive dialog interfaces, which can avoid the context break caused by page jumps and ensure the continuity of user operations and the integrity of information acquisition.
[0079] Additionally, information container components can include one or more of card components and list item components. Card components are primarily used to display structured information related to AI response text. List item components are suitable for presenting aggregations of multiple similar information items.
[0080] Specifically, for ease of understanding, Figure 7 This illustration shows an optional schematic diagram of an information container class component provided in an embodiment of this application. For example... Figure 7As shown, taking the information container component, including the card component, as an example, in the response content 605 to the first query question "Repayment not received," after the target object clicks on the information container component 6052 in the response content 605, the system can promptly respond to the click operation and display the aggregated information 609 of the AI response text 6051 corresponding to the "Repayment not received" question in the current interactive dialogue interface. From this aggregated information 609, it can be seen that it includes the multi-dimensional features of the AI response text 6051, such as: a confidence level of 98%, a generation time displayed as YYYY year MM month 8th 10:00, a data source labeled as "Bank Transaction System and User Agreement Document," and related knowledge graph nodes pointing to "Repayment Process Description" and "Receipt Time Limit Rules," and recommending the next step: "Contact bank customer service to confirm transaction status."
[0081] By binding the AI response text to an information container component and triggering the information container component, the multi-dimensional aggregated information of the AI response text can be automatically displayed. This allows the target audience to quickly find the contextual basis of the AI response text, thereby efficiently judging the credibility of the response and the execution path, and improving interaction efficiency and decision-making accuracy.
[0082] Optionally, in the foregoing Figure 5 Based on one or more of the described embodiments, in another embodiment provided by this application, after displaying aggregated information of the first AI response text in response to a triggering operation on an information container-type component, it may further include: In response to the configuration operation of the information filtering slider in the aggregated information for the first AI response text, the target filtering conditions are displayed; In response to the target filtering criteria, the first content is displayed, which is the information in the aggregated information of the first AI response text that meets the target filtering criteria.
[0083] In one or more embodiments, this embodiment provides an interactive method for dynamically filtering aggregated information using an information filtering slider. The aggregated information of the first AI response text includes an information filtering slider. The information filtering slider mentioned in this application mainly provides a configurable operation window for the target object to set filtering conditions to filter multi-dimensional features in the aggregated information.
[0084] Based on their query needs, the target user can configure one or more feature filtering conditions in the information filtering slider, thereby obtaining and displaying the target filtering conditions. Furthermore, in response to these target filtering conditions, information that meets the conditions can be filtered from the aggregated information of the first AI response text and presented as the primary content, thus achieving precise filtering and personalized display of aggregated information.
[0085] Specifically, for ease of understanding, Figure 8A This illustration shows a schematic diagram of an interactive interface based on an information filtering slider provided in an embodiment of this application. Figure 8A As shown above, Figure 7 Taking the first query question, "Repayment not received," as an example, from the aforementioned... Figure 7 It can be seen that the corresponding aggregated information 609 includes the multi-dimensional features of the AI response text 6051. Next, the target user slides the information filtering slider 6091 of the aggregated information 609 to select the keyword "data source labeling" as the target filtering condition. Thus, based on the target filtering condition of "data source labeling," all content fragments in the aggregated information 609 related to "data source labeling (i.e., the bank transaction system and user agreement document)" can be automatically filtered out, and this can be used as the first content 610. For example, the first content 610 includes the transaction record verification results from the "bank transaction system" and the terms and conditions related to the "user agreement document." Thus, in the overlay, the dynamically presented first content 610 is displayed in a highlighted form on top of the aggregated information 609, allowing users to intuitively view specific information fragments related to the "bank transaction system" and the "user agreement document."
[0086] Compared to the lengthy responses that require full-text browsing in traditional solutions, this embodiment uses an interactive design with an information filtering slider to allow users to quickly focus on information snippets that meet the configured target filtering criteria, thereby improving information acquisition efficiency and decision-making accuracy.
[0087] Optionally, in the foregoing Figure 5 Based on one or more of the described embodiments, in another embodiment provided by this application, after displaying aggregated information of the first AI response text in response to a triggering operation on an information container-type component, it may further include: In response to dragging the progress bar slider in the aggregated information of the first AI response text, the current information viewing progress is displayed; In response to the current information viewing progress, load the second content, which is an information fragment from the aggregated information of the first AI response text that matches the current information viewing progress.
[0088] In one or more embodiments, this embodiment provides an interactive method for dynamically filtering aggregated information using a progress bar slider. The aggregated information of the first AI response text includes a progress bar slider. The progress bar slider mentioned in this application mainly provides another operable window for the target object, which can present content fragments of different information dimensions in the aggregated information when sliding to different positions, so as to filter out information interference that is irrelevant to the current dimension.
[0089] The target object can be dragged to view the progress based on the query requirements. When the target object is dragged to a certain position, the current information viewing progress can be displayed, and based on the current information viewing progress, the second content corresponding to the current information viewing progress can be loaded from the aggregated information.
[0090] Specifically, for ease of understanding, Figure 8B This illustration shows a schematic diagram of an interactive interface based on a progress bar slider provided in an embodiment of this application. Figure 8B As shown above, Figure 7 Taking the first query question, "Repayment not received," as an example, from the aforementioned... Figure 7 It can be seen that the corresponding aggregated information 609 includes the multi-dimensional features of the AI response text 6051, and these multi-dimensional features are arranged in a structured manner. When the target object slides the progress bar slider 6092 in the aggregated information 609 and stops at position A, the progress bar slider 6092 displays the current information viewing progress as the "operation suggestion" dimension. Based on this, based on the "operation suggestion" dimension, specific content fragments related to "operation suggestion" are dynamically filtered from the aggregated information 609, which are then used as the second content 611 and displayed. For example, the second content 611 includes "Contact bank customer service to confirm transaction status".
[0091] Compared to the lengthy responses requiring full-text browsing in traditional solutions, this embodiment utilizes a progress bar slider design, allowing users to quickly focus on the information they are currently interested in, reducing interference from redundant information. Furthermore, it enables precise filtering and efficient browsing of aggregated information, improving user decision-making efficiency in complex scenarios.
[0092] Optionally, in the foregoing Figure 5 Based on one or more of the described embodiments, in another embodiment provided by this application, the interactive component includes an operation triggering class component. After displaying the first query question and M first responses to the first query question in the interactive dialog interface in response to a triggering operation for the first query question, the method further includes: In response to a trigger operation on an operation-triggered component, third content is displayed on the first information display page. The third content is related information about the second AI response text generated based on the business logic of the operation-triggered component. The second AI response text is the AI response text contained in any one of the M first response contents. The operation triggering component is associated with the second AI response text, and the second information display page is a related hierarchical page of the first information display page.
[0093] In one or more embodiments, an interaction method based on operation-triggered components is described. As mentioned above, the interaction components include operation-triggered components. In the embodiments of this application, operation-triggered components are mainly used to guide the target object to perform specific operations related to the AI response text. The operation-triggered components of this application include, but are not limited to, one or more of button components and link components. Both button components and link components can be used to trigger jumps to related hierarchical pages of the interactive dialogue interface, so as to display related information related to the AI response text on the related hierarchical pages, thereby achieving a seamless connection from information acquisition to operation execution.
[0094] Taking any one of the M first response contents as an example, its included AI response text is considered the second AI response text, and the operation triggering component is associated with this second AI response text. Thus, when the target object triggers the operation triggering component, it can respond to that triggering operation. Furthermore, based on the business logic of the operation triggering component, related information associated with the second AI response text is dynamically generated, and this related information is used as the third content. Therefore, this third content is displayed on the association hierarchy page of the interactive dialogue interface (i.e., the first information display page), realizing hierarchical expansion and deep interaction of information.
[0095] Specifically, for ease of understanding, Figure 9 This diagram illustrates an interface display of an operation-triggered component provided in an embodiment of this application. For example... Figure 9 As shown in section (a), taking an operation-triggered component including a button component and a link component as an example, for the first query question 900 "What are some recommended restaurants nearby?", the first response content generated after interacting with the AI interactive application displays AI response text including multiple recommended restaurants. For example, response content 901 includes AI response text 9011 (e.g., "A brief introduction to Sunshine Restaurant") and a button component 9012 (e.g., "Query Details" button). Similarly, for response content 902, it includes AI response text 9021 (e.g., "A brief introduction to Starlight Restaurant") and a link component 9022 (e.g., "Navigation Location" link).
[0096] As shown in section (b), when the target object clicks the "Query Details" button component 9012, the system responds to the trigger operation by dynamically generating third content related to the AI response text 9011 based on the business logic of the button component, and displays the third content on the information display page 903 (e.g., "Restaurant Details Page of Food Software"), such as detailed restaurant information, including business hours, menu recommendations, user reviews and reservation entry.
[0097] Alternatively, as shown in section (c), when the target object clicks the "Navigation Location" link component 9022, the system responds to this triggering operation by dynamically generating third content related to the AI response text 9021 based on the business logic of the link component, and displays the recommended navigation route for the "Starlight Restaurant" on the information display page 904 (e.g., the "Navigation Interface of the Map Application"). For example, Route 1, which is 30 kilometers away and takes 1 hour and 5 minutes, and Route 2, which is 38 kilometers away and takes 1 hour and 40 minutes. Through the above method, a coherent service loop from information query to actual arrival can be achieved.
[0098] By using the above method, the AI response text is dynamically associated with relevant information through operation-triggered components. After obtaining the initial information, users can directly switch from the current interactive dialogue interface to the corresponding information display page through operation-triggered components such as buttons and links to view deeper content. This achieves a seamless connection from query to operation, improves interaction and query efficiency, and effectively reduces the cognitive burden and operational cost for users switching between multiple applications.
[0099] Optionally, in the foregoing Figure 5 Based on one or more corresponding embodiments, in another embodiment provided by this application, the interactive component further includes a content display component; after displaying the first query question and M first answers to the first query question in the interactive dialogue interface in response to a trigger operation for the first query question, it may further include: In response to a trigger operation on a content display component, multimedia display information associated with the third AI response text is displayed. The multimedia display information associated with the third AI response text is used to present the target flow of the third AI response text. The target flow includes at least one of a business operation flow and a content generation flow. The third AI response text is the AI response text contained in any one of the M first response contents, and the content display component is associated with the third AI response text.
[0100] One or more embodiments introduce an interactive extension method based on content display components. As described above, the interactive components include content display components, which are primarily used to present multimedia information related to AI-generated response text, visually showcasing the complete path of business operation processes or content generation processes in the form of images, videos, etc. Furthermore, these content display components can exist in forms such as image preview components and video playback components.
[0101] Taking any one of the M first response responses as an example, its included AI response text is considered the third AI response text, and the content display component is associated with this third AI response text. When the target object triggers the content display component, the system dynamically retrieves the multimedia display information associated with the third AI response text according to the business logic of the content display component, and displays it through a floating layer, pop-up window, or full-screen mode to ensure that users can intuitively obtain the complete information of the target process without jumping to another page.
[0102] It should be noted that the target process may include at least one of a business operation process and a content generation process. In the embodiments of this application, the business operation process can be understood as the operational path surrounding the business matters involved in the third AI response text, such as approval processes and processing steps. The content generation process refers to the generation path that, based on requirements, structures and visualizes the information involved in the third AI response text, such as report generation and chart creation.
[0103] By deeply integrating content display components with third-party AI response text, as described above, complex processes can be presented intuitively and instantly accessed without departing from the original interaction scenario, significantly improving users' understanding efficiency and ease of operation of AI response text. Furthermore, by displaying multimedia information through the third-party AI response text, the execution flow of the target process can be clearly reconstructed, enhancing the interactive immersion and the accuracy of information delivery.
[0104] Optionally, in another embodiment, the content display component includes at least one of an image preview component and a video playback component. The image preview component of this application is mainly used to display static diagrams related to the target process, such as flowcharts, architecture diagrams, or screenshots of key nodes, and can achieve magnified display of thumbnails. The video playback component of this application is mainly used to present dynamic operation demonstrations or explanatory videos related to the target process, intuitively presenting the entire process of business operations or content generation.
[0105] The following sections will further explain the association mechanism and interaction implementation method between the image preview component and the video playback component and the third AI response text.
[0106] Scenario 1: Image preview component; Optionally, in the foregoing Figure 5 Based on one or more of the described embodiments, in another embodiment provided by this application, the content display component includes an image preview component; in response to a trigger operation on the content display component, multimedia display information associated with the third AI response text is displayed, including: In response to a trigger action on the image preview component, the first image associated with the third AI response text is enlarged and displayed. The first image is used to present the target flow associated with the third AI response text.
[0107] In one or more embodiments, when the content display component includes an image preview component, since the image preview component embeds a thumbnail of the target process, when the target object clicks on the image preview component, the system automatically loads and enlarges the thumbnail, thereby displaying the enlarged first image, which is associated with the third AI response text. Through this first image, the target process associated with the third AI response text can be clearly presented, allowing users to intuitively grasp the process structure and key node information. Optionally, the enlarged first image in this application supports gesture zooming and swiping browsing, facilitating users to view process details.
[0108] Specifically, for ease of understanding, Figure 10A This illustration shows a schematic diagram of an interface display based on an image preview component provided in an embodiment of this application. For example... Figure 10A As shown, taking the first query question 1000, "What is the handling fee discount?", as an example, after interacting with the AI interactive application, the generated first response content includes AI response text 1001 and a related image preview component 1002. AI response text 1001 could be something like, "Handling fee discounts refer to the XX platform's efforts to help users save money under specific conditions...; Business account repayment: If you are a merchant on the XX platform, using [business account] for repayment can be fee-free indefinitely. The specific operation is as follows:...". The image preview component 1002 displays a thumbnail corresponding to the business operation process associated with AI response text 1001: "Card payment → Business code payment → Repayment within XX." When the target user clicks on the image preview component 1002, the system responds to the click operation, displaying a magnified image corresponding to the thumbnail, i.e., the first image 1003, fully presenting the entire business process.
[0109] By utilizing the image preview component described above, the thumbnail of the target process can be enlarged into a high-resolution image, allowing users to clearly identify each step and logical relationship within the process. This helps users quickly understand the business operation path, reduces cognitive costs, and improves interaction efficiency. Especially in complex process guidance scenarios, this application only requires triggering the image preview component in the query results to clearly view the complete process diagram, avoiding the cumbersome operation and operational errors caused by frequent page switching in traditional solutions, significantly improving the convenience and accuracy of user operations.
[0110] Optionally, in another embodiment, the presented first image further includes multiple sequentially arranged sub-images, each sub-image describing a key stage in the target process. That is, these sequentially arranged sub-images can completely present the target process. Based on this, in another embodiment provided by this application, in response to a trigger operation on the image preview component, enlarging and displaying the first image associated with the third AI response text includes: In response to a trigger operation on the image preview component, the first sub-image is enlarged and displayed. The first sub-image is the first frame sub-image in the first image, and the first sub-image is used to describe the first stage in the target process. In response to the switching operation of the first sub-image, the second sub-image is enlarged and displayed. The second sub-image is used to describe the second stage in the target process. The first stage and the second stage are adjacent. The first sub-image and the second sub-image are two adjacent sub-images among multiple sub-images arranged in sequence. In response to a toggle operation on all sub-images, the last sub-image is displayed. This last sub-image is used to describe the final stage in the target process.
[0111] In one or more embodiments, an image preview method that switches frame by frame and zooms in is described. As described above, different key stages in the target process are treated as different frame images. Following the process sequence, in response to the target object's trigger operation on the image preview component, the first frame sub-image (i.e., the first sub-image) is zoomed in and displayed. Thus, the target object can clearly see the specific content of the first stage of the target process from this first sub-image. Next, the target object performs a switching operation such as sliding or clicking on the first sub-image to zoom in and display the second sub-image adjacent to it, thereby viewing the content of the second stage of the target process. This process continues until the target object performs a switching operation such as sliding or clicking on the second-to-last frame sub-image to zoom in and display the last frame sub-image, thereby viewing the end result of the target process.
[0112] Specifically, for ease of understanding, Figure 10B This illustration shows another interface display diagram based on the image preview component provided in an embodiment of this application. For example... Figure 10B As shown above, Figure 10A Taking the image preview component 1002 shown as an example, this component displays a thumbnail corresponding to the business operation process of "card payment → QR code payment → repayment within XX days" associated with the AI response text 1001. Moreover, it can be seen from the thumbnail that the process includes three stages of visual frame images: the first sub-image 10031 for the "card payment" stage, the second sub-image 10032 for the "QR code payment" stage, and the third sub-image 10033 for the "repayment within XX days" stage.
[0113] At this point, when the target user clicks on the image preview component 1002, in response to this click, the first sub-image 10031 is enlarged and displayed on the overlay, presenting detailed operation instructions for the "Card Payment" stage. Next, when the target user swipes left, the second sub-image 10032 is enlarged and displayed, presenting detailed operation instructions for the "Business Code Payment" stage. Further, when the target user continues to swipe left, the system responds with a switching action, enlarging and displaying the third sub-image 10033, fully presenting the operation details of the "Repayment within XX" stage, achieving a frame-by-frame progressive display of the overall target process.
[0114] Optionally, the target object can also switch to other sub-images while zooming in on any sub-image, allowing for free navigation between process nodes. Alternatively, it can switch to the previous frame to revisit the operation guidance from the previous stage.
[0115] In this application, the complex target process is broken down into multiple consecutive visual frame images through the image preview component, enabling the target object to browse the content of each stage frame by frame as needed, thereby clearly grasping the operation context of the entire process.
[0116] Scenario 2: Video playback component; Optionally, in the foregoing Figure 5 Based on one or more of the described embodiments, in another embodiment provided by this application, the content display component includes a video playback component; in response to a trigger operation on the content display component, multimedia display information associated with the third AI response text is displayed, including: In response to a trigger operation on the video playback component, the target video associated with the third AI response text is played. The target video is used to present the target flow associated with the third AI response text.
[0117] In one or more embodiments, an interaction method based on a video playback component is described. The video playback component provided in this application can dynamically present a target flow related to a third AI response text. Based on this, when the target object triggers the video playback component, the system automatically loads and plays the target video associated with the third AI response text. This target video can dynamically present the target flow associated with the third AI response text.
[0118] Specifically, for ease of understanding, Figure 11 This illustration shows a schematic diagram of an interface display based on a video playback component provided in an embodiment of this application. For example... Figure 11As shown, taking the first query question 1000, "What are the handling fee discounts?", as an example, after interacting with the AI interactive application, the generated first response content includes AI response text 1001 and a related video playback component 1102. This video playback component 1102 embeds the target video corresponding to the business operation process "card payment → business code payment → repayment within XX days" associated with the AI response text 1001. Therefore, when the target user clicks on the video playback component 1004, the system automatically plays the target video, displaying the operation details of each stage segment by segment.
[0119] By utilizing the video playback component described above, the target process associated with the third-party AI response text is dynamically presented. This allows the target user to intuitively view the complete execution path of the target process, enhancing the perceptibility and immersive experience of the operation guidance. Compared to the static frame-by-frame display presented by the image preview component, the video playback component more realistically recreates the actual operation scenario through continuous dynamic images, making it particularly suitable for process guidance that requires a high degree of temporal sequence or action continuity. For example, in the "Business Code Payment" process, the entire dynamic interaction from user scanning the code, entering the amount, confirming payment, to receiving the result can be clearly presented, improving understanding efficiency and execution accuracy.
[0120] Optionally, in the foregoing Figure 5 In one or more embodiments described, and in another embodiment provided in this application, the determination of the aforementioned M first responses can be implemented in the following manner. That is, Figure 12 This application illustrates a flowchart for determining the content of M first responses. Figure 12 As shown, it includes at least the following steps: S1201. Obtain the format syntax information and component configuration information of the first text format. The first text format is used to describe the text format of the AI response text. The component configuration information includes the component type and the component display method, or the component configuration information includes the component type, the component display method and the component triggering behavior.
[0121] In one or more embodiments, the first text format of this application can be understood as the text format that the generated AI response text needs to conform to. The format syntax information of the first text format includes paragraph structure, punctuation specifications, and text layout rules, etc., to ensure the consistency and standardization of the subsequent AI response text in terms of presentation.
[0122] For example, the first text format includes, but is not limited to, structured text formats such as Markdown. Taking Markdown as an example, its formatting syntax information includes heading levels, list identifiers, bold and italic markers, code block symbols, and other syntax rules, which are used to clearly define the text style and structural layout in the AI response text.
[0123] The component configuration information in this application includes component type and component display method. Alternatively, in addition to component type and display method, the component configuration information may further include component triggering behavior. Component type defines the specific type of interactive component embedded in the response content, such as button component, card component, image preview component, video playback component, etc. Component display method defines the presentation form of the interactive component in the response content, such as text, image, form, etc. Component triggering behavior sets the response logic when the target object interacts with the interactive component, such as URL redirection links, business interface calls, client interface invocation, etc.
[0124] S1202. Construct component protocol description information based on the format syntax information and component configuration information of the first text format.
[0125] In one or more embodiments, after obtaining the format syntax information and component configuration information of the first text format, component protocol description information can be constructed and generated based on the format syntax information and component configuration information of the first text format. Through this component protocol description information, the structured output specifications of the AI response text in the response content and the integration logic of interactive components can be defined, ensuring that the AI-generated content accurately embeds interactive components of specified types, display methods, and triggering behaviors while conforming to text format requirements. This component protocol description information, as the core basis for parsing and rendering, can be recognized by the front-end engine and converted into visual interface elements, realizing the mapping from semantics to the interface.
[0126] For example, Figure 13 This illustration shows an optional schematic diagram of component protocol description information provided in an embodiment of this application. For example... Figure 13 As shown, this interactive component is an image preview component, its component type is defined as "image", its display method is configured as the text "View Image Guide", and its trigger behavior is set as the interaction logic of "url: xxx.png". Based on this, the generated component protocol description information can be " <comp> {"desc":"",type":"image",data":{"url":"xxx.png",text":"Guide to view image"}}< / comp> The component protocol description information is encapsulated in a structured data format to ensure accurate identification and extraction during semantic parsing. After receiving the AI-generated response text, the front-end engine determines the component type based on the "type field" defined in the protocol, obtains the display content and interaction parameters through the "data field," and completes the interface rendering by combining auxiliary information such as the description (desc). For example, when the parsed type is "image," the engine will create an image preview component, embed it into the AI response text "View Image Guide," and bind a click event to trigger URL redirection or zoom in preview behavior, achieving seamless integration of content display and user interaction.
[0127] S1203. Based on the preset large language model, process the component protocol description information, candidate components and the first query question to generate the target text content.
[0128] In one or more embodiments, after the component protocol description information is constructed, the component protocol description information, candidate components, and the first query question are input into a preset large language model. The preset large language model uses the prompt engineering mechanism to perform semantic alignment and context fusion of the above three elements, generating target text content that conforms to the user's query intent, embeds component embedding logic, and conforms to the first text format.
[0129] The candidate components mentioned in this application can be selected from the list of available components based on the component protocol description information, choosing those that match the description information. Alternatively, components associated with historical question-and-answer pairs related to the first query question can be extracted from the QA knowledge base as candidate components. Alternatively, candidate components can be pre-configured within a specific process. For example, a specific process can be understood as a process configured for a specific business scenario, such as the return and exchange process or order query process in an intelligent customer service scenario. Components pre-configured within such processes can be selected as candidate components. By determining the source strategy for candidate components, it can be ensured that the generated target text content not only conforms to the current query context but also accurately associates with interactive components. During processing, the pre-configured large language model uses the first query question as semantic guidance, combining the structured parameters in the component protocol description information with the actual functional characteristics of the candidate components to perform multi-dimensional reasoning and content arrangement. The final output target text content maintains the fluency of natural language expression while embedding component calling logic, allowing users to directly trigger corresponding interactive operations while obtaining information, improving the overall practicality and operational efficiency of the response.
[0130] For example, Figure 14 This illustration shows a schematic diagram of the target text content provided in an embodiment of this application. For example... Figure 14 As shown, when the target user asks "Which repayment order would you like to inquire about?", based on the candidate components, component protocol description information, and the query question, the corresponding target text content can be generated, for example: { "result": "Which repayment order would you like to inquire about?" <comp> {"desc":"Repayment Records","type":"list","data":{"type":"recordlist","startDay":"30"}}< / comp> " } in," <comp>The structured data embedded in the tag describes the type and parameters of the interactive component. The "desc" field identifies the component's functional semantics, "type" being "list" indicates that its presentation format is a list selection, and "data" defines the specific data source type and time range parameters. This target text content can be automatically rendered into a visual selection control during front-end parsing. When the user clicks it, the corresponding interface is called to retrieve and display the repayment records for the past 30 days, achieving a seamless connection between semantic understanding and operation execution.
[0131] It should be noted that the above Figure 14 The target text content shown is for illustrative purposes only. In actual generation, the component type, parameter range, and interaction method can be dynamically adjusted according to specific business needs to ensure accurate response to user intent in diverse scenarios. Furthermore, the preset large language model in this application can be an LLM model, and this application does not impose any limitations on it.
[0132] S1204. Render the target text content to generate M first response contents.
[0133] In one or more embodiments, since the target text content contains parameters and other content related to interactive components, after the target text content is generated by calling the preset large language model, the target text content can be rendered, the embedded component tags can be parsed into interactive components that can be recognized by the front end, and each interactive component can be bound to the corresponding AI response text to generate M first response contents.
[0134] Optionally, for rendering scenarios that support the first text format, the ability to parse component tags may differ due to variations in the client environment of the AI interactive application. Therefore, format compatibility adaptation is required during rendering to ensure that interactive controls can be correctly presented on different terminals. Based on this, different embodiments will be described below.
[0135] Method 1: Rendering based on tree data structure; As an illustrative description, Figure 15A A schematic diagram of a rendering process provided in an embodiment of this application is shown. Figure 15A As shown, in the process of rendering the target text content to generate M first response contents, the target text content can first be parsed based on the format structure of the first text format to obtain M AI response texts and each interactive component. Then, each AI response text is treated as a text node, and each interactive component as a component node. Based on the logical relationship between the text nodes and component nodes, the text nodes and component nodes are constructed to generate a target tree data structure. Finally, the text nodes and component nodes in the target tree data structure are traversed and rendered to generate the M first response contents.
[0136] For example, Figure 15B This illustration shows another framework diagram of the rendering process provided in an embodiment of this application. (See diagram below.) Figure 15B As shown, the target text content includes multi-level nested content, such as: #1 heading 1. Ordered list item 1 2. Ordered list item 2 1. <comp> {"desc":"Click to open the official website of xx","type":"button","data":{"type":"navigate","url":"https: / / xx.com","text":"Go to the official website"}}< / comp> " 2. Sublist items Plain text.
[0137] In this structure, "#1 level heading" serves as the text node under the root node, "ordered list item 1" and "ordered list item 3" are child text nodes, and the "button component" generated after parsing the component tag serves as the component node under ordered list item 3. "Child list items" are sibling text nodes, and "plain text" content is the last-level text node, forming a tree-like hierarchical structure. By performing a depth-first traversal of this target tree data structure, styles are injected and layouts are applied to each text node and component node in sequence. Finally, the structured content containing interactive buttons is rendered on the front-end interface, achieving accurate presentation and functional binding of command controls in AI responses.
[0138] The above approach is applicable to scenarios such as mini-program platforms, enabling efficient parsing and consistent cross-platform display of interactive components in AI-generated content, especially ensuring the synchronization of rendering performance and user operation response in lightweight container environments. This solution abstracts the node structure, uniformly managing the hierarchical relationship between text and interactive components, effectively reducing the complexity of multi-platform adaptation and improving user experience consistency.
[0139] Method 2: Slot-based rendering; As an illustrative description, Figure 16 This illustration shows another framework diagram of the rendering process provided in an embodiment of this application. (See diagram below.) Figure 16 As shown, in the process of rendering the target text content to generate M first response contents, the target text content can first be parsed based on the format structure of the first text format to obtain M AI response texts and each interactive component. Then, based on the logical relationship between each AI response text and its corresponding interactive component, the corresponding interactive component is inserted into a preset slot in the AI response text to generate M first response contents.
[0140] The above method can be applied to lightweight operating environments such as mini-program agent platforms to achieve dynamic binding and partial updates of AI-generated content and interactive components.
[0141] For example, Figure 17 The image shows the rendering effects of the application on different application platforms. From Figure 17 As can be seen in part (a), in applications such as mini-programs, the rendering based on the aforementioned tree data structure yields responses including AI-generated text and embedded interactive components (e.g., a "Discount Card Zone" link, an "image preview component," etc.). Similarly, from... Figure 17 As can be seen in part (b), in applications such as mini-programs that support Agent mode, the aforementioned slot-based rendering results in AI-generated response text as well as embedded interactive components (e.g., "Discount Card Zone" link, "Image Preview Component," etc.). Parts (a) and (b) demonstrate that the mixed rendering effect remains consistent.
[0142] Optionally, in the foregoing Figure 12 Based on one or more embodiments described herein, this application Figure 18A In another embodiment provided, after processing the component protocol description information, candidate components, and the first query question based on a preset large language model to generate the target text content, the method may further include: Identify the message display format of the target application; When the message display format does not match the first text format, the message display format and the target text content are converted based on the preset large language model to obtain the converted target text content. The converted target text content is rendered to obtain M third response contents and an interactive component corresponding to each third response content. Each third response content includes AI response text. Displays M third-party responses and their corresponding interactive components.
[0143] In one or more embodiments, when certain target applications are limited by rendering capabilities or security policies and cannot directly parse the target interactive content in the first text format, it is necessary to identify the message display format of the target application and dynamically convert the target text content in the first text format into content adapted to the message display format of the target application through a preset large language model. That is, the text format of the converted target text content matches the message display format of the target application. In this way, rendering the converted target text content can ensure that M third response contents and the corresponding interactive components of each third response content can be displayed correctly in the target application.
[0144] It's worth noting that the M third responses here only include the corresponding AI response text and do not contain interactive components. That is, the interactive components are displayed independently of the third response content. Compared to the previous approach of embedding the AI response text and interactive components within the same response content, this approach decouples the interactive components from the AI response text, making it suitable for cross-platform applications or other application environments that do not support the first text format. This improves system compatibility and expands usage scenarios.
[0145] For example, Figure 18B This illustration shows another framework diagram of the rendering process provided in an embodiment of this application. (See diagram below.) Figure 18B As shown above, Figure 15B Taking the target text content as an example, after the message display format is recognized and the content format is converted by the target application, it can render three messages: text message 1701 is "Title & Content", text message 1702 is "Card Message: Click to open the xx official website", and text message 1703 is "Supplementary Content". Text message 1701 displays the core response content, text message 1702 presents an interactive link in card format that redirects to the xx official website, and text message 1703 provides supplementary information. The three messages are presented independently, with the interactive components and AI response text displayed separately, ensuring correct rendering across different applications.
[0146] Optionally, Figure 19 A schematic diagram of the overall processing framework of the interactive processing method provided in this application is shown. Figure 19 As shown, this processing framework includes at least a component definition stage, a model processing stage, and a front-end rendering stage. Optionally, it may also include a back-end data population stage.
[0147] In the component definition phase, it is mainly used to expand the format syntax information of the first text format. <comp>Tags are used to define component protocol description information for various interactive components.
[0148] Thus, during the model processing phase, the pre-defined large language model identifies candidate components from the list of available components, component information returned by specific processes, or the QA knowledge base. Based on the component protocol description information, and combined with the candidate components and the first query question, the pre-defined large language model is invoked to generate target text content containing component-related information.
[0149] Optionally, during the backend data population stage, component tags in the target text content are first identified through component protocol description information, and then the corresponding service is called according to business logic to populate the real-time business data into the predefined structure of the component, thereby achieving dynamic data binding.
[0150] Thus, during the front-end rendering stage, the target text content can be converted into a tree data structure of text nodes and component nodes. Rendering using this tree data structure allows for the mixed rendering of AI response text and interactive components. For details, please refer to Method 1 above; it will not be elaborated here. Alternatively, after generating the target text content, corresponding interactive components can be inserted into preset slots for the AI response text, thus rendering M first response contents. For details, please refer to Method 2 above; it will not be elaborated here.
[0151] Additionally, during the model processing stage, after generating the target text content, the format of the target text content can be converted to adapt to the message display formats of different application terminals. The converted target text content is then sent to the front-end rendering stage via message for rendering, resulting in a separate rendering result for the AI response text and interactive components; that is, the AI response text and interactive components are displayed using different messages. See the aforementioned section for details. Figure 18A or Figure 18B The content described is for your understanding and will not be elaborated upon here.
[0152] For example, taking Markdown format as an example, Figure 20 This diagram illustrates the framework for layered rendering and multi-platform adaptation provided in this application. Figure 20 As shown, the framework includes at least a rendering layer, a business service layer, an agent layer, a knowledge layer, and a component protocol layer. The rendering layer is used for content display across multiple platforms, supporting rendering functions for various terminals, including Markdown, mini-program rendering mode, mini-program agent rendering mode, custom components, and message rendering mode that does not support Markdown. The business service layer is mainly responsible for data population and message sending. The agent layer is mainly responsible for content generation and protocol compatibility. The knowledge layer mainly stores the QA knowledge base, a list of available components, fixed return information for specific processes, and a business process knowledge graph. The component protocol layer is mainly used to define the component types of interactive components, ensuring consistency in component calls across different platforms and scenarios.
[0153] Optionally, Figure 21 A schematic diagram of the overall process of the interactive processing method provided in this application is shown.
[0154] like Figure 21 As shown, firstly, the target object triggers the launch of the AI interactive application. In response to the interaction triggered by the AI interactive application, an interactive dialogue interface containing input controls is displayed. Subsequently, the target object performs an input operation in the input controls, and in response to the input operation in the input controls, the first query question is displayed.
[0155] Next, in response to the triggering operation for the first query question, the format syntax information and component configuration information of the first text format can be obtained. It should be noted that the first text format can reflect the text format of the AI's response text, and the component configuration information includes the component type and component display method, or the component configuration information includes the component type, component display method, and component triggering behavior. Then, based on the format syntax information and component configuration information of the first text format, component protocol description information is constructed. After obtaining this component protocol description information, a preset large language model is invoked to process the component protocol description information, candidate components, and the first query question, generating the target text content. Then, the rendering processing of this target text content is completed, generating M first response contents for the first query question. For details, please refer to the aforementioned... Figures 12 to 17 The content described herein will be understood in detail here, and will not be elaborated upon further.
[0156] Thus, after generating M first responses to the first query question, the first query question and the M first responses can be displayed in the same interactive dialog interface. Each first response includes the AI response text and the interactive components associated with that AI response text.
[0157] Furthermore, after displaying M first responses, different operations can be performed on these M first responses. The specifics can be understood by referring to the following different scenarios: Scenario 1: In response to a trigger operation on the first interactive component, a second query question is displayed in the interactive dialog interface. It should be noted that the first interactive component is the interactive component contained in any one of the M first response contents. The second query question is generated based on the business logic of the first interactive component and has a one-to-one association with it. Therefore, after displaying the second query question, the second response content matching the second query question can be directly displayed in the interactive dialog interface.
[0158] Scenario 2: Since the interactive component includes an information container component, this application can display aggregated information of the first AI response text in response to a trigger operation on the information container component. This aggregated information allows understanding of the multi-dimensional characteristics of the first AI response content. It should be understood that the first AI response text is the AI response text contained within any response content and is associated with the information container component.
[0159] Specifically, the aggregated information of the first AI response text also includes an information filtering slider or a progress bar slider. Based on this, in response to the configuration operation of the information filtering slider in the aggregated information, the target filtering conditions can be displayed, and then, in response to the target filtering conditions, information in the aggregated information that meets the target filtering conditions (i.e., the first content) can be displayed. Alternatively, in response to the dragging condition of the progress bar slider in the aggregated information, the current information viewing progress can be displayed, and then, in response to the current information viewing progress, information fragments in the aggregated information that match the current information viewing progress (i.e., the second content) can be loaded.
[0160] Scenario 3: Since interactive components can also include operation-triggered components, this application can also respond to trigger operations on operation-triggered components by displaying third content on the first information display page. The first information display page is a related hierarchical page of the interactive dialogue interface. It should be understood that the third content is related information generated based on the business logic of the operation-triggered component and related to the second AI response text. The second AI response text is the AI response text contained in any one of the M first response contents, and the operation-triggered component is associated with the second AI response text.
[0161] Scenario 4: Interactive components may also include content display components. Based on this, the application can respond to trigger operations on the content display component by displaying multimedia display information associated with the third AI response text. This multimedia display information can present the target flow of the third AI response content, such as one or more flows from business operation processes and content generation processes. It should be understood that the third AI response text is the AI response text contained in any response content and is associated with the content display component.
[0162] Specifically, the content display component can be either an image preview component or a video playback component. Based on this, in response to a trigger operation on the image preview component, the first image used to present the target flow associated with the third AI response text can be enlarged and displayed. Alternatively, in response to a trigger operation on the video playback component, the target video used to present the target flow associated with the third AI response text can be played.
[0163] It should be noted that the above Figure 21 The content described herein can also be referred to in the foregoing. Figures 5 to 20 The content described in any of the embodiments is for understanding purposes only, and will not be repeated here.
[0164] In this embodiment, a seamless intelligent interaction process integrating "intelligent answering" and "business processing" is achieved through multi-terminal compatible hybrid rendering technology. This process not only upgrades AI's text responses to a hybrid mode rich in operable components (such as buttons, cards, and jump links), significantly reducing the steps from user inquiry to completion and greatly improving interaction efficiency, but also completely solves the problem of multi-terminal rendering consistency through extended protocols and layered adaptation strategies, ensuring a unified and smooth cross-platform user experience. Furthermore, this solution possesses high reusability and scalability, can be reused in multiple interactive applications, providing a replicable benchmark practice for intelligent customer service systems in the financial and other industries, and promoting technological breakthroughs and comprehensive improvement in service efficiency of intelligent interactive systems.
[0165] The foregoing primarily describes the solutions provided by the embodiments of this application from a methodological perspective. It is understood that to achieve the above functions, corresponding hardware structures and / or software modules are included to execute each function. Those skilled in the art should readily recognize that, based on the modules and algorithm steps described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0166] The interactive processing apparatus in this application is described in detail below. Please refer to [link / reference]. Figure 22 , Figure 22 This illustration shows a schematic diagram of an embodiment of the interactive processing device in this application. The interactive processing device includes: The response display unit 2201 is used to display an interactive dialogue interface related to the AI interactive application in response to an interactive trigger operation of the AI interactive application. The interactive dialogue interface includes input controls. The response display unit 2201 is used to display a first query question in response to an input operation in the input control; The response display unit 2201 is used to display the first query question and M first response contents of the first query question in the interactive dialogue interface in response to the trigger operation for the first query question. Each first response content includes an artificial intelligence (AI) response text and an interactive component associated with the AI response text. The interactive component is used to guide the trigger to perform a target behavior on the corresponding AI response text. The AI response text is used to answer the first query question, and M is an integer greater than or equal to 1.
[0167] Optionally, in the above Figure 22 Based on the corresponding embodiments, in another embodiment of the interactive processing device provided in this application, the response display unit 2201 is further specifically used for: After displaying the first query question and M first responses to the first query question in the interactive dialog interface in response to a trigger operation for the first interactive component, a second query question is displayed in the interactive dialog interface in response to a trigger operation for the first interactive component. The first interactive component and the second query question have a one-to-one association relationship. The second query question is generated based on the business logic of the first interactive component. The first interactive component is an interactive component included in any one of the M first responses. In response to the second query question, a second answer matching the second query question is displayed in the interactive dialog interface.
[0168] Optionally, in the above Figure 22 Based on the corresponding embodiments, in another embodiment of the interactive processing device provided in this application, the interactive component includes an information container component; the response display unit 2201 is further specifically used for: In response to a trigger operation for the first query question, after displaying the first query question and M first responses to the first query question in the interactive dialogue interface, in response to a trigger operation for the information container class component, aggregated information of the first AI response text is displayed. The aggregated information of the first AI response text is used to present the multi-dimensional features of the first AI response text. Wherein, the first AI response text is the AI response text contained in any one of the M first response contents, and the information container class component is associated with the first AI response text.
[0169] Optionally, in the above Figure 22 Based on the corresponding embodiments, in another embodiment of the interactive processing device provided in this application, the response display unit 2201 is further specifically used for: After displaying aggregated information of the first AI response text in response to a trigger operation on an information container component, the target filter conditions are displayed in response to a configuration operation on the information filter slider in the aggregated information of the first AI response text. In response to the target filtering conditions, first content is displayed, which is information from the aggregated information of the first AI response text that meets the target filtering conditions.
[0170] Optionally, in the above Figure 22 Based on the corresponding embodiments, in another embodiment of the interactive processing device provided in this application, the response display unit 2201 is further specifically used for: After displaying aggregated information of the first AI response text in response to a trigger operation on an information container component, the current information viewing progress is displayed in response to a drag operation on the progress bar slider in the aggregated information of the first AI response text. In response to the current information viewing progress, second content is loaded, which is an information fragment in the aggregated information of the first AI reply text that matches the current information viewing progress.
[0171] Optionally, in the above Figure 22 Based on the corresponding embodiments, in another embodiment of the interactive processing device provided in this application, the information container component includes one or more of a card component and a list item component.
[0172] Optionally, in the above Figure 22 Based on the corresponding embodiments, in another embodiment of the interactive processing device provided in this application, the interactive component includes an operation triggering component; the response display unit 2201 is further specifically used for: After displaying the first query question and M first responses to the first query question in the interactive dialogue interface in response to the trigger operation for the operation triggering class component, third content is displayed on the first information display page in response to the trigger operation for the operation triggering class component. The third content is related information related to the second AI response text generated based on the business logic of the operation triggering class component. Wherein, the second AI response text is the AI response text contained in any one of the M first response contents, the operation triggering component is associated with the second AI response text, and the first information display page is the associated hierarchical page of the interactive dialogue interface.
[0173] Optionally, in the above Figure 22 Based on the corresponding embodiments, in another embodiment of the interactive processing device provided in this application, the operation triggering component includes one or more of a button component and a link component.
[0174] Optionally, in the above Figure 22 Based on the corresponding embodiments, in another embodiment of the interactive processing device provided in this application, the interactive component includes a content display component; the response display unit 2201 is further specifically used for: After displaying the first query question and M first responses to the first query question in the interactive dialogue interface in response to a trigger operation for the content display component, multimedia display information associated with the third AI response text is displayed in response to a trigger operation for the content display component. The multimedia display information associated with the third AI response text is used to present the target flow of the third AI response text. The target flow includes at least one of a business operation flow and a content generation flow. The third AI response text is the AI response text contained in any one of the M first response contents, and the content display component is associated with the third AI response text.
[0175] Optionally, in the above Figure 22 Based on the corresponding embodiments, in another embodiment of the interactive processing device provided in this application, the content display component includes an image preview component; the response display unit 2201 is specifically used for: In response to a trigger operation on the image preview component, the first image associated with the third AI response text is enlarged and displayed. The first image is used to present the target flow associated with the third AI response text.
[0176] Optionally, in the above Figure 22 Based on the corresponding embodiments, in another embodiment of the interactive processing device provided in this application, the first image includes multiple sub-images arranged in sequence; the response display unit 2201 is specifically used for: In response to a trigger operation on the image preview component, the first sub-image is enlarged and displayed. The first sub-image is the first frame sub-image in the first image and is used to describe the first stage in the target process. In response to the switching operation of the first sub-image, the second sub-image is enlarged and displayed. The second sub-image is used to describe the second stage in the target process. The first stage and the second stage are adjacent. The first sub-image and the second sub-image are two adjacent sub-images among the multiple sequentially arranged sub-images. In response to a switching operation for all said sub-images, the last sub-image is displayed, which is used to describe the end stage in the target process.
[0177] Optionally, in the above Figure 22 Based on the corresponding embodiments, in another embodiment of the interactive processing device provided in this application, the content display component includes a video playback component; the response display unit 2201 is specifically used for: In response to a trigger operation on the video playback component, a target video associated with the third AI response text is played, the target video being used to present the target flow associated with the third AI response text.
[0178] Optionally, in the above Figure 22 Based on the corresponding embodiments, in another embodiment of the interactive processing device provided in this application, the interactive processing device further includes an acquisition unit 2202 and a processing unit 2203. Specifically, the acquisition unit 2202 is used to acquire the format syntax information and component configuration information of the first text format. The first text format is used to describe the text format of the AI response text. The component configuration information includes the component type and the component display method, or the component configuration information includes the component type, the component display method and the component triggering behavior. Processing unit 2203 is specifically used for: Based on the format syntax information of the first text format and the component configuration information, construct the component protocol description information; Based on a preset large language model, the component protocol description information, candidate components, and the first query question are processed to generate target text content. The target text content is rendered to generate the M first response contents.
[0179] Optionally, in the above Figure 22 Based on the corresponding embodiments, in another embodiment of the interactive processing device provided in this application, the processing unit 2203 is specifically used for: Based on the format structure of the first text format, the target text content is parsed to obtain the M AI response texts and each of the interactive components; Each AI response text is treated as a text node, and each interactive component is treated as a component node. Based on the logical relationship between the text nodes and the component nodes, the text nodes and the component nodes are constructed to generate a target tree data structure. The text nodes and component nodes in the target tree data structure are traversed and rendered to generate the M first response contents.
[0180] Optionally, in the above Figure 22 Based on the corresponding embodiments, in another embodiment of the interactive processing device provided in this application, the processing unit 2203 is specifically used for: Based on the format structure of the first text format, the target text content is parsed to obtain the M AI response texts and each of the interactive components; Based on the logical relationship between each AI response text and the corresponding interactive component, the corresponding interactive component is inserted into a preset slot in the AI response text to generate the M first response contents.
[0181] Optionally, in the above Figure 22 Based on the corresponding embodiments, in another embodiment of the interactive processing device provided in this application, the processing unit 2203 is further specifically used for: After processing the component protocol description information, candidate components and the first query question based on the preset large language model to generate target text content, the message display format of the target application is identified. When the message display format does not match the first text format, the message display format and the target text content are converted based on the preset large language model to obtain the converted target text content. The converted target text content is rendered to obtain M third response contents and an interactive component corresponding to each third response content, wherein each third response content includes AI response text; Display the content of the M third responses and the corresponding interactive components.
[0182] The computer device in the embodiments of this application has been described above from the perspective of modular functional entities. The computer device in the embodiments of this application will now be described below from the perspective of hardware processing. Figure 23 A schematic diagram of the structure of a computer device provided in an embodiment of this application is shown. This computer device can vary considerably due to differences in configuration or performance, and may include, but is not limited to, the aforementioned terminal, server, etc.
[0183] like Figure 23 As shown, the computer device may include one or more central processing units (CPUs) 322 (e.g., one or more processors) and a memory 332, and one or more storage media 330 (e.g., one or more mass storage devices) for storing application programs 342 or data 344. The memory 332 and storage media 330 may be temporary or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), each module may include a series of instruction operations on the computer device. Furthermore, the CPU 322 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the computer device. Exemplarily, the CPU 322 is used to execute the application program 342 stored in the storage media 330, thereby implementing the interactive processing method provided in the above embodiments of this application.
[0184] The computer device may also include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0185] For example, Figure 23 The central processing unit 322 can invoke computer execution instructions stored in memory 332 to cause the computer device to perform actions such as... Figures 5 to 21 The method in one or more of the corresponding method embodiments.
[0186] The steps performed by the computer device in the above embodiments can be based on this Figure 23 The structure shown.
[0187] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0188] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0189] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0190] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0191] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0192] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.< / comp> < / comp>
Claims
1. A method for interactive processing, characterized in that, include: In response to an interactive trigger operation of an artificial intelligence (AI) interactive application, an interactive dialogue interface related to the AI interactive application is displayed, the interactive dialogue interface including input controls; In response to an input operation in the input control, a first query question is displayed; In response to a trigger operation for the first query question, the first query question and M first responses to the first query question are displayed in the interactive dialogue interface. Each first response includes an AI response text and an interactive component associated with the AI response text. The interactive component is used to guide the trigger to perform a target behavior on the corresponding AI response text. The AI response text is used to answer the first query question, and M is an integer greater than or equal to 1.
2. The method according to claim 1, characterized in that, After displaying the first query question and M first responses to the first query question in the interactive dialog interface in response to a triggering operation for the first query question, the method further includes: In response to a trigger operation on the first interactive component, a second query question is displayed in the interactive dialogue interface. The first interactive component and the second query question have a one-to-one association. The second query question is generated based on the business logic of the first interactive component. The first interactive component is any one of the M first response contents that contains the interactive component. In response to the second query question, a second answer matching the second query question is displayed in the interactive dialog interface.
3. The method according to claim 1, characterized in that, The interactive component includes an information container component; in response to a trigger operation for the first query question, after displaying the first query question and M first responses to the first query question in the interactive dialogue interface, the method further includes: In response to a trigger operation on the information container component, aggregated information of the first AI response text is displayed, which is used to present the multi-dimensional features of the first AI response text. Wherein, the first AI response text is the AI response text contained in any one of the M first response contents, and the information container class component is associated with the first AI response text.
4. The method according to claim 3, characterized in that, After displaying aggregated information of the first AI response text in response to a triggered operation on an information container component, the method further includes: In response to the configuration operation of the information filtering slider in the aggregated information of the first AI response text, the target filtering conditions are displayed; In response to the target filtering conditions, first content is displayed, which is information from the aggregated information of the first AI response text that meets the target filtering conditions.
5. The method according to claim 3, characterized in that, After displaying aggregated information of the first AI response text in response to a triggered operation on an information container component, the method further includes: In response to a drag operation of the progress bar slider in the aggregated information of the first AI response text, the current information viewing progress is displayed; In response to the current information viewing progress, second content is loaded, which is an information fragment in the aggregated information of the first AI reply text that matches the current information viewing progress.
6. The method according to claim 1, characterized in that, The interactive component includes an operation triggering component; in response to a triggering operation for the first query question, after displaying the first query question and M first responses to the first query question in the interactive dialogue interface, the method further includes: In response to a trigger operation on the operation triggering component, third content is displayed on the first information display page. The third content is related information about the second AI response text generated based on the business logic of the operation triggering component. Wherein, the second AI response text is the AI response text contained in any one of the M first response contents, the operation triggering component is associated with the second AI response text, and the first information display page is the associated hierarchical page of the interactive dialogue interface.
7. The method according to claim 1, characterized in that, The interactive component includes a content display component; in response to a trigger operation for the first query question, after displaying the first query question and M first responses to the first query question in the interactive dialogue interface, the method further includes: In response to a trigger operation on the content display component, multimedia display information associated with the third AI response text is displayed. The multimedia display information associated with the third AI response text is used to present the target flow of the third AI response text. The target flow includes at least one of a business operation flow and a content generation flow. The third AI response text is the AI response text contained in any one of the M first response contents, and the content display component is associated with the third AI response text.
8. The method according to claim 7, characterized in that, The content display components include an image preview component; In response to a triggered action on a content display component, display multimedia information associated with the third AI's response text, including: In response to a trigger operation on the image preview component, the first image associated with the third AI response text is enlarged and displayed. The first image is used to present the target flow associated with the third AI response text.
9. The method according to claim 8, characterized in that, The first image comprises multiple sub-images arranged in sequence; in response to a trigger operation on the image preview component, the first image associated with the third AI response text is enlarged and displayed, including: In response to a trigger operation on the image preview component, the first sub-image is enlarged and displayed. The first sub-image is the first frame sub-image in the first image and is used to describe the first stage in the target process. In response to the switching operation of the first sub-image, the second sub-image is enlarged and displayed. The second sub-image is used to describe the second stage in the target process. The first stage and the second stage are adjacent. The first sub-image and the second sub-image are two adjacent sub-images among the multiple sequentially arranged sub-images. In response to a switching operation for all said sub-images, the last sub-image is displayed, which is used to describe the end stage in the target process.
10. The method according to claim 7, characterized in that, The content display components include video playback components; In response to a triggered action on a content display component, display multimedia information associated with the third AI's response text, including: In response to a trigger operation on the video playback component, a target video associated with the third AI response text is played, the target video being used to present the target flow associated with the third AI response text.
11. The method according to any one of claims 1 to 10, characterized in that, The method further includes: Obtain format syntax information and component configuration information of a first text format, wherein the first text format is used to describe the text format of the AI response text, and the component configuration information includes component type and component display method, or the component configuration information includes component type, component display method and component triggering behavior; Based on the format syntax information of the first text format and the component configuration information, construct the component protocol description information; Based on a preset large language model, the component protocol description information, candidate components, and the first query question are processed to generate target text content. The target text content is rendered to generate the M first response contents.
12. The method according to claim 11, characterized in that, The target text content is rendered to generate the M first response contents, including: Based on the format structure of the first text format, the target text content is parsed to obtain the M AI response texts and each of the interactive components; Each AI response text is treated as a text node, and each interactive component is treated as a component node. Based on the logical relationship between the text nodes and the component nodes, the text nodes and the component nodes are constructed to generate a target tree data structure. The text nodes and component nodes in the target tree data structure are traversed and rendered to generate the M first response contents.
13. The method according to claim 11, characterized in that, The target text content is rendered to generate the M first response contents, including: Based on the format structure of the first text format, the target text content is parsed to obtain the M AI response texts and each of the interactive components; Based on the logical relationship between each AI response text and the corresponding interactive component, the corresponding interactive component is inserted into a preset slot in the AI response text to generate the M first response contents.
14. The method according to claim 11, characterized in that, After processing the component protocol description information, candidate components, and the first query question based on a preset large language model to generate the target text content, the method further includes: Identify the message display format of the target application; When the message display format does not match the first text format, the message display format and the target text content are converted based on the preset large language model to obtain the converted target text content. The converted target text content is rendered to obtain M third response contents and an interactive component corresponding to each third response content, wherein each third response content includes AI response text; Display the content of the M third responses and the corresponding interactive components.
15. An interactive processing device, characterized in that, include: A response display unit is used to respond to an interactive trigger operation of an artificial intelligence (AI) interactive application and display an interactive dialogue interface related to the AI interactive application, the interactive dialogue interface including input controls; The response display unit is used to display a first query question in response to an input operation in the input control; The response display unit is configured to respond to a trigger operation for the first query question by displaying the first query question and M first responses to the first query question in the interactive dialogue interface. Each first response includes an AI response text and an interactive component associated with the AI response text. The interactive component is configured to guide the triggering of a target behavior on the corresponding AI response text. The AI response text is used to answer the first query question, and M is an integer greater than or equal to 1.
16. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 14.
17. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 14.
18. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1 to 14.