Conversation processing method and device, equipment, storage medium and program product
By using non-textual visual content and generative models in the dialogue interface to help users clarify their needs, the problem of unclear user information acquisition needs is solved, achieving efficient and accurate clarification of information needs and adapting to the characteristics of different object categories.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING YOUZHUJU NETWORK TECH CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-01
AI Technical Summary
When users' information acquisition needs are broad and vague, existing technologies struggle to efficiently and accurately clarify these needs, resulting in low information acquisition efficiency. Furthermore, existing solutions cannot dynamically adapt to the characteristics of different object categories.
By supplementing user input with non-text interactive visual content in the dialogue interface, using generative models for semantic parsing and category recognition, providing multiple visual cues to help users clarify their needs, and combining visual and text interaction to reduce the interaction cost of clarifying needs.
It improves the clarity and efficiency of expressing information needs, reduces ambiguity in understanding, enhances the accuracy and clarification efficiency of information needs, and adapts to the characteristics of different object categories.
Smart Images

Figure CN121958503A_ABST
Abstract
Description
Dialogue processing methods, apparatus, devices, storage media, and program products Technical Field
[0001] This application relates to the field of human-computer interaction technology, and in particular to a dialogue processing method, apparatus, device, storage medium, and program product. Background Technology
[0002] With the development of internet technology, people frequently obtain the information they want online, such as through web searches. When a user's description of their information retrieval needs is broad, the feedback results obtained from the vast and complex databases will also be broad and general, potentially failing to meet the user's information needs. In this case, the user needs to continuously refine their information search description, a time-consuming and labor-intensive process that reduces the efficiency of clarifying their information retrieval needs and overall information acquisition efficiency. Summary of the Invention
[0003] To address the aforementioned technical problems, this application provides a dialogue processing method, apparatus, device, storage medium, and program product.
[0004] In a first aspect, this application provides a dialogue processing method, the method comprising: displaying a dialogue interface and displaying first input content in the dialogue interface; displaying first prompt content in the dialogue interface; the first prompt content comprising at least two non-textual first visual contents; each of the first visual contents having an interactive attribute; the first visual contents being used to supplement the first input content; and displaying second input content in the dialogue interface in response to a triggering operation on the first visual contents; the second input content being obtained based on the selected first visual contents.
[0005] Secondly, this application also provides a dialogue processing apparatus, comprising: a first input content display module for displaying a dialogue interface and displaying first input content in the dialogue interface; a first prompt content display module for displaying first prompt content in the dialogue interface; the first prompt content includes at least two non-textual first visual contents; each of the first visual contents has an interactive attribute; the first visual contents are used to supplement the first input content; and a second input content display module for displaying second input content in the dialogue interface in response to a triggering operation on the first visual contents; the second input content is obtained based on the selected first visual contents.
[0006] Thirdly, this application also provides an electronic device, which includes: a processor; and a memory for storing executable instructions; wherein the processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the dialogue processing method described in any embodiment of this application.
[0007] Fourthly, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the dialogue processing method described in any embodiment of this application.
[0008] Fifthly, this application also provides a computer program product for executing the dialogue processing method described in any embodiment of this application.
[0009] The dialogue processing method, apparatus, device, storage medium, and program product of this application are capable of displaying a dialogue interface and displaying first input content in the dialogue interface; displaying first prompt content in the dialogue interface, the first prompt content including at least two non-textual first visual contents, the first visual contents having interactive attributes and used to supplement the first input content; and displaying second input content in response to a triggering operation on the first visual contents; the second input content is obtained based on the selected first visual content. This achieves the goal of providing multiple non-textual visual contents as refinement or supplementation to the first input content when the user's submitted information acquisition needs—the first input content—are relatively broad and vague. This largely avoids misunderstandings caused by the inability of textual supplementary content to accurately express the user's true information needs, thereby reducing ambiguity in the supplementary information, enhancing the clarity and efficiency of the first prompt content in expressing information needs, and thus more accurately and efficiently assisting users in clarifying their true information needs, improving the accuracy of demand analysis for categories with strong style dependencies.
[0010] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. Attached Figure Description
[0011] The above and other features, advantages, and aspects of the embodiments of this application will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0012] Figure 1 is a schematic diagram of the scenario architecture applicable to a dialogue processing method provided in this application; Figure 2 is a flowchart of a dialogue processing method provided in this application; Figure 3 is a schematic diagram of a dialogue processing process for clarifying needs in a shopping scenario provided in this application; Figure 4 is a schematic diagram of a dialogue processing process for clarifying needs in another shopping scenario provided in this application; Figure 5 is a structural schematic diagram of a dialogue processing device provided in this application; Figure 6 is a structural schematic diagram of an electronic device provided in this application. Detailed Implementation
[0013] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this application. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.
[0014] It should be understood that the steps described in the method embodiments of this application may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.
[0015] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0016] It should be noted that the concepts of "first" and "second" mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0017] It should be noted that the terms "a" and "a plurality of" used in this application are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0018] The names of the messages or information exchanged between multiple devices in the embodiments of this application are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0019] The dialogue processing method provided in this application is mainly applicable to scenarios of information retrieval or content generation based on generative models. For example, it can be applicable to object retrieval scenarios (such as shopping scenarios) based on generative models.
[0020] In one scenario, Figure 1 is a schematic diagram of the scenario architecture of a dialogue processing method provided by this application. As shown in Figure 1, the scenario architecture provided by this application embodiment includes: a first device corresponding to the server and a second device corresponding to the client.
[0021] The server is configured to communicate with the client and execute core processing flows that clarify user queries in a dialog-like manner, such as semantic parsing, information completeness assessment, and ambiguity evaluation. These core processing flows can be implemented based on generative models or other algorithms. The first device is used to host and implement the server functions, and it can have various implementation forms, such as a server 110 or a server cluster.
[0022] The client is configured to provide a user interface (such as a dialog box) to receive user operations and display information. The second device carries and implements the client's functions and can take various forms, such as hardware or software. When the second device is hardware, it can be various electronic devices with human-computer interaction and display functions, including but not limited to smartphones 121, personal digital assistants (PDAs), tablet computers 122, laptops 123, or desktop computers 124. When the second device 120 is software, it can be installed in the aforementioned electronic devices and can be implemented as multiple software programs or software modules, or as a single software program or software module. No specific limitations are made here.
[0023] In another scenario, the dialogue processing method of this application can be executed by a first device. In this way, the first device can comprehensively perform the functions of both the aforementioned server and client.
[0024] In another scenario, the dialogue processing method of this application can be executed by a second device. In this way, the second device can comprehensively perform the functions of both the aforementioned server and client.
[0025] Related technologies may provide an information search function based on a generative model of a certain scale. Users can input their information needs through a human-computer interactive dialogue interface, and the information search function can then provide feedback to the user in a dialogue format. Because users are unsure what kind of input accurately expresses their information needs, their initial input is likely to be brief and unclear, lacking some key elements necessary for information searching. For example, in a shopping scenario, the user's initial input may lack elements such as budget, style, design, and usage scenario for the desired object (e.g., physical goods, virtual goods, or services). Such input makes it difficult for the information search function to accurately understand the user's information search needs, resulting in difficulty in accurately matching the desired object.
[0026] Information search functions can guide users to continuously input new content in various ways to help them clarify their information needs. One approach is through continuous text-based follow-up questions, which clarifies information needs by guiding users to repeatedly answer questions. However, this method requires users to repeatedly input text, resulting in high operational costs. Furthermore, it heavily relies on user input willingness, lacking a mechanism to judge the completeness of key information elements and easily deviating from the original information request during multiple rounds of dialogue. Another approach is to pre-configure structured information clarification templates, providing the information elements to be clarified in the form of forms or filters, requiring users to fill in multiple elements at once. This method easily disrupts the user flow and is difficult to integrate smoothly with conversational interactions. Moreover, when these methods are applied to shopping scenarios, they apply the same information clarification process to all product categories, easily ignoring the differences in product category characteristics, leading to relatively low efficiency in clarification during shopping. In short, none of the relevant clarification solutions can clarify users' true information needs more efficiently with lower user clarification costs (such as interaction costs), especially failing to dynamically adapt to the different product category characteristics in shopping scenarios.
[0027] Based on the above, this application provides a technical solution for dialogue processing. When it is determined that the user's key information needs are missing and that the missing information needs are difficult to clarify with text content, interactive first visual content is used as dialogue content to provide the user with information need clarification function. This maintains the continuous dialogue process while reducing the ambiguity of the need elements and the interaction cost of need clarification through visual content, thereby improving the efficiency and accuracy of clarifying the user's information needs.
[0028] The dialogue processing method provided in this application can be executed by a dialogue processing device, which can be implemented by software and / or hardware, and can be integrated into an electronic device with interactive and display functions. This electronic device can be the first device or the second device described in the foregoing embodiments.
[0029] Figure 2 shows a flowchart of a dialogue processing method provided in this application. As shown in Figure 2, the dialogue processing method may include the following steps: S210, displaying a dialogue interface and displaying the first input content in the dialogue interface.
[0030] The first input content is the information requirement initially entered into the dialog interface to trigger the execution of the information search function.
[0031] Specifically, the aforementioned information search function can be presented through a human-computer dialogue interface. Therefore, if a user launches a client (such as an application, mini-program, or webpage) and this function is accessed by default, or if the user triggers this function through other pages (such as pages introducing or promoting the function), the electronic device integrating this function can detect the trigger operation and then display the corresponding dialogue interface. Then, the electronic device can respond to the user's input or an upstream business request (i.e., an external request) to obtain the first input content and display it in the dialogue interface.
[0032] In one scenario, S210 includes: responding to an input operation applied to the dialog interface, displaying first input content in the dialog interface; the first input content corresponds to the input operation. This scenario corresponds to a user's need to input information themselves through an input control (such as an input box, voice control, etc.). After detecting the user's input operation in the dialog interface, the electronic device can obtain the result of the input operation, i.e., the first input content. Then, the electronic device displays the first input content in the dialog interface.
[0033] In another scenario, S210 includes: responding to an external request, displaying a dialog interface and showing first input content in the dialog interface; the first input content corresponds to content obtained by parsing the external request. This scenario corresponds to a user activating the information search function through an interactive operation other than the information search function itself. For example, a user triggers the activation of the information search function by performing an interactive operation on the page promoting the aforementioned function. The electronic device can detect the aforementioned interactive operation and obtain the corresponding external request. Then, the electronic device can parse the external request to obtain the text content corresponding to the aforementioned interactive operation as the first input content. Subsequently, the electronic device displays the first input content in the dialog interface.
[0034] S220. Display first prompt content in the dialog interface; the first prompt content includes at least two non-textual first visual contents; each first visual content has interactive attributes; the first visual contents are used to supplement the first input content.
[0035] The first prompt content is generated by the information search function to prompt the user to clarify their needs regarding the first input content. It includes at least multiple first visual content items corresponding to the same prompt dimension. The prompt dimension can be understood as a clarification dimension, referring to the information dimension necessary to supplement the response data corresponding to the input content. It may correspond to a part of the key requirement elements described above (i.e., missing requirement elements). The first visual content refers to the content presented in a non-textual visual form in relation to the first input content.
[0036] Specifically, electronic devices can invoke their corresponding generative models to obtain initial prompts for the first input content. First, the model performs semantic parsing on the first input content, identifying the user's information search request, detecting whether key demand elements in the corresponding scenario are missing, and assessing whether the first input content is ambiguous or conflicting. Then, the model compares the parsing results with convergence conditions for demand clarification to determine if the first input content triggers demand clarification. These convergence conditions can include thresholds for the confidence level of the demand classification, the completeness / completeness rate of demand elements, the confidence level of ambiguity / conflict in the demand, and the upper limit of clarification rounds, which can be obtained by the model learning from subdivided information search scenarios. If the parsing result triggers any convergence condition, demand clarification is triggered. Next, the model can obtain one or more prompt dimensions for this clarification based on the priority of each prompt dimension (e.g., required, high impact, optional) and / or the severity of demand ambiguity (e.g., high ambiguity or low ambiguity), as well as a pre-determined clarification granularity. If multiple prompt dimensions are obtained, they can be ordered according to the aforementioned priority or severity. For example, in the first round of clarification, mandatory and highly ambiguous cue dimensions can be identified. In subsequent clarification rounds, the priority and / or severity of ambiguity can be progressively reduced to achieve multi-level, orderly control of requirement clarification, further shortening the requirement clarification process and thus improving the efficiency of requirement clarification. The above clarification granularity can be understood as the number of cue dimensions covered in one round of clarification, which can be pre-configured or determined based on user preferences. Then, the model determines whether the object contained in the first input content has a strong visual feature dependency on a certain cue dimension. If so, it indicates that using text for requirement clarification is prone to ambiguity. In this case, the model can obtain multiple first visual contents corresponding to the object for each cue dimension obtained above, which can be used as candidate supplementary content for the first input content. Finally, the model can generate the first cue content using the cue dimension and its corresponding first visual contents. Alternatively, when the clarification granularity includes one cue dimension, the model can use each first visual content as the first cue content.
[0037] After receiving the initial prompt, the electronic device can display it in the dialog interface, within the display area following the initial input. To reduce interaction costs, the electronic device can present the initial visual content within the initial prompt in an interactive format. For example, interactive cards can be used to carry the initial visual content, giving it interactive attributes.
[0038] Referring to Figure 3(a), the first input content 310 is "Recommend a backpack for me, within xxx yuan," which clearly states that the target of the information search is "backpack" and the budget limit is "xxx yuan." Following the above process, the generative model can determine that the first input content will trigger demand clarification, and in the first round of clarification, it determines the following prompt dimensions: commuting method 320, items to pack 320, function 320, and style 320. The style dimension strongly depends on the visual characteristics of the backpack. The electronic device displays four first visual contents in the dialogue interface, represented by cards: image 321 for retro-artistic style, image 322 for outdoor functional style, image 323 for casual minimalist style, and image 324 for elegant commuting style.
[0039] In one scenario, the first visual content includes at least one of still images, moving images, and videos.
[0040] In this context, a still image can be understood as a single image. A moving image can be understood as visual data (such as an image sequence or animation) composed of multiple still images arranged in a time sequence, which can be played continuously to present a motion effect. Video can be understood as encoded and compressed streaming media or file data with a container format, which may contain an audio track; that is, an audio-visual file with encoding and containerization.
[0041] In one example, the first visual content can be implemented as a static image as shown in Figure 3(a), which can improve the universality and device compatibility of dialogue processing and reduce the device's bandwidth consumption to some extent.
[0042] In another example, the objects included in the first input content, in some cue dimensions, besides having a strong reliance on visual features, also require some changing or dynamic visual features to more clearly express the content of that cue dimension. For example, for cue dimensions such as the usage scenario or usage method of an object, dynamic visual features can more clearly demonstrate its specific usage scenario or usage method. Therefore, the first visual content can also be implemented as a dynamic image or video to further reduce the user's understanding cost of the first visual content.
[0043] For example, if the first input includes headphones, and if the defined prompt dimension is a usage scenario, the electronic device can display videos corresponding to subway usage scenarios, representing their noise reduction characteristics; videos corresponding to e-sports gaming scenarios, representing their low latency characteristics; and videos corresponding to sports scenarios, representing their sweat-proof and anti-fall-off characteristics. When the user triggers the video corresponding to the e-sports gaming scenario, the electronic device can obtain the text "gaming scenario, low latency," which serves as the basis for the subsequent second input.
[0044] In one scenario, the first visual content is obtained through the following steps A1 to A3.
[0045] Step A1: Perform semantic parsing and category identification based on the first input content to obtain the first category to which the object belongs and the corresponding prompt dimension.
[0046] The first input content contains the target information for the search. The first category refers to the category to which the target information in the first input content belongs.
[0047] Specifically, electronic devices can use generative models to perform semantic parsing and category recognition on the first input content, in order to parse out the objects contained therein, the prompt dimensions corresponding to the objects, and the first category to which the objects belong.
[0048] Step A2: Match the first category with each of the second categories to obtain the matching degree.
[0049] The second category is distinguished by visual characteristics. These visual characteristics can include at least one of the following: appearance, shape, size, style, usage, and structure. Appearance can be understood as the overall visible features of the object. Style can be understood as the object's external form, shape, format, or specific design. Style can refer to the object's overall design tone or aesthetic attributes, such as minimalism, retro style, sporty style, business style, or anime style. Structure can be understood as the object's component composition, the relative positions of these components, their connections, and their spatial arrangement and layout.
[0050] Specifically, for a given application scenario, generative models can be used to pre-identify the various second categories within that scenario. For example, in a shopping scenario, second categories such as clothing, bags, and home furnishings can be pre-identified.
[0051] After obtaining the first category, electronic devices can be matched with each of the second categories to obtain a matching degree. This matching degree is used to determine whether the first category belongs to a category that distinguishes objects based on visual features.
[0052] Step A3: If the detected matching degree is greater than the matching degree threshold, then obtain the first visual content based on the first category and the prompt dimension.
[0053] Specifically, the electronic device can compare each of the above matching scores with a pre-set matching score threshold. If all matching scores are less than the matching score threshold, then the objects included in the first input content do not require visual content clarification. If at least one matching score is greater than or equal to the matching score threshold, it indicates that the first category belongs to a category that distinguishes objects based on visual features, and the aforementioned objects can be clarified with the help of visual content. In this case, the electronic device can obtain each first visual content based on the first category and the obtained prompt dimensions.
[0054] In one scenario, step A3 above, "obtaining each first visual content based on the first category and cue dimension," includes: searching the resource library based on the first category and cue dimension to obtain the first visual content. In this case, the electronic device can use the first category and cue dimension as search indexes to search the resource library corresponding to the aforementioned application scenario (such as the product library in a shopping scenario), and the retrieved visual content can be used as the first visual content. This allows visual content to be obtained from an existing database, improving the efficiency of subsequent object recommendations.
[0055] In another scenario, step A3 above, "obtaining each first visual content based on the first category and cue dimension," includes: using a generative model to obtain the first visual content based on the first category and cue dimension. In this case, the electronic device can use the first category and cue dimension as input data, call the relevant generative model, and output the first visual content. This allows for the real-time generation of more suitable visual content, further improving the clarity of its expression of information needs.
[0056] In another scenario, the electronic device can first perform the steps described above to search the resource library. If the search result does not find matching visual content, then the step of generating visual content can be performed to obtain the first visual content. This approach can balance the content of the existing resource library with the compliance of the visual content.
[0057] Through steps A1 to A3 above, the category to which the object belongs can be used as the core decision factor for visual clarification, simplifying the acquisition process of the first visual content and thus further improving the efficiency and accuracy of visual clarification. Furthermore, in step S230, in response to a trigger operation on the first visual content, second input content is displayed in the dialog interface; the second input content is obtained based on the first visual content that has been selected.
[0058] The second input content refers to the input content automatically generated based on the user's interaction with the first visual content, which is used to supplement the first input content.
[0059] Specifically, the user performs a trigger action (such as clicking or selecting) on a displayed first visual content. The electronic device can detect this trigger action and obtain the prompt dimension corresponding to the selected first visual content. Then, based on the obtained prompt dimension, it obtains the second input content and displays it in the dialog interface. In this way, the user does not need to perform input or dialog operations; a request clarification can be completed through a simple interactive operation, greatly reducing the interaction cost of request clarification.
[0060] Referring again to Figure 3(a), if the user selects the image 322 corresponding to the outdoor functional style in the style dimension, the electronic device can use it as the first visual content to trigger the selection and obtain the text "outdoor functional style" as part of the second input content 330, which is displayed on the dialog interface as the user's re-input content.
[0061] In one scenario, displaying a second input in the dialog interface includes using the selected first visual content as the second input and displaying it on the dialog interface. In this case, the electronic device can directly display the selected first visual content as the second input in the dialog interface. For example, in the example of Figure 3(a), the electronic device can display an image of an "outdoor functional style" backpack as new input in the dialog interface (not shown in Figure 3(a)). This allows users to more clearly understand their supplementary information needs and improves user review efficiency.
[0062] In another scenario, displaying the second input content in the dialog interface includes: obtaining the second text content corresponding to the first visual content that triggered the selection, and displaying the second text content as the second input content in the dialog interface. In this scenario, the electronic device can obtain the second text content based on the specific prompt text of the prompt dimension associated with the first visual content that triggered the selection, and display it as the second input content in the dialog interface. The display effect of this scenario is shown in Figure 3(a) as an "outdoor functional style". This maintains the structure and consistency of the second input content and improves the simplicity of the dialog interface.
[0063] In another scenario, displaying a second input in the dialog interface includes: merging the first input and the first visual content selected by triggering the selection to obtain the second input, and then displaying the second input in the dialog interface. In this scenario, the electronic device can use both the first input and the first visual content selected by triggering the selection as the second input, giving it multimodal characteristics while maintaining the integrity of the second input, thus improving the user's efficiency in reviewing the input and the coherence of the dialogue.
[0064] The above merging process can be either by directly concatenating the first visual content that triggered the selection after the first input content to obtain the second input content, or by using a generative model to re-integrate the first visual content that triggered the selection and the first input content to obtain a second input content that is more fluent and in line with language logic.
[0065] In another scenario, displaying the second input content in the dialog interface includes: obtaining the second text content corresponding to the first visual content that triggered the selection, merging the first input content and the second text content to obtain the second input content, and displaying the second input content in the dialog interface. In this scenario, the electronic device can first obtain the second text content corresponding to the first visual content that triggered the selection, such as "outdoor functional style". Then, it merges it with the first input content to obtain the second input content. For example, in Figure 3(a), after each time the user makes a clarification selection, the electronic device can obtain the complete, text-based second input content 340. This maintains the integrity and structural consistency of the second input content, improving the user's efficiency in reviewing the input content and the coherence of the dialogue.
[0066] The dialogue processing method provided by the above embodiments can display a dialogue interface and display first input content in the dialogue interface; display first prompt content in the dialogue interface, the first prompt content including at least two non-textual first visual contents, the first visual contents being used to supplement the first input content; in response to a trigger operation on the first visual contents, display second input content in the dialogue interface; the second input content is obtained based on the selected first visual content; it realizes that when the information acquisition needs submitted by the user—the first input content—are relatively broad and vague, multiple non-textual visual contents are provided as a refinement or supplement to the first input content, which largely avoids the problem of misunderstanding caused by the inability of textual supplementary content to accurately express the user's true information needs, thereby reducing the ambiguity of the supplementary information, enhancing the clarity and efficiency of the first prompt content in expressing information needs, and thus more accurately and efficiently assisting users in clarifying their true information needs, improving the accuracy of demand analysis for categories with strong style dependence.
[0067] In one scenario, after S210, the method further includes steps B1 to B2.
[0068] Step B1: Display the second prompt content in the dialog interface; the second prompt content includes at least two first text contents; each first text content is displayed in a preset style; each first text content has interactive attributes.
[0069] The second prompt is generated by the information search function to prompt the user to clarify their needs regarding the first input content. It includes at least multiple first text contents corresponding to the same prompt dimension. The first text content refers to the text-based content presented in relation to the first input content, used to supplement it. The preset style refers to a pre-set display style, which differs from the default text display style.
[0070] Specifically, referring to the explanation of S220 above, after obtaining the prompt dimensions and clarification granularity corresponding to the objects contained in the first input content, if the model determines that the object does not have visual feature dependence on a certain prompt dimension, then it can achieve unambiguous clarification of the requirement through text. At this time, the model can obtain multiple first text contents corresponding to the object for that prompt dimension, which can be used as candidate supplementary content for the first input content. Then, the model can use the prompt dimension and its corresponding first text contents to generate second prompt content.
[0071] After receiving the second prompt, the electronic device can display it in the dialog interface, in the display area following the first input. To reduce interaction costs, the electronic device can present each piece of the first text in the second prompt in an interactive format. For example, interactive text can be used to display each piece of the first text, giving it interactive attributes.
[0072] Referring again to Figure 3(a), the first input 310 clearly states that the information search request is for a "backpack" and the budget limit is "xxx yuan". Following the above process, the generative model can determine that the commuting method prompt dimension 320, the items to be packed prompt dimension 320, and the function prompt dimension 320 all require no visual clarification. In the dialog interface, the electronic device displays multiple first text contents corresponding to each prompt dimension in a preset style with bold text and a "+" icon.
[0073] Step B2: In response to the triggering operation on the first text content, display the third input content in the dialog interface; the third input content is obtained based on the first text content that was triggered and selected.
[0074] The third input content refers to the input content automatically generated based on the user's interactive operation on the first text content, which is used to supplement the first input content.
[0075] Specifically, the user performs a trigger action (such as clicking or selecting) on a displayed first text content. The electronic device can detect this trigger action and obtain the prompt dimension corresponding to the first text content that was selected. Then, based on this first text content, a third input content is obtained and displayed in the dialog interface. In this way, no user input or dialog operation is required; a request clarification can be completed through a simple interactive operation, greatly reducing the interaction cost of request clarification.
[0076] Referring again to Figure 3(a), if the user selects “subway / bus” corresponding to the commuting mode prompt dimension 320, “notebook” corresponding to the items to carry prompt dimension 320, and “lightweight and does not strain the shoulders” corresponding to the function prompt dimension 320, then these texts are the first text content that triggers the selection. The electronic device can display them as part of the third input content (due to the coarser granularity, the third input content in Figure 3(a) is still exemplified as the first input content 330) on the dialog interface as the user's re-input content.
[0077] It should be noted that the process of generating the third input content can refer to the relevant part of S230 above, except that the first visual content and the second input content are replaced with the first text content and the third input content, respectively. In this way, the third input content can be the selected first text content alone, or it can be the combined result of the first input content and the selected first text content.
[0078] Through steps B1 to B2 above, requirements can be clarified in an interactive text format, further reducing the interaction cost of requirements clarification without interrupting the dialogue flow, and achieving a better integration effect of conversational interaction.
[0079] In one scenario, electronic devices can clarify requirements not only through visual or textual content alone, but also by combining both. As shown in Figure 3(a), when the initial clarification granularity is relatively coarse, the electronic device can simultaneously clarify requirements across four prompt dimensions, with three prompt dimensions presented as textual content and one as visual content. This allows for adaptive switching of clarification modalities, further enhancing the flexibility, clarity, and efficiency of requirements clarification.
[0080] In one scenario, step B2 includes: in response to a triggering operation on the first text content, displaying the first text content selected by the triggering operation in an input box within the dialog interface; in response to a modification operation on the first text content, obtaining the modified third text content, and displaying the third text content as the third input content in the dialog interface.
[0081] Specifically, after detecting a user's trigger action on a certain first text content, the electronic device can display it in the input box of the dialog interface, making it editable. If the user is not satisfied with the first text content, they can edit the input box, modifying the first text content. The electronic device can then obtain the modified first text content as the third text content. Then, in response to the user's send / submit interaction with this third text content, the electronic device can display this third text content as the third input content after the most recently displayed prompt, serving as new input.
[0082] Taking the text clarification in Figure 3(a) as an example, and referring to Figure 4(a), the electronic device displays the first text content selected by the user, such as "subway / bus," "notebook," and "lightweight and comfortable to carry," in the input box. The user can trigger the input box and continue typing "book" after "notebook," adding content to the prompt dimension 320 for items to be carried, thus obtaining the third text content—"subway / bus, notebook, book, lightweight and comfortable to carry." Then, after the user performs the send operation, as shown in Figure 4(b), the electronic device can display this content as the third input content in the dialog interface. This further improves the interactive flexibility and accuracy of the clarification content.
[0083] In one scenario, after S230, the method further includes the following steps C1 to C2.
[0084] Step C1: Display the third prompt content in the dialog interface; the third prompt content includes at least two partial contents; the partial contents have interactive attributes; the partial contents are used to supplement the second input content; the third prompt content is used to refine or supplement the prompt dimension corresponding to the first prompt content.
[0085] The third prompt is generated by the information search function to prompt users to clarify their needs regarding the second input content. It includes at least several partial contents corresponding to the same prompt dimension. These partial contents include non-textual second visual content or fourth text content. Second visual content refers to content presented in a non-textual visual form related to the second input content. Fourth text content refers to content presented in text form related to the second input content, used to supplement it.
[0086] Specifically, as described in the foregoing embodiments, after completing the first round of requirement clarification, the electronic device compares all input content (such as the first input content, the second input content / third input content) in the current information search process with the convergence conditions for requirement clarification described above to determine whether these input contents need to continue triggering requirement clarification. If these input contents trigger any convergence condition, then a new round of requirement clarification is triggered.
[0087] Referring to the explanation of S220 above, the model can obtain one or more cue dimensions for a new round of clarification based on the parsing results of these input contents. Then, after obtaining the new cue dimensions and clarification granularity, it can also determine whether each cue dimension is suitable for textual or visual cue content, and generate corresponding non-textual second visual content or fourth text content as candidate supplementary content to the second input content in this round. Next, the model can use the cue dimensions and their corresponding local content to generate third cue content. The electronic device can then display this third cue content in the dialog interface, within the display area following the second input content. Similarly, to reduce interaction costs, the electronic device can present the aforementioned local content in the interaction forms adapted to the aforementioned modalities.
[0088] Because the prompts in the previous rounds of clarification were ordered according to priority and / or the severity of ambiguity, the prompts in the new round of requirement clarification are a refinement or supplement to the prompts in the previous rounds. Accordingly, when the third prompt is used in subsequent rounds of requirement clarification, it is used to refine or supplement the prompts corresponding to the first prompt. For example, in the first round of clarification, mandatory and highly ambiguous prompts can be identified. In subsequent clarification rounds, the priority and / or severity of ambiguity can be gradually reduced, achieving multi-level, orderly requirement clarification control, avoiding over-clarification or under-clarification, further effectively and reasonably shortening the requirement clarification process, and thus further improving the efficiency of requirement clarification.
[0089] Step C2: In response to the triggering operation on the local content, display the fourth input content in the dialog interface; the fourth input content is obtained based on the selected local content.
[0090] The fourth input content refers to the input content automatically generated based on the user's interaction with local content, which is used to supplement the second input content.
[0091] Specifically, the user performs a trigger action (such as clicking or selecting) on a displayed section of content. The electronic device can detect this trigger action and obtain the prompt dimension corresponding to the selected section of content. Then, based on the obtained prompt dimension, a fourth input content is obtained and displayed in the dialog interface. In this way, no user input or dialog operation is required; a request clarification can be completed through a simple interactive operation, greatly reducing the interaction cost of request clarification.
[0092] Referring to Figure 3(b), after displaying the second input content 350, the electronic device determines that a further round of requirement clarification is needed, and its prompt dimension is the supplementary color dimension. The color dimension strongly depends on the visual characteristics of the backpack. The electronic device displays four second visual contents carried by cards in the dialog interface: an image corresponding to a stable dark color, an image corresponding to a bright and vivid color, an image corresponding to an elegant light color, and an image corresponding to a lively color block. If the user selects the image corresponding to a bright and vivid color, the electronic device can use it as the trigger for the selected second visual content and obtain the text "bright and vivid," which is displayed in the dialog interface as the user's re-input content.
[0093] The electronic device can repeat the aforementioned steps until all the input content fails to trigger any convergence condition, or until it detects that the user has interrupted the interaction or instruction to clarify the request. At this point, the clarification process of the current information search request can be terminated, and the subsequent object search can continue.
[0094] It should be noted that the process of generating the fourth input content can refer to the relevant part of S230 above, except that the first visual content, the first input content, and the second input content are replaced with local content, the aforementioned acquisition of each input content (such as the first input content, the second input content / third input content), and the fourth input content, respectively. In this way, the fourth input content can be a single selected local content, or it can be the result of merging the aforementioned input content with the selected local content.
[0095] The following are embodiments of the dialogue processing apparatus provided in this application. This apparatus and the dialogue processing methods of the above embodiments belong to the same inventive concept. For details not described in detail in the embodiments of the dialogue processing apparatus, please refer to the embodiments of the above dialogue processing methods.
[0096] Figure 5 shows a schematic diagram of a dialogue processing device. As shown in Figure 5, the dialogue processing device 500 may include: a first input content display module 510 for displaying a dialogue interface and displaying first input content in the dialogue interface; a first prompt content display module 520 for displaying first prompt content in the dialogue interface; the first prompt content includes at least two non-textual first visual contents; each first visual content has interactive attributes; the first visual content is used to supplement the first input content; and a second input content display module 530 for displaying second input content in the dialogue interface in response to a trigger operation on the first visual content; the second input content is obtained based on the selected first visual content.
[0097] The dialogue processing method provided by the above embodiments can display a dialogue interface and display first input content in the dialogue interface; display first prompt content in the dialogue interface, the first prompt content including at least two non-textual first visual contents, the first visual contents being used to supplement the first input content; in response to a trigger operation on the first visual contents, display second input content in the dialogue interface; the second input content is obtained based on the selected first visual content; it realizes that when the information acquisition needs submitted by the user—the first input content—are relatively broad and vague, multiple non-textual visual contents are provided as a refinement or supplement to the first input content, which largely avoids the problem of misunderstanding caused by the inability of textual supplementary content to accurately express the user's true information needs, thereby reducing the ambiguity of the supplementary information, enhancing the clarity and efficiency of the first prompt content in expressing information needs, and thus more accurately and efficiently assisting users in clarifying their true information needs, improving the accuracy of demand analysis for categories with strong style dependence.
[0098] In one scenario, the first visual content includes at least one of still images, moving images, and videos.
[0099] In one scenario, the dialogue processing device 500 further includes: a second prompt content display module, used to display second prompt content in the dialogue interface; the second prompt content includes at least two first text contents; each first text content is displayed in a preset style; each first text content has interactive attributes; the first text content is used to supplement the first input content; and a third input content display module, used to display third input content in the dialogue interface in response to a trigger operation on the first text content; the third input content is obtained based on the first text content that is selected by trigger.
[0100] Furthermore, the third input content display module is specifically used for: responding to a trigger operation on the first text content, displaying the first text content selected by the trigger in the input box; displaying the input box in the dialog interface; and responding to a modification operation on the first text content, obtaining the modified third text content, and displaying the third text content as the third input content in the dialog interface.
[0101] In one scenario, the second input content display module 530 is specifically configured to display the second input content in the dialog interface through at least one of the following: using the first visual content selected by trigger as the second input content and displaying it in the dialog interface; obtaining the second text content corresponding to the first visual content selected by trigger, and using the second text content as the second input content and displaying it in the dialog interface; merging the first input content and the first visual content selected by trigger to obtain the second input content, and displaying the second input content in the dialog interface; obtaining the second text content corresponding to the first visual content selected by trigger, merging the first input content and the second text content to obtain the second input content, and displaying the second input content in the dialog interface.
[0102] In one scenario, the dialogue processing device 500 further includes: a third prompt content display module for displaying third prompt content in the dialogue interface; the third prompt content includes at least two partial contents; the partial contents include non-textual second visual content or fourth text content; the partial contents have interactive attributes; the partial contents are used to supplement the second input content; the third prompt content is used to refine or supplement the prompt dimension corresponding to the first prompt content; and a fourth input content display module for displaying fourth input content in the dialogue interface in response to a triggering operation on the partial contents; the fourth input content is obtained based on the selected partial contents.
[0103] In one scenario, the dialogue processing device 500 further includes a first visual content acquisition module, configured to: perform semantic parsing and category recognition based on the first input content to obtain the first category to which the object belongs and the prompt dimension corresponding to the object; the first input content contains the object; match the first category and each second category to obtain each matching degree; the second category belongs to the category that distinguishes objects based on visual features; if the matching degree is detected to be greater than the matching degree threshold, then obtain each first visual content based on the first category and the prompt dimension.
[0104] Furthermore, the first visual content acquisition module is specifically used to acquire each first visual content based on the first category and the prompt dimension through at least one of the following: searching the material library based on the first category and the prompt dimension to obtain the first visual content; or using a generative model based on the first category and the prompt dimension to obtain the first visual content.
[0105] The dialogue processing apparatus provided in this application can execute the dialogue processing method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of executing the method.
[0106] It is worth noting that in the embodiments of the above-mentioned dialogue processing device, the modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional module are only for easy differentiation and are not used to limit the scope of protection of this application.
[0107] This application also provides an electronic device that may include a processor and a memory, the memory being used to store executable instructions. The processor can be used to read the executable instructions from the memory and execute the executable instructions to implement the dialogue processing method in the above embodiments.
[0108] Figure 6 shows a schematic diagram of an electronic device. As shown in Figure 6, the electronic device 600 may include a processing unit 601 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. The RAM 603 also stores various programs and data required for the operation of the electronic device 600. The processing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output interface (I / O interface) 605 is also connected to the bus 604.
[0109] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touch screens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data.
[0110] It should be noted that the electronic device 600 shown in Figure 6 is merely an example and should not impose any limitations on the functionality or scope of this application. That is, although Figure 6 shows an electronic device 600 with various devices, it should be understood that it is not required to implement or possess all of the shown devices. More or fewer devices may be implemented alternatively.
[0111] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 609, or installed from storage device 608, or installed from ROM 602. When the computer program is executed by processing device 601, it performs the functions defined in the dialogue processing method of any embodiment of this application.
[0112] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the dialogue processing method in any embodiment of this application.
[0113] It should be noted that the computer-readable medium described above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. Computer-readable storage media can be, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, radio frequency (RF), etc., or any suitable combination thereof.
[0114] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as Hypertext Transfer Protocol (HTTP), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0115] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0116] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the dialogue processing method described in any embodiment of this application.
[0117] In this application, computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0119] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0120] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.
[0121] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. Multitasking and parallel processing may be advantageous in certain environments. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this application. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0122] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A dialogue processing method, comprising: Display a dialog interface and display the first input content in the dialog interface; The first prompt message is displayed in the dialog interface; The first prompt content includes at least two non-textual first visual content items; each of the first visual content items has interactive attributes; the first visual content items are used to supplement the first input content; In response to a triggering operation on the first visual content, a second input content is displayed in the dialog interface; The second input content is obtained based on the first visual content that was triggered and selected.
2. The method according to claim 1, wherein the first visual content includes at least one of a static image, a dynamic image, and a video.
3. The method according to claim 1, further comprising: The second prompt is displayed in the dialog interface; The second prompt includes at least two pieces of first text content; Each of the first text contents is displayed in a preset style; Each of the first text contents has interactive attributes; the first text contents are used to supplement the first input content; In response to a triggering operation on the first text content, a third input content is displayed in the dialog interface; the third input content is obtained based on the first text content that was triggered and selected.
4. The method according to claim 3, wherein displaying third input content in the dialog interface in response to a triggering operation on the first text content includes: In response to the trigger operation on the first text content, the first text content that was triggered and selected is displayed in the input box; The input box is displayed in the dialog interface; In response to the modification operation on the first text content, the modified third text content is obtained, and the third text content is displayed on the dialog interface as the third input content.
5. The method according to claim 1 or 2, wherein displaying the second input content in the dialog interface includes at least one of the following: using the first visual content that triggers selection as the second input content and displaying it in the dialog interface; obtaining the second text content corresponding to the first visual content that triggers selection and displaying the second text content as the second input content in the dialog interface; The first input content and the first visual content that was triggered to be selected are combined to obtain the second input content, and the second input content is displayed in the dialog interface; Obtain the second text content corresponding to the first visual content that triggers the selection, merge the first input content and the second text content to obtain the second input content, and display the second input content in the dialog interface.
6. The method according to claim 1, further comprising: A third prompt is displayed in the dialog interface; the third prompt includes at least two partial contents; the partial contents include a second visual content or a fourth text content that is not text; the partial contents have interactive attributes; the partial contents are used to supplement the second input content; the third prompt is used to refine or supplement the prompt dimension corresponding to the first prompt; in response to a trigger operation on the partial contents, a fourth input is displayed in the dialog interface; the fourth input is obtained based on the selected partial contents.
7. The method according to claim 1 or 2, wherein the first visual content is obtained by: performing semantic parsing and category recognition based on the first input content to obtain the first category to which the object belongs and the prompt dimension corresponding to the object; the first input content includes the object; matching the first category and each second category to obtain each matching degree; the second category belongs to the category that distinguishes objects based on visual features; if the matching degree is detected to be greater than the matching degree threshold, then each of the first visual contents is obtained based on the first category and the prompt dimension.
8. The method according to claim 7, wherein obtaining each of the first visual contents based on the first category and the prompt dimension includes at least one of the following: searching a material library based on the first category and the prompt dimension to obtain the first visual content; and obtaining the first visual content using a generative model based on the first category and the prompt dimension.
9. A dialogue processing apparatus, comprising: The first input content display module is used to display a dialog interface and display the first input content in the dialog interface; The first prompt content display module is used to display the first prompt content in the dialog interface; The first prompt content includes at least two non-textual first visual content items; each of the first visual content items has interactive attributes; the first visual content items are used to supplement the first input content; The second input content display module is used to display the second input content in the dialog interface in response to the triggering operation of the first visual content; The second input content is obtained based on the first visual content that was triggered and selected.
10. An electronic device, comprising: processor; A memory for storing executable instructions; wherein the processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the dialogue processing method according to any one of claims 1-8.
11. A computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the dialogue processing method according to any one of claims 1-8.
12. A computer program product for implementing the dialogue processing method according to any one of claims 1-8.