Assistant dialogue method and device, electronic equipment and medium

By displaying the reply optimization control in the AI ​​assistant interface and using the result detection model to optimize the reply quality, the problem of users requiring multiple operations to obtain high-quality reply is solved, and efficient and simplified response quality improvement is achieved.

CN120583061APending Publication Date: 2025-09-02VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510722354.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The conversation information entered by the user is not detailed or the search results of the AI ​​assistant are inaccurate, resulting in low reply quality. Users need to operate multiple times to obtain high-quality reply, which is cumbersome and time-consuming.

Method used

The reply optimization control is displayed in the AI ​​assistant's session interface, receive user input, and detect reply quality through the result detection model. If the conditions do not meet the conditions, generate the optimized high-quality reply results.

Benefits of technology

Simplifies user operations, reduces the time to obtain high-quality replies, and improves the quality of replies without multiple user inputs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120583061A_ABST
    Figure CN120583061A_ABST
Patent Text Reader

Abstract

The invention discloses a hand dialogue method and device, electronic equipment and a medium, and belongs to the technical field of artificial intelligence. The assistant dialogue method provided by the embodiment of the invention comprises the following steps: under the condition that a first reply message is displayed on a dialogue interface of an AI assistant, receiving a first input for a reply optimization control; the session interface comprises a user message, the user message comprises dialogue information input by a user, and the first reply message comprises a first reply result used for replying the dialogue information; displaying a second reply message in response to the first input; the second reply message comprises a second reply result, the second reply result is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than that of the first reply result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to a hand dialogue method, device, electronic device and medium. Background Art

[0002] Typically, an artificial intelligence (AI) assistant is installed in an electronic device, so that when a user needs to communicate with the AI ​​assistant, he or she can enter conversation information in the conversation interface of the AI ​​assistant. The electronic device can then use the AI ​​assistant to search the Internet based on the conversation information and display the corresponding reply results based on the search results, so that the user can know the reply results.

[0003] However, since the conversation information entered by the user may not be detailed, or the search results obtained by the AI ​​assistant on the Internet may be inaccurate, the reply results displayed by the AI ​​assistant may have low quality. Therefore, the user may need to perform multiple operations to re-enter the conversation information, so that the AI ​​assistant can re-display reply results with higher quality based on the re-entered conversation information, resulting in the user's operation being cumbersome and time-consuming in the process of displaying reply results with higher quality. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a hand dialogue method, device, electronic device and medium, which can optimize the first reply result corresponding to the dialogue information in the conversation interface of the AI ​​assistant, and display a second reply message including a second reply result with higher reply quality, without requiring the user to perform multiple operations. Therefore, it can simplify the user's operation in displaying the reply result with higher reply quality and reduce time consumption.

[0005] In a first aspect, an embodiment of the present application provides an assistant dialogue method, the method comprising: receiving a first input to a reply optimization control when a first reply message is displayed on a conversation interface of an AI assistant; the conversation interface includes a user message, the user message includes conversation information input by the user, and the first reply message includes a first reply result for replying to the conversation information; in response to the first input, a second reply message is displayed; the second reply message includes a second reply result, the second reply result is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result.

[0006] In some embodiments of the present application, the above-mentioned reply quality includes at least one of the following: the accuracy of the reply result, the completeness of the reply result, the content detail of the reply result, and the image quality parameter of the reply result.

[0007] In some embodiments of the present application, before the above-mentioned receiving of the first input to the reply optimization control, the above-mentioned method also includes: inputting the user message and the first reply message into a result detection model, detecting the reply quality of the first reply result according to the conversation type of the conversation information through the result detection model, and outputting quality information; the quality information is used to indicate the reply quality of the first reply result; if the reply quality of the first reply result does not meet the reference conditions, displaying the reply optimization control.

[0008] In some embodiments of the present application, the reply quality of the above-mentioned first reply result includes the completeness rate of the reply result and the content detail of the reply result; the above-mentioned detection of the reply quality of the first reply result based on the dialogue type of the dialogue information includes: when the dialogue type is a text generation type, detecting the completeness rate and content detail of the first reply result; the above-mentioned display of the reply optimization control when the reply quality of the first reply result does not meet the reference conditions includes: when the completeness rate of the first reply result is less than or equal to the reference completeness rate, and the content detail of the first reply result is less than or equal to the reference content detail, displaying the reply optimization control; wherein, the completeness rate of the above-mentioned second reply result is greater than the reference completeness rate, and the content detail of the second reply result is greater than the reference content detail.

[0009] In some embodiments of the present application, when the above-mentioned dialogue information indicates that the AI ​​assistant generates an event notification, the above-mentioned detection of the completeness rate and content detail of the first reply result includes: detecting the event elements of the first reply result that are missing compared with the reference elements, and detecting the event detail information of the first reply result that is missing compared with the reference event detail information; the above-mentioned method also includes: when the number of missing event elements is greater than or equal to the event number threshold, determining that the completeness rate of the first reply result is less than or equal to the reference completeness rate; when the proportion of the missing event detail information in the reference event detail information is greater than or equal to the proportion threshold, determining that the content detail of the first reply result is less than or equal to the reference content detail.

[0010] In some embodiments of the present application, the reply quality of the first reply result includes the picture quality parameters of the reply result; the detection of the reply quality of the first reply result based on the conversation type of the conversation information includes: when the conversation type is a picture generation type, detecting the picture quality parameters of the first reply result; the display of the reply optimization control when the reply quality of the first reply result does not meet the reference conditions includes: when the parameter value of the picture quality parameter of the first reply result is less than or equal to the reference picture quality parameter threshold, displaying the reply optimization control; wherein the parameter value of the picture quality parameter of the second reply result is greater than the reference picture quality parameter threshold.

[0011] In some embodiments of the present application, when the above-mentioned dialogue information indicates that the AI ​​assistant generates a picture, the above-mentioned detection of the picture quality parameters of the first reply result includes: detecting the feature difference between the body features of the picture object of the first reply result and the reference body features; the above-mentioned method also includes: when the feature difference is greater than or equal to the feature difference threshold, determining that the parameter value of the picture quality parameter of the first reply result is less than or equal to the reference picture quality parameter threshold.

[0012] In some embodiments of the present application, the reply quality of the first reply result includes the accuracy of the reply result and the completeness of the reply result; the above-mentioned detection of the reply quality of the first reply result based on the dialogue type of the dialogue information includes: when the dialogue type is a skill call type, obtaining skill detection auxiliary information from the electronic device based on the dialogue information; wherein the skill detection auxiliary information includes at least one of the following: the above-mentioned context information of the dialogue information, and application data information matching the dialogue topic of the dialogue information; detecting the accuracy and completeness of the first reply result based on the dialogue information and the skill detection auxiliary information; the above-mentioned display of the reply optimization control when the reply quality of the first reply result does not meet the reference conditions includes: displaying the reply optimization control when the accuracy of the first reply result is less than or equal to the reference accuracy rate, and the completeness rate of the first reply result is less than or equal to the reference completeness rate; wherein the accuracy of the second reply result is greater than the reference accuracy rate, and the completeness rate of the second reply result is greater than the reference completeness rate.

[0013] In some embodiments of the present application, when the above-mentioned dialogue indication information instructs the AI ​​assistant to generate clothing suggestions and the application data information is weather information, the above-mentioned detection of the accuracy and completeness of the first reply result includes: detecting the suggestion difference between the suggestion content of the first reply result and the reference suggestion content, and detecting the suggestion elements of the first reply result that are missing compared with the reference suggestion elements, and the reference suggestion content is determined by the weather information; the above-mentioned method also includes: when the suggestion difference is greater than or equal to the suggestion difference threshold, determining that the accuracy of the first reply result is less than or equal to the reference accuracy rate; when the number of missing suggestion elements is greater than or equal to the missing suggestion number threshold, determining that the completeness of the first reply result is less than or equal to the reference completeness rate.

[0014] In some embodiments of the present application, the reply quality of the above-mentioned first reply result includes the accuracy of the reply result; the above-mentioned detection of the reply quality of the first reply result based on the dialogue type of the dialogue information includes: when the dialogue type is an online query type, obtaining Internet detection auxiliary information associated with the dialogue information from the Internet; detecting the accuracy of the first reply result based on the Internet detection auxiliary information; the above-mentioned display of the reply optimization control when the reply quality of the first reply result does not meet the reference conditions includes: when the accuracy of the first reply result is less than or equal to the reference accuracy, displaying the reply optimization control; wherein, the accuracy of the above-mentioned second reply result is greater than the reference accuracy.

[0015] In some embodiments of the present application, when the above-mentioned dialogue information instructs the AI ​​assistant to query the information of the specified object and the Internet detection auxiliary information is the reference content, the above-mentioned detection of the accuracy of the first reply result based on the Internet detection auxiliary information includes: detecting the content difference between the reply content of the first reply result and the reference content; the above-mentioned method also includes: when the content difference is greater than or equal to the content difference threshold, determining that the accuracy of the first reply result is less than or equal to the reference accuracy.

[0016] In some embodiments of the present application, the above method also includes: obtaining multiple training samples, the training samples including: sample conversation information, sample reply messages corresponding to the sample conversation information, conversation types of the sample conversation information, and sample quality information, the sample reply message including the sample reply result, and the sample quality information is used to indicate the reply quality of the sample reply result; based on the multiple training samples, the initial model is trained to obtain a result detection model.

[0017] In some embodiments of the present application, the quality information includes at least one of the following: status information, quality content information, content information, and test result information. The status information indicates whether the first reply is correct; the quality content information indicates at least one of the following: whether the first reply is complete, whether the content of the first reply is detailed, or whether the image quality of the first reply is acceptable; the content information indicates the difference between the first reply and a reference reply; and the test result information indicates whether the reply quality of the first reply is acceptable.

[0018] In some embodiments of the present application, before displaying the second reply message, the method further includes: generating optimized prompt words based on the quality information and the conversation information; inputting the optimized prompt words into the reply generation model, and outputting the second reply message.

[0019] In some embodiments of the present application, the second reply message is generated by a reply generation model, and the reply generation model is deployed on the first device; the method further includes: when the reply quality of the first reply result does not meet the reference conditions, sending log information to the first device, and the log information includes at least one of the following: conversation information, conversation type of the conversation information, and quality information; wherein the log information is used to train the reply generation model.

[0020] In the second aspect, an embodiment of the present application provides an assistant dialogue method, which includes: obtaining quality information when a user message is displayed on the conversation interface of the AI ​​assistant; the user message includes dialogue information input by the user, and the quality information is used to indicate the reply quality of the first reply result, and the first reply result is the reply result included in the first reply message; when the reply quality of the first reply result does not meet the reference conditions, a second reply message is displayed; the second reply message includes a second reply result, and the second reply result is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result.

[0021] In a third aspect, an embodiment of the present application provides an assistant dialogue device, which may include: a receiving module for receiving a first input to a reply optimization control when a first reply message is displayed on a conversation interface of the AI ​​assistant; the conversation interface includes a user message, the user message includes conversation information input by the user, and the first reply message includes a first reply result for replying to the conversation information. A display module for displaying a second reply message in response to the first input received by the receiving module; the second reply message includes a second reply result, the second reply result being obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result being higher than the reply quality of the first reply result.

[0022] In some embodiments of the present application, the above-mentioned reply quality includes at least one of the following: the accuracy of the reply result, the completeness of the reply result, the content detail of the reply result, and the image quality parameter of the reply result.

[0023] In some embodiments of the present application, the assistant dialogue device provided by the embodiments of the present application may further include: a processing module, which is used to input the user message and the first reply message into a result detection model before the receiving module receives the first input to the reply optimization control, and detect the reply quality of the first reply result according to the dialogue type of the dialogue information through the result detection model, and output quality information; the quality information is used to indicate the reply quality of the first reply result. The above-mentioned display module is also used to display the reply optimization control when the reply quality of the first reply result does not meet the reference conditions.

[0024] In some embodiments of the present application, the reply quality of the above-mentioned first reply result includes the completeness rate of the reply result and the content detail of the reply result. The above-mentioned processing module is specifically used to detect the completeness rate and content detail of the first reply result when the conversation type is a text generation type. The above-mentioned display module is specifically used to display the reply optimization control when the completeness rate of the first reply result detected by the processing module is less than or equal to the reference completeness rate, and the content detail of the first reply result detected by the processing module is less than or equal to the reference content detail; wherein, the completeness rate of the above-mentioned second reply result is greater than the reference completeness rate, and the content detail of the second reply result is greater than the reference content detail.

[0025] In some embodiments of the present application, when the dialogue information indicates that the AI ​​assistant generates an event notification, the processing module is specifically used to detect event elements that are missing from the first reply result compared to the reference event elements, and to detect event detail information that is missing from the first reply result compared to the reference event detail information. The processing module is also used to determine that the completeness rate of the first reply result is less than or equal to the reference completeness rate when the number of missing event elements is greater than or equal to the event number threshold; and to determine that the content detail of the first reply result is less than or equal to the reference content detail when the proportion of the missing event detail information in the reference event detail information is greater than or equal to the proportion threshold.

[0026] In some embodiments of the present application, the reply quality of the first reply result includes an image quality parameter of the reply result. The processing module is specifically configured to detect the image quality parameter of the first reply result when the conversation type is an image generation type. The display module is specifically configured to display a reply optimization control when the parameter value of the image quality parameter of the first reply result detected by the processing module is less than or equal to a reference image quality parameter threshold; wherein the parameter value of the image quality parameter of the second reply result is greater than the reference image quality parameter threshold.

[0027] In some embodiments of the present application, when the conversation information instructs the AI ​​assistant to generate an image, the processing module is specifically configured to detect a degree of difference between the body features of the image object in the first reply result and the reference body features. The processing module is further configured to determine, when the degree of difference is greater than or equal to a threshold value, that the parameter value of the image quality parameter of the first reply result is less than or equal to a reference image quality parameter threshold.

[0028] In some embodiments of the present application, the reply quality of the above-mentioned first reply result includes the accuracy rate of the reply result and the completeness rate of the reply result. The above-mentioned processing module is specifically used to obtain skill detection auxiliary information from the assistant dialogue device according to the dialogue information when the dialogue type is a skill call type; wherein the above-mentioned skill detection auxiliary information includes at least one of the following: the above-mentioned context information of the dialogue information, the application data information matching the dialogue topic of the dialogue information; and detect the accuracy rate and completeness rate of the first reply result according to the dialogue information and the skill detection auxiliary information. The above-mentioned display module is specifically used to display the reply optimization control when the accuracy rate of the first reply result detected by the processing module is less than or equal to the reference accuracy rate, and the completeness rate of the first reply result detected by the processing module is less than or equal to the reference completeness rate; wherein the accuracy rate of the above-mentioned second reply result is greater than the reference accuracy rate, and the completeness rate of the second reply result is greater than the reference completeness rate.

[0029] In some embodiments of the present application, when the dialogue indication information instructs the AI ​​assistant to generate clothing suggestions, and the application data information is weather information, the above-mentioned processing module is specifically used to detect the difference between the suggestion content of the first reply result and the reference suggestion content, and detect the suggestion elements of the first reply result that are missing compared to the reference suggestion elements, where the reference suggestion content is determined by the weather information. The above-mentioned processing module is also used to determine that the accuracy of the first reply result is less than or equal to the reference accuracy rate when the suggestion difference is greater than or equal to the suggestion difference threshold; and to determine that the completeness rate of the first reply result is less than or equal to the reference completeness rate when the number of missing suggestion elements is greater than or equal to the missing suggestion number threshold.

[0030] In some embodiments of the present application, the response quality of the first response result includes the accuracy of the response result. The processing module is specifically configured to, when the conversation type is an online query type, obtain internet detection auxiliary information associated with the conversation information from the internet; and detect the accuracy of the first response result based on the internet detection auxiliary information. The display module is specifically configured to display a response optimization control when the accuracy of the first response result detected by the processing module is less than or equal to a reference accuracy rate; wherein the accuracy of the second response result is greater than the reference accuracy rate.

[0031] In some embodiments of the present application, when the conversation information instructs the AI ​​assistant to query for information about a specified object, and the internet detection auxiliary information is reference content, the processing module is specifically configured to detect the content difference between the reply content of the first reply result and the reference content. The processing module is further configured to determine that the accuracy of the first reply result is less than or equal to the reference accuracy rate if the content difference is greater than or equal to a content difference threshold.

[0032] In some embodiments of the present application, the above-mentioned processing module is also used to obtain multiple training samples, and the training samples include: sample conversation information, sample reply messages corresponding to the sample conversation information, conversation types of the sample conversation information, and sample quality information. The sample reply message includes a sample reply result, and the sample quality information is used to indicate the reply quality of the sample reply result; and based on the multiple training samples, the initial model is trained to obtain a result detection model.

[0033] In some embodiments of the present application, the above-mentioned quality information includes at least one of the following: status information, quality content information, content information, and detection result information; wherein, the status information is used to indicate whether the first reply result is correct; the quality content information is used to indicate at least one of the following: whether the first reply result is complete, whether the content of the first reply result is detailed, and whether the picture quality of the first reply result is qualified; the content information is used to indicate the difference between the first reply result and the reference reply result; the detection result information is used to indicate whether the reply quality of the first reply result is qualified.

[0034] In some embodiments of the present application, the above-mentioned processing module is also used to generate optimized prompt words based on the quality information and dialogue information before the display module displays the second reply message; and input the optimized prompt words into the reply generation model to output the second reply message.

[0035] In some embodiments of the present application, the second reply message is generated by a reply generation model, and the reply generation model is deployed on the first device. The assistant dialogue device provided by the embodiment of the present application may also include: a sending module for sending log information to the first device when the reply quality of the first reply result does not meet the reference condition, and the log information includes at least one of the following: dialogue information, dialogue type of the dialogue information, and quality information; wherein the log information is used to train the reply generation model.

[0036] In a fourth aspect, an embodiment of the present application provides an assistant dialogue device, which may include: a processing module for obtaining quality information when displaying a user message on the conversation interface of the AI ​​assistant; the user message includes dialogue information input by the user, and the quality information is used to indicate the reply quality of a first reply result, which is the reply result included in the first reply message. A display module is used to display a second reply message when the reply quality of the first reply result does not meet the reference conditions; the second reply message includes a second reply result, which is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result.

[0037] In a fifth aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the method of the first aspect are implemented, or the steps of the method of the second aspect are implemented.

[0038] In a sixth aspect, an embodiment of the present application provides a readable storage medium, which stores a program or instruction. When the program or instruction is executed by a processor, it implements the steps of the method of the first aspect, or implements the steps of the method of the second aspect.

[0039] In the seventh aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the steps of the method of the first aspect, or to implement the steps of the method of the second aspect.

[0040] In an eighth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the steps of the method of the first aspect, or to implement the steps of the method of the second aspect.

[0041] In an embodiment of the present application, when a first reply message is displayed on a conversation interface of an AI assistant, the electronic device can display a second reply message based on a first input of a user to a reply optimization control; wherein the conversation interface includes a user message, the user message includes conversation information input by the user, the first reply message includes a first reply result for replying to the conversation information, the second reply message includes a second reply result, the second reply result is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result. Since, when the first reply message is displayed on the conversation interface, the electronic device can directly optimize the first reply result in the first reply message based on a single input of the user to obtain a second reply result with a higher reply quality, and display the second reply result with a higher reply quality in the conversation interface without the user having to perform multiple operations, thus simplifying the user's operations in displaying the reply result with a higher reply quality and reducing time consumption.

[0042] In an embodiment of the present application, when a user message is displayed on the conversation interface of the AI ​​assistant, the electronic device can obtain quality information, where the user message includes dialogue information input by the user, and the quality information is used to indicate the reply quality of the first reply result, which is the reply result included in the first reply message; and when the reply quality of the first reply result does not meet the reference conditions, a second reply message is displayed; the second reply message includes a second reply result, which is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result. Since the electronic device can directly obtain the quality information indicating the reply quality of the first reply result when displaying the user message on the conversation interface, and directly optimize the first reply result in the first reply message when the reply quality of the first reply result does not meet the reference conditions, obtain the second reply result with higher reply quality, and display the second reply result with higher reply quality in the conversation interface without the user having to perform multiple operations. Therefore, the user's operation in displaying the reply result with higher reply quality can be simplified and the time consumption can be reduced. Moreover, since the electronic device can obtain the first reply result corresponding to the conversation message, it does not display the first reply result in the conversation interface, and directly displays the second reply message when the reply quality of the first reply result does not meet the reference conditions, that is, displays the optimized second reply result. Therefore, the user can avoid determining the first reply result as the accurate reply result. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flowchart of an assistant dialogue method provided in some embodiments of the present application;

[0044] Figure 2A is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0045] Figure 2B is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0046] Figure 3A is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0047] Figure 3B is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0048] Figure 4A is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0049] Figure 4B is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0050] Figure 5A is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0051] Figure 5B is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0052] Figure 6 is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0053] Figure 7 is a flowchart of an assistant dialogue method provided in some embodiments of the present application;

[0054] Figure 8 is a flowchart of an assistant dialogue method provided in some embodiments of the present application;

[0055] Figure 9 is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0056] Figure 10 is a flowchart of an assistant dialogue method provided in some embodiments of the present application;

[0057] Figure 11 is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0058] Figure 12 is a flowchart of an assistant dialogue method provided in some embodiments of the present application;

[0059] Figure 13 is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0060] Figure 14 is a flowchart of an assistant dialogue method provided in some embodiments of the present application;

[0061] Figure 15 is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0062] Figure 16 is a schematic diagram of the data structure of sample data provided in some embodiments of the present application;

[0063] Figure 17 is a schematic diagram of a post-training judgment process provided by some embodiments of the present application;

[0064] Figure 18A is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0065] Figure 18B is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0066] Figure 19 is a flowchart of an assistant dialogue method provided in some embodiments of the present application;

[0067] Figure 20 It is a schematic diagram of log information uploading and subsequent application provided by some embodiments of the present application;

[0068] Figure 21 This is a schematic diagram of the model training effect provided by some embodiments of the present application;

[0069] Figure 22 is a flowchart of an assistant dialogue method provided in some embodiments of the present application;

[0070] Figure 23A is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0071] Figure 23B is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0072] Figure 23C is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0073] Figure 24A is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0074] Figure 24B is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0075] Figure 24C is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0076] Figure 25A is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0077] Figure 25B is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0078] Figure 25C is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0079] Figure 26A is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0080] Figure 26B is a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0081] Figure 26Cis a schematic diagram of a mobile phone interface provided by some embodiments of the present application;

[0082] Figure 27 is a schematic diagram of the structure of an assistant dialogue device provided in some embodiments of the present application;

[0083] Figure 28 is a schematic diagram of the structure of an assistant dialogue device provided in some embodiments of the present application;

[0084] Figure 29 is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of the present application;

[0085] Figure 30 This is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of the present application. DETAILED DESCRIPTION

[0086] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0087] The terms used in the implementation section of this application are only used to explain the specific embodiments of this application and are not intended to limit this application.

[0088] The following is an explanation of the professional terms used in the embodiments of this application.

[0089] An AI assistant is an intelligent assistance tool based on artificial intelligence technology that can automate tasks, provide search and information services, personalize responses, and optimize the user experience. It aims to help users efficiently complete various tasks, obtain information, or create content. AI assistants have natural language processing capabilities, understanding and processing natural language commands input by users via voice or text, and recognizing, parsing, and generating appropriate responses. Typically, users enter conversational information into the AI ​​assistant's conversational interface, which then displays responses based on the conversational information. Specifically, user conversational information can include queries or commands—requests or commands issued to the AI ​​assistant via typing on an input keyboard or voice input, such as asking about the weather, requesting reminders, or seeking information. The AI ​​assistant's responses can include responses—the answers or actions generated by the AI ​​assistant in response to the user's queries or commands.

[0090] On-device: This refers to data processing and computing performed locally on devices (such as mobile phones and Internet of Things (IoT) devices) rather than relying on cloud servers. This reduces data transmission latency, improves response speed, and enhances privacy protection.

[0091] Multimodality: refers to the simultaneous processing of multiple data types (text, images, voice, sensor data, etc.) and achieving cross-modal understanding to provide richer interaction and application experience.

[0092] An on-device large model refers to an artificial intelligence (AI) model with a large number of parameters deployed on the terminal. For example, the AI ​​model can be a large language model with billions of parameters, capable of processing multimodal data such as text and images, to achieve high-performance, low-latency intelligent services.

[0093] Intelligent Agent: An autonomous AI program running on a phone that can understand user intent, perform complex tasks, and answer complex semantic questions. In this context, this primarily refers to the AI ​​assistant installed on phones, which provides personalized assistance services.

[0094] User log: It records the user's interactive behavior data on the device. For example, the interactive behavior data may include clicks, voice commands, application usage time, etc. In this article, it mainly refers to the process data of the user obtaining answers after interacting with the mobile phone intelligent agent, which is used to optimize and improve the performance of the intelligent agent.

[0095] Screen recognition: It uses computer vision technology to analyze and understand the content on the screen, identify text, images and other information, and provide users with intelligent interaction and assistance.

[0096] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0097] The hand dialogue method, device, electronic device, and medium provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.

[0098] It should be noted that the assistant conversation method provided in the embodiments of the present application can be executed by electronic devices such as mobile phones, tablet computers, laptop computers, PDAs, and in-vehicle electronic devices. In some embodiments of the present application, the assistant conversation method provided in the embodiments of the present application is described by taking an electronic device as the execution subject to execute the assistant conversation method.

[0099] The assistant dialogue method provided in the embodiments of the present application can be applied to AI interaction scenarios.

[0100] Among them, one specific application scenario is a scenario in which a user uses an AI assistant to generate text, another specific application scenario is a scenario in which a user uses an AI assistant to generate pictures, another specific application scenario is a scenario in which a user uses an AI assistant to query clothing suggestions, and another specific application scenario is a scenario in which a user uses an AI assistant to query professional knowledge.

[0101] Figure 1 The figure shows a flow chart of the assistant dialogue method provided in the embodiment of the present application. Figure 1 As shown, the assistant dialogue method provided in the embodiment of the present application may include the following steps 101 and 102.

[0102] Step 101: When a first reply message is displayed on the conversation interface of the AI ​​assistant, the electronic device receives a first input to a reply optimization control.

[0103] In an embodiment of the present application, the above-mentioned conversation interface includes a user message, the user message includes conversation information input by the user, and the first reply message includes a first reply result for replying to the conversation information.

[0104] In some embodiments of the present application, the above-mentioned dialogue information can be used to request the AI ​​assistant to generate text, or can be used to request the AI ​​assistant to generate images, or can be used to request the AI ​​assistant to generate suggestions, or can be used to request the AI ​​assistant to query answers to questions. Of course, the dialogue information can also request the AI ​​assistant to perform other operations, which are not limited in the embodiments of the present application.

[0105] In some examples, the above-mentioned generated text may include, but is not limited to, generating event notifications, generating comment copy text, generating circle of friends copy text, generating comment text, and generating chat text.

[0106] In some examples, the above-mentioned generated pictures may include but are not limited to generated pictures of people, generated pictures of animals, and generated pictures of scenery.

[0107] In some examples, the above-mentioned generating suggestions may include, but is not limited to, generating clothing suggestions, generating shopping suggestions, generating travel planning suggestions, and generating learning planning suggestions.

[0108] In some examples, the query questions may include, but are not limited to, answers to online query questions and answers to query questions in an electronic device.

[0109] In some embodiments of the present application, the above-mentioned conversation information may include at least one of the following: text, voice, picture, video, etc. Of course, the question information may also include other information, which is not limited in the embodiments of the present application.

[0110] In some examples, when the conversation information includes text, the text can be used to describe the operation that the user requires the AI ​​assistant to perform.

[0111] For example, assuming that the conversation information includes text, the text may be "Help me write a meeting notice", and the "Help me write a meeting notice" text is used to describe the user's need for the AI ​​assistant to generate a meeting notice.

[0112] For another example, assuming that the conversation information includes text, the text may be "Help me draw a girl dancing ballet". The text "Help me draw a girl dancing ballet" is used to describe that the user needs the AI ​​assistant to generate a character picture, and the character in the character picture is a girl dancing ballet.

[0113] For another example, assuming that the conversation information includes text, the text may be "What is suitable to wear for tomorrow's weather in City A?" The text "What is suitable to wear for tomorrow's weather in City A" is used to describe that the user needs the AI ​​assistant to query the weather in City A and generate clothing suggestions.

[0114] For another example, assuming that the conversation information includes text, the text may be "What is the main virus of the current cold, and which vaccine is better?" The text "What is the main virus of the current cold, and which vaccine is better?" is used to describe the current cold virus that the user needs the AI ​​assistant to query and generate vaccination recommendations.

[0115] In some examples, where the conversation information includes speech, the speech can be used to describe the actions that the user wants the AI ​​assistant to perform.

[0116] In some examples, when the conversation information includes a picture, the picture may include questions that the user needs to ask the AI ​​assistant.

[0117] In some examples, where the conversation information includes a video, the video may include questions that the user needs the AI ​​assistant to inquire about.

[0118] In some embodiments of the present application, in the case of a conversation interface of an AI assistant displayed on an electronic device, a user can input conversation information in the conversation interface, so that the electronic device can generate a first reply result based on the conversation information through the AI ​​assistant, and display the first reply message and the reply optimization control in the conversation interface, so that the user can make a first input to the reply optimization control.

[0119] In some examples, the electronic device can input the conversation information into a response generation model corresponding to the AI ​​assistant, and generate a first response result based on the conversation information through the response generation model.

[0120] Optionally, the reply generation model can be deployed in an electronic device, so that the electronic device can directly input the conversation information into the reply generation model to obtain a first reply result generated by the reply generation model.

[0121] In another example, the reply generation model can be deployed on other devices, such as servers, other electronic devices, etc., so that the electronic device can send the conversation information to the other device, so that the other device can input the conversation information into the reply generation model to obtain a first reply result, and then the electronic device can receive the first reply result from the other device.

[0122] In some examples, the electronic device may directly display the first reply message and the reply optimization control after obtaining the first reply result; or, the electronic device may first display the first reply message and then display the reply optimization control after generating the first reply message.

[0123] For example, let’s take a mobile phone as an example. Figure 2A As shown, the mobile phone displays the conversation interface 10 of the AI ​​assistant, which includes an input box 11, so that the user can enter conversation information in the input box 11, such as the text "help me write a meeting notice". Figure 2B As shown, after the user enters the text "Help me write a meeting notice" in the input box 11, the mobile phone can output a first reply message based on the text "Help me write a meeting notice" through the AI ​​assistant. The first reply message includes the first reply result, such as the meeting notice 12, and displays the reply optimization control 13, so that the user can control the mobile phone to display the second reply message through the first input of the reply optimization control 13, and the second reply message includes the optimized meeting notice.

[0124] Another example is given, such as Figure 3A As shown, the mobile phone displays the conversation interface 10 of the AI ​​assistant, which includes an input box 11, so that the user can enter conversation information in the input box 11, such as the text "help me draw a girl dancing ballet". Figure 3B As shown, after the user enters the text "Help me draw a girl dancing ballet" in the input box 11, the mobile phone can use the AI ​​assistant to display a first reply message in the conversation interface 10 based on the text "Help me draw a girl dancing ballet". The first reply message includes a first reply result, such as picture 14, and displays a reply optimization control 13, so that the user can control the mobile phone to display a second reply message through the first input of the reply optimization control 13, and the second reply message includes an optimized picture.

[0125] Another example is given, such as Figure 4A As shown, the mobile phone displays the conversation interface 10 of the AI ​​assistant, which includes an input box 11, so that the user can enter conversation information in the input box 11, such as the text "What is suitable to wear in tomorrow's weather in City A". Figure 4B As shown, after the user enters the text "What is suitable to wear for tomorrow's weather in City A" in the input box 11, the mobile phone can use the AI ​​assistant to display a first reply message in the conversation interface 10 based on the text "What is suitable to wear for tomorrow's weather in City A". The first reply message includes a first reply result, such as weather and clothing suggestions 15, and displays a reply optimization control 13, so that the user can control the mobile phone to display a second reply message through the first input of the reply optimization control 13, and the second reply message includes optimized weather and clothing suggestions.

[0126] Another example is given, such as Figure 5A As shown, the mobile phone displays the conversation interface 10 of the AI ​​assistant, which includes an input box 11, so that the user can enter conversation information in the input box 11, such as "What is the main virus of the current cold, and which vaccine is better to take?" Figure 5B As shown, after the user enters the text "What is the main virus of the current cold, and what vaccine is better" in the input box 11, the mobile phone can use the AI ​​assistant to display a first reply message in the conversation interface 10 based on the text "What is the main virus of the current cold, and what vaccine is better". The first reply message includes a first reply result, such as virus and vaccination recommendations 16, and displays a reply optimization control 13, so that the user can control the mobile phone to display a second reply message through the first input of the reply optimization control 13, and the second reply message includes optimized virus and vaccination recommendations.

[0127] In some examples, when an electronic device displays an AI assistant's settings interface, the settings interface includes an optimization function control, so that the electronic device can enable a reply optimization function based on a user's click input on the optimization function control. Furthermore, after the electronic device generates a first reply message based on the conversation information through the AI ​​assistant, the reply optimization control can be displayed in the conversation interface. The optimization function control is used to enable or disable the reply optimization function, which is a function for the electronic device to optimize the reply results.

[0128] For example, Figure 6 As shown, the mobile phone displays an AI assistant settings interface 17, which includes an optimization function control 18. The user can click on the optimization function control 18 to enable the mobile phone to activate the reply optimization function. Furthermore, after the mobile phone generates a first reply message based on the conversation information through the AI ​​assistant, the mobile phone can display the first reply message and the reply optimization control in the above conversation interface.

[0129] In an embodiment of the present application, the above-mentioned reply optimization control is used to control the optimization of the first reply result.

[0130] In some embodiments of the present application, the first input is used to optimize the first reply result.

[0131] Exemplarily, the above-mentioned first input includes but is not limited to: touch input of the user to the display screen of the electronic device through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible input, which can be determined according to actual use needs and is not limited in the embodiment of the present invention. The specific gesture in the embodiment of the present application can be any one of a single-click gesture, a sliding gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double-press gesture, and a double-click gesture; the click input in the embodiment of the present application can be a single-click input, a double-click input, or any number of click inputs, etc., and can also be a long press input or a short press input. For example, the above-mentioned first input can be: a single-click input or a long press input of the reply optimization control by the user.

[0132] Step 102: The electronic device displays a second reply message in response to the first input.

[0133] In an embodiment of the present application, the above-mentioned second reply message includes a second reply result, which is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result.

[0134] In some embodiments of the present application, the above-mentioned reply quality includes at least one of the following: the accuracy of the reply result, the completeness of the reply result, the content detail of the reply result, and the image quality parameter of the reply result.

[0135] In some examples, the accuracy rate of the reply result is used to indicate the probability that the reply result is a correct result. The accuracy rate may be positively correlated with the probability that the reply result is a correct result, that is, the higher the accuracy rate, the higher the probability that the reply result is a correct result.

[0136] In some examples, the completeness rate of the response result is used to indicate the completeness of the elements in the response result compared to the reference elements. The completeness rate may be positively correlated with the completeness of the elements in the response result compared to the reference elements, that is, the higher the completeness rate, the more complete the elements in the response result are compared to the reference elements.

[0137] In some examples, the content detail level of the reply result is used to indicate the level of detail of the detail information in the reply result compared to the reference detail information. The content detail level may be positively correlated with the level of detail of the detail information in the reply result compared to the reference detail information; that is, the higher the content detail level, the more detailed the detail information in the reply result is compared to the reference detail information.

[0138] In some examples, the image quality parameter of the response result is used to indicate the image quality of the image. The image quality parameter is positively correlated with the image quality of the image, that is, the higher the image quality parameter, the higher the image quality of the image.

[0139] It can be seen that since the reply quality can include at least one of the accuracy of the reply result, the completeness of the reply result, the content detail of the reply result and the image quality parameters of the reply result, that is, the reply quality can include at least one information, and each information can reflect the quality of the reply result from different dimensions. Therefore, in subsequent steps, the electronic device can accurately optimize the first reply result based on the reply quality, thereby obtaining a second reply result with higher reply quality.

[0140] In some embodiments of the present application, the electronic device may directly display the second reply message in the conversation interface of the AI ​​assistant; or, the electronic device may update the first reply message in the conversation interface of the AI ​​assistant to the second reply message.

[0141] An embodiment of the present application provides an assistant conversation method, in which, when a first reply message is displayed on a conversation interface of an AI assistant, an electronic device can display a second reply message based on a first input of a user to a reply optimization control; wherein, the conversation interface includes a user message, the user message includes conversation information input by the user, the first reply message includes a first reply result for replying to the conversation information, the second reply message includes a second reply result, the second reply result is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result. Since, when the first reply message is displayed on the conversation interface, the electronic device can directly optimize the first reply result in the first reply message based on a single input of the user, obtain a second reply result with a higher reply quality, and display the second reply result with a higher reply quality in the conversation interface without the user having to perform multiple operations, thus simplifying the user's operation in displaying a reply result with a higher reply quality and reducing time consumption.

[0142] In some embodiments of the present application, Figure 1 ,like Figure 7 As shown, before "receiving the first input to the reply optimization control" in the above step 101, the assistant dialogue method provided by the embodiment of the present application can also include the following steps 201 and 202, and the above step 101 can be specifically implemented through the following step 101a.

[0143] Step 201: When the first reply message is displayed on the conversation interface of the AI ​​assistant, the electronic device inputs the user message and the first reply message into the result detection model, detects the reply quality of the first reply result according to the conversation type of the conversation information through the result detection model, and outputs quality information.

[0144] In some examples, the electronic device can extract the user message and the first reply message from the conversation interface of the AI ​​assistant, and input the user message and the first reply message into the result detection model; or, the electronic device can take a screenshot of the conversation interface to obtain a screenshot image, and input the screenshot image into the result detection model.

[0145] In the embodiment of the present application, the above-mentioned quality information is used to indicate the reply quality of the first reply result.

[0146] In some examples, the conversation type of the conversation information may include at least one of the following: a text generation type, an image generation type, a skill call type, and an online query type.

[0147] The text generation type can be understood as a type of conversation message used to request an AI assistant to generate text. The image generation type can be understood as a type of conversation message used to request an AI assistant to generate an image. The skill invocation type can be understood as a type of conversation message used to request an AI assistant to generate suggestions based on application data matching the conversation topic. The online query type can be understood as a type of conversation message used to request an AI assistant to perform an online query.

[0148] In some examples, after the electronic device inputs the user message and the first reply message into the result detection model, the result detection model can perform semantic analysis on the conversation information in the user message and determine the conversation type of the conversation information based on the semantics of the conversation information.

[0149] For example, assuming that the dialogue information is the text "Help me write a meeting notice", the result detection model can perform semantic analysis on the text "Help me write a meeting notice" and determine that the semantics of the text "Help me write a meeting notice" is to request the AI ​​assistant to generate a meeting notice, and the meeting notice is text. That is to say, the semantics of the text "Help me write a meeting notice" is to request the AI ​​assistant to generate text. Therefore, the result detection model can determine that the dialogue type of the text "Help me write a meeting notice" is a text generation type.

[0150] For another example, assuming the dialogue information is the text "Help me draw a girl dancing ballet", the result detection model can perform semantic analysis on the text "Help me draw a girl dancing ballet" and determine that the semantics of the text "Help me draw a girl dancing ballet" is to request the AI ​​assistant to generate a picture of a girl dancing ballet. In other words, the semantics of the text "Help me draw a girl dancing ballet" is to request the AI ​​assistant to generate a picture. Therefore, the result detection model can determine that the dialogue type of the text "Help me draw a girl dancing ballet" is a picture generation type.

[0151] For another example, assuming the dialogue information is the text "What to wear in City A tomorrow's weather", the result detection model can perform semantic analysis on the text "What to wear in City A tomorrow's weather" to determine that the semantics of the text "What to wear in City A tomorrow's weather" is to request the AI ​​assistant to query the weather in City A tomorrow and generate clothing suggestions based on the weather. In other words, the semantics of the text "What to wear in City A tomorrow's weather" is to request the AI ​​assistant to generate clothing suggestions based on application data information that matches the weather theme. Therefore, the result detection model can determine that the dialogue type of the text "What to wear in City A tomorrow's weather" is a skill call type.

[0152] For another example, assuming the dialogue information is the text "What is the main virus of the current cold, and which vaccine is better?", the result detection model can perform semantic analysis on the text "What is the main virus of the current cold, and which vaccine is better?" and determine that the semantics of the text "What is the main virus of the current cold, and which vaccine is better?" is to request the AI ​​assistant to query the current cold virus and vaccination recommendations. The current cold virus and vaccination recommendations involve more professional fields and have high real-time requirements. Therefore, the result detection model can determine that the dialogue type of the text "What is the main virus of the current cold, and which vaccine is better?" is an online query type.

[0153] In some examples, the above-mentioned quality information includes at least one of the following: status information, quality content information, content information, and detection result information; wherein, the status information is used to indicate whether the first reply result is correct; the quality content information is used to indicate at least one of the following: whether the first reply result is complete, whether the content of the first reply result is detailed, and whether the picture quality of the first reply result is qualified; the content information is used to indicate the difference between the first reply result and the reference reply result; the detection result information is used to indicate whether the reply quality of the first reply result is qualified.

[0154] Optionally, the above-mentioned status information may include at least one of the following: text, numbers, etc.

[0155] In which, when the status information includes text, the text may be "Is it abnormal: X", where X can be yes or no. When X is yes, the status information is used to indicate that the status of the first reply result is abnormal, that is, the first reply result is incorrect; when X is no, the status information is used to indicate that the status of the first reply result is normal, that is, the first reply result is correct.

[0156] It should be noted that the X in the above “Is it abnormal: X” is used to refer to a character, and the X can also be replaced by other letters or characters, which is not limited in this embodiment of the present application.

[0157] When the status information includes a number, when the number is a first preset value, the status information is used to indicate that the status of the first reply result is abnormal, that is, the first reply result is incorrect; when the number is a second preset value, the status information is used to indicate that the status of the first reply result is normal, that is, the first reply result is correct.

[0158] Optionally, the quality content information may include at least one of the following: text, numbers, etc.

[0159] Among them, when the quality content information includes text, the text can be "Is it high quality: Y", where Y can be yes or no. When Y is yes, the quality content information is used to indicate at least one of the following: the first reply result is complete, the content of the first reply result is detailed, and the picture quality of the first reply result is qualified; when Y is no, the quality content information is used to indicate at least one of the following: the first reply result is incomplete, the content of the first reply result is not detailed, and the picture quality of the first reply result is unqualified.

[0160] It should be noted that the Y in the above “Is it high quality: Y” is used to refer to a character, and the Y can also be replaced by other letters or characters, which is not limited in this embodiment of the present application.

[0161] When the quality content information includes a number, when the number is greater than or equal to a third preset value, the quality content information is used to indicate at least one of the following: the first reply result is complete, the content of the first reply result is detailed, and the picture quality of the first reply result is qualified; when the number is less than the third preset value, the quality content information is used to indicate at least one of the following: the first reply result is incomplete, the content of the first reply result is not detailed, and the picture quality of the first reply result is unqualified.

[0162] Optionally, the reference reply result may be a standard reply result corresponding to the dialogue information, and the reference reply result may be pre-set in the result detection model or pre-set in the electronic device.

[0163] Optionally, the above content information may include text, and the text is used to describe the difference between the first reply result and the reference reply result.

[0164] Optionally, the above-mentioned detection result information may include at least one of the following: text, numbers.

[0165] In which, when the test result information includes text, the text can be "Judgment result: Z", where Z can be qualified or unqualified. When Z is qualified, the test result information is used to indicate that the response quality of the first reply result is qualified. When Z is unqualified, the test result information is used to indicate that the response quality of the first reply result is unqualified.

[0166] It should be noted that the Z in the above “judgment result: Z” is used to refer to a character, and the Z can also be replaced by other letters or characters, which is not limited in this embodiment of the present application.

[0167] It can be seen that since the quality information can include at least one of status information, quality content information, content information and detection result information, that is, the quality information can include at least one information, and each information can indicate the response quality of the first response result from different dimensions, rather than including only one information, therefore, in subsequent steps, the electronic device can accurately determine whether the response quality of the first response result meets the reference conditions based on at least one of the at least one information.

[0168] In some examples, when the conversation type of the conversation information is a text generation type, the above-mentioned reply quality includes the completeness rate of the reply result and the content detail of the reply result; when the conversation type of the conversation information is an image generation type, the above-mentioned reply quality includes the image quality parameters of the reply result; when the conversation type of the conversation information is a skill call type, the above-mentioned reply quality includes the accuracy rate of the reply result and the completeness rate of the reply result; when the conversation type of the conversation information is an online query type, the above-mentioned reply quality includes the accuracy rate of the reply result.

[0169] It can be understood that the content included in the reply quality of the first reply result varies depending on the conversation type of the conversation information.

[0170] In some examples, a detection rule library can be set in the above-mentioned result detection model, so that the result detection model can select corresponding quality content from the detection rule library according to the dialogue type of the dialogue information, obtain the response quality required for detection, and select the detection method corresponding to the response quality from the detection rule library to detect the first response result.

[0171] For example, when the dialogue type of the dialogue information is a text generation type, the result detection model can select quality content corresponding to the text generation type from the detection rule library, that is, select the completeness rate and content detail from the detection rule library, and select a detection method corresponding to the completeness rate and content detail from the detection rule library, such as detecting the elements of the reply result to determine the completeness rate, and detecting the detailed information of the reply result to determine the content detail, and detecting the completeness rate and content detail of the first reply result.

[0172] The completeness rate of the first reply result can be a positive number, for example, the completeness rate of the first reply result can be any value between 0% and 100%. It is understood that the higher the completeness rate of the first reply result, the more complete the first reply result. The content detail of the first reply result can be a positive number, for example, the content detail of the first reply result can be any value between 0% and 100%. It is understood that the higher the content detail of the first reply result, the more detailed the content of the first reply result.

[0173] For another example, when the conversation type of the conversation information is a picture generation type, the result detection model can select the quality content corresponding to the picture generation type from the detection rule library, that is, select the picture quality parameters from the detection rule library, and select the detection method corresponding to the picture quality parameters from the detection rule library, such as detecting the object features of the picture object of the reply result to determine the picture quality parameters, and detecting the picture quality parameters of the first reply result.

[0174] Among them, the parameter value of the picture quality parameter of the first reply result can be a positive number. For example, the parameter value of the picture quality parameter of the first reply result can be any value from 0 to 100. It can be understood that the larger the parameter value of the picture quality parameter of the first reply result, the higher the picture quality of the first reply result.

[0175] For another example, when the dialogue type of the dialogue information is a skill call type, the result detection model can select quality content corresponding to the skill call type from the detection rule library, that is, select the accuracy and completeness from the detection rule library, and select a detection method corresponding to the accuracy and completeness from the detection rule library, such as detecting the content of the reply result to determine the accuracy, and detecting the elements of the reply result to determine the completeness, and detecting the accuracy and completeness of the first reply result.

[0176] Among them, the accuracy of the first reply result can be a positive number, for example, the accuracy of the first reply result can be any value between 0% and 100%. It can be understood that the higher the accuracy of the first reply result, the more likely it is that the first reply result is a correct reply result.

[0177] For another example, when the conversation type of the conversation information is a network query type, the result detection model can select the quality content corresponding to the network query type from the detection rule library, that is, select the accuracy from the detection rule library, and select the detection method corresponding to the accuracy from the detection rule library, such as detecting the content of the reply result to determine the accuracy, and detecting the accuracy of the first reply result.

[0178] Step 202: When the reply quality of the first reply result does not meet the reference condition, the electronic device displays a reply optimization control.

[0179] In some examples, the reference condition may include at least one of the following:

[0180] The completeness rate of the response results is less than or equal to the reference completeness rate;

[0181] The content detail of the response result is less than or equal to the reference content detail;

[0182] The parameter value of the picture quality parameter of the reply result is less than or equal to the reference picture quality parameter threshold;

[0183] The accuracy of the response result is less than or equal to the reference accuracy.

[0184] Optionally, the reference completeness rate, reference content detail, reference picture quality parameter, and reference accuracy rate may be pre-set in the result detection model or pre-set in the electronic device.

[0185] Among them, the above-mentioned reference completeness rate can be a positive number, and the value range of the reference completeness rate can be [10%, 100%], the above-mentioned reference content detail can be a positive number, and the value range of the reference content detail can be [10%, 100%], the reference image quality parameter threshold can be a positive number, for example [10, 100], the reference accuracy rate can be a positive number, and the value range of the reference accuracy rate can be [10%, 100%].

[0186] Optionally, when the conversation type of the conversation information is a text generation type, the above-mentioned reference conditions may include that the completeness rate of the reply result is less than or equal to the reference completeness rate and the content detail of the reply result is less than or equal to the reference content detail; when the conversation type of the conversation information is an image generation type, the above-mentioned reference conditions may include that the parameter value of the image quality parameter of the reply result is less than or equal to the reference image quality parameter threshold; when the conversation type of the conversation information is a skill call type, the above-mentioned reference conditions may include that the accuracy rate of the reply result is less than or equal to the reference accuracy rate and the completeness rate of the reply result is less than or equal to the reference completeness rate; when the conversation type of the conversation information is an online query type, the above-mentioned reference conditions may include that the accuracy rate of the reply result is less than or equal to the reference accuracy rate.

[0187] It can be understood that the content included in the reference conditions varies depending on the dialogue type of the dialogue information.

[0188] In some examples, when the reply quality of the first reply result does not meet the reference conditions, the status information in the above quality information can be "Is it abnormal: yes" or "Is it abnormal: no", the quality content information in the above quality information can be "Is it high quality: no", and the detection result information in the above quality information can be "Judgment result: qualified" or "Judgment result: unqualified".

[0189] It can be understood that if the reply quality of the first reply result does not meet the reference conditions, it can be considered that the reply quality of the first reply result is low, that is, the user may need to optimize the first reply result. Therefore, the electronic device can display the reply optimization control.

[0190] Step 101a: The electronic device receives a first input to a reply optimization control.

[0191] It should be noted that, for the description of the electronic device receiving the first input to the reply optimization control, reference may be made to the specific description in the above embodiment, and the embodiments of the present application will not be repeated here.

[0192] It can be seen that since the electronic device can first input the user message and the first reply message into the result detection model, and through the result detection model, detect the reply quality of the first reply result according to the conversation type of the conversation information, and obtain quality information, then when the reply quality of the first reply result indicated by the quality information does not meet the reference conditions, that is, when the reply quality of the first reply result is low, the reply optimization control is displayed instead of directly displaying the reply optimization control. Therefore, when the reply quality of the first reply result is high, that is, when the user may not need to optimize the first reply result, the electronic device may not display the reply optimization control, instead of displaying the reply optimization control, thereby making the conversation interface more concise.

[0193] Furthermore, when the result detection model is an end-to-end model, intervention can be made on the end-to-end to the first response result, allowing for timely optimization of the obtained first response result. Compared to related technologies that develop and launch the result generation model after offline analysis of user logs and output strategies, this can significantly shorten the response cycle and provide users with better results in a timely manner. For example, text generation results can be evaluated by the end-to-end large model and then reflected upon, with additional details added to improve the practicality of the results. Image generation results can be promptly improved through understanding of image information to address more obvious issues such as deformity and non-compliance with instructions. Skill execution results can be improved through multimodal understanding of image and text information to address issues such as missing cross-modal output information and inconsistent results. Knowledge information that requires an internet connection can also be determined by the end-to-end large model to determine whether it requires an internet connection, allowing for timely adjustments to the authenticity of the response results.

[0194] The following will take different types of conversation information as an example to illustrate the specific scheme of displaying reply optimization controls on electronic devices.

[0195] In some examples, the response quality of the first response result includes the completeness of the response result and the content detail of the response result. Figure 7 ,like Figure 8 As shown, the above step 201 can be specifically implemented through the following step 201a, and the above step 202 can be specifically implemented through the following step 202a.

[0196] Step 201a: When the first reply message is displayed on the conversation interface of the AI ​​assistant, the electronic device inputs the user message and the first reply message into the result detection model. Through the result detection model, when the conversation type is a text generation type, the completeness and content detail of the first reply result are detected, and quality information is output.

[0197] Optionally, the electronic device may first perform semantic analysis on the semantics of the first reply result through a result detection model to determine the requirements in the first reply result and the detailed information in the first reply result, and then determine the completeness rate of the first reply result based on the elements in the first reply result, and determine the content detail of the first reply result based on the detailed information in the first reply result.

[0198] Optionally, when the dialogue information indicates that the AI ​​assistant generates a matter notification, the above step 201a can be specifically implemented through the following step 201a1, and the assistant dialogue method provided in the embodiment of the present application can also include the following steps 301 and 302.

[0199] Step 201a1: When the first reply message is displayed on the conversation interface of the AI ​​assistant, the electronic device inputs the user message and the first reply message into the result detection model. Through the result detection model, when the dialogue type is a text generation type, the electronic device detects the missing matter elements of the first reply result compared with the reference elements, and detects the missing matter detail information of the first reply result compared with the reference matter detail information, and outputs quality information.

[0200] Exemplarily, the matter elements of the above-mentioned first reply result may include at least one of the following: meeting title, meeting content, meeting time, meeting participation method, meeting participants, meeting requirements, meeting leave process, etc.

[0201] Exemplarily, the matter detail information of the above-mentioned first reply result may include at least one of the following: detailed information on the meeting content, detailed information on the meeting time, detailed information on the meeting participation method, detailed information on the meeting participants, detailed information on the meeting requirements, and detailed information on the process of the meeting request process.

[0202] For example, combining Figure 2BThe mobile phone can input the meeting notice 12 and the text "Help me write a meeting notice" into the result detection model. The dialogue type of the text "Help me write a meeting notice" is a text generation type. Therefore, the mobile phone can use the result detection model to detect that the event elements of the meeting notice 12 are missing from the reference event elements, such as the meeting title, meeting time, and meeting leave application process. In addition, the mobile phone can use the result detection model to detect that the meeting content details of the meeting notice 12 are missing from the reference event details. For example, the missing content details include: the specific content of reviewing the work achievements and shortcomings of this month, and the specific content of discussing the work plan for next month. The mobile phone can also use the result detection model to detect that the meeting participation method details of the meeting notice 12 are missing from the reference event details. For example, the missing participation details include: the specific address of the meeting and the specific method of participating in the meeting. The mobile phone can also use the result detection model to detect that the meeting requirements details of the meeting notice 12 are missing from the reference event details. For example, the missing requirements details include: the specific materials to be brought to the meeting and the specific requirements for attending the meeting. Therefore, the result detection model can output quality information, which includes: status information "Is it abnormal: no", quality content information "Is it high-quality: no", content information "Insufficient description: missing details, unclear title, lack of specific date, single participation method, unclear meeting process, insufficient operability", and detection result information "Judgment result: qualified".

[0203] It can be understood that since the meeting notice 12 is a meeting notice, it can be understood that the meeting notice 12 is correct, so the status information can be "Is it abnormal: No". Since the meeting notice 12 is incomplete, for example, the meeting notice 12 lacks the meeting title, meeting time and meeting leave process, and the meeting notice 12 has missing content detail information, missing participation detail information and missing content detail information, that is, the meeting notice 12 is incomplete and the content of the meeting notice 12 is not detailed, so the quality content information can be "Is it high quality: No". Since the meeting notice 12 lacks the meeting title, meeting time and meeting leave process, and the meeting notice 12 has missing content detail information, missing participation detail information and missing content detail information, so the content information can be "Insufficient description: missing details, unclear title, missing specific date, unclear meeting content, single participation method, unclear meeting requirements, and insufficient operability". Since the meeting notice 12 is correct, the test result information can be "Judgment result: qualified".

[0204] Step 301: When the number of missing item elements is greater than or equal to the item quantity threshold, the electronic device determines that the completeness rate of the first reply result is less than or equal to the reference completeness rate.

[0205] Exemplarily, the above-mentioned threshold value of the number of items may be a positive integer, and the value range of the threshold value of the number of items may be [1, 100].

[0206] In some embodiments of the present application, if the number of missing matter elements is greater than or equal to the matter quantity threshold, it can be considered that the number of missing matter elements in the first reply result is large, and therefore, it can be considered that the completeness rate of the first reply result is less than or equal to the reference completeness rate.

[0207] For example, assuming the threshold for the number of items is 1 and the reference completeness is 90%, combined with Figure 2B The missing elements in the meeting notice 12 include the meeting title, meeting time, and meeting leave process. That is to say, the number of missing elements is 3, which is greater than the threshold of 1 item number. Therefore, the completeness rate of the first reply result is 75%. It can be considered that the completeness rate of the first reply result of 75% is lower than the reference completeness rate of 90%.

[0208] Step 302: When the proportion of the missing item detail information in the reference item detail information is greater than or equal to the proportion threshold, the electronic device determines that the content detail of the first reply result is less than or equal to the reference content detail.

[0209] Exemplarily, the above-mentioned percentage threshold may be a positive number, and the value range of the percentage threshold may be [1%, 100%].

[0210] In some embodiments of the present application, if the proportion of missing matter detail information in the reference matter detail information is greater than or equal to a proportion threshold, it can be considered that there is a large amount of missing matter detail information in the first reply result. Therefore, it can be considered that the content detail of the first reply result is less than or equal to the reference content detail.

[0211] For example, assuming the threshold is 80%, combined with Figure 2B , the item details information of meeting notice 12 is missing from the reference item details information, including: the specific content of reviewing the work results and shortcomings of this month, the specific content of discussing the work plan for next month, the specific address for participating in the meeting, the specific method for participating in the meeting, the specific materials to be brought to the meeting, and the specific requirements for attending the meeting. The proportion of this missing item details information in the reference item details information is greater than or equal to the proportion threshold of 80%. Therefore, it can be considered that the content detail of the first reply result can be 20%, which is less than the reference content detail of 60%.

[0212] It can be seen that, since the dialogue information instructs the AI ​​assistant to generate an event notification, the electronic device can use the result detection model to detect the event elements of the first reply result that are missing compared to the reference elements, that is, detect the missing event elements related to the completeness rate of the event notification, and detect the event detail information of the first reply result that is missing compared to the reference event detail information, that is, detect the missing event detail information related to the content detail of the event notification. Therefore, the electronic device can accurately determine the relationship between the completeness rate of the first reply result and the reference completeness rate based on the relationship between the missing event elements and the event quantity threshold, and accurately determine the content detail of the first reply result based on the proportion of the missing event detail information in the reference event detail information and the proportion threshold.

[0213] Step 202a: When the completeness rate of the first reply result is less than or equal to the reference completeness rate and the content detail of the first reply result is less than or equal to the reference content detail, the electronic device displays a reply optimization control.

[0214] In some embodiments of the present application, if the completeness rate of the first reply result is less than or equal to the reference completeness rate, and the content detail of the first reply result is less than or equal to the reference content detail, then the first reply result can be considered incomplete and not detailed. Therefore, the electronic device can display a reply optimization control to facilitate the user to control the electronic device to optimize the first reply result.

[0215] In an embodiment of the present application, the completeness rate of the second reply result is greater than the reference completeness rate, and the content detail of the second reply result is greater than the reference content detail.

[0216] It can be understood that the number of missing elements in the second reply result's matter elements compared to the reference elements is less than the matter quantity threshold, and the proportion of missing detail information in the second reply result's matter detail information compared to the reference matter information in the reference matter information is less than the proportion threshold.

[0217] For example, combining Figure 2B ,like Figure 9 As shown, after the user makes a first input to the reply optimization control 13, the mobile phone can display a second reply message in the conversation interface 10, such as an optimized meeting notification 19, which includes Figure 2B The meeting notice 12 in the optimization includes the following elements: the meeting title "I. Meeting Content" and "II. Meeting Requirements", the meeting time "March 28, 2025 (Friday) 10:00-12:00 AM" and the meeting leave application process "Leave application process: If you are unable to attend the meeting, please submit a leave application through the OA system before March 25 for approval by the department head". In addition, the optimized meeting notice 19 includes Figure 2B The meeting notice 12 in the article contains missing details, such as the missing content details "Summary of key work this month (including achievements and shortcomings): Deployment of work plan for next month: Discussion of cross-departmental collaboration issues (focus on: optimization of customer feedback response process)", the missing participation details "Offline (third floor conference room) + online (meeting number: 123, 10:00-12:00)" and the missing requirement details "Material preparation: Each department must submit an electronic summary report to the administrative department email address (123@xyz.com) before 17:00 on March 26; the project leader should prepare a 3-minute PPT presentation, focusing on the data results (see the attached template for details)". It can be understood that the completeness rate of the optimized meeting notice 19 is 100%, and the completeness rate of the meeting notice 12 is 75%, that is, the completeness rate of the optimized meeting notice 19 is higher than the completeness rate of the meeting notice 12, and the content detail of the optimized meeting notice 19 is 100%, and the content detail of the meeting notice 12 is 60%, that is, the content detail of the optimized meeting notice 19 is higher than the content detail of the meeting notice 12. Therefore, the reply quality of the optimized meeting notice 19 is higher than the reply quality of the meeting notice 12.

[0218] It can be seen that since users may be more concerned about the completeness and content details of the text, the electronic device can detect the completeness and content details of the first reply result when the conversation type is a text generation type, and when the completeness of the first reply result is less than or equal to the reference completeness rate, and the content details of the first reply result is less than or equal to the reference content details, that is, when the first reply result is incomplete and not detailed, that is, when the user may need to optimize the first reply result, the reply optimization control will be displayed instead of directly displaying the reply optimization control. Therefore, when the reply quality of the first reply result is high, that is, when the user may not need to optimize the first reply result, the electronic device may not display the reply optimization control, instead of displaying the reply optimization control, thereby making the conversation interface more concise.

[0219] In some examples, the reply quality of the first reply result includes a picture quality parameter of the reply result. Figure 7 ,like Figure 10 As shown, the above step 201 can be specifically implemented through the following step 201b, and the above step 202 can be specifically implemented through the following step 202b.

[0220] Step 201b: When the first reply message is displayed on the conversation interface of the AI ​​assistant, the electronic device inputs the user message and the first reply message into the result detection model. Through the result detection model, when the conversation type is the image generation type, the image quality parameters of the first reply result are detected and the quality information is output.

[0221] Optionally, the electronic device may first detect the image object in the first reply result using a result detection model to determine object features of the image object in the first reply result, and then determine the image quality parameter of the first reply result based on the object features of the image object in the first reply result. The object features may include at least one of the following: body features, facial features, color features, etc.

[0222] Among them, the body features can be used to indicate the characteristics of the physical form and physiological structure of the object, and the body features can include at least one of the following: body proportions, bone proportions, muscle distribution, body surface markings, etc.

[0223] The facial features are used to indicate the distribution of facial features and skin texture of the subject. The facial features may include at least one of the following: cheekbone height, nose bridge shape, eye distance ratio, lip shape, iris, facial blood vessel distribution, etc.

[0224] The color feature is used to indicate a feature of the object's color, and the color feature may include at least one of the following: hue, brightness, saturation, skin color, etc.

[0225] Optionally, when the dialogue information instructs the AI ​​assistant to generate a picture, the above step 201b can be specifically implemented through the following step 201b1, and the assistant dialogue method provided in the embodiment of the present application can also include the following step 303.

[0226] Step 201b1: When the first reply message is displayed on the conversation interface of the AI ​​assistant, the electronic device inputs the user message and the first reply message into the result detection model. Through the result detection model, when the conversation type is the image generation type, the feature difference between the body features of the image object of the first reply result and the reference body features is detected, and quality information is output.

[0227] For example, the body feature may include body proportions, and the reference body feature may include reference body proportions. The body proportions may include: the ratio of arm length to body length, the ratio of leg length to body length, and the reference body proportions may include: a reference hand-to-body ratio of arm length to body length, and a reference leg-to-body ratio of leg length to body length.

[0228] Among them, the reference hand-to-body ratio can be a positive number, and the value range of the reference hand-to-body ratio can be [0.4, 0.6]. The reference leg-to-body ratio can be a positive number, and the value range of the reference leg-to-body ratio can be [0.5, 0.7].

[0229] For example, combining Figure 3B, the mobile phone can input picture 14 and the text "Help me draw a girl dancing ballet" into the result detection model, and the dialogue type of the text "Help me draw a girl dancing ballet" is the picture generation type, so that the mobile phone can detect the body features of the person in picture 14 through the result detection model, such as body proportion, which includes the ratio of the length of the arm to the length of the whole body, which is 2, and the ratio of the length of the leg to the length of the whole body, which is 2, and detect the feature difference 1 between the ratio of the length of the arm to the length of the whole body 2 and the reference hand-to-body ratio 0.5, for example, the feature difference 1 is the difference between the ratio 2 and the reference hand-to-body ratio 0.5 of 1.5, and detect the feature difference 2 between the ratio of the length of the leg to the length of the whole body 2 and the reference leg-to-body ratio 0.618, for example, the feature difference 2 is the difference between the ratio 2 and the reference hand-to-body ratio 0.618 of 1.382, for example. Figure 3B Both feature difference 1 and feature difference 2 are large. Therefore, the result detection model can output quality information, which includes: status information "Is it abnormal: yes", quality content information "Is it high-quality: no", content information "Insufficient description: abnormal figure proportions, arms too long, legs too short", and detection result information "Judgment result: low quality".

[0230] It can be understood that since picture 14 is a person picture, and the characteristic difference between the body features of the person in the person picture and the reference body features is large, it can be understood that the person picture is incorrect, so the status information can be "Is it abnormal: yes". Since the characteristic difference between the body features of the person in the person picture and the reference body features is large, the quality content information can be "Is it high quality: no". Since the characteristic difference 1 between the ratio of the length of the arm to the length of the whole body and the reference hand-to-body ratio is large, and the characteristic difference 2 between the ratio of the length of the leg to the length of the whole body and the reference leg-to-body ratio is large, for example Figure 3B The character's arms are too long and their legs are too short. Therefore, the content information may be "Inadequate description: Character proportions are abnormal, arms are too long, and legs are too short." Since the character image is incorrect, the detection result information may be "Judgment result: low quality."

[0231] Step 303: When the feature difference is greater than or equal to the feature difference threshold, the electronic device determines whether the parameter value of the picture quality parameter of the first reply result is less than or equal to the reference picture quality parameter threshold.

[0232] Exemplarily, the feature difference threshold may be a positive number, and the value range of the feature difference threshold may be [0.1, 1].

[0233] In some embodiments of the present application, if the feature difference is greater than or equal to the feature difference threshold, it can be considered that the difference between the body features of the person in the first reply result and the reference body features is large, that is, the body features of the person in the first reply result are abnormal. Therefore, it can be considered that the parameter value of the picture quality parameter of the first reply result is less than or equal to the reference picture quality parameter threshold.

[0234] For example, assuming the feature difference threshold is 0.5 and the reference image quality parameter threshold is 60, combined with Figure 3B , feature difference 1 is 1.5, feature difference 2 is 1.382, both feature difference 1 and feature difference 2 are greater than the feature difference threshold of 0.5. Therefore, it can be considered that the parameter value of the picture quality parameter of the first reply result can be 0, that is, the parameter value of the picture quality parameter of the first reply result is less than the reference picture quality parameter threshold of 60.

[0235] It can be seen that since the dialogue information instructs the AI ​​assistant to generate a picture, the electronic device can detect the feature difference between the body features of the picture object of the first reply result and the reference body features through the result detection model, that is, the body features of the picture object related to the picture are detected. Therefore, the electronic device can accurately determine the picture quality parameters of the first reply result based on the size relationship between the feature difference between the body feature and the reference body feature and the feature difference threshold.

[0236] Step 202b: When the parameter value of the picture quality parameter of the first reply result is less than or equal to the reference picture quality parameter threshold, the electronic device displays a reply optimization control.

[0237] In an embodiment of the present application, if the parameter value of the picture quality parameter of the first reply result is less than or equal to the reference picture quality parameter threshold, it can be considered that the picture quality of the first reply result is low. Therefore, the electronic device can display a reply optimization control to facilitate the user to control the electronic device to optimize the first reply result.

[0238] In this embodiment of the present application, the parameter value of the picture quality parameter of the second reply result is greater than the reference picture quality parameter threshold.

[0239] It can be understood that the feature difference between the body feature of the picture object of the second reply result and the reference body feature is less than the feature difference threshold.

[0240] For example, combining Figure 3B ,like Figure 11As shown, after the user makes a first input to the reply optimization control 13, the mobile phone can display a second reply message in the conversation interface 10. For example, in picture 20, the feature difference 3 between the ratio of the arm length to the total length of the image object 21 in picture 20, 0.5, and the reference hand-to-body ratio of 0.5 is small. For example, feature difference 3 is the difference between this ratio of 0.5 and the reference thin ratio of 0.5, which is 0. The feature difference 4 between the ratio of the leg length to the total length of 0.6 and the reference leg-to-body ratio of 0.618 is small. Feature difference 4 is the difference between this ratio of 0.6 and the reference thin ratio of 0.618, which is 0.018. Both feature differences 3 and 4 are less than the feature difference threshold of 0.5. It can be understood that the parameter value of the image quality parameter of picture 20 is 100, while the parameter value of the image quality parameter of picture 13 is 0. In other words, the parameter value of the image quality parameter of picture 20 is greater than the parameter value of the image quality parameter of picture 13, i.e., the reply quality of picture 20 is higher than that of picture 13.

[0241] It can be seen that since users may be more concerned about the physical features of the picture object in the picture, the electronic device can detect the feature difference between the body features of the picture object of the first reply result and the reference body features when the conversation type is the picture generation type, and when the feature difference is greater than or equal to the feature difference threshold, that is, when the body features of the picture object of the first reply result are abnormal, that is, when the user may need to optimize the first reply result, the reply optimization control will be displayed instead of directly displaying the reply optimization control. Therefore, when the reply quality of the first reply result is high, that is, when the user may not need to optimize the first reply result, the electronic device may not display the reply optimization control, instead of displaying the reply optimization control, thereby making the conversation interface more concise.

[0242] In some examples, the response quality of the first response result includes the accuracy rate of the response result and the completeness rate of the response result. Figure 7 ,like Figure 12 As shown, the above step 201 can be specifically implemented through the following steps 201c and 201d, and the above step 202 can be specifically implemented through the following step 202c.

[0243] Step 201c: When the first reply message is displayed on the conversation interface of the AI ​​assistant, the electronic device inputs the user message and the first reply message into the result detection model. Through the result detection model, when the dialogue type is a skill call type, skill detection auxiliary information is obtained from the electronic device based on the dialogue information.

[0244] In the embodiment of the present application, the skill detection auxiliary information includes at least one of the following: context information of the dialogue information, and application data information matching the dialogue topic of the dialogue information.

[0245] Optionally, the above-mentioned context information may be information in the conversation interface, that is, the context information of the conversation information may be understood as the context information of the conversation information in the conversation interface.

[0246] Optionally, the application data information may be information in an application that matches the conversation topic of the conversation information.

[0247] For example, assuming that the dialogue information instructs the AI ​​assistant to generate clothing suggestions, the dialogue topic of the dialogue information may be a weather topic, and the application data information may be information in a weather application that matches the weather topic.

[0248] For another example, assuming that the dialogue information instructs the AI ​​assistant to query the location, the dialogue topic of the dialogue information may be the location topic, and the application data information may be information in a map application that matches the location topic.

[0249] For another example, assuming that the conversation information instructs the AI ​​assistant to query photos in the electronic device, the conversation topic of the conversation information can be a picture topic, and the application data information can be information in a picture application that matches the picture topic.

[0250] Step 201d: The electronic device detects the accuracy and completeness of the first reply result based on the dialogue information and the skill detection auxiliary information.

[0251] Optionally, the electronic device can first perform semantic detection on the semantics of the first reply result through a result detection model to determine the suggested content of the first reply result and the suggested elements of the first reply result, and then determine the accuracy of the first reply result based on the suggested content of the first reply result and the skill detection auxiliary information, and determine the completeness rate of the first reply result based on the suggested elements of the first reply result.

[0252] Optionally, when the dialogue indication information instructs the AI ​​assistant to generate clothing suggestions and the application data information is weather information, the above-mentioned step 201d can be specifically implemented through the following step 201d1, and the assistant dialogue method provided in the embodiment of the present application can also include the following steps 304 and 305.

[0253] Step 201d1: The electronic device detects the difference between the suggestion content of the first reply result and the reference suggestion content based on the dialogue information and the skill detection auxiliary information, and detects the suggestion elements missing from the reference suggestion elements.

[0254] In an embodiment of the present application, the above-mentioned reference suggestion content is determined by weather information.

[0255] Exemplarily, the suggestion elements may include suggested matters, and the reference suggestion elements may be determined by dialogue information.

[0256] Exemplarily, the electronic device can use a result detection model to generate reference suggestion content based on weather information and generate reference suggestion elements based on conversation information, then detect the difference in suggestion between the suggestion content of the first reply result and the reference suggestion content, and detect the suggestion elements that are missing from the suggestion elements of the first reply result compared to the reference suggestion elements.

[0257] For example, combining Figure 4B The mobile phone can input weather and clothing suggestions 15 and the text "What is suitable to wear for tomorrow's weather in City A" into the result detection model. The dialogue type of the text "What is suitable to wear for tomorrow's weather in City A" is a skill call type, so that the mobile phone can use the result detection model to detect the difference between the recommended content of weather and clothing suggestions 15 and the reference recommended content. The reference recommended content can be "Tomorrow's daytime temperature in City A: 15℃~23℃; nighttime temperature: 5℃~10℃; it is recommended to wear a thick down jacket and boots with good thermal insulation. If you are sensitive to cold, you can wear a scarf and a hat." Figure 4B The difference between the suggestion content of weather and clothing suggestion 15 "Tomorrow's daytime temperature in City A: 13℃~22℃ and nighttime temperature: 7~11℃" and the reference suggestion content is 2, that is, the difference is large; and the mobile phone can detect the suggestion elements missing from the suggestion elements of weather and clothing suggestion 15 compared with the reference suggestion elements, such as Figure 4B The text "What should I wear in tomorrow's weather in City A" instructs the AI ​​assistant to generate clothing suggestions. The reference suggestion elements may include the weather in City A "Tomorrow's weather in City A: 15℃~23℃ during the day and 5℃~10℃ at night" and the clothing suggestion "It is recommended to wear a thick down jacket and warm boots. If you are sensitive to cold, you can wear a scarf and a hat." Figure 4B Weather and Clothing Suggestions 15 in the example only includes the weather in City A: "Tomorrow in City A: Daytime: 13°C-22°C, Nighttime: 7°C-11°C." Therefore, the missing recommendation element in Weather and Clothing Suggestions 15 compared to the reference recommendation element is clothing advice. Therefore, the result detection model can output quality information, including: status information ("Abnormal: Yes"), quality content information ("High Quality: No"), content information ("Insufficient Description: No Clothing Recommendation Provided"), and detection result information ("Judgment Result: Low Quality").

[0258] Step 304: When the suggestion difference is greater than or equal to the suggestion difference threshold, the electronic device determines that the accuracy of the first reply result is less than or equal to the reference accuracy.

[0259] Exemplarily, the suggested difference threshold may be a positive number, and the value range of the suggested difference threshold may be [0.1, 1].

[0260] In some embodiments of the present application, if the suggested difference is greater than or equal to the suggested difference threshold, it can be considered that the first reply result may be incorrect. Therefore, it can be considered that the accuracy of the first reply result is less than or equal to the reference accuracy.

[0261] For example, assuming the recommended difference threshold is 0.5 and the reference accuracy is 80%, combined with Figure 4B , the recommended difference is 2, which is greater than the recommended difference threshold of 0.5. Therefore, it can be considered that the accuracy of the first reply result can be 0%, that is, the accuracy of the first reply result is less than the reference accuracy of 80%.

[0262] Step 305: When the number of missing suggestion elements is greater than or equal to the missing suggestion number threshold, the electronic device determines that the completeness rate of the first reply result is less than or equal to the reference completeness rate.

[0263] Exemplarily, the missing suggestion number threshold may be a positive integer, and the value range of the missing suggestion number threshold may be [1, 100].

[0264] In the embodiment of the present application, if the suggested difference is greater than or equal to the suggested difference threshold, it can be considered that the first reply result may be incomplete. Therefore, it can be considered that the completeness rate of the first reply result is less than or equal to the reference completeness rate.

[0265] For example, assuming the threshold for the number of missing suggestions is 1 and the reference completeness is 90%, combined with Figure 4B The missing suggestion elements in weather and clothing suggestions 15 include clothing suggestions, that is, the number of missing suggestion elements is 1, and the number of missing suggestion elements is equal to the missing suggestion number threshold. Therefore, the completeness rate of the first reply result is 50%. It can be considered that the completeness rate of the first reply result 50% is less than the reference completeness rate of 90%.

[0266] It can be seen that, since the dialogue information instructs the AI ​​assistant to generate clothing suggestions, the electronic device can use the result detection model to detect the suggestion difference between the suggestion content of the first reply result and the reference suggestion content, that is, detect the suggestion difference between the suggestion content related to clothing suggestions and the reference suggestion content, and detect the suggestion elements of the first reply result that are missing compared with the reference suggestion elements, that is, detect the suggestion elements related to clothing suggestions that are missing compared with the reference suggestion elements. Therefore, the electronic device can accurately determine the correctness of the first reply result based on the relationship between the suggestion difference and the suggestion difference threshold, and accurately determine the completeness rate of the first reply result based on the relationship between the number of missing suggestion elements and the missing suggestion number threshold.

[0267] Step 202c: When the accuracy rate of the first reply result is less than or equal to the reference accuracy rate, and the completeness rate of the first reply result is less than or equal to the reference completeness rate, the electronic device displays a reply optimization control.

[0268] In an embodiment of the present application, if the accuracy rate of the first reply result is less than or equal to the reference accuracy rate, and the completeness rate of the first reply result is less than or equal to the reference completeness rate, it can be considered that the accuracy rate and completeness rate of the first reply result are both low. Therefore, the electronic device can display a reply optimization control to facilitate the user to control the electronic device to optimize the first reply result.

[0269] In an embodiment of the present application, the accuracy rate of the above-mentioned second reply result is greater than the reference accuracy rate, and the completeness rate of the second reply result is greater than the reference completeness rate.

[0270] For example, combining Figure 4B ,like Figure 13As shown, after the user makes a first input to the reply optimization control 13, the mobile phone can display a second reply message in the conversation interface 10, such as the optimized weather and clothing suggestion 22 "Tomorrow's daytime temperature in City A: 15℃~23℃, and nighttime temperature: 5℃~10℃; it is recommended to wear a thick down jacket and boots with good warmth retention. If you are sensitive to the cold, you can wear a scarf and a hat." The difference between the suggestion content of the optimized weather and clothing suggestion 22 and the reference suggestion content is small, and the optimized weather and clothing suggestion 22 includes the weather of City A "Tomorrow's daytime temperature in City A: 15℃~23℃, and nighttime temperature: 5℃~10℃" and the clothing suggestion "It is recommended to wear a thick down jacket and boots with good warmth retention. If you are sensitive to the cold, you can wear a scarf and a hat." It can be understood that the accuracy rate of the optimized weather and clothing suggestions 22 is 100%, and the accuracy rate of the weather and clothing suggestions 15 is 0%, that is, the accuracy rate of the optimized weather and clothing suggestions 22 is higher than the accuracy rate of the weather and clothing suggestions 15; the completeness rate of the optimized weather and clothing suggestions 22 is 100%, and the completeness rate of the weather and clothing suggestions 15 is 50%, that is, the completeness rate of the optimized weather and clothing suggestions 22 is higher than the completeness rate of the weather and clothing suggestions 15. Therefore, the response quality of the optimized weather and clothing suggestions 22 is higher than the response quality of the weather and clothing suggestions 15.

[0271] It can be seen that since users may be more concerned about the accuracy and completeness of the reply results of skill-calling type conversation information, the electronic device can detect the accuracy and completeness when the conversation type is a skill-calling type, and when the accuracy of the first reply result is less than or equal to the reference accuracy rate, and the completeness of the first reply result is less than or equal to the reference completeness rate, that is, when the reply quality of the first reply result is poor, that is, when the user may need to optimize the first reply result, the reply optimization control will be displayed instead of directly displaying the reply optimization control. Therefore, when the reply quality of the first reply result is high, that is, when the user may not need to optimize the first reply result, the electronic device may not display the reply optimization control, instead of displaying the reply optimization control, thereby making the conversation interface more concise.

[0272] In some examples, the response quality of the first response result includes the accuracy of the response result. Figure 7 ,like Figure 14 As shown, the above step 201 can be specifically implemented through the following steps 201e and 201f, and the above step 202 can be specifically implemented through the following step 202d.

[0273] Step 201e: When the first reply message is displayed on the conversation interface of the AI ​​assistant, the electronic device inputs the user message and the first reply message into the result detection model. Through the result detection model, when the conversation type is an online query type, the electronic device obtains Internet detection auxiliary information associated with the conversation information from the Internet.

[0274] Optionally, the Internet detection auxiliary information may be reference content related to the conversation information that is searched on the Internet.

[0275] Optionally, the electronic device may search on the Internet using the result detection model, thereby obtaining Internet detection auxiliary information.

[0276] Step 201f: The electronic device detects the accuracy of the first reply result based on the auxiliary information detected by the Internet.

[0277] Optionally, the electronic device can perform semantic detection on the semantics of the first reply result through a result detection model, determine the reply content of the first reply result, and then determine the accuracy of the first reply result based on the reply content of the first reply result and Internet detection auxiliary information.

[0278] Optionally, when the dialogue indication information instructs the AI ​​assistant to query the specified object information and the Internet detection auxiliary information is the reference content, the above step 201f can be specifically implemented through the following step 201f1, and the assistant dialogue method provided in the embodiment of the present application can also include the following step 306.

[0279] Step 201f1: The electronic device detects the content difference between the reply content of the first reply result and the reference content based on the Internet detection auxiliary information.

[0280] Exemplarily, the above-mentioned specified object information may be question information that the user needs the AI ​​assistant to query.

[0281] For example, combining Figure 5B, the mobile phone can input the virus and vaccination recommendation 16 and the text "What is the main virus of the current cold, and which vaccine is better" into the result detection model. The dialogue type of the text "What is the main virus of the current cold, and which vaccine is better" is an online query type, so that the mobile phone can detect the content difference between the reply content of the virus and vaccination recommendation 16 and the reference content through the result detection model. The reference content can be "The current cold-related viruses are mainly influenza virus, respiratory syncytial virus, human rhinovirus, parainfluenza virus, pancreatic virus, etc. Among them, influenza vaccine: can prevent influenza A and B viruses, and it is recommended to be vaccinated from September to October every year; Nissvir monoclonal antibody: can prevent respiratory syncytial virus, suitable for infants aged 0-1 years; parainfluenza virus vaccine: there are trivalent and quadrivalent inactivated vaccines, which are voluntary vaccinations at one's own expense; pancreatic virus vaccine: there are vaccines for type 4 and type 7, which are suitable for children aged 4 years and above and children aged 6 months to 5 years respectively." Figure 5B The content difference between the response content for "Virus and Vaccination Recommendation 16" and the reference content is 9, indicating a significant content difference. Therefore, the result detection model can output quality information, including: status information "Abnormal: Yes," quality content information "High-quality: No," content information "Insufficient description: Response result error, incorrect list of currently prevalent cold-related viruses, incorrect vaccination recommendation," and test result information "Judgment result: Low quality."

[0282] Step 306: When the content difference is greater than or equal to the content difference threshold, the electronic device determines that the accuracy of the first reply result is less than or equal to the reference accuracy.

[0283] Exemplarily, the content difference threshold may be a positive number, and the value range of the feature difference threshold may be [0.1, 1].

[0284] In some embodiments of the present application, if the content difference is greater than or equal to the content difference threshold, it can be considered that the first reply result may be incorrect. Therefore, it can be considered that the accuracy of the first reply result is less than or equal to the reference accuracy.

[0285] For example, assuming the content difference threshold is 0.5 and the reference accuracy is 80%, combined with Figure 5B , the content difference is 9, which is greater than the content difference threshold of 0.5. Therefore, it can be considered that the accuracy of the first reply result can be 0%, that is, the accuracy of the first reply result is less than the reference accuracy of 80%.

[0286] It can be seen that since the dialogue indication information instructs the AI ​​assistant to query the specified object information, the electronic device can detect the content difference between the reply content of the first reply result and the reference content through the result detection model, that is, detect the content difference between the reply content related to the specified object information and the reference content. Therefore, the electronic device can accurately determine the accuracy of the first reply result based on the relationship between the content difference and the content difference threshold.

[0287] Step 202d: When the accuracy of the first reply result is less than or equal to the reference accuracy, the electronic device displays a reply optimization control.

[0288] In some embodiments of the present application, if the accuracy of the first reply result is less than or equal to the reference accuracy, it can be considered that the accuracy of the first reply result is low. Therefore, the electronic device can display a reply optimization control to facilitate the user to control the electronic device to optimize the first reply result.

[0289] In the embodiment of the present application, the accuracy rate of the above-mentioned second reply result is greater than the reference accuracy rate.

[0290] For example, combining Figure 5B ,like Figure 15 As shown, after the user makes a first input to the reply optimization control 13, the mobile phone can display a second reply message in the conversation interface 10, such as the optimized virus and vaccination suggestion 23. The content difference between the reply content of the optimized virus and vaccination suggestion 23 and the above-mentioned reference content is 0. It can be understood that the accuracy of the optimized virus and vaccination suggestion 23 is 100%, and the accuracy of the virus and vaccination suggestion 16 is 0%. That is to say, the accuracy of the optimized virus and vaccination suggestion 23 is higher than the accuracy of the virus and vaccination suggestion 16. Therefore, the reply quality of the optimized virus and vaccination suggestion 23 is higher than the reply quality of the virus and vaccination suggestion 16.

[0291] It can be seen that since users may be more concerned about the accuracy of the reply results of online query type conversation information, the electronic device can detect the accuracy when the conversation type is an online query type, and when the accuracy of the first reply result is less than or equal to the reference accuracy, that is, when the reply quality of the first reply result is poor, that is, when the user may need to optimize the first reply result, the reply optimization control will be displayed instead of directly displaying the reply optimization control. Therefore, when the reply quality of the first reply result is high, that is, when the user may not need to optimize the first reply result, the electronic device may not display the reply optimization control, instead of displaying the reply optimization control, thereby making the conversation interface more concise.

[0292] In some embodiments of the present application, before the above-mentioned step 201a, the assistant dialogue method provided by the embodiment of the present application may also include the following steps 401 and 402.

[0293] Step 401: The electronic device obtains multiple training samples.

[0294] In an embodiment of the present application, the above-mentioned training samples include: sample conversation information, a sample reply message corresponding to the sample conversation information, the conversation type of the sample conversation information, and sample quality information. The sample reply message includes a sample reply result, and the sample quality information is used to indicate the reply quality of the sample reply result.

[0295] In some examples, the information types of the multiple sample conversation information in the above-mentioned multiple training samples may be different, and the information type may include at least one of the following: plain text, text and picture, text and video, text and audio. For example, the multiple sample conversation information include sample conversation information 1, sample conversation information 2, sample conversation information 3 and sample conversation information 4, wherein the sample conversation information 1 is plain text, the sample conversation information 2 is text and picture, the sample conversation information 3 is text and video, and the sample conversation information 4 is text and audio.

[0296] In some examples, the sample reply results included in the multiple sample reply messages in the multiple training samples may have different result types, which may include at least one of the following: plain text, text and picture, text and video, or text and audio.

[0297] In some examples, the conversation type of the sample conversation information may include at least one of the following: a text generation type, an image generation type, a skill call type, and an online query type.

[0298] In some examples, the above-mentioned sample detection information may include at least one of the following: sample status information, used to indicate whether the status of the sample reply result is abnormal; sample quality content information, used to indicate at least one of the following: whether the sample reply result is complete, whether the content of the sample reply result is detailed, and whether the picture quality of the sample reply result is qualified; sample content information, used to indicate the difference between the content of the sample reply result and the reference reply result; sample detection result information, used to indicate whether the reply quality of the sample reply result is qualified.

[0299] For example, Figure 16 Figure 2 shows a schematic diagram of the data structure of a training sample. Figure 16As shown, the number of the above-mentioned multiple training samples can be N, where N is a positive integer. Each training sample in the N training samples can be regarded as a case. For example, the N training samples include training sample 24, training sample 25, and training sample 26. The training sample 25 can be regarded as case 1, and the case 1 includes input and output. For example, the input can be sample dialogue information 1, the sample reply message 1 corresponding to the sample dialogue information 1, and the dialogue type 1 of the sample dialogue information 1, and the output can be sample quality information 1. The training sample 26 can be regarded as case 2, and the case 2 includes input and output. For example, the input can be sample dialogue information 2, the sample reply message 2 corresponding to the sample dialogue information 2, and the dialogue type 2 of the sample dialogue information 2, and the output can be sample quality information 2. The training sample 27 can be regarded as case 3, and the case 3 includes input and output. For example, the input can be sample dialogue information 3, the sample reply message 3 corresponding to the sample dialogue information 3, and the dialogue type 3 of the sample dialogue information 3, and the output can be sample quality information 3. Among them, the dialogue type 1 of the sample dialogue information 1, the dialogue type 2 of the sample dialogue information 2, and the dialogue type 3 of the sample dialogue information 3 can cover the text generation type, the network query type, the image generation type, and the skill call type. The information types of the sample dialogue information 1, the sample dialogue information 2, the sample dialogue information 3, the sample reply message 1, the sample reply message 2, and the sample reply message 3 can cover plain text, text and picture, text and video, and text and audio, so that the sample quality information 1 in the training sample 24 can be used as the label of the training sample 24, the sample quality information 2 in the training sample 25 can be used as the label of the training sample 25, and the sample quality information 3 in the training sample 26 can be used as the label of the training sample 26 for annotation.

[0300] Step 402: The electronic device performs model training on the initial model based on multiple training samples to obtain a result detection model.

[0301] In some examples, the initial model may be a large language model or a Transformer model. Of course, the initial model may also be other models, which is not limited in the present application.

[0302] In some examples, for each training sample in multiple training samples, the electronic device can send multiple training samples to a cloud server, so that the cloud server can first generate a sample prompt word based on a sample conversation information and a prompt word template in a training sample, and the sample prompt word is used to request the initial model to detect the sample response result, and input the sample conversation information and the conversation type of the sample conversation information into the initial model, and the initial model detects the response quality of the sample response result based on the sample conversation information and the conversation type of the sample conversation information, outputs initial quality information, and then determines the loss function based on the initial quality information and the sample quality information, and uses the loss function to train the initial model to obtain a result detection model.

[0303] For example, combining Figure 16 ,like Figure 17 As shown, the cloud server can Figure 16 The training sample 24 is input into the initial model 27 to obtain the initial quality information 28 output by the initial model 27, so that the cloud server can determine the loss function based on the initial quality information 28 and the sample quality information 24, and use the loss function to train the initial model 27.

[0304] It can be understood that the trained result detection model can detect the response result and output accurate quality information after inputting any dialogue information and the response result generated based on the arbitrary dialogue information.

[0305] It can be seen that since the electronic device can obtain multiple training samples, the training samples include: sample conversation information, sample reply messages corresponding to the sample conversation information, conversation types of the sample conversation information, and sample quality information. In other words, each training sample includes a variety of different information corresponding to the sample conversation information. In this way, the result detection model trained based on the multiple training samples can output accurate quality information based on the conversation information and a variety of different information. Therefore, the accuracy of the quality information output by the result detection model can be improved.

[0306] In some embodiments of the present application, a result detection model is trained on a cloud server, and the cloud server can send an inquiry message to the electronic device, so that the electronic device can display the inquiry message and display a first interface based on the user's click input on the inquiry message. The first interface includes a download control, and based on the user's click input on the download control, the result detection model is downloaded and installed, so that the electronic device can use the result detection model to detect the response quality of the first reply result.

[0307] In some embodiments of the present application, after the installation of the result detection model is completed, the electronic device may display a prompt message, which is used to prompt the user whether to allow the reply result to be detected through the result detection model and optimize the reply result, so that the electronic device can determine whether to use the result detection model to detect the reply quality of the first reply result based on the user's input to the prompt information.

[0308] For example, Figure 18A As shown, the mobile phone displays a first interface 24, which includes a download control 25, so that the user can click and input the download control 25 so that the mobile phone can download and use the result detection model. Figure 18B As shown, after the installation of the result detection model is completed, the mobile phone can display a prompt message 26, which is used to prompt the user whether to allow the reply result to be detected by the result detection model and optimize the reply result. The prompt message 26 includes an "Agree" control 27 and a "Reject" control 28, so that the mobile phone can determine to use the result detection model to detect the reply quality of the first reply result based on the user's click input on the "Agree" control 27, or determine not to use the result detection model to detect the reply quality of the first reply result based on the user's click input on the "Reject" control 28.

[0309] In some examples, after the user clicks the “Agree” control 27, the phone may display Figure 6 The application setting interface 17 in the mobile phone can be used to open the reply optimization function control 18 in the application setting interface 17, so that the user can click and input the optimization function control 18 in the mobile phone.

[0310] In some embodiments of the present application, Figure 7 ,like Figure 19 As shown, before "displaying the second reply message" in the above step 102, the assistant dialogue method provided in the embodiment of the present application can also include the following steps 403 and 404, and the above step 102 can be specifically implemented through the following step 102a.

[0311] Step 403: The electronic device generates an optimization prompt word in response to the first input according to the quality information and the dialogue information.

[0312] In some examples, the optimization prompts are used to instruct the model to re-respond to the results based on the conversation information, and to avoid defects indicated by the quality information.

[0313] In some examples, the electronic device may combine the quality information, the conversation information, and the prompt word template to obtain an optimized prompt word.

[0314] Step 404: The electronic device inputs the optimized prompt word into the reply generation model and outputs a second reply message.

[0315] It should be noted that, for the description of generating the second reply message according to the optimized prompt words through the reply generation model, reference can be made to the specific description in the relevant technology, and the embodiments of the present application will not be repeated here.

[0316] Step 102a: The electronic device displays a second reply message.

[0317] It can be seen that since the electronic device can generate optimized prompt words based on quality information and conversation information, the optimized prompt words can be related to the quality information corresponding to the first reply result in addition to the object information. In this way, after the optimized prompt words are input into the reply generation model, the reply generation model can accurately regenerate the reply result based on the conversation information, and optimize the reply result based on the quality information. Therefore, the reply generation model can output a more accurate second reply message.

[0318] Of course, the electronic device can also report log information to other devices so that other devices can train the response generation model based on the log information, which will be explained below with examples.

[0319] In some examples, the second reply message is generated by a reply generation model, and the reply generation model is deployed on the first device. The assistant dialogue method provided in the embodiment of the present application may also include the following step 405.

[0320] Step 405: When the reply quality of the first reply result does not meet the reference condition, the electronic device sends log information to the first device.

[0321] In an embodiment of the present application, the above-mentioned log information includes at least one of the following: conversation information, conversation type of the conversation information, and quality information; the log information is used to train the reply generation model.

[0322] Optionally, the first device may be a cloud server. Of course, the first device may also be other devices, which is not limited in the present embodiment.

[0323] For example, combining Figure 2BAfter the mobile phone displays the meeting notice 12 in the conversation interface 10, the result detection model can output quality information, which includes: status information "Is it abnormal: no", quality content information "Is it high-quality: no", content information "Insufficient description: missing details, unclear title, lack of specific date, single participation method, unclear meeting process, insufficient operability", and detection result information "Judgment result: qualified". That is, the quality information indicates that the reply quality of the first reply result does not meet the reference conditions, so that the mobile phone can send log information to the first device. The log information may include "Help me write a meeting notice", "Text generation type" dialogue type, status information "Is it abnormal: no", quality content information "Is it high-quality: no", content information "Insufficient description: missing details, unclear title, lack of specific date, single participation method, unclear meeting process, insufficient operability", and detection result information "Judgment result: qualified".

[0324] Give another example, combined with Figure 3B After the mobile phone displays the picture 14 in the conversation interface 10, the result detection model can output quality information, which includes: status information "Is it abnormal: yes", quality content information "Is it high-quality: no", content information "Insufficient description: abnormal character proportions, arms too long, and legs too short", and detection result information "Judgment result: low quality". That is, the quality information indicates that the reply quality of the first reply result does not meet the reference conditions, so that the mobile phone can send a log message to the first device. The log information may include: "Help me draw a girl dancing ballet", "Picture generation type" dialogue type, status information "Is it abnormal: yes", quality content information "Is it high-quality: no", content information "Insufficient description: abnormal character proportions, arms too long, and legs too short", and detection result information "Judgment result: low quality".

[0325] It should be noted that the mobile phone may display the log information or not, and this embodiment of the present application does not limit this.

[0326] Optionally, after sending the log information to the first device, the first device can analyze the reasons for unqualified data in the uploaded log through offline analysis, and can assist the R&D team in formulating new strategic recommendations, thereby further optimizing the response generation model so that it can learn to output better quality results.

[0327] For example, Figure 20As shown, the first device can receive multiple log information, such as log information 29, log information 30 and log information 31, so that the first device can perform unqualified data analysis on the log information 29, log information 30 and log information 31, and make recommendations on the generation strategy of the reply results of the dialogue information corresponding to each dialogue type, such as text generation strategy recommendations, network query type strategy generation recommendations, image generation type strategy recommendations, and skill call type strategy recommendations, so that the reply generation model can be optimized and trained according to each strategy recommendation.

[0328] Optionally, after the reply generation model is optimized and trained, when the conversation information 32 is input into the optimized and trained reply generation model 33, the optimized and trained reply generation model 33 can output a better quality reply result 34, the effect of which is shown in the following figure: Figure 21 shown.

[0329] It can be seen that since the electronic device can also send log information including conversation information, conversation type of conversation information, and at least one of quality information to the first device, the first device can accurately know the conversation information, the conversation type of the conversation information, and at least one of the quality information through the log information. Therefore, the first device can optimize and train the reply generation model based on the conversation information, the conversation type of the conversation information, and at least one of the quality information, so that in the subsequent use process, the accuracy of the reply result output by the optimized and trained reply generation model can be improved.

[0330] As you can understand, in this example, when the reply detection model is an on-device model, the effectiveness of user input conversation information can be verified on the device side, enriching the originally single-information log data. This method of preprocessing and analyzing user data effectively aligns with data from traditional log reporting models, greatly improving the integrity of the data for subsequent analysis. Traditional log sampling methods only capture regular user behaviors, requiring manual annotation or the use of large models to determine the effectiveness of these behaviors. This is laborious to analyze and difficult to clearly align causal relationships. With this approach, user usage effectiveness can be verified on the device side before log collection, facilitating subsequent correlation with user profiles, behavior, and other data, providing high-quality input for developing product optimization strategies and improving user experience.

[0331] Figure 22 The figure shows a flow chart of the assistant dialogue method provided in the embodiment of the present application. Figure 22 As shown, the assistant dialogue method provided in the embodiment of the present application may include the following steps 501 and 502.

[0332] Step 501: When a user message is displayed on the conversation interface of the AI ​​assistant, the electronic device obtains quality information.

[0333] In an embodiment of the present application, the above-mentioned user message includes conversation information input by the user, and the above-mentioned quality information is used to indicate the reply quality of the first reply result, which is the reply result included in the first reply message.

[0334] It should be noted that for the description of the dialogue information, the first reply result and the quality information, reference can be made to the specific description in the above embodiment, and the embodiments of the present application will not be repeated here.

[0335] In some embodiments of the present application, when the conversation interface of the AI ​​assistant is displayed, the user can enter a user message in the conversation interface, so that the electronic device can generate a first reply message through the AI ​​assistant and obtain quality information.

[0336] It can be understood that in this example, when the electronic device generates the first reply message, it may not display the first reply message and obtain the quality information of the first reply result in the first reply message.

[0337] In some embodiments of the present application, the electronic device can input the user message and the first reply message into a result detection model, and through the result detection model, detect the reply quality of the first reply result according to the conversation type of the conversation information, and output quality information.

[0338] It should be noted that for the description of the result detection model detecting the reply quality of the first reply result according to the dialogue type of the dialogue information and outputting the quality information, please refer to the specific description in the above embodiment, and the embodiments of this application will not be repeated here.

[0339] Step 502: When the reply quality of the first reply result does not meet the reference condition, the electronic device displays a second reply message.

[0340] In an embodiment of the present application, the above-mentioned second reply message includes a second reply result, which is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result.

[0341] In some embodiments of the present application, the electronic device may first generate an optimized prompt word based on the quality information and the conversation information, input the optimized prompt word into a reply generation model, output a second reply message generated by the reply generation model, and then display the second reply message.

[0342] In some embodiments of the present application, the electronic device may display the second reply message in the above-mentioned conversation interface.

[0343] In some examples, the electronic device may display a prompt message in the conversation interface when the first reply message is displayed in the conversation interface. The prompt message is used to prompt that the electronic device is obtaining a reply result, and when the electronic device determines that the reply quality of the first reply result does not meet the reference conditions, the electronic device may display a second reply message in the conversation interface.

[0344] For example, Figure 23A As shown, the mobile phone displays the conversation interface 29 of the AI ​​assistant, which includes an input box 30, so that the user can enter conversation information in the input box 30, such as the text "help me write a meeting notice". Figure 23B As shown, after the user enters the text "Help me write a meeting notice" in the input box 30, the mobile phone can generate a first reply result based on the text "Help me write a meeting notice" through the AI ​​assistant, and display a prompt message 31 in the conversation interface 29, which is used to prompt the mobile phone that the reply result is being obtained. Figure 23C As shown, when the mobile phone determines that the reply quality of the first reply result does not meet the reference condition, the mobile phone displays a second reply message, such as text 32, in the conversation interface 29.

[0345] Another example is given, such as Figure 24A As shown, the mobile phone displays the conversation interface 29 of the AI ​​assistant, which includes an input box 30, so that the user can enter conversation information in the input box 30, such as the text "help me draw a girl dancing ballet". Figure 24B As shown, after the user enters the text "Help me draw a girl dancing ballet" in the input box 30, the mobile phone can generate a first reply result based on the text "Help me draw a girl dancing ballet" through the AI ​​assistant, and display a prompt message 33 in the conversation interface 29, which is used to prompt the mobile phone that the reply result is being obtained. Figure 24C As shown, when the mobile phone determines that the reply quality of the first reply result does not meet the reference condition, the mobile phone displays a second reply message in the conversation interface 29, such as picture 34.

[0346] Another example is given, such as Figure 25A As shown, the mobile phone displays the conversation interface 29 of the AI ​​assistant, which includes an input box 30, so that the user can enter conversation information in the input box 30, such as the text "What is suitable to wear in tomorrow's weather in City A". Figure 25B As shown, after the user enters the text "What is suitable to wear in tomorrow's weather in City A" in the input box 30, the mobile phone can generate a first reply result based on the text "What is suitable to wear in tomorrow's weather in City A" through the AI ​​assistant, and display a prompt message 35 in the conversation interface 29, which is used to prompt the mobile phone that the reply result is being obtained. Figure 25C As shown, when the mobile phone determines that the reply quality of the first reply result does not meet the reference condition, the mobile phone displays a second reply message, such as text 36, in the conversation interface 29.

[0347] Another example is given, such as Figure 26A As shown, the mobile phone displays the conversation interface 29 of the AI ​​assistant, which includes an input box 30, so that the user can enter conversation information in the input box 30, such as "What is the main virus of the current cold, and which vaccine is better to take?" Figure 26B As shown, after the user enters the text "What is the main virus of the current cold, and what vaccine is better" in the input box 30, the mobile phone can generate a first reply result based on the text "What is the main virus of the current cold, and what vaccine is better" through the AI ​​assistant, and display a prompt message 37 in the conversation interface 29, which is used to prompt the mobile phone that the reply result is being obtained. Figure 26C As shown, when the mobile phone determines that the reply quality of the first reply result does not meet the reference condition, the mobile phone displays a second reply message, such as text 38, in the conversation interface 29.

[0348] It can be understood that in this example, the reply result can be intervened in advance before it is displayed on the screen, so that the best reply result that the system can provide, that is, the second reply result, can be obtained without the user's perception.

[0349] An embodiment of the present application provides an assistant dialogue method, in which, when a user message is displayed on a conversation interface of an AI assistant, an electronic device can obtain quality information, where the user message includes dialogue information input by the user, and the quality information is used to indicate the reply quality of a first reply result, where the first reply result is the reply result included in the first reply message; and when the reply quality of the first reply result does not meet the reference conditions, a second reply message is displayed; the second reply message includes a second reply result, where the second reply result is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result. Since the electronic device can directly obtain the quality information indicating the reply quality of the first reply result when displaying the user message on the conversation interface, and directly optimize the first reply result in the first reply message when the reply quality of the first reply result does not meet the reference conditions, obtain the second reply result with higher reply quality, and display the second reply result with higher reply quality in the conversation interface without the user having to perform multiple operations. Therefore, the user's operation in displaying the reply result with higher reply quality can be simplified and the time consumption can be reduced. Moreover, since the electronic device can obtain the first reply result corresponding to the conversation message, it does not display the first reply result in the conversation interface, and directly displays the second reply message when the reply quality of the first reply result does not meet the reference conditions, that is, displays the optimized second reply result. Therefore, the user can avoid determining the first reply result as the accurate reply result.

[0350] In some embodiments of the present application, the second reply message is generated by a reply generation model, and the reply generation model is deployed on the first device. The assistant dialogue method provided in the embodiment of the present application may also include the following step 503.

[0351] Step 503: When the reply quality of the first reply result does not meet the reference condition, the electronic device sends log information to the first device.

[0352] In an embodiment of the present application, the above-mentioned log information includes at least one of the following: conversation information, conversation type of the conversation information, and quality information; the log information is used to train the reply generation model.

[0353] It should be noted that, for the description of the electronic device sending log information to the first device, reference can be made to the specific description in the above embodiment, and the embodiments of the present application will not be repeated here.

[0354] It can be seen that since the electronic device can also send log information including conversation information, conversation type of conversation information, and at least one of quality information to the first device, the first device can accurately know the conversation information, the conversation type of the conversation information, and at least one of the quality information through the log information. Therefore, the first device can optimize and train the reply generation model based on the conversation information, the conversation type of the conversation information, and at least one of the quality information, so that in the subsequent use process, the accuracy of the reply result output by the optimized and trained reply generation model can be improved.

[0355] The assistant dialogue method provided in the embodiment of the present application can be executed by an assistant dialogue device. In the embodiment of the present application, the assistant dialogue device provided in the embodiment of the present application is described by taking the assistant dialogue device executing the assistant dialogue method as an example.

[0356] Figure 27 The figure shows a schematic diagram of the structure of the assistant dialogue device provided in the embodiment of the present application. Figure 27 As shown, the assistant dialogue device 600 provided in an embodiment of the present application may include: a receiving module 601, which is used to receive a first input to a reply optimization control when a first reply message is displayed on the conversation interface of the AI ​​assistant; the conversation interface includes a user message, the user message includes conversation information input by the user, and the first reply message includes a first reply result for replying to the conversation information. A display module 602 is used to display a second reply message in response to the first input received by the receiving module 601; the second reply message includes a second reply result, the second reply result is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result.

[0357] An embodiment of the present application provides an assistant dialogue device. Since, when a first reply message is displayed on a conversation interface, the assistant dialogue device can directly optimize a first reply result in the first reply message based on a single input by the user, obtain a second reply result with a higher reply quality, and display the second reply result with a higher reply quality in the conversation interface without requiring the user to perform multiple operations. Therefore, the user's operations in displaying a reply result with a higher reply quality can be simplified and time consumption can be reduced.

[0358] In a possible implementation, the above-mentioned reply quality includes at least one of the following: the accuracy of the reply result, the completeness of the reply result, the content detail of the reply result, and the image quality parameter of the reply result.

[0359] In one possible implementation, the assistant dialogue device 600 provided in the embodiment of the present application may further include: a processing module for inputting the user message and the first reply message into a result detection model before the receiving module 601 receives the first input to the reply optimization control, and detecting the reply quality of the first reply result according to the dialogue type of the dialogue information through the result detection model, and outputting quality information; the quality information is used to indicate the reply quality of the first reply result. The above-mentioned display module 602 is also used to display the reply optimization control when the reply quality of the first reply result does not meet the reference conditions.

[0360] In one possible implementation, the reply quality of the first reply result includes the completeness rate of the reply result and the content detail of the reply result. The processing module is specifically used to detect the completeness rate and content detail of the first reply result when the conversation type is a text generation type. The display module 602 is specifically used to display the reply optimization control when the completeness rate of the first reply result detected by the processing module is less than or equal to the reference completeness rate, and the content detail of the first reply result detected by the processing module is less than or equal to the reference content detail; wherein the completeness rate of the second reply result is greater than the reference completeness rate, and the content detail of the second reply result is greater than the reference content detail.

[0361] In one possible implementation, when the dialogue information instructs the AI ​​assistant to generate an event notification, the processing module is specifically configured to detect event elements of the first reply result that are missing compared to reference event elements, and to detect event detail information of the first reply result that is missing compared to reference event detail information. The processing module is further configured to, when the number of missing event elements is greater than or equal to an event quantity threshold, determine that the completeness rate of the first reply result is less than or equal to a reference completeness rate; and, when the proportion of missing event detail information in the reference event detail information is greater than or equal to a proportion threshold, determine that the content detail level of the first reply result is less than or equal to the reference content detail level.

[0362] In one possible implementation, the reply quality of the first reply result includes an image quality parameter of the reply result. The processing module is specifically configured to detect the image quality parameter of the first reply result when the conversation type is an image generation type. The display module 602 is specifically configured to display a reply optimization control when the parameter value of the image quality parameter of the first reply result detected by the processing module is less than or equal to a reference image quality parameter threshold; wherein the parameter value of the image quality parameter of the second reply result is greater than the reference image quality parameter threshold.

[0363] In one possible implementation, when the conversation information instructs the AI ​​assistant to generate an image, the processing module is specifically configured to detect a degree of difference between the body features of the image object in the first reply result and the reference body features. The processing module is further configured to, when the degree of difference is greater than or equal to a feature difference threshold, determine that a parameter value of an image quality parameter of the first reply result is less than or equal to a reference image quality parameter threshold.

[0364] In one possible implementation, the reply quality of the first reply result includes the accuracy of the reply result and the completeness of the reply result. The processing module is specifically used to obtain skill detection auxiliary information from the assistant dialogue device according to the dialogue information when the dialogue type is a skill call type; wherein the skill detection auxiliary information includes at least one of the following: the above-mentioned context information of the dialogue information, application data information matching the dialogue topic of the dialogue information; and detect the accuracy and completeness of the first reply result based on the dialogue information and the skill detection auxiliary information. The display module 602 is specifically used to display the reply optimization control when the accuracy of the first reply result detected by the processing module is less than or equal to the reference accuracy rate, and the completeness of the first reply result detected by the processing module is less than or equal to the reference completeness rate; wherein the accuracy of the second reply result is greater than the reference accuracy rate, and the completeness of the second reply result is greater than the reference completeness rate.

[0365] In one possible implementation, when the dialogue indication information instructs the AI ​​assistant to generate clothing suggestions and the application data information is weather information, the above-mentioned processing module is specifically used to detect the difference between the suggestion content of the first reply result and the reference suggestion content, and detect the suggestion elements of the first reply result that are missing compared to the reference suggestion elements, where the reference suggestion content is determined by the weather information. The above-mentioned processing module is also used to determine that the accuracy of the first reply result is less than or equal to the reference accuracy rate when the suggestion difference is greater than or equal to the suggestion difference threshold; and to determine that the completeness rate of the first reply result is less than or equal to the reference completeness rate when the number of missing suggestion elements is greater than or equal to the missing suggestion number threshold.

[0366] In one possible implementation, the response quality of the first response result includes the accuracy of the response result. The processing module is specifically configured to, when the conversation type is an online query type, obtain Internet detection auxiliary information associated with the conversation information from the Internet; and detect the accuracy of the first response result based on the Internet detection auxiliary information. The display module 602 is specifically configured to display a response optimization control when the accuracy of the first response result detected by the processing module is less than or equal to a reference accuracy rate; wherein the accuracy of the second response result is greater than the reference accuracy rate.

[0367] In one possible implementation, when the conversation information instructs the AI ​​assistant to query for information about a specified object, and the internet detection auxiliary information is reference content, the processing module is specifically configured to detect a content difference between the reply content of the first reply result and the reference content. The processing module is further configured to determine, when the content difference is greater than or equal to a content difference threshold, that the accuracy of the first reply result is less than or equal to a reference accuracy rate.

[0368] In one possible implementation, the above-mentioned processing module is also used to obtain multiple training samples, which include: sample conversation information, sample reply messages corresponding to the sample conversation information, conversation types of the sample conversation information, and sample quality information. The sample reply message includes a sample reply result, and the sample quality information is used to indicate the reply quality of the sample reply result; and based on the multiple training samples, the initial model is trained to obtain a result detection model.

[0369] In one possible implementation, the above-mentioned quality information includes at least one of the following: status information, quality content information, content information, and detection result information; wherein, the status information is used to indicate whether the first reply result is correct; the quality content information is used to indicate at least one of the following: whether the first reply result is complete, whether the content of the first reply result is detailed, and whether the picture quality of the first reply result is qualified; the content information is used to indicate the difference between the first reply result and the reference reply result; the detection result information is used to indicate whether the reply quality of the first reply result is qualified.

[0370] In a possible implementation, the processing module is further configured to generate optimized prompt words based on the quality information and the conversation information before the display module 602 displays the second reply message; and input the optimized prompt words into the reply generation model to output the second reply message.

[0371] In one possible implementation, the second reply message is generated by a reply generation model, which is deployed on the first device. The assistant dialogue device 600 provided in the embodiment of the present application may also include: a sending module for sending log information to the first device when the reply quality of the first reply result does not meet the reference condition, the log information including at least one of the following: dialogue information, dialogue type of the dialogue information, and quality information; wherein the log information is used to train the reply generation model.

[0372] The assistant dialogue device in the embodiment of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a car-mounted electronic device, a mobile internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), an ATM or a self-service machine, etc., and the embodiment of the present application does not specifically limit it.

[0373] The assistant dialogue device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0374] The assistant dialogue device provided in the embodiment of the present application can achieve Figures 1 to 21 To avoid repetition, the various processes implemented in the method embodiment are not described here.

[0375] Figure 28 The figure shows a schematic diagram of the structure of the assistant dialogue device provided in the embodiment of the present application. Figure 28 As shown, the assistant dialogue device 700 provided in the embodiment of the present application may include: a processing module 701, which is used to obtain quality information when displaying a user message on the conversation interface of the AI ​​assistant; the user message includes dialogue information input by the user, and the quality information is used to indicate the reply quality of the first reply result, and the first reply result is the reply result included in the first reply message. A display module 702 is used to display a second reply message when the reply quality of the first reply result does not meet the reference conditions; the second reply message includes a second reply result, and the second reply result is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result.

[0376] An embodiment of the present application provides an assistant dialogue device. Since the assistant dialogue device can directly obtain quality information indicating the reply quality of a first reply result when displaying a user message on a conversation interface, and directly optimize the first reply result in the first reply message when the reply quality of the first reply result does not meet the reference conditions, obtain a second reply result with a higher reply quality, and display the second reply result with a higher reply quality in the conversation interface without the user having to perform multiple operations. Therefore, the user's operation in displaying a reply result with a higher reply quality can be simplified and time consumption can be reduced. Moreover, since the assistant dialogue device can obtain the first reply result corresponding to the conversation message, it does not display the first reply result in the conversation interface, and directly displays the second reply message when the reply quality of the first reply result does not meet the reference conditions, that is, displays the optimized second reply result. Therefore, it can avoid the user determining the first reply result as the accurate reply result.

[0377] The assistant dialogue device in the embodiment of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a car-mounted electronic device, a mobile internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), an ATM or a self-service machine, etc., and the embodiment of the present application does not specifically limit it.

[0378] The assistant dialogue device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0379] The assistant dialogue device provided in the embodiment of the present application can achieve Figure 22 To avoid repetition, the various processes implemented in the method embodiment are not described here.

[0380] In some embodiments of the present application, Figure 29 As shown, an embodiment of the present application also provides an electronic device 800, including a processor 801 and a memory 802, wherein the memory 802 stores a program or instruction that can be run on the processor 801, and when the program or instruction is executed by the processor 801, the various process steps of the above-mentioned assistant dialogue method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0381] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0382] Figure 30 A schematic diagram of the hardware structure of an electronic device implementing an embodiment of the present application.

[0383] The electronic device 100 includes but is not limited to components such as a radio frequency unit 101 , a network module 102 , an audio output unit 103 , an input unit 104 , a sensor 105 , a display unit 106 , a user input unit 107 , an interface unit 108 , a memory 109 , and a processor 110 .

[0384] Those skilled in the art will understand that the electronic device 100 may also include a power source (such as a battery) to power each component, and the power source may be logically connected to the processor 110 through a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system. Figure 30 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be repeated here.

[0385] In one example, the user input unit 107 is used to receive a first input to a reply optimization control when a first reply message is displayed on a conversation interface of the AI ​​assistant; the conversation interface includes a user message, the user message includes conversation information input by the user, and the first reply message includes a first reply result for replying to the conversation information.

[0386] The display unit 106 is used to display a second reply message in response to the first input received by the user input unit 107; the second reply message includes a second reply result, and the second reply result is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result.

[0387] An embodiment of the present application provides an electronic device. When a first reply message is displayed on a conversation interface, the electronic device can directly optimize a first reply result in the first reply message based on a single input from the user, obtain a second reply result with a higher reply quality, and display the second reply result with a higher reply quality in the conversation interface without requiring the user to perform multiple operations. Therefore, the user's operations in displaying a reply result with a higher reply quality can be simplified and time consumption can be reduced.

[0388] In a possible implementation, the above-mentioned reply quality includes at least one of the following: the accuracy of the reply result, the completeness of the reply result, the content detail of the reply result, and the image quality parameter of the reply result.

[0389] In one possible implementation, the processor 110 is used to input the user message and the first reply message into a result detection model before the user input unit 107 receives the first input to the reply optimization control, and detect the reply quality of the first reply result based on the conversation type of the conversation information through the result detection model, and output quality information; the quality information is used to indicate the reply quality of the first reply result.

[0390] The display unit 106 is further configured to display a reply optimization control when the reply quality of the first reply result does not meet the reference condition.

[0391] In a possible implementation, the response quality of the first response result includes the completeness rate of the response result and the content detail of the response result.

[0392] The processor 110 is specifically configured to detect the completeness and content detail of the first reply result when the conversation type is a text generation type.

[0393] The above-mentioned display unit 106 is specifically used to display the reply optimization control when the completeness rate of the first reply result detected by the above-mentioned processor 110 is less than or equal to the reference completeness rate, and the content detail of the first reply result detected by the above-mentioned processor 110 is less than or equal to the reference content detail; wherein, the completeness rate of the above-mentioned second reply result is greater than the reference completeness rate, and the content detail of the second reply result is greater than the reference content detail.

[0394] In one possible implementation, when the dialogue information indicates that the AI ​​assistant generates an event notification, the processor 110 is specifically used to detect event elements that are missing from the first reply result compared to the reference elements, and to detect event detail information that is missing from the first reply result compared to the reference event detail information.

[0395] The above-mentioned processor 110 is also used to determine that the completeness rate of the first reply result is less than or equal to the reference completeness rate when the number of missing matter elements is greater than or equal to the matter number threshold; and to determine that the content detail of the first reply result is less than or equal to the reference content detail when the proportion of the missing matter detail information in the reference matter detail information is greater than or equal to the proportion threshold.

[0396] In a possible implementation, the reply quality of the first reply result includes a picture quality parameter of the reply result.

[0397] The processor 110 is specifically configured to detect the image quality parameter of the first reply result when the conversation type is an image generation type.

[0398] The above-mentioned display unit 106 is specifically used to display the reply optimization control when the parameter value of the picture quality parameter of the first reply result detected by the above-mentioned processor 110 is less than or equal to the reference picture quality parameter threshold; wherein, the parameter value of the picture quality parameter of the above-mentioned second reply result is greater than the reference picture quality parameter threshold.

[0399] In one possible implementation, when the conversation information instructs the AI ​​assistant to generate an image, the processor 110 is specifically configured to detect a feature difference between the body features of the image object in the first reply result and the reference body features. The processor 110 is further configured to determine, when the feature difference is greater than or equal to a feature difference threshold, that a parameter value of an image quality parameter of the first reply result is less than or equal to a reference image quality parameter threshold.

[0400] In a possible implementation, the response quality of the first response result includes the accuracy rate of the response result and the completeness rate of the response result.

[0401] The above-mentioned processor 110 is specifically used to obtain skill detection auxiliary information from the electronic device based on the dialogue information when the dialogue type is a skill call type; wherein the above-mentioned skill detection auxiliary information includes at least one of the following: the above information of the dialogue information, and the application data information matching the dialogue topic of the dialogue information; and based on the dialogue information and the skill detection auxiliary information, detect the accuracy and completeness of the first reply result.

[0402] The above-mentioned display unit 106 is specifically used to display the reply optimization control when the accuracy rate of the first reply result detected by the above-mentioned processor 110 is less than or equal to the reference accuracy rate, and the completeness rate of the first reply result detected by the above-mentioned processor 110 is less than or equal to the reference completeness rate; wherein, the accuracy rate of the above-mentioned second reply result is greater than the reference accuracy rate, and the completeness rate of the second reply result is greater than the reference completeness rate.

[0403] In one possible implementation, when the dialogue indication information instructs the AI ​​assistant to generate clothing suggestions and the application data information is weather information, the above-mentioned processor 110 is specifically used to detect the difference in suggestions between the suggestion content of the first reply result and the reference suggestion content, and to detect the suggestion elements of the first reply result that are missing compared to the reference suggestion elements, where the reference suggestion content is determined by the weather information.

[0404] The above-mentioned processor 110 is also used to determine that the accuracy of the first reply result is less than or equal to the reference accuracy rate when the suggestion difference is greater than or equal to the suggestion difference threshold; and to determine that the completeness rate of the first reply result is less than or equal to the reference completeness rate when the number of missing suggestion elements is greater than or equal to the missing suggestion number threshold.

[0405] In a possible implementation, the response quality of the first response result includes the accuracy of the response result.

[0406] The processor 110 is specifically configured to obtain Internet detection auxiliary information associated with the conversation information from the Internet when the conversation type is an online query type; and detect the accuracy of the first reply result based on the Internet detection auxiliary information.

[0407] The display unit 106 is specifically configured to display a reply optimization control when the accuracy of the first reply result detected by the processor 110 is less than or equal to a reference accuracy rate; wherein the accuracy of the second reply result is greater than the reference accuracy rate.

[0408] In one possible implementation, when the dialogue information instructs the AI ​​assistant to query for information about a specified object and the Internet detection auxiliary information is reference content, the processor 110 is specifically used to detect the content difference between the reply content of the first reply result and the reference content.

[0409] The processor 110 is further configured to determine, when the content difference is greater than or equal to a content difference threshold, that the accuracy of the first reply result is less than or equal to a reference accuracy rate.

[0410] In one possible implementation, the processor 110 is further used to obtain a plurality of training samples, the training samples including: sample conversation information, a sample reply message corresponding to the sample conversation information, a conversation type of the sample conversation information, and sample quality information, the sample reply message including a sample reply result, and the sample quality information is used to indicate the reply quality of the sample reply result; and based on the plurality of training samples, the initial model is trained to obtain a result detection model.

[0411] In one possible implementation, the above-mentioned quality information includes at least one of the following: status information, quality content information, content information, and detection result information; wherein, the status information is used to indicate whether the first reply result is correct; the quality content information is used to indicate at least one of the following: whether the first reply result is complete, whether the content of the first reply result is detailed, and whether the picture quality of the first reply result is qualified; the content information is used to indicate the difference between the first reply result and the reference reply result; the detection result information is used to indicate whether the reply quality of the first reply result is qualified.

[0412] In one possible implementation, the processor 110 is further configured to generate optimized prompt words based on the quality information and the conversation information before the display unit 106 displays the second reply message; and input the optimized prompt words into the reply generation model to output the second reply message.

[0413] In a possible implementation, the second reply message is generated by a reply generation model, and the reply generation model is deployed on the first device.

[0414] The radio frequency unit 101 is used to send log information to the first device when the reply quality of the first reply result does not meet the reference conditions. The log information includes at least one of the following: conversation information, conversation type of the conversation information, and quality information; wherein the above log information is used to train the reply generation model.

[0415] In another example, the processor 110 is used to obtain quality information when a user message is displayed on the conversation interface of the AI ​​assistant; the user message includes dialogue information input by the user, and the quality information is used to indicate the reply quality of the first reply result, which is the reply result included in the first reply message.

[0416] Display unit 106 is used to display a second reply message when the reply quality of the first reply result does not meet the reference conditions; the second reply message includes a second reply result, and the second reply result is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result.

[0417] An embodiment of the present application provides an electronic device. Since, when a user message is displayed on a conversation interface, the electronic device can directly obtain quality information indicating the reply quality of a first reply result, and when the reply quality of the first reply result does not meet the reference conditions, the electronic device directly optimizes the first reply result in the first reply message to obtain a second reply result with higher reply quality, and displays the second reply result with higher reply quality in the conversation interface without the user having to perform multiple operations. Therefore, the user's operations in displaying reply results with higher reply quality can be simplified and time consumption can be reduced. Moreover, since the electronic device can obtain the first reply result corresponding to the conversation message, it does not display the first reply result in the conversation interface, and when the reply quality of the first reply result does not meet the reference conditions, it directly displays the second reply message, that is, displays the optimized second reply result. Therefore, it can avoid the user determining the first reply result as the accurate reply result.

[0418] It should be understood that in an embodiment of the present application, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042, and the graphics processor 1041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 106 may include a display panel 1061, and the display panel 1061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 107 includes a touch panel 1071 and at least one of other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include two parts: a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.

[0419] The memory 109 can be used to store software programs and various data. The memory 109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory 109 may include a volatile memory or a non-volatile memory, or the memory 109 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct memory bus random access memory (DRRAM). The memory 109 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0420] Processor 110 may include one or more processing units. Optionally, processor 110 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 110.

[0421] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned assistant dialogue method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0422] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0423] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned assistant dialogue method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0424] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0425] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned assistant dialogue method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0426] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0427] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0428] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. An assistant dialogue method, characterized in that: include: In a case where a conversation interface of the artificial intelligence (AI) assistant displays a first reply message, receiving a first input to a reply optimization control; the conversation interface includes a user message, the user message includes conversation information input by the user, and the first reply message includes a first reply result for replying to the conversation information; In response to the first input, a second reply message is displayed; the second reply message includes a second reply result, the second reply result is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result.

2. The method according to claim 1, characterized in that The reply quality includes at least one of the following: the accuracy of the reply result, the completeness of the reply result, the content detail of the reply result, and the image quality parameter of the reply result.

3. The method according to claim 1, characterized in that Before receiving the first input to the reply optimization control, the method further includes: Inputting the user message and the first reply message into a result detection model, detecting the response quality of the first reply result based on the conversation type of the conversation information through the result detection model, and outputting quality information; the quality information is used to indicate the response quality of the first reply result; When the reply quality of the first reply result does not meet the reference condition, a reply optimization control is displayed.

4. The method according to claim 3, characterized in that The response quality of the first response result includes the completeness rate of the response result and the content detail of the response result; The detecting, according to the conversation type of the conversation information, the reply quality of the first reply result includes: In a case where the dialogue type is a text generation type, detecting the completeness and content detail of the first reply result; The displaying of a reply optimization control when the reply quality of the first reply result does not meet the reference condition includes: When the completeness rate of the first reply result is less than or equal to the reference completeness rate, and the content detail of the first reply result is less than or equal to the reference content detail, displaying a reply optimization control; The completeness rate of the second reply result is greater than the reference completeness rate, and the content detail of the second reply result is greater than the reference content detail.

5. The method according to claim 4, characterized in that When the conversation information indicates that the AI ​​assistant generates a notification of an event, detecting the completeness and content detail of the first reply result includes: Detecting missing item elements of the first reply result compared to the reference item elements, and detecting missing item detail information of the first reply result compared to the reference item detail information; The method further comprises: When the number of the missing item elements is greater than or equal to the item quantity threshold, determining that the completeness rate of the first reply result is less than or equal to the reference completeness rate; When the proportion of the missing matter detail information in the reference matter detail information is greater than or equal to a proportion threshold, it is determined that the content detail of the first reply result is less than or equal to the reference content detail.

6. The method according to claim 3, characterized in that The response quality of the first response result includes a picture quality parameter of the response result; The detecting, according to the conversation type of the conversation information, the reply quality of the first reply result includes: In a case where the conversation type is a picture generation type, detecting a picture quality parameter of the first reply result; The displaying of a reply optimization control when the reply quality of the first reply result does not meet the reference condition includes: In a case where the parameter value of the picture quality parameter of the first reply result is less than or equal to the reference picture quality parameter threshold, displaying a reply optimization control; The parameter value of the picture quality parameter of the second reply result is greater than the reference picture quality parameter threshold.

7. The method according to claim 6, characterized in that When the conversation information instructs the AI ​​assistant to generate an image, detecting the image quality parameter of the first reply result includes: Detecting a degree of difference between a body feature of the image object of the first reply result and a reference body feature; The method further comprises: In a case where the feature difference is greater than or equal to a feature difference threshold, it is determined that a parameter value of the picture quality parameter of the first reply result is less than or equal to the reference picture quality parameter threshold.

8. The method according to claim 3, characterized in that The response quality of the first response result includes the accuracy rate of the response result and the completeness rate of the response result; The detecting, according to the conversation type of the conversation information, the reply quality of the first reply result includes: When the conversation type is a skill invocation type, obtaining skill detection auxiliary information from the electronic device based on the conversation information; wherein the skill detection auxiliary information includes at least one of the following: context information of the conversation information, and application data information matching the conversation topic of the conversation information; detecting the accuracy and completeness of the first reply result based on the dialogue information and the skill detection auxiliary information; The displaying of a reply optimization control when the reply quality of the first reply result does not meet the reference condition includes: When the accuracy rate of the first reply result is less than or equal to the reference accuracy rate, and the completeness rate of the first reply result is less than or equal to the reference completeness rate, displaying a reply optimization control; The accuracy of the second reply result is greater than the reference accuracy rate, and the completeness rate of the second reply result is greater than the reference completeness rate.

9. The method according to claim 8, characterized in that When the dialogue indication information instructs the AI ​​assistant to generate clothing suggestions, and the application data information is weather information, detecting the accuracy and completeness of the first reply result includes: detecting a difference between suggestion content of the first reply result and reference suggestion content, and detecting suggestion elements missing from suggestion elements of the first reply result compared to reference suggestion elements, wherein the reference suggestion content is determined by the weather information; The method further comprises: When the suggestion difference is greater than or equal to a suggestion difference threshold, determining that the accuracy of the first reply result is less than or equal to the reference accuracy; When the number of the missing suggestion elements is greater than or equal to a missing suggestion number threshold, it is determined that the completeness rate of the first reply result is less than or equal to the reference completeness rate.

10. The method according to claim 3, characterized in that The response quality of the first response result includes the accuracy rate of the response result; The detecting, according to the conversation type of the conversation information, the reply quality of the first reply result includes: In a case where the conversation type is an online query type, obtaining Internet detection auxiliary information associated with the conversation information from the Internet; detecting the accuracy of the first reply result according to the Internet detection auxiliary information; The displaying of a reply optimization control when the reply quality of the first reply result does not meet the reference condition includes: When the accuracy of the first reply result is less than or equal to the reference accuracy rate, displaying a reply optimization control; The accuracy of the second reply result is greater than the reference accuracy.

11. The method according to claim 10, characterized in that When the conversation information instructs the AI ​​assistant to query for information about a specified object, and the Internet detection auxiliary information is reference content, detecting the accuracy of the first reply result based on the Internet detection auxiliary information includes: detecting a content difference between the reply content of the first reply result and the reference content; The method further comprises: When the content difference is greater than or equal to a content difference threshold, it is determined that the accuracy of the first reply result is less than or equal to the reference accuracy.

12. The method according to claim 3, characterized in that The method further comprises: Acquire multiple training samples, the training samples including: sample conversation information, a sample reply message corresponding to the sample conversation information, a conversation type of the sample conversation information, and sample quality information, wherein the sample reply message includes a sample reply result, and the sample quality information is used to indicate the reply quality of the sample reply result; Based on the multiple training samples, the initial model is trained to obtain a result detection model.

13. The method according to claim 3, characterized in that The quality information includes at least one of the following: status information, quality content information, content information, and test result information; The status information is used to indicate whether the first reply result is correct; The quality content information is used to indicate at least one of the following: whether the first reply result is complete, whether the content of the first reply result is detailed, and whether the image quality of the first reply result is qualified; The content information is used to indicate the difference between the first reply result and the reference reply result; The detection result information is used to indicate whether the response quality of the first response result is qualified.

14. The method according to claim 3 or 13, characterized in that Before displaying the second reply message, the method further includes: generating optimized prompt words according to the quality information and the conversation information; The optimized prompt word is input into the reply generation model, and the second reply message is output.

15. The method according to claim 3, characterized in that The second reply message is generated by a reply generation model, and the reply generation model is deployed on the first device; The method further comprises: When the response quality of the first response result does not meet the reference condition, sending log information to the first device, the log information including at least one of the following: the conversation information, the conversation type of the conversation information, and the quality information; The log information is used to train a response generation model.

16. An assistant dialogue method, characterized in that: include: Acquiring quality information when a user message is displayed on a conversation interface of the AI ​​assistant; the user message includes conversation information input by the user, and the quality information is used to indicate the reply quality of a first reply result, where the first reply result is the reply result included in the first reply message; When the reply quality of the first reply result does not meet the reference conditions, a second reply message is displayed; the second reply message includes a second reply result, the second reply result is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result.

17. An assistant dialogue device, characterized in that: include: A receiving module, configured to receive a first input to a reply optimization control when a first reply message is displayed on a conversation interface of the AI ​​assistant; the conversation interface includes a user message, the user message includes conversation information input by the user, and the first reply message includes a first reply result for replying to the conversation information; A display module is used to display a second reply message in response to the first input received by the receiving module; the second reply message includes a second reply result, the second reply result is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result.

18. An assistant dialogue device, characterized in that: include: A processing module, configured to obtain quality information when a user message is displayed on a conversation interface of the AI ​​assistant; the user message includes conversation information input by the user, and the quality information is used to indicate the reply quality of a first reply result, where the first reply result is the reply result included in the first reply message; A display module is used to display a second reply message when the reply quality of the first reply result does not meet the reference conditions; the second reply message includes a second reply result, the second reply result is obtained by optimizing the reply content of the first reply result, and the reply quality of the second reply result is higher than the reply quality of the first reply result.

19. An electronic device, characterized in that: It includes a processor and a memory, the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the assistant dialogue method as described in any one of claims 1 to 15 are implemented.

20. An electronic device, characterized in that: It includes a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the assistant dialogue method as described in claim 16 are implemented.

21. A readable storage medium, characterized in that The readable storage medium stores a program or instruction, which, when executed by a processor, implements the steps of the assistant dialogue method as described in any one of claims 1 to 15.

22. A readable storage medium, characterized in that The readable storage medium stores programs or instructions, which, when executed by the processor, implement the steps of the assistant dialogue method as described in claim 16.