Reply generation method and electronic device

By combining instructional information and displayed images with the reply answers to generate labeled images, the problem of low reply answer comprehension efficiency in existing technologies is solved, thereby improving user operation efficiency and interface viewing experience.

WO2026040614A1PCT designated stage Publication Date: 2026-02-26HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/104395
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-22
Filing Date
2025-06-27
Publication Date
2026-02-26

AI Technical Summary

Technical Problem

In existing technologies, text-based responses require users to read step by step to understand the operation steps, which is time-consuming. Furthermore, different users may have different understandings of the operation instructions, leading to low efficiency and repeated questions.

Method used

By combining instructional information with operation demonstration images, labeled images are generated, forming intuitive response answers, simplifying the user's understanding process and reducing the need for searching.

Benefits of technology

It improves the readability and efficiency of replies, reduces the amount of time users have to search through the chat reply interface, and lowers the likelihood of asking the same question repeatedly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025104395_26022026_PF_FP_ABST
    Figure CN2025104395_26022026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a reply generation method and an electronic device, relating to the technical field of human-machine interaction. The method comprises: acquiring query information describing a first question; acquiring an intermediate answer to the first question on the basis of the query information, the intermediate answer comprising at least one piece of operation description information and at least one operation display picture, the operation description information being used for describing an operation step of solving the first question, there being a correspondence between the operation description information and the operation display picture; annotating the operation display picture on the basis of the operation description information to acquire at least one annotated picture; and generating a reply answer to the first question on the basis of the at least one annotated picture.
Need to check novelty before this filing date? Find Prior Art

Description

Reply generation method and electronic device

[0001] The present application claims priority to the Chinese patent application No. 202411163031.9, filed on August 22, 2024, and entitled "Reply generation method and electronic device", the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to the field of human-computer interaction, and in particular to a reply generation method and an electronic device. BACKGROUND

[0003] With the development of artificial intelligence (AI) technology, the dialogue reply technology is becoming mature, and is gradually widely applied in multiple fields.

[0004] For example, after an electronic device receives inquiry information input by a user, the electronic device sends the inquiry information to a dialogue reply system. The dialogue reply system can automatically output a reply answer corresponding to the inquiry information based on dialogue reply technology, and feed back the reply answer to the electronic device, triggering the electronic device to display the reply answer. The dialogue reply system generates the reply answer based on the inquiry information by using AI model technology, and feeds back the reply answer to the electronic device used by the user. In some examples, the inquiry information can be an operation type question, and the reply answer includes operation instruction information.

[0005] For the reply answer including only operation instruction information, the user needs to read the reply answer first to understand each operation step. It can be seen that this form of reply answer will consume a long time for reading. SUMMARY

[0006] To solve the above technical problems, the present application provides a reply generation method and an electronic device. The technical scheme provided by the present application is that in the process of generating a reply answer based on inquiry information of an operation question, operation instruction information corresponding to an operation step is used to label an operation display picture, text and pictures are combined to generate a more intuitive reply answer, thereby improving the readability of the reply answer, helping to provide a picture-text reference answer for the user, shortening the time consumed by the user for viewing the reply answer, thereby helping to help the user efficiently solve the current operation problem, and reducing the possibility of the user repeatedly asking the same question.

[0007] To achieve the above technical purposes, the present application provides the following technical scheme:

[0008] In a first aspect, a reply generation method applied to a server is provided. The method comprises: obtaining inquiry information describing a first question; obtaining an intermediate answer of the first question according to the inquiry information, the intermediate answer comprising at least one operation description information and at least one operation display picture, the operation description information being used to describe an operation step for solving the first question, and the operation description information and the operation display picture having a corresponding relationship; labeling the operation display picture according to the operation description information to obtain at least one labeled picture; and generating a reply answer of the first question according to the at least one labeled picture.

[0009] In this way, the reply answer in the form of a combination of text and pictures is more intuitive and easy to understand, and the reading efficiency of the reply answer is improved. The user can be familiar with the environment (such as a page to be operated, an instrument panel, etc.) of the operation step in advance before performing the step, which helps to eliminate the understanding ambiguity after the user reads the step description information.

[0010] In a possible implementation, the labeling of the operation display picture according to the operation description information to obtain the at least one labeled picture comprises: for a first operation display picture in the at least one operation display picture, determining a labeling position of the first operation description information according to the first operation display picture and the first operation description information, the first operation display picture being any one of the at least one operation display picture, and the first operation description information and the first operation display picture having a corresponding relationship; and labeling the first operation display picture according to the labeling position of the first operation description information to obtain a first labeled picture.

[0011] In this way, the operation description information is labeled in the operation display picture, so that the operation description information is combined with the operation display picture, and all contents about one operation step can be obtained through one labeled display picture.

[0012] In a possible implementation, the determining of the labeling position of the first operation description information according to the first operation display picture and the first operation description information comprises: performing text recognition on the first operation display picture to obtain a text set of the first operation display picture, the text set comprising at least one candidate text information displayed in the first operation display picture; determining first text information matched with the first operation description information from the at least one candidate text information; and determining the labeling position of the first operation description information according to a display position of the first text information in the first operation display picture.

[0013] In this way, the operation description information is displayed near the operation object corresponding to the operation step, the operation description information attracts the attention of the user, the user can quickly notice the display position of the operation object in the labeled picture, and the prominence of the operation object in the labeled picture is improved.

[0014] In a possible implementation, the reply generation method further includes: performing a sentence component analysis on the operation instruction information to obtain a component analysis result, the component analysis result being used to represent a sentence component to which a word included in the operation instruction information belongs; and reducing the operation instruction information according to the component analysis result to obtain reduced operation instruction information, the reduced operation instruction information being used to obtain the labeled picture.

[0015] In this way, on one hand, text information with a shorter character length is obtained, thereby helping to reduce the amount of calculation in the semantic matching process; and on the other hand, the deleted word usually does not have a semantic association with the operation object, and removing the verb in the operation instruction information is equivalent to removing irrelevant interference information in the operation instruction information, thereby improving the accuracy of the matched text information.

[0016] In a possible implementation, the reply answer of the first question is generated according to the at least one labeled picture; the operation demonstration video is generated according to the at least one labeled picture, and the operation demonstration video is used to display the at least one labeled picture in the execution order of the execution steps; and the operation demonstration video is taken as the reply answer of the first question.

[0017] In this way, it is helpful to generate a reply answer with the shortest length, to reduce the display area occupied by the reply answer in the dialogue reply interface, to help reduce the searching operation of the user in the process of viewing the reply answer, and to help improve the viewing convenience of the reply answer.

[0018] In a possible implementation, the reply answer of the first question includes the at least one labeled picture.

[0019] In this way, the at least one labeled picture is included in the reply answer, which helps to simplify the generation logic of the reply answer and to reduce the amount of calculation in the process of generating the reply answer by the server, thereby improving the generation speed of the reply answer.

[0020] In a possible implementation, before the at least one labeled picture is obtained according to the operation instruction information and the operation demonstration picture, the method further includes: for second operation instruction information in the intermediate answer of the first question that does not correspond to an operation demonstration picture, determining a second operation demonstration picture based on the second operation instruction information; and adding the second operation demonstration picture to the intermediate answer of the first question to obtain an updated intermediate answer, the updated intermediate answer being used to generate the at least one labeled picture.

[0021] In this way, the missing operation demonstration picture is supplemented by searching the picture based on the operation instruction information, which helps to improve the comprehensiveness of the reply answer generated based on the operation instruction information and the operation demonstration picture, thereby achieving the effect of helping the user to quickly solve the to-be-solved problem and reducing the situation that the user repeatedly asks the same question multiple times.

[0022] In a possible implementation, for the second operation description information in the intermediate answer to the first question that does not correspond to the operation display picture, the second operation display picture is determined based on the second operation description information, including: in the case that the second operation description information and the third operation description information are included in the at least one operation description information, the second operation display picture is determined from the at least one candidate picture based on the third operation display picture; wherein the at least one candidate picture is searched based on the second operation description information, the third operation description information corresponds to the third operation display picture, and the execution sequence numbers of the third operation description information and the second operation description information are adjacent.

[0023] In this way, by executing the previous or subsequent operation display picture in time sequence, more reference information can be provided for the process of selecting the operation display picture from the candidate picture, thereby helping to select the operation display picture that meets the operation description information from the candidate picture.

[0024] In a second aspect, a dialogue reply method applied to an electronic device is provided, including: displaying a dialogue reply interface, the dialogue reply interface being used for dialogue interaction of at least two dialogue subjects; in response to a question obtaining operation, displaying inquiry information describing a first question in the dialogue reply interface; in response to a dialogue request operation, displaying a reply answer to the first question, the reply answer to the first question including at least one labeled picture composed of operation description information and operation display information, the operation description information and the operation display information being used to describe operation steps for solving the first question.

[0025] In a possible implementation, the answer style of the reply answer to the first question includes at least one of the following: a picture set composed of the at least one labeled picture, and / or an operation demonstration video used to play the at least one labeled picture.

[0026] In a possible implementation, the method further includes: in response to a style adjustment operation, displaying adjusted style prompt information, the style prompt information being used to prompt an answer style adopted by a subsequent reply answer; and in response to the dialogue request operation, displaying the reply answer to the first question, including: displaying the reply answer with a style that meets the displayed style prompt information.

[0027] In a third aspect, a reply generation apparatus is provided, which comprises: an information acquisition module configured to acquire inquiry information describing a first question; an answer acquisition module configured to acquire an intermediate answer to the first question according to the inquiry information, the intermediate answer comprising at least one operation description information and at least one operation display picture, the operation description information being configured to describe an operation step for solving the first question, and the operation description information and the operation display picture having a corresponding relationship; a picture acquisition module configured to acquire at least one labeled picture by labeling the operation display picture according to the operation description information; and an answer generation module configured to generate a reply answer to the first question according to the at least one labeled picture.

[0028] In a possible implementation, the picture acquisition module is configured to, for a first operation display picture in the at least one operation display picture, determine a labeling position of the first operation description information according to the first operation display picture and the first operation description information, the first operation display picture being any one of the at least one operation display picture, and the first operation description information and the first operation display picture having a corresponding relationship; and label the first operation description information in the first operation display picture according to the labeling position of the first operation description information, to obtain a first labeled picture.

[0029] In a possible implementation, the picture acquisition module is configured to perform text recognition on the first operation display picture to obtain a text set of the first operation display picture, the text set comprising at least one candidate text information displayed in the first operation display picture; determine first text information matched with the first operation description information from the at least one candidate text information; and determine the labeling position of the first operation description information according to a display position of the first text information in the first operation display picture.

[0030] In a possible implementation, the reply generation apparatus further comprises an information reduction module configured to perform a sentence component analysis on the operation description information to obtain a component analysis result, the component analysis result being configured to represent a sentence component to which a word included in the operation description information belongs; and reduce the operation description information according to the component analysis result to obtain reduced operation description information, the reduced operation description information being configured to acquire the labeled picture.

[0031] In a possible implementation, the picture acquisition module is configured to generate an operation demonstration video according to the at least one labeled picture, the operation demonstration video being configured to display the at least one labeled picture according to an execution order of the execution steps; and take the operation demonstration video as the reply answer to the first question.

[0032] In a possible implementation, the reply answer to the first question comprises the at least one labeled picture.

[0033] In a possible implementation, the reply generation apparatus further includes: an answer updating module, configured to, for second operation description information in the intermediate answer of the first question that does not correspond to the operation display picture, determine a second operation display picture based on the second operation description information, and add the second operation display picture to the intermediate answer of the first question to obtain an updated intermediate answer, and the updated intermediate answer is used to generate the at least one labeled picture.

[0034] In a fourth aspect, a dialogue reply apparatus is provided, and the apparatus includes: an interface display module, configured to display a dialogue reply interface, the dialogue reply interface being used for dialogue interaction of at least two dialogue subjects; a question receiving module, configured to receive a first question input by a user in the dialogue reply interface; and an information sending module, configured to send inquiry information corresponding to the first question to a server. The interface display module is further configured to display a reply answer of the first question sent by the server, the reply answer of the first question including at least one labeled picture, and the at least one labeled picture being obtained by marking a corresponding operation display picture with operation description information, and the operation description information being used to describe operation steps for solving the first question.

[0035] In a possible implementation, the answer style of the reply answer of the first question includes at least one of the following: a picture set composed of the at least one labeled picture, and / or an operation demonstration video used to play the at least one labeled picture.

[0036] In a possible implementation, the dialogue reply apparatus further includes: a style selection module, configured to, in response to a style adjustment operation, display adjusted style prompt information, the style prompt information being used to prompt an answer style adopted by a subsequent reply answer; and the interface display module is configured to display the reply answer of which the style conforms to the displayed style prompt information.

[0037] In a fifth aspect, a server is provided, and the server includes a processor and a memory, the memory being coupled to the processor, and the memory being used to store computer program codes, the computer program codes including computer instructions, and when the processor reads the computer instructions from the memory, the server is caused to perform: obtaining inquiry information describing a first question; obtaining an intermediate answer of the first question according to the inquiry information, the intermediate answer including at least one operation description information and at least one operation display picture, the operation description information being used to describe operation steps for solving the first question, and a corresponding relationship existing between the operation description information and the operation display picture; marking the operation display picture according to the operation description information to obtain at least one labeled picture; and generating a reply answer of the first question according to the at least one labeled picture.

[0038] In a possible implementation, the operation display picture is marked according to the operation instruction information, and the at least one marked picture is obtained, including: for a first operation display picture in the at least one operation display picture, a marking position of the first operation instruction information is determined according to the first operation display picture and the first operation instruction information, the first operation display picture is any one of the at least one operation display picture, and the first operation instruction information and the first operation display picture have a corresponding relationship; the first operation instruction information is marked in the first operation display picture according to the marking position of the first operation instruction information, and a first marked picture is obtained.

[0039] In a possible implementation, the marking position of the first operation instruction information is determined according to the first operation display picture and the first operation instruction information, including: text recognition is performed on the first operation display picture, a text set of the first operation display picture is obtained, and the text set includes at least one candidate text information displayed in the first operation display picture; first text information matched with the first operation instruction information is determined from the at least one candidate text information; and the marking position of the first operation instruction information is determined according to a display position of the first text information in the first operation display picture.

[0040] In a possible implementation, when the processor reads the computer instruction from the memory, the server further performs: performing sentence component analysis on the operation instruction information to obtain a component analysis result, the component analysis result being used to represent a sentence component to which a word included in the operation instruction information belongs; reducing the operation instruction information according to the component analysis result to obtain reduced operation instruction information, the reduced operation instruction information being used to obtain the marked picture.

[0041] In a possible implementation, the reply answer package of the first question is generated according to the at least one marked picture, including: generating an operation demonstration video according to the at least one marked picture, the operation demonstration video being used to display the at least one marked picture according to an execution order of the execution steps; and taking the operation demonstration video as the reply answer of the first question.

[0042] In a possible implementation, the reply answer includes the at least one marked picture.

[0043] In a possible implementation, before the operation display picture is marked according to the operation instruction information, and the at least one marked picture is obtained, the method further includes: for second operation instruction information in the intermediate answer of the first question and not corresponding to an operation display picture, determining a second operation display picture based on the second operation instruction information; adding the second operation display picture to the intermediate answer of the first question to obtain an updated intermediate answer, and the updated intermediate answer is used to generate the at least one marked picture.

[0044] In a sixth aspect, an electronic device is provided, the electronic device comprising: a display screen configured to display a conversation reply interface, the conversation reply interface configured to display inquiry information describing a first question and a reply answer to the first question; a processor configured to receive a first question input by a user in the conversation reply interface; a transceiver configured to transmit, to a server, inquiry information corresponding to the first question; the display screen further configured to display a reply answer to the first question transmitted by the server, the reply answer to the first question comprising at least one annotated picture, the at least one annotated picture being obtained by annotating an operation demonstration picture with operation instruction information, the operation instruction information being configured to describe operation steps for solving the first question; a memory configured to store computer program instructions; and the processor configured to execute the computer program instructions to support the electronic device to implement the method according to any one of the possible implementation manners of the second aspect.

[0045] In a possible implementation manner, the answer style of the reply answer to the first question comprises at least one of: a picture set composed of the at least one annotated picture, and / or an operation demonstration video for playing the at least one annotated picture.

[0046] In a possible implementation manner, when the processor reads the computer program instructions from the memory, the electronic device is further caused to perform: in response to a style adjustment operation, displaying adjusted style prompt information, the style prompt information being configured to prompt an answer style to be adopted by a subsequent reply answer; and in response to a conversation request operation, displaying the reply answer to the first question, comprising: displaying a reply answer whose style conforms to the displayed style prompt information.

[0047] In a seventh aspect, a computer readable storage medium is provided, the computer readable storage medium storing computer program instructions, the computer program instructions being executed by a processing circuit to implement the method according to any one of the possible implementation manners of the first aspect or the method according to any one of the possible implementation manners of the second aspect.

[0048] In an eighth aspect, a conversation reply system is provided, the conversation reply system comprising a server and an electronic device, the server comprising a first module configured to implement the method according to any one of the possible implementation manners of the first aspect, and the electronic device comprising a second module configured to implement the method according to any one of the possible implementation manners of the second aspect.

[0049] In a ninth aspect, a chip system is provided, the chip system comprising a processing circuit and a storage medium, the storage medium storing computer program instructions, the computer program instructions being executed by the processing circuit to implement the method according to any one of the possible implementation manners of the first aspect or the method according to any one of the possible implementation manners of the second aspect.

[0050] In a tenth aspect, a computer program product including instructions, which when the computer program product is run on a computer, cause the computer to perform the method according to any possible implementation of the first aspect or the method according to any possible implementation of the second aspect.

[0051] The technical effects of any possible implementation of the third aspect to the tenth aspect can refer to the technical effects of the first aspect and any possible implementation of the first aspect, or refer to the technical effects of the first aspect and any possible implementation of the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0052] FIG. 1 is a schematic diagram of a reply answer display effect in a conversation reply interface in the related art;

[0053] FIG. 2 is a schematic diagram of a communication system to which a reply generation method provided by an embodiment of the present application is applied;

[0054] FIG. 3 is a schematic diagram of a hardware structure of a server provided by an embodiment of the present application;

[0055] FIG. 4 is a schematic diagram of a hardware structure of an electronic device provided by an embodiment of the present application;

[0056] FIG. 5 is a schematic diagram of interaction between functional modules in a server provided by an embodiment of the present application;

[0057] FIG. 6 is a schematic diagram of a flow of a reply generation method provided by an embodiment of the present application;

[0058] FIG. 7 is a display effect diagram of a conversation reply interface provided by an embodiment of the present application;

[0059] FIG. 8 is a schematic diagram of a structure of a semantic matching model provided by an embodiment of the present application;

[0060] FIG. 9 is a schematic diagram of a labeled picture provided by an embodiment of the present application;

[0061] FIG. 10 is a schematic diagram of a labeled picture provided by an embodiment of the present application;

[0062] FIG. 11 is a schematic diagram of a flow of synthesizing a labeled picture provided by an embodiment of the present application;

[0063] FIG. 12 is a schematic diagram of a structure of a text detection model provided by an embodiment of the present application;

[0064] FIG. 13 is a schematic diagram of an effect after text region recognition provided by an embodiment of the present application;

[0065] FIG. 14 is a schematic diagram of a determination method of first text information provided by an embodiment of the present application;

[0066] FIG. 15 is a schematic diagram of a question reply model provided by an embodiment of the present application;

[0067] FIG. 16 is an interactive timing diagram of a reply generation method provided by an embodiment of the present application;

[0068] FIG. 17 is a schematic diagram of a reply answer display effect provided by an embodiment of the present application;

[0069] FIG. 18 is a schematic diagram of a reply answer display effect provided by an embodiment of the present application;

[0070] FIG. 19 is a schematic diagram of a server structure provided by an embodiment of the present application;

[0071] FIG. 20 is a schematic diagram of an electronic device structure provided by an embodiment of the present application. DETAILED DESCRIPTION

[0072] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application.

[0073] The terms "comprising" and "having" and any variations thereof in the description of the embodiments of the present application are intended to cover the non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally include other steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device.

[0074] Hereinafter, the terms "first", "second", and the like are used only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second", and the like can explicitly or implicitly include one or more of the features.

[0075] In the embodiments of the present application, the words "exemplarily" or "for example" and the like are used to represent as an example, illustration or description. Any embodiment or design scheme described as "exemplarily" or "for example" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words "exemplarily" or "for example" and the like are intended to present the relevant concept in a specific manner.

[0076] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "multiple" is two or more. "And / or" in this paper is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone.

[0077] With the development of AI technology, the dialogue reply technology is applied in more and more fields. The dialogue reply technology is used to analyze the inquiry information put forward by the user to obtain a reply answer, and through the dialogue reply technology, the user message can be automatically replied. The dialogue reply technology is widely used in customer service reply, smart education and other fields. In these scenarios, the dialogue reply technology is applied, the inquiry information sent by the user can be replied at any time, thereby helping to improve the reply speed of the inquiry information.

[0078] In an implementation manner, the automatic reply technology searches based on the inquiry information put forward by the user, or uses an AI model to process the inquiry information, so as to obtain a reply answer for solving a to-be-solved problem described in the inquiry information. Taking the to-be-solved problem as a computer device software operation problem as an example, the reply answer of the to-be-solved problem usually includes operation instruction information for describing operation steps. The automatic reply of the inquiry information is realized by displaying the reply answer in a dialogue reply interface.

[0079] However, the reply answer in the form of pure text needs to be read in order to understand each operation step. For example, the operation steps for solving an operation problem include n, and the reply answer includes operation instruction information corresponding to the n operation steps, n is a positive integer. For the case that the value of n is large, the reply answer needs to occupy a long length in the dialogue reply interface, and the step explanation information in the form of text is far less intuitive than the picture effect, and the user needs to spend a long time to read and understand the reply answer. In addition, since the user needs to read and understand the operation instruction information to determine the operation step, different users may have different understandings after reading the same operation instruction information, which leads to the fact that some users cannot perform the correct operation step according to the operation instruction information, thereby affecting the efficiency of the user in solving the to-be-solved problem.

[0080] In another implementation manner, the reply answer includes operation display pictures in addition to operation instruction information. In this case, the operation display pictures and the operation instruction information need to be displayed synchronously in the dialogue reply interface, and the user often needs to read the operation instruction information to understand the execution mode of the operation step, and then find the operation object of the operation step from the operation display pictures. It can be seen that in this form of reply answer, the combination degree of the operation instruction information and the operation display pictures is poor, which leads to poor readability of the reply answer, the viewing steps of the reply answer are complicated, and it is easy to cause the user to ask the same question multiple times, occupy the computing power of the dialogue reply system, and increase the reply pressure of the dialogue reply system.

[0081] In addition, since the reply information generated in the related art is long in length, but the display area of the dialogue reply interface is limited, it can result in that the operation instruction information and the operation display picture corresponding to the operation step that needs to be performed first can not be completely displayed in the dialogue reply interface. Since there is an execution order between multiple operation steps, the user needs to start from the earliest operation step corresponding to the problem to be solved, and needs to find the operation instruction information and / or the operation display picture corresponding to the operation step that needs to be performed first in the dialogue reply interface by means of a searching operation such as a drag operation, a slide operation, and a page turning operation, which increases the operation complexity of the process of viewing the reply answer.

[0082] Exemplarily, FIG. 1 is a display effect diagram of a reply answer in a dialogue reply interface. As shown in FIG. 1, the inquiry information 10 raised by the user in the dialogue reply interface is "What is the setting step of the dialing program?", and the reply answer is: operation instruction information 21, operation display picture 22, operation instruction information 23, and so on. The reply answer is long in length, and even the dialogue reply interface can not display all the operation instruction information and the operation display picture at the same time.

[0083] Based on the above problems, the embodiment of the present application provides a reply generation method. After determining the step instruction information and the operation display picture based on the inquiry information of the problem to be solved, the operation instruction text and the operation display picture corresponding to the same operation step are combined to generate a marked picture, the execution mode of the operation step is displayed intuitively through the marked picture, and the execution effect of the operation step is displayed. Then, the reply answer is generated according to the marked picture.

[0084] In the automatic dialogue scene, by showing the user the reply answer generated by the marked picture, on the one hand, compared with the reply answer in the form of pure text, the reply answer including the picture is more vivid. Through the picture, the user can be helped to familiarize with the environment (such as the page to be operated, the instrument panel, etc.) of the operation step in advance before performing the step, which is also helpful to eliminate the understanding ambiguity generated after the user reads the step instruction information.

[0085] On the other hand, compared with the reply answer in which the text and the picture are separated from each other, the marked picture combines the operation instruction information and the operation display picture, can provide the user with all the information of the picture and the text about an operation step through the marked picture, so that the user does not need to read the operation information and then find the corresponding operation display picture from multiple operation display pictures to view, which is helpful to improve the reading experience of the reply answer.

[0086] In addition, the scheme provided by the embodiments of the present application does not need to simultaneously display the operation instruction information and the operation display picture in the dialogue reply interface, compared with the display form of the dialogue reply interface in the above, the length of the reply answer is shortened, thereby reducing the display area occupied by the dialogue reply interface in the reply answer, which helps to reduce unnecessary searching operations in the process of searching and browsing the reply answer in the dialogue reply interface.

[0087] The reply generation scheme provided by the embodiments of the present application can not only be used to generate a reply answer in real time in a dialogue reply system and provide a reference answer for a user to solve a problem in time, but also can be used to pre-generate a reply answer. For example, a developer of a system or an application program pre-generates a reply answer by using the scheme, and saves the pre-generated reply answer in an official website, so that a user of the system or the application program can directly obtain the reply answer from the official website. In this way, it is helpful to help the developer of the system or the application program reduce the operation and maintenance cost, and it is also helpful to help the user quickly solve an operation problem encountered in the process of using the system or the application program.

[0088] Exemplarily, FIG. 2 is a schematic diagram of a communication system to which the reply generation method provided by the embodiments of the present application is applied. As shown in FIG. 2, the communication system includes a server 100 and an electronic device 200.

[0089] Optionally, the server 100 can be a device or a server with a computing function, such as a cloud server or a network server. The server can be a server, a server cluster composed of multiple servers, or a cloud computing service center. The server 100 can be a background server of an application program with a dialogue reply function, or a background server of a system tool provided by an operating system of the electronic device 200.

[0090] Optionally, the electronic device 200 can be a terminal device such as a mobile phone, a tablet computer, a notebook computer, a smart screen, a wearable device, a vehicle-mounted terminal, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), an AI device, etc. The operating system installed on the electronic device 200 includes but is not limited to or other operating systems. The specific type of the electronic device 200 and the operating system installed thereon are not limited in the present application.

[0091] In some embodiments, the dialogue reply scenario to which the communication system described above is applied is, for example, a scenario of intelligent customer service, intelligent voice assistant, chat robot, etc. In these scenarios, the user interacts with the robot or dialogue reply system, and the robot or dialogue reply system usually needs to generate a reply answer for the user. The server 100 obtains the inquiry information sent by the electronic device 200 and describing the problem to be solved, and obtains an intermediate answer to the problem to be solved based on the inquiry information, generates a labeled picture based on the intermediate answer, and generates a reply answer based on the labeled picture. Subsequently, the server 100 feeds back the reply answer to the electronic device 200, so that the electronic device 200 displays the reply answer to the problem to be solved on the dialogue reply interface.

[0092] Exemplarily, FIG. 3 shows the structure of the server 100. As shown in FIG. 3, the server 100 includes at least one processor 301, a communication line 302, a memory 303, and at least one communication interface 304. The memory 303 can also be included in the processor 301.

[0093] The processor 301 can be a general central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of programs of the solutions of the present application.

[0094] The communication line 302 can include at least one channel for transmitting information between the above-mentioned components.

[0095] The memory 303 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer, but is not limited thereto. The memory 303 can exist independently and be connected to the processor 301 through the communication line 302. The memory 303 can also be integrated with the processor 301.

[0096] The memory 303 is configured to store computer-executable instructions for implementing the solutions of the present application, and the processor 301 is configured to execute the computer-executable instructions stored in the memory 303. The processor 301 is configured to execute the computer-executable instructions stored in the memory 303, so as to implement the reply generation method provided in the following embodiments of the present application.

[0097] The communication interface 304 is configured to communicate with other devices. In the embodiments of the present application, the communication interface 304 can be a module, a circuit, a bus, an interface, a transceiver or other devices capable of realizing the communication function. Optionally, when the communication interface 304 is a transceiver, the transceiver can be a separately arranged transmitter, which can be configured to send information to other devices. The transceiver can also be a separately arranged receiver, which can be configured to receive information from other devices. The transceiver can also be a component integrating the functions of sending and receiving information. The specific implementation of the transceiver is not limited in the embodiments of the present application.

[0098] Optionally, the computer-executable instructions in the embodiments of the present application can also be referred to as application program codes, instructions, computer programs or other names, which are not limited in the embodiments of the present application.

[0099] In specific implementation, as an embodiment, the processor 301 can include one or more CPUs, for example, CPU0 and CPU1 in FIG. 3. In specific implementation, as an embodiment, the server 100 can include multiple processors, for example, the processor 301 and the processor 307 in FIG. 3. Each of the processors can be a single-CPU processor or a multi-CPU processor. The processor herein can refer to one or more devices, circuits and / or processing cores for processing data (for example, computer program instructions).

[0100] It can be understood that the structure as shown in FIG. 3 does not constitute a specific limitation on the structure implementation of the server 100. In other embodiments of the present application, the server 100 can include more or fewer components than those shown, or combine certain components, or split certain components, or different arrangement of components. The components shown can be implemented in hardware, software or a combination of software and hardware.

[0101] Exemplarily, FIG. 4 shows a structural schematic diagram of the electronic device 200. As shown in FIG. 4, the electronic device 200 can include a processor 110, a memory 120, a communication module 130 and a display screen 140, etc.

[0102] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU), etc. Among them, different processing units can be independent devices, or can be integrated in one or more processors.

[0103] The memory 120 includes an external memory card and an internal memory. The external memory card communicates with the processor 110 through an external memory interface to realize data storage functions. For example, files such as music, videos, etc. are saved in the external memory card. The internal memory can be used to store computer executable program codes, including instructions. The internal memory can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc. The data storage area can store data created during the use of the electronic device 200 (such as audio data, a phone book, etc.), etc. In addition, the internal memory can include a high-speed random access memory, and can also include a non-volatile memory, for example, at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various function applications and data processing of the electronic device 200 by running the instructions stored in the internal memory and / or the instructions stored in the memory provided in the processor.

[0104] The communication module 130 can include a mobile communication module and a wireless communication module, for realizing data transceiving between the electronic device 200 and the server 100.

[0105] The display screen 140 is configured to display images, videos, and the like. The display screen 140 includes a display panel. The display panel can be a liquid crystal display (LCD). For example, the display panel can be manufactured by using an organic light emitting diode (OLED), an active-matrix organic light emitting diode (AMOLED), a flex light-emitting diode (FLED), a Mini-led, a Micro-led, a Micro-ole, and a quantum dot light emitting diode (QLED), and the like. In some embodiments, the electronic device 200 can include one or N display screens 140, where N is a positive integer greater than 1.

[0106] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 can include more or fewer components than shown, or combine certain components, or split certain components, or different arrangement of components. The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0107] In some embodiments, the electronic device 200 displays, through the display screen 140, inquiry information describing a problem to be solved, sends the inquiry information to the server 100 through the communication module 130, and receives a reply answer generated by the server 100 according to the inquiry information. Then, the electronic device 200 displays the reply answer through the display screen 140. The inquiry information can be input by the user through the electronic device 200.

[0108] As shown in FIG. 5, the server 100 provided by the embodiments of the present application can include the following software functional modules: a reply generation module 510 and a question identification and processing module 512. The reply generation module 510 is configured to, after obtaining the inquiry information describing the problem to be solved through the information acquisition channel, obtain the intermediate answer of the problem to be solved based on the inquiry information. The question identification and processing module 512 is configured to identify the type of the problem to be solved, and in the case of determining that the problem to be solved described by the inquiry information belongs to an operation type problem, label the operation display picture using the operation instruction information included in the intermediate answer, so as to obtain a labeled picture. In addition, the question identification and processing module 512 is further configured to generate a reply answer according to the labeled picture. Optionally, the question identification and processing module sends at least one labeled picture to the electronic device as the reply answer. Optionally, the question identification and processing module generates an operation demonstration video based on the labeled picture. For example, the server 100 sends the operation demonstration video to the electronic device 200 as the reply answer of the first question.

[0109] It can be understood that the structure shown in FIG. 5 does not constitute a specific limitation on the server 100. In other embodiments of the present application, the server 100 can include more or fewer modules than those shown, or combine certain modules, or split certain modules, or different arrangement of modules. The modules shown can be implemented in hardware, software or a combination of software and hardware.

[0110] Optionally, the server 100 can further include a plug-in engine module 514, configured to, in the case of lacking an operation display picture corresponding to a certain operation instruction information, find an operation display picture matching the operation instruction information from a search engine.

[0111] In the embodiments of the present application, at least one AI model can be used in the process of generating the reply answer. The functional modules in the server complete the functions possessed by the modules by calling the corresponding AI models. Optionally, the reply generation module includes a semantic matching model 1. The question identification and processing module includes a question classification model and a semantic matching model 2. The semantic matching model is also called a semantic similarity matching model.

[0112] The semantic matching model 1 and the semantic matching model 2 can be natural language processing (NLP) models or generation augmented retrieval (GAR) models. For example, the semantic matching model 1 is configured to retrieve the intermediate answer of the first question based on the inquiry information, and the intermediate answer includes operation instruction information for solving the first question. For more information about the intermediate answer, please refer to the following embodiments.

[0113] The question classification model can be trained by a model base through training data. The model base can be a generative pre-training (GPT) model or a generative language model (GLM), etc. For example, the question classification model is a bidirectional encoder representation of transformer (BERT) model.

[0114] In the embodiments of the present application, the GLM model base can be fine-tuned by different types of training data to obtain a corresponding model. For example, the GLM model base is trained by pre-collected type training samples to obtain a question classification model.

[0115] It can be understood that the type training sample refers to the data used when training the question classification model, which can include sample inquiry information and sample labels. The sample label is used to indicate the question type of the question described by the sample inquiry information. For example, the type label is represented by 1-bit data, "1" represents an operation class question, and "0" represents a non-operation class question. The GLM model base is trained by the type training sample, so that the trained question classification model. After the server receives the inquiry information in the process of generating a reply answer, the trained question classification model is used to determine the question type of the problem to be solved by the user, so as to subsequently determine the reply information generation method used in the process of generating a reply answer according to the question type of the problem to be solved.

[0116] Compared with the existing dialog reply system, the operation instruction information in the embodiments of the present application is annotated to display pictures, and the annotated pictures are synthesized. The synthesized annotated pictures can display the operation steps required to solve the problem and reflect the operation objects corresponding to the operation steps, so that the user can determine the operation steps to be performed and the operation objects corresponding to the operation steps by observing the annotated pictures, which helps to improve the readability of the reply answer and improve the user experience of obtaining the reply answer through the dialog reply system.

[0117] The reply generation method provided by the embodiments of the present application will be described in detail below taking a dialog reply system as an example, wherein the dialog reply system includes a server and an electronic device. The server includes a reply generation module and a question recognition and processing module, and can also include a plug-in engine module; the electronic device includes an interaction module and a display module. Please refer to FIG. 6, which is a flowchart of the reply generation method provided by the embodiments of the present application. As shown in FIG. 6, the reply generation method provided by the embodiments of the present application can include steps S610 to S670.

[0118] S610: The electronic device receives inquiry information describing the first question.

[0119] In some embodiments, the first question refers to a question raised by the user to inquire about a reference solution to the first question. Illustratively, the user raises the first question to the dialogue reply system through the electronic device to obtain a method to solve the first question.

[0120] Optionally, the type of the first question includes, but is not limited to, an operation type question and an object recognition type question. The operation type question is used to inquire about at least one operation step required to achieve a target effect. The object recognition type question is used to determine the display style of a target object. For example, an object recognition type question is used to request to identify the position of a target object in a picture.

[0121] Illustratively, the first question is an operation type question related to a computer device, and the first question is used to inquire about how to call a function of an application program or configure a parameter in an operating system. The computer device is the electronic device, or other device that can be manipulated by the user. For example, the computer device is the electronic device, and the user needs to configure the network connection type of the electronic device, and the first question is: how to configure the network connection type of the computer device.

[0122] Illustratively, the first question is an operation type question related to a physical device, a physical tool or an instrument, and the first question is used to inquire about the use method or fault repair method of the physical device, the physical tool or the instrument. For example, the first question is used to inquire about the driving method of a vehicle.

[0123] In some embodiments, the inquiry information is used to describe the first question. Optionally, the inquiry information is a textual description of the first question. For example, the inquiry information refers to a natural language description of the first question. The natural language includes Chinese, English and Japanese, etc., and the type of the natural language is determined according to the region to which the user belongs, the user's usage habit, etc., which is not limited herein.

[0124] In some embodiments, the inquiry information is obtained by the electronic device. Optionally, the way in which the electronic device obtains the inquiry information includes, but is not limited to, text input, copy and paste, and transmission by other devices, etc.

[0125] Illustratively, an application program supporting the dialogue reply function is running in the electronic device, or the operating system of the electronic device integrates a tool program for implementing the dialogue reply function; the inquiry information of the first question is obtained by the application program running in the electronic device or the tool program integrated in the electronic device.

[0126] In one example, the electronic device displays a dialogue reply interface, and the electronic device obtains the inquiry information for describing the first question from the dialogue reply interface. The dialogue reply interface is an interface provided by an application program running in the electronic device or an interface provided by a tool program. FIG. 7 is a display effect diagram of the dialogue reply interface provided by an embodiment of the present application. As shown in FIG. 7, the electronic device displays a dialogue reply interface 700, which includes an information input field 710 and an inquiry control 720. The user can input the inquiry information for describing the first question in the information input field 710, and after completing the input of the inquiry information, the user triggers the inquiry control 720. The electronic device obtains the inquiry information in the information input field 710 in response to the control triggering operation. The control triggering operation includes, but is not limited to, at least one of the following: a click operation, a long press operation, a sliding operation, a key operation, a gesture operation, a voice operation, and the like.

[0127] In some embodiments, the user inputs the inquiry information for describing the first question in another device other than the electronic device, and the other device sends the inquiry information to the electronic device in response to the inquiry information sending operation, so that the electronic device obtains the inquiry information.

[0128] S620: The electronic device sends the inquiry information to the server.

[0129] Optionally, the electronic device generates a dialogue request based on the inquiry information and sends the dialogue request to the server. The dialogue request includes the identification information of the electronic device and the inquiry information. The identification information of the electronic device is used to uniquely identify the electronic device, and the identification information of the electronic device includes at least one of the following: a serial number of the electronic device, an identification code.

[0130] In some embodiments, the electronic device sends the dialogue request to the server in response to the information sending operation. Optionally, the electronic device displays a dialogue reply interface, and the dialogue reply interface displays a sending control. The electronic device generates a dialogue request based on the inquiry information and sends the dialogue request to the server in response to the operation of triggering the sending control. The operation of triggering the sending control includes, but is not limited to, a click operation, a long press operation, a sliding operation, and the like.

[0131] Optionally, the sending control and the inquiry control in the above embodiment can be the same control. That is, the electronic device obtains the inquiry information for describing the first question from the information input field in response to the operation of triggering the inquiry control, and then generates a dialogue request based on the inquiry information and sends the dialogue request to the server, so that the server obtains the inquiry information from the dialogue request. Of course, the electronic device can also send the inquiry information directly to the server.

[0132] S630: The server obtains the intermediate answer to the first question according to the inquiry information.

[0133] In some embodiments, an intermediate answer to the first question is used to indicate the solution to the first question.

[0134] Optionally, the intermediate answer to the first problem includes at least one step-by-step description, which is a textual description of the steps required to solve the first problem. For example, one step-by-step description corresponds to one operation step, and different step-by-step descriptions correspond to different operation steps. The user can solve the first problem by following the corresponding operation steps step by step according to at least one step-by-step description. For example, the server decomposes the intermediate answer to the first problem to obtain at least one operation description. For a method of decomposing the intermediate answer to obtain the operation description, please refer to the following embodiment.

[0135] Optionally, the intermediate answer to the first question includes multiple step-by-step descriptions. In this case, the multiple step-by-step descriptions are arranged in a first order, which relates to the execution order of the operation steps corresponding to each of the multiple step-by-step descriptions. Here, "multiple" refers to two or more steps. For example, the first order may be either from first to last or from last to first, according to the execution order of the multiple operation steps.

[0136] For example, the intermediate answer to the first question includes: Step Description Information 1, Step Description Information 2, and Step Description Information 3. Step Description Information 1 corresponds to Operation Step 1, Step Description Information 2 corresponds to Operation Step 2, and Step Description Information 3 corresponds to Operation Step 3. The three operation steps are executed in the following order: Operation Step 1, Operation Step 2, Operation Step 3. Therefore, the order of these three operation description information in the intermediate result of the first question is: Step Description Information 1, Step Description Information 2, Step Description Information 3.

[0137] In some embodiments, the intermediate answer to the first question further includes at least one operation demonstration image. One operation demonstration image corresponds to one or more step descriptions. Optionally, the operation demonstration image is used to characterize the execution effect or execution method of the operation step. For example, if operation description information 1 corresponds to operation step 1, and operation step 1 is clicking the first control on the page to be operated, then the operation demonstration image corresponding to operation description information 1 is used to demonstrate: the display effect of the page to be operated after clicking the first control. As another example, if operation description information 1 corresponds to operation step 1, and operation step 1 is clicking the first control on the page to be operated, then the operation demonstration image corresponding to operation description information 1 is used to demonstrate the display position of the first control on the page to be operated.

[0138] Optionally, the operation display picture can also be used to represent the position distribution of each operable control in the operation panel. For example, when the first question is an operation-related question about a physical device, the operation display picture is used to represent the control panel of the physical device. The operation display picture can be pre-recorded by the manufacturer of the physical device or obtained from the Internet through a search engine.

[0139] Optionally, for the first step description information of the at least one step description information, the first step description information corresponds to a first operation display picture. The first operation display picture is obtained based on the inquiry information of the first question, or the first operation display picture is obtained by searching according to the first step description information. The first step description information is any one of the at least one step description information. For details of the specific acquisition method of the operation display picture, please refer to the following embodiments.

[0140] In some embodiments, the server implements a semantic search operation based on the inquiry information to obtain the intermediate answer of the first question. Optionally, the search range of the semantic search operation includes but is not limited to at least one of the following: description documents, answer libraries, question and answer posts, etc. The description documents can be referred to as technical documents in the software field, and can be referred to as user manuals, operation guides, etc. in the field of physical devices, physical instruments, tools, etc.

[0141] The process of retrieving the intermediate answer of the first question will be introduced below by taking the first question as an operation-related question about an application program or a system as an example. It can be understood that the search range of the semantic search operation corresponding to different types of first questions is different, and the semantic search principle is similar, so the semantic search process is universal for different types of first questions, and will not be described one by one here.

[0142] Optionally, the application program or the system has a technical document provided by a developer. The technical document is used to introduce the use method of the application program or the system. The system includes but is not limited to: an operating system, a management system, a data storage and management system, etc.

[0143] Optionally, the server performs a semantic search in the technical document based on the inquiry information to obtain the intermediate answer of the first question. For example, the intermediate answer of the first question is recorded in the technical document, and the server finds the intermediate answer of the first question in the technical document through the inquiry information.

[0144] As can be known from the above embodiments, the server can include a reply generation module, and the semantic matching model is set in the reply generation module. In one example, the server determines the target text content matched with the inquiry question from the technical document through the semantic matching model; and obtains the intermediate answer of the first question according to the target text content.

[0145] In some embodiments, the semantic matching model is an NLP model or an RAG model. Optionally, the semantic matching model is designed based on a BERT model. For example, the semantic matching model is obtained by fine-tuning a Sentence-BERT model.

[0146] FIG. 8 is a structural diagram of a semantic matching model according to an embodiment of the present application. As shown in FIG. 8, the semantic matching model includes two branch networks and a matching degree prediction layer 830. The two branch networks are a first branch network 810 and a second branch network 820, and the first branch network 810 and the second branch network 820 have the same network structure and can generate sentence vectors of the same dimension. The first branch network 810 is used to calculate the sentence vector of the query information, and the second branch network 820 is used to calculate the sentence vector of the candidate text content. The matching degree prediction layer 830 is used to calculate the matching result between the query information and the candidate text content based on the sentence vector of the query information and the sentence vector of the candidate text content, and the matching result is used to represent the matching degree between the query information and the candidate text content.

[0147] Taking the first branch network as an example, the first branch network includes a feature extraction layer (such as a BERT constructed feature extraction layer) and a pooling layer. The candidate text content is part of the text content in the technical document. For example, the server determines the length and position of each content section in the technical document according to the directory in the technical document, so as to take each content section as a candidate text content, thereby obtaining a plurality of candidate text contents.

[0148] For example, the matching degree prediction layer calculates the vector cosine between the sentence vector of the candidate text content and the sentence vector of the query information to obtain the matching result between the candidate text content and the query information. The matching result is used to reflect the similarity between the candidate text content and the query information.

[0149] For the matching result calculated by the vector cosine, the numerical range of the matching result is [0, 1], and the larger the numerical value of the matching result is, the greater the similarity between the candidate text content and the query information is; the smaller the numerical value of the matching result is, the greater the difference between the candidate text content and the query information is. For the matching result calculated by the Hamming distance and the Euclidean distance, the smaller the numerical value of the matching result is, the more similar the candidate text content and the query information are; the larger the numerical value of the matching result is, the greater the difference between the candidate text content and the query information is.

[0150] After obtaining the target text content, the server obtains at least one operation instruction information from the target text content. The following describes the process through several embodiments.

[0151] The server selects a first matching result from the plurality of matching results, and takes the candidate text content corresponding to the first matching result as target text content. The server obtains an intermediate answer of the first question according to the target text content. Optionally, the first matching result is the maximum value in the plurality of matching results, or the minimum value in the plurality of matching results. For example, in the case where the matching result is determined by vector cosine, the first matching result is the matching result with the maximum value in the plurality of matching results. For another example, in the case where the matching result is determined by Hamming distance or Euclidean distance, the first matching result is the matching result with the minimum value in the plurality of matching results.

[0152] In some embodiments, the server splits the target text content to obtain the at least one operation instruction information. The at least one operation instruction information can constitute part or all of the content in the target text content.

[0153] Optionally, the server splits the target text information according to the execution sequence number to obtain the at least one operation instruction information. The target text content includes the at least one operation instruction information and the execution sequence number of each operation instruction information, and the execution sequence number is used to indicate the execution order of the operation step corresponding to the operation instruction information.

[0154] For example, the server identifies the at least one execution sequence number from the target text content, and divides the target text information according to the execution sequence number to obtain the at least one operation instruction information. The execution sequence number is used to divide the at least one operation step.

[0155] Optionally, the execution sequence number can be a fixed sequence number format. For example, the sequence number format of the execution sequence number is a number. For another example, the sequence number format of the execution sequence number is a number+identifier. For another example, the execution sequence number is a literal instruction, and the sequence number format of the execution sequence number is “step x”, where x is a positive integer.

[0156] In some embodiments, the server identifies the at least one execution sequence number from the target text content based on the sequence number format of the execution sequence number. For example, in the case where at least two execution sequence numbers are identified from the target text content, the statements between any two adjacent execution sequence numbers in the target text content correspond to an operation instruction information.

[0157] For example, the target text content is: 1. Click the desktop icon to open the network configuration page; 2. Click the "Network Settings" button, check the "Wireless Network Connection" in the pop-up window, and click "Next"; 3. Click the "Save" control to complete the network configuration. Among them, the sequence number format of the execution sequence number is: a combination of numbers and ".". The server can identify 3 execution sequence numbers from the target text information, which are "1.", "2." and "3.", respectively. The statements between adjacent execution sequence numbers correspond to an operation instruction information. Through the target text content, 3 statements can be disassembled, one of which is "Click the desktop icon to open the network configuration page".

[0158] Alternatively, the execution sequence number can also be a punctuation mark. For example, a semicolon or a period in the target text content is used as an execution sequence number. In this case, if the target text content includes 3 execution sequence numbers in the form of punctuation marks, the target text content is divided into 4 statements, and each statement corresponds to an operation instruction information. It can be understood that in the case where the target text content includes only one execution sequence number or does not include an execution sequence number, the entire target text content corresponds to an operation instruction information.

[0159] For example, after dividing the target text content by the execution sequence number to obtain at least one statement, for any one of the at least one statement, the server takes the statement as an operation instruction information, or performs component recognition on the statement, and deletes a number of characters (such as prepositions, adverbials, etc.) in the statement that have relatively low relevance to the operation steps to obtain the operation instruction information corresponding to the statement.

[0160] In this way, at least one operation instruction information can be extracted from the text content. Compared with the complete text content, the sentence of the operation instruction information is shorter, which helps users quickly find the operation action and operation object from the operation instruction information.

[0161] The following describes the method for obtaining the operation display picture through several embodiments.

[0162] In some embodiments, in the process of generating the intermediate answer to the first question, the server can obtain at least one operation display picture. Alternatively, the server obtains at least one operation instruction information corresponding to the operation display picture.

[0163] In one possible implementation, the source of the operation instruction information is the same as the source of the operation display picture corresponding thereto. Alternatively, the operation instruction information and the operation display picture corresponding thereto are both obtained based on semantic matching of the inquiry information. In this embodiment, the function of obtaining the operation display picture can be realized by the question classification and processing module in the server.

[0164] Taking the operation instruction information being acquired from the technical document as an example, generally, the technical document includes a picture serial number, if the picture serial number is included in the target text content, the server can take the picture with the picture serial number in the technical document as an operation display picture.

[0165] For example, after the server divides the target text content into multiple sentences, for any one sentence, the server extracts the operation instruction information from the sentence, and in the case that the sentence also includes a picture serial number, the server determines the picture with the picture serial number in the technical document as the operation display picture corresponding to the operation instruction information. In this case, the operation instruction information and the operation display picture can be acquired in parallel or sequentially. For example, the server creates an information acquisition thread and a picture acquisition thread, acquires the operation instruction information from the technical document through the information acquisition thread, and simultaneously acquires the operation display picture from the technical document through the picture acquisition thread.

[0166] For example, the target text content includes the following content: "as shown in FIG. 3, display control a", which indicates that the picture serial number of an operation display picture is FIG. 3 in the technical document. The server reads the picture with the picture serial number of FIG. 3 from the technical document to obtain an operation display picture.

[0167] The scheme provided in this embodiment ensures that the operation instruction information and the operation display picture have the same source, which helps to improve the matching degree of the operation instruction information and the operation display picture, thereby helping to improve the reliability of the generated reply answer in the case that the generated reply answer includes a picture and text that do not correspond to each other.

[0168] In another possible implementation, the source of the operation instruction information is different from the source of the operation display picture corresponding thereto. In some embodiments, after the operation instruction information is acquired through semantic query, the server searches for at least one picture related to the operation instruction information through a search engine, and selects the operation display picture corresponding to the operation instruction information therefrom. In this case, the function of acquiring the operation display picture can be implemented through a plug-in engine module in the server. The plug-in engine module is used to send the operation instruction information to the search engine to acquire at least one candidate picture. Illustratively, the server is provided with a plug-in engine module for providing reliable operation display pictures.

[0169] The following introduces the operation display picture acquisition method through several embodiments to ensure that the server reduces the occurrence of operation instruction information without corresponding operation display pictures during the generation of the reply result of the to-be-solved problem. Moreover, the following embodiments help to find more accurate and reliable operation display pictures, thereby avoiding the problem of using incorrect operation display pictures during the generation of the reply answer. The present embodiment enables the reply generation method to cope with the situation that the operation display picture cannot be directly obtained through the inquiry information or the number of operation display pictures is insufficient, and helps to improve the universality of the reply generation method in various problems and scenarios.

[0170] In some embodiments, the server sends a picture search request to the search engine. The picture search request includes k operation instruction information, where k is a positive integer. Each of the k operation instruction information does not have a corresponding operation display picture. The search engine receives at least one candidate picture corresponding to each of the k operation instruction information. For each of the k operation instruction information, one of the at least one candidate picture is selected as the corresponding operation display picture.

[0171] It can be understood that the server can generate a picture search request for each operation instruction information (one picture search request includes one operation instruction information), or control multiple operation instruction information to share one picture search request (one picture search request includes multiple operation instruction information). Here, the number of operation instruction information carried by one picture search request is not limited.

[0172] Optionally, the step is performed by a plug-in engine module of the server. For example, the plug-in engine module selects the operation display information corresponding to the operation instruction information from the at least one candidate picture according to the recall order of the at least one candidate picture. Generally, in order to improve the hit rate of the returned result, the search engine determines the recall order of multiple candidate pictures. Then, the earlier the recall order of the candidate picture, the greater the degree of association between the candidate picture and the operation instruction information considered by the search engine. Therefore, preferentially selecting the candidate picture with the earlier recall order as the operation display picture can make the operation display picture and the operation instruction information have a certain relevance. At the same time, it does not increase the additional calculation amount of the process of determining the operation display picture, thereby ensuring the reply generation speed, and helping the dialogue reply system to quickly respond to the inquiry information raised by the user, so as to improve the response speed of the dialogue reply system.

[0173] In some embodiments, the plug-in engine module selects, based on other operation instruction information, an operation display picture corresponding to the current operation instruction information from the candidate pictures. Taking the first question as a question about an operation related to an application or a system, and taking the operation instruction information obtained from a technical document as an example, in a case where the second operation instruction information and the third operation instruction information are included in the at least one operation instruction information, if there is a third operation display picture corresponding to the third operation instruction information in the technical document, but there is no second operation display picture corresponding to the second operation instruction information, the server determines the second operation display picture from the at least one candidate picture based on the third operation display picture.

[0174] Considering that the operation steps related to an application or a system to be operated are performed in a page to be operated, operation steps adjacent in execution sequence have a certain probability of being performed in the same page to be operated. Even if operation steps adjacent in execution sequence need to be performed in different pages to be operated, since the same version of an application or a system usually uses the same display element style, border, and the like in interface window design, the operation display pictures of the preceding or subsequent operation steps adjacent in execution sequence can provide reference information for selecting an operation display picture from candidate pictures, and help to select an operation display picture corresponding to the operation instruction information from the candidate pictures.

[0175] The second operation instruction information and the third operation instruction information have execution sequence numbers adjacent to each other. For example, the execution sequence number of the second operation instruction information is n, and the execution sequence number of the third operation instruction information is n+1, where n is a positive integer. For another example, the execution sequence number of the second operation instruction information is m, and the execution sequence number of the third operation instruction information is m-1, where m is a positive integer greater than 1.

[0176] For example, the server determines the key point information of the third operation display picture based on a scale invariant feature transform (SIFT) method, and the at least one candidate picture respectively has key point information; the server respectively calculates the similarity between the at least one candidate picture and the third operation display picture based on the key point information, and takes the candidate picture with the highest similarity to the third operation display picture as the second operation display picture.

[0177] Taking the key point information of the third operation display picture as an example, the key point information is composed of at least one key point descriptor. The key point refers to a point in the third operation display picture that has rotation invariance. The key point descriptor is used to locate the position, scale, and neighborhood direction of the key point in the third candidate picture. Exemplarily, the key point descriptor can be represented in the form of a vector.

[0178] Optionally, the server performs key point matching based on the key point information of the first operation display picture and the key point information of the candidate picture, to determine a similarity between the third operation display picture and the candidate picture. The greater the value of the similarity is, the more similar the third operation display picture is to the candidate picture; the smaller the value of the similarity is, the greater the difference between the third operation display picture and the candidate picture is; and the server determines the candidate picture with the greatest value of the similarity as the second operation display picture.

[0179] In one possible embodiment, the server determines the second operation display picture corresponding to the fourth operation instruction information from at least one candidate picture according to the third operation display picture corresponding to the third operation instruction information and the fourth operation display picture corresponding to the fourth operation instruction information. The server calculates a first similarity between each of the at least one candidate picture and the third operation display picture, and a second similarity between each of the at least one candidate picture and the fourth operation display picture. For any one of the at least one candidate picture, the server performs weighted summation on the first similarity and the second similarity of the candidate picture to obtain a cumulative similarity of the candidate picture. Subsequently, the server selects the candidate picture with the greatest value of the cumulative similarity from the at least one candidate picture as the second operation display picture. For details of the similarity calculation, please refer to the previous embodiment, which will not be repeated here.

[0180] In this embodiment, the third operation instruction information and the fourth operation instruction information are adjacent to the second operation instruction information in the execution sequence. For example, the execution sequence number of the third operation instruction information is p-1, the execution sequence number of the second operation instruction information is p, and the execution sequence number of the fourth operation instruction information is p+1, where p is a positive integer greater than 1. In this example, the third operation display picture is a previous operation display picture of the second operation display picture, and the fourth operation display picture is a subsequent operation display picture of the second operation display picture.

[0181] In this embodiment, the operation display picture corresponding to the operation instruction information adjacent in the execution sequence is determined from the at least one candidate picture recalled from the search engine, which provides additional reference information for the process of screening candidate pictures and helps to improve the consistency between the selected operation display picture and the operation instruction information.

[0182] The scheme is applied to a server with strong computing power, and good results can be achieved. In addition, it can also be applied in a dialogue reply scene with low timeliness requirement. For example, in the application scenario of pre-generated reply answers, the matching degree between operation instruction information and its corresponding operation display picture is improved through this way. Since operation instruction information and its corresponding operation display picture are needed in the generation process of reply answers, improving the matching degree between operation instruction information and its corresponding operation display picture helps to improve the accuracy and reliability of generated reply answers, thereby improving the quality of dialogue replies to help users clearly and intuitively determine the solution to the problem to be solved based on the reply answers.

[0183] In one example, the intermediate answer of the first question includes at least two operation display pictures, and a part of the pictures are obtained by semantic retrieval based on the query information, and the other part is obtained by searching the search engine for pictures related to the operation instruction information. For example, taking the operation instruction information from a technical document as an example, for any operation instruction information, the server first tries to obtain the operation display picture corresponding to the operation instruction information from the technical document. If the technical document includes the operation display picture corresponding to the operation instruction information, the server searches for the corresponding operation display picture for the next operation instruction information; if the technical document does not include the operation display picture corresponding to the operation instruction information, the server searches the search engine for pictures related to the operation instruction information to find the corresponding operation display picture. Then, the server searches for the corresponding operation display picture for the next operation instruction information until each operation instruction information has a corresponding operation display picture. The server obtains at least one operation instruction information and at least one operation display picture as the intermediate answer of the first question.

[0184] Searching for missing operation display pictures through operation instruction information helps to improve the comprehensiveness of reply answers generated based on operation instruction information and operation display pictures, thereby helping users to quickly solve the problem to be solved and reducing the occurrence of repeated questioning of the same problem by users.

[0185] Through the above S630, the server can obtain operation instruction information and operation display pictures according to the inquiry information uploaded by the electronic device, and thereby obtain the intermediate answer of the first question. Then, the server can perform the following S640 to obtain corresponding labeled pictures according to the intermediate answer. The process of obtaining the labeled pictures is described in detail below.

[0186] S640: The server obtains at least one labeled picture according to the intermediate answer of the first question.

[0187] In some embodiments, the marked picture refers to an operation demonstration picture marked with operation instruction information. Optionally, the marked picture is obtained based on the operation demonstration picture, and the server marks the operation demonstration picture based on the operation instruction information to obtain a marked picture. For example, for operation instruction information 1, the operation instruction information 1 corresponds to an operation demonstration picture 1, and the server marks the operation instruction information 1 on the operation demonstration picture 1 to obtain a marked picture 1.

[0188] In some embodiments, the server marks the operation demonstration picture based on the operation instruction information to obtain a synthesized marked picture, or the server processes the operation instruction information and marks the operation demonstration picture based on the processed operation instruction information to obtain a synthesized marked picture.

[0189] Optionally, the server reduces the operation instruction information to obtain reduced operation instruction information, and marks the reduced operation instruction information on the operation demonstration picture to obtain a marked picture. For example, the server analyzes the sentence components of the operation instruction information to obtain a component analysis result, and deletes content such as pronouns that have no actual meaning based on the component analysis result to obtain reduced operation instruction information. For example, the reduced operation instruction information includes a verb and an object in the original operation instruction information. The component analysis result is used to indicate the components to which each word in the operation instruction information belongs.

[0190] In this embodiment, the reduced operation instruction information is obtained by reducing the operation instruction information. Since the number of characters in the reduced operation instruction information is less than the number of characters in the operation instruction information before reduction, it is helpful to reduce the area occupied by marking the reduced operation instruction information in the operation demonstration picture, thereby reducing the occlusion of other display content in the operation demonstration picture caused by marking the operation instruction information. In order to combine the operation instruction information and the corresponding operation demonstration picture while preserving the original appearance of the operation demonstration picture as much as possible, it is convenient to locate the operation object of the operation step from the marked picture.

[0191] In some embodiments, the operation step has an operation object, the operation instruction information corresponding to the operation step records attribute information of the operation object, the operation demonstration picture corresponding to the operation instruction information displays the operation object and the attribute information of the operation object, the server determines a display area of the operation object in the operation demonstration picture by recognizing the attribute information of the operation object in the operation demonstration picture, and determines a marking position of the operation instruction information in the operation demonstration picture based on the display area of the operation object in the operation demonstration picture.

[0192] Optionally, the operation object refers to an action object of the operation step. For example, the operation object is a control, a button, a text input box, an icon, a joystick, a key, or the like in the to-be-operated page. The attribute information of the operation object includes at least one of the following: identification information of the operation object, interpretation information for interpreting the function of the operation object, and the like. The identification information of the operation object can be a name, a number, a serial number, or the like of the operation object. For example, the operation object is a control with a name of "connect to the Internet", the operation display picture includes the to-be-operated page, and the server determines the display area of the operation object in the operation display picture by identifying the name "connect to the Internet" of the operation display page from the operation display page. The server determines the marking position of the operation instruction information based on the display area of the operation object.

[0193] Optionally, the marking position of the operation instruction information is used to indicate the display position of the operation instruction information on the operation display picture. For example, the marking position of the operation instruction information is represented by a set of coordinates on a picture coordinate system. The picture coordinate system is used to indicate a two-dimensional coordinate system of each pixel position in the operation display picture. For example, the image coordinate system has a left-bottom vertex of the operation display picture as a vertex, a positive direction of an x-axis as a direction from the left-bottom vertex to a right-bottom vertex, and a positive direction of a y-axis as a direction from the left-bottom vertex to a left-top vertex. For example, the marking position of the operation instruction information includes the left-top vertex coordinate and the right-bottom vertex coordinate of the display area of the operation instruction information. FIG. 9 is one of the schematic diagrams of the marking picture provided by the embodiments of the present application. As shown in FIG. 9, the operation instruction information 920 is displayed in the marking picture 910.

[0194] Optionally, in addition to the operation instruction information, the marking picture also displays a prompt mark, and the prompt mark is used to indicate the display position of the operation object in the operation display picture. The prompt mark includes, but is not limited to, a closed line surrounding the operation object, an icon, or the like. For example, the display position of the prompt mark is determined according to the display position in the operation display area of the operation object, and details are described in the following embodiments.

[0195] For example, the display style of the prompt mark, such as the color and the line thickness, can be pre-set by the user. For example, the electronic device displays a display style setting interface, responds to a setting operation of the prompt mark style, determines and stores the style information of the prompt mark, and sends the query information and the prompt style information to the server, so that the server sets the display style of the prompt mark according to the prompt style information in the process of generating the marking picture. FIG. 10 is another schematic diagram of the marking picture provided by the embodiments of the present application. As shown in FIG. 10, the operation instruction information 920 and the prompt mark 930 are displayed in the marking picture 910, and the prompt mark 930 surrounds the operation object.

[0196] In this embodiment, the prompt mark is used to improve the prominence of the operation object in the operation display picture, and the display style of the prompt mark is customized by the user, which helps the user to become familiar with the prompt mark, and thus helps the user to quickly determine the display position of the operation object in the operation display picture.

[0197] Optionally, the operation instruction information is located above, below, left or right of the display position of the operation object.

[0198] For example, the distance between the annotation position of the operation instruction information and the display area of the operation object in the operation display picture is less than or equal to a distance threshold. The distance threshold is pre-set, for example, the value range of the distance threshold is [0, X / Y] pixels, X is the pixel length of the shortest side of the operation display picture, and Y is a natural number greater than or equal to 1 (for example, Y=2). For example, the display area of the operation object in the operation display picture is a rectangular area A, the operation display picture has a rectangular area B, the rectangular area A and the rectangular area B share a center point (the intersection of the two diagonals of the rectangle), the length of the rectangular area A is a, the height is b, the length of the rectangular area B is a+2k, and the height is b+2k, k is the distance threshold, and then any position in the intersection of the rectangular area B and the rectangular area A c may be used as the annotation position of the operation instruction information. For the determination method of the annotation position of the operation instruction information, please refer to the following embodiments.

[0199] In the operation display picture, the operation instruction information is annotated near the operation object, the operation object is highlighted by the annotated operation instruction information, and after the user's attention is attracted by the operation instruction information displayed in the annotated picture, the user can quickly notice the display position of the operation object in the operation display picture, so that the user can directly determine the position of the operation object in the page to be operated or the control panel after referring to the annotated picture, thereby helping the user to quickly solve the problem.

[0200] In some embodiments, the intermediate answer to the first question includes at least one operation instruction information and at least one operation display picture.

[0201] Optionally, in the case where the number of operation instruction information is large or the number of operation display pictures is large, the server can store the correspondence between the operation instruction information and the operation display picture in a table. The following several embodiments introduce the management method of the operation instruction information and the operation display picture in the process of generating the reply answer.

[0202] In some embodiments, after obtaining the intermediate answer of the first question, the server generates a step text list and a picture list to record the mapping relationship between the operation instruction information and the operation display picture. The step text list includes at least one operation instruction information arranged according to the execution sequence, and the picture list includes at least one operation display picture arranged according to the execution sequence. Optionally, for any operation instruction information, if the operation instruction information corresponds to an operation display picture, the position of the operation instruction information in the step text list is the same as the position of the operation display information corresponding to the operation instruction information in the picture list.

[0203] For example, the server establishes an operation step record table, and the operation step record table includes a step label list, a step composition list, a step text list, and a picture list. The step label list is used to represent the execution sequence of the operation step. The step composition list is used to represent the composition of the operation step.

[0204] For example, an operation step includes multiple operation sub-steps, and the step composition list records each operation sub-step included in the operation step in the list unit corresponding to the operation step. If an operation step does not include multiple operation sub-steps, the list unit corresponding to the operation step in the step display list only records the operation step, or the list unit corresponding to the operation step in the step display list only records that the operation step is empty.

[0205] In one example, the above four lists are arranged side by side to form an operation step record table, and the list elements at the same position in the above four lists are all related to the same operation step.

[0206] Table 1: Operation step record table

[0207] As shown in Table 1, for operation step 1, the step text list includes operation instruction information 1 corresponding to operation step 1, and the picture list includes operation display picture corresponding to operation instruction information 1. For operation step 2, operation step 2 includes operation sub-step 2.1 and operation sub-step 2.2, so the position corresponding to operation step 2 in the step text list records operation instruction text 2.1 corresponding to operation sub-step 2.1 and operation instruction text 2.2 corresponding to operation sub-step 2.2; the position corresponding to operation step 2 in the picture list records operation display picture img2.1 corresponding to operation sub-step 2.1 and operation display picture img2.2 corresponding to operation sub-step 2.2.

[0208] In the embodiment, the mapping relationship between the operation description information and the operation display picture can be recorded concisely by setting the table, and the operation display picture corresponding to the operation description information can be quickly found through the table, thereby helping to shorten the time consumption for generating the reply answer based on the operation description information and the operation display picture.

[0209] After obtaining the intermediate answer of the first question, the server needs to label the operation display picture using the operation description information, so as to obtain a labeled picture combined with text and picture. In this process, the operation description information can be labeled at any position in the operation display picture, or the labeling position of the operation description information in the operation display picture can be determined through text recognition, and the operation description information is displayed at the labeling position. The following several embodiments will introduce the method for labeling the operation description information in the operation display picture.

[0210] FIG. 11 is a flowchart of a synthesis labeled picture provided by an embodiment of the application. As shown in FIG. 11, in some implementations, S640: the server generates at least one labeled picture according to the intermediate answer of the first question, which can include the following sub-steps, and the execution subject of the following sub-steps is the server. Sub-step 641: according to the intermediate answer of the first question, at least one operation description information and at least one operation display picture are obtained; sub-step 643: for each operation description information in the at least one operation description information, text recognition is performed on the operation display picture corresponding to the operation description information to obtain a text set of the operation display picture; sub-step 645: first text information matching the operation description information is determined from the text set; sub-step 647: according to the display position of the first text information in the operation display picture, the labeling position of the operation description information is determined; sub-step 649: the operation description information is displayed on the operation display picture according to the labeling position, and a labeled picture is obtained.

[0211] In some embodiments, the text set of a certain operation display picture includes the characters displayed in the operation display picture. The text set of the operation display picture is obtained by performing text recognition on the operation display picture. The process of obtaining the text set will be described in the following embodiments.

[0212] The first text information is related to the operation object involved in the operation step. The first text information can be used to describe the operation object. For example, for the operation description information 1, the operation description information 1 corresponds to the operation step 1 and the operation display picture 1, and the first text information 1 determined from the operation display picture 1 is used to describe the operation object involved in the operation step 1.

[0213] Optionally, the first text information comprises attribute information of the operation object. The attribute information of the operation object comprises at least one of the following: a name of the operation object, an icon of the operation object, an explanation of a function of the operation object, etc. For example, the first question is an operation-related question related to an application program, and the operation object can be a control. The first text information used for describing the control comprises a name of the control and an explanation of a function of the control.

[0214] For example, the operation object is a button, and the name of the button is displayed on the button. The display position of the name of the button in the operation display picture overlaps with the display position of the button in the operation display picture. Therefore, by the display position of the text information in the operation display picture, the server can determine the display position of the operation object in the operation display picture, and further determine the marking position of the operation instruction information in the area close to the operation object.

[0215] In the embodiments of the present application, by text recognition and semantic matching, the display area of the operation object in the operation display picture is located, so that when the operation instruction information is marked on the operation display picture, the operation instruction information is marked near the operation object, which helps to improve the prominence of the operation object and helps to improve the speed of the user observing the operation object in the marked picture. Compared with independently displaying the operation instruction information and the operation display picture, marking the operation instruction information near the operation object in the operation display picture helps to improve the efficiency of the user reading the reply answer.

[0216] In the embodiments of the present application, by text recognition and semantic matching, the display area of the operation object in the operation display picture is located, so that when the operation instruction information is marked on the operation display picture, the operation instruction information is marked near the operation object, which helps to improve the prominence of the operation object and helps to improve the speed of the user observing the operation object in the marked picture. Compared with independently displaying the operation instruction information and the operation display picture, marking the operation instruction information near the operation object in the operation display picture helps to improve the efficiency of the user reading the reply answer.

[0217] In some embodiments, in substep 643, for each operation instruction information in the at least one operation instruction information, the server identifies the operation display picture corresponding to the operation instruction information, obtains a text set of the operation display picture, comprising: performing text area recognition on the operation display picture to obtain at least one text area information, wherein the text area information is used to represent the position and size of the text display area, and the text display area refers to the picture area displaying text in the operation display picture; for each text display area in the at least one text area information, performing text recognition on the text display area to determine the text information displayed in the text display area, and after completing the text recognition on each text display area, generating a text set based on the obtained multiple text information.

[0218] The text display area can be understood as a text box identified from the operation display picture. Optionally, the text display area is a rectangular area. For example, the text display area is a minimum circumscribed rectangular area surrounding a piece of text in the operation display picture. Illustratively, the text area information includes coordinates of the text display area and an arrangement serial number of the text display area. The coordinates of the text display area can be represented by a lower-left vertex and an upper-right vertex of the text display area. For example, the coordinates of a certain text display area are represented as (x1, y1), (x2, y2).

[0219] Illustratively, the arrangement serial number of the text display area is related to the arrangement of the plurality of text display areas in the operation display picture. For example, in the case where the operation display picture includes a plurality of text display areas, the server determines the arrangement serial number of each of the plurality of text display areas in a manner from top to bottom and from left to right.

[0220] Optionally, the step of generating the text set is completed by a question identification and processing module in the server, the question identification and processing module including a text detection model. The text detection model is used to predict at least one text area information based on the operation display picture by inputting the operation display picture into the text detection model. Optionally, the text detection model belongs to a deep learning model, and the text detection model includes but is not limited to a Textboxes model, a Textboxes++ model, and the like.

[0221] FIG. 12 is a structural schematic diagram of a text detection model provided by an embodiment of the present application. As shown in FIG. 12, the server processes the operation display picture by using the text detection model to obtain the coordinates of the text display area. Specifically, a backbone structure of a visual geometry group (VGG-16) in the text detection model is used to extract features of the operation display picture to obtain a plurality of feature information with different feature dimensions. The plurality of feature information correspond to different receptive fields, respectively. For example, the plurality of feature information are output by different convolution layers, and the plurality of convolution layers are arranged in series. The receptive field of the feature information output by a high-level convolution layer is greater than or equal to the receptive field of the feature information output by a low-level convolution layer.

[0222] Exemplarily, as shown in FIG. 12, the plurality of feature information includes: feature information obtained by performing convolution processing on the third layer of the fourth convolution block (conv4_3), the dimension of the feature information being 512, feature information obtained by performing convolution processing on the fifth convolution block, the dimension of the feature information being 1024, and the like. After obtaining the plurality of feature information, the prediction layer in the text detection model predicts candidate text display areas with a high probability of existing text, removes part of the candidate text display areas based on the overlapping degree between the plurality of candidate text display areas, and determines at least one text display area; and determines the coordinates of each text display area according to the distribution of the text display areas on the operation display picture.

[0223] In the problem identification and processing module, the at least one candidate text display area can be removed based on the overlapping degree between the plurality of candidate text display areas by using a non-maximum suppression (NMS) algorithm.

[0224] It should be understood that the convolution blocks and the number of convolution layers for generating the feature information in the above embodiments are only used as examples to introduce how the text detection model obtains the plurality of feature information, and do not limit the convolution layer outputting the feature information in the text detection model.

[0225] FIG. 13 is a schematic diagram of the effect of text area identification according to an embodiment of the present application. As shown in FIG. 13, by implementing text area identification on the operation display picture, the server determines at least one text display area (such as the position of the text display area "1310" in FIG. 13), and determines the arrangement serial number of each text display area according to the arrangement of the plurality of text display areas on the operation display picture. For example, as shown in FIG. 13, the arrangement serial number of the text display area "1310" is "0" as shown in 1320 in FIG. 13.

[0226] In some embodiments, the problem identification and processing module further includes a text recognizer, which is configured to identify the picture in the text display area to obtain the text information displayed in the text display area. The text recognizer is based on optical character recognition (OCR) technology. Optionally, the server pre-processes the picture of the text display area by using the text recognizer to obtain a pre-processed picture, performs character positioning segmentation and character recognition on the pre-processed picture, and obtains the text information displayed in the text area.

[0227] Exemplarily, the pre-processing includes at least one of the following: color image graying, binarization processing, image change angle detection, and correction processing. The pre-processing is used to correct the picture deflection and distortion of the operation display. Exemplarily, the character positioning and segmentation is used to perform horizontal projection on the picture of the text display area, to determine the upper limit and lower limit of each line of text; based on the upper limit and lower limit, at least one line of text is cut from the picture of the text display area; for each line of text in the at least one line of text, vertical projection is performed on the line of text to determine the left boundary and right boundary of each character included in the line of text, and based on the left boundary and right boundary of each character, an independent character is cut from the line of text. The horizontal projection and vertical projection operations can be realized by calling the function functions in the open source computer vision library (OPENCV).

[0228] In some embodiments, the sub-step 645 of determining the first text information matching the operation description information from the text set can be implemented as: performing component analysis on the operation description information to determine the verb component and the noun component in the operation description information; then filtering out the verb component in the operation description information to obtain the to-be-matched information, wherein the to-be-matched information includes the noun component in the operation description information; calculating the matching result between the to-be-matched information and the text information included in the text combination; and determining the first text information matching the operation description information from the text set based on the matching result. The method of determining the matching result is the same as the method of determining the matching result between the query information and the candidate text content in the above embodiments, and details are referred to the above embodiments.

[0229] FIG. 14 is a schematic diagram of a method for determining the first text information provided in the embodiments of the present application. As shown in FIG. 14, the operation description information is analyzed by the question classification and processing model: selecting connecting to the Internet, to obtain the verb component “select” and the noun component “connecting to the Internet”; the electronic device takes “connecting to the Internet” as the to-be-matched information, calculates the matching result between the to-be-matched information and the text information included in the text combination through the semantic matching model; and determines the first text information matching the operation description information from the text set based on the matching result. In this embodiment, the first text information is “connecting to the Internet (C) connecting to the network of my workplace (O) network connection type”.

[0230] In the embodiment, by filtering the verb components in the operation description information, on one hand, shorter character length of the to-be-matched information is obtained, thereby helping to reduce the calculation amount in the semantic matching process; on the other hand, the filtered verbs usually have no semantic association with the operation objects, and removing the verbs in the operation description information is equivalent to removing the irrelevant interference information in the operation description information, thereby helping to improve the accuracy of the matched text information.

[0231] In one example, the semantic matching model is included in the question classification and processing module, and the first text information matched with the operation description information is determined by the semantic matching model. Exemplarily, the model structures of the semantic matching models respectively included in the reply generation module and the question classification and processing module are the same, the two semantic matching models are trained by different training data, and the model parameters contained in the two semantic matching models are different.

[0232] In some examples, the total number of operation description information included in the intermediate answer of the first question and the total number of operation display pictures can be the same or different. In order to reduce the calculation pressure of the server, before step 640, the operation display pictures can be deduplicated to reduce the number of operation display pictures that need to be processed by the server in step 640. Alternatively, if a plurality of operation description information with adjacent execution serial numbers correspond to the same operation display picture, the server deduplicates the plurality of operation display pictures, and finally only retains one, so that the plurality of operation description information share one operation display picture.

[0233] Exemplarily, for the operation display picture shared by m operation description information, the operation display picture is stored only once in the server; after the text set of the operation display picture is determined for the first time, the server determines m corresponding first text information in the text set, and m is a positive integer.

[0234] Alternatively, the server calculates the similarity values respectively possessed by a plurality of operation display pictures; if the similarity values of two operation display pictures are greater than or equal to a similarity threshold value, the two operation display pictures are determined to be the same operation display picture; if the similarity values of two operation display pictures are less than the similarity threshold value, the two operation display pictures are determined to be different operation display pictures. The similarity threshold value is pre-set, for example, the similarity threshold value is equal to 0.96.

[0235] Alternatively, after the picture deduplication is completed, the server determines whether at least two operation display pictures in the plurality of operation display pictures are completely the same operation display picture by the picture identifiers of the operation display pictures. Exemplarily, if the picture identifiers of two operation display pictures are consistent, the two operation display pictures are completely the same; if the picture identifiers of two operation display pictures are inconsistent, the two operation display pictures are different.

[0236] In the process of labeling operation demonstration pictures with operation instruction information, text recognition and other operations need to be performed on the operation demonstration pictures. By determining the similarity between multiple operation demonstration pictures, only one operation demonstration picture is retained for multiple identical operation demonstration pictures, which helps to avoid multiple image recognition of the same operation demonstration picture, thereby reducing the computing pressure of the server, and further helps to shorten the time consumption in generating a reply answer, so as to improve the response speed of the dialogue reply system to inquiry information.

[0237] In some embodiments, before performing step 640, the server further includes: determining type information of the first question based on the inquiry information; in a case where the type information of the first question meets a processing condition, performing step 640; in a case where the type information of the first question does not meet the processing condition, determining the intermediate answer of the first question as the reply answer of the first question; and sending the reply answer to the electronic device. Optionally, the processing condition is used to limit the question type of the first question.

[0238] For example, the processing condition is that the first question belongs to an operation type question or the first question belongs to an object recognition question.

[0239] As shown in FIG. 15, the question classification and processing module includes a question classification model. Before generating at least one labeled picture based on the intermediate answer of the first question, the question classification and processing module first determines the question type to which the first question belongs based on the inquiry information through the question classification model. Optionally, the question classification model is a binary classification model, the input information of the question classification model is the inquiry information, and the output result of the question classification model is a classification result. The classification result includes two cases. If the first question belongs to an operation type question or an object recognition question, the question classification model outputs a first answer, otherwise the question classification model outputs a second answer. In a case where the classification result is the first answer, the server performs step 640 based on the intermediate answer.

[0240] After step S640 is performed, the server obtains at least one labeled picture. In subsequent step S650, the server generates the reply answer of the first question based on the at least one labeled picture. The following describes the process through several embodiments.

[0241] S650: The server generates the reply answer of the first question according to the at least one labeled picture.

[0242] In some embodiments, the reply answer to the first question refers to the solution to the first question fed back to the user after the user proposes the first question. The reply answer to the first question and the intermediate answer to the first question are used to describe the same solution, but the reply answer to the first question and the intermediate answer to the first question describe the solution to the first question in different ways. Among them, the reply answer to the first question describes the solution to the first question in a combination of text and pictures, the intermediate answer to the first question only includes operation instruction information, or includes operation instruction information and its corresponding operation display picture, but the operation instruction information and the operation display picture are independent of each other.

[0243] In some embodiments, when the server generates multiple labeled pictures, each of the multiple labeled pictures has a feedback serial number, and the feedback serial number is used to determine the display order of the multiple labeled pictures on the electronic device. Multiple refers to two or more. For example, the smaller the feedback serial number, the earlier the display order of the labeled picture; the larger the feedback serial number, the later the display order of the labeled picture. The display order is manifested as the appearance order of the labeled picture in the dialogue reply interface, or the arrangement order of the multiple labeled pictures in the dialogue reply interface.

[0244] It can be understood that since the labeled picture is obtained by labeling the operation instruction information in the operation display picture, the execution step corresponding to the labeled picture is the same as the operation step corresponding to the operation display picture used to generate the labeled picture, and therefore the feedback serial number of the labeled picture is the same as the execution serial number of the operation display picture used to generate the labeled picture. For example, the execution serial number of an operation demonstration picture is 1, and the feedback serial number of the labeled picture generated by the operation demonstration picture is also 1.

[0245] Optionally, for the case where the operation step includes at least two operation sub-steps, the at least two operation sub-steps have a relative execution order within the operation step, and have independent operation instruction information and operation display pictures, and each operation sub-step has a corresponding labeled picture. The operation sub-step can be understood as an operation step.

[0246] In some embodiments, the reply answer to the first question can be a picture combination including the at least one labeled picture. For example, in this case, after generating the at least one labeled picture, the server sends the at least one labeled picture to the electronic device as the reply answer to the first question.

[0247] Optionally, if there is operation instruction information in the intermediate answer that does not correspond to the operation demonstration picture, the server generates a simulation picture that displays the operation instruction information, and takes the simulation picture as the label picture corresponding to the operation instruction information. For example, the server labels the operation instruction information on a background picture to obtain a simulation picture. The background picture can be pre-set. The operation instruction information can be labeled at any position in the background picture.

[0248] In some other embodiments, in the case where there are multiple label pictures, the reply answer to the first question is an operation demonstration video generated according to the multiple label pictures. Optionally, the server determines the playing time period or the number of continuous frames (i.e., the video frames in which the same label picture is continuously displayed) of the multiple label pictures in the operation demonstration video according to the feedback serial numbers of the multiple label pictures. During the playing process of the operation demonstration video, any label picture in the multiple label pictures is continuously displayed in its playing time period. For example, the multiple label pictures include label picture 1, label picture 2 and label picture 3, the playing time period of label picture 1 is [0-5) seconds, the playing time period of label picture 2 is [5-10) seconds, and the playing time period of label picture 3 is [10-15) seconds, i.e., label picture 1 is displayed in [0-5) seconds, label picture 2 is displayed in [5-10) seconds, and label picture 3 is displayed in [10-15) seconds during the playing process of the operation demonstration video.

[0249] For example, an operation demonstration picture needs to be displayed for two seconds in the operation demonstration video, and the frame rate of the operation demonstration video is 60 frames, so the length of the playing time period of the operation demonstration picture is two seconds, and the number of continuous frames of the operation demonstration picture is 120 frames.

[0250] In the embodiments of the present application, the multiple label pictures are integrated into an operation demonstration video, so that the display area occupied by the multiple label pictures is converted into the time length consumed by playing the operation demonstration video. Compared with directly displaying the multiple label pictures and / or a large amount of text information, the operation demonstration video occupies a smaller area. In the conversation reply scenario, simultaneously displaying the multiple label pictures and / or a large amount of text information can cause the label picture with the earliest display order to disappear from the conversation reply page, and the user often needs to drag a window slider or the like to search upwards in the multiple label pictures to find the label picture with the earliest display order, which is a complex user operation. In the case of the operation demonstration video, the user can directly watch the label picture with the earliest display order by playing the operation demonstration video, so as to refer to the label picture to perform the first operation step for solving the problem, which helps to simplify the user operation.

[0251] S660: The server sends the reply answer to the first question to the electronic device.

[0252] Optionally, in the case that the reply answer of the first question is a picture set including at least one labeled picture, the server sends the at least one labeled picture to the electronic device; in the case that the reply answer of the first question is an operation demonstration video, the server sends the operation demonstration video to the electronic device.

[0253] Exemplarily, in the case that the reply answer of the first question is an operation demonstration video, the server can also send at least one labeled picture to the electronic device. In this case, the electronic device plays the labeled pictures in the conversation reply interface according to the corresponding playing time period of each labeled picture, achieving the same effect as playing the operation demonstration video.

[0254] S670: The electronic device displays the reply answer of the first question.

[0255] In some embodiments, the electronic device displays the reply answer of the first question in the conversation reply interface. Optionally, in the case that the representation form of the reply answer of the first question is at least one labeled picture, the electronic device displays the at least one labeled picture through a chat bubble. Exemplarily, the at least one labeled picture is displayed in the same chat bubble, or each of the at least one labeled picture is displayed in a chat bubble.

[0256] In other embodiments, in the case that the reply answer of the first question is an operation demonstration video, the electronic device displays the operation demonstration video and a video playing control in the conversation reply interface, and the electronic device plays the video demonstration operation through a video player in response to an operation of triggering the video playing control.

[0257] Optionally, during the playing of the operation demonstration video, the electronic device also displays a playing progress bar of the operation demonstration video, and the playing progress bar includes at least one picture switching identifier, which is used to represent the starting playing time of the labeled picture corresponding to the operation step in the operation demonstration video. Exemplarily, the display form of the picture switching identifier includes but is not limited to an icon and a number.

[0258] For example, a plurality of icons are displayed on the playing progress bar, and the playing time corresponding to each labeled picture on the playing progress bar is the starting playing time of the labeled picture. For example, a plurality of numbers are displayed on the playing progress bar, and the playing time corresponding to the number i on the playing progress bar is the starting playing time of the labeled picture i, where i is a natural number.

[0259] In the embodiment of the present application, after obtaining the intermediate answer of the problem to be solved based on the inquiry information of the problem to be solved, the intermediate answer is not directly fed back to the user as the reply answer of the problem to be solved, but the operation instruction information used to describe the operation steps and the operation display information used to show the operation effect of the operation steps are combined based on the intermediate answer; the operation instruction information marked in the operation display picture is generated as a marked picture, and then the reply answer of the problem to be solved is obtained.

[0260] Compared with the automatic reply method in the related art, the embodiment combines the operation instruction information and the operation display picture by generating the marked picture, can more intuitively provide the reference solution of the problem to be solved, does not need the user to first understand the text meaning and then determine the operation object by referring to the picture, helps to improve the probability that the user successfully handles the problem to be solved through one question, thereby reducing the user's multiple questions in the dialogue reply system for the same problem, and achieving the effect of improving the efficiency of dialogue reply.

[0261] FIG. 16 is an interactive timing diagram of a reply generation method provided by an embodiment of the present application. The dialogue reply function of the server is cooperatively implemented by a reply generation module, a question identification module and a plug-in engine module. In the embodiment, the reply generation method is introduced by taking the generation of the reply answer of the operation type question as an example. As shown in FIG. 16, the dialogue reply system generates the reply answer including the following steps:

[0262] S1601: The electronic device sends inquiry information describing a first question to the server. Optionally, the inquiry information can be input by the user on the electronic device, or provided to the electronic device through copying and pasting or other means. Optionally, the reply generation module in the server obtains the inquiry information through an information acquisition channel.

[0263] S1602: The reply generation module obtains an intermediate answer of the first question according to the inquiry information.

[0264] Optionally, the reply generation module uses an NLP semantic matching method or a method based on RAG to obtain the intermediate answer of the first question. For example, the inquiry information is: “How to set up a dial-up program” and the intermediate answer content obtained by the reply generation module is as follows:

[0265] “According to your question, the answer generated for you is as follows:

[0266] Step 1: Click “Start” on the desktop -> “Program” -> “Accessories” -> “Communication” -> “New Connection Wizard”, enter the “New Connection Wizard” interface as shown in the following figure, select “Connect to Internet”, and click “Next”;

[0267] Step 2.1: Select "Manually set up my connection" and click "Next";

[0268] Step 2.2: Select "Broadband connection that requires a user name and password (Point-to-Point Protocol over Ethernet, PPPOE) to connect" and click "Next";

[0269] Step 3: Fill in the connection name, for example, "Fiber Broadband", and click "Next";

[0270] Step 4.1: Fill in the account name in the "User Name" field, which is in the format "139xxxxxxxx", and fill in the customer password in the "Password" and "Confirm Password" fields. Note that the username and password are case-sensitive, and the information entered must be correct, otherwise you will not be able to log in successfully. Then, you can choose whether to check the options according to your needs, and click "Next";

[0271] Step 4.2: Check "Add a shortcut to this connection on my desktop", so that there will be a dial-up connection icon called "Fiber Broadband" on your computer desktop, and click "Finish" to complete the installation.

[0272] S1603: The reply generation module requests the question identification and processing module to identify the question type of the first question.

[0273] Optionally, the reply generation module carries the inquiry information and the intermediate answer in the request sent to the question identification and processing module. Illustratively, in step S1603, the reply generation module sends the inquiry information to the question identification and processing module; if in step S1604, the question identification and processing module determines that the first question is an operation type question, the reply generation module sends the intermediate answer to the question identification and processing module. Otherwise, the reply generation module does not need to send the intermediate answer to the question identification and processing module.

[0274] S1604: The question identification and processing module determines whether the first question belongs to the operation type question according to the inquiry information.

[0275] Optionally, the question identification and processing module inputs the inquiry information into the question classification model, and determines the question type to which the first question belongs based on the classification result output by the type classification model. Illustratively, the identification result of the type classification model can be represented by 1bit character; if the classification result is 1, it means that the first question belongs to the operation type question; if the classification result is 0, it means that the first question belongs to the non-operation type question.

[0276] Optionally, if the first question belongs to the operation type, step S1605 is performed, otherwise, the intermediate answer is taken as the reply answer of the first question, and the intermediate answer is directly returned to the electronic device.

[0277] S1605: The question recognition and processing module determines at least one operation instruction information and at least one operation display picture based on the intermediate answer.

[0278] Optionally, the question recognition and processing module generates a step text list for storing the at least one operation instruction information and a picture list for storing the at least one operation display picture. The step text list and the picture list are used to represent the corresponding relationship between the operation instruction information and the operation display picture. For details of the step text list and the picture list, please refer to the above embodiments, which will not be described here.

[0279] S1606: The question recognition and processing module requests the plug-in engine module to find the operation display picture corresponding to the second operation instruction information.

[0280] The second operation instruction information is the operation instruction information for which the operation display picture is missing. It can be understood that step S1606 is an optional step. In the case where the operation display picture is missing, the server can perform step S1606 to obtain the second operation display picture corresponding to the second operation instruction information through the plug-in engine module.

[0281] Optionally, the question recognition and processing module sends the second operation instruction information to the plug-in engine module and receives the second operation display picture returned by the plug-in engine module.

[0282] S1607: The question recognition and processing module obtains a marked picture according to the operation instruction information and the operation display picture.

[0283] Optionally, for any one operation instruction information, the question recognition and processing module determines the marking position of the operation instruction information in the operation display picture based on the display position of the operation object corresponding to the operation step in the operation display picture; and obtains a marked picture by marking the operation instruction information into the corresponding operation display picture. For details, please refer to the above embodiments.

[0284] S1608: The question recognition and processing module generates a reply answer according to the marked picture.

[0285] In some embodiments, the reply answer includes at least one marked picture, and each marked picture corresponds to at least one operation instruction information.

[0286] Optionally, the server sends a picture set composed of the at least one annotated picture as the reply answer. For example, the reply answer includes multiple annotated pictures. Correspondingly, as shown in FIG. 17, the subsequent electronic device can display the multiple annotated pictures in the dialogue reply interface after obtaining the reply answer.

[0287] Optionally, after generating the annotated pictures, the question identification and processing module generates an operation demonstration video or assembles a hyper text markup language (HTML) page according to the at least one annotated picture, to obtain the operation demonstration video. For example, the reply answer includes the operation demonstration video, and the operation demonstration video is played in sequence by the multiple annotated pictures. Correspondingly, as shown in FIG. 18, the subsequent electronic device can display the operation demonstration video in the dialogue reply interface after obtaining the reply answer. Optionally, in response to the operation of instructing the user to play the operation demonstration video, the operation demonstration video can also be played.

[0288] S1609: The server sends the reply answer to the electronic device.

[0289] S1610: The electronic device displays the reply answer.

[0290] Optionally, the electronic device displays the dialogue reply interface, and after receiving the reply answer, the electronic device displays the reply answer in the dialogue reply interface. For example, the dialogue reply interface includes a video reply control, and the electronic device requests the server to generate an operation demonstration video for solving the first question in response to an operation of triggering the video reply control. The operation of triggering the video reply control includes, but is not limited to, at least one of the following: a click operation, a touch operation, a sliding operation, and the like. The operation of triggering the video reply control can be regarded as a style adjustment operation, and the style prompt information includes a display effect of the video reply control.

[0291] In the embodiments of the present application, after obtaining the intermediate answer of the to-be-solved question based on the inquiry information of the to-be-solved question, the intermediate answer is not directly fed back to the user as the reply answer of the to-be-solved question, but the operation instruction information for describing the operation steps and the operation display information for showing the operation effect of the operation steps are combined based on the intermediate answer. The operation instruction information marked in the operation display picture is annotated to generate an annotated picture, and then the reply answer of the to-be-solved question is obtained.

[0292] Compared with the automatic reply method in the related art, the embodiment can more intuitively provide a reply answer about the problem to be solved by generating a marked picture to combine operation description information and an operation display picture, the reply answer does not require the user to first understand the text meaning and then determine the operation object by referring to the picture, so that the reply answer is easier to understand, thereby helping to improve the probability of successfully processing the problem to be solved by the user through one question, thereby reducing the user's multiple questions in the dialogue reply system for the same problem, and achieving the effect of improving the efficiency of dialogue reply.

[0293] The reply generation method provided by the embodiment of the application is described in detail above in combination with FIGS. 6-18. The server and the electronic device provided by the embodiment of the application are described in detail below in combination with FIGS. 19 and 20.

[0294] All or part of any feature of any embodiment in the application can be freely combined. The combined technical solution also belongs to the scope described in the application.

[0295] In a possible design, FIG. 19 is a structural schematic diagram of a server provided by the embodiment of the application. As shown in FIG. 19, the server 1900 can include a transceiver unit 1901 and a processing unit 1902. The server 1900 can be used to implement the functions of the server involved in the above method embodiments.

[0296] Optionally, the transceiver unit 1901 is configured to support the server 1900 to perform S620 and S660 in FIG. 6; and / or is configured to support the server 1900 to perform S1602 and S1609 in FIG. 16.

[0297] Optionally, the processing unit 1902 is configured to support the server 1900 to perform S630 to S650 in FIG. 6; and / or is configured to support the server 1900 to perform S1603 to S1608 in FIG. 16.

[0298] The transceiver unit can include a receiving unit and a sending unit, can be implemented by a transceiver or a transceiver related circuit component, and can be a transceiver or a transceiver module. The operations and / or functions of each unit in the server 1900 are respectively used to implement the corresponding processes of the reply generation method described in the above method embodiments. All related contents of each step involved in the above method embodiments can be referred to the function description of the corresponding functional unit, and will not be described herein for brevity.

[0299] Optionally, the server 1900 shown in FIG. 19 can further include a storage unit (not shown in FIG. 19), and the storage unit stores programs or instructions. When the transceiver unit 1901 and the processing unit 1902 execute the programs or instructions, the server 1900 shown in FIG. 19 can execute the reply generation method described in the above method embodiments.

[0300] The technical effects of the server 1900 shown in FIG. 19 can refer to the technical effects of the reply generation method described in the above method embodiments, which are not described here again.

[0301] In addition to the form of the server 1900, the technical solutions provided in the present application can also be a functional unit or a chip in the server, or a device matched with the server.

[0302] The embodiments of the present application also provide a chip system, comprising: a processor coupled with a memory, the memory being used to store programs or instructions, when the programs or instructions are executed by the processor, the chip system implements the method in any of the above method embodiments.

[0303] In a possible design, FIG. 20 is a structural schematic diagram of an electronic device provided by the embodiments of the present application. As shown in FIG. 20, the electronic device 2000 can include: a transceiver unit 2001 and a processing unit 2002. The electronic device 2000 can be used to implement the functions of the electronic device involved in the above method embodiments.

[0304] Optionally, the transceiver unit 2001 is configured to support the electronic device 2000 to perform S620 and S660 in FIG. 6; and / or, is configured to support the electronic device 2000 to perform S1602 and S1606 in FIG. 16.

[0305] Optionally, the processing unit 2002 is configured to support the electronic device 2000 to perform S670 in FIG. 6; and / or, is configured to support the electronic device 2000 to perform S1601 and S1610 in FIG. 16.

[0306] The transceiver unit can include a receiving unit and a sending unit, can be implemented by a transceiver or a transceiver related circuit component, and can be a transceiver or a transceiver module. The operations and / or functions of each unit in the electronic device 2000 are respectively used to implement the corresponding procedures of the reply generation method described in the above method embodiments. All related contents of each step involved in the above method embodiments can be referred to the function description of the corresponding functional unit, which is not described here again for brevity.

[0307] Optionally, the electronic device 2000 shown in FIG. 20 can further include a storage unit (not shown in FIG. 20), which stores programs or instructions. When the transceiver unit 2001 and the processing unit 2002 execute the programs or instructions, the electronic device 2000 shown in FIG. 20 can execute the reply generation method described in the above method embodiments.

[0308] The technical effects of the electronic device 2000 shown in FIG. 20 can refer to the technical effects of the reply generation method described in the above method embodiments, which are not described here again.

[0309] In addition to the form of the electronic device 2000, the technical solutions provided in the present application can also be a functional unit or a chip in the electronic device, or a device matched with the electronic device.

[0310] Optionally, the processor in the chip system can be one or more. The processor can be implemented by hardware or software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented by software, the processor can be a general-purpose processor, which is implemented by reading software codes stored in a memory.

[0311] Optionally, the memory in the chip system can also be one or more. The memory can be integrated with the processor, or can be separately arranged from the processor, and the embodiments of the present application do not limit the same. Illustratively, the memory can be a non-transient processor, for example, a read-only memory (ROM), which can be integrated on the same chip as the processor, or can be separately arranged on different chips, and the embodiments of the present application do not limit the type of the memory and the arrangement manner of the memory and the processor.

[0312] Illustratively, the chip system can be a field programmable gate array (FPGA), can be an application specific integrated chip (ASIC), can be a system chip (SoC), can be a central processor unit (CPU), can be a network processor (NP), can be a digital signal processing circuit (DSP), can be a micro controller unit (MCU), can be a programmable logic device (PLD), or can be other integrated chips.

[0313] It should be understood that each step in the above method embodiments can be completed by the integrated logic circuit of hardware in the processor or the instruction in the form of software. The method steps disclosed in conjunction with the embodiments of the present application can be directly embodied as hardware processor execution completion, or executed by a combination of hardware and software modules in the processor.

[0314] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program. When the computer program is run on a computer, the computer is caused to execute the related steps to realize the reply generation method in the above embodiment.

[0315] The embodiment of the present application further provides a computer program product, which causes a computer to execute the related steps to realize the reply generation method in the above embodiment when the computer program product is run on the computer.

[0316] In addition, the embodiment of the present application further provides an apparatus. The apparatus can be specifically a component or a module, and the apparatus can include one or more processors and memories connected thereto. The memories are configured to store computer programs. When the computer programs are executed by the one or more processors, the apparatus is caused to execute the reply generation method in the above method embodiments.

[0317] The apparatus, the computer readable storage medium, the computer program product or the chip provided by the embodiment of the present application are used to execute the corresponding method provided above. Therefore, the beneficial effects achieved thereby can refer to the beneficial effects in the corresponding method provided above, which will not be described herein again.

[0318] The steps of the method or the algorithm described in connection with the disclosure of the embodiment of the present application can be implemented in the form of hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in a random access memory, a flash memory, a read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, a register, a hard disk, a mobile hard disk, a read-only optical disc or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in a special integrated circuit.

[0319] From the above description of the embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration. In actual application, the above functions can be completed by different functional modules according to needs; that is, the internal structure of the apparatus is divided into different functional modules to complete all or part of the functions described above. The specific working process of the above-described system, apparatus and unit can refer to the corresponding process in the foregoing method embodiments, which will not be described herein again.

[0320] In several embodiments provided in the present application, it should be understood that the disclosed method can be implemented in other ways. The above-described device embodiments are only illustrative. For example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, modules or units, which can be electrical, mechanical or other forms.

[0321] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit. The computer readable storage medium includes, but is not limited to, any one of the following: a U disk, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk or an optical disk, and various media that can store program codes.

[0322] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any change or replacement within the technical scope disclosed in the present application should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A reply generation method characterized by comprising: Applied to a server, the method comprises: Obtaining inquiry information describing a first question; According to the inquiry information, obtaining an intermediate answer of the first question, the intermediate answer including at least one operation description information and at least one operation display picture, the operation description information being used to describe an operation step for solving the first question, and the operation description information and the operation display picture having a corresponding relationship; According to the operation description information, marking the operation display picture to obtain at least one marked picture; According to the at least one marked picture, generating a reply answer of the first question.

2. The method of claim 1, wherein, The operation description information is marked on the operation display picture to obtain at least one marked picture, comprising: For a first operation display picture in the at least one operation display picture, according to the first operation display picture and first operation description information, determining a marking position of the first operation description information, the first operation display picture being any one of the at least one operation display picture, and the first operation description information and the first operation display picture having a corresponding relationship; According to the marking position of the first operation description information, marking the first operation description information in the first operation display picture to obtain a first marked picture.

3. The method of claim 2, wherein, According to the first operation display picture and the first operation description information, the marking position of the first operation description information is determined, comprising: Text recognition is performed on the first operation display picture to obtain a text set of the first operation display picture, the text set including at least one candidate text information displayed in the first operation display picture; From the at least one candidate text information, determining first text information matched with the first operation description information; According to the display position of the first text information in the first operation display picture, determining the marking position of the first operation description information.

4. The method according to any one of claims 1 to 3, characterized in that, The method further comprises: Performing sentence component analysis on the operation description information to obtain a component analysis result, the component analysis result being used to represent a sentence component to which a word included in the operation description information belongs; According to the component analysis result, reducing the operation description information to obtain reduced operation description information, the reduced operation description information being used to obtain the marked picture.

5. The method according to any one of claims 1 to 4, characterized in that, According to the at least one marked picture, the method further comprises: According to the at least one marked picture, generating an operation demonstration video, the operation demonstration video being used to display the at least one marked picture in an execution order of the execution steps; Taking the operation demonstration video as the reply answer of the first question.

6. The method according to any one of claims 1 to 5, characterized in that, The reply answer of the first question includes the at least one marked picture.

7. The method according to any one of claims 1 to 6, characterized in that, Before the operation description information is marked on the operation display picture to obtain at least one marked picture, the method further comprises: For second operation description information in the intermediate answer of the first question that does not correspond to an operation display picture, determining a second operation display picture based on the second operation description information; Add the second operation display picture to the intermediate answer of the first question to obtain an updated intermediate answer, and the updated intermediate answer is used to generate the at least one labeled picture.

8. A method of replying to a conversation, the method comprising: The method is applied to an electronic device, and the method comprises: displaying a dialogue reply interface, the dialogue reply interface being used for dialogue interaction of at least two dialogue subjects; receiving a first question input by a user in the dialogue reply interface; sending inquiry information corresponding to the first question to a server; displaying a reply answer of the first question sent by the server, the reply answer of the first question comprising at least one labeled picture, the at least one labeled picture being obtained by labeling a corresponding operation display picture by operation instruction information, the operation instruction information being used for describing operation steps for solving the first question.

9. The method of claim 8, wherein, The answer style of the reply answer of the first question comprises at least one of the following: a picture set composed of the at least one labeled picture, and / or an operation demonstration video used for playing the at least one labeled picture.

10. The method of claim 9, wherein, The method further comprises: in response to a style adjustment operation, displaying adjusted style prompt information, the style prompt information being used for prompting an answer style adopted by a subsequent reply answer; in response to a dialogue request operation, displaying the reply answer of the first question, comprising displaying a reply answer with a style conforming to the displayed style prompt information.

11. A reply generation apparatus characterized by comprising: The device comprises: an information acquisition module configured to acquire inquiry information describing a first question; an answer acquisition module configured to acquire an intermediate answer of the first question according to the inquiry information, the intermediate answer comprising at least one operation instruction information and at least one operation display picture, the operation instruction information being used for describing operation steps for solving the first question, and the operation instruction information and the operation display picture having a corresponding relationship; a picture acquisition module configured to acquire at least one labeled picture by labeling the operation display picture according to the operation instruction information; an answer generation module configured to generate a reply answer of the first question according to the at least one labeled picture.

12. A dialogue reply device, characterized by The device comprises: an interface display module configured to display a dialogue reply interface, the dialogue reply interface being used for dialogue interaction of at least two dialogue subjects; a question receiving module configured to receive a first question input by a user in the dialogue reply interface; an information sending module configured to send inquiry information corresponding to the first question to a server; the interface display module is further configured to display a reply answer of the first question sent by the server, the reply answer of the first question comprising at least one labeled picture, the at least one labeled picture being obtained by labeling a corresponding operation display picture by operation instruction information, the operation instruction information being used for describing operation steps for solving the first question.

13. A server, characterized by The server comprises: a transceiver configured to send and receive radio signals; a memory configured to store computer program instructions; a processor configured to execute the computer program instructions to support the server to implement the method according to any one of claims 1 to 7.

14. An electronic device, comprising: The electronic device comprises: a display screen configured to display an interface. a transceiver for transmitting and receiving radio signals; a memory for storing computer program instructions; a processor for executing the computer program instructions to support the electronic device to implement the method of any one of claims 8 to 10.

15. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and when the computer program runs on a computer, causes the method of any one of claims 1 to 7, or executes the method of any one of claims 8 to 10.

16. A computer program product comprising instructions, characterized in that, The computer program product, when running on a computer, causes the computer to execute the method of any one of claims 1 to 7, or execute the method of any one of claims 8 to 10.

Citation Information

Patent Citations

  • Question and answer method and related device

    CN115269895A

  • Visual question and answer method and device, electronic equipment and storage medium

    CN116186315A

  • Question and answer method and question and answer model training method

    CN116561270A

  • Medical visual question and answer method and system based on cross-modal data fusion

    CN116932722A

  • Question and answer pair generation method and device and electronic equipment

    CN117421413A