Man-machine conversation method and device, electronic equipment and storage medium
By obtaining the session messages entered by the user, determining the candidate reply text and target constraint information, the problem of the language model output reply content being separated from the scene is solved, and a more accurate and effective user session experience is achieved.
Patent Information
- Application Number
- CN202311403281.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-26
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, the reply content output by the language model is prone to detachment from the user's session scenario, resulting in the user being unable to obtain the desired information, resulting in a poor user's session experience.
By obtaining the session message entered by the user, the candidate reply text and target constraint information are determined, and the target reply message is generated. The method includes converting the conversation message into a session vector, retrieving the preset vector database to obtain the associated text, filtering the candidate reply text, and generating the reply text based on the target constraint information.
Ensure that the final output target reply message does not deviate from the user's session scenario, improve the accuracy and effectiveness of the conversation, and improve the user's satisfaction with the conversation.
Smart Images

Figure CN119938813A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a method, device, electronic device and storage medium for human-computer interaction. Background Art
[0002] With the continuous development of computer technology, various language models have been born, which can complete tasks with objective answers, such as continuous dialogue and question-and-answer, and tasks without objective answers, such as writing copy. However, in some scenarios, the response content output by the language model will be out of touch with the user's conversation context, resulting in the user being unable to obtain the expected information, resulting in a poor user conversation experience. Summary of the invention
[0003] In order to overcome the problems existing in the related art, the present disclosure provides a method, device, electronic device and storage medium for human-computer dialogue.
[0004] According to a first aspect of an embodiment of the present disclosure, a method for human-computer dialogue is provided, the method comprising:
[0005] Get the conversation message entered by the user;
[0006] Determine candidate reply texts and target constraint information according to the conversation message; the target constraint information is used to indicate a scenario corresponding to the conversation message;
[0007] Outputting a target reply message corresponding to the conversation message according to the candidate reply text and the target constraint information.
[0008] Optionally, determining the candidate reply text and target constraint information according to the conversation message includes:
[0009] According to the conversation message, obtaining a plurality of associated texts related to the conversation message;
[0010] Determining the candidate reply text according to the plurality of associated texts;
[0011] The target constraint information is determined according to the candidate reply text and the conversation message.
[0012] Optionally, the associated text includes a first associated text, and acquiring a plurality of associated texts related to the conversation message according to the conversation message includes:
[0013] Converting the session message into a corresponding session vector;
[0014] According to the conversation vector, a corresponding plurality of first associated texts are obtained from a preset vector database.
[0015] Optionally, converting the session message into a corresponding session vector includes:
[0016] Determining a target semantic slot corresponding to the conversation message;
[0017] The target semantic slot is converted into a corresponding semantic slot vector, and the semantic slot vector is used as the session vector.
[0018] Optionally, the associated text further includes a second associated text, and acquiring, according to the conversation message, a plurality of associated texts related to the conversation message includes:
[0019] According to the target semantic slot, a corresponding plurality of second associated texts are acquired from a preset text database.
[0020] Optionally, the method further comprises:
[0021] Obtain historical conversations with the user within a historical period;
[0022] The acquiring, according to the conversation message, a plurality of associated texts related to the conversation message comprises:
[0023] According to the historical conversation and the conversation message, a plurality of associated texts related to the conversation message are acquired.
[0024] Optionally, the method further comprises:
[0025] Determining the user's conversation intention according to the conversation message;
[0026] Determining the candidate reply text according to the plurality of associated texts includes:
[0027] According to the conversation intention, a plurality of target texts are selected from the plurality of associated texts;
[0028] The plurality of target texts are used as the candidate reply texts.
[0029] Optionally, outputting a target reply message corresponding to the conversation message according to the candidate reply text and the target constraint information includes:
[0030] Generate a corresponding target reply text through a language model according to the candidate reply text and the target constraint information;
[0031] According to the target reply text, a target reply message corresponding to the conversation message is output.
[0032] Optionally, outputting a target reply message corresponding to the conversation message according to the target reply text includes:
[0033] Generate multiple stream texts according to the target reply text;
[0034] According to the generation time of each of the streaming texts, multiple streaming texts are output in sequence.
[0035] Optionally, outputting the plurality of streaming texts in sequence according to the generation time of each streaming text comprises:
[0036] Each time a stream text is generated, determining whether the stream text contains characters of a preset type;
[0037] In a case where the stream text does not contain characters of the preset type, the stream text is output.
[0038] Optionally, the method further comprises:
[0039] When the stream text contains characters of the preset type, caching the current stream text;
[0040] Determining whether to generate new streaming text within a preset time period;
[0041] If no new streaming text is generated within the preset time period, the cached streaming text is output; or,
[0042] Generate new streaming text within a preset time period, cache the new streaming text, and return to the step of determining whether to generate new streaming text within the preset time period when the new streaming text contains characters of a preset type.
[0043] Optionally, the method further comprises:
[0044] In the case where the new streaming text does not contain characters of the preset type, the cached streaming text is output. Optionally, if the cached streaming text includes multiple characters, in the case where the new streaming text does not contain characters of the preset type, the cached streaming text is output including:
[0045] When the new streaming text does not contain characters of the preset type, multiple cached streaming texts are concatenated to obtain the target streaming text;
[0046] Output the target streaming text.
[0047] According to a second aspect of an embodiment of the present disclosure, a device for human-computer dialogue is provided, the device comprising:
[0048] An acquisition module is configured to acquire a conversation message input by a user;
[0049] A determination module, configured to determine candidate reply texts and target constraint information according to the conversation message; the target constraint information is used to indicate a scenario corresponding to the conversation message;
[0050] The output module is configured to output a target reply message corresponding to the conversation message according to the candidate reply text and the target constraint information.
[0051] Optionally, the determination module is configured to obtain a plurality of associated texts related to the conversation message according to the conversation message; determine the candidate reply text according to the plurality of associated texts; and determine the target constraint information according to the candidate reply text and the conversation message.
[0052] Optionally, the associated text includes a first associated text, and the determining module is configured to convert the conversation message into a corresponding conversation vector; and obtain corresponding multiple first associated texts from a preset vector database according to the conversation vector.
[0053] Optionally, the determination module is configured to determine a target semantic slot corresponding to the conversation message; convert the target semantic slot into a corresponding semantic slot vector, and use the semantic slot vector as the conversation vector.
[0054] Optionally, the associated text further includes second associated text, and the determining module is configured to obtain a corresponding plurality of second associated texts from a preset text database according to the target semantic slot.
[0055] Optionally, the acquisition module is further configured to acquire historical conversations with the user within a historical period;
[0056] The determining module is configured to obtain a plurality of associated texts related to the conversation message according to the historical conversation and the conversation message.
[0057] Optionally, the determination module is configured to determine the user's conversation intention based on the conversation message; filter out multiple target texts from the multiple associated texts based on the conversation intention; and use the multiple target texts as the candidate reply texts.
[0058] Optionally, the output module is configured to generate a corresponding target reply text through a language model according to the candidate reply text and the target constraint information; and output a target reply message corresponding to the conversation message according to the target reply text.
[0059] Optionally, the output module is configured to generate a plurality of streaming texts according to the target reply text; and output the plurality of streaming texts in sequence according to the generation time of each of the streaming texts.
[0060] Optionally, the output module is configured to determine whether the streaming text contains characters of a preset type each time the streaming text is generated; if the streaming text does not contain characters of the preset type, output the streaming text.
[0061] Optionally, the output module is further configured to, if the streaming text contains the preset type of characters, cache the current streaming text; determine whether to generate new streaming text within a preset time period; if the preset time period is reached and no new streaming text is generated, output the cached streaming text; or, generate new streaming text within the preset time period, cache the new streaming text, and if the new streaming text contains the preset type of characters, return to the step of determining whether to generate new streaming text within the preset time period.
[0062] Optionally, the output module is further configured to output the cached streaming text when the new streaming text does not contain characters of a preset type.
[0063] Optionally, if the cached streaming texts include multiple ones, the output module is configured to perform text splicing on the multiple cached streaming texts to obtain a target streaming text when the new streaming text does not contain characters of a preset type; and output the target streaming text.
[0064] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor executable instructions; wherein the processor is configured to implement the steps of the human-computer dialogue method provided in the first aspect of the present disclosure when calling the executable instructions stored on the memory.
[0065] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the human-computer dialogue method provided in the first aspect of the present disclosure are implemented.
[0066] The technical solution provided by the embodiment of the present disclosure may include the following beneficial effects: first, a conversation message input by a user is obtained. Then, according to the conversation message, a candidate reply text and target constraint information are determined; the target constraint information is used to indicate the scenario corresponding to the conversation message. Finally, according to the candidate reply text and the target constraint information, a target reply message corresponding to the conversation message is output. Through the above method, the candidate reply text and the target constraint information can be determined according to the conversation message input by the user. Then, when the target reply message is subsequently generated, the scenario corresponding to the user conversation message can be well combined according to the target constraint information. In this way, the target reply message finally output can not be separated from the scenario of the user conversation, thereby ensuring the accuracy and effectiveness of the entire conversation, and indirectly improving the user's satisfaction with the conversation. At the same time, the implementation method of this solution is simple and can be implemented without complex algorithms. The solution has strong executable capabilities and can adapt to various application scenarios.
[0067] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0069] Figure 1 The diagram is a diagram of a conversation flow according to an exemplary embodiment.
[0070] Figure 2 The present invention is a flowchart of a method for human-computer dialogue according to an exemplary embodiment.
[0071] Figure 3 The figure is a flow chart showing another method for human-computer dialogue according to an exemplary embodiment.
[0072] Figure 4 The figure is a flow chart showing another method for human-computer dialogue according to an exemplary embodiment.
[0073] Figure 5 The figure is a flow chart showing another method for human-computer dialogue according to an exemplary embodiment.
[0074] Figure 6 The present invention is a schematic diagram showing a conversation interface according to an exemplary embodiment.
[0075] Figure 7 The present invention is a block diagram showing a human-computer dialogue device according to an exemplary embodiment.
[0076] Figure 8It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0077] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0078] The terms "first", "second", etc. in the specification and claims of this application and the above drawings are used to distinguish similar objects and do not have to be understood as a specific order or sequence. In addition, in the description with reference to the drawings, the same symbols in different drawings represent the same elements.
[0079] In the description of the present disclosure, unless otherwise specified, "multiple" means two or more than two, and other quantifiers are similar thereto; "at least one item", "one item or multiple items" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one item a can represent any number of a; for another example, one item or multiple items among a, b and c can represent: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple; "and / or" is a kind of description of the association relationship of associated objects, indicating that there can be three kinds of relationships, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " indicates that the associated objects before and after are in an "or" relationship.
[0080] Although operations or steps are described in a specific order in the drawings in the embodiments of the present disclosure, it should not be understood that it is required to perform these operations or steps in the specific order shown or in a serial order, or to perform all the operations or steps shown to obtain the desired results. In the embodiments of the present disclosure, these operations or steps can be performed in series; these operations or steps can also be performed in parallel; or some of these operations or steps can be performed.
[0081] Before introducing a method, device, electronic device and storage medium for human-computer dialogue provided by the present disclosure, the application scenarios involved in each embodiment of the present disclosure are first introduced. This solution is applied to the scenario of human-computer dialogue. In some scenarios, the reply content output by the language model will be out of the scenario of the user conversation, resulting in the user being unable to obtain the expected information, resulting in a poor user conversation experience. For example, the user initiates a conversation in the dialog box and asks about the relevant parameters of mobile phone A, but the reply content finally output by the language model is about the after-sales solution of mobile phone A, that is, the user's question and the model's reply do not belong to the same scenario. The user's question belongs to the product consultation scenario, and the model's answer content belongs to the after-sales customer service scenario. At this time, the user will not be able to obtain the information they want to know, thereby affecting the user's conversation experience.
[0082] In order to solve the above technical problems, the present invention provides a method, device, electronic device and storage medium for human-computer dialogue, which can determine candidate reply texts and target constraint information according to the conversation message input by the user. Then, when the target reply message is subsequently generated, the scenario corresponding to the user conversation message can be well combined according to the target constraint information. In this way, the target reply message finally output can not deviate from the scenario of the user conversation, thereby ensuring the accuracy and effectiveness of the entire conversation, and indirectly improving the user's satisfaction with the conversation. At the same time, the implementation method of this solution is simple and can be implemented without complex algorithms. The solution has strong executable capabilities and can adapt to various application scenarios.
[0083] For ease of understanding, the entire session implementation process is briefly described below. Figure 1 As shown, the user can input the message text that needs to be communicated by voice or text. It should be noted that when the user inputs voice, the voice can be converted into corresponding text by a related voice recognition model, such as an ASR (Automatic Speech Recognition) model. Then, the database is searched according to the acquired message text to obtain one or more possible candidate reply texts, which can be understood as a reference text for the large model (i.e., language model) to reply. After that, a complete reply text set is constructed through prompt, that is, one or more candidate reply texts are spliced and combined. After that, the spliced reply text set is input into the large model to obtain the result of the model output, i.e., the target reply text. Finally, according to the target reply text output by the model, the target reply message corresponding to the input is output to the user. In this way, a conversation process is completed. If the user needs multiple conversations, the above conversation process can be initiated multiple times.
[0084] The specific implementation modes of the present invention are described in detail below with reference to the accompanying drawings.
[0085] Figure 2 This is a flowchart of a method for human-computer dialogue according to an exemplary embodiment. The method is applied to a terminal. The terminal may be, for example, a mobile terminal such as a smart phone, a tablet computer, a smart TV, a smart watch, a PDA (Personal Digital Assistant), a portable computer, or a smart home device such as a sweeping robot, an air purifier, an air conditioner, a lighting lamp, a speaker, a robot, etc. Figure 2 As shown, the method may include the following steps:
[0086] In step S101, a conversation message input by a user is obtained.
[0087] The conversation message may be generated by the user through text input or voice input. In the case where the conversation message is input through voice, the semantics may be converted into corresponding text through the relevant voice recognition model, so as to subsequently determine the candidate reply text according to the conversation message in text format. The conversation message may represent the relevant information of the target entity that the user wants to obtain. The target entity may be an item, a commodity, etc., and the relevant information of the target entity may be parameters, prices, attributes, etc. of the target entity. For example, if the user wants to know the price of a certain mobile phone, the conversation message may include "What is the price of mobile phone A".
[0088] In step S102, candidate reply texts and target constraint information are determined according to the conversation message.
[0089] The candidate reply text may be understood as a specific parameter value obtained for the conversation message. For example, if the conversation message includes "how much is the price of mobile phone A", the candidate reply text may include that the price of mobile phone A is 3999. The target constraint information is used to indicate the scenario corresponding to the conversation message.
[0090] For example, the preset database can be searched through the conversation message to obtain the corresponding candidate reply text, which may be one or more. Secondly, according to the conversation message, the current user's conversation intention, such as purchase intention, after-sales intention, etc., can be determined, and the corresponding target constraint information can be determined according to the conversation intention and the candidate reply text.
[0091] In step S103, a target reply message corresponding to the conversation message is output according to the candidate reply text and the target constraint information.
[0092] In this step, a corresponding target reply text may be generated through a language model according to the candidate reply text and the target constraint information, and then, according to the target reply text, a target reply message corresponding to the conversation message is output.
[0093] The language model may be any language model in related question-answering technologies, such as but not limited to an LLM model (Large Language Model), a miniMax model, and the like.
[0094] In some embodiments, the target reply message may be output in the form of text, or in the form of voice, or in a combination of text and voice, and the present disclosure does not make any specific limitation on this.
[0095] By adopting the above method, the candidate reply text and target constraint information can be determined according to the conversation message input by the user. Then, when the target reply message is subsequently generated, the scenario corresponding to the user conversation message can be well combined according to the target constraint information. In this way, the target reply message finally outputted can be consistent with the scenario of the user conversation, thereby ensuring the accuracy and effectiveness of the entire conversation, and indirectly improving the user's satisfaction with the conversation. At the same time, the implementation method of this solution is simple and can be implemented without complex algorithms. The solution has strong executable capabilities and can adapt to various application scenarios.
[0096] The above step S102 is described in detail below.
[0097] In some embodiments, Figure 3 As shown, in the above step S102, determining the candidate reply text and the target constraint information according to the conversation message may include the following steps:
[0098] In step S1021, a plurality of associated texts related to the conversation message are acquired according to the conversation message.
[0099] In this step, a database may be pre-established, and the database may store relevant information of multiple target entities. For example, for mobile phones, the names of all models of mobile phones of a certain brand and relevant parameter information corresponding to the mobile phones may be pre-stored. In this way, after obtaining the conversation message input by the user, the conversation message may be searched through the database to obtain multiple associated texts associated with the conversation message. For example, the associated text may include mobile phone, mobile phone A, and price 3999.
[0100] In the related technology, a graph database is usually pre-built. The data in the graph database is a relational network obtained by connecting different types of information together. The data in the graph database is stored in text format. However, due to the different data sources and data types, the data formats stored in the same application scenario are often very different. On the one hand, it will lead to poor data migration capabilities. On the other hand, due to the different data formats, during the information retrieval process, it is necessary to convert the data in different formats, which indirectly affects the efficiency of retrieval.
[0101] Considering the above problems, in a possible implementation, the database may include a preset vector database, in which data are stored after vectorization processing. For example, a text may be formed in the form of product name + product attribute name + product attribute value, such as "the price of mobile phone A is 3999 yuan", and then the text is converted into a vector using a vector conversion model and stored in the preset vector database. The associated text includes a first associated text, and obtaining multiple associated texts related to the conversation message in step S1021 may include: first, converting the conversation message into a corresponding conversation vector. Then, according to the conversation vector, obtaining multiple corresponding first associated texts from the preset vector database.
[0102] That is to say, the conversation message can be vectorized first, that is, the conversation message can be converted into a corresponding conversation vector. For example, the conversation message can be converted into a corresponding conversation vector through a vector conversion model in the related art. After that, a vector matching and / or having a high similarity with the conversation vector is retrieved through a preset vector database, and the vector is converted into a text form to obtain a plurality of first associated texts. For example, after the conversation message including "how much is mobile phone A" is converted into a vector, the vector similarity with the vector corresponding to "the price of mobile phone A is 3999 yuan" in the preset vector database is the highest, so the corresponding first associated text can be obtained. By means of vector retrieval, the problem of semantic similarity can be effectively solved. For example, although the texts of "how much is mobile phone A" and "the price of mobile phone A" are different, the semantics are similar. By converting into a vector, the corresponding associated text can be effectively recalled. In addition, by converting data into a vector, the scalability of the data can be enhanced. At the same time, since there is no need to convert multiple data formats, the efficiency of retrieval is also indirectly improved.
[0103] In addition, in order to achieve better conversion effects of the vector conversion model, the vector conversion model may be trained and optimized in advance using sample sets under specified scenarios, so that the trained vector conversion model can be more adaptable to vector conversions in various scenarios.
[0104] In another possible implementation, the database includes a preset vector database, the associated text includes a first associated text, and in step S1021, obtaining multiple associated texts related to the session message according to the session message may include: first, determining the target semantic slot corresponding to the session message. The semantic slot may be understood as an attribute that has been clearly defined by an entity (query). For example, the attributes in the product name slot, product price slot, and product attribute slot in the commodity consultation scenario are "product name", "product price", and "product attribute", respectively. By determining the target semantic slot corresponding to the session message, the user's current needs can be preliminarily determined. For example, the target semantic slot corresponding to the session message may be determined by rule slot extraction (inverse maximum matching algorithm) or model slot extraction (text to sql model). The rule slot extraction may be understood as matching the session message with multiple preset semantic slots to obtain a successfully matched target semantic slot; the model slot extraction may be understood as determining a target semantic slot with a high similarity to the session message through a model. Then, the target semantic slot may be converted into a corresponding semantic slot vector, and the semantic slot vector may be used as the session vector. Afterwards, a plurality of corresponding first associated texts are obtained from a preset vector database according to the conversation vector.
[0105] In addition, when there are multiple target semantic slots, the multiple target semantic slots can be concatenated to obtain corresponding semantic texts, and then the semantic texts are converted into corresponding semantic text vectors, and the semantic text vectors are used as the session vectors. Afterwards, the corresponding multiple first associated texts are obtained from the preset vector database according to the session vectors.
[0106] In another possible implementation, considering the integrity and comprehensiveness of data retrieval, in order to obtain all possible associated texts as much as possible, the database may further include a preset text database. The data in the preset text database is stored in the format of the data itself. For example, data may be stored in a multi-level mapping manner of "product category->product name->attribute->attribute value". The associated text may further include a second associated text. In step S1021, obtaining multiple associated texts related to the conversation message according to the conversation message may include: first, determining the target semantic slot corresponding to the conversation message. Then, according to the target semantic slot, obtaining multiple corresponding second associated texts from the preset text database. That is, after obtaining multiple first associated texts, the second associated text corresponding to the target semantic slot may also be retrieved from the preset text database by matching. In this way, the associated text may be obtained by converting into a vector, and the corresponding associated text may be effectively recalled; the associated text may also be obtained by text retrieval, and the associated text may be accurately recalled.
[0107] It should be noted that the above three possible implementations can be implemented separately or in combination, and the present disclosure does not make any specific limitations on this.
[0108] In step S1022, the candidate reply text is determined based on the multiple associated texts.
[0109] In a possible implementation, multiple associated texts may be used as candidate reply texts, so that as many associated texts related to the conversation message as possible can be retained. The candidate reply text is the prompt text.
[0110] Specifically, multiple associated texts may be concatenated to obtain the candidate reply text. For example, when the associated text includes multiple parameters of a product, the following candidate reply text may be obtained by concatenating:
[0111] Parameters of {product name}:(
[0112] {parameter 1}:({value1};{value2})
[0113] {parameter 2}: ({value1}; {value2}) ...)
[0115] The () symbol can be used to separate the complete text content. Each parameter uses a key-value structure, and different parameters can be separated by line breaks. The value of the parameter is a natural text expression. When there are multiple values, the ; symbol can be used to separate them.
[0116] The above text concatenation method can effectively help the subsequent voice model understand the prompt text, and can give the same understanding attention to different product parameters in the text, making it easier for the voice model to generate reply text.
[0117] In another possible implementation, the user's conversation intention may also be determined based on the conversation message. For example, the user's conversation intention reflected by the conversation message may be determined by an intention recognition model in the related art. Then, based on the conversation intention, multiple target texts may be screened out from the multiple associated texts. Finally, the multiple target texts may be used as the candidate reply texts.
[0118] For example, if the conversation message includes "recommend a mobile phone", it can be determined through identification that the user's conversation intention is a recommendation intention. At this time, according to the above step S1021, a large amount of information about mobile phones of different brands and models may be obtained. If all of the large amount of information is directly displayed to the user, it cannot effectively help the user make a choice. Therefore, for the recommendation intention, the obtained multiple associated texts can be filtered. For example, the target text corresponding to the latest preset number of mobile phones released can be determined from multiple associated texts, and the target text can be used as a candidate reply text. In this way, the associated text can be flexibly filtered according to different conversation intentions.
[0119] It is understandable that different screening conditions may be set for different session intents to meet the needs of users under different session intents as much as possible.
[0120] In step S1023, the target constraint information is determined according to the candidate reply text and the conversation message.
[0121] For example, the user's conversation intent can be determined based on the conversation message, and then the target constraint information can be determined based on the candidate reply text and the conversation intent. In other words, the target constraint information is used to constrain the language model in the subsequent process of generating the reply text, so as to prevent the language model from making arbitrary replies out of the conversation scenario, and ensure that the language model outputs a reply text that better meets the requirements under the constraints.
[0122] The target constraint information may include, for example, reply scenario constraint information, reply content constraint information, and reply speech constraint information.
[0123] The reply scenario constraint information can be understood as the background persona of the reply system. For example, when the conversation intent includes a recommendation intent and the candidate reply text contains relevant information about a mobile phone, the reply scenario constraint information may include: "You are a mall assistant for a certain brand. Your name is XX. You will be responsible for the product shopping guide and customer service of the brand. Please actively and positively reply to my questions. If I need you to recommend a mobile phone, you will only recommend products of a certain brand to me. If you are solving customer service problems, you do not need to recommend a mobile phone." The reply scenario constraint information can be used to limit the scenario in which the current reply system is located and the identity of the role simulated, so that when the language model subsequently generates a target reply message, it can not deviate from the current conversation scenario, thereby more accurately replying to the user's conversation questions.
[0124] The reply content constraint information can be understood as supplementary information to the reply content. For example, when the conversation intent includes a recommendation intent and the candidate reply text contains relevant information about the mobile phone, the reply content constraint information may include: "If your answer contains a link, you must use the link in the candidate reply text above and cannot modify it. If my question does not conform to normal logic, you will give detailed reasons and let me ask again. If my question has a definite answer, you will reply to me with the definite answer." The reply content constraint information can be used to constrain the content of the reply system so that when the language model subsequently generates the target reply message, it will not deviate from the content in the candidate reply text, thereby more accurately replying to the user's conversational questions.
[0125] The reply word constraint information can be understood as a guiding template for the reply word. For example, when the conversation intent includes a recommendation intent and the candidate reply text includes relevant information about the mobile phone, the reply word constraint information may include:
[0126] “Me: Who are you?
[0127] Student XX: I am the smart shopping guide assistant of Brand A.
[0128] Me: Is the mobile phone of brand B easy to use?
[0129] Student XX: Sorry, I am not in a position to comment on products of other brands.
[0130] Me: How much is mobile phone A?
[0131] Student XX: The price of mobile phone A is 3999.
[0132] Me: How about the photo taking function of mobile phone B?
[0133] Student XX: Mobile phone B is not on the market yet.
[0134] Me: Is mobile phone C available?
[0135] Student XX: I found that mobile phone C is out of stock in brand A mall.
[0136] In actual application scenarios, when the user's conversation message includes "Is mobile phone C in stock?", if the reply text constraint information is not added, the reply text generated by the subsequent language model may include "Sorry, I am an AI assistant and cannot query the inventory information of mobile phone C." After adding the reply text constraint information, the reply text generated by the subsequent language model will include "I found that mobile phone C is out of stock in the brand A mall." In this way, in the scenario of a single round of conversation, the reply text style of the language model can be directly corrected through the reply text constraint information, so that the reply text finally output by the language model can meet the expected problem. At the same time, if the conversation message of the user's question does not conform to the usual logic, the reply text constraint information can also be used to guide the user to re-initiate a new conversation.
[0137] In addition, considering that in some scenarios, the question-answering system in the related technology is not capable of corresponding to multi-round task-based dialogue scenarios, and has limited processing capabilities for context-dependent commodity knowledge questions and answers. For example, if a user initiates multiple conversations, the first conversation initiated by the user includes "Recommend a mobile phone"; the reply includes "OK, I recommend that you choose mobile phone A." The second conversation initiated by the user includes "How fast does this phone charge?" If the specific mobile phone model cannot be obtained based only on the second conversation initiated by the user, the generated reply message cannot accurately correspond to the mobile phone model that the user needs to know. Therefore, in response to the above problems, Figure 4 As shown, the method may also include the following steps:
[0138] In step S1024, historical conversations with the user within a historical period are obtained.
[0139] Correspondingly, the above step S1021 of acquiring multiple associated texts related to the conversation message according to the conversation message includes: acquiring multiple associated texts related to the conversation message according to the historical conversation and the conversation message.
[0140] Wherein, according to the historical conversation and the conversation message, obtaining multiple associated texts related to the conversation message may refer to the three implementation methods shown in the above step S1021, which will not be repeated here.
[0141] In this way, historical conversations can be combined during the retrieval process, thereby improving the accuracy of replies and the coherence of reply content, thereby improving the user's conversation experience.
[0142] The above step S103 is described in detail below.
[0143] In some embodiments, Figure 5As shown, in the above step S103, outputting the target reply message corresponding to the conversation message according to the candidate reply text and the target constraint information may include the following steps:
[0144] In step S1031, a corresponding target reply text is generated through a language model according to the candidate reply text and the target constraint information.
[0145] For example, the candidate reply text and the target constraint information can be packaged and input into the language model to obtain the target reply text output by the language model. For example, when the target constraint information includes reply scenario constraint information, reply content constraint information, and reply speech constraint information, the candidate reply text and the target constraint information can be packaged in the following manner:
[0146] {Reply scene constraint information}
[0147] ---
[0148] {Candidate response text}
[0149] ---
[0150] {Reply content constraint information}
[0151] ---
[0152] {Reply to the restriction information}
[0153] ---
[0154] Different parts of information can be distinguished by "---" separators. After encapsulation, the candidate reply text and the target constraint information can be input into the language model together.
[0155] In step S1032, a target reply message corresponding to the conversation message is output according to the target reply text.
[0156] In actual application scenarios, since the reasoning of language models often takes a long time, in order to improve the efficiency of replies, a streaming text output method is often used. Specifically, multiple streaming texts can be generated based on the target reply text. Then, according to the generation time of each streaming text, multiple streaming texts are output in sequence. That is, the generated streaming text can be output in real time, and the interface seen by the user is the text displayed line by line and word by word. Among them, the streaming text can be, for example, a stream text.
[0157] In order to improve the display effect and diversity of the reply text, the database may store the purchase link, picture and other information corresponding to a certain product. In this way, the generated target reply text may contain links (such as website links, picture links, etc.) or phone numbers. If the real-time output is continued according to the above streaming text output method, the link may be truncated. For example, if three streaming texts are generated according to the target reply text, streaming text 1 includes "XX TV's after-sales service policy is as follows: \n\n1.7 days no-reason return:", streaming text 2 includes "Warranty policy: For specific warranty policies, please refer to the following figure: https / / abc.", and streaming text 3 includes "ice.api / images / 12345678.jpeg Tips: The remaining warranty time can be [queried by SN number]". If streaming text 1, streaming text 2 and streaming text 3 are output in sequence according to the generation time, the links contained therein will not be rendered normally, that is, the links finally displayed are pure strings that cannot be jumped.
[0158] Considering the above problems, in some embodiments, each time a stream text is generated, it can be determined whether the stream text contains characters of a preset type. And if the stream text does not contain characters of the preset type, the stream text is output. That is, if the current stream text does not contain characters of the preset type, it is considered that there is no link or telephone number in the current stream text, and at this time, the stream text can be directly output. For example, the stream text 1 in the above example can be directly output.
[0159] Furthermore, in the case where the streaming text contains characters of the preset type, the current streaming text can be cached first, that is, the streaming text is not directly output. Then, it is determined whether a new streaming text is generated within a preset time period. If the preset time period is reached and no new streaming text is generated, the cached streaming text is output; or, a new streaming text is generated within the preset time period, the new streaming text is cached, and in the case where the new streaming text contains characters of the preset type, the step of determining whether a new streaming text is generated within the preset time period is returned. The above steps can be understood as, if there are characters of the preset type in the current streaming text, the streaming text is cached, and the subsequently generated streaming texts are cached in sequence, until no new streaming text is generated within the preset time, or the generated new streaming text does not contain characters of the preset type, then all cached streaming texts can be output. In addition, in the case where the new streaming text does not contain characters of the preset type, all cached streaming texts are output.
[0160] For ease of understanding, taking the three streaming texts in the above example as an example, since streaming text 1 does not contain characters of the preset type, streaming text 1 can be directly output. Streaming text 2 contains characters of the preset type, so streaming text 2 can be cached. After generating streaming text 3, since streaming text 3 still includes characters of the preset type, streaming text 3 is cached and continues to wait for the newly generated streaming text. After waiting for the preset time period, since no new streaming text is generated, the cached streaming text 2 and streaming text 3 can be output.
[0161] In some embodiments, if the cached streaming texts include multiple ones, the above-mentioned outputting the cached streaming texts may include: performing text splicing on the multiple cached streaming texts to obtain a target streaming text, and outputting the target streaming text.
[0162] For example, taking the three streaming texts in the above example as an example, after concatenating streaming text 2 and streaming text 3, the target streaming text can be obtained as "Warranty Policy: For specific warranty policy, please refer to the figure below: https: / / abc.ice.api / images / 12345678.jpeg Tips: The remaining warranty time can be [queried by SN number]". In this way, a complete link can be output through the same streaming text, ensuring that the link will not be truncated.
[0163] In addition, after the text is spliced, the spliced target streaming text can be rendered and output. For example, the target streaming text can be converted into a markdown text format to achieve the rendering of the target streaming text. In this way, after the client receives the target streaming text in the markdown text format, it can display the GUI for pictures, jumpable links, jumpable phone numbers, cards, etc.
[0164] For example, Figure 6 A schematic diagram showing a conversation interface is shown, such as Figure 6As shown, when the conversation message input by the user includes conversation message 1: "Can you recommend a mobile phone?", reply 1: "Of course, do you have any special requirements for the performance, photography, screen, battery life, etc. of the mobile phone?"; conversation message 2: "Well, I want a mobile phone of about 4,000 yuan", reply 2: "Based on your budget and needs, I recommend mobile phone A to you. This mobile phone is equipped with XX mobile platform and has powerful performance. The rear camera supports anti-shake function and has excellent photography effect. The screen uses OLED screen and has excellent display effect. The battery capacity is 4500mah and the battery life is also good. The reference price of the 8GB+256GB version is 3999 yuan, which is very suitable for your budget."; conversation message 3: "What is the pixel of this mobile phone?", reply 3: "The pixel of the rear camera of mobile phone A is 50 million + 50 million + 50 million + 50 million, and the pixel of the front camera is 32 million."; and, after reply 2 and reply 3, the purchase link of mobile phone A is displayed respectively, so that the user can purchase mobile phone A through the purchase link.
[0165] By adopting the above method, the candidate reply text and target constraint information can be determined according to the conversation message input by the user. Then, when the target reply message is subsequently generated, the scenario corresponding to the user conversation message can be well combined according to the target constraint information. In this way, the target reply message finally outputted can be consistent with the scenario of the user conversation, thereby ensuring the accuracy and effectiveness of the entire conversation, and indirectly improving the user's satisfaction with the conversation. At the same time, the implementation method of this solution is simple and can be implemented without complex algorithms. The solution has strong executable capabilities and can adapt to various application scenarios.
[0166] Figure 7 is a block diagram of a device for human-computer dialogue according to an exemplary embodiment. Figure 7 The device 200 comprises:
[0167] The acquisition module 201 is configured to acquire a conversation message input by a user;
[0168] The determination module 202 is configured to determine a candidate reply text and target constraint information according to the conversation message; the target constraint information is used to indicate a scenario corresponding to the conversation message;
[0169] The output module 203 is configured to output a target reply message corresponding to the conversation message according to the candidate reply text and the target constraint information.
[0170] Optionally, the determination module 202 is configured to obtain a plurality of associated texts related to the conversation message according to the conversation message; determine the candidate reply text according to the plurality of associated texts; and determine the target constraint information according to the candidate reply text and the conversation message.
[0171] Optionally, the associated text includes a first associated text, and the determining module 202 is configured to convert the conversation message into a corresponding conversation vector; and acquire a corresponding plurality of first associated texts from a preset vector database according to the conversation vector.
[0172] Optionally, the determination module 202 is configured to determine a target semantic slot corresponding to the conversation message; convert the target semantic slot into a corresponding semantic slot vector, and use the semantic slot vector as the conversation vector.
[0173] Optionally, the associated text further includes second associated texts, and the determination module 202 is configured to obtain a corresponding plurality of second associated texts from a preset text database according to the target semantic slot.
[0174] Optionally, the acquisition module 201 is further configured to acquire historical conversations with the user within a historical period;
[0175] The determination module 202 is configured to obtain a plurality of associated texts related to the conversation message according to the historical conversation and the conversation message.
[0176] Optionally, the determination module 202 is configured to determine the user's conversation intention based on the conversation message; filter out multiple target texts from the multiple associated texts based on the conversation intention; and use the multiple target texts as the candidate reply texts.
[0177] Optionally, the output module 203 is configured to generate a corresponding target reply text through a language model according to the candidate reply text and the target constraint information; and output a target reply message corresponding to the conversation message according to the target reply text.
[0178] Optionally, the output module 203 is configured to generate a plurality of flow texts according to the target reply text; and output the plurality of flow texts in sequence according to the generation time of each flow text.
[0179] Optionally, the output module 203 is configured to determine whether the streaming text contains characters of a preset type each time the streaming text is generated; if the streaming text does not contain characters of the preset type, output the streaming text.
[0180] Optionally, the output module 203 is further configured to cache the current streaming text if the streaming text contains characters of the preset type; determine whether to generate new streaming text within a preset time period; if the preset time period is reached and no new streaming text is generated, output the cached streaming text; or, generate new streaming text within the preset time period, cache the new streaming text, and if the new streaming text contains characters of the preset type, return to the step of determining whether to generate new streaming text within the preset time period.
[0181] Optionally, the output module 203 is further configured to output the cached streaming text when the new streaming text does not contain characters of a preset type.
[0182] Optionally, if the cached streaming texts include multiple ones, the output module 203 is configured to perform text splicing on the multiple cached streaming texts to obtain a target streaming text when the new streaming text does not contain characters of a preset type; and output the target streaming text.
[0183] By adopting the above device, it is possible to determine the candidate reply text and target constraint information according to the conversation message input by the user. Then, when the target reply message is subsequently generated, the scenario corresponding to the user conversation message can be well combined according to the target constraint information. In this way, the target reply message finally outputted can be consistent with the scenario of the user conversation, thereby ensuring the accuracy and effectiveness of the entire conversation, and indirectly improving the user's satisfaction with the conversation. At the same time, the implementation method of this solution is simple and can be implemented without complex algorithms. The solution has strong executable capabilities and can adapt to various application scenarios.
[0184] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0185] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the steps of the human-computer dialogue method provided by the present disclosure.
[0186] Figure 8 is a block diagram of an electronic device 300 according to an exemplary embodiment. For example, the electronic device 300 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0187] Reference Figure 8, the electronic device 300 may include one or more of the following components: a processing component 302 , a memory 304 , a power component 306 , a multimedia component 308 , an audio component 310 , an input / output interface 312 , a sensor component 314 , and a communication component 316 .
[0188] The processing component 302 generally controls the overall operation of the electronic device 300, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 302 may include one or more processors 320 to execute instructions to complete all or part of the steps of the above-mentioned human-computer dialogue method. In addition, the processing component 302 may include one or more modules to facilitate the interaction between the processing component 302 and other components. For example, the processing component 302 may include a multimedia module to facilitate the interaction between the multimedia component 308 and the processing component 302.
[0189] The memory 304 is configured to store various types of data to support operations on the electronic device 300. Examples of such data include instructions for any application or method operating on the electronic device 300, contact data, phone book data, messages, pictures, videos, etc. The memory 304 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0190] The power supply component 306 provides power to the various components of the electronic device 300. The power supply component 306 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 300.
[0191] The multimedia component 308 includes a screen that provides an output interface between the electronic device 300 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 308 includes a front camera and / or a rear camera. When the electronic device 300 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0192] The audio component 310 is configured to output and / or input audio signals. For example, the audio component 310 includes a microphone (MIC), and when the electronic device 300 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 304 or sent via the communication component 316. In some embodiments, the audio component 310 also includes a speaker for outputting audio signals.
[0193] The input / output interface 312 provides an interface between the processing component 302 and the peripheral interface modules, which may be keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0194] The sensor assembly 314 includes one or more sensors for providing various aspects of status assessment for the electronic device 300. For example, the sensor assembly 314 can detect the open / closed state of the electronic device 300, the relative positioning of components, such as the display and keypad of the electronic device 300, and the sensor assembly 314 can also detect the position change of the electronic device 300 or a component of the electronic device 300, the presence or absence of user contact with the electronic device 300, the orientation or acceleration / deceleration of the electronic device 300, and the temperature change of the electronic device 300. The sensor assembly 314 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 314 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 314 may also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0195] The communication component 316 is configured to facilitate wired or wireless communication between the electronic device 300 and other devices. The electronic device 300 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 316 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 316 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0196] In an exemplary embodiment, the electronic device 300 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above-mentioned human-computer dialogue method.
[0197] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 304 including instructions, and the instructions can be executed by a processor 320 of an electronic device 300 to complete the above-mentioned human-computer dialogue method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0198] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program that can be executed by a programmable device, and the computer program has a code portion for executing the above-mentioned human-computer dialogue method when executed by the programmable device.
[0199] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the present disclosure. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0200] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A method for human-computer dialogue, characterized in that: The method comprises: Get the conversation message entered by the user; Determine candidate reply texts and target constraint information according to the conversation message; the target constraint information is used to indicate a scenario corresponding to the conversation message; Outputting a target reply message corresponding to the conversation message according to the candidate reply text and the target constraint information.
2. The method according to claim 1, characterized in that: The determining, according to the conversation message, candidate reply texts and target constraint information comprises: According to the conversation message, obtaining a plurality of associated texts related to the conversation message; Determining the candidate reply text according to the plurality of associated texts; The target constraint information is determined according to the candidate reply text and the conversation message.
3. The method according to claim 2, characterized in that The associated text includes a first associated text, and acquiring a plurality of associated texts related to the conversation message according to the conversation message includes: Converting the session message into a corresponding session vector; According to the conversation vector, a corresponding plurality of first associated texts are obtained from a preset vector database.
4. The method according to claim 3, characterized in that The converting the session message into a corresponding session vector comprises: Determining a target semantic slot corresponding to the conversation message; The target semantic slot is converted into a corresponding semantic slot vector, and the semantic slot vector is used as the session vector.
5. The method according to claim 4, characterized in that The associated text further includes a second associated text, and acquiring a plurality of associated texts related to the conversation message according to the conversation message includes: According to the target semantic slot, a corresponding plurality of second associated texts are acquired from a preset text database.
6. The method according to claim 2, characterized in that The method further comprises: Obtain historical conversations with the user within a historical period; The acquiring, according to the conversation message, a plurality of associated texts related to the conversation message comprises: According to the historical conversation and the conversation message, a plurality of associated texts related to the conversation message are acquired.
7. The method according to claim 2, characterized in that The method further comprises: Determining the user's conversation intention according to the conversation message; Determining the candidate reply text according to the plurality of associated texts includes: According to the conversation intention, a plurality of target texts are selected from the plurality of associated texts; The plurality of target texts are used as the candidate reply texts.
8. The method according to any one of claims 1 to 7, characterized in that The outputting the target reply message corresponding to the conversation message according to the candidate reply text and the target constraint information comprises: Generate a corresponding target reply text through a language model according to the candidate reply text and the target constraint information; According to the target reply text, a target reply message corresponding to the conversation message is output.
9. The method according to claim 8, characterized in that Outputting the target reply message corresponding to the conversation message according to the target reply text includes: Generate multiple stream texts according to the target reply text; According to the generation time of each of the streaming texts, multiple streaming texts are output in sequence.
10. The method according to claim 9, characterized in that Outputting the plurality of streaming texts in sequence according to the generation time of each streaming text comprises: Each time a stream text is generated, determining whether the stream text contains characters of a preset type; In a case where the stream text does not contain characters of the preset type, the stream text is output.
11. The method according to claim 10, characterized in that The method further comprises: When the stream text contains characters of the preset type, caching the current stream text; Determining whether to generate new streaming text within a preset time period; If no new streaming text is generated within the preset time period, the cached streaming text is output; or, Generate new streaming text within a preset time period, cache the new streaming text, and return to the step of determining whether to generate new streaming text within the preset time period when the new streaming text contains characters of a preset type.
12. The method according to claim 11, characterized in that The method further comprises: If the new streaming text does not contain characters of the preset type, the cached streaming text is output.
13. The method according to claim 12, characterized in that If the cached streaming text includes multiple characters, in the case where the new streaming text does not include characters of the preset type, outputting the cached streaming text includes: When the new streaming text does not contain characters of the preset type, multiple cached streaming texts are concatenated to obtain the target streaming text; Output the target streaming text.
14. A device for human-computer dialogue, characterized in that: The device comprises: An acquisition module is configured to acquire a conversation message input by a user; A determination module, configured to determine candidate reply texts and target constraint information according to the conversation message; the target constraint information is used to indicate a scenario corresponding to the conversation message; The output module is configured to output a target reply message corresponding to the conversation message according to the candidate reply text and the target constraint information.
15. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the steps of the method described in any one of claims 1 to 13 when calling the executable instructions stored in the memory.
16. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the steps of the method described in any one of claims 1 to 13 are implemented.