Information recommendation method, device, equipment and storage medium

By analyzing multiple rounds of conversations with a large language model, the intelligent agent identifies user intent and generates matching reference information, solving the problem of intent recognition in intelligent customer service when user input is inaccurate, and improving service quality.

CN119691160BActive Publication Date: 2025-10-10BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411751346.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-10-10
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

Existing intelligent customer service cannot accurately identify user intentions when user input is not accurate enough, resulting in inaccurate reference content provided, which affects service quality.

Method used

By analyzing multiple rounds of conversations based on a large language model, an intelligent agent can determine the user's intention, retrieve reference multi-round conversations that match the intention from the database, and generate reference recommendation information.

Benefits of technology

The accuracy of user intent recognition and service quality have been improved, ensuring that the reference information provided is more in line with user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119691160B_ABST
    Figure CN119691160B_ABST
Patent Text Reader

Abstract

The present disclosure provides an information recommendation method and device, equipment and a storage medium, relates to the technical field of computers, in particular to the technical field of artificial intelligence such as large models, intelligent recommendation and agents, and can be used in the field of smart finance. The specific implementation scheme is as follows: in response to receiving a current multi-round conversation between a first object and a second object, determining the intention of the first object based on the current multi-round conversation, wherein the first object is a service receiver and the second object is a service provider; based on the intention of the first object, determining a plurality of candidate items and a plurality of candidate item associated reference multi-round conversations from a database; based on the intention of the first object and the plurality of candidate item associated reference multi-round conversations, determining one or more reference items from the plurality of candidate items; taking the one or more reference items as reference recommendation information, and sending the reference recommendation information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, in particular to artificial intelligence technologies such as large models, intelligent recommendations, and intelligent agents, and can be used in fields such as smart finance. Background Art

[0002] In related technologies, intelligent customer service determines user intent based on a question entered by the user, and then searches a database for reference content that has the highest similarity to the user's intent. However, if the user's question is not accurate enough, the intelligent customer service often fails to accurately identify the user's intent. This results in inaccurate reference content based on the inaccurate intent, making it impossible to provide users with accurate reference content, thus affecting the quality of service provided to users. Therefore, how to accurately identify the user's actual intent and then obtain reference content that matches the user's intent, thereby providing users with quality service, becomes a technical problem that needs to be solved. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, device, and storage medium for information recommendation.

[0004] According to one aspect of the present disclosure, there is provided an information recommendation method, comprising:

[0005] In response to receiving a current multi-round conversation between a first object and a second object, determining an intention of the first object based on the current multi-round conversation, wherein the first object is a service recipient and the second object is a service provider;

[0006] Based on the intention of the first object, determining a plurality of candidate entries and a plurality of reference rounds of dialogue associated with the plurality of candidate entries from a database;

[0007] Determining one or more reference items from the multiple candidate items based on the intention of the first object and the reference multiple-round conversations associated with the multiple candidate items;

[0008] The one or more reference items are used as reference recommendation information, and the reference recommendation information is sent.

[0009] According to another aspect of the present disclosure, there is provided an information recommendation device, comprising:

[0010] an intent analysis module, configured to, in response to receiving a current multi-round conversation between a first object and a second object, determine an intent of the first object based on the current multi-round conversation, wherein the first object is a service recipient and the second object is a service provider;

[0011] retrieving a plurality of candidate entries and reference multi-turn dialogues associated with the plurality of candidate entries from a database based on the intention of the first object;

[0012] screening the plurality of candidate entries based on the intention of the first object and the reference multi-turn dialogues associated with the plurality of candidate entries to determine one or more reference entries from the plurality of candidate entries;

[0013] sending the one or more reference entries as reference recommendation information.

[0014] According to another aspect of the present disclosure, an electronic device is provided, comprising:

[0015] at least one processor; and

[0016] a memory connected with the at least one processor in communication; wherein,

[0017] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any of the embodiments of the present disclosure.

[0018] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to perform the method according to any of the embodiments of the present disclosure.

[0019] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to any of the embodiments of the present disclosure.

[0020] By adopting the method provided in the present embodiment, the intention of the first object can be obtained more accurately according to the analysis of the current multi-turn dialogue between the first object as the service recipient and the second object as the service provider. Further, by the more accurate intention of the first object, it can be ensured that the obtained candidate entries can provide questions and answers closer to the intention of the first object. In addition, in combination with the intention of the first object and the reference multi-turn dialogues associated with the plurality of candidate entries, reference recommendation information closer to the intention of the first object can be obtained more accurately from the plurality of candidate entries and sent, so as to obtain reference information matching the intention of the first object, thereby improving the quality of providing services for the first object.

[0021] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0022] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.

[0023] Figure 1 is a flowchart of an information recommendation method according to an embodiment of the present disclosure;

[0024] Figure 2 is a flowchart of an information recommendation method according to another embodiment of the present disclosure;

[0025] Figure 3 is a schematic block diagram of an information recommendation device according to an embodiment of the present disclosure;

[0026] Figure 4 is a schematic block diagram of an information recommendation device according to another embodiment of the present disclosure;

[0027] Figure 5 is a block diagram of an electronic device for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0029] Figure 1 It is a schematic flow chart of an information recommendation method provided according to an embodiment of the present disclosure. The method is applicable to situations where information recommendation is performed, and is particularly applicable to situations where information recommendation is performed in the financial field. The method can be performed by an information recommendation device, which can be implemented in software and / or hardware, and can be integrated into an electronic device that carries the function of recommending information, such as a server. It should be noted that the information recommendation method of the present disclosure can be performed by a large-model-based intelligent agent in a server, wherein an intelligent agent refers to a computer program based on a large language model, which has planning and thinking capabilities, memory capabilities, and the ability to use tool functions, and can independently complete a given task.

[0030] Figure 1 As shown, the information recommendation method proposed in the embodiment of the present disclosure includes:

[0031] S110 , in response to receiving a current multi-round conversation between a first object and a second object, determining an intention of the first object based on the current multi-round conversation, wherein the first object is a service recipient and the second object is a service provider.

[0032] The receiving the current multi-turn conversation between the first object and the second object can refer to that the server receives the current multi-turn conversation between the first object and the second object from the second device.

[0033] The first object can be a first user, the first user can be a user currently in need of service, and the first user can use a first device.

[0034] The second object is a service provider, that is, the second object can be a service provider providing service for the first object; specifically, the second object can be a customer service, for example, the second object can be a manual customer service or an intelligent customer service.

[0035] Optionally, the second object can be a manual customer service, and the second device is an electronic device used by the second object. Optionally, the second object is an intelligent customer service, and the intelligent customer service can be set as a module or a function in the second device.

[0036] The device type of the first device can be any one of a mobile phone, a tablet computer, a desktop computer, a notebook computer, etc., and the device type of the second device can be any one of a mobile phone, a tablet computer, a desktop computer, a notebook computer, etc., which will not be enumerated here for the sake of brevity.

[0037] The determining the intention of the first object based on the current multi-turn conversation can refer to that an agent in the server determines the intention of the first object based on the current multi-turn conversation.

[0038] S120, determining a plurality of candidate entries and reference multi-turn conversations associated with the plurality of candidate entries from a database based on the intention of the first object.

[0039] The reference multi-turn conversations associated with the plurality of candidate entries include one or more reference multi-turn conversations associated with each candidate entry in the plurality of candidate entries.

[0040] Each candidate entry includes a candidate question and a corresponding candidate answer.

[0041] The determining the plurality of candidate entries and the reference multi-turn conversations associated with the plurality of candidate entries from the database based on the intention of the first object can refer to that the agent in the server determines the plurality of candidate entries and the reference multi-turn conversations associated with the plurality of candidate entries from the database based on the intention of the first object.

[0042] S130, determining one or more reference entries from the plurality of candidate entries based on the intention of the first object and the reference multi-turn conversations associated with the plurality of candidate entries.

[0043] Among them, the determining of one or more reference items from the multiple candidate items based on the intention of the first object and the reference multi-round dialogues associated with the multiple candidate items can be: the intelligent agent in the server determines one or more reference items from the multiple candidate items based on the intention of the first object and the reference multi-round dialogues associated with the multiple candidate items.

[0044] S140: Use the one or more reference items as reference recommendation information, and send the reference recommendation information.

[0045] The taking of one or more reference items as reference recommendation information may be: the agent in the server takes the one or more reference items as reference recommendation information.

[0046] The sending of the reference recommendation information may be: the server sending the reference recommendation information to the second device, wherein the reference recommendation information is used by the second object corresponding to the second device to provide a service for the first object.

[0047] By adopting the method provided in this embodiment, based on the analysis of the current multiple rounds of conversations between the first object as the service recipient and the second object as the service provider, the intention of the first object can be obtained more accurately. Furthermore, through a more accurate intention of the first object, it can be ensured that the obtained candidate entries can provide questions and answers that are closer to the intention of the first object. In addition, combined with the intention of the first object and the reference multiple rounds of conversations associated with the multiple candidate entries, reference recommendation information that is closer to the intention of the first object can be more accurately obtained from the multiple candidate entries and sent, thereby obtaining reference information that matches the intention of the first object, thereby improving the quality of service provided to the first object.

[0048] In one embodiment, determining the intention of the first object based on the current multi-round conversations includes: determining multi-round conversations to be analyzed based on the current multi-round conversations; generating input information based on the multi-round conversations to be analyzed; inputting the input information into an intention analysis model to obtain an analysis result corresponding to the multi-round conversations to be analyzed output by the intention analysis model, wherein the analysis result includes the intention of the first object.

[0049] In one case, the current multi-round conversation between the first object and the second object received by the server may be the voice recognition result of all voice conversations between the first object and the second object recorded on the second device side before the trigger moment, which is received by the server from the second device. Wherein, in the case where the second object is a manual customer service, the trigger moment is the moment corresponding to the second object performing the trigger operation on the second device; in the case where the second object is an intelligent customer service, the trigger moment is the moment corresponding to the second device determining that the trigger condition is met. The trigger operation may be the operation of the second object clicking the corresponding physical trigger button or virtual trigger button on the second device according to actual needs, which is not limited in this embodiment. The trigger condition may be preset in the second device, which is not limited in this embodiment.

[0050] The current multi-round dialogue may be a speech recognition result corresponding to the speech dialogue between the first object and the second object, and the speech recognition result may be a character recognition result or a text recognition result.

[0051] The current multiple rounds of conversation have a sequence. The sequence of the current multiple rounds of conversation may be determined by one of the following methods: sorting each round of conversation based on the time stamps of the speech recognition results corresponding to each round of conversation from earliest to latest; or sorting each round of conversation based on the position of the speech recognition results corresponding to each round of conversation in the current multiple rounds of conversation.

[0052] Each round of the current multi-round conversation includes speech recognition results corresponding to two voices, and the speech recognition results corresponding to the two voices included in each round of conversation also have a sequence. The sequence of the speech recognition results corresponding to the two voices included in each round of conversation is determined by one of the following methods: sorting the speech recognition results corresponding to the two voices based on the time stamps of the speech recognition results corresponding to the two voices from earliest to latest; or sorting the speech recognition results corresponding to the two voices based on the positions of the speech recognition results corresponding to the two voices.

[0053] Optionally, determining the multi-round dialogue to be analyzed based on the current multi-round dialogue may include: the agent in the server determining whether the number of rounds in the current multi-round dialogue is greater than Z; if the number of rounds in the current multi-round dialogue is greater than Z, using the dialogue ranked as the last Z rounds in the current multi-round dialogue as the multi-round dialogue to be analyzed; if the number of rounds in the current multi-round dialogue is less than or equal to Z, using the current multi-round dialogue as the multi-round dialogue to be analyzed. Z may be a positive integer greater than or equal to 2 preset according to actual circumstances. In one example, Z may be equal to 5.

[0054] Optionally, the determining the to-be-analyzed multi-turn conversation based on the current multi-turn conversation can include that the agent in the server takes the current multi-turn conversation as the to-be-analyzed multi-turn conversation. That is, in this case, whether the number of turns of the current multi-turn conversation is greater than Z is not judged, and the current multi-turn conversation is directly taken as the to-be-analyzed multi-turn conversation.

[0055] In one case, the current multi-turn conversation between the first object and the second object received by the server can be a voice recognition result of the last Z-turn voice conversation between the first object and the second object recorded by the second device before the trigger time.

[0056] In this case, the determining the to-be-analyzed multi-turn conversation based on the current multi-turn conversation can include that the agent in the server takes the current multi-turn conversation as the to-be-analyzed multi-turn conversation.

[0057] It should be pointed out that if the number of turns of the voice recognition result corresponding to the voice conversation between the first object and the second object recorded by the second device before the trigger time is less than Z, the previous case can be used for processing, that is, the voice recognition result of all the voice conversations between the first object and the second object recorded before the trigger time is sent to the server, which is not repeated here.

[0058] All the to-be-analyzed multi-turn conversations must include the voice recognition result corresponding to one or more voices of the first object and the voice recognition result corresponding to one or more voices of the second object. Each turn of the to-be-analyzed multi-turn conversation includes two voice recognition results corresponding to two voices. It should be understood that although all the to-be-analyzed multi-turn conversations must include the voice recognition results corresponding to the voices of the two objects respectively, a turn of the to-be-analyzed multi-turn conversation can include the voice recognition results corresponding to two voices of the same object, or the voice recognition results corresponding to two voices of the two objects respectively.

[0059] In one example, the generating the input information based on the to-be-analyzed multi-turn conversation can include that the agent in the server takes the to-be-analyzed multi-turn conversation as the input information.

[0060] In one example, the generating the input information based on the to-be-analyzed multi-turn conversation can include that the agent in the server takes the to-be-analyzed multi-turn conversation and a prompt as the input information. The prompt is used to guide the intent analysis model to understand the task corresponding to the input information and supplement background knowledge. The prompt can be set according to actual conditions, which is not limited here.

[0061] Optionally, the analysis result can include an intention of the first object and a content summary.

[0062] The content summary is used to represent a question that the first object wants to ask by analyzing the multi-turn conversation to be analyzed. The content summary can also be referred to as a session summary of the multi-turn conversation to be analyzed, or a question that the first object wants to ask.

[0063] The intention of the first object is to express the question that the first object wants to ask in a standard question form. The intention of the first object can be a standard question or a question in a standard question form (or format) obtained by rewriting the content summary. The intention of the first object can also be referred to as a rewriting of the content summary, or a rewriting of the session summary, or a standard question, etc. The standard question form can be a standard question form of a specified business domain, which can include one of the following: a financial domain, a medical domain, an education domain, a logistics domain, etc.

[0064] Optionally, the analysis result can only include the intention of the first object.

[0065] In this way, the intention analysis model is used to analyze the multi-turn conversation to be analyzed in the current multi-turn conversation to obtain the intention of the first object, so that the efficiency and accuracy of obtaining the intention of the first object can be improved.

[0066] In an embodiment, the input information is generated based on the multi-turn conversation to be analyzed, including: inputting the multi-turn conversation to be analyzed into a keyword extraction model to obtain target keywords output by the keyword extraction model; and in a case where one or more target reference questions matching the target keywords exist in a plurality of preset entries contained in the database, taking the multi-turn conversation to be analyzed and the one or more target reference questions as the input information.

[0067] Each of the plurality of preset entries includes a preset question and a corresponding preset answer.

[0068] The keyword extraction model can be obtained by training a second preset model, where the second preset model can be a large language model. The specific training method of the second preset model is not limited in the present application.

[0069] The database contains a plurality of preset entries, where each of the plurality of preset entries includes a preset question and a corresponding preset answer, and different preset entries in the plurality of preset entries include different preset questions and / or different preset answers. The plurality of preset entries at least include a plurality of preset entries of a specified business domain; in addition, the plurality of preset entries can or can not include preset entries of other business domains, which is not limited here.

[0070] In addition to containing a plurality of preset entries, the database further includes at least one of the following: a vector of each preset entry in the plurality of preset entries, a preset service type associated with each preset entry, one or more reference multi-turn dialogues associated with each preset entry, and a vector of each reference multi-turn dialogue in the one or more reference multi-turn dialogues.

[0071] The vector of each preset entry in the plurality of preset entries can be obtained in one of the following ways: an agent in the server inputs each preset entry in the plurality of preset entries into a vector extraction model to obtain a vector of each preset entry in the plurality of preset entries output by the vector extraction model; or the agent in the server calculates each preset entry using a preset vector calculation method to obtain a vector of each preset entry. The vector extraction model can be a BGE-large embedding (Bag of Global Embeddings-Large Embedding) model, and the vector calculation method is not limited in this embodiment.

[0072] The one or more reference multi-turn dialogues associated with each preset entry can be obtained in the following way: the agent in the server inputs each preset entry into a dialogue generation model to obtain one or more reference multi-turn dialogues associated with each preset entry output by the dialogue generation model. The dialogue generation model can be trained based on a third preset model, and the third preset model can be a large language model. The training method for the third preset model is not limited in this application.

[0073] Determining whether there is a reference question matching the target keyword in the plurality of preset entries included in the database can include: determining, by the agent in the server, whether the target keyword is included in a preset question of each preset entry in the plurality of preset entries included in the database; in a case where the target keyword is included in a preset question of one or more first entries in the plurality of preset entries, taking the preset question in the one or more first entries as the one or more target reference questions; and in a case where the target keyword is not included in a preset question of each preset entry in the plurality of preset entries, determining that there is no reference question matching the target keyword in the plurality of preset entries included in the database.

[0074] In an example, the input information can include the one or more target reference questions and the multi-turn dialogue to be analyzed.

[0075] In addition, the input information can further include the multi-turn dialogue to be analyzed in a case where the database does not contain a reference question matching the target keyword.

[0076] In this way, the target reference question related to the multi-turn dialogue to be analyzed can be retrieved through the target keyword corresponding to the multi-turn dialogue to be analyzed, and the input information generated through the target reference question and the multi-turn dialogue to be analyzed can enable the intent analysis model to obtain more reference information, thereby improving the accuracy of the intent analysis model in analyzing the intent of the first object.

[0077] In an embodiment, the training manner of the intent analysis model can include: inputting a training sample into a first preset model to obtain a prediction analysis result output by the first preset model, wherein the training sample includes a historical multi-turn dialogue between a third object and a fourth object, the third object is a service receiver, and the fourth object is a service provider, and the prediction analysis result includes a predicted intent of the third object; obtaining a target loss based on the prediction analysis result and a plurality of sample labels corresponding to the training sample, wherein the plurality of sample labels corresponding to the training sample are obtained based on a plurality of large models, a parameter quantity of each large model in the plurality of large models is different from a parameter quantity of the first preset model, and a parameter quantity of each large model in the plurality of large models is different from a parameter quantity of another large model in the plurality of large models; and training the first preset model based on the target loss to obtain the intent analysis model.

[0078] The parameter quantity of each large model in the plurality of large models can be greater than the parameter quantity of the first preset model.

[0079] The number of the training samples can be one or more.

[0080] For example, the number of the training samples is one.

[0081] Optionally, the training sample can include the historical multi-turn dialogue between the third object and the fourth object.

[0082] The third object can be a second user, and the second user can be a user who has been historically served. The second user and the first user can be the same or different, and the present application does not make any limitation.

[0083] The fourth object is a service provider, i.e., the fourth object can be a service provider providing services for the third object; specifically, the fourth object can be a customer service, and the fourth object and the second object can be the same customer service or different customer services, which is not limited in the present application.

[0084] Optionally, the training sample can include a historical multi-turn dialogue between the third object and the fourth object, and one or more historical reference questions associated with the historical multi-turn dialogue.

[0085] The manner of obtaining the one or more historical reference questions associated with the historical multi-turn dialogue can include: inputting the historical multi-turn dialogue between the third object and the fourth object into a keyword extraction model to obtain keywords corresponding to the historical multi-turn dialogue output by the keyword extraction model; determining whether there is a reference question matching the keywords corresponding to the historical multi-turn dialogue in a plurality of preset entries contained in the database; and in the case that there is one or more historical reference questions matching the keywords corresponding to the historical multi-turn dialogue in the plurality of preset entries contained in the database, extracting the one or more historical reference questions associated with the historical multi-turn dialogue. The manner of determining whether there is a reference question matching the keywords corresponding to the historical multi-turn dialogue in the plurality of preset entries contained in the database is the same as the manner of determining whether there is a reference question matching the target keywords in the plurality of preset entries contained in the database, which will not be repeated here.

[0086] For example, the number of training samples is a plurality.

[0087] Optionally, each training sample in the plurality of training samples can include a historical multi-turn dialogue between the third object and the fourth object.

[0088] Optionally, each training sample in the plurality of training samples can include a historical multi-turn dialogue between the third object and the fourth object, and one or more historical reference questions associated with the historical multi-turn dialogue. The one or more historical reference questions associated with the historical multi-turn dialogue included in different training samples in the plurality of training samples are different, and the manner of obtaining the one or more historical reference questions associated with the historical multi-turn dialogue included in each training sample is the same as the above example, which will not be repeated here.

[0089] In the scenario of the plurality of training samples, the historical multi-turn dialogue included in different training samples in the plurality of training samples is different; and the third object corresponding to different training samples in the plurality of training samples is different, and / or the fourth object corresponding to different training samples in the plurality of training samples is different, wherein the related descriptions of the third object and the fourth object are similar to the above examples, which will not be repeated here.

[0090] The determination of the multiple sample labels corresponding to each training sample is the same, and taking any one training sample as an example, the determination of the multiple sample labels corresponding to the training sample can include: inputting the qth training sample into the multiple large models respectively to obtain the sample labels of the qth training sample output by the multiple large models respectively. The large model can also be referred to as a large language model.

[0091] The multiple large models are exemplarily illustrated: assuming that the number of the multiple large models is three, the parameter quantity of the first large model can be a model of the order of magnitude of ten billion parameters, the parameter quantity of the second large model can be a model of the order of magnitude of one hundred billion parameters, and the parameter quantity of the third large model can be a model of the order of magnitude of one thousand billion parameters. The parameter quantity of the first preset model can be less than ten billion parameters.

[0092] The input of the training sample into the first preset model to obtain the prediction analysis result output by the first preset model is the same for each training sample. Still taking any one training sample as an example, the input of the qth training sample into the first preset model to obtain the prediction analysis result output by the first preset model includes: inputting the qth training sample and the prompt word into the first preset model to obtain the prediction analysis result of the qth training sample output by the first preset model. The prompt word is the same as described above, and will not be described herein.

[0093] Optionally, the prediction analysis result of the qth training sample includes the predicted intention of the third object of the qth training sample and the predicted content summary of the qth training sample. The predicted content summary is similar to the related description of the content summary described above, and the related description of the predicted intention of the third object is similar to the related description of the intention of the first object described above, and will not be described herein.

[0094] Optionally, the prediction analysis result of the qth training sample includes the predicted intention of the third object of the qth training sample.

[0095] In an implementation, the obtaining of the target loss based on the prediction analysis result and the multiple sample labels corresponding to the training sample includes: calculating the probability distribution between the prediction analysis result and the training sample, and the probability distribution between the multiple sample labels and the training sample; and obtaining the target loss based on the probability distribution between the prediction analysis result and the training sample, and the probability distribution between the multiple sample labels and the training sample.

[0096] In an example, the number of training samples is one.

[0097] In the case where the predicted intention of the third object and the predicted content summary of the training sample are included in the predicted output result, the obtaining the target loss based on the probability distribution between the predicted analysis result and the training sample and the probability distribution between the plurality of sample labels and the training sample can include: calculating a loss corresponding to each sample label in the plurality of sample labels based on the probability distribution between the predicted analysis result and the training sample and the probability distribution between the plurality of sample labels and the training sample; and adding the loss corresponding to each sample label to obtain the target loss.

[0098] The plurality of sample labels are obtained by processing the training sample by each of the plurality of large models, and thus each sample label can also be referred to as a sample label corresponding to each large model.

[0099] Optionally, taking any one of the plurality of sample labels as the i th sample label (or the sample label corresponding to the i th large model) as an example, the calculating the loss corresponding to each sample label in the plurality of sample labels based on the probability distribution between the predicted analysis result and the training sample and the probability distribution between the plurality of sample labels and the training sample includes: obtaining an i th first sub-loss based on the probability distribution between the predicted content summary in the predicted analysis result and the training sample and the probability distribution between the content summary label included in the i th sample label and the training sample; obtaining an i th second sub-loss based on the probability distribution between the predicted intention in the predicted analysis result and the training sample and the probability distribution between the intention label included in the i th sample label and the training sample; and taking the sum of the i th first sub-loss and the i th second sub-loss as the loss corresponding to the i th sample label. The i is a positive integer.

[0100] The obtaining the i th first sub-loss based on the probability distribution between the predicted content summary in the predicted analysis result and the training sample and the probability distribution between the content summary label included in the i th sample label and the training sample includes: calculating, based on a divergence function, the probability distribution between the predicted content summary in the predicted analysis result and the training sample, the probability distribution between the content summary label included in the i th sample label in the plurality of sample labels and the training sample, and a first weight corresponding to the i th large model to obtain the i th first sub-loss. The divergence function can be a Kullback-Leibler (KL) divergence function. The first weight corresponding to different large models in the plurality of large models is different, and the first weight corresponding to each large model in the plurality of large models is related to the parameter amount of each large model, for example: the larger the parameter amount of a certain large model, the higher the first weight corresponding to the large model, and vice versa.

[0101] The method of obtaining the i-th second sub-loss based on the probability distribution between the predicted intent in the prediction analysis result and the training sample, and the probability distribution between the intent annotation contained in the i-th sample annotation and the training sample, includes: calculating the probability distribution between the predicted intent in the prediction analysis result and the training sample, the probability distribution between the intent annotation contained in the i-th sample annotation and the training sample, and the second weight corresponding to the i-th large model based on the divergence function to obtain the i-th second sub-loss. The second weights corresponding to different large models in the multiple large models are different, and the second weight corresponding to each large model in the multiple large models is related to the parameter amount of each large model. For example, the larger the parameter amount of a large model, the higher the corresponding second weight.

[0102] Optionally, taking any one of the multiple sample annotations as the i-th sample annotation as an example, the calculation of the loss corresponding to each sample annotation in the multiple sample annotations based on the probability distribution between the prediction analysis result and the training sample, and the probability distribution between the multiple sample annotations and the training sample includes: obtaining the i-th second sub-loss based on the probability distribution between the predicted intent in the prediction analysis result and the training sample, and the probability distribution between the intent annotation included in the i-th sample annotation and the training sample; and using the i-th second sub-loss as the loss corresponding to the i-th sample annotation. The calculation method of the i-th second sub-loss is the same as that in the above example and will not be repeated here.

[0103] In one example, the number of training samples is multiple.

[0104] In the case where the prediction output result includes the prediction intention of the third object and the prediction content summary of the training sample, obtaining the target loss based on the probability distribution between the prediction analysis result and the training sample, and the probability distribution between the multiple sample annotations and the training sample may include: calculating the loss corresponding to each sample annotation in the multiple sample annotations of each training sample in the multiple training samples based on the probability distribution between the prediction analysis results corresponding to the multiple training samples and each training sample in the multiple training samples, and the probability distribution between the multiple sample annotations of each training sample and each training sample; adding the loss corresponding to each sample annotation in the multiple sample annotations of each training sample to obtain the total loss corresponding to each training sample; and adding the total loss corresponding to each training sample to obtain the target loss.

[0105] Among them, the multiple sample annotations of each training sample are obtained by each large model in the multiple large models processing the training sample, so each sample annotation of each training sample can also be called the sample annotation corresponding to each large model under the training sample.

[0106] Optionally, taking any one of the plurality of training samples as a qth training sample, and any one of the plurality of sample labels of the qth training sample as an ith sample label (or a sample label corresponding to the ith large model under the qth training sample) as an example, the calculating of the loss corresponding to each sample label in the plurality of sample labels of each training sample in the plurality of training samples based on the probability distribution between the prediction analysis result corresponding to each training sample in the plurality of training samples and each training sample, and the probability distribution between the plurality of sample labels of each training sample and each training sample includes: obtaining an ith first sub-loss corresponding to the ith sample label of the qth training sample based on the probability distribution between the predicted content summary in the prediction analysis result corresponding to the qth training sample and the qth training sample, and the probability distribution between the content summary label included in the ith sample label of the qth training sample and the qth training sample; obtaining an ith second sub-loss corresponding to the ith sample label of the qth training sample based on the probability distribution between the predicted intent in the prediction analysis result corresponding to the qth training sample and the qth training sample, and the probability distribution between the intent label included in the ith sample label of the qth training sample and the qth training sample; and taking the sum of the ith first sub-loss and the ith second sub-loss as the loss corresponding to the ith sample label of the qth training sample. Wherein q is a positive integer. The calculation method of the ith first sub-loss and the ith second sub-loss is similar to the above example, which will not be described here.

[0107] Optionally, taking any one of the plurality of training samples as a qth training sample, and any one of the plurality of sample labels of the qth training sample as an ith sample label as an example, the calculating of the loss corresponding to each sample label in the plurality of sample labels of each training sample in the plurality of training samples based on the probability distribution between the prediction analysis result corresponding to each training sample in the plurality of training samples and each training sample, and the probability distribution between the plurality of sample labels of each training sample and each training sample includes: obtaining an ith second sub-loss corresponding to the ith sample label of the qth training sample based on the probability distribution between the predicted intent in the prediction analysis result corresponding to the qth training sample and the qth training sample, and the probability distribution between the intent label included in the ith sample label of the qth training sample and the qth training sample; and taking the ith second sub-loss as the loss corresponding to the ith sample label of the qth training sample. The calculation method of the ith second sub-loss is similar to the above example, which will not be described here.

[0108] Still taking any one of the plurality of training samples as the qth training sample as an example, the adding of the loss corresponding to each sample label in the plurality of sample labels of each training sample to obtain the total loss corresponding to each training sample comprises: adding the loss corresponding to each sample label in the plurality of sample labels of the qth training sample to obtain the total loss corresponding to the qth training sample.

[0109] In the case where the plurality of large models is three, the total loss is described in combination with the formula.

[0110]

[0111] wherein, is a target loss. i can be used to refer to different large models, i∈b,tb,hb respectively represent a large model of a billion parameter order (billion, b), a large model of a ten billion parameter order (ten billion, tb), and a large model of a hundred billion parameter order (hundred billion, hb).

[0112] s q is a prediction content summary in the prediction analysis result corresponding to the qth training sample, x q is the qth training sample, s i q is a content summary label of the qth training sample output by the ith large model, q takes a value in a range from 1 to N, and θ represents a parameter of the first preset model; P θ (s q |x q ) is a probability distribution between the content summary in the prediction analysis result corresponding to the qth training sample and the qth training sample, P i (s i q |x q ) is a probability distribution between the content summary label of the sample label corresponding to the ith large model under the qth training sample and the qth training sample; D KL is a divergence function, α i is the first weight corresponding to the ith large model, α i ·D KL (P θ (s q |x q )||P i (s i q |x q )) is the ith first sub-loss corresponding to the ith sample label of the qth training sample.

[0113] r q is a prediction intention in the prediction analysis result corresponding to the qth training sample, ri q an intent label in the sample label corresponding to the i-th large model under the q-th training sample, β i is a second weight corresponding to the i-th large model; β i ·D KL (P θ (r q |x q )||P i (r i q |x q )) is the i-th second sub-loss corresponding to the i-th sample label of the q-th training sample.

[0114] When the training sample is one (i.e., N = 1), only the sum of the losses corresponding to each sample label needs to be calculated, that is, the target loss is obtained by adding the losses corresponding to each sample label. When the training sample is multiple (i.e., N is a positive integer greater than 1), the sum of the losses corresponding to each sample label in the multiple sample labels of the N training samples needs to be calculated, that is, the total loss corresponding to each training sample is obtained by adding the losses corresponding to each sample label in the multiple sample labels of the N training samples; the target loss is obtained by adding the total losses corresponding to the N training samples.

[0115] The training of the first preset model based on the target loss to obtain the intent analysis model can include: adjusting the parameters of the first preset model based on the target loss; in the case of meeting the convergence condition, obtaining the trained first preset model, and taking the trained first preset model as the intent analysis model. The convergence condition can be at least one of the following: the target loss is equal to zero; the target loss is less than or equal to a threshold; the number of iterations of training reaches a preset number, and the like.

[0116] In this way, the sample labels output by the multiple large models with large parameter quantities are used to train the first preset model with small parameter quantities, and through such distillation training, the first preset model with small parameter quantities can learn the capabilities of the large models. In addition, the target loss is calculated based on the probability distribution between the prediction analysis result and the training sample and the probability distribution between the multiple sample labels and the training sample, which can ensure the accuracy of the target loss and further ensure that the trained intent analysis model can learn the capabilities of the multiple large models more accurately, thereby improving the accuracy of the trained intent analysis model in predicting intents.

[0117] ​​In addition, the scenarios to which this application applies are particularly suitable for voice channels, so there are higher requirements for real-time performance. Although the model with a larger number of parameters is more accurate in summarizing, it cannot meet the real-time requirements in actual use. Although the first preset model with a relatively small number of parameters is fast, it lacks in the accuracy of the reply content and the standardization of the reply format. Therefore, by training the first preset model in the above manner, the above problems can be solved, taking into account speed while ensuring the accuracy of the predicted intention.

[0118] In one example, the method further includes: fine-tuning the intent analysis model based on the fine-tuning sample to obtain a fine-tuned intent analysis model. The fine-tuning method can be LoRA (Low-Rank Adaptation of Large Language Models). The specific fine-tuning process is not limited in this application.

[0119] The fine-tuning sample includes at least one of the following: multiple preset entries in the database; multiple general domain question-answer pairs. The number of general domain question-answer pairs can be a specified multiple of the multiple preset entries. The specified multiple can be set based on actual circumstances and is not limited in this application. For example, it can be 5 times. The method for obtaining the general domain question-answer pairs is not limited in this application.

[0120] Since there are many proper nouns in the designated business field, and for the intent analysis model, some professional vocabulary and expressions are not easily triggered and expressed during the reasoning process of the intent analysis model, in order to solve this problem, multiple preset entries in the database can be used as fine-tuning samples. In this way, the summary content of the intent analysis model can be closer to the preset entries being retrieved to improve the retrieval effect. In addition, in order to make the fine-tuned intent analysis model more general, general domain question and answer pairs of specified multiples of multiple preset entries are used together with multiple preset entries as fine-tuning samples to fine-tune the intent analysis model, and the method can also reduce the cost of fine-tuning (time cost, computing resources).

[0121] In one embodiment, the method of determining multiple candidate entries and reference multi-round conversations associated with the multiple candidate entries from a database based on the intention of the first object includes: calculating the similarity between multiple preset entries included in the database and the intention of the first object, wherein the similarity includes text similarity and / or vector similarity; selecting multiple candidate entries from the multiple preset entries based on the similarity between the multiple preset entries and the intention of the first object; and extracting reference multi-round conversations associated with the multiple candidate entries from the database.

[0122] Each of the plurality of preset entries includes a preset question and a corresponding preset answer.

[0123] The text similarity can be calculated by a BM25 (Best Match 25) algorithm, and the vector similarity can be calculated by one of a cosine distance, a Manhattan algorithm, and the like.

[0124] Optionally, the selecting the plurality of candidate entries from the plurality of preset entries based on the similarity between the plurality of preset entries and the intent of the first object includes: in a case where the similarity includes a text similarity, the agent in the server sorts, from high to low, the text similarity between the plurality of preset entries included in the database and the intent of the first object to obtain a first sorting result of the plurality of preset entries; and selecting a first quantity of second entries from the plurality of preset entries based on the first sorting result, the first quantity being a positive integer greater than or equal to 2, and the first quantity being set according to an actual situation.

[0125] Optionally, the selecting the plurality of candidate entries from the plurality of preset entries based on the similarity between the plurality of preset entries and the intent of the first object includes: in a case where the similarity includes a vector similarity, the agent in the server sorts, from high to low, the vector similarity between the plurality of preset entries included in the database and the intent of the first object to obtain a second sorting result of the plurality of preset entries; and selecting a second quantity of third entries from the plurality of preset entries based on the second sorting result, the second quantity being a positive integer greater than or equal to 2, and the second quantity being set according to an actual situation.

[0126] Optionally, the selecting the plurality of candidate entries from the plurality of preset entries based on the similarity between the plurality of preset entries and the intent of the first object includes: in a case where the similarity includes a text similarity and a vector similarity, the agent in the server sorts, from high to low, the text similarity between the plurality of preset entries included in the database and the intent of the first object to obtain a first sorting result of the plurality of preset entries; selecting a first quantity of second entries from the plurality of preset entries based on the first sorting result; sorting, from high to low, the vector similarity between the plurality of preset entries included in the database and the intent of the first object to obtain a second sorting result of the plurality of preset entries; selecting a second quantity of third entries from the plurality of preset entries based on the second sorting result; and taking a union of the first quantity of second entries and the second quantity of third entries as the plurality of candidate entries.

[0127] In one example, before selecting multiple candidate entries from the multiple preset entries based on the similarity between the multiple preset entries and the intention of the first object, the method further includes: an agent in the server obtains a target service type corresponding to the first object. The target service type corresponding to the first object may be a service type selected by the first object by pressing a button when making a call using the first device, where the button is a physical button or a virtual button; or the target service type corresponding to the first object may be a service type provided by the second device or a default, etc. The method for obtaining or determining the target service type corresponding to the first object is not limited herein.

[0128] The method of selecting multiple candidate entries from the multiple preset entries based on the similarity between the multiple preset entries and the intention of the first object includes: the intelligent agent in the server selects multiple preset entries corresponding to the target business type from all preset entries in the database, and selects multiple candidate entries from the multiple preset entries corresponding to the target business type based on the similarity between each preset entry in the multiple preset entries corresponding to the target business type and the intention of the first object. The method of selecting multiple candidate entries from the multiple preset entries corresponding to the target business type based on the similarity between each preset entry in the multiple preset entries corresponding to the target business type and the intention of the first object is the same as the method of selecting multiple candidate entries from the multiple preset entries based on the similarity between each preset entry and the intention of the first object, and will not be repeated here.

[0129] In this way, multiple candidate entries are selected from the multiple preset entries included in the database based on the similarity between each preset entry and the first subject's intent, and reference multi-turn conversations associated with the multiple candidate entries are extracted from the database. This allows for more accurate identification of multiple candidate entries similar to the first subject's intent and their associated reference multi-turn conversations.

[0130] In one embodiment, the method of determining one or more reference items from the multiple candidate items based on the intention of the first object and the reference multi-round conversations associated with the multiple candidate items includes: determining initial ranking scores of the multiple candidate items based on the similarity between the intention of the first object and the multiple candidate items; adjusting the initial ranking scores of the multiple candidate items based on the similarity between the intention of the first object and the reference multi-round conversations associated with the multiple candidate items to obtain target ranking scores of the multiple candidate items; and determining one or more reference items from the multiple candidate items based on the target ranking scores of the multiple candidate items.

[0131] Optionally, the determining the initial ranking scores of the plurality of candidate entries based on the similarity between the intent of the first object and the plurality of candidate entries comprises: in the case that the similarity is text similarity, the agent in the server taking the text similarity between each candidate entry in the plurality of candidate entries and the intent of the first object as the initial ranking score of the each candidate entry.

[0132] Optionally, the determining the initial ranking scores of the plurality of candidate entries based on the similarity between the intent of the first object and the plurality of candidate entries comprises: in the case that the similarity is vector similarity, the agent in the server taking the vector similarity between each candidate entry in the plurality of candidate entries and the intent of the first object as the initial ranking score of the each candidate entry.

[0133] Optionally, the determining the initial ranking scores of the plurality of candidate entries based on the similarity between the intent of the first object and the plurality of candidate entries comprises: in the case that the similarity comprises text similarity and vector similarity, the agent in the server determining the initial ranking score of each candidate entry in the plurality of candidate entries based on the text similarity and / or the vector similarity between the each candidate entry and the intent of the first object.

[0134] In combination with the foregoing embodiments, in the case that the similarity comprises text similarity and vector similarity, the plurality of candidate entries can comprise a first number of second entries selected based on the highest text similarity, a second number of third entries selected based on the highest vector similarity, any one second entry can be the same as or different from one of the second number of third entries, and any one third entry can be the same as or different from one of the first number of second entries. Therefore, the composition of the plurality of candidate entries can comprise at least one of the following cases: a part of the candidate entries only have corresponding text similarity, a part of the candidate entries only have corresponding vector similarity, and a part of the candidate entries have both text similarity and vector similarity.

[0135] Taking a kth candidate item in the plurality of candidate items as an example, the determining, based on the text similarity and / or the vector similarity between each candidate item in the plurality of candidate items and the text of the intention of the first object, of an initial ranking score of the each candidate item can include one of: in a case where the kth candidate item has a corresponding vector similarity and a corresponding text similarity, determining, by an agent in the server, the initial ranking score of the kth candidate item based on the vector similarity and the text similarity corresponding to the kth candidate item; in a case where the kth candidate item has only the corresponding vector similarity, taking the vector similarity corresponding to the kth candidate item as the initial ranking score of the kth candidate item; and in a case where the kth candidate item has only the corresponding text similarity, taking the text similarity corresponding to the kth candidate item as the initial ranking score of the kth candidate item. The k is a positive integer greater than or equal to 1.

[0136] The determining, based on the vector similarity and the text similarity corresponding to the kth candidate item, of the initial ranking score of the kth candidate item can include one of: taking, by the agent in the server, an average of the vector similarity and the text similarity corresponding to the kth candidate item as the initial ranking score of the kth candidate item; and taking, by the agent in the server, a maximum value of the vector similarity and the text similarity corresponding to the kth candidate item as the initial ranking score of the kth candidate item.

[0137] In an example, the adjusting, based on the similarity between the intention of the first object and the reference multi-turn dialogues associated with the plurality of candidate items, of the initial ranking scores of the plurality of candidate items to obtain target ranking scores of the plurality of candidate items includes: determining adjustment parameters corresponding to the plurality of candidate items based on the similarity between the intention of the first object and the reference multi-turn dialogues associated with the plurality of candidate items, a weight value of the reference multi-turn dialogues associated with the plurality of candidate items, and an influence factor of the reference multi-turn dialogues associated with the plurality of candidate items on the plurality of candidate items; and adjusting, based on the adjustment parameters corresponding to the plurality of candidate items, the initial ranking scores of the plurality of candidate items to obtain the target ranking scores of the plurality of candidate items.

[0138] The number of the reference multi-turn dialogues associated with each candidate item in the plurality of candidate items can be one or more. The similarity between the intention of the first object and the reference multi-turn dialogues associated with the plurality of candidate items can be a similarity between a vector of the intention of the first object and a vector of each of one or more reference multi-turn dialogues associated with each candidate item in the plurality of candidate items.

[0139] The method for determining the weight value of the reference multi-round dialogue associated with the multiple candidate entries includes: the intelligent agent in the server determines the weight value of the reference multi-round dialogue associated with each candidate entry in the multiple candidate entries based on the specified features of each reference multi-round dialogue in one or more reference multi-round dialogues associated with each candidate entry in the multiple candidate entries. Among them, the specified features of any reference multi-round dialogue may include at least one of the following: the length of any reference multi-round dialogue, the number of key information contained in any reference multi-round dialogue, and the key information may be the relevant information in the candidate entry. The specific calculation method is not limited in this application. It can be understood that the longer the length of any reference multi-round dialogue, the greater its weight, and the more key information contained in any reference multi-round dialogue, the greater its weight.

[0140] The influence factor of the reference multi-round dialogue associated with each candidate entry in the multiple candidate entries on each candidate entry is used to indicate the degree of association between the reference multi-round dialogue associated with each candidate entry and each candidate entry. The specific setting can be based on actual conditions and is not limited in this application.

[0141] Taking the kth candidate entry among multiple candidate entries as an example, the adjustment parameters corresponding to the multiple candidate entries are determined based on the similarity between the intention of the first object and the reference multi-round conversations associated with the multiple candidate entries, the weight values ​​of the reference multi-round conversations associated with the multiple candidate entries, and the influence factors of the reference multi-round conversations associated with the multiple candidate entries on the multiple candidate entries. It can include: the intelligent agent in the server obtains one or more first values ​​based on the similarity between the intention of the first object and one or more reference multi-round conversations associated with the kth candidate entry, the weight values ​​of each reference multi-round conversation in the one or more reference multi-round conversations associated with the kth candidate entry, and the influence factor of each reference multi-round conversation associated with the kth candidate entry on the kth candidate entry; and determines the adjustment parameter corresponding to the kth candidate entry based on the one or more first values ​​and the first adjustment weight.

[0142] In the case that there is one reference multi-round dialogue associated with the k-th candidate entry, one or more first values ​​are obtained based on the similarity between the intention of the first object and one or more reference multi-round dialogues associated with the k-th candidate entry, the weight value of each reference multi-round dialogue in the one or more reference multi-round dialogues associated with the k-th candidate entry, and the influence factor of each reference multi-round dialogue associated with the k-th candidate entry on the k-th candidate entry, including: the intelligent agent in the server multiplies the similarity between the intention of the first object and the reference multi-round dialogue associated with the k-th candidate entry, the weight value of the reference multi-round dialogue associated with the k-th candidate entry, and the influence factor of the reference multi-round dialogue associated with the k-th candidate entry to obtain the first value.

[0143] The determining the adjustment parameter corresponding to the kth candidate entry based on the one or more first values and the first adjustment weight includes: an agent in the server multiplying the first value and the first adjustment weight to obtain the adjustment parameter corresponding to the kth candidate entry.

[0144] In a case where the reference multi-turn dialogues associated with the kth candidate entry are multiple, taking an arbitrary one of the reference multi-turn dialogues as the jth reference multi-turn dialogue as an example, the obtaining the one or more first values based on the similarity between the intent of the first object and the one or more reference multi-turn dialogues associated with the kth candidate entry, the weight value of each of the one or more reference multi-turn dialogues associated with the kth candidate entry, and the influence factor of each of the reference multi-turn dialogues associated with the kth candidate entry on the kth candidate entry includes: an agent in the server multiplying the similarity between the intent of the first object and the jth reference multi-turn dialogue associated with the kth candidate entry, the weight value of the jth reference multi-turn dialogue associated with the kth candidate entry, and the influence factor of the jth reference multi-turn dialogue associated with the kth candidate entry on the kth candidate entry to obtain the jth first value.

[0145] The determining the adjustment parameter corresponding to the kth candidate entry based on the one or more first values and the first adjustment weight includes: an agent in the server summing the one or more first values to obtain a second value; and multiplying the second value and the first adjustment weight to obtain the adjustment parameter corresponding to the kth candidate entry.

[0146] Still taking the kth candidate entry in the multiple candidate entries as an example, the adjusting the initial ranking scores of the multiple candidate entries based on the adjustment parameters corresponding to the multiple candidate entries to obtain target ranking scores of the multiple candidate entries includes: an agent in the server multiplying the initial ranking score of the kth candidate entry and a second adjustment weight to obtain an intermediate ranking score of the kth candidate entry; and adding the intermediate ranking score of the kth candidate entry and the adjustment parameter corresponding to the kth candidate entry to obtain the target ranking score of the kth candidate entry.

[0147] The first adjustment weight and the second adjustment weight are related, and are both used to adjust the degree of influence of the reference multi-turn dialogue on the initial ranking score. The second adjustment weight can be equal to 1 minus the first adjustment weight, and the value of the first adjustment weight can be set according to actual conditions, as long as it is in the range of 0 to 1 (including 0 and 1). For example: assuming that the first adjustment weight is represented as λ, when λ = 0, the second adjustment weight is equal to 1, at this time, the initial ranking score is completely ranked according to the initial ranking score (that is, the initial ranking score is taken as the target ranking score), and the influence of the reference multi-turn dialogue is not considered. When λ = 1, the second adjustment weight is equal to 0, at this time, the target ranking score is completely determined according to the similarity between the reference multi-turn dialogue and the intent of the first object, and the initial ranking score is ignored. When 0 < λ < 1, the first adjustment weight and the second adjustment weight are both greater than 0 and less than 1, at this time, both the similarity between the reference multi-turn dialogue and the intent of the first object and the initial ranking score are considered.

[0148] Taking the kth candidate item in the plurality of candidate items as an example, combining the formula, the initial ranking score of the plurality of candidate items is adjusted based on the adjustment parameter corresponding to the plurality of candidate items to obtain the target ranking score of the plurality of candidate items, and the adjustment is described as follows:

[0149]

[0150] Wherein, λ is the first adjustment weight, (1-λ) is the second adjustment weight, b k is the kth candidate item in the plurality of candidate items; OriginalScore(b k ) (original score (b k )) is the initial ranking score of the kth candidate item, d j k is the jth reference multi-turn dialogue associated with the kth candidate item (the value of j is in the range of 1 to m); S(d j k ) is the similarity between the intent of the first object and the jth reference multi-turn dialogue associated with the kth candidate item; W(d j k ) is the weight value of the jth reference multi-turn dialogue corresponding to the kth candidate item; Influence(b k ,d j k ) (influence (b k ,d j k ) is the influence factor of the jth reference multi-turn dialogue corresponding to the kth candidate item on the kth candidate item; Score(b k ) (score (b k ) is the target ranking score of the kth candidate item.

[0151] The determining the one or more reference entries from the plurality of candidate entries based on the target ranking scores of the plurality of candidate entries can include that the agent in the server ranks the plurality of candidate entries in descending order of the target ranking scores, to obtain a ranking result of the plurality of candidate entries; and the agent in the server takes the first H candidate entries in the ranking result of the plurality of candidate entries as the one or more reference entries, where H can be a positive integer greater than or equal to 1 and can be set according to actual conditions. For example, the first three candidate entries in the plurality of candidate entries in terms of the target ranking scores can be selected as the three reference entries.

[0152] In this way, the reference multi-turn dialogue can represent the relevant scene of the associated candidate entry, so that the initial ranking score is adjusted by the similarity between the intent of the first object and the reference multi-turn dialogue associated with the plurality of candidate entries, and the parameter related to the reference multi-turn dialogue is further combined in the adjustment process, so that the adjusted target ranking score is more accurate.

[0153] In an embodiment, the method further includes: inputting, by the agent in the server, the current multi-turn dialogue, the intent of the first object, and the one or more reference entries into an enhanced retrieval model to obtain a question reference reply output by the enhanced retrieval model, and sending the question reference reply to the second device. The enhanced retrieval model can be obtained by training a third preset model, and the specific training method is not limited in the present application. The third preset model can be a RAG (Retrieval Augmented Generation) model.

[0154] In combination with Figure 2 The above embodiments are described, including:

[0155] S201, the second device sends a current multi-turn dialogue between a first object and a second object to a server.

[0156] The agent in the server performs the processing of S202 to S210, and the details are as follows:

[0157] S202, receiving a current multi-turn dialogue between a first object and a second object.

[0158] S203, determining a multi-turn dialogue to be analyzed based on the current multi-turn dialogue.

[0159] S204, inputting the multi-turn dialogue to be analyzed into a keyword extraction model to obtain a target keyword output by the keyword extraction model.

[0160] S205: Using the multiple rounds of conversations to be analyzed and the one or more target reference questions as the input information; or using the multiple rounds of conversations to be analyzed as the input information. The one or more target reference questions are target reference questions that match the target keyword from a plurality of preset entries included in a database.

[0161] S206: Input the input information into an intention analysis model to obtain analysis results corresponding to the multiple rounds of conversations to be analyzed output by the intention analysis model, wherein the analysis results include the intention of the first object.

[0162] S207: Extract the intention of the first object from the analysis result.

[0163] S208: Obtain the target service type corresponding to the first object from the second device. It is understandable that as long as the service type corresponding to the first object is obtained before S209, it is within the scope of this application.

[0164] S209, selecting multiple preset entries corresponding to the target business type from all preset entries in the database, and selecting multiple candidate entries from the multiple preset entries corresponding to the target business type based on the similarity between the multiple preset entries corresponding to the target business type and the intention of the first object; determining one or more reference entries from the multiple candidate entries based on the intention of the first object and the reference multi-round dialogue associated with the multiple candidate entries; using the one or more reference entries as reference recommendation information, and sending the reference recommendation information to the second device.

[0165] S210, optionally, inputting the current multi-round conversation, the intention of the first object, and one or more reference items into an enhanced retrieval model, obtaining a reference answer to the question output by the enhanced retrieval model, and sending the reference answer to the question to the second device.

[0166] Figure 3 FIG. 1 shows a schematic block diagram of an information recommendation device provided by an embodiment of the present disclosure. Figure 3 Shown, including:

[0167] An intent analysis module 301 is configured to, in response to receiving a current multi-round conversation between a first object and a second object, determine the intent of the first object based on the current multi-round conversation, wherein the first object is a service recipient and the second object is a service provider;

[0168] A retrieval module 302 is configured to determine, from a database, a plurality of candidate entries and reference multi-round conversations associated with the plurality of candidate entries based on the intention of the first object;

[0169] The screening module 303 is configured to determine one or more reference entries from the plurality of candidate entries based on the intention of the first object and reference multi-turn dialogues associated with the plurality of candidate entries.

[0170] The reference recommendation information sending module 304 is configured to send the one or more reference entries as reference recommendation information.

[0171] The screening module is configured to determine initial ranking scores of the plurality of candidate entries based on similarities between the intention of the first object and the plurality of candidate entries, adjust the initial ranking scores of the plurality of candidate entries based on similarities between the intention of the first object and reference multi-turn dialogues associated with the plurality of candidate entries, to obtain target ranking scores of the plurality of candidate entries, and determine one or more reference entries from the plurality of candidate entries based on the target ranking scores of the plurality of candidate entries.

[0172] The screening module is configured to determine adjustment parameters corresponding to the plurality of candidate entries based on similarities between the intention of the first object and reference multi-turn dialogues associated with the plurality of candidate entries, weight values of the reference multi-turn dialogues associated with the plurality of candidate entries, and influence factors of the reference multi-turn dialogues associated with the plurality of candidate entries on the plurality of candidate entries, and adjust the initial ranking scores of the plurality of candidate entries based on the adjustment parameters corresponding to the plurality of candidate entries to obtain target ranking scores of the plurality of candidate entries.

[0173] The retrieval module is configured to calculate similarities between a plurality of preset entries included in the database and the intention of the first object, wherein the similarities include text similarities and / or vector similarities, select a plurality of candidate entries from the plurality of preset entries based on the similarities between the plurality of preset entries and the intention of the first object, and extract reference multi-turn dialogues associated with the plurality of candidate entries from the database.

[0174] The intention analysis module is configured to determine a multi-turn dialogue to be analyzed based on the current multi-turn dialogue, generate input information based on the multi-turn dialogue to be analyzed, and input the input information into an intention analysis model to obtain an analysis result corresponding to the multi-turn dialogue to be analyzed output by the intention analysis model, wherein the analysis result includes the intention of the first object.

[0175] The intention analysis module is configured to input the multi-turn dialogue to be analyzed into a keyword extraction model to obtain target keywords output by the keyword extraction model, and in a case where one or more target reference questions matching the target keywords exist in a plurality of preset entries included in the database, input the multi-turn dialogue to be analyzed and the one or more target reference questions as the input information.

[0176] As Figure 4 The device further comprises:

[0177] The model training module 401 is configured to input a training sample into a first preset model to obtain a prediction analysis result output by the first preset model, wherein the training sample comprises historical multi-round dialogues between a third object and a fourth object, the third object is a service receiver, the fourth object is a service provider, and the prediction analysis result comprises a predicted intention of the third object; based on the prediction analysis result and a plurality of sample labels corresponding to the training sample, a target loss is obtained, wherein the plurality of sample labels corresponding to the training sample are obtained based on a plurality of large models, the plurality of large models have a parameter quantity greater than that of the first preset model, and different large models in the plurality of large models have different parameter quantities; and the first preset model is trained based on the target loss to obtain the intention analysis model.

[0178] The model training module is configured to calculate a probability distribution between the prediction analysis result and the training sample and a probability distribution between the plurality of sample labels and the training sample; and based on the probability distribution between the prediction analysis result and the training sample and the probability distribution between the plurality of sample labels and the training sample, the target loss is obtained.

[0179] The specific functions and examples of the modules and sub-modules of the device of the embodiments of the present disclosure are described above in the corresponding steps of the method embodiments, and will not be described here.

[0180] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.

[0181] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.

[0182] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.

[0183] As Figure 5As shown, the electronic device 500 includes a computing unit 501 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 502 or a computer program loaded into a random access memory (RAM) 503 from a storage unit 508. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0184] A plurality of components in the electronic device 500 are connected to the I / O interface 505, including an input unit 506 such as a keyboard, a mouse, and the like, an output unit 507 such as various types of displays, a speaker, and the like, a storage unit 508 such as a magnetic disk, an optical disk, and the like, and a communication unit 509 such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0185] The computing unit 501 can be various general and / or special-purpose processing components having processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The computing unit 501 performs the various methods and processes described above. For example, in some embodiments, the above-described methods can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, at least one step of the above-described methods can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the above-described methods by any other appropriate means, such as by means of firmware.

[0186] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0187] The program code for implementing the method of the present disclosure can be written in any combination of at least one programming language. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0188] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on at least one line, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0189] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0190] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0191] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0192] It should be understood that various forms of flow shown above can be used, with steps reordered, added, or removed. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, without limitation herein, so long as the desired results of the technology disclosed in the present disclosure are achieved.

[0193] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the principles of the disclosure. Accordingly, modifications, equivalent alternatives and improvements should be included within the scope of the disclosure.

Claims

1. An information recommendation method, comprising: In response to receiving a current multi-round conversation between a first object and a second object, determining an intention of the first object based on the current multi-round conversation, wherein the first object is a service recipient and the second object is a service provider; Based on the intent of the first object, determining from a database a plurality of candidate entries and a reference multi-round dialogue associated with the plurality of candidate entries, wherein the plurality of candidate entries are selected from a plurality of preset entries based on similarities between the plurality of preset entries and the intent of the first object, each of the preset entries including a preset question and a corresponding preset answer; Determining one or more reference items from the multiple candidate items based on the intention of the first object and the reference multiple-round conversations associated with the multiple candidate items; using the one or more reference items as reference recommendation information, and sending the reference recommendation information; Among them, the method of determining one or more reference items from the multiple candidate items based on the intention of the first object and the reference multi-round dialogues associated with the multiple candidate items includes: determining the initial ranking scores of the multiple candidate items based on the similarity between the intention of the first object and the multiple candidate items; determining the adjustment parameters corresponding to the multiple candidate items based on the similarity between the intention of the first object and the reference multi-round dialogues associated with the multiple candidate items, the weight values ​​of the reference multi-round dialogues associated with the multiple candidate items, and the influence factors of the reference multi-round dialogues associated with the multiple candidate items on the multiple candidate items, wherein the weight values ​​are determined based on the specified features of the reference multi-round dialogues, and the influence factors represent the degree of association between the reference multi-round dialogues and the multiple candidate items; adjusting the initial ranking scores of the multiple candidate items based on the adjustment parameters corresponding to the multiple candidate items to obtain the target ranking scores of the multiple candidate items; and determining one or more reference items from the multiple candidate items based on the target ranking scores of the multiple candidate items.

2. The method according to claim 1, wherein The determining, based on the intention of the first object, a plurality of candidate entries and a reference multi-round dialogue associated with the plurality of candidate entries from a database includes: Calculating similarities between a plurality of preset entries included in the database and the intention of the first object, wherein the similarities include text similarity and / or vector similarity; selecting a plurality of candidate entries from the plurality of preset entries based on similarities between the plurality of preset entries and the intention of the first object; Extracting reference multi-turn dialogues associated with the plurality of candidate entries from the database.

3. The method according to claim 1, wherein The determining the intention of the first object based on the current multi-round dialogue includes: Determining multiple rounds of dialogue to be analyzed based on the current multiple rounds of dialogue; Generating input information based on the multiple rounds of conversations to be analyzed; The input information is input into an intention analysis model to obtain analysis results corresponding to the multiple rounds of conversations to be analyzed output by the intention analysis model, wherein the analysis results include the intention of the first object.

4. The method according to claim 3, wherein: Generating input information based on the multiple rounds of conversations to be analyzed includes: Inputting the multiple rounds of conversations to be analyzed into a keyword extraction model to obtain target keywords output by the keyword extraction model; In the case that there are one or more target reference questions matching the target keyword in the plurality of preset entries included in the database, the plurality of rounds of dialogue to be analyzed and the one or more target reference questions are used as the input information.

5. The method according to claim 3 or 4, further comprising: Inputting a training sample into a first preset model to obtain a prediction analysis result output by the first preset model, wherein the training sample includes multiple rounds of historical conversations between a third object and a fourth object, the third object being the service recipient and the fourth object being the service provider, and the prediction analysis result includes a predicted intention of the third object; Obtaining a target loss based on the prediction analysis result and a plurality of sample annotations corresponding to the training sample, wherein the plurality of sample annotations corresponding to the training sample are obtained based on a plurality of large models, the plurality of large models having a greater number of parameters than the first preset model, and different large models in the plurality of large models having different number of parameters; The first preset model is trained based on the target loss to obtain the intention analysis model.

6. The method according to claim 5, wherein: Obtaining a target loss based on the prediction analysis result and a plurality of sample labels corresponding to the training samples includes: Calculating the probability distribution between the prediction analysis result and the training sample, and the probability distribution between the multiple sample annotations and the training sample; The target loss is obtained based on the probability distribution between the prediction analysis result and the training samples, and the probability distribution between the multiple sample annotations and the training samples.

7. An information recommendation device, comprising: an intent analysis module, configured to, in response to receiving a current multi-round conversation between a first object and a second object, determine an intent of the first object based on the current multi-round conversation, wherein the first object is a service recipient and the second object is a service provider; a retrieval module, configured to determine, from a database, a plurality of candidate entries and reference multi-turn conversations associated with the plurality of candidate entries based on the intent of the first object, wherein the plurality of candidate entries are selected from a plurality of preset entries based on similarities between the plurality of preset entries and the intent of the first object, each of the preset entries including a preset question and a corresponding preset answer; a screening module, configured to determine one or more reference items from the plurality of candidate items based on the intention of the first object and the reference multi-round dialogues associated with the plurality of candidate items; a reference recommendation information sending module, configured to use the one or more reference items as reference recommendation information and send the reference recommendation information; The screening module is used to determine one or more reference items from the multiple candidate items based on the intention of the first object and the reference multi-round conversations associated with the multiple candidate items, including: determining initial ranking scores of the multiple candidate items based on the similarity between the intention of the first object and the multiple candidate items; determining adjustment parameters corresponding to the multiple candidate items based on the similarity between the intention of the first object and the reference multi-round conversations associated with the multiple candidate items, weight values ​​of the reference multi-round conversations associated with the multiple candidate items, and influence factors of the reference multi-round conversations associated with the multiple candidate items on the multiple candidate items, wherein the weight values ​​are determined based on specified features of the reference multi-round conversations, and the influence factors represent the degree of association between the reference multi-round conversations and the multiple candidate items; adjusting the initial ranking scores of the multiple candidate items based on the adjustment parameters corresponding to the multiple candidate items to obtain target ranking scores of the multiple candidate items; and determining one or more reference items from the multiple candidate items based on the target ranking scores of the multiple candidate items.

8. The device according to claim 7, wherein The retrieval module is configured to calculate the similarity between a plurality of preset entries included in the database and the intent of the first object, wherein the similarity includes text similarity and / or vector similarity; select a plurality of candidate entries from the plurality of preset entries based on the similarity between the plurality of preset entries and the intent of the first object; and extract a reference multi-round dialogue associated with the plurality of candidate entries from the database.

9. The device according to claim 7, wherein The intention analysis module is used to determine multiple rounds of dialogue to be analyzed based on the current multiple rounds of dialogue; Based on the multiple rounds of conversations to be analyzed, input information is generated; the input information is input into an intention analysis model to obtain an analysis result corresponding to the multiple rounds of conversations to be analyzed output by the intention analysis model, wherein the analysis result includes the intention of the first object.

10. The device according to claim 9, wherein The intention analysis module is used to input the multiple rounds of conversations to be analyzed into a keyword extraction model to obtain the target keywords output by the keyword extraction model; when there are one or more target reference questions matching the target keywords in the multiple preset entries contained in the database, the multiple rounds of conversations to be analyzed and the one or more target reference questions are used as the input information.

11. The apparatus according to claim 9 or 10, further comprising: A model training module is used to input training samples into a first preset model to obtain a predictive analysis result output by the first preset model, wherein the training samples include historical multiple rounds of conversations between a third object and a fourth object, the third object is the service recipient, and the fourth object is the service provider, and the predictive analysis result includes the predicted intention of the third object; based on the predictive analysis result and multiple sample annotations corresponding to the training samples, a target loss is obtained, wherein the multiple sample annotations corresponding to the training samples are obtained based on multiple large models, the parameter amounts of the multiple large models are greater than the parameter amounts of the first preset model, and the parameter amounts of different large models in the multiple large models are different; the first preset model is trained based on the target loss to obtain the intention analysis model.

12. The device according to claim 11, wherein The model training module is used to calculate the probability distribution between the prediction analysis results and the training samples, and the probability distribution between the multiple sample labels and the training samples; based on the probability distribution between the prediction analysis results and the training samples, and the probability distribution between the multiple sample labels and the training samples, the target loss is obtained.

13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.

15. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-round dialogue method and device in knowledge question-answering system

    CN112199473A

  • Intent recognition model training and user intent recognition

    WO2023246393A1