Digital employee reply generation method, system, device and storage medium

By constructing a training sample dataset and a first response generation model, and combining a pattern prediction unit and a response generation unit, dynamically selecting generative and inferential patterns, the problem of low accuracy in digital employee response generation is solved, achieving higher accuracy.

CN121009181BActive Publication Date: 2026-02-27HUNAN ZHITONG STAR TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511536083.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2026-02-27
Estimated Expiration
2045-10-27

AI Technical Summary

Technical Problem

The existing digital employees have difficulty achieving dynamic adaptation in response generation, resulting in low accuracy in response generation.

Method used

A training sample dataset and a first response generation model are constructed, including a pattern prediction unit and a response generation unit. By determining the total advantage value and performing iterative optimization, dynamic selection of generative and inferential patterns is achieved.

Benefits of technology

The accuracy of response generation has been improved by considering multiple factors, including the prediction domain and generation mode, and dynamic selection and optimization of the two modes have been achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009181B_ABST
    Figure CN121009181B_ABST
Patent Text Reader

Abstract

The application discloses a digital employee reply generation method, system, device and storage medium. The digital employee reply generation method comprises the following steps: inputting a training sample data set into a mode prediction unit to obtain a first predicted field, a first predicted reply generation mode and a confidence value of the first predicted reply generation mode output by the mode prediction unit; inputting the training sample data set into a reply generation unit to obtain first reply generation data and second reply generation data of each training sample data output by the reply generation unit; determining an advantage total value based on the first predicted field, the first predicted reply generation mode, the first reply generation data, the second reply generation data and the confidence value; and iteratively optimizing the first reply generation model based on the advantage total value until the number of iterations reaches a preset maximum number of iterations, thereby obtaining a trained reply generation model and improving the accuracy of reply generation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of digital employee reply generation, and in particular to a digital employee reply generation method, system, device and storage medium. BACKGROUND

[0002] With the deep integration of generative artificial intelligence and natural language processing technology, digital employees have become the core interactive carriers in enterprise service, technical consultation, intelligent customer service and other scenarios, and their core capabilities are embodied in generating accurate replies that meet user interaction requests. Currently, the reply generation of digital employees mainly relies on two modes: reasoning mode and generative mode. The reasoning mode takes rule engine and structured knowledge graph as the core, and derives replies through preset logic chain, and its accuracy depends on the completeness of rules and knowledge. The generative mode generates replies by learning the language rules of massive text data, and shows flexibility in open scenarios, but its accuracy is restricted by data quality and model characteristics.

[0003] Currently, the selection of the above two modes by digital employees generally adopts a static binding or manual rule presetting method, but the above method is difficult to realize a dynamic adaptive decision mechanism according to business scenarios, thereby resulting in low accuracy of reply generation. SUMMARY

[0004] The present application aims to at least solve the technical problems existing in the prior art. To this end, the present application proposes a digital employee reply generation method, system, device and storage medium, which can realize dynamic adaptation of the two modes, thereby improving the accuracy of reply generation.

[0005] In a first aspect of the present application, a digital employee reply generation method is provided, comprising the following steps:

[0006] A training sample data set and a first reply generation model are constructed, wherein the training sample data set includes the real field and the real reply generation mode of each training sample data, the first reply generation model includes a mode prediction unit and a reply generation unit, the real field is at least one of customer service, finance or medical treatment, and the real reply generation mode is at least one of generative mode and reasoning mode;

[0007] The training sample data set is input into the mode prediction unit to obtain a first predicted field, a first predicted reply generation mode and a confidence value of the first predicted reply generation mode output by the mode prediction unit;

[0008] inputting the training sample data set into the reply generation unit to obtain first reply generation data and second reply generation data of each training sample data output by the reply generation unit, wherein the first reply generation data is a plurality of pieces of reply generation data generated by a generative mode, and the second reply generation data is a plurality of pieces of reply generation data generated by an inferential mode;

[0009] determining an advantage total value based on the first predicted field, the first predicted reply generation mode, the first reply generation data, the second reply generation data and the confidence value, wherein the advantage total value is a scalar value used to evaluate the correctness of the first predicted reply generation mode;

[0010] iteratively optimizing the first reply generation model based on the advantage total value until the number of iterations reaches a preset maximum number of iterations, to obtain a trained reply generation model.

[0011] According to the reply generation method of the digital employee, at least the following beneficial effects are achieved:

[0012] The method comprises the following steps: constructing a training sample data set and a first reply generation model, wherein the training sample data set comprises a real field and a real reply generation mode of each training sample data, the first reply generation model comprises a mode prediction unit and a reply generation unit, the real field is at least one of customer service, finance or medical treatment, and the real reply generation mode is at least one of a generative mode and an inferential mode; inputting the training sample data set into the mode prediction unit to obtain a first predicted field, a first predicted reply generation mode and a confidence value of the first predicted reply generation mode output by the mode prediction unit; inputting the training sample data set into the reply generation unit to obtain first reply generation data and second reply generation data of each training sample data output by the reply generation unit, wherein the first reply generation data is a plurality of pieces of reply generation data generated by the generative mode, and the second reply generation data is a plurality of pieces of reply generation data generated by the inferential mode; determining an advantage total value based on the first predicted field, the first predicted reply generation mode, the first reply generation data, the second reply generation data and the confidence value, wherein the advantage total value is a scalar value used to evaluate the correctness of the first predicted reply generation mode; iteratively optimizing the first reply generation model based on the advantage total value until the number of iterations reaches a preset maximum number of iterations, to obtain a trained reply generation model, and the first predicted reply generation mode is obtained by the mode prediction unit, and the first reply generation data and the second reply generation data of each training sample data are output by the reply generation unit, the first predicted field, the first predicted reply generation mode and the first reply generation data and the second reply generation data are trained separately, so that the advantage total value can be calculated by combining the first predicted field, the first predicted reply generation mode and the reply generation data in the future, the prediction field, the prediction generation mode and the generated data are considered multiple times, and then the first reply generation model is iteratively optimized based on the advantage total value, so that the dynamic selection of the two modes is realized, and the accuracy of reply generation is improved.

[0013] According to some embodiments of the present application, the advantage total value is determined based on the first predicted field, the first predicted reply generation mode, the first reply generation data, the second reply generation data and the confidence value, and comprises:

[0014] obtaining a first data length of each piece of reply generation data in the first reply generation data, a second data length of each piece of reply generation data in the second reply generation data and a first preset reward threshold value;

[0015] determining a quality reward value based on the first predicted field, the first reply generation data, the second reply generation data and a preset field rule, wherein the quality reward value is a constant value used to represent the accuracy of the reply generation data;

[0016] determine an efficiency reward value based on the first data length, the second data length and a preset length threshold, wherein the efficiency reward value is a scalar value used to represent the expression conciseness of the reply generation data;

[0017] determine a first total reward value of each reply generation data in each of the training sample data based on the confidence value, the quality reward value and the efficiency reward value;

[0018] calculate a mean value of the first total reward values of all the reply generation data in each of the first reply generation data as a first reward mean value of each of the first reply generation data, and calculate a mean value of the first total reward values of all the reply generation data in each of the second reply generation data as a second reward mean value of each of the second reply generation data;

[0019] calculate an absolute value of a difference between the first reward mean value and the second reward mean value;

[0020] determine a second total reward value of each reply generation data in the first reply generation data and the second reply generation data based on the absolute value of the difference, the first total reward value, the first reward mean value and the second reward mean value;

[0021] calculate a mean value of the second total reward values of all the reply generation data in each of the first reply generation data as a third reward mean value of each of the first reply generation data, and calculate a mean value of the second total reward values of all the reply generation data in each of the second reply generation data as a fourth reward mean value of each of the second reply generation data;

[0022] determine the advantage total value based on the first predicted reply generation mode, the second total reward value, the third reward mean value and the fourth reward mean value.

[0023] According to some embodiments of the present application, the determination of the second total reward value of each reply generation data in the first reply generation data and the second reply generation data based on the absolute value of the difference, the first total reward value, the first reward mean value and the second reward mean value comprises:

[0024] screen out a reply generation data with the largest first total reward value in each of the first reply generation data as a third reply generation data, and add a preset same-mode advantage reward value to the first total reward value of each reply generation data in the third reply generation data as a third total reward value of each reply generation data in the third reply generation data;

[0025] screening out the reply generation data with the maximum first total reward value in each of the second reply generation data as fourth reply generation data; and adding the first total reward value of each of the fourth reply generation data to a preset same-mode advantage reward value as a third total reward value of each of the fourth reply generation data;

[0026] screening out the reply generation data with the first total reward value not being the maximum in each of the first reply generation data and the reply generation data with the first total reward value not being the maximum in each of the second reply generation data as fifth reply generation data; and taking the first total reward value of the fifth reply generation data as a third total reward value of each of the fifth reply generation data;

[0027] screening out all of the first reply generation data with the absolute value of the difference being greater than the first preset reward threshold and the first reward average being greater than the second reward average as sixth reply generation data; and adding the third total reward value of each of the sixth reply generation data to a preset cross-mode advantage reward value to obtain a second total reward value of each of the sixth reply generation data;

[0028] screening out all of the second reply generation data with the absolute value of the difference being greater than the first preset reward threshold and the second reward average being greater than the first reward average as seventh reply generation data; and adding the third total reward value of each of the seventh reply generation data to the preset cross-mode advantage reward value to obtain a second total reward value of each of the seventh reply generation data;

[0029] screening out all of the first reply generation data and all of the second reply generation data with the absolute value of the difference being less than or equal to the first preset reward threshold as eighth reply generation data; and taking the third total reward value of each of the eighth reply generation data as a second total reward value of each of the eighth reply generation data.

[0030] According to some embodiments of the present application, the determining of the advantage total value based on the first predicted reply generation mode, the second total reward value, the third reward average and the fourth reward average comprises:

[0031] a difference between the second total reward value of each of the first reply generation data and the corresponding third reward mean value is calculated as a same-mode advantage value of each of the first reply generation data;

[0032] a difference between the third reward mean value of each of the training sample data and the corresponding fourth reward mean value is obtained as a cross-mode advantage value of the first reply generation data of each of the training sample data; a difference between the fourth reward mean value of each of the training sample data and the corresponding third reward mean value is obtained as a cross-mode advantage value of the second reply generation data of each of the training sample data;

[0033] a sum of the same-mode advantage value and the cross-mode advantage value of each of the reply generation data is calculated as a first sum of each of the reply generation data;

[0034] training sample data whose first predicted reply generation mode is the generative mode is filtered out from the training sample data set as a first training sample data set; training sample data whose first predicted reply generation mode is the inferential mode is filtered out from the training sample data set as a second training sample data set;

[0035] a second sum of each of the first reply generation data of the first training sample data set is obtained by adding a preset additional reward value to the first sum of each of the first reply generation data of the first training sample data set; a first sum of each of the second reply generation data of the first training sample data set is taken as the corresponding second sum;

[0036] a second sum of each of the second reply generation data of the second training sample data set is obtained by adding a preset additional reward value to the first sum of each of the second reply generation data of the second training sample data set; a first sum of each of the first reply generation data of the second training sample data set is taken as the corresponding second sum;

[0037] a sum of the second sums of all the reply generation data is calculated as the advantage total value.

[0038] According to some embodiments of the present application, the method further comprises:

[0039] In the case that the user input data is received and the user input data includes voice data, image data and text data, the voice data is converted into first text data by an automatic speech recognition method, and the image data is converted into second text data by an optical character recognition method; the second text data is converted by an automatic speech recognition method;

[0040] The text data of the user input data, the first text data and the second text data are encoded by a UTF-8 encoding method to obtain encoded data;

[0041] A feature vector of the encoded data is extracted as a to-be-predicted feature vector;

[0042] The to-be-predicted feature vector is input into the trained reply generation model to obtain reply generation data output by the trained reply generation model;

[0043] The reply generation data is sent to the user by the digital employee.

[0044] According to some embodiments of the present application, the training sample data set and the first reply generation model are constructed, comprising:

[0045] An initial training sample data set and an initial reply generation model are constructed;

[0046] Each initial training sample data in the initial training sample data set is encoded by a UTF-8 encoding method to obtain the training sample data set;

[0047] The training sample data set is input into the initial reply generation model to obtain a second predicted field and a second predicted reply generation mode output by the initial reply generation model;

[0048] A cross-entropy loss value of the second predicted field and the real field is calculated by a cross-entropy loss function;

[0049] A mean square error loss value of the second predicted reply generation mode and the real reply generation mode is calculated by a mean square error loss function;

[0050] A triple pattern comparison loss value of the second predicted reply generation mode and the real reply generation mode is calculated by a triple pattern comparison loss function;

[0051] A total loss value is determined based on the cross-entropy loss value, the mean square error loss value and the triple pattern comparison loss value;

[0052] The initial reply generation model is iteratively optimized based on the total loss value until the number of iterations reaches the preset maximum number of iterations to obtain the first reply generation model.

[0053] According to some embodiments of the present application, the total loss value is determined based on the cross-entropy loss value, the mean square error loss value and the triplet mode comparison loss value, comprising:

[0054] The cross-entropy loss value is multiplied by a first preset weight value to obtain a first product value;

[0055] The mean square error loss value is multiplied by a second preset weight value to obtain a second product value;

[0056] The triplet mode comparison loss value is multiplied by a third preset weight value to obtain a third product value;

[0057] The first product value, the second product value and the third product value are added to obtain the total loss value.

[0058] According to a second aspect of the present application, a digital employee reply generation system is provided, comprising:

[0059] A data and model construction module is configured to construct a training sample data set and a first reply generation model, wherein the training sample data set comprises a real field and a real reply generation mode of each training sample data, and the first reply generation model comprises a mode prediction unit and a reply generation unit, the real field is at least one of customer service, finance or medical treatment, and the real reply generation mode is at least one of a generative mode and an inferential mode;

[0060] A mode prediction module is configured to input the training sample data set into the mode prediction unit to obtain a first predicted field, a first predicted reply generation mode and a confidence value of the first predicted reply generation mode output by the mode prediction unit;

[0061] A reply generation data output module is configured to input the training sample data set into the reply generation unit to obtain first reply generation data and second reply generation data of each training sample data output by the reply generation unit, wherein the first reply generation data is a plurality of reply generation data generated by a generative mode, and the second reply generation data is a plurality of reply generation data generated by an inferential mode;

[0062] An advantage total value determination module is configured to determine an advantage total value based on the first predicted field, the first predicted reply generation mode, the first reply generation data, the second reply generation data and the confidence value, wherein the advantage total value is a scalar value for evaluating the correctness of the first predicted reply generation mode;

[0063] An iteration module is configured to iteratively optimize the first reply generation model based on the advantage total value until the number of iterations reaches a preset maximum number of iterations, to obtain a trained reply generation model.

[0064] The system constructs a training sample data set and a first reply generation model. The training sample data set includes a real field and a real reply generation mode of each training sample data. The first reply generation model includes a mode prediction unit and a reply generation unit. The real field is at least one of customer service, finance, or medical treatment. The real reply generation mode is at least one of a generative mode and an inferential mode. The training sample data set is input into the mode prediction unit to obtain a first predicted field, a first predicted reply generation mode, and a confidence value of the first predicted reply generation mode output by the mode prediction unit. The training sample data set is input into the reply generation unit to obtain first reply generation data and second reply generation data of each training sample data output by the reply generation unit. The first reply generation data is a plurality of pieces of reply generation data generated by the generative mode. The second reply generation data is a plurality of pieces of reply generation data generated by the inferential mode. An advantage total value is determined based on the first predicted field, the first predicted reply generation mode, the first reply generation data, the second reply generation data, and the confidence value. The advantage total value is a scalar value used to evaluate the correctness of the first predicted reply generation mode. The first reply generation model is iteratively optimized based on the advantage total value until the number of iterations reaches a preset maximum number of iterations, to obtain a trained reply generation model. The first predicted reply generation mode is obtained by the mode prediction unit. The first reply generation data and the second reply generation data of each training sample data are output by the reply generation unit. The first predicted field, the first predicted reply generation mode, and the first reply generation data and the second reply generation data are trained separately, so that the advantage total value can be calculated based on the first predicted field, the first predicted reply generation mode, and the reply generation data in the future. The prediction field, the predicted generation mode, and the generated data are considered multiple times. The first reply generation model is iteratively optimized based on the advantage total value, so that the dynamic selection of the two modes is realized, and the accuracy of reply generation is improved.

[0065] In a third aspect, the present application provides a digital employee reply generation electronic device, including at least one control processor and a memory connected in communication with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to perform the digital employee reply generation method described above.

[0066] In a fourth aspect of the present application, a computer readable storage medium is provided, and the computer readable storage medium stores computer executable instructions for causing a computer to execute the digital employee reply generation method.

[0067] It should be noted that the beneficial effects between the second aspect to the fourth aspect of the present application and the prior art are the same as the beneficial effects between the digital employee reply generation system and the prior art, which will not be described here.

[0068] Additional aspects and advantages of the present application will be in part apparent and in part pointed out below in the description of the application. BRIEF DESCRIPTION OF DRAWINGS

[0069] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the description of the embodiments, given below with reference to the following drawings, wherein:

[0070] Figure 1 is a flow diagram of an embodiment of the digital employee reply generation method provided by the present application;

[0071] Figure 2 is a structural diagram of an embodiment of the digital employee reply generation system provided by the present application;

[0072] Figure 3 is a structural diagram of an embodiment of the electronic device provided by the present application. DETAILED DESCRIPTION

[0073] Embodiments of the present application are described in detail below with reference to the accompanying drawings, in which examples of the embodiments are shown, wherein the same or similar notations are used to represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as limiting the present application.

[0074] In the description of the present application, if there is a description of first, second, etc., it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features or the sequence of indicated technical features.

[0075] In the description of the present application, it should be understood that the orientation description, such as the orientation or position relationship indicated by up, down, etc., is based on the orientation or position relationship shown in the drawings, and is only for the purpose of facilitating the description of the present application and simplifying the description, and therefore cannot be understood as indicating or implying that the device or element indicated must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application.

[0076] In the description of the present application, it should be noted that, unless otherwise explicitly defined, the words such as setting, installing, connecting, etc. should be understood broadly, and the person skilled in the art can determine the specific meaning of the above words in the present application in combination with the specific content of the technical solution.

[0077] With the deep integration of generative artificial intelligence and natural language processing technology, digital employees have become the core interactive carriers in enterprise service, technical consultation, intelligent customer service and other scenarios, and their core capabilities are embodied in generating accurate replies that meet user interaction requests. Currently, the reply generation of digital employees mainly relies on two modes: reasoning mode and generative mode. The reasoning mode takes rule engine and structured knowledge graph as the core, and derives the reply through preset logic chain, and its accuracy depends on the completeness of rules and knowledge. The generative mode generates replies by learning the language rules of massive text data, and shows flexibility in open scenarios, but its accuracy is restricted by data quality and model characteristics.

[0078] Currently, the selection of the above two modes by digital employees generally adopts a static binding or manual rule presetting method, but the above method is difficult to realize a dynamic adaptive decision mechanism according to the business scenario, thereby resulting in low accuracy of reply generation.

[0079] In order to solve the above technical defects, the embodiments of the present application provide a digital employee reply generation method, system, device and storage medium.

[0080] Please refer to Figure 1 It is a flowchart of a digital employee reply generation method provided by the embodiments of the present application, which is applied to an electronic device, which can be a server, etc. As Figure 1 The digital employee reply generation method comprises the following steps:

[0081] Step S101, constructing a training sample data set and a first reply generation model, wherein the training sample data set comprises a real field and a real reply generation mode of each training sample data, the first reply generation model comprises a mode prediction unit and a reply generation unit, the real field is at least one of customer service, finance or medical treatment, and the real reply generation mode is at least one of generative mode and reasoning mode;

[0082] Step S102, inputting the training sample data set into the mode prediction unit to obtain the first predicted field, the first predicted reply generation mode, and the confidence value of the first predicted reply generation mode output by the mode prediction unit;

[0083] Step S103, input the training sample data set into the reply generation unit to obtain the first reply generation data and the second reply generation data of each training sample data output by the reply generation unit, wherein the first reply generation data is a plurality of reply generation data generated by the generative mode, and the second reply generation data is a plurality of reply generation data generated by the inferential mode;

[0084] Step S104, determine an advantage total value based on the first predicted field, the first predicted reply generation mode, the first reply generation data, the second reply generation data and the confidence value, wherein the advantage total value is a scalar value for evaluating the correctness of the first predicted reply generation mode;

[0085] Step S105, iteratively optimize the first reply generation model based on the advantage total value until the number of iterations reaches a preset maximum number of iterations to obtain a trained reply generation model.

[0086] The first reply generation model can be a deep learning neural network model.

[0087] The training sample data set includes a plurality of training sample data, each training sample data corresponds to a first reply generation data and a second reply generation data, each first reply generation data corresponds to a plurality of reply generation data generated by the generative mode, and each second reply generation data corresponds to a plurality of reply generation data generated by the inferential mode.

[0088] The preset maximum number of iterations can be a value set in advance according to actual needs.

[0089] In step S105, the first reply generation model is iteratively optimized based on the advantage total value until the number of iterations reaches a preset maximum number of iterations to obtain a trained reply generation model, which can be iteratively optimized based on the advantage total value to make the model generate reply generation data with higher advantage total value, until the number of iterations reaches a preset maximum number of iterations, and a trained reply generation model is obtained.

[0090] The method comprises the following steps: constructing a training sample data set and a first reply generation model, wherein the training sample data set comprises a real field and a real reply generation mode of each training sample data, the first reply generation model comprises a mode prediction unit and a reply generation unit, the real field is at least one of customer service, finance or medical treatment, and the real reply generation mode is at least one of a generative mode and an inferential mode; inputting the training sample data set into the mode prediction unit to obtain a first predicted field, a first predicted reply generation mode and a confidence value of the first predicted reply generation mode output by the mode prediction unit; inputting the training sample data set into the reply generation unit to obtain first reply generation data and second reply generation data of each training sample data output by the reply generation unit, wherein the first reply generation data is a plurality of pieces of reply generation data generated by the generative mode, and the second reply generation data is a plurality of pieces of reply generation data generated by the inferential mode; determining an advantage total value based on the first predicted field, the first predicted reply generation mode, the first reply generation data, the second reply generation data and the confidence value, wherein the advantage total value is a scalar value used to evaluate the correctness of the first predicted reply generation mode; iteratively optimizing the first reply generation model based on the advantage total value until the number of iterations reaches a preset maximum number of iterations, to obtain a trained reply generation model, wherein the first predicted reply generation mode is obtained by the mode prediction unit, and the first reply generation data and the second reply generation data of each training sample data are output by the reply generation unit, the first predicted field, the first predicted reply generation mode and the first reply generation data and the second reply generation data are trained separately, so that the advantage total value can be calculated based on the first predicted field, the first predicted reply generation mode and the reply generation data in the future, the prediction field, the prediction generation mode and the generation data are considered multiple times, the first reply generation model is iteratively optimized based on the advantage total value, so that the dynamic selection of the two modes is realized, and the accuracy of reply generation is improved.

[0091] In some embodiments, step S104 can include, but is not limited to, steps S201 to S209:

[0092] Step S201, acquiring a first data length of each piece of reply generation data in the first reply generation data, a second data length of each piece of reply generation data in the second reply generation data and a first preset reward threshold;

[0093] Step S202, determining a quality reward value based on the first predicted field, the first reply generation data, the second reply generation data and a preset field rule, wherein the quality reward value is a constant value used to represent the accuracy of the reply generation data;

[0094] Step S203: Based on the first data length, the second data length and the preset length threshold, determine the efficiency reward value, wherein the efficiency reward value is a scalar value used to characterize the conciseness of the expression of the response generated data;

[0095] Step S204: Determine the first total reward value for each response generated for each training sample based on the confidence value, quality reward value, and efficiency reward value;

[0096] Step S205: Calculate the average of the first total reward values ​​of all reply generation data in each first reply generation data set, and use it as the first reward average of each first reply generation data set; calculate the average of the first total reward values ​​of all reply generation data in each second reply generation data set, and use it as the second reward average of each second reply generation data set;

[0097] Step S206: Calculate the absolute value of the difference between the average of the first reward and the average of the second reward;

[0098] Step S207: Based on the absolute value of the difference, the first total reward value, the first average reward value, and the second average reward value, determine the second total reward value for each reply generated in the first reply generated data and the second reply generated data;

[0099] Step S208: Calculate the average of the second total reward values ​​of all reply generation data in each first reply generation data, and use it as the third reward average of each first reply generation data; calculate the average of the second total reward values ​​of all reply generation data in each second reply generation data, and use it as the fourth reward average of each second reply generation data;

[0100] Step S209: Determine the total advantage value based on the first predicted response generation mode, the second total reward value, the third average reward value, and the fourth average reward value.

[0101] The aforementioned forecast areas are at least one of customer service, finance, or healthcare.

[0102] The aforementioned preset domain rules can be a pre-built domain rule library. The pre-built domain rule library may include, but is not limited to, verifying whether the current ratio calculation in the financial field complies with the Basel Accords or checking whether the disease judgment in the medical field matches the ICD-10 coding system.

[0103] The aforementioned preset length threshold can be a value set in advance according to actual needs.

[0104] In the step S201, the first data length of each reply generation data in the first reply generation data and the second data length of each reply generation data in the second reply generation data can be that, in the case that each reply generation data in the first reply generation data and each reply generation data in the second reply generation data are in text format, the string length of each reply generation data in the first reply generation data is taken as the first data length, and the string length of each reply generation data in the second reply generation data is taken as the second data length.

[0105] In the step S201, the first preset reward threshold can be that, in the case that the iteration number of the first reply generation model is the first time, the first preset reward threshold is a value set in advance according to actual needs, and in the case that the iteration number of the first reply generation model is the kth time, the mean and the standard deviation of all first total reward values of the k-1th iteration are calculated, and the first preset reward threshold of the kth iteration is calculated through the following formula: wherein k is a constant value greater than 1, is the first preset reward threshold of the kth iteration, is the mean of all first total reward values of the k-1th iteration, is the standard deviation of all first total reward values of the k-1th iteration.

[0106] In the step S202, the quality reward value is determined based on the first predicted field, the first reply generation data, the second reply generation data and the preset field rule, which can be that the first reply generation data, the second reply generation data and the first predicted field are input into the pre-constructed field rule library to match, and the quality reward value of each reply generation data output by the field rule library is obtained.

[0107] In the step S203, the efficiency reward value is determined based on the first data length, the second data length and the preset length threshold, which can be that, in the case that the first data length or the second data length is within the preset length threshold range, the efficiency reward value is equal to the first preset efficiency reward value, and in the case that the first data length or the second data length is out of the preset length threshold range, the efficiency reward value is equal to the second preset efficiency reward value, wherein the first preset efficiency reward value is a value set in advance according to actual needs, the second preset efficiency reward value is a value set in advance according to actual needs, and the first preset efficiency reward value is greater than the second preset efficiency reward value.

[0108] The step S204 can include but is not limited to steps S2041 to S2044:

[0109] In the step S2041, the confidence value is multiplied by the first preset reward weight value to obtain a first reward product value.

[0110] Step S2042, multiplying the quality reward value by a second preset reward weight value to obtain a second reward product value;

[0111] Step S2043, multiplying the efficiency reward value by a third preset reward weight value to obtain a third reward product value;

[0112] Step S2044, adding the first reward product value, the second reward product value and the third reward product value to obtain a first total reward value of each reply generation data of each training sample data.

[0113] The first preset reward weight value can be a value set in advance according to actual needs, and can be 0.6.

[0114] The second preset reward weight value can be a value set in advance according to actual needs, and can be 0.3.

[0115] The third preset reward weight value can be a value set in advance according to actual needs, and can be 0.1.

[0116] The sum of the first preset reward weight value, the second preset reward weight value and the third preset reward weight value is 1.

[0117] The application calculates the first total reward value by combining the confidence value, the field, the quality of the reply generation data and the length of the reply generation data, and then calculates the advantage total value according to the reply generation mode, thereby providing data basis for subsequent model iteration and update, and improving the accuracy of reply generation.

[0118] In some embodiments, step S207 can include but is not limited to steps S301 to S306:

[0119] Step S301, filtering out the reply generation data with the maximum first total reward value in each first reply generation data as third reply generation data; and adding the first total reward value of each reply generation data in the third reply generation data to a preset same mode advantage reward value as a third total reward value of each reply generation data in the third reply generation data;

[0120] Step S302, filtering out the reply generation data with the maximum first total reward value in each second reply generation data as fourth reply generation data; and adding the first total reward value of each reply generation data in the fourth reply generation data to a preset same mode advantage reward value as a third total reward value of each reply generation data in the fourth reply generation data;

[0121] Step S303, screening out, as the seventh reply generation data, all the second reply generation data whose absolute value of the difference is greater than the first preset reward threshold and whose second reward average is greater than the first reward average; and adding the third total reward value of each piece of reply generation data in the seventh reply generation data to the preset cross-mode advantage reward value to obtain the second total reward value of each piece of reply generation data in the seventh reply generation data.

[0122] Step S304, screening out, as the sixth reply generation data, all the first reply generation data whose absolute value of the difference is greater than the first preset reward threshold and whose first reward average is greater than the second reward average; and adding the third total reward value of each piece of reply generation data in the sixth reply generation data to the preset cross-mode advantage reward value to obtain the second total reward value of each piece of reply generation data in the sixth reply generation data.

[0123] Step S305, screening out, as the seventh reply generation data, all the second reply generation data whose absolute value of the difference is greater than the first preset reward threshold and whose second reward average is greater than the first reward average; and adding the third total reward value of each piece of reply generation data in the seventh reply generation data to the preset cross-mode advantage reward value to obtain the second total reward value of each piece of reply generation data in the seventh reply generation data.

[0124] Step S306, screening out, as the eighth reply generation data, all the first reply generation data and all the second reply generation data whose absolute value of the difference is less than or equal to the first preset reward threshold; and taking the third total reward value of each piece of reply generation data in the eighth reply generation data as the second total reward value of each piece of reply generation data in the eighth reply generation data.

[0125] The preset same-mode advantage reward value can be a value set in advance according to actual needs.

[0126] The preset cross-mode advantage reward value can be a value set in advance according to actual needs.

[0127] The preset same-mode advantage reward value and the preset cross-mode advantage reward value are used to calculate the second total reward value, so that the reply generation data between the same mode and the cross mode is selected, the second total reward value is used as a data basis for subsequent calculation of the advantage total value, and the accuracy of the advantage total value data is improved.

[0128] In some embodiments, step S209 can include, but is not limited to, steps S401 to S407:

[0129] Step S401, calculate the difference between the second total reward value of each reply generation data in the first reply generation data and the corresponding third reward mean value as the same mode advantage value of each reply generation data in the first reply generation data; calculate the difference between the second total reward value of each reply generation data in the second reply generation data and the corresponding fourth reward mean value as the same mode advantage value of each reply generation data in the second reply generation data.

[0130] Step S402, subtract the corresponding fourth reward mean value from the third reward mean value of each training sample data to obtain the cross-mode advantage value of the first reply generation data of each training sample data; subtract the corresponding third reward mean value from the fourth reward mean value of each training sample data to obtain the cross-mode advantage value of the second reply generation data of each training sample data.

[0131] Step S403, calculate the sum of the same mode advantage value and the cross-mode advantage value of each reply generation data as the first sum of each reply generation data.

[0132] Step S404, filter the training sample data whose first predicted reply generation mode is the generative mode from the training sample data set as the first training sample data set; filter the training sample data whose first predicted reply generation mode is the inferential mode from the training sample data set as the second training sample data set.

[0133] Step S405, add the first sum of each reply generation data of the first reply generation data of the first training sample data set to the preset additional reward value to obtain the second sum of each reply generation data of the first reply generation data of the first training sample data set; the first sum of each reply generation data of the second reply generation data of the first training sample data set is taken as the corresponding second sum.

[0134] Step S406, add the first sum of each reply generation data of the second reply generation data of the second training sample data set to the preset additional reward value to obtain the second sum of each reply generation data of the second reply generation data of the second training sample data set; the first sum of each reply generation data of the first reply generation data of the second training sample data set is taken as the corresponding second sum.

[0135] Step S407, calculate the sum of the second sums of all reply generation data as the total advantage value.

[0136] The above predicted reply generation mode can be at least one of the generative mode and the inferential mode.

[0137] The cross-modal advantage values of all the reply generation data in the same first reply generation data are the same; and the cross-modal advantage values of all the reply generation data in the same second reply generation data are the same.

[0138] The preset additional reward value can be a value set in advance according to actual needs.

[0139] The application considers the advantage value of the first predicted reply generation mode by calculating the sum of the same-mode advantage value and the cross-mode advantage value, improves the accuracy of the total advantage value data, and provides more accurate data basis for subsequent model training.

[0140] In some embodiments, the method can further include, but is not limited to, steps S501 to S505:

[0141] Step S501, in the case that the user input data is received and the user input data includes voice data, image data and text data, the voice data is converted into first text data by an automatic speech recognition method, the image data is converted into second text data by an optical character recognition method, and the voice data is converted into second text data by an automatic speech recognition method;

[0142] Step S502, encoding the text data, the first text data and the second text data of the user input data by a UTF-8 encoding method to obtain encoded data;

[0143] Step S503, extracting a feature vector of the encoded data as a to-be-predicted feature vector;

[0144] Step S504, inputting the to-be-predicted feature vector into the trained reply generation model to obtain reply generation data output by the trained reply generation model;

[0145] Step S505, sending the reply generation data to the user by a digital employee.

[0146] In the above step S503, the feature vector of the encoded data can be extracted by an attention mechanism.

[0147] In the above step S504, the reply generation data output by the trained reply generation model can be the reply generation data with the maximum total advantage value.

[0148] The application converts the voice data, image data and text data into text data, and then performs encoding and feature extraction, thereby realizing feature extraction and reply generation of multi-modal data and improving the accuracy of reply generation.

[0149] In some embodiments, step S101 can include, but is not limited to, steps S601 to S608:

[0150] Step S601, constructing an initial training sample data set and an initial reply generation model;

[0151] Step S602, encoding each initial training sample data in the initial training sample data set by a UTF-8 encoding method to obtain a training sample data set;

[0152] Step S603, inputting the training sample data set into the initial reply generation model to obtain a second predicted field and a second predicted reply generation mode output by the initial reply generation model;

[0153] Step S604, calculating a cross-entropy loss value of the second predicted field and the real field by a cross-entropy loss function;

[0154] Step S605, calculating a mean square error loss value of the second predicted reply generation mode and the real reply generation mode by a mean square error loss function;

[0155] Step S606, calculating a triple pattern comparison loss value of the second predicted reply generation mode and the real reply generation mode by a triple pattern comparison loss function;

[0156] Step S607, determining a total loss value based on the cross-entropy loss value, the mean square error loss value and the triple pattern comparison loss value;

[0157] Step S608, iteratively optimizing the initial reply generation model based on the total loss value until the number of iterations reaches a preset maximum number of iterations to obtain a first reply generation model.

[0158] In the above step S608, the initial reply generation model is iteratively optimized based on the total loss value until the number of iterations reaches the preset maximum number of iterations to obtain the first reply generation model. The initial reply generation model can be iteratively optimized based on the total loss value to make the model generate reply generation data with a smaller total loss value, and the first reply generation model is obtained when the number of iterations reaches the preset maximum number of iterations.

[0159] The present application determines the total loss value based on the cross-entropy loss value, the mean square error loss value and the triple pattern comparison loss value, and iteratively optimizes the initial reply generation model based on the total loss value, thereby improving the accuracy and efficiency of model training.

[0160] In some embodiments, step S607 can include but is not limited to steps S701 to S704:

[0161] Step S701, multiplying the cross-entropy loss value by a first preset weight value to obtain a first product value;

[0162] Step S702, multiply the mean square error loss value with the second preset weight value to obtain a second product value;

[0163] Step S703, multiply the triplet pattern comparison loss value with a third preset weight value to obtain a third product value;

[0164] Step S704, add the first product value, the second product value and the third product value to obtain a total loss value.

[0165] The first preset weight value can be a value set in advance according to actual needs.

[0166] The second preset weight value can be a value set in advance according to actual needs.

[0167] The third preset weight value can be a value set in advance according to actual needs.

[0168] The sum of the first preset weight value, the second preset weight value and the third preset weight value is 1.

[0169] The application adjusts the contribution of different loss values by setting different weight values and diversified loss values, and improves the generalization ability of the model.

[0170] In addition, with reference to Figure 2 An embodiment of the application provides a digital employee reply generation system, which comprises a data and model construction module 1100, a pattern prediction module 1200, a reply generation data output module 1300, an advantage total value determination module 1400 and an iteration module 1500, wherein:

[0171] The data and model construction module 1100 is used for constructing a training sample data set and a first reply generation model, wherein the training sample data set comprises a real field and a real reply generation pattern of each training sample data, and the first reply generation model comprises a pattern prediction unit and a reply generation unit, the real field is at least one of customer service, finance or medical treatment, and the real reply generation pattern is at least one of a generative pattern and an inferential pattern;

[0172] The pattern prediction module 1200 is used for inputting the training sample data set into the pattern prediction unit to obtain a first predicted field, a first predicted reply generation pattern and a confidence value of the first predicted reply generation pattern output by the pattern prediction unit;

[0173] The reply generation data output module 1300 is used for inputting the training sample data set into the reply generation unit to obtain first reply generation data and second reply generation data of each training sample data output by the reply generation unit, wherein the first reply generation data is a plurality of pieces of reply generation data generated through the generative pattern, and the second reply generation data is a plurality of pieces of reply generation data generated through the inferential pattern.

[0174] The advantage total value determination module 1400 is configured to determine an advantage total value based on the first predicted field, the first predicted reply generation mode, the first reply generation data, the second reply generation data, and the confidence value, where the advantage total value is a scalar value used to evaluate the correctness of the first predicted reply generation mode.

[0175] The iteration module 1500 is configured to iteratively optimize the first reply generation model based on the advantage total value until the number of iterations reaches a preset maximum number of iterations, to obtain a trained reply generation model.

[0176] The system is configured to construct a training sample data set and a first reply generation model, where the training sample data set includes a real field and a real reply generation mode of each training sample data, the first reply generation model includes a mode prediction unit and a reply generation unit, the real field is at least one of customer service, finance, or medical treatment, and the real reply generation mode is at least one of a generative mode and an inferential mode; the training sample data set is input into the mode prediction unit to obtain a first predicted field, a first predicted reply generation mode, and a confidence value of the first predicted reply generation mode output by the mode prediction unit; the training sample data set is input into the reply generation unit to obtain first reply generation data and second reply generation data of each training sample data output by the reply generation unit, where the first reply generation data is a plurality of pieces of reply generation data generated by the generative mode, and the second reply generation data is a plurality of pieces of reply generation data generated by the inferential mode; an advantage total value is determined based on the first predicted field, the first predicted reply generation mode, the first reply generation data, the second reply generation data, and the confidence value, where the advantage total value is a scalar value used to evaluate the correctness of the first predicted reply generation mode; the first reply generation model is iteratively optimized based on the advantage total value until the number of iterations reaches a preset maximum number of iterations, to obtain a trained reply generation model. The system is configured to obtain the first predicted reply generation mode by the mode prediction unit, and obtain the first reply generation data and the second reply generation data of each training sample data output by the reply generation unit, train the first predicted field, the first predicted reply generation mode, and the first reply generation data and the second reply generation data separately, so that the advantage total value can be calculated by combining the first predicted field, the first predicted reply generation mode, and the reply generation data in the future, the prediction field, the predicted generation mode, and the generated data are considered multiple times, the first reply generation model is iteratively optimized based on the advantage total value, so that the dynamic selection of the two modes is realized, and the accuracy of reply generation is improved.

[0177] It should be noted that the system embodiment and the method embodiment described above are based on the same inventive concept, and therefore the related content of the method embodiment described above is also applicable to the system embodiment, which will not be described here again.

[0178] Figure 3 A hardware structure diagram of the digital employee reply generation provided by the embodiments of the present application is shown.

[0179] The digital employee reply generation device can include a processor 301 and a memory 302 storing computer program instructions.

[0180] Specifically, the processor 301 described above can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement one or more embodiments of the present application.

[0181] The memory 302 can include a mass storage for data or instructions. By way of example and not limitation, the memory 302 can include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. Where appropriate, the memory 302 can include removable or non-removable (or fixed) media. Where appropriate, the memory 302 can be internal or external to the integrated gateway disaster recovery device. In some embodiments, the memory 302 is a non-volatile solid-state memory.

[0182] In some embodiments, the memory 302 can include read-only memory (ROM), random access memory (RAM), a disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software that, when executed (e.g., by one or more processors), is operable to perform operations described with reference to the methods according to an aspect of the present disclosure.

[0183] The processor 301 implements any of the digital employee reply generation methods described above by reading and executing computer program instructions stored in the memory 302.

[0184] In one example, the digital employee reply generation device can further include a communication interface 303 and a bus 310. As shown, the processor 301, the memory 302, and the communication interface 303 are connected through the bus 310 and complete communication with each other. Figure 3

[0185] ​The communication interface 303 is mainly configured to realize the communication between the modules, devices, units and / or equipment in the embodiments of the present application.

[0186] Bus 310 includes hardware, software, or both, that couples components of the digital employee response generation device to each other. As an example and not by way of limitation, the bus can include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand (IB) interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus or a combination of two or more of these. Where appropriate, bus 310 can include one or more buses. Although this application describes and shows a particular bus, this application contemplates any suitable bus or interconnect.

[0187] The digital employee response generation device can perform the digital employee response generation method in the embodiments of the present application based on the three-dimensional design model, thereby realizing the digital employee response generation method and system described in combination Figure 1 and Figure 2 with the above embodiments.

[0188] In addition, in combination with the digital employee response generation method in the above embodiments, the embodiments of the present application can provide a computer storage medium for implementation. The computer storage medium has computer program instructions stored thereon; the computer program instructions are executed by a processor to implement any of the digital employee response generation methods in the above embodiments.

[0189] It needs to be clear that the present application is not limited to the specific configurations and processes described above and shown in the drawings. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between steps, after understanding the spirit of the present application.

[0190] The functional blocks shown in the structural block diagrams above can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, and the like. When implemented in software, the elements of the present application are program or code segments that are used to perform the required tasks. The program or code segments can be stored in a machine-readable medium, or transmitted through a data signal carried in a carrier wave over a transmission medium or communication link. A "machine-readable medium" includes any medium that can store or transport information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, and the like. The code segments can be downloaded via computer networks such as the Internet, intranets, and the like.

[0191] It is also important to note that the examples mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the examples, or in an order different from the examples, or several steps can be performed simultaneously.

[0192] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other processing device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other processing device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer program instructions can also be stored in a computer readable medium that can direct a computer, other programmable data processing apparatus, or other processing device to operate in a particular manner, such that the instructions stored in the computer readable medium produce an article of manufacture including instructions which implement the function / act specified in the flowchart and / or block diagram block or blocks.

[0193] The above merely describes a specific implementation of the present application. Those skilled in the art can clearly understand the specific working processes of the system, modules and units described above for the convenience and brevity of description, and can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein again. It should be understood that the protection scope of the present application is not limited to this, and any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application.

Claims

1. A method for digital employee reply generation, characterized by, The digital employee reply generation method comprises: constructing a training sample data set and a first reply generation model, wherein the training sample data set comprises a real field and a real reply generation mode of each training sample data, the first reply generation model comprises a mode prediction unit and a reply generation unit, the real field is at least one of customer service, finance or medical treatment, and the real reply generation mode is at least one of a generative mode and an inferential mode; inputting the training sample data set into the mode prediction unit to obtain a first predicted field, a first predicted reply generation mode and a confidence value of the first predicted reply generation mode output by the mode prediction unit; inputting the training sample data set into the reply generation unit to obtain first reply generation data and second reply generation data of each training sample data output by the reply generation unit, wherein the first reply generation data is a plurality of reply generation data generated by the generative mode, and the second reply generation data is a plurality of reply generation data generated by the inferential mode; determining an advantage total value based on the first predicted field, the first predicted reply generation mode, the first reply generation data, the second reply generation data and the confidence value, wherein the advantage total value is a scalar value for evaluating the correctness of the first predicted reply generation mode; iteratively optimizing the first reply generation model based on the advantage total value until the number of iterations reaches a preset maximum number of iterations to obtain a trained reply generation model.

2. The method of claim 1, wherein, The determination of the advantage total value based on the first predicted field, the first predicted reply generation mode, the first reply generation data, the second reply generation data and the confidence value comprises: obtaining a first data length of each reply generation data in the first reply generation data, a second data length of each reply generation data in the second reply generation data and a first preset reward threshold value; determining a quality reward value based on the first predicted field, the first reply generation data, the second reply generation data and a preset field rule, wherein the quality reward value is a constant value for representing the accuracy of the reply generation data; determining an efficiency reward value based on the first data length, the second data length and a preset length threshold value, wherein the efficiency reward value is a scalar value for representing the conciseness of the expression of the reply generation data; determining a first total reward value of each reply generation data of each training sample data based on the confidence value, the quality reward value and the efficiency reward value; calculating the mean value of the first total reward value of all reply generation data in each first reply generation data as a first reward mean value of each first reply generation data, and calculating the mean value of the first total reward value of all reply generation data in each second reply generation data as a second reward mean value of each second reply generation data; calculating the absolute value of the difference between the first reward mean value and the second reward mean value; determine, based on the absolute value of the difference, the first total reward value, the first reward mean value and the second reward mean value, a second total reward value of each of the first reply generation data and the second reply generation data; calculate a mean value of the second total reward value of all of the first reply generation data as a third reward mean value of each of the first reply generation data, and calculate a mean value of the second total reward value of all of the second reply generation data as a fourth reward mean value of each of the second reply generation data; determine the advantage total value based on the first predicted reply generation mode, the second total reward value, the third reward mean value and the fourth reward mean value.

3. The method of claim 2, wherein, The determining, based on the absolute value of the difference, the first total reward value, the first reward mean value and the second reward mean value, a second total reward value of each of the first reply generation data and the second reply generation data comprises: selecting, as third reply generation data, a reply generation data with the maximum first total reward value in each of the first reply generation data, and adding the first total reward value of each of the reply generation data in the third reply generation data to a preset same-mode advantage reward value to obtain a third total reward value of each of the reply generation data in the third reply generation data; selecting, as fourth reply generation data, a reply generation data with the maximum first total reward value in each of the second reply generation data, and adding the first total reward value of each of the reply generation data in the fourth reply generation data to the preset same-mode advantage reward value to obtain a third total reward value of each of the reply generation data in the fourth reply generation data; selecting, as fifth reply generation data, a reply generation data with a first total reward value that is not the maximum in each of the first reply generation data and a reply generation data with a first total reward value that is not the maximum in each of the second reply generation data, and taking the first total reward value of the fifth reply generation data as a third total reward value of each of the reply generation data in the fifth reply generation data; selecting, as sixth reply generation data, all of the first reply generation data with the absolute value of the difference greater than the first preset reward threshold value and the first reward mean value greater than the second reward mean value, and adding the third total reward value of each of the reply generation data in the sixth reply generation data to a preset cross-mode advantage reward value to obtain a second total reward value of each of the reply generation data in the sixth reply generation data; selecting, as seventh reply generation data, all of the second reply generation data with the absolute value of the difference greater than the first preset reward threshold value and the second reward mean value greater than the first reward mean value, and adding the third total reward value of each of the reply generation data in the seventh reply generation data to the preset cross-mode advantage reward value to obtain a second total reward value of each of the reply generation data in the seventh reply generation data; and Screening all the first reply generation data and all the second reply generation data whose absolute value of the difference is less than or equal to the first preset reward threshold as eighth reply generation data; and taking the third total reward value of each piece of reply generation data in the eighth reply generation data as the second total reward value of each piece of reply generation data in the eighth reply generation data.

4. The method of claim 3, wherein, The determining the advantage total value based on the first predicted reply generation mode, the second total reward value, the third reward mean value and the fourth reward mean value comprises: calculating the difference between the second total reward value of each piece of reply generation data in each of the first reply generation data and the corresponding third reward mean value as the same-mode advantage value of each piece of reply generation data in the first reply generation data; and calculating the difference between the second total reward value of each piece of reply generation data in each of the second reply generation data and the corresponding fourth reward mean value as the same-mode advantage value of each piece of reply generation data in the second reply generation data; subtracting the corresponding fourth reward mean value from the third reward mean value of each of the training sample data to obtain the cross-mode advantage value of the first reply generation data of each of the training sample data; and subtracting the corresponding third reward mean value from the fourth reward mean value of each of the training sample data to obtain the cross-mode advantage value of the second reply generation data of each of the training sample data; calculating the sum of the same-mode advantage value and the cross-mode advantage value of each piece of reply generation data as the first sum of each piece of reply generation data; screening the training sample data whose first predicted reply generation mode is the generative mode from the training sample data set as a first training sample data set; and screening the training sample data whose first predicted reply generation mode is the inferential mode from the training sample data set as a second training sample data set; adding the first sum of each piece of reply generation data of the first reply generation data of the first training sample data set to a preset additional reward value to obtain the second sum of each piece of reply generation data of the first reply generation data of the first training sample data set; and taking the first sum of each piece of reply generation data of the second reply generation data of the first training sample data set as the corresponding second sum; adding the first sum of each piece of reply generation data of the second reply generation data of the second training sample data set to a preset additional reward value to obtain the second sum of each piece of reply generation data of the second reply generation data of the second training sample data set; and taking the first sum of each piece of reply generation data of the first reply generation data of the second training sample data set as the corresponding second sum; calculating the sum of the second sums of all the reply generation data as the advantage total value.

5. The method of claim 1, wherein, The method further comprises: In the case that the user input data is received and the user input data includes voice data, image data and text data, the voice data is converted into first text data by an automatic speech recognition method, and the image data is converted into second text data by an optical character recognition method; the second text data is converted into second text data by an automatic speech recognition method; The text data of the user input data, the first text data and the second text data are encoded by a UTF-8 encoding method to obtain encoded data; A feature vector of the encoded data is extracted as a to-be-predicted feature vector; The to-be-predicted feature vector is input into the trained reply generation model to obtain reply generation data output by the trained reply generation model; The reply generation data is sent to the user by the digital employee.

6. The method of claim 1, wherein, The construction of the training sample data set and the first reply generation model comprises: An initial training sample data set and an initial reply generation model are constructed; Each initial training sample data in the initial training sample data set is encoded by a UTF-8 encoding method to obtain the training sample data set; The training sample data set is input into the initial reply generation model to obtain a second predicted field and a second predicted reply generation mode output by the initial reply generation model; A cross-entropy loss value of the second predicted field and the real field is calculated by a cross-entropy loss function; A mean square error loss value of the second predicted reply generation mode and the real reply generation mode is calculated by a mean square error loss function; A triple pattern comparison loss value of the second predicted reply generation mode and the real reply generation mode is calculated by a triple pattern comparison loss function; A total loss value is determined based on the cross-entropy loss value, the mean square error loss value and the triple pattern comparison loss value; The initial reply generation model is iteratively optimized based on the total loss value until the number of iterations reaches the preset maximum number of iterations to obtain the first reply generation model.

7. The method of claim 6, wherein, The determination of the total loss value based on the cross-entropy loss value, the mean square error loss value and the triple pattern comparison loss value comprises: The cross-entropy loss value is multiplied by a first preset weight value to obtain a first product value; The mean square error loss value is multiplied by a second preset weight value to obtain a second product value; The triple pattern comparison loss value is multiplied by a third preset weight value to obtain a third product value; The first product value, the second product value and the third product value are added to obtain the total loss value.

8. A digital employee reply generation system characterized by, The reply generation system of the digital employee comprises: A data and model construction module is configured to construct a training sample data set and a first reply generation model, wherein the training sample data set comprises a real field and a real reply generation mode of each training sample data, the first reply generation model comprises a mode prediction unit and a reply generation unit, the real field is at least one of customer service, finance or medical treatment, and the real reply generation mode is at least one of a generative mode and an inferential mode. a pattern prediction module, configured to input the training sample data set into the pattern prediction unit to obtain a first predicted domain, a first predicted reply generation pattern and a confidence value of the first predicted reply generation pattern output by the pattern prediction unit; a reply generation data output module, configured to input the training sample data set into the reply generation unit to obtain first reply generation data and second reply generation data of each of the training sample data output by the reply generation unit, wherein the first reply generation data is a plurality of pieces of reply generation data generated by a generative pattern, and the second reply generation data is a plurality of pieces of reply generation data generated by an inferential pattern; an advantage total value determination module, configured to determine an advantage total value based on the first predicted domain, the first predicted reply generation pattern, the first reply generation data, the second reply generation data and the confidence value, wherein the advantage total value is a scalar value used to evaluate correctness of the first predicted reply generation pattern; an iteration module, configured to iteratively optimize the first reply generation model based on the advantage total value until a preset maximum iteration number is reached to obtain a trained reply generation model.

9. A digital employee's reply generation device, characterized by, The computer readable storage medium stores computer executable instructions for causing a computer to perform a digital employee reply generation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer readable storage medium stores computer executable instructions for causing a computer to perform a digital employee reply generation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for training dialogue model, dialogue implementation method and related device

    CN118536572A

  • Inference type dialogue response method and system based on large model and storage medium

    CN120450039A