Artificial intelligence-based reply script generation method and device, equipment and medium
Patent Information
- Application Number
- CN202310789872.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-28
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2043-06-28
AI Technical Summary
[0003]基于此,有必要针对现有技术的在采用人工智能生成回复话术时,话术库包含能适应各种情况的话术,导致匹配话术时的噪音过多,使匹配话术的准确性不高的技术问题,提出了一种基于人工智能的回复话术生成方法、装置、设备及介质
[0016] This application's AI-based response script generation method involves inputting initial text into a preset script type recognition model to identify the script type, obtaining a target script type, determining a target script library based on the target script type and the target disease corresponding to the target medical consultation dialogue, and matching scripts from the target script library based on the initial text as the target response script. This dynamically determines the target script library, ensuring it contains only scripts applicable to the target script type and the target disease, reducing noise in the target script library, and improving the accuracy of script matching.
Smart Images

Figure CN116860931B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and healthcare technology, and in particular to a method, apparatus, device, and medium for generating response scripts based on artificial intelligence. Background Technology
[0002] With the development of computer technology, online medical consultations have become widely used. To improve the efficiency of consultations, artificial intelligence is used to generate response scripts, enabling patients to respond quickly. Currently, when using AI to generate response scripts, scripts are matched from a fixed script library. This requires the script library to contain scripts adaptable to various situations, resulting in excessive noise during script matching and low accuracy. Summary of the Invention
[0003] Based on this, it is necessary to address the technical problem in existing technologies where the script library contains scripts that can adapt to various situations, resulting in excessive noise during script matching and low accuracy. Therefore, an AI-based method, device, equipment, and medium for generating response scripts are proposed.
[0004] Firstly, an artificial intelligence-based method for generating response scripts is provided, the method comprising:
[0005] Obtain the initial text corresponding to the target consultation dialogue, wherein the target consultation dialogue is a dialogue between the target consultation subject and the target consultation subject, and the initial text is the text input by the target consultation subject;
[0006] The initial text is input into a preset speech type recognition model to identify the speech type and obtain the target speech type.
[0007] A target script library is determined based on the target script type and the target disease corresponding to the target consultation dialogue;
[0008] Based on the initial text, a script is matched from the target script library to serve as the target response script.
[0009] Secondly, an artificial intelligence-based response script generation device is provided, the device comprising:
[0010] The text acquisition module is used to acquire the initial text corresponding to the target consultation dialogue, wherein the target consultation dialogue is a dialogue between the target consultation subject and the target consultation subject, and the initial text is the text input by the target consultation subject.
[0011] The script type recognition module is used to input the initial text into a preset script type recognition model to recognize the script type and obtain the target script type.
[0012] The target dialogue database determination module is used to determine the target dialogue database based on the target dialogue type and the target disease corresponding to the target consultation dialogue;
[0013] The target response script determination module is used to match scripts from the target script library based on the initial text, and use them as target response scripts.
[0014] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described artificial intelligence-based response script generation method.
[0015] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described artificial intelligence-based response script generation method.
[0016] This application's AI-based response script generation method involves inputting initial text into a preset script type recognition model to identify the script type, obtaining a target script type, determining a target script library based on the target script type and the target disease corresponding to the target medical consultation dialogue, and matching scripts from the target script library based on the initial text as the target response script. This dynamically determines the target script library, ensuring it contains only scripts applicable to the target script type and the target disease, reducing noise in the target script library, and improving the accuracy of script matching. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] in:
[0019] Figure 1 This is an application environment diagram of an AI-based response script generation method in one embodiment;
[0020] Figure 2 This is a flowchart of an AI-based response script generation method in one embodiment;
[0021] Figure 3 This is a structural block diagram of an AI-based response script generation device in one embodiment;
[0022] Figure 4 This is a structural block diagram of a computer device in one embodiment;
[0023] Figure 5 This is another structural block diagram of a computer device in one embodiment. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] The AI-based response script generation method provided in this invention can be applied to, for example... Figure 1 In this application environment, client 110 communicates with server 120 via a network. Server 120 can receive and obtain the initial text corresponding to the target consultation dialogue from client 110. The target consultation dialogue is a dialogue between the target patient and the target patient, and the initial text is the text input by the target patient. Server 120 inputs the initial text into a preset dialogue type recognition model for dialogue type recognition to obtain the target dialogue type. Based on the target dialogue type and the target disease corresponding to the target consultation dialogue, a target dialogue library is determined. Based on the initial text, dialogue is matched from the target dialogue library as the target response dialogue. Server 120 sends the target response dialogue to client 110, thereby dynamically determining the target dialogue library. This ensures that the target dialogue library only contains dialogue applicable to the target dialogue type and the target disease, reducing noise in the target dialogue library and improving the accuracy of dialogue matching.
[0026] In another embodiment of this application, client 110 obtains the initial text corresponding to the target consultation dialogue, where the target consultation dialogue is a dialogue between a target patient and a target patient. The initial text is the text input by the target patient. The initial text is input into a preset dialogue type recognition model for dialogue type recognition to obtain the target dialogue type. Based on the target dialogue type and the target disease corresponding to the target consultation dialogue, a target dialogue library is determined. Based on the initial text, dialogue is matched from the target dialogue library as the target response dialogue. Specifically, client 110 downloads the total dialogue library from server 120. Client 110 matches the dialogue library from the total dialogue library based on the target dialogue type and the target disease corresponding to the target consultation dialogue as the target dialogue library.
[0027] The client 110 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server 120 can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will now be described in detail through specific embodiments.
[0028] Please see Figure 2 As shown, Figure 2 A flowchart illustrating an AI-based response script generation method provided in an embodiment of the present invention includes the following steps:
[0029] S1: Obtain the initial text corresponding to the target consultation dialogue, wherein the target consultation dialogue is a dialogue between the target consultation subject and the target consultation subject, and the initial text is the text input by the target consultation subject;
[0030] The target patient is the person being consulted in the targeted consultation dialogue. The patient is the person providing the consultation service, such as a doctor.
[0031] The target patient is the person being consulted in the intended consultation dialogue. The patient is the recipient of the consultation service, such as a patient.
[0032] The initial text is the text input by the target patient, intended for use in responding to the target patient.
[0033] Specifically, when the target patient conducts a consultation with another target patient, the target patient enters initial text in the dialog box on the patient's end. The client (i.e., the patient corresponding to the target consultation dialogue) retrieves the text from the dialog box and uses the retrieved text as the initial text.
[0034] Optionally, the initial text corresponding to the target consultation dialogue sent by the third-party application can be obtained.
[0035] S2: Input the initial text into a preset speech type recognition model to identify the speech type and obtain the target speech type;
[0036] Specifically, the initial text input is used to perform a preset speech type recognition model to classify and predict the speech type. The vector element with the largest value is obtained from the predicted vector, and this vector element is taken as the hit element. The speech type corresponding to the hit element is taken as the target speech type.
[0037] The dialogue type recognition model is a pre-trained multi-classification model. The dialogue type recognition model is based on a neural network trained on it.
[0038] The range of possible dialogue types includes, but is not limited to: dialogue about the cause of illness, lifestyle advice, the dangers of the disease, information about medications, treatment suggestions, and meaningless dialogue. Meaningless dialogue has no specific dialogue type.
[0039] S3: Determine the target script library based on the target script type and the target disease corresponding to the target consultation dialogue;
[0040] Specifically, based on the target dialogue type and the target disease corresponding to the target consultation dialogue, a dialogue library is matched from the main dialogue library, and the matched dialogue library is used as the target dialogue library. In other words, the target dialogue library only contains dialogues applicable to the target dialogue type and the target disease.
[0041] The target disease is the disease for which the consultation is conducted during the target consultation dialogue.
[0042] The target script library includes: script type, disease to be diagnosed, and script set. The script set contains multiple scripts.
[0043] Optionally, if the target speech type is meaningless, then the default speech library corresponding to the target disease in the target speech library shall be used as the target speech library.
[0044] S4: Based on the initial text, match a script from the target script library as the target response script.
[0045] Specifically, a similarity calculation is performed on the initial text and each dialogue in the target dialogue library. The most similarity is extracted from each of the calculated similarities, and the dialogue corresponding to the extracted similarity is used as the target response dialogue.
[0046] The target response is a text.
[0047] In other words, this application will replace the initial text with the target response script, and finally use the target response script to reply to the target consultation recipient.
[0048] Optionally, a similarity calculation is performed between the initial text and each script in the target script library. If the larger the similarity, the more similar the text, then the largest similarity that is greater than a first value is extracted from the calculated similarities. If the similarity is found, the script corresponding to the extracted similarity is used as the target response script. If the similarity is not found, then the initial text is used as the target response script. If the smaller the similarity, the more similar the text, then the smallest similarity that is less than a second value is extracted from the calculated similarities. If the similarity is found, then the script corresponding to the extracted similarity is used as the target response script. If the similarity is not found, then the initial text is used as the target response script.
[0049] This embodiment identifies the target dialogue type by inputting the initial text into a preset dialogue type recognition model. Based on the target dialogue type and the target disease corresponding to the target consultation dialogue, a target dialogue library is determined. Dialogue is then matched from the target dialogue library based on the initial text to serve as the target response dialogue. This achieves dynamic determination of the target dialogue library, ensuring that it contains only dialogue applicable to the target dialogue type and the target disease, reducing noise in the target dialogue library and improving the accuracy of dialogue matching.
[0050] In one embodiment, the step of determining the target script library based on the target script type and the target disease corresponding to the target consultation dialogue includes:
[0051] S31: Based on the target script type, the target disease, and the object identifier corresponding to the target consultation object, filter the first script library from each first script library to obtain the filtering result;
[0052] Specifically, the first dialogue library is selected from each first dialogue library to correspond to the object identifier of the target dialogue type, the target disease, and the target consultation object. If the first dialogue library is selected, the selection result is determined to be successful; if the first dialogue library is not selected, the selection result is determined to be unsuccessful.
[0053] The first script library includes: script type, disease to be diagnosed, object identifier corresponding to the diagnostic subject, and script set.
[0054] An object identifier can be data that uniquely identifies an object, such as an object name or an object ID.
[0055] S32: If the filtering result is successful, then the first script library corresponding to the filtering result shall be used as the target script library;
[0056] Specifically, if the screening result is successful, that is, a personalized script library is provided for the target patient, then the first script library corresponding to the screening result is used as the target script library.
[0057] S33: If the screening result is unsuccessful, then according to the target speech type and the target disease, the second speech library is selected from each second speech library as the target speech library.
[0058] Specifically, if the screening result is unsuccessful, that is, there is no personalized script library for the target patient, then the second script library corresponding to the target script type and the target disease is selected from each second script library, and the selected second script library is used as the target script library.
[0059] The second script library includes: script type, disease to be diagnosed, and script collection. The script collection contains multiple scripts.
[0060] In another embodiment of this application, the step of determining the target dialogue library based on the target dialogue type and the target disease corresponding to the target consultation dialogue includes: selecting second dialogue libraries from various second dialogue libraries based on the target dialogue type and the target disease, and using these second dialogue libraries as the target dialogue library. This provides a dialogue library specifically for the target dialogue type and the target disease, ensuring that the target dialogue library contains only dialogues applicable to the target dialogue type and the target disease, reducing noise in the target dialogue library, and improving the accuracy of dialogue matching.
[0061] In this embodiment, when a first script library corresponding to the target script type, the target disease, and the object identifier corresponding to the target patient exists, this first script library is used as the target script library. This ensures the target script library contains only scripts applicable to the target patient, the target script type, and the target disease, reducing noise in the target script library, improving the accuracy of matching scripts, providing personalized consultations for the target patient, and increasing patient satisfaction. Conversely, when a first script library corresponding to the target script type, the target disease, and the object identifier corresponding to the target patient does not exist, a second script library corresponding to the target script type and the target disease is used as the target script library. This provides a script library specifically for the target script type and the target disease, ensuring the target script library contains only scripts applicable to the target script type and the target disease, reducing noise in the target script library, and improving the accuracy of matching scripts.
[0062] In one embodiment, before the step of obtaining the initial text corresponding to the target consultation dialogue, the method further includes:
[0063] S51: Obtain historical consultation data, wherein each consultation data in the historical consultation data includes: consultation object identifier, consultation disease, and historical dialogue text;
[0064] Historical consultation data refers to consultation data formed from historical consultation dialogues.
[0065] The patient identification is the identifier for the patient being consulted. Historical dialogue text contains multi-turn dialogue text.
[0066] Specifically, it can obtain historical consultation data input by the user, historical consultation data can be obtained from a preset storage space, or historical consultation data can be obtained from third-party applications.
[0067] S52: Extract each round of dialogue text of the consulted subject from each of the historical dialogue texts, as the target single-round dialogue text;
[0068] Specifically, each round of dialogue text of the consulted subject is extracted from each of the historical dialogue texts, and each extracted round of dialogue text is used as a target single-round dialogue text.
[0069] S53: Obtain the speech type annotation value corresponding to each target single-turn dialogue text;
[0070] Specifically, the user obtains the wording type annotation value corresponding to each target single-turn dialogue text.
[0071] Optionally, the annotation value of the speech type output by the annotation model for each target single-turn dialogue text can also be obtained.
[0072] The labeled model is a multi-classification model obtained by training a neural network.
[0073] The script type label value is the accurate result of the script type.
[0074] S54: Combine the target single-turn dialogue text, the consultation object identifier, the consultation disease, and the dialogue type label value corresponding to the target single-turn dialogue text into a data combination;
[0075] Specifically, the target single-turn dialogue text, the consultation object identifier, the consultation disease, and the script type label value corresponding to the target single-turn dialogue text are associated, and the associated data is combined as a data set.
[0076] S55: Train the initial model according to each of the data combinations, and use the trained initial model as the speech type recognition model.
[0077] Specifically, the initial model is trained based on the target single-turn dialogue text, the disease being diagnosed, and the speech type labeling value in each of the data combinations.
[0078] When the loss value of the initial model converges to a preset value, or when the number of training iterations of the initial model reaches a preset number, it means that the training of the initial model is over, and the initial model at this time is used as the speech type recognition model.
[0079] Specifically, the target single-turn dialogue text and the disease being diagnosed are input into the initial model for prediction. The loss value of the initial model is calculated based on the predicted data and the speech type label value. The network parameters of the initial model are updated based on the loss value to achieve one training of the initial model.
[0080] This embodiment determines each of the aforementioned data combinations based on historical consultation data, thereby enabling the initial model to be trained based on each of the aforementioned data combinations. This allows the initial model to learn the features of historical consultation data, improving the accuracy of dialogue type recognition based on historical consultation data.
[0081] In one embodiment, the initial model includes: a TinyBERT model and a classification prediction layer connected in sequence, wherein the classification prediction layer is a network layer using the softmax activation function.
[0082] TinyBERT is a compressed version of BERT (Bidirectional Encoder Representation from Transformers). TinyBERT is mainly used for model distillation compression, which can retain 96% of the performance of BERT-base, but is 7 times smaller and 9 times faster than BERT.
[0083] The softmax activation function is a normalization function. Network layers using the softmax activation function perform classification predictions on features extracted by the TinyBERT model.
[0084] Since this application generates response scripts in real time from the text input by the target patient, the real-time requirements are relatively high. This embodiment improves the speed of the script type recognition model trained based on the initial model by using the TinyBERT model and the classification prediction layer connected in sequence as the initial model, thereby making the script type recognition model suitable for scenarios with high real-time requirements.
[0085] In one embodiment, after the step of combining the target single-turn dialogue text, the consultation object identifier, the consultation disease, and the dialogue type label value corresponding to the target single-turn dialogue text as a data combination, the method further includes:
[0086] S561: Using a preset clustering method, cluster each of the data combinations under the first dimension to obtain multiple first cluster sets corresponding to the first dimension, wherein the first dimension is the dimension under the disease being consulted, the text type label value, and the object being consulted;
[0087] Specifically, a preset clustering method is used to cluster each of the data combinations in the first dimension, and each set obtained by clustering is taken as a first cluster set.
[0088] S562: Select the first cluster set with the highest frequency of use from each of the first cluster sets corresponding to the first dimension, as the first candidate set;
[0089] Specifically, the first cluster set with the highest frequency of use is selected from each of the first cluster sets corresponding to the first dimension, and the selected first cluster set is used as the first candidate set.
[0090] Optionally, the first cluster set with the highest total frequency of use is selected from each of the first cluster sets corresponding to the first dimension as the first candidate set, wherein the total frequency of use is the sum of all usage counts (the usage count is the number of times the target single-turn dialogue text is used) corresponding to all the target single-turn dialogue texts in the first cluster set.
[0091] Optionally, the first cluster set with the highest average frequency of use is selected from each of the first cluster sets corresponding to the first dimension as the first candidate set, wherein the average frequency of use is the average of all usage times (the usage times are the usage times of the target single-turn dialogue text) corresponding to all the target single-turn dialogue texts in the first cluster set.
[0092] S563: Select the K most frequently used target single-turn dialogue texts from the first candidate set as the first dialogue library corresponding to the first dimension.
[0093] K is a positive integer.
[0094] Specifically, the K most frequently used target single-turn dialogue texts are selected from the first candidate set, and the selected K target single-turn dialogue texts are used as the first speech script library corresponding to the first dimension. The most frequently used text is the number of times a target single-turn dialogue text is used. The selected target single-turn dialogue texts are used as the speech scripts in the first speech script library.
[0095] Understandably, there are multiple first dimensions, and steps S561 to S563 are used to determine the first speech library for each first dimension. Each first dimension is different.
[0096] This embodiment determines the first dialogue library based on the dimensions of the diagnosed disease, the dialogue type label value, and the diagnosis object identifier. This ensures that the first dialogue library contains only the target single-turn dialogue text under the same first dimension, reducing noise in the first dialogue library. Furthermore, it performs clustering on the dimensions of the diagnosed disease, the dialogue type label value, and the diagnosis object identifier, and selects the K most frequently used target single-turn dialogue texts from the first cluster with the highest usage frequency as the first dialogue library. This ensures that the more popular target single-turn dialogue texts are used as dialogues in the first dialogue library, improving the accuracy of responses based on the first dialogue library.
[0097] In one embodiment, after the step of combining the target single-turn dialogue text, the consultation object identifier, the consultation disease, and the dialogue type label value corresponding to the target single-turn dialogue text as a data combination, the method further includes:
[0098] S571: Using a preset clustering method, cluster each of the data combinations under the second dimension to obtain multiple second cluster sets corresponding to the second dimension, wherein the second dimension is the dimension under the diagnosis disease and the type of dialogue;
[0099] Specifically, a preset clustering method is used to cluster each of the data combinations in the second dimension, and each set obtained by clustering is used as a second cluster set.
[0100] S572: Select the second cluster set with the highest frequency of use from each of the second cluster sets corresponding to the second dimension, as the second candidate set;
[0101] Specifically, the second cluster set with the highest frequency of use is selected from each of the second cluster sets corresponding to the second dimension, and the selected second cluster set is used as the second candidate set.
[0102] Optionally, the second cluster set with the highest total frequency of use is selected from each of the second cluster sets corresponding to the second dimension as the second candidate set, wherein the total frequency of use is the sum of all usage counts (the usage count is the number of times the target single-turn dialogue text is used) for all the target single-turn dialogue texts in the second cluster set.
[0103] Optionally, the second cluster set with the highest average frequency of use is selected from each of the second cluster sets corresponding to the second dimension as the second candidate set, wherein the average frequency of use is the average of all usage times (the usage times are the usage times of the target single-turn dialogue text) corresponding to all the target single-turn dialogue texts in the second cluster set.
[0104] S573: Select the M most frequently used target single-turn dialogue texts from the second candidate set as the second speech library corresponding to the second dimension.
[0105] M is a positive integer.
[0106] Specifically, the M target single-turn dialogue texts with the highest usage frequency are selected from the second candidate set, and the selected M target single-turn dialogue texts are used as the second speech library corresponding to the second dimension. The highest usage frequency is the number of times the two target single-turn dialogue texts are used. The selected target single-turn dialogue texts are used as speech in the second speech library.
[0107] Understandably, there are multiple second dimensions, and steps S571 to S573 are used to determine the second speech library for each second dimension. Each second dimension is different.
[0108] This embodiment determines the second dialogue library based on the dimensions of the diagnosed disease and the dialogue type label value, so that the second dialogue library contains only the target single-turn dialogue text under the same second dimension, reducing noise in the second dialogue library; moreover, it performs clustering on the dimensions of the diagnosed disease, the dialogue type label value, and the diagnosis object identifier, and selects the M target single-turn dialogue texts with the highest usage frequency in the second cluster set as the second dialogue library, thereby using the more popular target single-turn dialogue texts as the dialogue in the second dialogue library, improving the accuracy of responses based on the second dialogue library.
[0109] In one embodiment, the preset clustering method employs the DBSCAN density clustering method, and the distance equation of the preset clustering method, distance(x,y), is expressed as:
[0110]
[0111] Wherein, log is a logarithmic function, x is the target single-turn dialogue text in the data combination, y is the target single-turn dialogue text in the data combination, editDistance is the edit distance function, jaccardDistance is the Jaccard distance function, max is the maximum value to be calculated, and length is the number of characters in the calculated text.
[0112] DBSCAN, Density-Based Spatial Clustering of Applications with Noise, is a density-based clustering algorithm.
[0113] This embodiment combines the edit distance function, Jaccard distance function, normalization, and logarithmic smoothing in its distance equation, making the distance less sensitive to text length and focusing more on the relative distance between two text segments, thus improving the accuracy of clustering.
[0114] Please see Figure 3 As shown, in one embodiment, an artificial intelligence-based response script generation device is provided, the device comprising:
[0115] The text acquisition module 801 is used to acquire the initial text corresponding to the target consultation dialogue, wherein the target consultation dialogue is a dialogue between the target consultation subject and the target consultation subject, and the initial text is the text input by the target consultation subject.
[0116] The script type recognition module 802 is used to input the initial text into a preset script type recognition model to perform script type recognition and obtain the target script type.
[0117] The target dialogue database determination module 803 is used to determine the target dialogue database based on the target dialogue type and the target disease corresponding to the target consultation dialogue;
[0118] The target response script determination module 804 is used to match a script from the target script library based on the initial text, and use it as the target response script.
[0119] This embodiment identifies the target dialogue type by inputting the initial text into a preset dialogue type recognition model. Based on the target dialogue type and the target disease corresponding to the target consultation dialogue, a target dialogue library is determined. Dialogue is then matched from the target dialogue library based on the initial text to serve as the target response dialogue. This achieves dynamic determination of the target dialogue library, ensuring that it contains only dialogue applicable to the target dialogue type and the target disease, reducing noise in the target dialogue library and improving the accuracy of dialogue matching.
[0120] In one embodiment, the step of determining the target dialogue library based on the target dialogue type and the target disease corresponding to the target medical consultation dialogue in the target dialogue library determination module 803 includes:
[0121] Based on the target script type, the target disease, and the object identifier corresponding to the target consultation object, the first script library is filtered from each first script library to obtain the filtering results;
[0122] If the filtering result is successful, then the first script library corresponding to the filtering result is taken as the target script library;
[0123] If the screening result is unsuccessful, then based on the target speech type and the target disease, the second speech library is selected from each of the second speech libraries as the target speech library.
[0124] In one embodiment, the apparatus further includes a data combination generation module and a model training module;
[0125] The data combination generation module is used to acquire historical consultation data, wherein each consultation data in the historical consultation data includes: consultation object identifier, consultation disease, and historical dialogue text; extract each round of dialogue text of the consultation object from each of the historical dialogue texts as target single-round dialogue text; acquire the corresponding speech type label value for each target single-round dialogue text; and combine the target single-round dialogue text, as well as the consultation object identifier, consultation disease, and speech type label value corresponding to the target single-round dialogue text, into a data combination;
[0126] The model training module is used to train the initial model based on each of the data combinations, and use the trained initial model as the speech type recognition model.
[0127] In one embodiment, the initial model of the model training module includes: a TinyBERT model and a classification prediction layer connected in sequence, wherein the classification prediction layer is a network layer using the softmax activation function.
[0128] In one embodiment, the apparatus further includes a first script generation module, the first script generation module being used for:
[0129] A preset clustering method is used to cluster each of the data combinations under the first dimension to obtain multiple first cluster sets corresponding to the first dimension, wherein the first dimension is the dimension under the diagnosis disease, the script type label value and the diagnosis object identifier;
[0130] Select the first cluster set with the highest frequency of use from each of the first cluster sets corresponding to the first dimension, and use it as the first candidate set;
[0131] Select the K most frequently used target single-turn dialogue texts from the first candidate set as the first dialogue library corresponding to the first dimension.
[0132] In one embodiment, the apparatus further includes a second script generation module, the second script generation module being used for:
[0133] A preset clustering method is used to cluster each of the data combinations under the second dimension to obtain multiple second cluster sets corresponding to the second dimension, wherein the second dimension is the dimension under the diagnosis disease and the type of dialogue.
[0134] Select the second cluster set with the highest frequency of use from each of the second cluster sets corresponding to the second dimension, and use it as the second candidate set;
[0135] Select the M target single-turn dialogue texts with the highest frequency of use from the second candidate set as the second speech library corresponding to the second dimension.
[0136] In one embodiment, the preset clustering method employs the DBSCAN density clustering method, and the distance equation of the preset clustering method, distance(x,y), is expressed as:
[0137]
[0138] Wherein, log is a logarithmic function, x is the target single-turn dialogue text in the data combination, y is the target single-turn dialogue text in the data combination, editDistance is the edit distance function, jaccardDistance is the Jaccard distance function, max is the maximum value to be calculated, and length is the number of characters in the calculated text.
[0139] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a server-side method for generating response dialogue based on artificial intelligence.
[0140] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements client-side functions or steps of an artificial intelligence-based response dialogue generation method.
[0141] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, performs the following steps:
[0142] Obtain the initial text corresponding to the target consultation dialogue, wherein the target consultation dialogue is a dialogue between the target consultation subject and the target consultation subject, and the initial text is the text input by the target consultation subject;
[0143] The initial text is input into a preset speech type recognition model to identify the speech type and obtain the target speech type.
[0144] A target script library is determined based on the target script type and the target disease corresponding to the target consultation dialogue;
[0145] Based on the initial text, a script is matched from the target script library to serve as the target response script.
[0146] This embodiment identifies the target dialogue type by inputting the initial text into a preset dialogue type recognition model. Based on the target dialogue type and the target disease corresponding to the target consultation dialogue, a target dialogue library is determined. Dialogue is then matched from the target dialogue library based on the initial text to serve as the target response dialogue. This achieves dynamic determination of the target dialogue library, ensuring that it contains only dialogue applicable to the target dialogue type and the target disease, reducing noise in the target dialogue library and improving the accuracy of dialogue matching.
[0147] In one embodiment, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that, when executed by a processor, performs the following steps:
[0148] Obtain the initial text corresponding to the target consultation dialogue, wherein the target consultation dialogue is a dialogue between the target consultation subject and the target consultation subject, and the initial text is the text input by the target consultation subject;
[0149] The initial text is input into a preset speech type recognition model to identify the speech type and obtain the target speech type.
[0150] A target script library is determined based on the target script type and the target disease corresponding to the target consultation dialogue;
[0151] Based on the initial text, a script is matched from the target script library to serve as the target response script.
[0152] This embodiment identifies the target dialogue type by inputting the initial text into a preset dialogue type recognition model. Based on the target dialogue type and the target disease corresponding to the target consultation dialogue, a target dialogue library is determined. Dialogue is then matched from the target dialogue library based on the initial text to serve as the target response dialogue. This achieves dynamic determination of the target dialogue library, ensuring that it contains only dialogue applicable to the target dialogue type and the target disease, reducing noise in the target dialogue library and improving the accuracy of dialogue matching.
[0153] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0154] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0155] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0156] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for generating response scripts based on artificial intelligence, the method comprising: Obtain the initial text corresponding to the target consultation dialogue, wherein the target consultation dialogue is a dialogue between the target consultation subject and the target consultation subject, and the initial text is the text input by the target consultation subject; The initial text is input into a preset speech type recognition model to identify the speech type and obtain the target speech type. A target script library is determined based on the target script type and the target disease corresponding to the target consultation dialogue; Based on the initial text, a script is matched from the target script library as the target response script; Before the step of obtaining the initial text corresponding to the target consultation dialogue, the method further includes: Obtain historical consultation data, wherein each consultation data in the historical consultation data includes: consultation subject identifier, consultation disease, and historical dialogue text; Extract each round of dialogue text of the person being consulted from each of the historical dialogue texts, as the target single-round dialogue text; Obtain the speech type annotation value corresponding to each target single-turn dialogue text; The target single-turn dialogue text, along with the corresponding consultation object identifier, consultation disease, and dialogue type label value, are combined into a single data set. Based on each of the data combinations, the initial model is trained, and the trained initial model is used as the speech type recognition model; wherein, during training, the target single-turn dialogue text and the disease consultation are input into the initial model for prediction; the initial model includes: a TinyBERT model and a classification prediction layer connected in sequence, wherein the classification prediction layer is a network layer using the softmax activation function.
2. The AI-based response script generation method according to claim 1, characterized in that, The step of determining the target script library based on the target script type and the target disease corresponding to the target consultation dialogue includes: Based on the target script type, the target disease, and the object identifier corresponding to the target consultation object, the first script library is filtered from each first script library to obtain the filtering results; If the filtering result is successful, then the first script library corresponding to the filtering result is taken as the target script library; If the screening result is unsuccessful, then based on the target speech type and the target disease, the second speech library is selected from each of the second speech libraries as the target speech library.
3. The artificial intelligence-based response script generation method according to claim 1, characterized in that, After combining the target single-turn dialogue text, the corresponding consultation object identifier, the consultation disease, and the dialogue type label value into a data set, the method further includes: A preset clustering method is used to cluster each of the data combinations under the first dimension to obtain multiple first cluster sets corresponding to the first dimension, wherein the first dimension is the dimension under the diagnosis disease, the script type label value and the diagnosis object identifier; Select the first cluster set with the highest frequency of use from each of the first cluster sets corresponding to the first dimension, and use it as the first candidate set; Select the K most frequently used target single-turn dialogue texts from the first candidate set as the first dialogue library corresponding to the first dimension.
4. The AI-based response script generation method according to claim 1, characterized in that, After combining the target single-turn dialogue text, the corresponding consultation object identifier, the consultation disease, and the dialogue type label value into a data set, the method further includes: A preset clustering method is used to cluster each of the data combinations under the second dimension to obtain multiple second cluster sets corresponding to the second dimension, wherein the second dimension is the dimension under the diagnosis disease and the type of dialogue. Select the second cluster set with the highest frequency of use from each of the second cluster sets corresponding to the second dimension, and use it as the second candidate set; Select the M most frequently used target single-turn dialogue texts from the second candidate set as the second dialogue library corresponding to the second dimension.
5. The artificial intelligence-based response script generation method according to claim 3 or 4, characterized in that, The preset clustering method adopts the DBSCAN density clustering method, and the distance equation of the preset clustering method is expressed as follows: for: Where log is the logarithmic function. It is the target single-turn dialogue text in the data combination. It is the target single-turn dialogue text in the data combination. It is an edit distance function. It is the Jaccard distance function. It calculates the maximum value. It calculates the number of characters in the text.
6. A response script generation device based on artificial intelligence, characterized in that, The device includes: The text acquisition module is used to acquire the initial text corresponding to the target consultation dialogue, wherein the target consultation dialogue is a dialogue between the target consultation subject and the target consultation subject, and the initial text is the text input by the target consultation subject. The script type recognition module is used to input the initial text into a preset script type recognition model to recognize the script type and obtain the target script type. The target dialogue database determination module is used to determine the target dialogue database based on the target dialogue type and the target disease corresponding to the target consultation dialogue; The target response script determination module is used to match a script from the target script library based on the initial text, and use it as the target response script; Before the step of acquiring the initial text corresponding to the target consultation dialogue, the text acquisition module is further configured to: Obtain historical consultation data, wherein each consultation data in the historical consultation data includes: consultation subject identifier, consultation disease, and historical dialogue text; Extract each round of dialogue text of the person being consulted from each of the historical dialogue texts, as the target single-round dialogue text; Obtain the speech type annotation value corresponding to each target single-turn dialogue text; The target single-turn dialogue text, along with the corresponding consultation object identifier, consultation disease, and dialogue type label value, are combined into a single data set. Based on each of the data combinations, the initial model is trained, and the trained initial model is used as the speech type recognition model; wherein, during training, the target single-turn dialogue text and the disease consultation are input into the initial model for prediction; the initial model includes: a TinyBERT model and a classification prediction layer connected in sequence, wherein the classification prediction layer is a network layer using the softmax activation function.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the artificial intelligence-based response script generation method as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the artificial intelligence-based response script generation method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Online inquiry data processing method and device and computer equipment
CN112035615A
Dialogue method, device and equipment, computer readable storage medium and program product
CN114048299A