Electronic device and control method thereof

By automatically training the neural network model of the chatbot and utilizing voice data from customer and consultant conversations, the problem of high training costs for administrators was solved, and the accuracy of the chatbot's understanding and response to customer inquiries was improved.

CN114830229BActive Publication Date: 2025-11-28SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080085442.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-01-23
Filing Date
2020-05-22
Publication Date
2025-11-28
Estimated Expiration
2040-05-22

AI Technical Summary

Technical Problem

Administrators need to spend a lot of effort and time training and maintaining chatbots to improve their responsiveness to customer inquiries.

Method used

By using voice data from conversations between clients and consultants, a neural network model included in the chatbot is automatically trained to distinguish between client and consultant voices and to train the neural network model based on different categories of information.

Benefits of technology

It reduces the cost and time of training and maintaining chatbots, and improves the accuracy of chatbots in understanding customer inquiries and the accuracy of category information in response content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114830229B_ABST
    Figure CN114830229B_ABST
Patent Text Reader

Abstract

An electronic device is provided. The electronic device includes a memory storing recording data including conversation content and at least one instruction, and a processor configured to, by executing the at least one instruction, input first data corresponding to a first speech in the conversation content into a first neural network model, and acquire category information of the first data, and acquire category information of second data corresponding to a second speech in the conversation content. The processor is configured to, based on the category information of the first data and the category information of the second data being different, train the first neural network model based on the category information of the second data and the first data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The disclosure relates to an electronic device that identifies a category of a customer's voice and a control method thereof. BACKGROUND

[0002] Recently, as the robot industry is developing, technology for a service robot to provide consultation and respond to a customer is being developed.

[0003] For example, a chat robot of a call center provides a response to a customer's inquiry instead of a human consultant.

[0004] In the past, in order to train such a chat robot, an administrator needed to find a situation in which the chat robot incorrectly recognized a customer's inquiry, extract learning data by himself or herself, and update the chat robot. SUMMARY

[0005] TECHNICAL PROBLEM

[0006] In this case, there is a problem in which the administrator needs to spend a lot of effort and time to train and maintain the chat robot.

[0007] SOLUTION TO THE PROBLEM

[0008] Embodiments of the disclosure address the above-described needs and provide an electronic device that automatically trains an artificial intelligence model included in a chat robot using voice data including conversation content between a customer and a consultant and a control method thereof.

[0009] The electronic device according to an example embodiment of the disclosure includes a memory that stores recorded data including conversation content and at least one instruction, and a processor that is configured to, by executing the at least one instruction, input first data corresponding to a first voice in the conversation content into a first neural network model and acquire category information of the first data, and acquire category information of second data corresponding to a second voice in the conversation content. Based on the category information of the first data and the category information of the second data being different, the processor can train the first neural network model based on the category information of the second data and the first data.

[0010] The control method of the electronic device that stores recorded data including conversation content according to an example embodiment of the disclosure includes inputting first data corresponding to a first voice in the conversation content into a first neural network model and acquiring category information of the first data, acquiring category information of second data corresponding to a second voice in the conversation content, and based on the category information of the first data and the category information of the second data being different, training the first neural network model based on the category information of the second data and the first data. BRIEF DESCRIPTION OF DRAWINGS

[0011] The above and other aspects, features, and advantages of certain embodiments of the present disclosure will be more apparent from the following detailed description, taken in conjunction with the accompanying drawings, in which:

[0012] Figure 1 FIG. 1 is a diagram illustrating an example electronic device that replaces a consultant according to an embodiment of the present disclosure;

[0013] Figure 2 FIG. 2 is a flowchart illustrating an example training of a neural network model according to an embodiment of the present disclosure;

[0014] Figure 3 FIG. 3 is a flowchart illustrating an example process of acquiring second category information according to an embodiment of the present disclosure;

[0015] Figure 4 FIG. 4 is a diagram illustrating an example of a correspondence relationship between a first voice and a second voice according to an embodiment of the present disclosure;

[0016] Figure 5 FIG. 5 is a block diagram illustrating an example configuration of an example electronic device according to an embodiment of the present disclosure;

[0017] Figure 6 FIG. 6 is a diagram illustrating an example process of training a second neural network model according to an embodiment of the present disclosure; and

[0018] Figure 7 FIG. 7 is a flowchart illustrating an example method of controlling an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] Example embodiments of the present disclosure address the above-described needs and provide an electronic device that automatically trains an artificial intelligence model included in a chat robot using voice data including contents of a conversation between a customer and a consultant and a control method thereof.

[0020] Hereinafter, the present disclosure will be described in greater detail with reference to the accompanying drawings.

[0021] The terms used in the present disclosure will be described briefly, and then the present disclosure will be described in greater detail.

[0022] As terms used in the embodiments of the present disclosure, the general terms which are currently and widely used are selected in consideration of the functions in the present disclosure described in the present disclosure, so that the meaning of the terms can be obviously understood by a person having ordinary skill in the art. However, the terms used herein should be defined based on the meaning of the terms and the overall content of the present disclosure, not the names of the terms, and the terms used herein have the same or similar meanings as those currently widely used.

[0023] Various modifications can be made to various embodiments of the present disclosure, and various types of embodiments can exist. Therefore, various embodiments will be shown in the accompanying drawings, and these embodiments will be described in detail in the DETAILED DESCRIPTION. However, it should be noted that various example embodiments are not intended to limit the scope of the present disclosure to specific embodiments, but are understood to encompass all modifications, equivalents, or alternatives of embodiments within the scope of the ideas and technologies disclosed herein. When it is determined that a detailed explanation of related known technologies can unnecessarily obscure or confuse the gist of the present disclosure in describing embodiments, detailed explanation can be omitted.

[0024] The singular expression includes the plural expression, unless it conflicts with the context. In the present disclosure, terms such as "include" and "consist of" should be understood to specify that such features, numbers, steps, operations, elements, components, or combinations thereof described in the specification are described, but should not be understood to preclude the presence or possibility of adding one or more other features, numbers, steps, operations, elements, components, or combinations thereof.

[0025] The expression "at least one of A and / or B" should be interpreted to include any one of "A" or "B" or "A and B".

[0026] The expressions "first", "second", and the like used in the present disclosure can be used to describe various elements regardless of any order and / or degree of importance. Also, such expressions are used only to distinguish one element from another element, not to limit the elements.

[0027] The expression that one element (for example, a first element) described in the present disclosure is "(operatively or communicatively) coupled to" or "connected to" or "coupled with" another element (for example, a second element) should be understood to include a case where the one element is directly coupled to the other element, and a case where the one element is coupled to the other element through yet another element (for example, a third element).

[0028] In the present disclosure, a "module" or a "component" performs at least one function or operation, and can be implemented as hardware or software or a combination of hardware and software. Also, except for a "module" or a "component" that needs to be implemented as a specific hardware, a plurality of "modules" or "components" can be integrated into at least one module and implemented as at least one processor (not shown). Also, in the present specification, the term "user" can refer to a person using a terminal device or a device (for example: an artificial intelligence electronic device) using a terminal device.

[0029] Hereinafter, various example embodiments of the present disclosure will be described in greater detail with reference to the accompanying drawings. However, it should be noted that the present disclosure can be implemented in various different forms and is not limited to the embodiments described herein. In the drawings, portions irrelevant to the explanation can be omitted for the sake of clarity in explaining the present disclosure, and like reference numerals can be used to refer to like elements throughout the disclosure.

[0030] Hereinafter, example embodiments of the present disclosure will be described in greater detail with reference to the accompanying drawings.

[0031] Figure 1 FIG. 1 is a diagram illustrating an example electronic device that replaces a consultant according to an embodiment of the present disclosure.

[0032] The electronic device 100 can be a device that understands a user voice transmitted from a terminal device or the like and provides a response. For example, the electronic device 100 can be implemented as, for example, and without limitation, a chatbot, a smartphone, a server, or the like. The chatbot can, for example, include a device that understands a user voice and provides a response, and can, for example, include a device that simulates a human, and can, for example, be a chatbot that replaces a consultant at a call center. In the case where the electronic device 100 is implemented as a chatbot, the electronic device 100 can, for example, and without limitation, provide response information in the form of a voice or response information in the form of a visual UI corresponding to a customer inquiry content to a terminal device through the chatbot.

[0033] As described above, according to various example embodiments of the present disclosure, a process of training a chatbot can be automated. Accordingly, a cost spent for training and maintaining a chatbot can be reduced, and a time spent can also be shortened.

[0034] Because the chatbot is trained in distinguishing a voice of a customer from a voice of a consultant in voice data including a conversation content between the customer and the consultant, an accuracy of the chatbot understanding a customer inquiry content is improved. Further, because a process of determining whether a voice of a consultant is a distinguished response or a combined response is included, an accuracy of category information of a response content can be improved. Accordingly, an accuracy of the chatbot understanding a customer inquiry content can be improved.

[0035] For the electronic device 100, it is important to accurately recognize an inquiry content to provide an appropriate response to a customer inquiry content. Hereinafter, a method for the electronic device 100 to train a neural network model for recognizing an inquiry content will be described in greater detail with reference to the accompanying drawings according to various example embodiments of the present disclosure.

[0036] Figure 2 FIG. 2 is a flowchart illustrating an example training of a neural network model according to an embodiment of the present disclosure.

[0037] The training of the neural network model that classifies the category of the inquiry content can be performed at the electronic device 100 according to an embodiment of the disclosure.

[0038] Referring to Figure 2 At operation S205, voice data (e.g., recording data) can be input into the electronic device 100. The voice data can include, for example, recording data including conversation content. For example, the recording data can be recording data including conversation content between a customer and a counselor, and can include a customer's voice and a counselor's voice. As an example, the customer's voice included in the recording data can include, for example, inquiry content, and the counselor's voice can include, for example, response content.

[0039] According to an embodiment of the disclosure, the recording data can include, for example, data including all conversation content during one call between a customer and a counselor. However, the disclosure is not limited thereto, and the recording data can be data including a customer's inquiry content and response content corresponding thereto. The electronic device 100 can identify an inquiry form in a customer's voice or text data corresponding to the customer's voice, and identify a response form in a counselor's voice or text data corresponding to the counselor's voice. The inquiry form can include, for example, but not limited to, not only a general question form, but also various cases in which a response is requested from the other party, and the response form can include, for example, but not limited to, not only a general statement form, but also various cases in which information is provided. In other words, the recording data can be data in which a customer's voice including an inquiry form and a counselor's voice including a response form are edited. Accordingly, information about greetings, customer information, etc. unnecessary for training a neural network model that classifies the category of inquiry content can be excluded from the recording data. However, hereinafter, for ease and convenience of explanation, the recording data will be assumed to be data including all conversation content during one call between a customer and a counselor.

[0040] The recording data can be data in which conversation content between a customer and a counselor is recorded in advance, or data in which conversation content between a customer and a counselor is recorded in real time.

[0041] The recording data can be input from an external device (e.g., an external server, an external database, etc.). The electronic device 100 can distinguish different voices included in the input recording data. In other words, at operation S210, the electronic device 100 can distinguish a first voice and a second voice included in the input recording data. For example, the first voice can correspond to a customer's voice, and the second voice can correspond to a counselor's voice. Hereinafter, for ease and convenience of explanation, an explanation will be made based on an example in which the first voice is a customer's voice and the second voice is a counselor's voice.

[0042] According to an embodiment of the disclosure, the electronic device 100 can identify the second voice in the recording data based on pre-stored voice profile information related to the second voice. For example, the voice profile information can include, for example, and without limitation, waveform information about the voice of the counselor, voice identification information including the name of the counselor, etc.

[0043] For example, the electronic device 100 can compare the frequency waveform information of the voice included in the voice profile information and the frequency waveform information of the voice included in the recording data, and identify the second voice in the recording data. The electronic device 100 can identify the remaining voice other than the second voice among the voice included in the recording data as the first voice, which can be, for example, the voice of the client. However, the disclosure is not limited thereto, and the second voice that is the voice of the counselor can be identified based on the content of the voice. For example, a voice corresponding to the content of the elicitation question, such as "What can I do for you?", or the content that identifies personal information can be identified as the second voice that is the voice of the counselor. The information for identifying the content of the second voice as above can be pre-stored in the electronic device 100 or received from the outside. If at least some of the second voice is identified based on the content of the recording data, the electronic device 100 can acquire the voice profile information corresponding to the voice of the counselor based on the feature information of the identified second voice.

[0044] In operation S215, the electronic device 100 can perform voice recognition on the first voice and acquire text data corresponding to the first voice (hereinafter referred to as first text data). For example, the electronic device 100 can perform voice recognition on the first voice using a neural network model for voice recognition. However, the disclosure is not limited thereto, and voice recognition on the first voice can be performed at an external device, an external server, etc., and the electronic device 100 can receive the first text data corresponding to the first voice from the outside.

[0045] In operation S220, the electronic device 100 can identify whether the first text data is in the form of a question. For example, the form of a question can include not only a general interrogative form, but also various cases in which a response is requested from the other party. In the case where the first text data is not in the form of a question at operation S220-N, the electronic device 100 can perform voice recognition on another first voice.

[0046] Further, in operation S225, the electronic device 100 can perform voice recognition on the second voice and acquire text data corresponding to the second voice (hereinafter referred to as second text data). The voice recognition on the second voice can be performed in the same / similar manner as the voice recognition on the first voice.

[0047] The electronic device 100 can identify whether the second text data is a response form at operation S230. For example, the response form can include not only a general statement form but also various cases of providing information. In a case where the second text data is not a response form at operation S230-N, the electronic device 100 can perform speech recognition on another second speech.

[0048] In a case where the first text data is an inquiry form at operation S220-Y and the second text data is a response form at operation S230-Y, the electronic device 100 can identify whether the second speech corresponding to the second text data is uttered after the first speech corresponding to the first text data at operation S235. This is because the second speech of the consultant uttered after the first speech of the customer inquiry can be a response corresponding to the customer inquiry. This will be described below with reference to FIG. 2B. Figure 4 This will be described in detail.

[0049] In a case where the second speech is not uttered after the first speech at operation S235-N, in other words, in a case where the second speech is uttered before the first speech, the electronic device 100 can identify that the first text data and the second text data do not match. Accordingly, the electronic device 100 can perform speech recognition on the first speech and the second speech different from the previous second speech, thereby identifying the matched first text data and the second text data.

[0050] If it is identified that the second speech corresponding to the second text data is uttered after the first speech corresponding to the first text data at operation S235-Y, the electronic device 100 can input the first text data into the first neural network model. When the first text data is input into the first neural network model, the first neural network model can output at least one category information and a probability value of the first text data at operation S240. The probability value can be a probability value of the accuracy of the data of the first speech being classified as the category information. According to an embodiment of the disclosure, the first neural network model can be trained using text data corresponding to the first speech and category information of the text data as an input data and an output data pair. For example, if the first text data corresponding to the first speech is input, the first neural network model can be trained to output probability information corresponding to each of a plurality of predefined categories.

[0051] For example, based on the first text data corresponding to the first speech "weak air conditioner", the first neural network model can output category information and a probability value of the first text data corresponding to the first speech as "air conditioner failure (category ID: 2000), probability value 0.9", "window open (category ID: 1010), probability value 0.1", and the like.

[0052] In this example, the electronic device 100 can acquire category information (first category information) of data corresponding to the first voice based on the highest probability value at operation S245. For example, the electronic device 100 can acquire "air conditioner failure (category ID: 2000)".

[0053] The category information can include, for example, information classified based on frequently asked questions (FAQs), and ranges having the same features can be classified as one category information. The features can include the type of the device, functions, etc. The category information can be classified according to a predetermined depth. For example, in a wide range, questions related to an air conditioner can be classified as the same category information, and in a narrow range, different functions of the air conditioner can be classified as separate category information. Such category information can be different for each business operator.

[0054] If it is identified at operation S235-Y that the second voice corresponding to the second text data is uttered after the first voice corresponding to the first text data, the electronic device 100 can input the second text data into the second neural network model.

[0055] When the second text data is input into the second neural network model, the second neural network model can output at least one category information and a probability value of the second text data at operation S250. According to an embodiment of the disclosure, the second neural network model can be trained using text data corresponding to the second voice and category information of the text data as an input data and an output data pair. For example, if the second text data corresponding to the second voice is input, the second neural network model can be trained to output probability information corresponding to each of a plurality of predefined categories.

[0056] For example, based on data corresponding to the second voice "select the desired temperature as the minimum temperature", the second neural network model can output category information and a probability value of data corresponding to the second voice as "air conditioner failure (category ID: 2000), probability value 0.9", "window open (category ID: 1010), probability value 0.1", etc.

[0057] In this example, the electronic device 100 can acquire category information (second category information) of data corresponding to the second voice based on the highest probability value at operation S255. For example, the electronic device 100 can acquire "air conditioner failure (category ID: 2000)". The operation S255 of acquiring category information of data corresponding to the second voice will be described below with reference to FIG. 2B. Figure 3 The operation S255 of acquiring category information of data corresponding to the second voice is described in detail.

[0058] Hereinafter, for convenience of explanation, the category information of data corresponding to the first voice can be referred to as first category information, and the category information of data corresponding to the second voice can be referred to as second category information.

[0059] At operation S260, the electronic device 100 can compare the first category information and the second category information output from the first neural network model and the second neural network model. The input data of the first neural network model is the first text data corresponding to the first voice, and the input data of the second neural network model is the second text data corresponding to the second voice, and thus the category information output from the first neural network model and the second neural network model can be separate independent information.

[0060] In the case where the first category information and the second category information are the same at operation S260-Y, the first neural network model recognizes that correct category information corresponding to the record data is output, and thus there is no need to use the record data as learning data for the first neural network model. Accordingly, at operation S265, the electronic device 100 can exclude the record data from the training of the first neural network model.

[0061] However, in the case where the first category information and the second category information are different at operation S260-N, the first neural network model determines that proper training based on the record data is not performed, and at operation S270, the record data can be used for the training of the first neural network model.

[0062] The electronic device 100 can train the first neural network model based on the first text data corresponding to the first voice and the second category information. In other words, based on the premise that the training of the second neural network model has been well performed and the second neural network model outputs more accurate category information than the first neural network model, the electronic device 100 can determine that the accuracy of the second category information is relatively higher than the first category information output from the first neural network model, and train the first neural network model using the first text data corresponding to the first voice and the second category information as input data and output data, respectively. For example, it will be assumed that the record data includes the inquiry content "cold wind does not suddenly come out from the air conditioner" and the response content "select the desired temperature as the minimum temperature." In this case, the data corresponding to the first voice can be "cold wind does not suddenly come out from the air conditioner," and the data corresponding to the second voice can be "select the desired temperature as the minimum temperature."

[0063] In the above example, it is assumed that text data corresponding to the first utterance "cold wind does not come out from the air conditioner suddenly" is input into the first neural network model and category information "air conditioner fan problem (category ID: 2050)" is acquired, and that text data corresponding to the first utterance "select the desired temperature as the minimum temperature" is input into the second neural network model and category information "air conditioner failure (category ID: 2000)" is acquired. The above case can fall into a case in which the first category information about the inquiry content and the second category information about the response content are different. In this case, it is assumed that the second neural network model is pre-trained to classify the category of the response content well and output category information having high reliability for the response content included in the record data. Also, in a case in which the category information acquired from the first neural network model and the category information acquired from the second neural network model are different, the first neural network model can be trained based on the second category information.

[0064] In other words, the electronic device 100 can train the first neural network model based on data corresponding to the first utterance and the second category information, and thereafter, in a case in which the utterance "cold wind does not come out from the air conditioner suddenly" is included in the first utterance, the first neural network model can output category information corresponding to "air conditioner failure (category ID: 2000)" instead of category information corresponding to "air conditioner fan problem (category ID: 2050)". Thus, the first neural network model can be trained so that the inquiry "cold wind does not come out from the air conditioner suddenly" is classified as "air conditioner failure (category ID: 2000)" not only in a case in which the customer makes an inquiry "there is a problem with the air conditioner" which is an inquiry directly mentioning a keyword corresponding to "air conditioner failure (category ID: 2000)", but also in a case in which the customer makes an inquiry "cold wind does not come out from the air conditioner suddenly". Thus, for the inquiry "cold wind does not come out from the air conditioner suddenly", a response corresponding to "air conditioner failure (category ID: 2000)" can be provided.

[0065] Each of the aforementioned first and second neural network models can include a plurality of neural network layers. Each of the plurality of neural network layers can have a plurality of weight values, and can perform a neural network operation through an operation between an operation result of a previous layer and the plurality of weight values. The plurality of weight values that the plurality of neural network layers have can be optimized and / or improved through a learning result of the neural network model. For example, the plurality of weight values can be updated such that a loss value or a cost value obtained from the neural network model is reduced or minimized in a learning process. The artificial neural network can include, for example, and without limitation, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or a deep Q-network, etc., but the disclosure is not limited to the aforementioned examples.

[0066] The output part of the first and second neural network models can be implemented so that a softmax process is possible. The softmax is a function that normalizes all input values to a value between 0 and 1 and always makes the sum of the output values 1, and can perform a function of outputting a probability value for each class. The output part of the first and second neural network models can be implemented so that an argmax process is possible. The argmax is a function of selecting a label having the highest probability among a plurality of labels, and here, it can perform a function of selecting a ratio having the maximum value among the probability values for each class. For example, in the case where the argmax process has been performed on the output part of each of the first and second neural network models, only one class information having the highest probability value can be output.

[0067] The first and second neural network models can have been trained by the electronic device 100 or a separate server / system through various learning algorithms. The learning algorithm can include, for example, a method of training a specific object device using a plurality of learning data, thereby enabling the specific object device to make a decision or a prediction by itself. Examples of the learning algorithm include, for example, and without limitation, supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, etc., and the learning algorithm in the disclosure is not limited to the above-described examples (excluding specific cases).

[0068] After the first neural network model is sufficiently trained through the above-described example, the first neural network model can be used for a chatbot function, as an example. For example, a device that performs a chatbot function can identify class information corresponding to a customer inquiry content through the first neural network model, and acquire and provide response content corresponding to the identified class information from a storage. In this case, in the storage, a sample of the response content corresponding to each class information can have been pre-stored.

[0069] As described above, the first neural network model trained according to the embodiment of the disclosure can be used when included in the chatbot function, but the second neural network model can not be used for the chatbot function. Hereinafter, the method of acquiring the second category information using the second neural network model will be described with reference to FIGS. 10 to 13. Figure 3 The method of acquiring the second category information using the second neural network model will be described in more detail.

[0070] Figure 3 is a flowchart illustrating an example process of acquiring the second category information according to an embodiment of the disclosure.

[0071] Figure 3 is a flowchart illustrating Figure 2 is a flowchart of the operation S255 of acquiring the second category information in

[0072] As described above, various embodiments of the disclosure are based on the premise that the response content included in the second text data corresponding to the second voice (e.g., the voice of the counselor) is a correct answer to the inquiry content, and the second category information of the second voice is information having a high degree of accuracy compared to the first category information. Although the form of making inquiries, expressions for making inquiries, etc. are different for each customer, the response content of the counselor is provided according to predetermined human data, and thus the form of the response content or the diversity of the expression of the response content is relatively small compared to the inquiry, and thus it can be assumed that the second category information is information having a high degree of accuracy compared to the first category information. In other words, the first neural network model is trained based on the second category information, and thus it is important to acquire accurate second category information. The second category information can be acquired through various operations.

[0073] When the second text data corresponding to the second voice is acquired at operation S305, the electronic device 100 can distinguish the sentences included in the second text data at operation S310. The sentence can refer to, for example, the smallest unit expressing complete content. However, the disclosure is not limited thereto, and it is obvious that the second text data can be divided by various standards, such as division based on intent. Hereinafter, for convenience of explanation, the division of the second text data based on the sentence will be described.

[0074] The electronic device 100 can sequentially input the text data corresponding to each of the plurality of sentences included in the second text data corresponding to the second voice to the second neural network model.

[0075] When each of the divided sentences is input to the second neural network model, the second neural network model can output the category information of each sentence and the probability value corresponding to the category information at operation S315.

[0076] For example, assume that the examples of "Please ventilate" and "Please open the air drum" are included in the second text data. The electronic device 100 can input each of "Please ventilate" and "Please open the air drum" included in the second text data into the second neural network model, and acquire the category information and the probability value of each of them. For example, the category information and the probability value of the sentence "Please ventilate" can be classified as "Air leakage problem (category ID: 3000), probability value 0.78", "Air conditioner odor problem, probability value 0.19", etc. In this case, the electronic device 100 can acquire the category information of the sentence "Please ventilate" as "Air leakage problem (category ID: 3000)" based on the highest probability value.

[0077] As described above, at operation S320, based on the acquired category information and the probability value of each sentence, the electronic device 100 can identify whether the probability value of the category information corresponding to each sentence is greater than or equal to a threshold value. If each of the acquired probability values is greater than or equal to the threshold value at operation S320-Y, at operation S325, the electronic device 100 can acquire each of the acquired category information as the second category information of each corresponding sentence. In other words, the electronic device 100 can identify the acquired category information as the second category information.

[0078] As an example, assume a case in which the threshold value is 0.7, "Please ventilate" is classified as the category information of "Air leakage problem" and the probability value is 0.78, and "Please open the air drum" is classified as the category information of "Air conditioner fan problem" and the probability value is 0.9. In this case, because the probability value of each sentence is all greater than or equal to the threshold value, the electronic device 100 can identify "Please ventilate" as the category information of "Air leakage problem (category ID: 3000)" and "Please open the air drum" as the category information of "Air conditioner fan problem (category ID: 2050)", and identify them as different category information from each other.

[0079] If the probability value of the category information corresponding to each sentence acquired at operation S320-N is less than the threshold value, at operation S330, the electronic device 100 can combine at least some of the plurality of sentences and input the combined sentence into the second neural network model.

[0080] For example, assume a case in which "Please ventilate" included in the second text data is classified as the category information of "Air conditioner odor problem" and the probability value is 0.5, and "Please open the air drum" is also classified as the category information of "Air conditioner odor problem" and the probability value is 0.65. In this case, because the probability value of each sentence is all less than the threshold value, the electronic device 100 can combine the sentences, such as "Please ventilate. Please open the air drum", and input the combined sentence into the second neural network model.

[0081] The second neural network model can output the category information and the probability value corresponding to the combined sentence at operation S335.

[0082] The electronic device 100 can identify whether the probability value of the category information corresponding to the combined sentence is greater than or equal to a threshold at operation S340. If the probability value of the combined sentence is greater than or equal to the threshold, the electronic device 100 can acquire the acquired category information as the second category information of the combined sentence at operation S345. In other words, the electronic device 100 can identify the acquired category information as the second category information. If the probability value of the combined sentence is less than the threshold, the electronic device 100 can acquire each of the acquired category information as the second category information of each corresponding sentence at operation S325.

[0083] For example, it is assumed that the combined sentence "Please ventilate. Please turn on the drum" is classified as the category information of the "air conditioner odor problem (category ID: 1500)" and the probability value is 0.85. In this case, because the probability value of the combined sentence is greater than the threshold (for example, 0.7), the electronic device 100 can identify the acquired category information (air conditioner odor problem (category ID: 1500)) as the second category information. In other words, "Please ventilate. Please turn on the drum" is a combination of each sentence, and can be identified as one response content corresponding to one inquiry content.

[0084] According to an embodiment of the disclosure, the electronic device 100 can combine only the sentences whose category information acquired at operation S325 is the same, input the combined sentence into the second neural network model, and acquire the probability value. For example, in the case where the category information of the sentence "Please ventilate" is classified as "air leakage problem (category ID: 3000)" and the category information of the sentence "Please turn on the drum" is classified as "air conditioner fan problem (category ID: 2050)" and each sentence has category information different from each other, the sentences can not be an object for combination. As described above, even in the case where the category information of different sentences is the same, but the probability value of the category information of each sentence is less than the threshold, and the text data is excluded from learning, if the probability value of the category information of the sentence of the combined sentence is greater than or equal to the threshold, the category information can be used in learning of the first neural network model using the first category related to the category information of the combined sentence.

[0085] However, depending on the situation, even in the case where the category information of each sentence is different, the sentences can be combined, and in the case where the probability value of the category information of the combined sentence is greater than the probability value of the category information of each sentence before the combination, the electronic device 100 can use the category information in learning of the first neural network model using the first category related to the category information of the combined sentence.

[0086] In a case where a conjunction is included in a sentence acquired by combining at least some of each of a plurality of sentences, the electronic device 100 can apply a weight value to a probability value of category information corresponding to the acquired sentence based on at least one of a type or a number of the conjunction. The conjunction can be, for example, a word that functions to naturally connect each of the sentences, and for example, the conjunction can be expressed as a correlative. Among the types of the conjunction, a causal conjunction such as "therefore, so" and a transitional conjunction such as "still, however" can be included. In other words, the conjunction indicates that two sentences are associated with each other, and thus, in a case where the two sentences are combined, it is expected that the electronic device 100 adds a weight value to a probability value of category information corresponding to the two sentences.

[0087] Depending on the situation, in a case where a plurality of sentences connected by a conjunction is included in the second text data, the electronic device 100 can not acquire category information and a probability value of each of the sentences, but can directly acquire category information and a probability value of the plurality of sentences (i.e., combined sentences) connected by the conjunction.

[0088] Based on feedback information of the customer on the response information, the electronic device 100 can determine whether to train the first neural network model using the first speech or the text data associated with the response information as learning data. For example, in a case where the first speech or the second speech matching the second category information of the first text data or the second text data is output through the second neural network model, and response information according to the second category information is provided, but the customer has negative feedback on the response information, the electronic device 100 can identify that the second category information does not correspond to the first speech or the first text data, and can not train the first neural network model based on the first speech or the first text data and the second category information. In other words, the electronic device 100 can determine learning data for training the first neural network model based on the feedback information of the customer.

[0089] Figure 4 FIG. 1 is a diagram illustrating an example corresponding relationship between a first speech and a second speech according to an embodiment of the disclosure.

[0090] As described above, according to various embodiments of the disclosure, the electronic device 100 can train the first neural network model by comparing first category information corresponding to the first speech and second category information corresponding to the second speech. In this case, the first speech of the first category information and the second speech of the second category information should be inquiry content and response content matching each other. In other words, in a case where the second speech as a comparison object is a speech including response information corresponding to the first speech, the first speech and the text can be used as learning data for training the first neural network model. Hereinafter, an embodiment in which the electronic device 100 identifies the second speech matching the first speech will be described.

[0091] Figure 4 is a diagram in which the first voice and the second voice included in the recorded data are indicated when they are distinguished from each other. In this regard, assume a case in which the voices are indicated according to the lapse of time, and are indicated in solid lines during the time when each voice is uttered, and the first voice is filtered as a question form and the second voice is filtered as a response form.

[0092] According to an embodiment of the disclosure, the electronic device 100 can match the first voice and the second voice based on the point of utterance of each voice. For example, the electronic device 100 can recognize the second voice uttered after the first voice as a candidate group matched with the first voice. This is because the response information of the consultant is uttered after the inquiry of the client.

[0093] According to an embodiment of the disclosure, the electronic device 100 can recognize a plurality of second voices 411, 412, 413 uttered between the first first voice 410 and the second first voice 420 in Figure 4 as second voices matched with the first first voice 410. In addition, the electronic device 100 can recognize a plurality of second voices 421, 422 uttered after the second first voice 420 in Figure 4 as second voices matched with the second first voice 420. Thereafter, as in the operation in Figure 3 , the electronic device 100 can identify whether the plurality of second voices 411, 412, 413 are voices having one category information and being a combination of one response content, or are separate voices having respective category information.

[0094] According to another embodiment of the disclosure, the electronic device 100 can identify the point of utterance of each voice and whether there is a correlation between each voice, and acquire the first voice and the second voice matched with each other. For example, the electronic device 100 can recognize a plurality of second voices 411, 412, 413 uttered between the first first voice 410 and the second first voice 420 in Figure 4 as a candidate group of second voices matched with the first first voice 410.

[0095] The electronic device 100 can identify whether there is a correlation between the first utterance 410 and the second utterance candidate group 411, 412, 413. As an example, the electronic device 100 can extract a keyword included in the first utterance 410 or first text data corresponding to the first utterance, and extract a keyword included in the second utterance candidate group 411, 412, 413 or text data corresponding to the second utterance candidate group, and identify whether there is a correlation between the two keywords. For example, in the case where the keyword of the first utterance is "smell" and the keyword of at least one of the second utterance candidate group is "ventilation", the electronic device 100 can identify "smell" and "ventilation" as a relevant keyword, and determine that the first utterance and the second utterance match. The correlation between the keywords can be identified by an artificial intelligence model, or by an external server. The electronic device 100 can identify whether there is a correlation between the keywords based on information about the correlation between the words stored in the memory.

[0096] The electronic device 100 can identify a plurality of second utterances 421, 422 uttered after the second first utterance 420 as a candidate group of second utterances matching the second first utterance 420. If it is identified that there is no correlation between the second first utterance 420 and the second utterance candidate group 421, 422, the electronic device 100 can exclude the second first utterance 420 and the second utterance candidate group 421, 422 from the learning data used to train the first neural network model.

[0097] If it is identified that the second first utterance 420 is not in the form of a question, the electronic device 100 can identify a plurality of second utterances 411, 412, 413, 421, 422 uttered after the first first utterance 410 as a candidate group matching the first first utterance 410.

[0098] Figure 5 FIG. 1 is a block diagram illustrating an example configuration of an example electronic device according to an embodiment of the disclosure.

[0099] The electronic device 100 can include a memory 110 and a processor (e.g., including processing circuitry) 120.

[0100] The memory 110 can be electrically connected with the processor 120, and store data required for various embodiments of the disclosure.

[0101] The memory 110 can be implemented in the form of a memory embedded in the electronic device 100, or in the form of a memory that can be attached to or detached from the electronic device 100, according to the purpose of the stored data. For example, in the case of data for operating the electronic device 100, the data can be stored in a memory embedded in the electronic device 100, and in the case of data for expanding the function of the electronic device 100, the data can be stored in a memory that can be attached to or detached from the electronic device 100. Meanwhile, in the case of a memory embedded in the electronic device 100, the memory can be implemented as at least one of a volatile memory (for example, a dynamic RAM (DRAM), a static RAM (SRAM), or a synchronous dynamic RAM (SDRAM), etc.) or a non-volatile memory (for example, a one time programmable ROM (OTPROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a mask ROM, a flash ROM, a flash memory (for example, a NAND flash or a NOR flash, etc.), a hard drive, or a solid state drive (SSD)). In the case of a memory that can be attached to or detached from the electronic device 100, the memory can be implemented in the form of a memory card (for example, a compact flash (CF), a secure digital (SD), a micro secure digital (Micro-SD), a mini secure digital (Mini-SD), an extreme digital (xD), a multimedia card (MMC), etc.) and an external memory (for example, a USB memory) that can be connected to a USB port, etc.

[0102] According to an embodiment of the disclosure, the memory 110 can store a first neural network model, a second neural network model, and a voice recognition model.

[0103] The first neural network model can be a model outputting at least one category information and a probability value of the first speech or the first text data corresponding to the first speech, and the second neural network model can be a model outputting at least one category information and a probability value of the second speech or the second text data corresponding to the second speech. The first neural network model and the second neural network model can output the category information as unknown. Here, the category information is "unknown". The speech recognition model can be a model recognizing a user's speech and outputting the speech as text, and can be implemented as, for example, an automatic speech recognition (ASR) model. For example, the speech recognition model can output the first speech as the first text data corresponding to the first speech, and output the second speech as the second text data corresponding to the second speech. However, the disclosure is not limited thereto, and in a case where the speech recognition model is not included in the memory 110, it is possible to perform speech recognition of the first speech and the second speech at an external device, an external server, or the like, and the electronic device 100 receives the first text data and the second text data corresponding to the first speech and the second speech from the outside.

[0104] The memory 110 can store speech profile information related to the second speech. Here, the speech profile information can include waveform information about the speech of the counselor, speech identification information including the name of the counselor, or the like.

[0105] In addition, the memory 110 can include sample information of response contents corresponding to each category information. For example, in a case where the category information of the first speech is identified by the trained first neural network model, the response content stored in the memory 110 can be provided through the speaker of the terminal device according to the control of the processor 120.

[0106] The processor 120 can include various processing circuits, and is electrically connected with the memory 110, and controls the overall operation of the electronic device 100. The processor 120 controls the overall operation of the electronic device 100 using various instructions or programs stored in the memory 110. For example, according to an embodiment of the disclosure, the main CPU can copy the program in the RAM according to the instruction stored in the ROM, and access the RAM and execute the program. Here, the program can include an artificial intelligence model or the like.

[0107] The processor 120 can load the first neural network model, the second neural network model, the speech recognition model, etc. stored in the memory 110 into the processor 120. For example, to train the first neural network model, the processor 120 can load the first neural network model stored in the memory 110 external to the processor 120 into a memory (not shown) internal to the processor 120. Also, the processor 120 can load the speech recognition model stored in the memory 110 into the memory internal to the processor 120, and access the loaded speech recognition model and perform speech recognition.

[0108] Figure 6 FIG. 2 is a diagram illustrating an example process of training a second neural network model according to an embodiment of the disclosure.

[0109] The second neural network model can be a model that outputs at least one of category information and a probability value of second text data. In other words, the second neural network model can be a model that outputs category information of text data corresponding to the counselor's speech (second speech), and thus can train the second neural network model based on data associated with the counselor's speech.

[0110] For example, the electronic device 100 can train the second neural network model based on short text, long text, etc. included in the second text data corresponding to the second speech.

[0111] For example, it is assumed that the electronic device 100 trains the second neural network model with respect to category information of "air conditioner smell problem (category ID: 1500)".

[0112] As an example, the electronic device 100 can use the short text "used after ventilation" included in the second text data as input data, and train the second neural network model with output data of "air conditioner smell problem (category ID: 1500)" consistent therewith. The electronic device 100 can use the long text "Air circulation method in which the air conditioner absorbs indoor air and exchanges the air with cold air, and then discharges the air. Then, what about the discharged air?" included in the second text data as input data, and train the second neural network model with output data of "air conditioner smell problem (category ID: 1500)" consistent therewith. Thus, in the case where the input second text data includes "used after ventilation" or "Air circulation method in which the air conditioner absorbs indoor air and exchanges the air with cold air, and then discharges the air. Then, what about the discharged air?", the second neural network model can be trained to output "air conditioner smell problem (category ID: 1500)" as category information corresponding thereto.

[0113] As another example, the electronic device 100 can train the second neural network model using various forms of data provided by the counselor as response information in addition to the second text data corresponding to the second voice. For example, the electronic device 100 can train the second neural network model using an image about a method of replacing an air conditioner filter as input data and using output data of "air conditioner smell problem (category ID: 1500)" consistent therewith. Accordingly, in the case where the input data is an image about a method of replacing an air conditioner filter, the second neural network model can output "air conditioner smell problem (category ID: 1500)" as category information corresponding thereto.

[0114] As still another example, in the case where the counselor provides a moving image about a method of handling an air conditioner smell problem based on the inquiry of the client, the electronic device 100 can train the second neural network model using video image data. For example, the electronic device 100 can extract a subtitle (text) from the moving image. For example, in the case where the moving image includes text such as "remember, if you end the driving of the air conditioner, then use it after ventilating the room, you can reduce the occurrence of smell", the electronic device 100 can extract such text and use the text as input data of the second neural network model. Accordingly, if text such as "remember, if you end the driving of the air conditioner, then use it after ventilating the room, you can reduce the occurrence of smell" or a moving image including such text is input as input data, the second neural network model can output "air conditioner smell problem (category ID: 1500)" as category information corresponding thereto.

[0115] In other words, even in the case where data input into the second neural network model is not an image or a moving image in the form of text data, the second neural network model can output category information corresponding thereto.

[0116] Accordingly, in the case where text such as "use after ventilation" and "air circulation method in which the air conditioner sucks indoor air and exchanges the air with cold air and then discharges the air. Then, the discharged air?", an image about a method of replacing an air conditioner filter, or a moving image including text such as "remember, if you end the driving of the air conditioner, then use it after ventilating the room, you can reduce the occurrence of smell" is input into the second neural network model, the second neural network model can be trained to output "air conditioner smell problem (category ID: 1500)" as category information corresponding thereto.

[0117] The second neural network model can be trained based on at least one of predetermined manual data or Frequently Asked Questions (FAQ) data about the response, in addition to the data related to the voice of the counselor. Since the predetermined manual data or the FAQ data is data whose category information has been classified, a separate step of acquiring the category information is not required. Accordingly, the electronic device 100 can train the second neural network model using the response manual data as input data and using the category information consistent therewith as output data. As described above, the category information of the manual data or the FAQ data has been acquired, and thus the second neural network model can be trained without a separate process of matching the response content and the category information.

[0118] Figure 7 FIG. 1 is a flowchart illustrating an example method of controlling an electronic device according to an embodiment of the disclosure.

[0119] At operation S710, the electronic device 100 can input first data corresponding to a first voice included in a conversation content included in recording data into a first neural network model, and acquire category information of the first data as a result of inputting the first data into the first neural network model. Here, the recording data is data including counseling content between a client and a counselor, and the first voice can be a voice of the client, and a second voice to be described below can be a voice of the counselor.

[0120] The electronic device 100 can identify the second voice in the recording data based on voice profile information related to pre-stored data corresponding to the second voice.

[0121] At operation S720, the electronic device 100 can acquire category information of second data corresponding to the second voice in the conversation content.

[0122] The electronic device 100 can input data corresponding to the second voice into a second neural network model, and acquire category information of the data corresponding to the second voice.

[0123] For example, the electronic device 100 can input a sentence acquired by combining at least some of each of a plurality of sentences included in the data corresponding to the second voice into the second neural network model, and acquire category information of the data corresponding to the second voice.

[0124] If at least one of text data or image data related to the second voice is acquired from at least one of moving image data, image data, or text data, the electronic device 100 can input the acquired data into the second neural network model, and acquire category information of the data corresponding to the second voice.

[0125] In operation S730, if the category information of the first data and the category information of the second data are different, the electronic device 100 can train the first neural network model based on the category information of the second data and the first data.

[0126] The electronic device 100 can sequentially input text data corresponding to each of a plurality of sentences included in data corresponding to the second voice into the second neural network model, acquire at least one category information corresponding to each of the plurality of sentences and a first probability value corresponding to the at least one category information from the second neural network model, input text data corresponding to a sentence acquired by combining at least some of each of the plurality of sentences into the second neural network model, acquire at least one category information corresponding to the acquired sentence and a second probability value corresponding to the at least one category information from the second neural network model, and train the first neural network model based on a sentence selected according to the first probability value and the second probability value and category information corresponding to the selected sentence.

[0127] As an example, the electronic device 100 can select a sentence having a probability value greater than or equal to a threshold value between the first probability value and the second probability value, and train the first neural network model based on the selected sentence and category information corresponding to the selected sentence.

[0128] In the case where a conjunction is included in the sentence acquired by combining at least some of each of the plurality of sentences, the electronic device 100 can apply a weight value to a probability value corresponding to the acquired sentence based on at least one of a type or a number of the conjunction.

[0129] After the first neural network model is trained according to the above-described steps, if user voice data is input, the electronic device 100 can input text data corresponding to the user voice data into the trained first neural network model, and acquire category information corresponding to the user voice data, and acquire response information corresponding to the user voice data based on the category information.

[0130] The method according to the aforementioned various example embodiments of the disclosure can be implemented in the form of an application that can be installed on a conventional electronic device.

[0131] The method according to the aforementioned various example embodiments of the disclosure can be implemented through software upgrade or hardware upgrade of a conventional electronic device.

[0132] The aforementioned various example embodiments of the disclosure can be executed through an embedded server provided on an electronic device or at least one external server of the electronic device.

[0133] According to an embodiment of the disclosure, the above-described various embodiments can be implemented as software including instructions stored in a machine-readable storage medium, which can be read by a machine (e.g., a computer). The machine can refer to, for example, a device that invokes instructions stored in a storage medium and can operate according to the invoked instructions, and the device can include an electronic device according to the aforementioned embodiments. In the case where the instructions are executed by a processor, the processor can perform functions corresponding to the instructions by itself or using other components under its control. The instructions can include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium can be provided in the form of a non-transitory storage medium. The "non-transitory" storage medium does not include a signal and is tangible, but does not indicate whether data is stored in the storage medium semi-permanently or temporarily. For example, the "non-transitory storage medium" can include a buffer that temporarily stores data.

[0134] According to an embodiment of the disclosure, the method according to the aforementioned various embodiments can be provided when the method according to the aforementioned various embodiments is included in a computer program product. The computer program product refers to a product that can be traded between a seller and a buyer. The computer program product can be distributed online through an application store (e.g., Play Store TM) or in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)). In the case of online distribution, at least part of the computer program product can be temporarily stored in a storage medium such as a memory of a manufacturer's server, an application store's server, and a relay server, or can be temporarily generated.

[0135] In addition, according to an embodiment of the disclosure, the above-described various embodiments can be implemented in a recording medium that can be read by a computer or a device similar to a computer using software, hardware, or a combination thereof. In some cases, the embodiments described in the specification can be implemented by a processor itself. According to the software implementation, the embodiments such as processes and functions described in the disclosure can be implemented by separate software modules. Each of the software modules can perform one or more functions and operations described in the specification.

[0136] Computer instructions for performing processing operations of a machine according to the aforementioned various embodiments can be stored in a non-transitory computer-readable medium. When the computer instructions stored in such a non-transitory computer-readable medium are executed by a processor of a specific machine, the instructions cause processing operations at the machine according to the aforementioned various embodiments to be performed by the specific machine.

[0137] The non-transitory computer-readable medium can include a CD, a DVD, a hard disk, a Blu-ray disc, a USB, a memory card, a ROM, etc.

[0138] Each of the components (for example, modules or programs) according to the aforementioned various embodiments can include a single object or a plurality of objects. Also, some of the aforementioned corresponding sub-components can be omitted, or other sub-components can be further included in the various embodiments. Generally or additionally, some of the components (for example, modules or programs) can be integrated as one object, and perform functions performed by each component prior to the integration in the same or similar manner. Operations performed by the modules, programs, or other components according to the various embodiments can be executed sequentially, in parallel, repeatedly, or heuristically. Or, at least some of the operations can be executed in different orders or omitted, or other operations can be added.

[0139] While various example embodiments of the present disclosure have been illustrated and described, it will be understood by those of ordinary skill in the art that various example embodiments are intended to be illustrative only and not restrictive, and that modifications can be made by those skilled in the art without departing from the true spirit and full scope of the present disclosure, including the appended claims.

Claims

1. An electronic device, comprising: A memory that stores recorded data including the content of the conversation and at least one instruction; as well as The processor is configured to execute at least one of the instructions as follows: First data corresponding to the first voice in the dialogue content is input into a first neural network model, and category information of the first data is obtained, wherein the first voice is a question, and Obtain category information of second data corresponding to the second voice in the dialogue content, wherein the second voice is emitted after the first voice and forms a query-response pair with the first voice, and The first neural network model is trained based on the difference between the category information of the first data and the category information of the second data, and the category information of the second data and the first data. The category information of the second data is information with higher accuracy than the category information of the first data.

2. The electronic device according to claim 1, in, The memory stores voice profile information related to the data corresponding to the second voice, and The processor is configured to: The second voice in the recorded data is identified based on the stored voice profile information.

3. The electronic device according to claim 1, in, The processor is configured to: The data corresponding to the second speech is input into the second neural network model, and the category information of the data corresponding to the second speech is obtained.

4. The electronic device according to claim 3, in, The processor is configured to: A sentence obtained by combining at least a portion of each of a plurality of sentences included in the data corresponding to the second speech is input into a second neural network model, and category information of the data corresponding to the second speech is obtained.

5. The electronic device according to claim 4, in, The processor is configured to: Text data corresponding to each of a plurality of sentences included in the data corresponding to the second speech is sequentially input into a second neural network model, and at least one category information corresponding to each of the plurality of sentences and a first probability value corresponding to the at least one category information are obtained from the second neural network model. Text data corresponding to sentences obtained by combining at least a portion of each of a plurality of sentences is input into a second neural network model, and at least one category information corresponding to the obtained sentences and a second probability value corresponding to the at least one category information are obtained from the second neural network model. The first neural network model is trained based on sentences selected according to a first probability value and a second probability value, as well as first category information corresponding to the selected sentences.

6. The electronic device according to claim 5, in, The processor is configured to: A sentence with a probability value greater than or equal to a threshold is selected, and a first neural network model is trained based on the selected sentence and the first category information corresponding to the selected sentence, wherein the threshold is between a first probability value and a second probability value.

7. The electronic device according to claim 5, in, The processor is configured to: Based on sentences obtained by combining at least a portion of each of multiple sentences, including conjunctions, weight values ​​are applied to probability values ​​corresponding to the obtained sentences based on at least one of the types or frequencies of the conjunctions.

8. The electronic device according to claim 3, in, The processor is configured to: Based on obtaining at least one of the motion image data, text data, or image data related to the second speech from at least one of motion image data, image data, or text data, the obtained data is input into a second neural network model, and category information of the data corresponding to the second speech is obtained.

9. The electronic device according to claim 1, in, The processor is configured to: Based on the input user voice data, text data corresponding to the user voice data is input into a trained first neural network model, and category information corresponding to the user voice data is obtained. Response information corresponding to the user voice data is then obtained based on the category information.

10. The electronic device according to claim 1, in, The recorded data includes data containing consultation content between clients and consultants, and The first voice is the customer's voice, and The second voice is the consultant's voice.

11. A method for controlling an electronic device storing recorded data including conversation content, the method comprising: The first data corresponding to the first voice in the dialogue content is input into the first neural network model, and the category information of the first data is obtained, wherein the first voice is a question; Obtain category information of second data corresponding to the second speech in the dialogue content, wherein the second speech is emitted after the first speech and forms a query-response pair with the first speech; and Based on the difference between the category information of the first data and the category information of the second data, a first neural network model is trained based on the category information of the second data and the first data. The category information of the second data is information with higher accuracy than the category information of the first data.

12. The method of claim 11, further comprising: The second speech in the recorded data is identified based on speech profile information associated with pre-stored data corresponding to the second speech.

13. The method according to claim 11, in, The category information for obtaining the second data includes: The data corresponding to the second speech is input into the second neural network model, and the category information of the data corresponding to the second speech is obtained.

14. The method according to claim 13, in, The category information for obtaining the second data includes: A sentence obtained by combining at least a portion of each of a plurality of sentences included in the data corresponding to the second speech is input into a second neural network model, and category information of the data corresponding to the second speech is obtained.

15. The method according to claim 14, in, Training the first neural network model includes: Text data corresponding to each of a plurality of sentences included in the data corresponding to the second speech is sequentially input into a second neural network model, and at least one category information corresponding to each of the plurality of sentences and a first probability value corresponding to the at least one category information are obtained from the second neural network model; Text data corresponding to sentences obtained by combining at least a portion of each of a plurality of sentences is input into a second neural network model, and at least one category information corresponding to the obtained sentences and a second probability value corresponding to the at least one category information are obtained from the second neural network model; and The first neural network model is trained based on sentences selected according to a first probability value and a second probability value, as well as first category information corresponding to the selected sentences.

Citation Information

Patent Citations

  • Voice recognition apparatus and method

    US20180061394A1

  • Method and apparatus for classifying class, to which sentence belongs, using deep neural network

    WO2018212584A2