Systems and methods for modifying speech recognition results

By interoperating between devices and servers and using text modification models to adjust speech recognition results, the problem of insufficient contextual adaptability of existing systems in different fields is solved, and more accurate speech recognition and text output are achieved.

CN112397063BActive Publication Date: 2025-10-10SAMSUNG ELECTRONICS CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010814764.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-09
Filing Date
2020-08-13
Publication Date
2025-10-10
Estimated Expiration
2040-10-29

AI Technical Summary

Technical Problem

Existing automatic speech recognition systems have deficiencies in recognition accuracy and adaptability to contextual understanding in different fields, resulting in inaccurate output results.

Method used

Through interoperability between the device and the server, the server's text modification model is used to modify the device's automatic speech recognition model output, and multiple domain-related text modification models are used to identify and adjust the output text to ensure that it matches the context of a specific domain.

Benefits of technology

The accuracy and adaptability of speech recognition results have been improved, enabling better understanding and processing of speech input in different fields and providing more precise text output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112397063B_ABST
    Figure CN112397063B_ABST
Patent Text Reader

Abstract

A system and method for modifying speech recognition results are provided. The method includes receiving, from a device, text output from an automatic speech recognition (ASR) model of the device; identifying at least one domain related to the received text; selecting, from a plurality of text modification models included in a server, at least one text modification model corresponding to the identified at least one domain; and modifying the received text by using the selected at least one text modification model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is based upon and claims the benefit of U.S. Provisional Patent Application No. 62 / 886,027 filed in the U.S. Patent and Trademark Office on August 13, 2019, and claims priority to Korean Patent Application No. 10-2019-0162921 filed in the Korean Intellectual Property Office on December 9, 2019, the disclosures of which are incorporated herein by reference in their entirety. Technical Field

[0003] The present disclosure relates to a system and method for modifying speech recognition results, and more particularly, to a system and method for modifying speech recognition results through interoperation of a device and a server. Background Art

[0004] Automatic speech recognition (ASR) is a technology used to receive human speech and convert it into text. Speech recognition is used in various electronic devices, such as smartphones, air conditioners, refrigerators, and artificial intelligence (AI) home assistants. For example, a device detects human speech as input, recognizes the received speech using a speech recognition model trained to recognize speech, and converts the recognized speech into text. Text can be the final output of the device.

[0005] In recent years, deep neural network (DNN) algorithms have been used in various machine learning fields, and the performance of speech recognition has been improved. For example, in the field of speech recognition, the performance has been greatly improved by using neural networks, and speech recognition models (e.g., ASR models) for speech recognition have been developed. As AI systems have improved, recognition rates have increased and user preferences have been understood more accurately, so existing rule-based intelligent systems have gradually been replaced by AI systems based on deep learning. Summary of the Invention

[0006] Provided are a system and method for providing an output value of an automatic speech recognition (ASR) model of a device to a server and for modifying the output value of the ASR model by using an artificial intelligence (AI) model of the server.

[0007] Provided are a system and method for modifying a speech recognition result by using a text modification model corresponding to a domain related to an output value of an ASR model of a device.

[0008] A system and method are provided for efficiently applying text received by a server from a device to text modification models associated with multiple domains.

[0009] A system and method are provided by which a server can efficiently identify domains associated with a text by using a plurality of domain identification modules associated with the plurality of domains.

[0010] Additional aspects will be set forth in part in the description which follows and, in part, will be apparent from the description, or may be learned by practice of the presented embodiments of the disclosure.

[0011] According to an embodiment of the present disclosure, a method for modifying a speech recognition result provided from a device, performed by a server, includes: receiving output text from an automatic speech recognition (ASR) model of the device from the device; identifying at least one domain related to the subject of the output text; selecting at least one text modification model of at least one domain from a plurality of text modification models included in the server, wherein the at least one text modification model is an artificial intelligence (AI) model trained to analyze text related to the subject; and using the at least one text modification model to modify the output text to generate modified text.

[0012] According to another embodiment of the present disclosure, a server for modifying a speech recognition result provided from a device includes: a communication interface; a storage device storing a program including one or more instructions; a processor configured to run one or more instructions of the program stored in the storage device to receive output text from an automatic speech recognition (ASR) model of the device from the device, identify at least one domain related to the subject of the output text, select at least one text modification model of at least one domain from a plurality of text modification models included in the server, wherein the at least one text modification model is an artificial intelligence (AI) model trained to analyze text related to the subject, use the at least one text modification model to modify the output text to generate modified text, and provide the modified text to the device. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The above and other aspects, features and advantages of certain embodiments of the present disclosure will become more apparent through the following description in conjunction with the accompanying drawings, in which:

[0014] Figure 1 is a block diagram illustrating a speech recognition system according to an embodiment of the present disclosure;

[0015] Figure 2 is a block diagram illustrating a speech recognition system including text modification models associated with multiple domains according to an embodiment of the present disclosure;

[0016] Figure 3is a flowchart illustrating a method for a device and a server in a speech recognition system to recognize a voice input and obtain a modified text according to an embodiment of the present disclosure;

[0017] Figure 4 is a diagram illustrating a server for identifying a domain related to a text and selecting a text modification model for the domain related to the text according to an embodiment of the present disclosure;

[0018] Figure 5 is a diagram illustrating an apparatus for identifying a domain associated with a text and a server for selecting a text modification model for the domain associated with the text according to an embodiment of the present disclosure;

[0019] Figure 6 is a diagram illustrating a server and a device for identifying a domain associated with a text and a server for selecting a text modification model for the domain associated with the text according to an embodiment of the present disclosure;

[0020] Figure 7 is a flowchart illustrating a method of selecting a domain related to a text by using domain reliability obtained by a device and domain reliability obtained by a server, performed by a server according to an embodiment of the present disclosure;

[0021] Figure 8 is a diagram illustrating a server selecting a text modification model by using a domain identification module selected from a plurality of domain identification modules in the server according to an embodiment of the present disclosure;

[0022] Figure 9 is a flowchart illustrating a method in which a server selects a domain for text modification by using a domain identification module selected from a plurality of domain identification modules according to an embodiment of the present disclosure;

[0023] Figure 10 is a diagram illustrating a first domain identification module, a second domain identification module, and a text modification model related to hierarchically classified domains according to an embodiment of the present disclosure;

[0024] Figure 11 is a diagram illustrating a server that modifies text by using a plurality of text modification models according to an embodiment of the present disclosure;

[0025] Figure 12 is a flowchart illustrating a method in which a server accumulates and calculates domain reliability of texts of a plurality of sections according to an embodiment of the present disclosure;

[0026] Figure 13 is a diagram illustrating a server for obtaining domain reliability of a text flow accumulated in units of grammatical words according to an embodiment of the present disclosure;

[0027] Figure 14is a flowchart illustrating a method in which a server divides a text into a plurality of sections and selects a domain of the text of each of the plurality of sections according to an embodiment of the present disclosure;

[0028] Figure 15 is a diagram illustrating an example of a server of a text modification model that compares domain reliability of text according to a plurality of domains and selects and modifies text of each section according to an embodiment of the present disclosure;

[0029] Figure 16 is a diagram illustrating a server that modifies text received from a device by using modified text output from a plurality of text modification models according to an embodiment of the present disclosure;

[0030] Figure 17 is a block diagram illustrating a server according to an embodiment of the present disclosure; and

[0031] Figure 18 is a block diagram illustrating a device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0032] Hereinafter, embodiments of the present disclosure will be described in detail so that those skilled in the art can easily embody and practice the present disclosure with reference to the accompanying drawings. However, the present disclosure can be implemented in many different forms and should not be construed as limited to the embodiments of the present disclosure set forth herein. In the accompanying drawings, for the sake of clarity of the present disclosure, parts that are not necessary for the description are omitted in the drawings, and the same reference numerals represent the same elements throughout.

[0033] Throughout the specification, it will be understood that when a component is referred to as being “connected” to another component, it may be “directly connected” to the other component or “electrically connected” to the other component through an intervening element. It will also be understood that when a component “includes” or “comprising” an element, unless otherwise defined, the component may also include other elements without excluding other elements.

[0034] Throughout this disclosure, the expression "at least one of a, b, or c" means only a, only b, only c, both a and b, both a and c, both b and c, all three of a, b, and c, or variations thereof.

[0035] The present disclosure will now be described more fully with reference to the accompanying drawings, in which embodiments of the disclosure are shown.

[0036] Figure 1 is a block diagram illustrating a speech recognition system according to an embodiment of the present disclosure.

[0037] refer to Figure 1 , a speech recognition system according to an embodiment of the present disclosure includes a device 1000 and a server 2000 .

[0038] The device 1000 can include an automatic speech recognition (ASR) model, and the server 2000 can include a text modification model. The device 1000 can recognize a user's voice input by using the ASR model and can output text. The server 2000 can modify the text generated by the device 1000.

[0039] The ASR model is a speech recognition model for recognizing a voice by using an integrated neural network, and can output text from a user's voice input. The ASR model can be an artificial intelligence (AI) model including, for example, an acoustic model, a pronunciation dictionary, and a language model. Alternatively, the ASR model can be an end-to-end speech recognition model having a structure including an integrated neural network without separately including, for example, an acoustic model, a pronunciation dictionary, and a language model.

[0040] Because the end-to-end ASR model uses an integrated neural network, the end-to-end ASR model can convert a voice into text without a process of recognizing phonemes from the voice and then converting the phonemes into text. The text can include at least one character. A character refers to a symbol used to describe a human language in a visible form. Examples of characters can include Korean letters (Korean), an alphabet, Chinese characters, numbers, phonetic symbols, punctuation, and other symbols.

[0041] In addition, for example, the text can include a string. A string refers to a sequence of characters. For example, the text can include at least one grapheme. A grapheme is the smallest unit including at least one character and representing a sound. For example, in an alphabetic writing system, one character can become a grapheme, and a string can refer to a sequence of graphemes. For example, the text can include a morpheme or a word. A morpheme is the smallest meaningful unit including at least one grapheme. A word is an independent unit of language having a grammatical function and including at least one morpheme grammatical function.

[0042] The device 1000 can receive a user's voice input, can recognize the received voice input by using the ASR model, and can provide text that is an output value of the ASR model to the server 2000. In addition, the server 2000 can receive text that is an output value of the ASR model from the device 1000, and can modify the received text. The server 2000 can identify a degree to which the text that is an output value of the ASR model relates to a specific domain registered in the server 2000. The server can modify the text by using a text modification model of the identified domain. In addition, the server 2000 can provide the modified text to the device 1000.

[0043] The text modification model, which is an AI model trained to modify at least a portion of text that is a speech recognition result, can include, for example, a sequence-to-sequence mapper. The text modification model can be an AI model trained by using text output from the ASR model and preset ground truth text. The text modification model can be an AI model trained for each domain. For example, a text modification model for a first domain can be trained by using text that is an output value of the ASR model and ground truth text specific to the first domain. Also, for example, a text modification model for a second domain can be trained by using text that is an output value of the ASR model and ground truth text specific to the second domain.

[0044] The text modification model can be an AI model trained by using a plurality of pieces of text output from a plurality of types of ASR models and a plurality of pieces of preset ground truth text. In this case, because the text modification model is trained by using a plurality of pieces of text output from a plurality of types of ASR models, accurate output values (e.g., modified text) can be provided regardless of which type of ASR model the text input to the text modification model is output from.

[0045] For example, the modified text can include at least one of a modified character, a modified grapheme, a modified morpheme, or a modified word. For example, when the text output from the ASR model includes an error, the modified text can include a character corrected from the error. Also, for example, when the text output from the ASR model includes a word having an incorrect meaning that is not suitable for a context, the modified text can include a word having a correct meaning that replaces the word having the incorrect meaning. Furthermore, for example, the modified text can be generated by replacing a specific word in the text output from the ASR model with a similar word.

[0046] Examples of the device 1000 can include, but are not limited to, a smart phone, a tablet personal computer (PC), a PC, a smart TV, a mobile phone, a personal digital assistant (PDA), a laptop computer, a media player, a micro server, a global positioning system (GPS) device, an electronic book terminal, a digital broadcasting terminal, a navigation, a kiosk, an MP3 player, a digital camera, a home appliance, and other mobile or non-mobile computing devices. Furthermore, the device 1000 can be a wearable device having a communication function and a data processing function, such as a watch, glasses, a headband, or a ring. However, the present disclosure is not limited thereto, and the device 1000 can include any type of device capable of transmitting and receiving data to and from the server 2000 through the network 200 to perform speech recognition.

[0047] Examples of the network 200 include a local area network (LAN), a wide area network (WAN), a value added network (VAN), a mobile radio communication network, a satellite communication network, and a combination thereof. The network 200 is a data communication network in an integrated sense for enabling the network constituent elements to communicate smoothly with each other, and includes a wired Internet, a wireless Internet, and a mobile wireless communication network. Figure 1 Examples of the network 200 include a local area network (LAN), a wide area network (WAN), a value added network (VAN), a mobile radio communication network, a satellite communication network, and a combination thereof. The network 200 is a data communication network in an integrated sense for enabling the network constituent elements to communicate smoothly with each other, and includes a wired Internet, a wireless Internet, and a mobile wireless communication network.

[0048] Figure 2 is a block diagram illustrating a speech recognition system including a text modification model related to a plurality of domains according to an embodiment of the disclosure.

[0049] Referring to Figure 2 , the device 1000 can include an ASR model. The server 2000 can include a text modification model for modifying text output from the ASR model. For example, the server 2000 can include a first text modification model corresponding to a first domain and a second text modification model corresponding to a second domain.

[0050] The device 1000 can obtain a feature vector by extracting a feature from a voice input, and can provide the obtained feature vector as an input to the ASR model. The device 1000 can provide an output value output from the ASR model to the server 2000. The output value output from the ASR model can include various forms of text. For example, the device 1000 can provide a sentence-unit text to the server 2000, or can provide a text stream to the server 2000, but the system is not limited thereto.

[0051] The server 2000 can receive text from the device 1000, and can select a (text) domain related to the received text. The domain indicates a field related to a topic of an input voice, and can be preset according to, for example, a meaning of the input voice or a property of the input voice. When the topic of the input voice is a service, the domain can be classified according to, for example, a service related to the input voice. In addition, a text modification model can be trained for each domain, and in this case, the text modification model trained for each domain can be a model trained by using input text related to the domain and real text corresponding to the input text. The server 2000 can select at least one of a plurality of preset domains, and can select at least one of text modification models corresponding to the selected domain. In addition, the server 2000 can obtain modified text by inputting the text received from the device 1000 to the text modification model of the selected domain. The server 2000 can provide the modified text to the device 1000.

[0052] The server 2000 can provide various types of voice assistant services to the device 1000 by using modified text. The voice assistant service can be a service that provides a dialogue with a user who provides commands or questions to the voice assistant service. In the voice assistant service, a response message can be provided to the user, just like a person talking directly to the user taking into account the user's situation, the condition of the device, etc. In addition, in the voice assistant service, the information required by the user can be appropriately generated and provided to the user, just like the user's personal assistant provides the information. The voice assistant service can provide the user with the information or function requested by the user in combination with various services such as broadcast services, content sharing services, content provision services, power management services, game services, chat services, document creation services, search services, call services, photography services, traffic recommendation services, and video playback services.

[0053] Figure 3 The present invention is a flowchart illustrating a method of a device and a server in a speech recognition system for recognizing a voice input and obtaining a modified text according to an embodiment of the present disclosure.

[0054] The device 1000 may execute the instructions stored in the memory of the device 1000. Figure 3 For example, the device 1000 may execute the following operations: Figure 18 The speech recognition evaluation module 1430, the domain identification module 1440, the natural language understanding (NLU) determination module 1450, the domain registration module 1460, the ASR model 1410 or the NLU model 1420 are executed by at least one of Figure 3 However, the present disclosure is not limited thereto, and the device 1000 may execute other programs stored in the memory to perform other operations associated with the programs stored in the memory.

[0055] In addition, the server 2000 can execute the instructions stored in the memory of the server 2000. Figure 3 For example, the server 2000 can be operated by running the following Figure 17 The server 2000 may execute an operation by executing at least one of the domain management module 2310, the speech analysis management module 2340, the text modification module 2320, or the NLU module 2330. However, the present disclosure is not limited thereto, and the server 2000 may execute other programs stored in the memory to execute a certain operation of the server 2000.

[0056] In operation S300, the device 1000 may obtain a feature vector from a voice signal. The device 1000 may receive a user's voice input (e.g., speech) through a microphone and may generate a feature vector indicating the characteristics of the voice signal using the voice signal obtained through the microphone. When the voice signal includes noise, the device 1000 may remove the noise from the voice signal and obtain a feature vector from the noise-removed voice signal. In addition, for example, the device 1000 may extract a feature vector indicating the characteristics of the voice signal from the voice signal. For example, the device 1000 may receive data indicating the feature vector of the voice signal from an external device.

[0057] In operation S305, the device 1000 may obtain text from the feature vector by using an ASR model. The device 1000 may provide the feature vector as input to the ASR model in the device 1000 to recognize the user's voice. When the device 1000 includes multiple ASR models, the device 1000 may select one of the multiple ASR models and may convert the feature vector into a format suitable for the selected ASR model. The ASR model of the device 1000 may be an AI model including, for example, an acoustic model, a pronunciation dictionary, and a language model. Alternatively, the ASR model of the device 1000 may be an end-to-end speech recognition model having a structure including an integrated neural network without having to separately include, for example, an acoustic model, a pronunciation dictionary, and a language model.

[0058] In operation S310, the device 1000 may obtain the reliability of the text. The reliability of the text may be a value indicating the degree to which the text output from the ASR model matches the input speech, and may include, for example, but not limited to, a confidence score. In addition, the reliability of the text may be related to the probability that the text will match the input speech. For example, the reliability of the text may be calculated based on at least one of the likelihood of a plurality of estimated texts output from the ASR model of the device 1000 or the posterior probability that at least one character in the text will be replaced by another character. For example, the device 1000 may calculate the reliability based on the likelihood output as a result of Viterbi decoding. Or, for example, the device 1000 may calculate the reliability based on the posterior probability output from the softmax layer in the end-to-end ASR model. For example, the device 1000 may determine a plurality of estimated texts estimated during the speech recognition process of the ASR model of the device 1000, and may calculate the reliability of the text based on the correlation of the characters in the plurality of estimated texts. In addition, for example, the device 1000 may calculate the reliability of the text based on the posterior probability output from the softmax layer in the end-to-end ASR model. Figure 18 The speech recognition evaluation module 1430 is used to obtain the reliability of the text.

[0059] In operation S315, the device 1000 may determine whether to send the text to the server 2000. The device 1000 may determine whether to send the text to the server 2000 by comparing the reliability of the text with a preset threshold. When the reliability of the text is equal to or greater than the preset threshold, the device 1000 may determine not to send the text to the server 2000. In addition, when the reliability of the text is less than the preset threshold, the device 1000 may determine to send the text to the server 2000.

[0060] In addition, the device 1000 can determine whether to send the text to the server 2000 based on at least one text with high reliability among the multiple estimated texts in the speech recognition process of the ASR model. For example, the multiple estimated texts estimated in the speech recognition process of the ASR model include a first estimated text with high reliability and a second estimated text with high reliability. If the difference between the reliability of the first estimated text and the reliability of the second estimated text is equal to or less than a certain threshold, the device 1000 can determine to send the text to the server 2000. In addition, for example, the multiple estimated texts estimated in the speech recognition process of the ASR model include a first estimated text with high reliability and a second estimated text with high reliability. If the difference between the reliability of the first estimated text and the reliability of the second estimated text is greater than a certain threshold, the device 1000 can determine not to send the text to the server 2000.

[0061] When it is determined to transmit the text to the server 2000 in operation S315, the device 1000 may request the server 2000 to modify the text in operation S320.

[0062] The device 1000 may transmit the text to the server 2000 and may request the modified text from the server 2000. In this case, for example, the device 1000 may transmit the type of the ASR model in the device 1000 and the identification value of the ASR model to the server 2000 while requesting the modified text from the server 2000, but the present disclosure is not limited thereto.

[0063] Additionally, for example, device 1000 may provide domain information associated with the text output from the ASR model of device 1000 to server 2000 while requesting the modified text from server 2000. Domain information used to identify a domain may include, but is not limited to, a domain name and a domain identifier. Device 1000 may identify the domain associated with the text using domain identification module 1440 in device 1000. For example, device 1000 may identify the domain associated with the text based on the domain reliability of the text output from the ASR model of device 1000. Domain reliability may be a value indicating the degree to which at least a portion of the text is associated with a specific domain. For example, device 1000 may calculate a confidence score indicating the degree to which the text output from the ASR model is associated with a domain pre-registered in device 1000. Furthermore, device 1000 may identify the domain associated with the text based on the calculated domain reliability. Device 1000 may identify the domain associated with the text based on rules, or may obtain the domain reliability associated with the text using an AI model trained for domain identification. Additionally, for example, the AI ​​model for domain identification can be part of the NLU model, or a separate model from the NLU model.

[0064] When it is determined in operation S315 that the text is not to be sent to the server 2000, the device 1000 can provide a voice assistant service by using the text output from the ASR model. For example, when the reliability of the text output from the ASR model is equal to or greater than a preset threshold, the device 1000 can perform an operation for the voice assistant service by using the text output from the ASR model. In addition, for example, the multiple estimated texts estimated during the speech recognition process of the ASR model include a first estimated text with the highest reliability and a second estimated text with the second highest reliability. If the difference between the reliability of the first estimated text and the reliability of the second estimated text is greater than a certain threshold, the device 1000 can provide a voice assistant service by using the first estimated text with the highest reliability.

[0065] For example, the device 1000 may display text output from the ASR model on the screen. For example, the device 1000 may perform an operation for communicating with the user based on the text output from the ASR model. In addition, for example, the device 1000 may provide various services such as broadcast services, content sharing services, content provision services, power management services, game services, chat services, document creation services, search services, call services, photography services, traffic recommendation services, and video playback services by communicating with the user based on the text output from the ASR model.

[0066] In operation S325, server 2000 may identify a domain for text modification. When server 2000 receives domain information from device 1000, server 2000 may identify the domain for text modification based on the domain information. Alternatively, when server 2000 does not receive domain information from device 1000, server 2000 may identify the domain associated with the text received from device 1000 by using domain identification module 2312 in server 2000. For example, in this case, server 2000 may identify the domain associated with the text based on the domain reliability of the text received from device 1000. For example, server 2000 may calculate a confidence score indicating the degree to which the text received from device 1000 is associated with a domain pre-registered for text modification. Additionally, server 2000 may identify the domain associated with the text received from device 1000 based on the domain reliability calculated for the pre-registered domain. Server 1000 may identify the domain associated with the text based on rules, or may obtain the domain reliability associated with the text by using an AI model trained for domain identification. Additionally, for example, the AI ​​model for domain identification can be part of the NLU model, or a separate model from the NLU model.

[0067] In operation S330, the server 2000 may modify the text by using a text modification model corresponding to the determined domain. The server 2000 may include a plurality of text modification models corresponding to a plurality of domains and may select a text modification model corresponding to the domain identified in operation S325 from the plurality of text modification models.

[0068] Server 2000 may select a domain corresponding to the domain identified in operation S325 from the domains registered in server 2000 and may select a text modification model for the selected domain. The domain corresponding to the domain identified in operation S325 among the domains registered in server 2000 may be a domain that is the same as or similar to the identified domain. For example, if the multiple domains registered in server 2000 are "movie," "location," and "region name," and the domain identified in operation S325 is "movie," server 2000 may select "movie." For example, if the multiple domains in server 2000 are "video content," "location," and "region name," and the domain identified in operation S325 is "movie," server 2000 may select "video content." In this case, information regarding identification values ​​similar to the identification values ​​of each domain of server 2000 may be stored in server 2000.

[0069] In addition, the server 2000 can generate a modified text by using the selected text modification model. The server 2000 can input the text into the selected text modification model and can obtain the modified text output from the text modification model. In this case, the server 2000 can pre-process the format of the text received from the device 1000 to make it suitable for the text modification model, and can input the processed value into the text modification model.

[0070] When the text received from the device 1000 is related to multiple domains, the server 2000 can select multiple text modification models corresponding to the multiple domains to perform text modification. In this case, the server 2000 can obtain the modified text to be provided to the device 1000 from the multiple modified texts output from the multiple text modification models. For example, when the server 2000 generates multiple modified texts by using multiple text modification models, the server 2000 can compare the reliability of the multiple modified texts and can determine the modified text with high or highest reliability as the modified text to be provided to the device 1000. The reliability of the modified text can be a value indicating the degree to which the modified text matches the input voice, and can include, for example, but not limited to, a confidence score.

[0071] In addition, for example, the server 2000 may extract several pieces of text from the multiple pieces of modified text output from the multiple text modification models and combine the extracted several pieces of text to obtain the modified text to be provided to the device 1000. For example, when the server 2000 generates the first modified text and the second modified text by using the multiple text modification models and the reliability of a part of the first modified text and the reliability of a part of the second modified text are high, the server 2000 may obtain the modified text to be provided to the device 1000 by combining the part of the first modified text and the part of the second modified text to generate a text having a higher reliability than the reliability of the first modified text or the second modified text.

[0072] In operation S335 , the server 2000 may provide the modified text to the device 1000 .

[0073] Despite Figure 3In the embodiment, device 1000 requests the modified text from server 2000 and server 2000 provides the modified text to device 1000, but the present disclosure is not limited to this. Server 2000 can provide various types of voice assistant services to device 1000 by using the modified text. A voice assistant service can be a service for providing a dialogue with a user. In the voice assistant service, a response message can be provided to the user by a voice assistant, just like a person talking directly to the user taking into account the user's situation and the status of the device. In addition, in the voice assistant service, the information required by the user can be appropriately generated and provided to the user, just like the user's personal assistant provides the information. The voice assistant service can provide the user with the information or function requested by the user in combination with various services such as broadcast services, content sharing services, content provision services, power management services, game services, chat services, document creation services, search services, call services, photography services, traffic recommendation services and video playback services.

[0074] In this case, the server 2000 can provide information for performing a dialogue with the user to the device 1000 by using a natural language understanding (NLU) model, a dialog manager (DM) model, a natural language generating (NLG), etc. in the server 2000. Therefore, the server 2000 can provide a voice assistant service based on the text. In addition, the server 2000 can directly control another device based on the result obtained by interpreting the text. In addition, the server 2000 can generate control information for enabling the device 1000 to control another device based on the result obtained by interpreting the modified text, and can provide the generated control information to the device 1000.

[0075] Figure 4 is a diagram illustrating a server that identifies a domain related to a text and selects a text modification model of the domain related to the text according to an embodiment of the present disclosure.

[0076] refer to Figure 4, the text output from the ASR model 40 of the device 1000 can be provided to the domain identification module 2312 in the server 2000. The server 2000 can use the domain identification module 2312 in the server 2000 to identify the domain related to the text received from the device 1000. In this case, the server 2000 can identify the domain related to the text based on the domain reliability of the text received from the device 1000. For example, the server 2000 calculates a confidence score that indicates the degree to which the text received from the device 1000 is related to the domain pre-registered for text modification. The domain identification module 2312, as an AI model trained for domain identification, can output domain reliability using the text as an input value. In addition, for example, the domain identification module 2312 can be part of the NLU model or a model separate from the NLU model. Alternatively, the domain identification module 2312 can identify the domain related to the text based on rules.

[0077] exist Figure 4 In the embodiment, the domain identification module 2312 may obtain, for example, a first domain reliability of a first domain, a second domain reliability of a second domain, and a third domain reliability of a third domain.

[0078] In addition, the model selection module 2313 of the server 2000 can select a text modification model for text modification. For example, the model selection module 2313 can compare the first domain reliability of the first domain, the second domain reliability of the second domain, and the third domain reliability of the third domain, and can determine that the first domain reliability of the first domain is the highest domain reliability. In addition, the model selection module 2313 can select the text modification model 41 for the first domain from the multiple text modification models 41, 42, and 43 in the server 2000.

[0079] The server 2000 may input the text received from the device 1000 to the text modification model 41 of the first domain, and may obtain the modified text output from the text modification model 41. Next, the server 2000 may provide the modified text to the device 1000.

[0080] Figure 5 is a diagram illustrating an apparatus for identifying a domain associated with a text and a server for selecting a text modification model for the domain associated with the text according to an embodiment of the present disclosure.

[0081] refer to Figure 5, the text output from the ASR model 50 of the device 1000 can be provided to the domain identification module 1440 in the device 1000. The device 1000 can use the domain identification module 1440 in the device 1000 to identify the domain related to the text output from the ASR model 50. In this case, the device 1000 can identify the domain related to the text based on the domain reliability of the text output from the ASR model 50. For example, the device 1000 can calculate a confidence score that indicates the degree to which the text output from the ASR model 500 is related to a pre-registered domain. The domain identification module 1440, as an AI model trained for domain identification, can output domain reliability by using the text as an input value. In addition, for example, the domain identification module 1440 can be part of the NLU model or a model separate from the NLU model. Alternatively, the domain identification module 1440 can identify the domain related to the text based on rules.

[0082] exist Figure 5 In the embodiment, the domain identification module 1440 may obtain, for example, a first domain reliability of the first domain, a second domain reliability of the second domain, and a third domain reliability of the third domain.

[0083] In addition, the device 1000 may provide the text output from the ASR model 50 to the server 2000. In addition, the device 1000 may provide the domain reliability obtained by the domain identification module 1440 to the server 2000. Alternatively, the device 1000 may identify a domain related to the text based on the domain reliability obtained by the domain identification module 1440, and may provide identification information of the identified domain to the server 2000.

[0084] The model selection module 2313 of the server 2000 can select a text modification model for text modification. When the device 1000 provides the domain reliability to the server 2000, the model selection module 2313 can compare, for example, a first domain reliability of the first domain, a second domain reliability of the second domain, and a third domain reliability of the third domain, and can determine that the domain reliability of the first domain is the highest domain reliability. In addition, the model selection module 2313 can select the text modification model 51 for the first domain corresponding to the domain with the highest domain reliability from the multiple text modification models 51, 52, and 53 in the server 2000.

[0085] Alternatively, when the device 1000 provides an identification value of a domain associated with the text to the server 2000 , the model selection module 2313 may select the text modification model 51 of the first domain according to the domain identification value received from the device 1000 .

[0086] The server 2000 may input the text received from the device 1000 to the text modification model 51 of the first domain, and may obtain the modified text output from the text modification model 51. Next, the server 2000 may provide the modified text to the device 1000.

[0087] Figure 6 is a diagram illustrating a server and a device for identifying a domain associated with a text and a server for selecting a text modification model for the domain associated with the text according to an embodiment of the present disclosure.

[0088] refer to Figure 6 , the text output from the ASR model 60 of the device 1000 may be provided to the first domain identification module 61 in the device 1000. The first domain identification module 61 may be the identification module 1440. The device 1000 may obtain a first domain reliability of the text output from the ASR model 60 by using the first domain identification module 61 in the device 1000. In addition, the device 1000 may provide the text output from the ASR model 60 and the first domain reliability obtained from the first domain identification module 61 to the server 2000.

[0089] The server 2000 may receive text from the device 1000 and may provide the received text to the second domain identification module 62 in the server 2000. The second domain identification module 62 may be the domain identification module 2312. The server 2000 may obtain the second domain reliability of the text received from the device 1000 by using the second domain identification module 62 in the server 2000.

[0090] Next, the model selection module 2313 of the server 2000 can select a text modification model for text modification based on the first domain reliability and the second domain reliability. For example, the model selection module 2313 can select a first domain related to the text from the domains registered in the server 2000 based on the weighted sum of the first domain reliability and the second domain reliability, and can select the text modification module 63 of the selected domain. In addition, for example, in this case, because the first domain reliability and the second domain reliability are normalized, the weight values ​​can be reflected in the first domain reliability and the second domain reliability, respectively, but the present disclosure is not limited thereto. Reference will be made to Figure 7 The method of selecting a domain related to a text based on the first domain reliability and the second domain reliability, performed by the server 2000, is described in more detail.

[0091] The server 2000 may input the text received from the device 1000 to the text modification model 63 of the first domain, and may obtain the modified text output from the text modification model 63. Next, the server 2000 may provide the modified text to the device 1000.

[0092] Figure 7 is a flowchart illustrating a method of selecting a domain related to text by using a domain reliability obtained by a device and a domain reliability obtained by a server, performed by a server according to an embodiment of the disclosure.

[0093] In operation S700, the server 2000 can receive, from the device 1000, a first domain reliability of the text calculated by the first domain identification module 61 of the device 1000. The first domain identification module 61 of the device 1000 can calculate a confidence score indicating a degree to which the text output from the ASR model 60 is related to a pre-registered domain. In this case, for example, the domain identification module 61, which is an AI model trained for domain identification, can output the first domain reliability by using the text as an input value. When a plurality of domains are registered in the device 1000, the first domain identification module 61 can obtain a first plurality of domain reliabilities each indicating a degree to which the text is related to each of the plurality of domains.

[0094] In operation S710, the server 2000 can calculate a second domain reliability of the text received from the device 1000 by using the second domain identification module 62. The second domain identification module 62 of the server 2000 can calculate a confidence score indicating a degree to which the text received from the device 1000 is related to a pre-registered domain. In this case, for example, the domain identification module 62, which is an AI model trained for domain identification, can output the second domain reliability by using the text as an input value. When a plurality of domains are registered in the server 2000, the second domain identification module 62 can obtain a second plurality of domain reliabilities each indicating a degree to which the text is related to each of the plurality of domains.

[0095] In operation S720, the server 2000 may select a domain related to the text based on the first domain reliability and the second domain reliability. The server 2000 may select the domain related to the text from a plurality of pre-registered domains based on a weighted sum of the first domain reliability and the second domain reliability. For example, when a domain with a high first domain reliability and a domain with a high second domain reliability are different from each other, the server 2000 may assign a preset first weight value to the first domain reliability and a preset second weight value to the second domain reliability. Alternatively, the server 2000 may select a domain related to the text from a plurality of pre-registered domains based on the first domain reliability assigned the first weight value and the second domain reliability assigned the second weight value. In this case, for example, because the first domain reliability and the second domain reliability are normalized, the weight values ​​may be reflected in the first domain reliability and the second domain reliability, respectively, but the present disclosure is not limited thereto. For example, when the first domain reliability is the reliability of a higher-level domain and the second domain reliability is the reliability of a lower-level domain, a low weight value may be assigned to the first domain reliability and a high weight value may be assigned to the second domain reliability.

[0096] According to an embodiment of the present disclosure, the server 2000 may consider the reliability of the text output from the ASR model of the device to select a domain. In this case, the reliability of the text output from the ASR model may be obtained by the device 1000 and may be provided to the server 2000, but the present disclosure is not limited thereto. In addition, for example, the server 2000 may assign a preset third weight value to the reliability of the text output from the device 1000, and may select a domain related to the text from a plurality of pre-registered domains based on the first domain reliability assigned the first weight value, the second domain reliability assigned the second weight value, and the reliability of the text assigned the third weight value.

[0097] For example, when a domain having high first domain reliability and a domain having high second domain reliability are the same, the server 2000 may select the domain having high first domain reliability from among the plurality of domains without considering the weight value.

[0098] Figure 8 is a diagram illustrating a server selecting a text modification model by using a domain identification module selected from a plurality of domain identification modules in the server according to an embodiment of the present disclosure.

[0099] refer to Figure 8The text output from the ASR model 80 of the device 1000 can be provided to the first domain identification module 81 in the device 1000. The first domain identification module 81 can be the domain identification module 1440. The device 1000 can obtain the first domain reliability of the text output from the ASR model 80 by using the first domain identification module 81 in the device 1000. In addition, the device 1000 can provide the text output from the ASR model 80 and the first domain reliability obtained from the first domain identification module 81 to the server 2000.

[0100] The server 2000 can select one of a plurality of second domain identification modules 82 in the server 2000 based on the first domain reliability received from the device 1000. The domains for speech recognition in the speech recognition system can be hierarchically set. The domains for speech recognition can include, for example, a first layer domain, a second layer domain that is a sub-domain of the first layer domain, a third layer domain that is a sub-domain of the second layer domain, and a fourth layer domain that is a sub-domain of the third layer domain. In addition, for example, the second domain identification module 82 can include, for example, at least one second layer domain identification module 82-1, at least one third layer domain identification module 82-2, and at least one fourth layer domain identification module 82-3. In addition, for example, the first layer domain can correspond to the first domain identification module 81, the second layer domain can correspond to the second layer domain identification module 82-1, the third layer domain can correspond to the third layer domain identification module 82-2, the fourth layer domain can correspond to the fourth layer domain identification module 82-3, and the fifth layer domain can correspond to the text modification model. In this case, the server 2000 can identify a second layer domain having high reliability from among a plurality of second layer domains according to the first domain reliability calculated from the first domain identification module 81 by using the domain identification module selection module 2311. In addition, the domain identification module selection module 2311 of the server 2000 can select the second layer domain identification module 82-1 corresponding to the identified second layer domain.

[0101] In addition, for example, the server 2000 can obtain the second domain reliability of the text by using the selected second layer domain identification module 82-1. For example, the domain identification module selection module 2311 of the server 2000 can identify a third layer domain having high reliability from among a plurality of third layer domains based on the second domain reliability, and can select the third layer domain identification module 82-2 corresponding to the identified third layer domain.

[0102] In addition, for example, the server 2000 can obtain the third domain reliability of the text by using the selected third layer domain identification module 82-2. For example, the domain identification module selection module 2311 of the server 2000 can identify a fourth layer domain having high reliability from among a plurality of fourth layer domains based on the third domain reliability, and can select the fourth layer domain identification module 82-3 corresponding to the identified fourth layer domain.

[0103] In addition, for example, the server 2000 can obtain the fourth domain reliability of the text by using the selected fourth-layer domain identification module 82-3. The model selection module 2313 of the server 2000 can select the text modification model 85 corresponding to the third domain from multiple text modification models based on the fourth domain reliability.

[0104] Next, the server 2000 may input the text received from the device 1000 to the text modification model 85 of the third domain, and may obtain the modified text output from the text modification model 85. The server 2000 may provide the modified text to the device 1000.

[0105] Despite Figure 8 The layers of the second domain identification module 82 in the server 2000 include the second layer, the third layer and the fourth layer, and the server 2000 sequentially selects the second layer domain identification module 82-1, the third layer domain identification module 82-2 and the fourth layer domain identification module 82-3, but the present disclosure is not limited thereto.

[0106] The server 2000 may select a domain for text modification by taking into account the second domain reliability calculated from the second-layer domain identification module 82-1, the third domain reliability calculated from the third-layer domain identification module 82-2, and the fourth domain reliability calculated from the fourth-layer domain identification module 82-3. In this case, the server 2000 may normalize the second domain reliability calculated from the second-layer domain identification module 82-1, the third domain reliability calculated from the third-layer domain identification module 82-2, and the fourth domain reliability calculated from the fourth-layer domain identification module 82-3, and may select a domain for text modification by comparing the normalized values ​​of the domain reliability scores.

[0107] For example, the first domain identification module 81 can calculate the domain reliability of the first layer domain, the second layer domain identification module 82-1 can calculate the domain reliability of the second layer domain, the third layer domain identification module 82-2 can calculate the domain reliability of the third layer domain, and the fourth layer domain identification module 82-3 can calculate the domain reliability of the fourth layer domain.

[0108] The layers of the second domain identification module 82 in the server 2000 may include only the second layer. Alternatively, the layers of the second domain identification module 82 in the server 2000 may include other layers in addition to the second to fourth layers. In this case, according to the layers of the second domain identification module 82 in the server 2000, the server 2000 may include a domain identification module corresponding to each layer.

[0109] For example, the first domain identification module 81 can calculate the reliability of the domain of "location" as the first layer domain as 60%, and can calculate the reliability of "weather" as the first layer domain as 30%. The second layer domain identification module 82-1 can calculate the reliability of "Canada" as the second layer domain as 40%, can calculate the reliability of "USA" as the second layer domain as 20%, and can calculate the reliability of "rain" as the second layer domain as 25%. The third layer domain identification module 82-2 can calculate the reliability of "British Columbia" as the third layer domain as 20%, can calculate the reliability of "Ontario" as the third layer domain as 30%, can calculate the reliability of "New York" as the third layer domain as 10%, and can calculate the reliability of "precipitation" as the third layer domain as 5%.

[0110] In addition, for example, the domain selection module 2313 can assign a first weight value to the reliability of the domain of "location" as the first layer domain and the reliability of "weather" as the first layer domain. The domain selection module 2313 can assign a second weight value to the reliability of "Canada" as the second layer domain, the reliability of "USA" as the second layer domain, and the reliability of "rain" as the second layer domain. The domain selection module 2313 can assign a third weight value to the reliability of "British Columbia" as the third layer domain, the reliability of "Ontario" as the third layer domain, the reliability of "New York" as the third layer domain, and the reliability of "precipitation" as the third layer domain. In this case, the second weight value can be greater than the first weight value and can be less than the third weight value. In addition, the domain selection module 2313 can select a domain for text modification in consideration of the reliability assigned with the first weight value, the reliability assigned with the second weight value, and the reliability assigned with the third weight value.

[0111] In addition, for example, the domain selection module 2313 can calculate a first weighted sum of the reliability of "location", the reliability of "Canada", and the reliability of "Columbia". The domain selection module 2313 can calculate a second weighted sum of the reliability of "location", the reliability of "Canada", and the reliability of "Ontario". The domain selection module 2313 can calculate a third weighted sum of the reliability of "location", the reliability of "USA", and the reliability of "New York". In addition, for example, the domain selection module 2313 can calculate a fourth weighted sum of the reliability of "weather", the reliability of "rain", and the reliability of "precipitation".

[0112] For example, the domain selection module 2313 can determine that the first weighted sum is the highest by comparing the calculated weighted sums, and can determine "British Columbia" as a domain for text modification.

[0113] In addition, for example, the domain selection module 2313 may select a second-layer domain based on the domain reliability of the first-layer domain and the reliability of the second-layer domain. The domain selection module 2313 may select a subdomain related to the selected second-layer domain and may modify the text by using a text modification model corresponding to the selected subdomain.

[0114] In addition, despite Figure 8 The first domain identification module 81 in the device 1000 corresponds to the first layer, and the second domain identification module 82 in the server 2000 corresponds to the second to fourth layers, but the present disclosure is not limited thereto. For example, the first domain identification module 81 in the device 1000 may correspond to the first layer, and the second domain identification module 82 in the server 2000 may correspond to the first to third layers.

[0115] The weight value of the domain reliability may be assigned based on context information related to the device 1000. The context information may include, but is not limited to, at least one of the surrounding environment information of the device 1000, the status information of the device 1000, the status information of the user, the device usage history information of the user, or the schedule information of the user. The surrounding environment information of the device 1000, which is the environmental information within a certain radius of the device 1000, may include, for example, but is not limited to, weather information, temperature information, humidity information, illumination information, noise information, and sound information. The status information of the device 1000 may include, but is not limited to, information about the mode of the device 1000 (e.g., sound mode, vibration mode, silent mode, power saving mode, cut-off mode, multi-window mode, and automatic rotation mode), the location information of the device 1000, time information, activation information of the communication module (e.g., Wi-Fi is on, Bluetooth is off, GPS is on, or NFC is on), the network connection status information of the device 1000, and information about applications running in the device 1000 (e.g., application identification information, application type, application usage time, or application usage cycle). The user's state information, which is information about the user's movement and lifestyle, may include but is not limited to information about the user's walking state, exercise state, driving state, sleeping state, and emotional state. The user's device usage history information, which is information about events in which the user uses the device 1000, may include but is not limited to information about running applications, functions of running applications, the user's phone conversations, and the user's text messages.

[0116] For example, the weight value of domain reliability can be determined based on contextual information about an application running on device 1000. For example, when the user's voice input is input to an application running on device 1000, a high weight value can be assigned to the domain reliability of the domain associated with the application. Alternatively, the domain associated with the application can be directly determined as the domain used for text modification. For example, when a voice input saying "Acrovista" is input while a map application is running on device 1000, a high weight value can be assigned to the map domain, or the map domain can be directly determined as the domain used for text modification.

[0117] For example, the weight value of the domain reliability may be determined based on a conversation history of a user of a voice assistant service provided through the device 1000. For example, when a voice input saying "search IU" is input to the device 1000 while the user is talking about music to the device 1000 through the voice assistant service, a high weight value may be assigned to the music domain, or the music domain may be directly determined as the domain for text modification.

[0118] For example, a weight value of domain reliability may be determined based on sensing information collected by device 1000. A weight value may be assigned to a domain based on location information (e.g., GPS information) obtained by device 1000. For example, when device 1000 is located near a movie theater, a high weight value may be assigned to the movie domain. For example, when a user's voice input is input to device 1000 while searching for a restaurant in device 1000, a high weight value may be assigned to a domain related to the location of device 1000.

[0119] A weight value for domain reliability can be assigned based on trend information. For example, a high weight value can be assigned to a domain of major news or a domain of a real-time search term through a portal site.

[0120] Figure 9 is a flowchart illustrating a method in which a server selects a domain for text modification by using a domain identification module selected from a plurality of domain identification modules according to an embodiment of the present disclosure.

[0121] Figure 10 is a diagram illustrating a first domain identifying module, a second domain identifying module, and a text modifying module in relation to hierarchically classified domains according to an embodiment of the present disclosure.

[0122] For example, in Figure 9 and 10 In [1], the domains used for speech recognition can be classified into the first, second, and third layers.

[0123] In operation S900, the server 2000 may receive, from the device 1000, a first domain reliability of a text calculated by the first domain identification module 81 of the device 1000. The server 2000 may receive, from the device 1000, a first domain reliability of a text output from the ASR model of the device 1000. For example, referring to Figure 10 , the first domain identification module 100 of the device 1000 may correspond to "all" as a first-level domain related to "location", and the first domain reliability calculated from the first domain identification module 100 may be the domain reliability of the second-level domain related to "country". For example, the first domain reliability may include the domain reliability of the domain "Canada" and the domain reliability of the domain "USA". In addition, for example, the second domain identification module 101 may correspond to the domain "Canada", and the second domain identification module 102 may correspond to the domain "USA".

[0124] In operation S910, the server 2000 may select at least one of the plurality of second domain identification modules 82 based on the first domain reliability. The domain identification module selection module 2311 of the server 2000 may select the second layer domain identification module 82-1 from the plurality of second domain identification modules 82 based on the first domain reliability. Figure 10 For example, server 2000 may compare the domain reliability of "Canada" with the domain reliability of "USA" and determine that the domain reliability of "Canada" is higher than a certain threshold. Additionally, server 2000 may select second domain identification module 101 corresponding to "Canada" from second domain identification module 101 corresponding to "Canada" and second domain identification module 102 corresponding to "USA." In this case, second domain identification module 101 and second domain identification module 102 may be second domain identification modules corresponding to the second layer.

[0125] In operation S920, the server 2000 may calculate the second domain reliability of the text by using the selected second-layer domain identification module 82-1. The second-layer domain identification module 82-1 may calculate the second domain reliability by using the text as input. Figure 10 The second domain reliability calculated by the second domain identification module 101 may be the domain reliability of the third-level domain related to "province or state". For example, the second domain reliability may include the domain reliability of the domain "British Columbia", the domain reliability of the domain "Ontario", the domain reliability of the domain "New York", and the domain reliability of the domain "Illinois". In addition, for example, the text modification model 103 may correspond to the domain "British Columbia", the text modification model 104 may correspond to the domain "Ontario", the text modification model 105 may correspond to the domain "New York", and the text modification model 106 may correspond to the domain "Illinois".

[0126] In operation S930, the server 2000 may select a domain related to text modification based on the second domain reliability. The model selection module 2313 of the server 2000 may select one of the plurality of text modification models 83, 84, and 85 based on the second domain reliability. Figure 10 For example, server 2000 may compare the domain reliability of the domain "British Columbia," the domain reliability of the domain "Ontario," the domain reliability of the domain "New York," and the domain reliability of the domain "Illinois," and may determine that the domain reliability of the domain "British Columbia" is greater than a certain threshold. Furthermore, server 2000 may select the domain "British Columbia" as the domain for text modification. Therefore, the text may be modified by text modification model 103 corresponding to the domain "British Columbia."

[0127] Despite Figure 9 and 10 In the embodiment, the first domain identification module 81 corresponds to the first layer, the second domain identification module 82 corresponds to the second layer, and the text modification models 83, 84, and 85 correspond to the third layer, but the present disclosure is not limited thereto. For example, the second domain identification module 82 may correspond to more layers. For example, the first domain identification module 81 may correspond to the first layer, the second domain identification module 82 may correspond to the second layer, the third layer, and the fourth layer, and the text modification models 83, 84, and 85 may correspond to the fifth layer, but the present disclosure is not limited thereto.

[0128] exist Figure 10 , the server 2000 may normalize the domain reliability calculated from the first domain identification module 100, the domain reliability calculated from the second domain identification module 101, and the domain reliability calculated from the second domain identification module 102, and may select a text modification model for text modification by comparing the normalized values.

[0129] Figure 11 is a diagram illustrating a server that modifies text by using a plurality of text modification models according to an embodiment of the present disclosure.

[0130] refer to Figure 11 , the text output from the ASR model 110 of the device 1000 may be provided to the domain identification module 2312 in the server 2000. The server 2000 may identify a domain related to the text received from the device 1000 by using the domain identification module 2312 in the server 2000. In this case, the server 2000 may identify a domain related to the text based on the domain reliability of the text received from the device 1000.

[0131] The domain identification module 2312 can obtain, for example, domain reliability of a first domain, domain reliability of a second domain, and domain reliability of a third domain. For example, the domain identification module 2312 can divide the text into multiple sections and obtain domain reliability of a first domain, domain reliability of a second domain, and domain reliability of a third domain for each section. For example, the domain identification module 2312 can classify the text into a first section, a second section, and a third section and obtain domain reliability of a first domain, domain reliability of a second domain, and domain reliability of a third domain for each of the first section, the second section, and the third section. For example, when the server 2000 receives a text stream, the server 2000 can obtain a first domain reliability for the text of the first section, a second domain reliability for the text of the second section, and a third domain reliability for the text of the third section. In this case, while receiving the text stream, the server 2000 can modify the text that has been received in real time without waiting for text to be received later, and can identify domains related to the text at a higher speed.

[0132] Alternatively, for example, when the server 2000 receives a text stream, the server 2000 may accumulate and calculate domain reliability for texts of multiple sections. For example, while receiving the text stream, the server 2000 may divide the text of a sentence into multiple sections and obtain the domain reliability of the text of the first section, the domain reliability of the text of the first and second sections, and the domain reliability of the text of the first to third sections. In this case, because the server 2000 calculates domain reliability on a sentence-by-sentence basis by accumulating multiple sections, the server 2000 can more efficiently identify domains associated with the text.

[0133] Or, for example, for a sentence unit of text, the domain identification module 2312 may obtain a first domain reliability of a first domain, a second domain reliability of a second domain, and a third domain reliability of a third domain.

[0134] According to an embodiment of the present disclosure, the model selection module 2313 of the server 2000 can select a text modification model for text modification. For example, when the text is divided into multiple sections, the server 2000 can select different text modification models according to the sections of the text. For example, when the text is divided into the first section, the second section, and the third section, the server 2000 can select the text modification model 111 to modify the first section of the text, can select the text modification model 112 to modify the second section of the text, and can select the text modification model 113 to modify the third section of the text.

[0135] Alternatively, for example, the server 2000 may select multiple text modification models to modify the text of a sentence unit. For example, the server 2000 may select the text modification model 111, the text modification model 112, and the text modification model 113 to modify the text of a sentence unit.

[0136] Next, the server 2000 may obtain a modified text of the text received from the device 1000 by using the first modified text output from the text modification model 111 , the second modified text output from the text modification model 112 , and the third modified text output from the text modification model 113 .

[0137] For example, the server 2000 may select at least one of the first modified text, the second modified text, or the third modified text, and may obtain the modified text of the text received from the device 1000 by using any portion of the selected modified text. Alternatively, for example, the server 2000 may obtain the modified text of the text received from the device 1000 by selecting one of the first modified text, the second modified text, and the third modified text. Alternatively, for example, the server 2000 may obtain the modified text of the text received from the device 1000 by combining at least a portion of the first modified text, at least a portion of the second modified text, and at least a portion of the third modified text.

[0138] Next, the server 2000 may provide the modified text to the device 1000 .

[0139] Figure 12 is a flowchart illustrating a method in which a server accumulates and calculates domain reliabilities of texts of a plurality of sections according to an embodiment of the present disclosure.

[0140] In operation S1200, the server 2000 may obtain a first section of the text. The text may be divided into a plurality of sections, and each section of the text may be divided, for example, into units of syntactic words, words, or phrases. The server 2000 may receive the text from the device 1000 as a text stream. In this case, the server 2000 may obtain the first section of the text while receiving the text stream in real time. Alternatively, the server 2000 may receive the text as a sentence from the device 1000 and extract the text of the first section from the received text.

[0141] In operation S1210, the server 2000 may calculate the domain reliability of the text of the first section. For the domains registered in the server 2000, the server 2000 may calculate the domain reliability of the text of the first section.

[0142] In operation S1220, the server 2000 may obtain the second section of the text. When the server 2000 receives the text as a text stream from the device 1000, the server 2000 may obtain the second section of the text while receiving the text stream in real time. Alternatively, the server 2000 may receive the text as a sentence from the device 1000 and extract the text of the second section from the received text.

[0143] In operation S1230, the server 2000 may calculate domain reliabilities of the texts of the first section and the second section. The server 2000 may accumulate the texts of the first section and the second section, and may calculate the domain reliabilities of the texts of the first section and the second section.

[0144] In operation S1240, the server 2000 may obtain the nth section of the text. When the server 2000 receives the text as a text stream from the device 1000, the server 2000 may obtain the nth section of the text while receiving the text stream in real time. Alternatively, the server 2000 may receive the text as a sentence from the device 1000 and extract the text of the nth section from the received text.

[0145] In operation S1250, the server 2000 may calculate domain reliabilities of texts of the first to nth sections. The server 2000 may accumulate texts of the first to nth sections and calculate domain reliabilities of texts of the first to nth sections.

[0146] In operation S1260, the server 2000 may determine a domain for modifying the text received from the device 1000 based on the domain reliabilities of the texts of the first through nth sections.

[0147] Figure 13 is a diagram illustrating a server that obtains domain reliability of a text stream accumulated in units of syntactic words according to an embodiment of the present disclosure.

[0148] refer to Figure 13 The domain identification module 2312 of the server 2000 may divide the text in units of syntactic words, may accumulate the divided text, and may obtain domain reliability of the accumulated text.

[0149] For example, when "new" as the text of the first section is input to the domain identification module 2312, the domain identification module 2312 may output "reject" as the domain identification value because the domain reliability associated with "new" is "0.1" which is a low value.

[0150] Next, when "twice" as the text of the second section is input into the domain identification module 2312, the domain identification module 2312 can accumulate "new" as the text of the first section and "twice" as the text of the second section, and can output "music" as the domain identification value related to "new twice" as the accumulated text, and output "0.7" as the domain reliability.

[0151] Next, when "yes or no" as the text of the third section is input into the domain identification module 2312, the domain identification module 2312 can accumulate "new" as the text of the first section, "twice" as the text of the second section, and "yes or no" as the text of the third section, and can output "music" as the domain identification value related to "new twice yes or no" as the accumulated text, and output "0.9" as the domain reliability.

[0152] Next, when "play" as the text of the fourth section is input into the domain identification module 2312, the domain identification module 2312 can accumulate "new" as the text of the first section, "twice" as the text of the second section, "yes or no" as the text of the third section, and "play" as the text of the fourth section, and can output "music" as the domain identification value related to "play new twice yes or no" as the accumulated text, and output "1.0" as the domain reliability.

[0153] Therefore, the server 2000 may select the domain modification model of the domain "music" as the domain modification model for modifying "play new twice yes or no".

[0154] Despite Figure 13 The domain identification module 2312 outputs the domain and the domain reliability having the highest value, but the present disclosure is not limited thereto. The domain identification module 2312 may output the domain reliability of each of the plurality of domains registered in the server 2000.

[0155] Despite Figure 13 The domain reliabilities of the texts of the multiple sections are accumulated and calculated, but the present disclosure is not limited to this. For example, the server 2000 can sequentially select text modification models for modifying the texts of the multiple sections. For example, when the server 2000 receives a text stream, the server 2000 can select a text modification model for modifying the text of the first section by calculating the domain reliability of the text of the first section while receiving the text stream, can select a text modification model for modifying the text of the second section by calculating the domain reliability of the text of the second section, and can select a text modification model for modifying the text of the nth section by calculating the domain reliability of the text of the nth section.

[0156] Figure 14 is a flowchart illustrating a method in which a server divides a text into a plurality of sections and selects a domain of the text for each of the plurality of sections according to an embodiment of the present disclosure.

[0157] In operation S1400 , the server 2000 may calculate domain reliability of the text for each of the plurality of domains. The server 2000 may calculate domain reliability of the plurality of domains registered in the server 2000 for the text received from the device 1000 .

[0158] In operation S1410, the server 2000 may divide the text into a plurality of sections by comparing the calculated domain reliabilities. The server 2000 may divide the text into a plurality of sections by identifying a text section having high domain reliability for each domain.

[0159] In operation S1420, the server 2000 may select a domain for text modification for each section text. The server 2000 may select a domain with the highest domain reliability for each section text as a domain corresponding to the text of each section.

[0160] Figure 15 is a diagram illustrating an example of a server of a text modification model that compares domain reliability of text according to a plurality of domains and selects and modifies text of each section according to an embodiment of the present disclosure.

[0161] refer to Figure 15 , the server 2000 can divide the text into multiple sections while receiving the text stream, and can calculate the domain reliability of the text of each section in real time. For example, the server 2000 can receive a text stream saying "Today near Yeomtong Station I will meet Gil Hong to watch Avengers: Love Dream". The server 2000 can identify "Today near Yeomtong Station" as the first section, "I will meet Gil Hong" as the second section, and "To watch Avengers: Love Dream" as the third section while receiving the text stream. In addition, while receiving the text stream, the server 2000 can sequentially calculate the domain reliability related to "Today near Yeomtong Station", the domain reliability related to "I will meet Gil Hong", and the domain reliability related to "Today near Yeomtong Station".

[0162] For example, for each section of the text saying “Today near Yeomtong Station I will meet Gil Hong to watch Avengers: Age of Ultron”, the domain identification module 2312 of the server 2000 can calculate the domain reliability of the domain “movie”, the domain reliability of the domain “location”, and the domain reliability of the domain “contact”.

[0163] For example, the domain selection module 2313 of the server 2000 can compare the domain reliability of the domain "movie", the domain reliability of the domain "location", and the domain reliability of the domain "contact". The domain selection module 2313 can determine that the domain reliability of the domain "location" is high for "today near Yeomtong station". The domain selection module 2313 can determine that the domain reliability of the domain "contact" is high for "I will meet Gil Hong". The domain selection module 2313 can determine that the domain reliability of the domain "movie" is high for "go see Avengers: Infinity War". Accordingly, the domain selection module 2313 can sequentially identify "today near Yeomtong station" as a first section, can identify "I will meet Gil Hong" as a second section, and can identify "go see Avengers: Infinity War" as a third section from the text stream saying "today near Yeomtong station I will meet Gil Hong go see Avengers: Infinity War".

[0164] In addition, for example, the domain selection module 2313 can select a domain related to "today near Yeomtong station" as the domain "location", can select a domain related to "I will meet Gil Hong" as the domain "contact", and can select a domain related to "go see Avengers: Infinity War" as the domain "movie".

[0165] The text modification model of the domain "location" can modify "today near Yeomtong station" to "today near Yeongtong station", the text modification model of the domain "contact" can modify "I will meet Gil Hong" to "I will meet Gil Dong", and the text modification model of the domain "movie" can modify "go see Avengers: Infinity War" to "go see Avengers Alliance". At least one of the text modification operation of the text modification model of the domain "location", the text modification operation of the text modification model of the domain "contact", or the text modification operation of the text modification model of the domain "movie" can be sequentially performed while the text stream is received.

[0166] Figure 16 FIG. 1 is a diagram illustrating a server that modifies text received from a device by using modified text output from a plurality of text modification models according to an embodiment of the disclosure.

[0167] For example, referring to Figure 16 , the server 2000 can provide the text 160 saying "today near Yeomtong station I will meet Gil Hong go see Avengers: Infinity War" to the text modification model of the domain "location", the text modification model of the domain "contact", and the text modification model of the domain "movie".

[0168] Therefore, the text modification model of the domain "location" can output modified text 161 saying "Today near Yeongtong Station I will meet Gil Hong to see Avengers: Age of Ultron", the text modification model of the domain "movie" can output modified text 162 saying "Today near Yeomtong Station I will meet Gil Hong to see Avengers: Age of Ultron", and the text modification model of the domain "contact" can output modified text 163 saying "Today near Yeomtong Station I will meet Gil Dong to see Avengers: Age of Ultron".

[0169] Next, the server 2000 can recognize "Yeongtong" as a modified word in the modified text 161, "Avengers" as a modified word in the modified text 162, and "GilDong" as a modified word in the modified text 163, and can generate "Today near Yeongtong Station I will meet Gil Dong to see the Avengers" as a modified text 164 to be provided to the device 1000.

[0170] Figure 17 is a block diagram illustrating a server according to an embodiment of the present disclosure.

[0171] refer to Figure 17 , the server 2000 according to an embodiment of the present disclosure may include a communication interface 2100, a processor 2200 and a storage device 2300, and the storage device 2300 may include a domain management module 2310, a text modification module 2320, an NLU module 2330 and a speech analysis management module 2340.

[0172] The communication interface 2100 may include at least one element for communicating with the device 1000 and another server. The communication interface 2100 may send and receive information for speech recognition and voice assistant services to and from the device 1000 and another server. The communication interface 2100 may perform communication via, for example, but not limited to, a local area network (LAN), a wide area network (WAN), a value-added network (VAN), a mobile radio communication network, a satellite communication network, or a combination thereof.

[0173] The processor 2200 controls the overall operation of the server 2000. The processor 2200 can generally control the operation of the server 2000 described herein by executing a program stored in the storage device 2300.

[0174] The storage device 2300 may store programs used by the processor 2200 to perform processing and control, and may store data input to or output from the server 2000. The storage device 2300 may include, but is not limited to, at least one type of storage medium selected from a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., secure digital (SD) or extreme digital (XD) memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable ROM (EEPROM), a programmable ROM (PROM), a magnetic memory, a magnetic disk, or an optical disk.

[0175] Programs stored in the storage device 2300 may be classified into a plurality of modules according to their functions, for example, a domain management module 2310 , a text modification module 2320 , an NLU module 2330 , and a speech analysis management module 2340 .

[0176] The domain management module 2310 provides the text received from the device 1000 to the text modification module 2320. The domain management module 2310 may include a domain identification module selection module 2311, at least one domain identification module 2312, and a domain selection module 2313.

[0177] The domain identification module selection module 2311 may select a domain identification module 2312. When a plurality of domain identification modules 2312 exist, the domain identification module selection module 2311 may select at least some of the plurality of domain identification modules 2312.

[0178] The domain identification module selection module 2311 can select one of the multiple domain identification modules 2312 in the server 2000 based on the first domain reliability received from the device 1000. The domains used for speech recognition in the speech recognition system can be arranged in layers. The domains used for speech recognition can include, for example, a first-layer domain, a second-layer domain that is a subdomain of the first-layer domain, and a third-layer domain that is a subdomain of the second-layer domain. Furthermore, for example, the first-layer domain can correspond to the domain identification module 1440 of the device 1000, the second-layer domain can correspond to the domain identification module 2312 of the server 2000, and the third-layer domain can correspond to the text modification module 2320. In this case, the domain identification module selection module 2311 can identify a second-layer domain with high reliability from the multiple second-layer domains based on the first domain reliability calculated from the domain identification module 1440 of the device 1000. Furthermore, the domain identification module selection module 2311 can select the domain identification module 2312 corresponding to the identified second-layer domain.

[0179] Domain identification module 2312 can identify domains for text modification. When server 2000 receives domain information from device 1000, domain identification module 2312 can identify domains for text modification based on the domain information. Alternatively, when server 2000 does not receive domain information from device 1000, domain identification module 2312 can identify domains related to the text based on the domain reliability of the text received from device 1000. For example, domain identification module 2312 can calculate a confidence score indicating the degree to which the text received from device 1000 is related to domains pre-registered for text modification. Additionally, domain identification module 2312 can identify domains related to the text received from the device based on domain reliability calculated for pre-registered domains. Domain identification module 2312 can identify domains related to the text based on rules, or can obtain domain reliability related to the text using an AI model trained for domain identification. Additionally, for example, the AI ​​model used for domain identification can be part of a NLU model or a separate model from the NLU model.

[0180] The domain identification module 2312 may identify domains associated with the text by accumulating texts of multiple sections and calculating domain reliability of the accumulated texts. Alternatively, the domain identification module 2312 may divide the text into multiple sections and identify domains associated with the text of each section.

[0181] The domain selection module 2313 may select a text modification model corresponding to the domain identified by the domain identification module 2312 from among the plurality of text modification models 2321 , 2322 , and 2323 .

[0182] The domain selection module 2313 may select a domain corresponding to the domain identified by the domain identification module 2312 from among the domains registered in the server 2000 and may select a text modification model of the selected domain.

[0183] When the domain identification module 2312 divides the text into a plurality of sections and identifies a domain associated with the text of each section, the domain selection module 2313 may select the domain of the text of each section.

[0184] The text modification module 2320 modifies the text received from the device 1000. The text modification module 2320 can modify the text by using a text modification model corresponding to the determined domain. The text modification module 2320 may include a text modification model 2321 for the first domain, a text modification model 2322 for the second domain, and a text modification model 2323 for the third domain.

[0185] The text modification module 2320 can generate modified text by using the selected text modification model. The text modification module 2320 can input the text to the selected text modification model, and can obtain the modified text output from the text modification model. In this case, the text modification module 2320 can pre-process the format of the text received from the device 1000 to be suitable for processing the text modification model, and can input the pre-processed value to the text modification model.

[0186] When the text received from the device 1000 relates to a plurality of domains, the text modification module 2320 can select a plurality of text modification models corresponding to the plurality of domains to modify the text. In this case, the text modification module 2320 can obtain the modified text to be provided to the device 1000 from among a plurality of pieces of modified text output from the plurality of text modification models. For example, when the text modification module 2320 generates a plurality of pieces of modified text by using a plurality of text modification models, the text modification module 2320 can compare the reliability of the plurality of pieces of modified text, and can determine the modified text having high or highest reliability as the modified text to be provided to the device 1000. The reliability of the modified text can be a value indicating the degree to which the modified text matches the input speech, and can include, for example, and without limitation, a confidence score.

[0187] In addition, for example, when the text modification module 2320 generates a plurality of pieces of modified text by using a plurality of text modification models, the text modification module 2320 can extract a modified portion from among the plurality of pieces of modified text, and can obtain the modified text to be provided to the device 1000 by using the extracted modified portion.

[0188] In addition, for example, the text modification module 2320 can obtain the modified text to be provided to the device 1000 by extracting a plurality of pieces of text from among a plurality of pieces of modified text output from a plurality of text modification models and combining the extracted plurality of pieces of text. For example, when the text modification module 2320 generates a first modified text and a second modified text by using a plurality of text modification models and the reliability of a portion of the first modified text and the reliability of a portion of the second modified text are high, the text modification module 2320 can obtain the modified text to be provided to the device 1000 by combining the portion of the first modified text and the portion of the second modified text to generate modified text having higher reliability than the reliability of the first modified text and the second modified text.

[0189] In addition, for example, when a domain related to the text of each section is selected, the text modification module 2320 can provide the text of each section to a corresponding domain modification model. In this case, the text modification module 2320 can obtain the modified text to be provided to the device 1000 by combining a plurality of pieces of modified text of a plurality of sections output from the respective domain modification models.

[0190] NLU module 2330 can interpret the modified text output from text modification module 2320. NLU module 2330 can include multiple NLU models for multiple domains, such as a first NLU model 2331 and a second NLU model 2332. The result values ​​generated when NLU module 2330 interprets text can include, for example, intent and parameters. Intent is information determined by interpreting text using an NLU model, which can indicate, for example, the user's utterance intention. Intent can include information indicating the user's utterance intention (hereinafter referred to as intent information) and a numerical value corresponding to the information indicating the user's intent. The numerical value can indicate the probability that the text is associated with information indicating a specific intent. When multiple pieces of intent information indicating the user's intent are obtained as a result of interpreting text using an NLU model, the intent information with the largest numerical value among the multiple pieces of intent information can be determined as the intent. In addition, parameters can indicate detailed information related to the intent. Parameters can be information related to the intent, and multiple types of parameters can correspond to one intent.

[0191] Furthermore, the result value generated when the NLU module 2330 interprets the text may be used to provide some kind of voice assistant service to the device 1000 .

[0192] The speech analysis management module 2340 may evaluate the text modified by the text modification module 2320 and may determine whether to perform NLU processing on the modified text. The speech analysis management module 2340 may include a speech recognition evaluation module 2341 and an NLU determination module 2342.

[0193] The speech recognition evaluation module 2341 can calculate the reliability of the text modified by the text modification module 2320. The reliability of the modified text can be a value indicating the probability that the modified text matches the input speech, and can include, for example, but not limited to, a confidence score. Furthermore, the speech recognition evaluation module 2341 can calculate the domain reliability of the modified text. The speech recognition evaluation module 2341 can calculate the domain reliability, which indicates the degree to which the modified text is related to a domain pre-registered in the server 2000 for NLU processing.

[0194] The NLU determination module 2342 may determine whether to perform NLU processing on the modified text in the server 2000. The NLU determination module 2342 may determine whether to perform NLU processing in the server 2000 based on the reliability of the modified text and the domain reliability of the modified text. The NLU determination module 2342 may determine whether the domain associated with the modified text is a domain for performing NLU processing in the device 1000 or a domain for performing NLU processing in the server 2000.

[0195] Figure 18is a block diagram illustrating a device according to an embodiment of the present disclosure.

[0196] refer to Figure 18 , the device 1000 according to an embodiment of the present disclosure may include a communication interface 1100, an input / output interface 1200, a processor 1300 and a memory 1400, and the memory 1400 may include at least one ASR model 1410, at least one NLU model 1420, a speech recognition evaluation module 1430, a domain recognition module 1440 and an NLU determination module 1450.

[0197] The communication interface 1100 may include at least one component for communicating with the server 2000 and external devices. The communication interface 1100 may send and receive information for speech recognition and voice assistant services to and from the server 2000 and external devices. The communication interface 1100 may communicate via, for example, but not limited to, a local area network (LAN), a wide area network (WAN), a value-added network (VAN), a mobile radio communication network, a satellite communication network, or a combination thereof.

[0198] The input / output interface 1200 may receive data input to the device 1000 and may output data from the device 1000. The input / output interface 1200 may include a user input interface, a camera, a microphone, a display, and an audio output interface. The user input interface may include, but is not limited to, a keyboard, a dome switch, a touchpad (e.g., a capacitive cover type, a resistive cover type, an infrared beam type, an integral strain gauge type, a surface acoustic wave type, a piezoelectric type, etc.), a scroll wheel, or a scroll wheel switch.

[0199] The display can display and output information processed by the device 1000. For example, the display can display a graphical user interface (GUI) for a voice assistant service. When the display forms a layer structure together with the touchpad to construct a touch screen, the display can be implemented as an input device as well as an output device. The display can include at least one of a liquid crystal display (LCD), a thin film transistor-liquid crystal display (TFT-LCD), an organic light-emitting diode (OLED), a flexible display, a three-dimensional (3D) display, or an electrophoretic display.

[0200] The audio output interface may output audio data and may include, for example, a speaker and a buzzer.

[0201] The camera can obtain image frames such as still images or moving pictures by using the image sensor in a video call mode or a shooting mode. The image captured by the image sensor can be processed by the processor 1300 or a separate image processor.

[0202] The microphone may receive the user's speech and may process the user's speech into electronic voice data.

[0203] The processor 1300 controls the overall operation of the device 1000. The processor 1300 may control the overall operation of the device 1000 described herein by executing a program stored in the memory 1400.

[0204] The memory 1400 may store programs used by the processor 1300 to perform processing and control, and may store data input to or output from the device 1000. The memory 1400 may include, but is not limited to, at least one type of storage medium selected from the group consisting of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., an SD or XD memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable ROM (EEPROM), a programmable ROM (PROM), a magnetic memory, a magnetic disk, or an optical disk.

[0205] The programs stored in the memory 1400 may be classified into a plurality of modules according to their functions, for example, an ASR model 1410 , an NLU model 1420 , a speech recognition evaluation module 1430 , a domain recognition module 1440 , and an NLU determination module 1450 .

[0206] The ASR model 1410 can obtain text from a feature vector generated from the user's voice input. The processor 1300 of the device 1000 can input the feature vector into the ASR model 1410 to recognize the user's voice. When the device 1000 includes multiple ASR models 1410, the processor 1300 of the device 1000 can select one of the multiple ASR models 1410 and convert the feature vector into a format suitable for the selected ASR model 1410. The ASR model 1410 can be an AI model including, for example, an acoustic model, a pronunciation dictionary, and a language model. Alternatively, the ASR model 1410 can be an end-to-end speech recognition model having a structure including an integrated neural network without having to separately include, for example, an acoustic model, a pronunciation dictionary, and a language model.

[0207] The NLU model 1420 may interpret the text output from the ASR model 1410. Alternatively, the NLU model 1420 may interpret the modified text provided from the server 2000. The result value generated when the NLU model 1420 interprets the text or the modified text may be used to provide a certain voice assistant service to the user.

[0208] The speech recognition evaluation module 1430 can obtain the reliability of the text output from the ASR model 1410. The reliability of the text can be a value indicating the degree to which the text output from the ASR model 1410 correlates with the input speech, and can include, for example, but not limited to, a confidence score. Furthermore, the reliability of the text can be related to the probability that the text matches the input speech. For example, the reliability of the text can be calculated based on at least one of the likelihood of multiple estimated texts output from the ASR model 1410 of the device 1000 or the posterior probability that at least one character in the text will be replaced by another character. For example, the speech recognition evaluation module 1430 can calculate the reliability based on the likelihood output as a result of Viterbi decoding. Alternatively, for example, the speech recognition evaluation module 1430 can calculate the reliability based on the posterior probability output from the softmax layer in the end-to-end ASR model. Alternatively, for example, the speech recognition evaluation module 1430 can determine multiple estimated texts estimated during the speech recognition process of the ASR model 1410 of the device 1000 and calculate the reliability of the text based on the correlation of the characters in the multiple estimated texts.

[0209] In addition, the speech recognition evaluation module 1430 may determine whether to transmit the text output from the ASR model 1410 to the server 2000. The speech recognition evaluation module 1430 may determine whether to transmit the text to the server 2000 by comparing the reliability of the text with a preset threshold. When the reliability of the text is equal to or greater than the preset threshold, the speech recognition evaluation module 1430 may determine not to transmit the text to the server 2000. In addition, when the reliability of the text is less than the preset threshold, the speech recognition evaluation module 1430 may determine to transmit the text to the server 2000.

[0210] In addition, the speech recognition evaluation module 1430 may determine whether to send the text to the server 2000 based on at least one text with high reliability among the multiple estimated texts estimated during the speech recognition process of the ASR model 1410. For example, when the multiple estimated texts estimated during the speech recognition process of the ASR model 1410 include a first estimated text with high reliability and a second estimated text with high reliability, and the difference between the reliability of the first estimated text and the reliability of the second estimated text is equal to or less than a certain threshold, the speech recognition evaluation module 1430 may determine to send the text to the server 2000.

[0211] The domain identification module 1440 can identify domains related to the text output from the ASR model 1410. The domain identification module 1440 can identify domains related to the text based on the domain reliability of the text output from the ASR model 1410. For example, the domain identification module 1440 can calculate a confidence score indicating the degree to which the text output from the ASR model 1410 is related to a pre-registered domain. As an AI model trained for domain identification, the domain identification module 1440 can output domain reliability using the text as an input value. In addition, for example, the domain identification module 1440 can be part of the NLU model or a model separate from the NLU model. Alternatively, the domain identification module 1440 can identify domains related to the text based on rules.

[0212] Domains used for speech recognition in the speech recognition system may be set hierarchically, and the domain recognized by the domain recognition module 1440 of the device 1000 may be a higher-level domain than the domain recognized by the domain recognition module 2312 of the server 2000 .

[0213] The NLU determination module 1450 may determine whether to perform NLU processing on the text output from the ASR model 1410 in the device 1000 or the server 2000. The NLU determination module 1450 may determine whether the domain associated with the text output from the ASR model 1410 is a domain for performing NLU processing in the device 1000. When the domain associated with the text output from the ASR model 1410 is a domain pre-registered in the device 1000, the NLU determination module 1450 may determine that the device 1000 will perform NLU processing. In addition, when the domain associated with the text output from the ASR model 1410 is not a domain pre-registered in the device 1000, the NLU determination module 1450 may determine that the device 1000 will not perform NLU processing.

[0214] The AI-related functions according to the present disclosure are performed by a processor and a memory. The processor may include at least one processor. In this case, the at least one processor may include a general-purpose processor (such as a central processing unit (CPU), an access point (AP) or a digital signal processor (DSP)), a graphics processor (such as a graphics processing unit (GPU) or a vision processing unit (VPU)) or an AI processor (such as a neural processing unit (NPU)). At least one processor controls the input data to be processed according to predefined operating rules or an AI model stored in a memory. Alternatively, when the at least one processor is an AI processor, the AI ​​processor may be designed to have a hardware structure dedicated to processing a specific AI model.

[0215] Creating predefined operating rules or AI models through learning and training. When creating predefined operating rules or AI models through learning, it means that when a basic AI model is trained using multiple training data using a learning algorithm, a set of predefined operating rules or AI models for achieving desired characteristics (or purposes) is created. This learning can be performed by the device itself using the AI ​​according to the present disclosure, or it can be performed by a separate server and / or system. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning.

[0216] The AI ​​model may include multiple neural network layers. The multiple neural network layers may respectively have multiple weight values, and each layer performs a neural network operation by calculating the multiple weight values ​​and the calculation results of the previous layer. The multiple weight values ​​of the multiple neural network layers can be optimized by the training results of the AI ​​model. For example, during the learning process, the multiple weight values ​​can be refined to reduce or minimize the loss value or cost value obtained by the AI ​​model. The artificial neural network may include a deep neural network (neural network, DNN), and examples of artificial neural networks may include but are not limited to convolutional neural networks (CNN), DNN, recurrent neural networks (RNN), restricted Boltzmann machines (RBM), deep belief networks (DBN), bidirectional recurrent deep neural networks (BRDNN) and deep Q networks.

[0217] Some embodiments of the present disclosure may be implemented as a recording medium including computer-executable instructions, such as a computer-executable program module. Computer-readable media can be any available media that is accessible to a computer, and examples thereof include all volatile and non-volatile media and detachable and non-detachable media. In addition, examples of computer-readable media may include computer storage media and communication media. Examples of computer storage media include all volatile and non-volatile media and detachable and non-detachable media for storing information such as computer-readable instructions, data structures, program modules, or other data, which are implemented by any method or technology. Communication media typically include computer-readable instructions, data structures, program modules, or other data that modulates a data signal.

[0218] In addition, the term "unit" used herein may be a hardware component such as a processor or a circuit and / or a software component running in a hardware component such as a processor.

[0219] Throughout the disclosure, the expression "at least one of a, b, or c" means only a, only b, only c, both a and b, both a and c, both b and c, all three of a, b, and c, or variations thereof.

[0220] Although the present disclosure has been specifically shown and described with reference to the embodiments thereof, it should be understood that the embodiments thereof are merely illustrative of the present disclosure and that various changes in form and detail may be made without departing from the spirit and scope of the present disclosure. It should be understood that the embodiments of the present disclosure are to be considered in a descriptive sense only and not for purposes of limitation. For example, each component described as a single type may be run in a distributed manner, and components described in a distributed form may also be run in an integrated form.

[0221] The scope of the present disclosure is defined not by the detailed description of the present disclosure but by the claims, and all modifications or substitutions derived from the scope and spirit of the claims and their equivalents fall within the scope of the present disclosure.

Claims

1. A method for a server to modify a speech recognition result provided by a slave device, the method comprising: receiving, from the device, output text from an automatic speech recognition (ASR) model of the device; receiving, from the device, a first domain reliability of the output text calculated by the device, wherein the first domain reliability is a value indicating the extent to which at least a portion of the output text is related to a first domain; calculating a second domain reliability of the output text, wherein the second domain reliability is a value indicating the extent to which at least a portion of the output text is related to a second domain; identifying at least one domain related to a subject of the output text based on a weighted sum of the first domain reliability and the second domain reliability; selecting at least one text modification model for the at least one domain from a plurality of text modification models included in the server, wherein the at least one text modification model is an artificial intelligence (AI) model trained to analyze text related to the topic; modifying the output text using the at least one text modification model to generate modified text; and The modified text is provided to the device.

2. The method according to claim 1, wherein The at least one text modification model is an AI model trained by using output text from the ASR model and real text of a preset domain.

3. The method according to claim 2, wherein: The at least one text modification model is an AI model trained to analyze text related to the topic by using a plurality of texts output from different ASR models.

4. The method according to claim 1, wherein The multiple text modification models correspond to multiple domains respectively, and The selecting of the at least one text modification model includes selecting the at least one text modification model corresponding to the at least one domain from the plurality of text modification models corresponding to the plurality of domains.

5. The method according to claim 1, wherein Receiving the text includes: receiving a text stream output from the ASR model, and Here, identifying the at least one domain includes identifying the at least one domain by accumulating the text flow in units of sections.

6. The method according to claim 5, wherein: Identifying the at least one domain includes identifying the at least one domain by calculating domain reliability of the text flow accumulated in units of sections.

7. The method according to claim 1, wherein Identifying the at least one domain includes classifying the output text into a plurality of sections and identifying a plurality of domains respectively associated with the plurality of sections. wherein selecting the at least one text modification model comprises selecting the plurality of text modification models corresponding to the plurality of domains, and Wherein, modifying the output text includes respectively modifying the plurality of texts in the plurality of sections by using the plurality of text modification models.

8. The method of claim 1 , further comprising selecting a domain identification module from a plurality of domain identification modules in the server based on the first domain reliability, in, Calculating the second domain reliability includes calculating the second domain reliability of the output text by using the domain identification module.

9. The method according to claim 1, wherein The ASR model of the device is an end-to-end ASR model, and Wherein, each of the multiple text modification models included in the server is a sequence-to-sequence model for speech recognition.

10. A server for modifying a speech recognition result provided by a slave device, the server comprising: Communication interface; a memory storing one or more instructions; and a processor configured to execute the one or more instructions to control the server to: receiving, from the device, output text from an automatic speech recognition (ASR) model of the device, receiving, from the device, a first domain reliability of the output text calculated by the device, wherein the first domain reliability is a value indicating the degree to which at least a portion of the output text is related to a first domain, calculating a second domain reliability of the output text, wherein the second domain reliability is a value indicating the degree to which at least a portion of the output text is related to a second domain, identifying at least one domain related to a subject of the output text based on a weighted sum of the first domain reliability and the second domain reliability, selecting at least one text modification model for the at least one domain from a plurality of text modification models included in the server, wherein the at least one text modification model is an artificial intelligence (AI) model trained to analyze text related to the topic, modifying the output text using the at least one text modification model to generate modified text, and The modified text is provided to the device.

11. The server according to claim 10, wherein: The at least one text modification model is an AI model trained by using output text from the ASR model and real text of a preset domain.

12. The server according to claim 11, wherein: The at least one text modification model is an AI model trained to analyze text related to the topic by using a plurality of texts output from different ASR models.

13. The server according to claim 10, wherein: The multiple text modification models correspond to multiple domains respectively, and The processor is further configured to execute the one or more instructions to select the at least one text modification model corresponding to the at least one domain from the plurality of text modification models corresponding to the plurality of domains.

14. The server according to claim 10, wherein: The processor is further configured to execute the one or more instructions to receive a text stream output from the ASR model and identify the at least one domain by accumulating and using the text stream in units of sections.

Citation Information

Patent Citations

  • Speech recognition method and device, computer readable storage medium and computer device

    CN108711422A

  • User recognition for speech processing systems

    US10032451B1

  • Device, Method, and Program for Performing Interaction Between User and Machine

    US20100131277A1

  • Methods and apparatus for hybrid speech recognition processing

    US20180197545A1