Methods executed by computers, server devices, information processing systems, programs, and client terminals

JP7915010B2Active Publication Date: 2026-09-03SOUNDHOUND INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021171822
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-10-20
Publication Date
2026-09-03
Estimated Expiration
2039-10-23

Smart Images

  • Figure 0007915010000001
    Figure 0007915010000001
  • Figure 0007915010000002
    Figure 0007915010000002
  • Figure 0007915010000003
    Figure 0007915010000003
Patent Text Reader

Abstract

To provide a technology for enabling a mobile terminal to function as a digital assistant even when the mobile terminal is in a state where it cannot communicate with a server device. [Solution] In response to query A input from a user, a user terminal 200 transmits query A to a server 100. The server 100 interprets the meaning of query A using grammar A. The server 100 obtains a response to query A based on the meaning of query A and transmits the response to the user terminal 200. The server 100 further transmits grammar A to the user terminal 200. In other words, the server 100 transmits the grammar used to interpret the query received from the user terminal 200 to the user terminal 200.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to natural language interpretation, and in particular to grammar management for natural language interpretation.

Background Art

[0002] Conventionally, digital assistants that operate in response to queries input by users have been used. When a user uses a digital assistant, the user inputs a query to a client terminal such as a smartphone. In typical use of a digital assistant, the client terminal transmits the query to a server device. The server device determines the meaning of the query by performing utterance interpretation and natural language interpretation on the query. Then, the server device retrieves or generates a response to the query in a database corresponding to the determined meaning, and / or transmits the query to an API (Application Programming Interface) corresponding to the determined meaning, thereby obtaining a response to the query. The server device obtai ned response is transmitted to the client terminal. The client terminal outputs the response. That is, the client terminal behaves as a part of the digital assistant by communicating with the server device.

[0003] Paragraph

[0037] and the like of US Patent Application Publication No. 2007 / 0276651 (Patent Document 1) discloses a system in which a mobile terminal performs natural language interpretation. In this system, when a mobile terminal receives an utterance from a user, the mobile terminal attempts to perform natural language interpretation on the utterance. If the mobile terminal fails in the natural language interpretation, it requests a server device to execute natural language interpretation on the utterance.

[0004] Users may want to use a digital assistant in locations where their mobile device cannot communicate with the server (for example, inside a tunnel). It is necessary to enable the mobile device to function as a digital assistant even when it cannot communicate with the server. [Prior art documents] [Patent Documents]

[0005] [Patent Document 1] U.S. Patent Application Publication No. 2007 / 0276651 [Overview of the Initiative]

[0006] This disclosure provides a technical solution to the aforementioned problems of conventional systems by enabling mobile devices to function as digital assistants even when they cannot access server equipment.

[0007] In accordance with certain aspects of this disclosure, a method is provided which is performed by a computer, comprising the steps of: receiving a query input from a client terminal; performing natural language interpretation of the query using a grammar; outputting a response to the query after performing natural language interpretation; and sending the grammar to the client terminal.

[0008] The method may further include a step of determining whether the client terminal does not store the grammar before sending the grammar to the client terminal. The step of sending the grammar to the client terminal may be performed on the condition that the client terminal does not store the grammar.

[0009] The method may further include a step of determining whether the client terminal is configured to perform functions using the grammar in an offline state, without communicating with the computer, before sending the grammar to the client terminal. The step of sending the grammar to the client terminal may be performed if it is determined that the client terminal is configured to perform functions using the grammar in an offline state.

[0010] The step of sending a grammar to a client terminal may also include sending other grammars belonging to the domain to which the grammar belongs, along with the grammar itself, to the client terminal.

[0011] The method may further include a step of counting the number of times the grammar has been used for natural language interpretation of queries from a client terminal. The step of sending the grammar to the client terminal may be performed only if the counted number exceeds a threshold.

[0012] Counting may include counting the number of times all grammars belonging to the domain to which a grammar belongs have been used in the natural language interpretation of a query.

[0013] The method may further include the steps of predicting the type of data required to respond to future queries based on the input query, and sending the type of data to the client terminal.

[0014] The step of sending data of a certain type may also include sending an expiration date for the data of that type.

[0015] The step of receiving query input from a client terminal may include receiving voice input from the client terminal. The method may further include the steps of training a speech recognition model using the user's utterances to suit the user of the client terminal, and sending the trained speech recognition model to the client terminal.

[0016] In accordance with other aspects of this disclosure, a server device is provided which has one or more processors and further comprises a storage device that stores a program, which is executed by the one or more processors, causing the server device to perform the above method.

[0017] In accordance with yet another aspect of this disclosure, an information processing system is provided, comprising a client terminal and a server device that sends responses to queries entered from the client terminal to the client terminal, wherein the server device includes one or more processors that perform natural language interpretation of queries using a grammar, and the one or more processors send the grammar to the client terminal.

[0018] A method is provided, performed by a computer, comprising: sending a first query to a server device; receiving a grammar used for natural language interpretation of the first query from the server device; storing the received grammar in memory; receiving input for a second query; and, when the computer is not connected to the server device, performing natural language interpretation of the second query using the grammar.

[0019] The method includes the steps of: accepting input for a third query; performing natural language interpretation of the third query when the computer is not connected to a server device; determining that the natural language interpretation of the third query has failed; storing the third query in memory; and, depending on the failure, when the computer is connected to a server device, the third query The system may further include the step of sending the ERI to the server device.

[0020] The method may further include the steps of receiving data related to a first query from a server device, storing the data related to the first query in memory, and obtaining a response to a second query using the data related to the first query.

[0021] Data associated with the first query may include metadata indicating an expiration date. The method may further comprise the step of deleting, from memory, data associated with the first query after the expiration date has passed.

[0022] The method may further comprise the step of acquiring position information of the computer. The step of performing natural language interpretation on the second query may comprise selecting, from one or more grammars stored in memory, a grammar to be used based on the position information.

[0023] The method may further comprise the step of acquiring time information indicating a time at which the second query is input. The step of performing natural language interpretation on the second query may comprise selecting, from one or more grammars stored in memory, a grammar to be used based on the time information.

[0024] The step of receiving an input of the second query may comprise receiving an audio input. The method may further comprise: the step of receiving, from a server device, a speech recognition model trained to be adapted for a user of the computer; and the step of performing speech recognition on the input audio using the speech recognition model when the computer is not connected to the server device.

[0025] According to still another aspect of the present disclosure, there is provided a method executed by a computer, the method comprising: when the computer is connected to a server device, the step of receiving an input of a first query; the step of transmitting the first query to the server device; the step of receiving a response to the first query from the server device; when the computer is not connected to the server device, the step of receiving an input of a second query; the step of storing the second query in memory together with time information indicating a time at which the second query is input; and when the computer is connected to the server device, the step of transmitting the second query in the memory to the server device together with the time information.

[0026] Storing the second query in a memory may include storing, in the memory together with the second query, positional information of the computer obtained when an input of the second query is received. Transmitting the second query to a server device may include transmitting the positional information to the server device together with the second query.

[0027] According to still another aspect of the present disclosure, there is provided a computer program that, when executed by one or more processors of a client terminal, causes the client terminal to perform the above method.

[0028] According to still another aspect of the present disclosure, there is provided a client terminal including one or more processors, the client terminal comprising a memory that stores a program which, when executed by the one or more processors, causes the client terminal to perform the above method. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] [Figure 1] It is a diagram illustrating a schematic configuration of a query processing system. [Figure 2] It is a diagram illustrating an implementation example of query processing in the query processing system. [Figure 3] It is a diagram illustrating a specific example of an aspect for query processing using a user terminal. [Figure 4] It is a diagram illustrating a specific example of an aspect for query processing using a user terminal. [Figure 5] It is a diagram illustrating a specific example of an aspect for query processing using a user terminal. [Figure 6] It is a diagram illustrating a specific example of an aspect for query processing using a user terminal. [Figure 7] It is a diagram illustrating a specific example of an aspect for query processing using a user terminal. [Figure 8] It is a diagram illustrating a specific example of an aspect for query processing using a user terminal. [Figure 9] It is a diagram illustrating a specific example of an aspect for query processing using a user terminal. [Figure 10] This is a diagram showing the server's hardware configuration. [Figure 11] This diagram shows the data structure of the grammar library. [Figure 12] This diagram shows the data structure of user information. [Figure 13] This diagram shows the hardware configuration of the user terminal. [Figure 14] This figure shows an example of a data structure for location information. [Figure 15] This is a flowchart of the processes performed on the server to output a response to a query. [Figure 16] This is a flowchart of the processes performed on the server to output a response to a query. [Figure 17] This is a flowchart of the processes executed on the user's terminal. [Figure 18] This is a flowchart of the processes executed on the user's terminal. [Modes for carrying out the invention]

[0030] An embodiment of the information processing system will be described below with reference to the drawings. In the following description, the same parts and components are denoted by the same reference numerals. Their names and functions are also the same. Therefore, these descriptions will not be repeated.

[0031] 1. Overview of the query processing system Figure 1 is a diagram illustrating the configuration of a query processing system. The query processing system includes a server and user terminals. In Figure 1, the server is shown as "Server 100," and the user terminals are shown as "User Terminals 200A to 200G" depending on the context in which they are used.

[0032] Figure 2 illustrates an example of query processing in a query processing system. In Figure 2, all types of user terminals are shown as "User Terminal 200". A user terminal is an example of a client terminal.

[0033] As shown as step (1) in Figure 2, the user terminal 200 sends query A to the server 100 in response to the user's input of query A. The user terminal 200 may receive the input of query A as voice via a microphone, as text data via a keyboard or touch panel, or as an image or video representing an object or gesture via a camera.

[0034] As shown in step (2), the server 100 interprets the meaning of query A using grammar A.

[0035] As shown in step (3), the server 100 generates a response to query A based on the meaning of query A and sends the response to the user terminal 200.

[0036] In the query processing system shown in Figure 2, as indicated by step (4), the server 100 further sends grammar A to the user terminal 200. That is, the server 100 sends the grammar used to interpret the query received from the user terminal 200 to the user terminal 200.

[0037] User terminal 200 records grammar A received from server 100 in its memo. The data is stored in the log. When a query is entered by the user terminal 200 while it is offline, the terminal interprets the meaning of the query using grammar A (step (1X)), generates a response to the query based on that meaning (step (2X)), and displays (and / or outputs audibly) the response (step (3X)).

[0038] 2. Specific examples of responses to queries Figures 3 to 9 are diagrams illustrating specific examples of scenarios for processing queries using the user terminal 200.

[0039] 2-1. Figure 3 Figure 3 shows an example of a user terminal 200, specifically user terminal 200A used in a car. This is shown. User terminal 200A is, for example, an information processing terminal installed in a car.

[0040] The user inputs the utterance "Turn on the radio!" as a query into the user terminal 200A. The user terminal 200A sends the utterance "Turn on the radio!" as a query to the server 100. In one implementation example, the user may input the above query after pressing a given button on the user terminal 200A. The user terminal 200A may accept the query input within a given time after the button is pressed and send the input query to the server 100.

[0041] Server 100 interprets the meaning of the query input from user terminal 200A using a grammar for functions that operate on the elements of the automobile, and obtains a response to the query based on that meaning. Then, as a response to the query input from user terminal 200A, server 100 sends a control signal to user terminal 200A to turn on the radio. In response, user terminal 200A turns on the radio of the automobile in which it is installed.

[0042] Server 100 may also send an instruction to user terminal 200A to output the voice message "Turning on the radio" as a response to a query input from user terminal 200A. User terminal 200A may output the voice message "Turning on the radio" upon receiving the instruction.

[0043] 2-2. Figure 4 Figure 4 shows a user terminal 200B, used in an operating room, as an example of a user terminal 200. The user terminal 200B is, for example, an information processing terminal that can be attached to the head of a physician, who is the user.

[0044] The user inputs the utterance "Show me the medical record!" as a query into user terminal 200B. User terminal 200B sends the utterance "Show me the medical record!" as a query to server 100. In one implementation example, the user may say a predetermined message (for example, "OK!") to initiate query input before inputting the above query. User terminal 200B may accept the query input upon receiving the above message and send the entered query to server 100.

[0045] Server 100 interprets the meaning of a query entered from user terminal 200B using a grammar for the function of providing information to the user in the operating room, and generates a response to the query based on that meaning. Server 100 may generate a response based only on processing performed within Server 100. Alternatively or additionally, Server 100 may generate a response by obtaining data from an external service provider or website. Then, as a response to the query entered from user terminal 200B, Server 100 The medical record of a patient in the operating room is transmitted to the user terminal 200B. In response, the user terminal 200B displays the medical record on the display to which it is connected. The server 100 may also transmit the medical record directly to the display (or the computer to which the display is connected).

[0046] Server 100 may also send an instruction to user terminal 200B to output the voice message "This is Mr. Yamada's (example of a patient name) medical record." in response to a query input from user terminal 200B. User terminal 200B may output the voice message "This is Mr. Yamada's (example of a patient name) medical record." upon receiving the instruction.

[0047] 2-3. Figure 5 Figure 5 shows an example of a user terminal 200, specifically user terminal 200C, which is used in an office. User terminal 200C is, for example, a smartphone.

[0048] The user inputs the utterance "Check Company A!" as a query into user terminal 200C. User terminal 200C then sends the utterance "Check Company A!" as a query to server 100.

[0049] Server 100 interprets the meaning of the query input from user terminal 200C using a grammar for a function that provides information on stock prices, and generates a response to the query based on that meaning. Then, as a response to the query input from user terminal 200C, server 100 sends the stock price of company A to user terminal 200C. In response, user terminal 200C displays the stock price of company A on its display and / or outputs the stock price of company A audibly from the speaker of user terminal 200C.

[0050] 2-4. Figure 6 Figure 6 shows an example of a user terminal 200, specifically a user terminal 200D used in a home environment. The user terminal 200D is, for example, a smart speaker.

[0051] The user inputs the utterance "Call Grandma!" as a query into user terminal 200D. User terminal 200D then sends the utterance "Call Grandma!" as a query to server 100.

[0052] Server 100 interprets the meaning of the query input from user terminal 200D using a grammar for the call function and generates a response to the query based on that meaning. Then, as a response to the query input from user terminal 200D, server 100 sends an instruction to user terminal 200D to call the phone number registered as "Grandma" on user terminal 200D. In response, user terminal 200D makes a call to the phone number registered as "Grandma" on user terminal 200D.

[0053] 2-5. Figure 7 Figure 7 shows an example of a user terminal 200, specifically a user terminal 200E used in a kitchen. The user terminal 200E is, for example, a smartphone.

[0054] The user inputs the utterance "Tell me the recipe for pot-au-feu!" as a query into user terminal 200E. User terminal 200E then sends the utterance "Tell me the recipe for pot-au-feu!" as a query to server 100.

[0055] Server 100 interprets the meaning of queries entered from user terminal 200E using a grammar for providing information about cooking, and generates a response to the query based on that meaning. The server 100 then sends a pot-au-feu recipe to the user terminal 200E as a response to the query entered from the user terminal 200E. In response, the user terminal 200E displays the pot-au-feu recipe on the display to which the user terminal 200E is connected. Alternatively, the server 100 may send the user terminal 200E a list of links to websites that provide pot-au-feu recipes as a response to the above query. In this case, the user terminal 200E displays the list. Depending on whether the user has selected a link from the list, the user terminal 200E connects to the selected link.

[0056] 2-6. Figure 8 Figure 8 shows an example of a user terminal 200, specifically user terminal 200F, which is used by a user sitting in front of a television. User terminal 200F is, for example, a smartphone.

[0057] The user inputs the utterance "What's on TV tonight?" as a query into user terminal 200F. User terminal 200F then sends the utterance "What's on TV tonight?" as a query to server 100.

[0058] Server 100 interprets the meaning of the query input from user terminal 200F using a grammar for a function that provides information about television programs, and generates a response to the query based on that meaning. Then, as a response to the query input from user terminal 200F, server 100 sends the nighttime television program schedule for the day the query was entered to user terminal 200F. In response, user terminal 200F displays the television program schedule sent from server 100.

[0059] 2-7. Figure 9 Figure 9 shows user terminal 200G as an example of user terminal 200. User terminal 200G is, for example, a smartphone.

[0060] The user inputs the utterance "What's the weather like today?" as a query into user terminal 200G. User terminal 200G sends the utterance "What's the weather like today?" and the location information of the smartphone (user terminal 200G) to server 100 as a query.

[0061] Server 100 interprets the meaning of the query input from user terminal 200G using a grammar for a function that provides weather information, and generates a response to the query based on that meaning. Then, as a response to the query input from user terminal 200G, server 100 sends to user terminal 200G the weather forecast for the location where the query was entered on the day the query was entered. In response, user terminal 200G outputs the weather forecast sent from server 100 on display and / or by voice.

[0062] 3. Hardware configuration (Server 100) Figure 10 shows the hardware configuration of server 100.

[0063] Referring to Figure 10, the server 100 includes, as its main hardware elements, a processing unit 11, memory 12, input / output (I / O) interface 14, network controller 15, and storage 16.

[0064] The processing unit 11 is a computing entity that performs the processing necessary for realizing the server 100 by executing various programs as described later. The processing unit 11 consists of, for example, one or more CPUs (Central Processing Units) and / or GPUs. This is a Graphics Processing Unit. The processing unit 11 has multiple cores. It may be a CPU or GPU. The processing unit 11 may be an NPU (Neural Network Processing Unit) suitable for training processing to generate a trained model.

[0065] Memory 12 provides a storage area for temporarily storing program code, work memory, and other data when the processing unit 11 executes a program. Memory 12 may be a volatile memory device such as DRAM (Dynamic Random Access Memory) or SRAM (Static Random Access Memory).

[0066] The processing unit 11 can accept data input from devices (keyboard, mouse, etc.) connected via the I / O interface 14, and can also output data to devices (display, speaker, etc.) via the I / O interface 14.

[0067] The network controller 15 sends and receives data to and from any information processing device, including the user terminal 200, via public lines and / or LAN (Local Area Network). The network controller 15 may be, for example, a network interface card. The server 100 uses the network controller 15 to obtain responses to queries from external Web APIs (Application Programming Interfaces). A request may be sent to the network controller 15. The network controller 15 may support any communication method, such as Ethernet®, wireless LAN, or Bluetooth®.

[0068] The storage 16 may be, for example, a hard disk drive or a non-volatile memory device such as an SSD (Solid State Drive). The single unit 11 stores the learning program 16A, the preprocessing program 16B, the application program 16C, and the OS (Operating System) 16D, which are executed in the single unit 11.

[0069] The processing unit 11 executes a learning program 16A, a preprocessing program 16B, an application program 16C, and an OS (operating system) 16D. In this implementation example, the server 100 may receive voice data from different users and may use the voice data from each user to build a trained model 16G for each user. The trained model 16G may be downloaded to the user's terminal 200 so that the user terminal 200 can locally perform speech recognition of the user's utterances.

[0070] To that end, the server 100 may have a training program 16A for training a pre-trained model 16G used for speech recognition of queries. In the implementation example, the training program 16A may have the structure of a neural network. The pre-processing program 16B is a program for generating a training dataset 16H for each user by collecting and pre-processing the speech data input from each user for training the pre-trained model 16G. If only speech data collected from a specific user is used for training the pre-trained model 16G, the pre-trained model 16G can be trained individually for that specific user. The application program 16C is a program for sending a response to a query to the client terminal 200 in response to the query input from the client terminal 200. The OS 16D is the basic software program for processing on the server 100.

[0071] Storage 16 further contains the grammar library 16E, user information 16F, and trained models. It stores 16GB, a training dataset 16H, and audio data 16X.

[0072] Grammar library 16E stores information about grammar used to interpret the meaning of queries. The data structure of grammar library 16E will be described later with reference to Figure 11.

[0073] User information 16F stores information about each user registered in the query processing system. The data structure of user information 16F will be described later with reference to Figure 12.

[0074] The trained model 16G is used for query speech recognition, as described above. The training dataset 16H is the dataset used to train the trained model 16G. In the training dataset 16H, each dataset may be tagged with the user who uttered the corresponding speech, the phonetic transcription of the word or phrase the user intended to utter, the user's characteristics (age, gender, occupation, etc.), and / or the circumstances in which the user uttered the corresponding speech (location, time, etc.).

[0075] 4. Data structure of Grammar Library 16E Figure 11 shows the data structure of grammar library 16E.

[0076] Grammar Library 16E stores grammars used to interpret the meaning of queries, as well as information related to each grammar.

[0077] The information associated with each grammar includes categories (domains) for classifying grammars. Figure 11 shows domains A through G. All domains in Figure 11 contain multiple grammars, but some domains may contain only one grammar.

[0078] Domain A includes grammars A1, A2, A3, etc. Grammar A1 defines the word combination "turn on the radio" ("radio", "o", "turn on"). Grammars A2 and A3 define the word combinations "close the window" and "open the window," respectively. The grammars belonging to Domain A are mainly used to interpret queries entered inside the vehicle in order to realize the function of manipulating elements of the automobile.

[0079] Domain B includes grammar B1, etc. Grammar B1 defines the word combination "Show me the medical record." Grammar belonging to Domain B is mainly used to interpret queries entered in the operating room in order to realize the function of providing information to the user in the operating room.

[0080] Domain C includes grammar C1, etc. Grammar C1 is a slot representing the company name (in Figure 11).<name of company> ), and the combination of the words "check". The grammar belonging to Domain C is primarily used to interpret queries entered in the office, in order to implement a function that provides information about stock prices.

[0081] Domain D includes grammars D1, D2, etc. Grammar D1 is a slot representing a name registered in the address book (in Figure 11). <name>Grammar D2 defines the combination of words, and "call me". Grammar D2 is a slot that represents the name of a musician (in Figure 11). <musician>), the word "no", a slot representing the song title (in Figure 11) <title> ), and< / title> This defines the word combination "wo kakete". The grammar belonging to Domain D is primarily used to interpret queries entered at home or in the office in order to realize general information provision and communication functions.

[0082] Domain E includes grammar E1, etc. Grammar E1 is a slot that represents the name of a dish (in Figure 11). <dish>), and the combination of the words "tell me the recipe". Domain E The grammar belonging to this category is primarily used to interpret queries entered in the kitchen in order to implement a function that provides information about cooking.

[0083] Domain F includes grammar F1, etc. Grammar F1 is a slot that represents time or date (in Figure 11). <time>), and the combination of the words "What's on TV?". The grammar belonging to Domain F is primarily used to interpret queries entered in front of a home television in order to realize the function of providing information about television programs.

[0084] Domain G includes grammars G1, G2, etc. Grammar G1 defines the word combination "Tell me the weather today." Grammar G2 is a slot that represents the name of a city (in Figure 11). <city>), and the combination of the words "Tell me the weather for today." The grammar belonging to Domain G is mainly used to interpret queries entered by users asking for weather forecasts in order to implement the function of providing weather-related information.

[0085] The "Offline Settings" item in Grammar Library 16E specifies whether functions that utilize grammars belonging to each domain can be used while the user terminal 200 is offline. An example of a user terminal 200 being offline is when it is unable to communicate with the server 100 because it is not connected to a Wi-Fi network. A value of "ON" indicates that the function can be used even while the user terminal 200 is offline. A value of "OFF" indicates that the function cannot be used while the user terminal 200 is offline. It represents that.

[0086] The "Expected Data Type" entry in Grammar Library 16E defines the type of data that a user terminal is expected to need in the future when Server 100 receives a query from that terminal.

[0087] For example, when server 100 receives the query "Play The Beatles' Yesterday!", it will use grammar D2(" <musician>of <title> The meaning of the query is interpreted using ")< / title> The value for the "Expected data type" item in grammar D2 is, <musician>This is a list of song titles.

[0088] Of the queries "Play The Beatles' Yesterday!", "The Beatles" is a slot <musician>In response to this, "Yesterday" is a slot <title> This corresponds to one implementation example.< / title> When the above query "Play Yesterday by The Beatles!" is entered into Server 100, it may further identify "List of Beatles song titles" as the "Expected data type". Server 100 may then retrieve the identified type of data, namely the list of Beatles song titles, as related data. In addition to the response to the query "Play Yesterday by The Beatles!", Server 100 may also send the list of Beatles song titles as related data to the query "Play Yesterday by The Beatles!" to the user terminal that sent the query "Play Yesterday by The Beatles!".

[0089] The "Expiration Date" item in grammar library 16E defines the expiration date assigned to the associated data. The expiration date is primarily used on user terminal 200 and defines the period (such as days) during which the query is maintained in the memory of user terminal 200.

[0090] For example, in Figure 11, the predicted data for grammar G1, "7-day weather forecast for the current location," has an expiration date of "7 days." In this case, when the server 100 receives the query "Tell me today's weather" from a user terminal 200 located in Osaka on October 1st, it sends the weather forecast for Osaka on October 1st as a response to the query, and further sends related data. The system transmits the weather forecast for Osaka for the seven days from October 2nd to October 8th. This related data is given an expiration date of "7 days". User terminal 200 retains the related data transmitted from server 100 for seven days after October 1st, i.e., until October 8th. User terminal 200 can use the related data to output responses to queries entered by the user while user terminal 200 is offline. After the expiration date, i.e., after October 8th, user terminal 200 deletes the related data from it.

[0091] The "Count (1)" item in Grammar Library 16E represents the number of times each grammar has been used to interpret the meaning of a query. The "Count (2)" item represents the total number of times grammars belonging to each domain have been used to interpret the meaning of a query. For example, when grammar A1 is used to interpret the meaning of a query, the Count (1) value for grammar A1 increases by 1, and the Count (2) value for domain A also increases by 1.

[0092] 5. User Information 16F Data Structure Figure 12 shows the data structure of user information 16F. User information 16F associates "User ID," "Terminal ID," and "Sent Grammar." "Sent Grammar" defines the name of the domain to which the grammar sent to each terminal belongs.

[0093] The User ID specifies a value assigned to each user. The Terminal ID represents a value assigned to each user terminal 200. The Sent Grammar represents the domain to which the grammar sent to each user terminal 200 belongs.

[0094] In the example in Figure 12, terminal ID "SP01" is associated with user ID "0001" and domains A and C. This means that a user assigned user ID "0001" sent one or more queries to server 100 using the terminal assigned terminal ID "SP01", and that server 100 sent grammars belonging to domains A and C to the terminal assigned terminal ID "SP01".

[0095] In one implementation example, the server 100 may send only a portion of the domain's grammar to each user terminal 200. In this case, the user information 16F may specify the grammar sent to each user terminal 200 (grammar A1, etc.) as the value of "Sent Grammar".

[0096] 6. Hardware configuration (User terminal 200) Figure 13 shows the hardware configuration of user terminal 200.

[0097] Referring to Figure 13, the user terminal 200 has the following main hardware elements: CPU 201, display 202, microphone 203, speaker 204, GPS (Global Positioning System) receiver 205, and communication interface (I / F) 206. It includes storage 207 and memory 210.

[0098] CPU201 is the computing unit that executes various programs to perform the processing necessary for realizing the user terminal 200.

[0099] The display 202 may be, for example, a liquid crystal display device. The CPU 201 may display the results of its processing on the display 202.

[0100] The microphone 203 receives audio input and outputs a signal corresponding to the received audio to the storage 207 for access by the CPU 201. The speaker 204 outputs audio. The CPU 201 outputs the result of processing as audio from the speaker 204. You may use force.

[0101] The GPS receiver 205 receives signals from GPS satellites and outputs these signals to the storage 207 for access by the CPU 201. The CPU 201 may determine the current location of the user terminal 200 based on the signals from the GPS receiver 205.

[0102] The communication interface 206 transmits and receives data to and from any information processing device, including the server 100, via public lines and / or LANs. The communication interface 206 may be, for example, a mobile network interface.

[0103] Storage 207 may be, for example, a hard disk drive or a non-volatile memory device such as an SSD (Solid State Drive). Storage 207 stores the grammar area 2072, the related data area 2073, the failure data area 2074, the location information 2075, and the trained model 2076.

[0104] Application program 2071 is a program that receives query input from the user and outputs a response to that query. Application program 2071 may be, for example, a car navigation program or an assistant program.

[0105] The grammar area 2072 is an area for storing grammar sent from the server 100. The related data area 2073 is an area for storing related data sent from the server 100. The failure data area 2074 is an area for storing the query in the case where the interpretation of the query's meaning is performed on the user terminal 200 and the interpretation fails.

[0106] Location information 2075 represents the location of the user terminal 200 and is a region for storing information that can be used to select the type of grammar used to interpret the meaning of the query. The data structure of location information 2075 is described later with reference to Figure 14. The trained model 16G is sent from the server 100 and stored in storage 207 as trained model 2076.

[0107] 7. Data structure of location information 2075 Figure 14 shows an example of the data structure of location information 2075. Location information 2075 associates the location "home" with grammar "domains D, E, F, G", and the location "office" with grammar "domain C".

[0108] The user terminal 200 may determine the grammar to use in interpreting the meaning of offline queries based on the information shown in Figure 14. For example, the user terminal 200 may use different grammars depending on its location.

[0109] The user terminal 200 may determine its own location in response to the query input. In one example, the location of the user terminal 200 may be determined based on GPS data received by the GPS receiver 205. In another example, the location of the user terminal 200 may be determined based on the type of beacon signal received by the user terminal 200. In yet another example, the location of the user terminal 200 may be determined based on network information such as an IP address or a mobile phone base station ID.

[0110] In one implementation example, if the identified location is a place pre-registered as "home," the user terminal 200 attempts to interpret the query using the grammar belonging to domains D, E, F, and G, and other It is not necessary to attempt to interpret the query using a grammar belonging to a particular domain. If the specified location is a location pre-registered as an "office," the user terminal 200 may attempt to interpret the query using a grammar belonging to domain C, and may not attempt to interpret the query using a grammar belonging to another domain.

[0111] 8. Processing on Server 100 Figures 15 and 16 are flowcharts of the processes performed by the server 100 to output a response to a query. In one implementation, the processes in Figures 15 and 16 are realized by the processing unit 11 executing the application program 16C. In one implementation, the server 100 starts the processes in Figures 15 and 16 when it receives data from the user terminal 200 declaring the submission of a query.

[0112] First, referring to Figure 15, in step S100, the server 100 receives a query from the user terminal 200.

[0113] In step S102, the server 100 generates a transcript of the query by performing speech recognition on the query sent from the user terminal 200. If the query is sent from the user terminal 200 in a format other than speech, step S102 may be omitted. The user can input the query in text format into the user terminal 200. The user terminal 200 can send the query to the server 100 in text format. If the query sent from the user terminal 200 is in text format, the server 100 may omit step S102.

[0114] In step S104, the server 100 performs natural language interpretation on the transcription generated in step S102 (or the text data sent from the user terminal 200). This interprets the meaning of the query.

[0115] In step S104, the server 100 may select a grammar from among several grammars that can be used to interpret the meaning of the query, and interpret the meaning of the query using the selected grammar.

[0116] In one implementation example, if a single speech by a user is expected to include multiple transcriptions, the steps S102 and S104 for those multiple transcriptions can be combined.

[0117] In step S106, the server 100 increments the count in the grammar library 16E for the grammars used to interpret the meaning of the query in step S104. More specifically, the server 100 updates the count (1) by incrementing by 1 for the grammars used, and updates the count (2) to which the grammars used belong by incrementing by 1.

[0118] In step S108, the server 100 generates a response to the query based on the interpretation in step S104.

[0119] In one example, server 100 receives an instruction to turn on the radio installed in the car as a response to the query "Turn on the radio". In another example, server 100 may query the stock price of company A by sending at least part of the query (company A) to an API that provides stock prices in order to obtain a response to the query "Check company A". Server 100 obtains the stock price of company A, which it has obtained from the API as a response to that query, as a response to the query "Check company A".

[0120] In yet another example, server 100 receives, in response to the query "Play The Beatles' Yesterday," instructions to search for the audio file of The Beatles' Yesterday and instructions to play the audio file.

[0121] In step S110, the server 100 sends the response obtained in step S108 to the user terminal 200.

[0122] In step S112, the server 100 determines whether the grammar used in step S104 is stored in the user terminal 200. In one implementation example, the server 100 refers to the sent grammars of the user terminal 200, which is the source of the query sent in step S100, in the user information 16F (Figure 12). More specifically, if the grammars belonging to the domain stored as sent grammars include the grammar used in step S104, the server 100 determines that the grammar is stored in the user terminal 200.

[0123] If the server 100 determines that the grammar used in step S104 is stored in the user terminal 200 (YES in step S112), it proceeds to step S120 (Figure 16); otherwise (NO in step S112), it proceeds to step S114.

[0124] There may be implementations that do not include step S112. In such implementations, the process may proceed directly from step S110 to step S114. In such implementations, if the user terminal 200 receives the same grammar multiple times, it ignores (or deletes) copies of the same grammar. Such implementations will utilize more communication bandwidth due to increased network traffic, but they avoid the complexity of accurately managing information on which grammars are stored on the user terminal on the server.

[0125] In step S114, the server 100 determines whether the value of the "online setting" for the grammar used in step S104 is set to ON in the grammar library 16E. If the value of the "online setting" for the grammar used in step S104 is ON (YES in step S114), the server 100 proceeds to step S116; otherwise (NO in step S114), the server 100 proceeds to step S120 (Figure 16).

[0126] In one implementation, the user terminal 200 may be configured to receive only grammars that have been used a given number of times for a given query. In this case, the download of less frequently used grammars can be avoided. Depending on such an implementation, in step S116, the server 100 determines whether the count value associated with the grammar used in step S104 exceeds a given threshold. The "count value" in step S116 may be the value of count(1), the value of count(2), or both of the values ​​of count(1) and count(2) in the grammar library 16E. If the server 100 determines that the count value exceeds a given threshold (YES in step S116), it proceeds to step S118; otherwise (NO in step S116), it proceeds to step S120 (Figure 16).

[0127] In step S118, the server 100 sends the grammar used in step S104 to the user terminal 200. In step S118, the server 100 may also send other grammars belonging to the same domain as the grammar used in step S104 to the user terminal 200. After that, the control proceeds to step S120 (Figure 16).

[0128] Referring to Figure 16, in step S120, the server 100, similar to step S114, determines whether the value of "online setting" for the grammar used in step S104 is set to ON in the grammar library 16E. If the value of "online setting" is ON (YES in step S120), the server 100 proceeds to step S122; otherwise (NO in step S120), it terminates the process.

[0129] In step S122, the server 100 identifies the "predicted data type" in the grammar library 16E that corresponds to the grammar used in step S104.

[0130] If grammar D2 is used in step S104, server 100 identifies the list of song titles by musicians included in the query as the "expected data type". More specifically, if the meaning of the query "Play Yesterday by The Beatles" is interpreted using grammar D2, server 100 identifies the "list of song titles by The Beatles" as the "expected data type".

[0131] If grammar G2 is used in step S104, server 100 identifies the 7-day weather forecast for the city included in the query as the "expected data type". More specifically, if the meaning of the query "Tell me the weather in Osaka" is interpreted using grammar G2, server 100 identifies "the 7-day weather forecast for Osaka starting from the day after the query was entered" as the "expected data type".

[0132] In step S124, server 100 retrieves data of the type identified in step S122 as related data. For example, if "list of Beatles song titles" is identified as the "predicted data type," server 100 retrieves the data for that title list. If "7 days of weather forecast for Osaka starting from the day after the query was entered" is identified as the "predicted data type," server 100 requests the weather forecast API for the 7 days of Osaka weather and retrieves the weather forecast data obtained in response to the request.

[0133] In step S126, the server 100 sends the relevant data acquired in step S124 to the user terminal 200.

[0134] In step S128, the server 100 sends the trained model 16G corresponding to the user of the user terminal 200 to the user terminal 200. In one implementation example, the server 100 identifies the user ID associated with the user terminal 200, which is the communication partner, by referring to the user information 16F (Figure 12), and sends the trained model 16G stored in the storage 16 corresponding to the identified user to the user terminal 200. After that, the server 100 terminates the process.

[0135] As described above with reference to Figures 15 and 16, the server 100 sends the grammar used for natural language interpretation of the query sent from the user terminal 200 back to the user terminal 200.

[0136] As described in relation to step S112, the server 100 may send the grammar to the user terminal 200, provided that the grammar is not stored in the user terminal 200. As described in relation to step S114, the server 100 may send the grammar to the user terminal 200, provided that the user terminal 200 is configured to perform functions using the grammar in an offline state (offline setting value is ON). As described in relation to step S116, the server 100 may determine the number of times the grammar used in step S104 has been used (count (1)), or whether it belongs to the same domain as the grammar. The above grammar may be sent to the user terminal 200 if the number of times the grammar has been used (count (2)) exceeds a given threshold.

[0137] As described in relation to steps S122 to S126, the server 100 may predict the type of data (predicted data type) required to respond to future queries based on the input query, and send the predicted type of data (relevant data) to the user terminal 200.

[0138] In step S126, server 100 may further transmit the expiration date of the associated data. The expiration date may be specified for each “expected data type,” as shown in Figure 11. The expiration date may be transmitted as metadata for the associated data.

[0139] The server 100 may perform processing for training the trained model 16G. Training may be performed for each user. As described in relation to step S128, the server 100 may send the trained model 16G of the user of the user terminal 200 to the user terminal 200.

[0140] In one implementation example, training the trained model 16G on server 100 may utilize speech data of one or more users and corresponding text data as the training dataset 16H. The training data may further include information about each of the one or more users (for example, the name in the "Contacts" file stored on each user's terminal). For training, for example, reference 1 ("Robust i-vector based Adaptation of DNN Acoustic Model for Speech Recognition"),<URL: http: / / www1.icsi.berkeley.edu / ~sparta / 2015_ivector_paper.pdf > ), Reference 2 ("PERSONALIZED SPEECH RECOGNITION ON MOBILE DEVICES",<URL: https: / / arxiv.org / pdf / 1603.03185.pdf> ) Reference 3 ("Speech Recognition Based on Unified Model of Acoustic and Language Aspects of Speech") <url: https: www.ntt-review.jp archive ntttechnical.php?contents="ntr201312fa4.pdf&mode=show_pdf">), and Reference 4 (Integration of Sound and Language) Speech recognition technology based on type learning,<URL: https: / / www.ntt.co.jp / journal / 1309 / files / jn201309022.pdf> The technologies described in ) may be used.

[0141] 9. Processing on user terminal 200 Figures 17 and 18 are flowcharts of the processes executed on the user terminal 200. In one implementation example, the processes in Figures 17 and 18 are realized when the CPU 201 of the user terminal 200 executes a given program. The processes in Figures 17 and 18 are started, for example, at regular intervals.

[0142] Referring to Figure 17, in step S200, the user terminal 200 deletes related data that has expired from the related data stored in the related data area 2073.

[0143] In step S202, the user terminal 200 determines whether it is online (i.e., whether it can communicate with the server 100). If the user terminal 200 determines that it is online (YES in step S202), it proceeds to step S204; otherwise (NO in step S202), it proceeds to step S226 (Figure 18).

[0144] In step S204, the user terminal 200 sends the query stored in the failure data area 2074 (see step S242 below) to the server 100. If the query is associated with time information and / or location information, the time information and / or location information may also be sent to the server 100 in step S204. After sending the query, the user terminal 200 The submitted query may be removed from the failure data area 2074.

[0145] In step S206, the user terminal 200 receives a query. In one example, the query is input by voice via the microphone 203. In another example, the query is input as text data by operating the touch sensor (not shown) of the user terminal 200.

[0146] In step S208, the user terminal 200 sends the query obtained in step S206 to the server 100. The sent query may be received by the server 100 in step S100 (Figure 15).

[0147] In step S210, the user terminal 200 receives a response to the query from the server 100. The response may be sent from the server 100 in step S110 (Figure 15).

[0148] In step S212, the user terminal 200 outputs the response sent from the server 100. An example of the output of the response is an action taken in accordance with the instructions contained in the response. For example, if the response includes the instruction to "turn on the radio," the user terminal 200 turns on the radio installed in the car in which the user terminal 200 is located.

[0149] In step S214, the user terminal 200 receives the grammar sent from the server 100. The grammar may be sent from the server 100 in step S118 (Figure 15).

[0150] In step S216, the user terminal 200 stores the grammar received in step S214 in the grammar area 2072.

[0151] In step S218, the user terminal 200 receives the relevant data. The relevant data may be transmitted from the server 100 in step S126 (Figure 16).

[0152] In step S220, the user terminal 200 stores the relevant data received in step S218 in the relevant data area 2073.

[0153] In step S222, the user terminal 200 receives the trained model 16G. The trained model 16G may be transmitted from the server 100 in step S128 (Figure 16).

[0154] In step S224, the user terminal 200 stores the trained model 16G received in step S222 in the storage 207 as trained model 2076. After that, the user terminal 200 terminates the process.

[0155] The query sent to the server 100 in step S208 may or may not have an effect on at least one of the following: the grammar in step S214, the related data in step S218, and the trained model in step S222.

[0156] In accordance with the aspects of the present technology, when it is determined in step S202 whether the user terminal 200 is offline (not connected to the server 100), the user terminal 200 uses speech recognition to analyze utterances and generate queries. A personalized, trained model 2076 (generated by the server 100 and downloaded to the user terminal 200) may be used for this purpose. Once a query is recognized, the technology may use one or more grammars that have been downloaded and stored locally to interpret the natural language meaning of the query. The technology may apply any given criteria, as described below, to the selection of whether to use a locally stored grammar or a grammar stored on the server 100.

[0157] Referring to Figure 18, in step S226, the user terminal 200 obtains a query in the same manner as in step S206.

[0158] In step S228, the user terminal 200 obtains the time information when the query was obtained. The obtained time information may be associated with the query obtained in step S226 and stored in the storage 207.

[0159] In step S230, the user terminal 200 obtains its location information at the time the query was obtained. The obtained location information may be stored in storage 207 in association with the query obtained in step S226.

[0160] In step S231, the user terminal 200 identifies one user from among one or more expected users who is using the user terminal 200. In this implementation, the identification of one user may be performed by reading the profile of the login account. In this implementation, the identification of one user may be performed by using a voice fingerprinting algorithm for voice queries. Based on the user identification, the user terminal 200 selects one trained model from one or more models as a trained model 2076 optimized for the individual.

[0161] In step S232, the user terminal 200 performs speech recognition of the query acquired in step S226. In speech recognition, the user terminal 200 may use the trained model 2076. The control in step S232 may be performed on the condition that the query was input as voice. If the query was input as text data, the control in step S232 may be omitted.

[0162] In step S234, the user terminal 200 performs natural language interpretation on the query obtained in step S226. This interprets the meaning of the query. In natural language interpretation, the user terminal 200 may use grammars stored in the grammar area 2072. The grammars stored in the grammar area 2072 include the grammar used to interpret the meaning of the query obtained in step S206. That is, the user terminal 200 can use the grammar used to interpret the meaning of the query obtained in step S206 (the first query) to interpret the meaning of the query obtained in step S226 (the second query).

[0163] The grammar used for natural language interpretation may be selected according to the circumstances under which the user terminal 200 obtained the query. The differences in the grammar used depending on the circumstances are explained below. By restricting the available grammars depending on the circumstances, it is possible to more reliably avoid the use of incorrect grammars in the natural language interpretation of queries. In other words, the accuracy of interpreting the meaning of queries can be improved.

[0164] In one example, the user terminal 200 may refer to location information 2075 (Figure 14) and select a range of grammars that can be used for natural language interpretation according to the location information obtained when the query was acquired. In the example in Figure 14, if the location information when the query was acquired is included in a location registered as "home," the grammar used for natural language interpretation will be selected from domains D, E, F, and G. If the location information when the query was acquired is included in a location registered as "office," the grammar used for natural language interpretation will be selected from domain C.

[0165] In another example, user terminal 200 follows the time information obtained when the query was retrieved. In other words, you may select the range of grammars that can be used for natural language interpretation. More specifically, if the time information falls within the time period registered as the user's working hours, the grammar used for natural language interpretation will be selected from domain C. If the time information does not fall within the time period registered as the user's working hours, the grammar used for natural language interpretation will be selected from domains D, E, F, and G.

[0166] In another example, the user terminal 200 may select a range of grammars that can be used for natural language interpretation according to the combination of location and time information obtained when the query is acquired. More specifically, if the location information when the query is acquired is within the location registered as "home" and the time information is within the time period registered as the user's cooking time, the grammar used for natural language interpretation is selected from domain E; otherwise, the grammar used for natural language interpretation is selected from domains D, F, and G.

[0167] In one implementation, instead of location and / or time information being strict conditions for selecting a range of grammars, one or more grammars that are assumed to be correct for interpreting a transcription of a given text data or utterance may be weighted. In one example, application program 16C outputs a probability for each grammar that it is applicable to the above text data or utterance transcription. Application program 16C selects the grammar whose sum of the output probability and the weights described above is the largest as the grammar to be used for interpreting the above text data or utterance transcription.

[0168] In step S236, the user terminal 200 determines whether the natural language interpretation in step S234 was successful. If the query is interpreted by one or more grammars stored in the grammar area 2072, the user terminal 200 determines that the natural language interpretation was successful. If the query is not interpreted by one or more grammars stored in the grammar area 2072, the user terminal 200 determines that the natural language interpretation failed.

[0169] If the user terminal 200 determines that the natural language interpretation in step S234 was successful (YES in step S236), it proceeds to step S238. If the user terminal 200 determines that the natural language interpretation in step S234 failed (NO in step S236), it proceeds to step S242.

[0170] In step S238, the user terminal 200 obtains the response to the query obtained in step S226 based on the interpretation in step S234. The user terminal 200 may also obtain the response to the query from the relevant data in the relevant data area 2073.

[0171] In one example, when the query "Play Yesterday by The Beatles" is received in step S206, the user terminal 200 stores a list of Beatles song titles as related data in the related data area 2073 in step S220. Then, when the query "Tell me a list of Beatles songs" is received in step S226, the user terminal 200 retrieves the list of Beatles songs stored in the related data area 2073 as a response to that query.

[0172] In step S240, the user terminal 200 outputs the response obtained in step S238. After that, the user terminal 200 terminates the process.

[0173] In step S242, the user terminal 200 stores the query obtained in step S226 in the failure data area 2074. The query stored in the failure data area 2074 is sent to the server 100 in step S204. After that, the user terminal 200 Terminate the process.

[0174] In the process described with reference to Figures 17 and 18, queries that failed to be interpreted in natural language while the user terminal 200 is offline are sent to the server 100. Queries that were successfully interpreted in natural language while the user terminal 200 is offline may also be sent to the server 100. This allows the server 100 to obtain queries entered into the user terminal 200 while it was offline. The server 100 may identify the grammar used to interpret the meaning of these queries. The server 100 may add and update counts (1) and (2) related to the identified grammar. This allows the values ​​of counts (1) and (2) to also reflect information about queries entered into the user terminal 200 while it was offline.

[0175] The disclosed features can be summarized as a computer-readable recording medium that stores non-temporarily recorded data containing methods, systems, computer software (programs), and / or instructions for performing such methods. For example, according to one aspect of this disclosure, a computer-readable recording medium that stores non-temporarily recorded data stores instructions for performing a method which includes receiving query input from a client terminal, performing natural language interpretation of the query using a grammar, outputting a response to the query after the natural language interpretation, and transmitting the grammar to the client terminal.

[0176] The embodiments disclosed herein should be considered in all respects as illustrative and not restrictive. The scope of the invention is indicated by the claims rather than by the foregoing description, and all modifications within the meaning and scope of the claims are intended to be included. Furthermore, the inventions described in the embodiments and each variation are intended to be practiced individually or in combination as far as possible. [Explanation of symbols]

[0177] 100 servers, 200, 200A~200G user terminals, 1500 estimated models.< / url:> < / musician> < / musician> < / musician> < / city> < / time> < / dish> < / musician> < / name>

Claims

1. A method performed by a computer, The steps include: accepting query input from the client terminal, The steps include: performing natural language interpretation of the query using grammar, After the execution of natural language interpretation, the step of outputting a response to the query, The steps include sending the grammar to the client terminal, The method includes the step of transmitting a speech recognition model used for speech recognition in the client terminal to the client terminal, Accepting the input of the aforementioned query includes accepting the input of a query generated on the client terminal when the client terminal is not connected to the computer. The step of sending the aforementioned grammar is, The number of times the aforementioned grammar is used for natural language interpretation of queries entered from the client terminal is counted, This includes sending grammars whose count value exceeds a threshold to the client terminal, The speech recognition model is trained using a training dataset that includes audio input from the user of the client terminal. The aforementioned training dataset includes information about the user stored in the client terminal as training data. A method wherein, in the training dataset, one or more audio recordings are tagged at the location where the user input the audio.

2. The method according to claim 1, further comprising the step of generating a training dataset for the speech recognition model using audio input from the user of the client terminal.

3. The method according to claim 1 or claim 2, wherein the information relating to the user is included in a contact file.

4. The method according to claim 1 or 2, wherein the information relating to the user includes the name in the contact file.

5. One or more processors, A server device comprising: a storage device storing a computer program which, when executed by the one or more processors, causes the one or more processors to carry out the method described in any one of claims 1 to 4; and a storage device storing a computer program.

6. A method performed by a computer that includes a storage device for storing information about a user, The steps include sending a first query to the server device, The steps include receiving a grammar from the server device that has been used for natural language interpretation of the first query and whose number of uses exceeds a threshold, The steps include receiving a speech recognition model used for speech recognition from the server device, The steps include storing the received grammar in memory, Steps to accept voice input, The steps include generating a second query by performing speech recognition on the aforementioned voice input using the aforementioned speech recognition model, The system includes the step of performing natural language interpretation of the second query using the grammar when the computer is not connected to the server device, The step of performing the aforementioned natural language interpretation is: The second query is stored in the memory, The computer sends the second query to the server device when it is reconnected to the server device, The speech recognition model is trained using a training dataset that includes audio input from the user. The training dataset includes the user information stored in the storage device as training data. A method wherein, in the training dataset, one or more audio recordings are tagged at the location where the user input the audio.

7. A computer program that, when executed by one or more processors of a computer, causes the computer to carry out the method described in claim 6.

8. Client terminal and The system includes a server device that sends a response to a query input from the client terminal to the client terminal, The server device includes one or more processors that perform natural language interpretation of the query using a grammar, The one or more processors transmit the speech recognition model used for speech recognition and the grammar to the client terminal. The aforementioned client terminal is When the client terminal is not connected to the server device, the speech recognition model is used to generate a second query. The client terminal stores the second query, When the client terminal is reconnected to the server device, the second query is sent to the server device. The one or more processors described above are: The number of times the aforementioned grammar is used for natural language interpretation of queries entered from the client terminal is counted. When the count value of the grammar exceeds the threshold, the grammar is sent to the client terminal. The speech recognition model is trained using a training dataset that includes audio input from the user of the client terminal. The aforementioned training dataset includes information about the user stored in the client terminal as training data. An information processing system in which, in the aforementioned training dataset, one or more audio recordings are tagged with the location where the user input the audio.

Citation Information

Patent Citations

  • Offline voice recognition model updating method, household appliance, and server

    CN110277089A

  • Local maintenance of data for selective offline-enabled voice actions in voice-enabled electronic devices

    JP2018523143A

  • Facilitating offline semantic processing on resource-constrained devices

    JP2019516150A

  • Methods and apparatus for generating, updating and distributing speech recognition models

    US20020065656A1

  • Grammar adaptation through cooperative client and server based speech recognition

    US20070276651A1