Voice response system and response method using thereof

The voice response system addresses the challenge of identifying non-standard place names by using a specialized database and inference units to efficiently search and respond to taxi dispatch requests, reducing operator burden and enhancing user experience.

JP2025162928APending Publication Date: 2025-10-28ETOMO WIRELESS BUSINESS COLLABORATION GROUP +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024066437
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-16
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing voice recognition systems struggle to accurately identify and respond to non-standard place names and map coordinates in taxi dispatch requests, especially when multiple names with similar spellings, pronunciations, or incomplete pronunciations are involved, leading to increased operator burden.

Method used

A voice response system with a voice recognition engine, speech engine, and integrated control device that utilizes a difficult-to-choose place name database and location database to identify and respond to non-standard place names, incorporating a map search unit and location identification inference unit to efficiently narrow down search results.

Benefits of technology

Enables efficient voice-based taxi dispatch by accurately identifying and responding to non-standard place names, reducing operator burden and improving user experience for elderly and tech-savvy individuals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025162928000001_ABST
    Figure 2025162928000001_ABST
Patent Text Reader

Abstract

To provide a voice response system capable of searching for place names and location information upon receiving a voice utterance and responding via voice.SOLUTION: A voice response system 1 comprises: a voice recognition engine 3 that converts voice information into text information; a speech engine 4 that converts text information into voice information; a difficult-to-select place name database 5 that stores difficult-to-select place names 14, which are difficult to search from a mapping system 10 (101, 102) registered with standardized place names 13 or difficult to reduce to a manageable number of easily voice-describable place names; a location database 6 that stores location information 15 of the difficult-to-select place names 14 of the difficult-to-select place name database 5; and an integrated control unit 1000 that searches the difficult-to-select place names 14 and location information 15 of the difficult-to-select place names 14 from the difficult-to-select place name database 5 and the location database 6 to reduce to the manageable number of easily voice-describable place names. This facilitates the voice response system 1 to search the location information of a difficult-to-select place name and respond with voice.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a voice response system that uses a computer to respond to voice calls, and a response method that uses the voice response system. [Background technology]

[0002] In recent years, when calling a taxi from a mobile device such as a smartphone or tablet PC, if you know the name of the store, company, house, place, or address of your current location, you can call a taxi by speaking to an operator over the telephone network. Furthermore, when you are out and about and do not know your exact location, or when it is difficult to describe it verbally, you can use a taxi dispatch app. By using such a taxi dispatch app, you can connect to the Internet from your mobile device and request a taxi to your current location on a web map displayed on the screen without going through an operator. However, elderly people and children who have difficulty using mobile devices have no choice but to call a taxi by speaking over the telephone network, as in the past, and taxi companies have been unable to reduce the number of operator requests.

[0003] One technology that can reduce the burden on operators is voice ordering using a computer. For example, Patent Document 1 (JP 8-320696 A) discloses a method for automatic speech recognition of arbitrarily spoken words in an automatic speech recognition (ASR) device having a first database storing word models and correlation data on which recognition decisions are at least partially based, and for enhancing the capabilities of the ASR device using information stored in an auxiliary second database, the method comprising the steps of receiving an input from a user having first and second portions, storing the input obtained from the user in the ASR device, recognizing the first portion of the input collected from the user by the ASR device, identifying and retrieving auxiliary data stored in the auxiliary database associated with the first portion of the input, creating a template derived from the information retrieved from the auxiliary database, and using the template to recognize the second portion of the input as spoken by the user. The auxiliary database is accessed to retrieve auxiliary document information, such as proper nouns. The text / speech means is used to generate a phonemic representation of the text retrieved from the auxiliary database, which phonemic representation can be used as a speaker-independent template in an automatic speech recognition (ASR) device to recognize spoken words, allowing the ASR device to recognize arbitrarily spoken words without speaker-specific training.

[0004] Furthermore, the automatic call system shown in Patent Document 2 (JP Patent Publication No. 2023-2650) discloses a method, a system, and an apparatus for an automatic call system. Some implementations are directed to using a bot to initiate a telephone call and conduct a telephone conversation with a user. The bot may be interrupted while providing synthesized speech during the telephone call. The interruption can be classified into one of a plurality of different interruption types, and the bot can react to the interruption based on the interruption type. Some implementations are directed to determining that a first user has been put on hold by a second user during a telephone conversation and maintaining the telephone call active in response to determining that the first user has hung up the telephone call. The first user can be notified when the second user rejoins the call, and a bot associated with the first user can notify the first user that the second user has rejoined the telephone call. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Publication No. 8-320696 [Patent Document 2] Japanese Patent Application Publication No. 2023-2650 Summary of the Invention [Problem to be solved by the invention]

[0006] However, the automatic speech recognition method of arbitrarily spoken words in Patent Document 1 relates to a method of using auxiliary information read from a database to assist an automatic speech recognition device (ASR) in recognizing words spoken by a user over a telephone network, but it is not capable of searching for and responding to target place names and map coordinates such as the boarding location and destination based on the user's voice information when receiving a taxi dispatch request over a telephone network.

[0007] The automatic call system of Patent Document 2 can collect important information such as the date and time, reservation details, number of people, etc., even when multiple users are making voice calls. However, like Patent Document 1, it cannot search for and respond to a request for a taxi dispatch by searching for the target place name and map coordinates.

[0008] When a taxi dispatch app receives a taxi dispatch request from a user, it detects the user's current location as the pickup location using a mapping system, or obtains location information such as an address and latitude and longitude through character input by the user, and then the destination can be selected on a map in the mapping system or obtained through character input by the user. However, if elderly people, children, or others who are unfamiliar with operating smartphones or tablet PCs are unable to use the taxi dispatch app properly, they will have to request a taxi by phone.

[0009] With previous computer technology, it was difficult to accept vehicle dispatch requests via voice call. Voice information was converted into text, place names were extracted from the text, and those place names were then searched for in a mapping system. Consumer mapping systems only registered official place names, and could not identify the location on a map for non-official place names such as local place names, shop names, and nicknames that were significantly different from the official place names. Furthermore, when multiple addresses were detected, such as place names with the same spelling but different pronunciations, place names with different spellings, and place names with incomplete pronunciations, it was difficult to efficiently narrow down the search to a single location via voice call. For these reasons, it was difficult to reduce the burden on operators.

[0010] Furthermore, it is expected that in the future, taxi orders will not only be accepted through the telephone network of taxi companies, but also that users in vehicles equipped with autonomous driving technology will speak their destinations (place names) to the AI ​​engine installed in the vehicle through a voice interface and indicate their destination. The AI ​​engine will convert the user's voice information into text information and search for the text information of the place name in the onboard car navigation system and a web mapping system connected to the onboard car navigation system. In this case, too, if the user utters an unofficial place name that is not registered in the onboard car navigation system or web mapping system, the problem of not being able to detect location information will arise. Furthermore, when multiple addresses are detected, such as place names with different spellings, different pronunciations, or place names with incomplete pronunciation, it remains difficult to efficiently narrow down the search to a single location via voice communication.

[0011] In view of the above circumstances, the present invention provides a voice response system that can search for place names and location information without an operator when receiving a voice utterance and respond by voice.The present invention provides a voice response system that can search for location information of place names that are difficult to choose between, such as irregular place names that are not on maps such as existing consumer mapping systems, and irregular place names that are detected multiple times, and respond by voice.The present invention provides a voice response system that can, when receiving a voice utterance, more efficiently narrow down the search to one location while conversing with the user if multiple addresses of place names with the same spelling but different pronunciations, place names with different spellings, or place names with incomplete pronunciations are detected. [Means for solving the problem]

[0012] The present invention relates to a voice response system having a voice recognition engine that converts voice information received via a voice interface into text information, a speech engine that converts text information into voice information, a difficult-to-choose place name database that stores difficult-to-choose place names that are difficult to search in a mapping system in which regular place names are registered, or that are difficult to narrow down to a number of place names that are easy to explain aloud, a location database that stores location information of the difficult-to-choose place names in the difficult-to-choose place name database, and an integrated control device that searches the difficult-to-choose place name database and the location database for the difficult-to-choose place names and location information of the difficult-to-choose place names, and narrows down to a number of place names that are easy to explain aloud.

[0013] The voice interface is interposed between the user and the voice recognition engine and speech engine, allowing the user's voice to be input to the voice recognition engine and transmitting the voice output from the speech engine to the user. The voice interface can be a microphone connected to the voice recognition engine, a speaker or headphones connected to the speech engine, etc. The voice interface can be a telephone, mobile terminal, tablet PC, etc. connected to a telephone network, a mobile communication network, the Internet, a telephone switching system, etc. The voice interface can be a car navigation system, AI engine, in-vehicle AI engine, etc. connected to a telephone network, a mobile communication network, the Internet, a telephone switching system, etc. and having a voice input / output unit such as a microphone and a speaker. The telephone network can be a communication network capable of voice calls, including a public switched telephone network, a mobile communication network, the Internet, etc. The telephone can be a communication device capable of voice calls, such as a landline telephone, a mobile phone, or a mobile terminal.

[0014] The place name is information that the voice recognition engine can determine as a place name or landmark name by converting the user's voice information into text information, and can be, for example, a building name, facility name, shop name, station name, route name, bus stop name, intersection name, street name, area name, utility pole number, bridge name, tunnel name, river name, coast name, latitude and longitude, or other information linked to a specific location on a map.

[0015] The legitimate place names can be defined as place names registered in the consumer-oriented mapping system. The difficult-to-alternate place names can be defined as place names extracted from the speech of multiple users that partially, almost completely, or completely match the legitimate place names and are not detected by the mapping system. Place names not registered in the mapping system can also be treated as difficult-to-alternate place names until they are registered. The difficult-to-alternate place names can be defined as place names that exist multiple times, regardless of whether they are legitimate or non-legitimate, and for which multiple addresses are found by a map search, such as the names of chain stores or companies within the same affiliation. As the mapping system evolves, if the difficult-to-alternate place names are detected by the mapping system, a maintenance unit having a cleaner can be provided to delete them from the difficult-to-alternate place name database and the location database accordingly. The cleaner can be software that deletes unnecessary data to increase the search speed of the difficult-to-alternate place name database and the location database.

[0016] In the integrated control device, the voice recognition engine converts the voice information uttered by the user into text information and outputs it, and the speech engine converts the text information received from the integrated control device into voice and outputs it. The integrated control device may have the difficult-to-choose place name database. The integrated control device can be customized to recognize words with identification symbols, identifiers, identification codes, extensions, or other index information.

[0017] The voice recognition engine, speech engine, and location inference unit described later are software that can be implemented in devices that include a CPU, such as a personal computer, a cloud server, a local server, etc. The communication processing of voice signals by the voice recognition engine and speech engine for the voice interface can also be offline processing using a local engine.

[0018] The voice recognition engine converts the voice information uttered by the user into text information. For example, the voice recognition engine converts the voice information uttered by the user into two types of text information: text information including kanji and text information consisting only of hiragana. The integrated control device extracts and acquires place names from the text information recognized by the voice recognition engine, for example, text information including kanji, searches the acquired place names using a mapping system, acquires location information (address, latitude, and longitude) for the acquired place names, and outputs the location information to a guidance system, a vehicle dispatch system, or the like. The integrated control device may also include a map search unit and the location identification inference unit. If the integrated control device is unable to search using the mapping system, it may search the difficult-to-choose place name database and the location database for the place names and location information (address, latitude, and longitude). Furthermore, if the integrated control device is unable to search using the difficult-to-choose place name database and the location database, it may search the mapping system for the place names and location information (address, latitude, and longitude).

[0019] The integrated control device searches the mapping system for a place name recognized from the user's words by the map search unit, and when the place name is not detected by the mapping system, determines the place name to be a difficult-to-alternate place name. The map search unit searches the difficult-to-alternate place name in the difficult-to-alternate place name database, and when the number of detected place names that are easy to explain by voice (for example, 1 to 3) cannot be narrowed down, or when no place names are detected, the integrated control device transfers processing from the map search unit to the location identification inference unit. The location identification inference unit can request additional geographic information from the user and narrow down the number of detected place names based on the difficult-to-alternate place name and the additional geographic information.

[0020] The integrated control device establishes a voice call with the user through the voice recognition engine and speech engine, searches for place names in the difficult-to-choose place name database and the location database, and enables the map search unit and / or the location identification inference unit to narrow down the number of detected place names. The integrated control device controls the map search unit and / or the location identification inference unit to search for place names in the difficult-to-choose place name database and the location database. The integrated control device may be, for example, an AI commander. The AI ​​commander may include the voice recognition engine, speech engine, map search unit, and location identification inference unit. Furthermore, the AI ​​commander may include a map search unit and an inference selection unit. The number of detected place names that are easy to explain by voice is the number of place names that can be efficiently conveyed without error in a voice call or voice conversation. Generally, since a larger number can confuse the user, it is desirable to limit the number of place names to approximately 1 to 5, preferably 1 to 3.

[0021] The mapping system accumulates the legitimate place names and their location information, and can output the latitude, longitude, and address of the place names by searching for text information such as the place names and latitude, longitude, etc., and can also output the place names by searching for the text information of the latitude, longitude, or address. The mapping system can be, for example, a web mapping platform or map engine published on the Internet, or a terminal device equipped with a map application. The mapping system can be, for example, a microcomputer or software installed in a car navigation system, a car navigation system for an autonomous driving system, an in-vehicle AI, a wearable device worn by a user, or the like. The mapping system can be, for example, Increment P (registered trademark: GeoTechnologies Inc., hereinafter omitted), Google Maps (registered trademark: Google LLC), Yahoo! Maps (LINE Yahoo! Corporation), etc.

[0022] The mapping system has a geocoding function that returns an address, latitude, and longitude from a place name, and a reverse geocoding function that returns a place name from an address, latitude, and longitude.The mapping system is a target that the integrated control device of the present invention searches for location information (address, latitude, and longitude) of land corresponding to a place name based on the place name.The mapping system can be one or more web mapping systems that access place names and their location information (address, latitude, and longitude) via the Internet and can obtain necessary map information.The mapping system can be one or more map information databases that are not publicly available on the Internet.The mapping system can be one or more web mapping systems, map engines, and one or more map information databases that are not publicly available on the Internet.

[0023] The integrated control device can search a difficult-to-choose place name database (described later) for place names that are found in the mapping system in excess of the number of place names that are easy to explain by voice (for example, 1 to 3).The integrated control device can search the location information (address, latitude and longitude) of place names found in the difficult-to-choose place name database (described later) or the mapping system (search for legitimate place names).

[0024] The voice recognition engine converts voice information into text information such as characters including kanji or hiragana. The speech engine converts text information emitted by the integrated control device into voice and outputs it. The text information can be phonograms, phonetic characters, phonemic characters, ideograms, or logographic characters. The text information can be, for example, kanji, Roman characters, alphabetic characters, Hangul characters, Arabic characters, Greek characters, numbers, or a mixture of these. The integrated control device or the voice recognition engine can have a language control unit equipped with a language discrimination unit that determines whether the received user's voice information is Japanese or a foreign language, and if it is a foreign language, which country's language it is, based on the received user's voice information, and the language control unit can control the speech engine to respond in the language determined by the language discrimination unit.

[0025] The map search unit and location identification inference unit convert the user's voice information into text information using the voice recognition engine, and based on the place names extracted from the text information, search the difficult-to-choose place name database and the location database for the place names and location information (address, latitude and longitude) and output them.

[0026] The map search unit and the location determination reasoning unit can be incorporated as software into the integrated control device. The map search unit and the location determination reasoning unit can be incorporated as software into at least one of the voice recognition engine or the speech engine. The voice recognition engine or the speech engine can be incorporated as software into the map search unit or the location determination reasoning unit. The speech recognition engine can be incorporated as software into the map search unit or the location determination reasoning unit. The speech engine can be incorporated as software into the map search unit or the location determination reasoning unit.

[0027] The place name can be the name of a place such as the user's boarding location or destination. The place name can be a legitimate place name stored in the mapping system together with location information such as address, latitude, and longitude. The place name can be a difficult-to-choose place name that is not stored in the mapping system. The place name can be a mispronunciation or multiple place names with the same spelling but different pronunciations. The place name can be multiple place names with different spellings. The place name can be a place name with different spellings that is not clear enough.

[0028] The place name can be the name of a location that can be identified as a single point on a map. The location information can be location information (address, latitude, and longitude) of the land, building, etc. of the place name. The location information can be an address or a telephone number linked to an address. The location information can also be an address linked to a store name, company name, name, trade name, or telephone number. The location information can be, for example, latitude and longitude, a location information code, or an ID (identification) set on a map to uniquely identify a specific point. Location information codes other than latitude and longitude may have low accuracy and may be less practical.

[0029] The voice response system is configured such that the integrated control device outputs place names and location information (addresses, latitude and longitude) received from the voice recognition engine, map search unit, location identification inference unit, or inference selection unit (described later). The voice response system can be incorporated into a mobile terminal equipped with an electronic map that guides pedestrians, vehicles, etc., or a car navigation system. The voice response system can be incorporated into an autonomous driving system or a vehicle dispatch system that manages the operation, route selection, and driving management of autonomous vehicles. The voice response system can be a vehicle dispatch system that manages vehicle operations, for example, dispatching taxis, managing bus operations, managing collection and delivery routes for delivery vehicles, and managing routes for day care vehicles.

[0030] The difficult-to-alternate place name database includes, for example, a common place name dictionary, a homograph-variant pronunciation place name dictionary, and a homophone-variant spelling place name dictionary, and the common place name dictionary, the homograph-variant pronunciation place name dictionary, and the homophone-variant spelling place name dictionary store distinguishable difficult-to-alternate place names obtained by assigning identifiers to the difficult-to-alternate place names in text information. The common place name dictionary can store distinguishable common place names as distinguishable difficult-to-alternate place names, which are obtained by assigning identifiers indicating that they are common place names to common place names in text information as the difficult-to-alternate place names, in association with the regular place names that match the distinguishable common place names. The homograph-variant pronunciation place name dictionary can store mispronounced place names as the difficult-to-alternate place names, and distinguishable homograph-variant pronunciation place names as distinguishable difficult-to-alternate place names, which are obtained by assigning identifiers indicating that they are homograph-variant pronunciation place names to place names in text information with the same pronunciation, in association with the regular place names that match the distinguishable homograph-variant pronunciation place names. The homophone place name dictionary can store, in association with the difficult-to-choose place names, place names with homophone text information, and place names with insufficient text information, identified homophone place names as difficult-to-choose place names, with an identifier indicating that they are homophone place names, and the regular place names that match the identified homophone place names.

[0031] The difficult-to-identify place names registered in the common place name dictionary, the homophonic place name dictionary, and the homophonic place name dictionary in the difficult-to-identify place name database, and the location information (address, latitude, longitude) associated with the difficult-to-identify place names in each dictionary in the location database, match the official place names and location information (address, latitude, longitude) of the mapping system. The difficult-to-identify place names can be stored in both the difficult-to-identify place name database and the location database in duplicate so that they match each other, and a multi-search for the difficult-to-identify place names in the location database can be performed based on the identifiers of the difficult-to-identify place names in the difficult-to-identify place name database to detect the location information (address, latitude, longitude) of the difficult-to-identify place names in the location database. The location database can also store the official place names and location information (address, latitude, longitude) of the mapping system that match the difficult-to-identify place names.

[0032] The difficult-to-choose place name database and the location database record identity information (clone information, copy information, etc.) of the difficult-to-choose place names (difficult-to-choose place names with identifiers: _e, _n, _d, etc.) registered in the common place name dictionary, the homophonic and different pronunciation place name dictionary, the homophonic and different spelling place name dictionary, and each place name dictionary. The linking keywords that enable multi-index search of the difficult-to-choose place name database and the location database can be said to be the identifiers (_e, _n, _d, etc. added to the difficult-to-choose place names of difficult-to-choose place names).

[0033] The common place name dictionary stores searchable place names such as local nicknames and nicknames that are difficult to choose from, such as those that are not detected by the mapping system or that are detected multiple times, unlike the regular place names stored in the mapping system. The common place names registered in the common place name dictionary are assigned identifiers that are different from other difficult-to-choose place names, making them distinguishable common place names, and therefore can be distinguished more efficiently during searches, unlike other distinguishable homophonic, different pronunciation place names and distinguishable homophonetic place names. The common place name dictionary has a smaller number of search targets than the homophonic, different pronunciation place name dictionary and the homophonetic place name dictionary, allowing for faster place name searches.

[0034] The homograph-different pronunciation place name dictionary stores, in a searchable manner, homograph-different pronunciation place names, such as misreadings of kanji characters or different pronunciations of English, among the difficult-to-choose place names. The homograph-different pronunciation place names registered in the homograph-different pronunciation place name dictionary are assigned identifiers that are different from other difficult-to-choose place names, making them distinguishable homograph-different pronunciation place names, and therefore can be distinguished more efficiently during searches, unlike the other distinguishable common place names and the distinguishable homophone-different spelling place names.

[0035] The homonymous place name dictionary stores searchable homonymous place names that cannot be distinguished by phonetic means. The homonymous place names registered in the homonymous place name dictionary are assigned an identifier that is different from other difficult-to-choose place names, making them distinguishable homonymous place names. Therefore, they can be distinguished more efficiently during searches, unlike other distinguishable commonly used place names and distinguishable homonymous place names with different pronunciations.

[0036] The difficult-to-choose place names may be, for example, place names based on the speech of each user, such as local names, nicknames, abbreviated names, and alternative names resulting from pronunciations or misrememberings unique to each user. To efficiently search for the difficult-to-choose place names, it is necessary that the character information output by the speech recognition engine be stored in each dictionary as identified difficult-to-choose place names with the identifier (e, n, d, etc.). The difficult-to-choose place name database stores the difficult-to-choose place names that are not included in the mapping system in which the legitimate place names are registered, as well as many difficult-to-choose place names that are included in the mapping system. If the difficult-to-choose place names can be detected in the difficult-to-choose place name database, the burden on the operator can be reduced.

[0037] The difficult-to-identify place names stored in the common place name dictionary, the homophonic place name dictionary, and the homophonic place name dictionary are assigned identifiers indicating the classification of each dictionary, making them difficult-to-identify place names. This enables a double index search of the difficult-to-identify place names and their identifiers, improving search efficiency and speed. The difficult-to-identify place names are place names written in kana characters, and a double index search using the place names written in kanji characters and / or kana characters and their identifiers enables faster and more thorough searches. Each dictionary can store the difficult-to-identify place names and the legitimate place names that match the difficult-to-identify place names. Because the difficult-to-identify place names recorded in the difficult-to-identify place name database and the location database match each other, a multi-search of location information (address, latitude, longitude) can be performed from the difficult-to-identify place names. By linking and recording the difficult-to-identify alternative place names and the legitimate place names in at least one of the difficult-to-alternate place name database or the location database, it becomes possible to search for more accurate location information (address, latitude and longitude).

[0038] The present invention can be the voice response system in which the location database stores the difficult-to-alternate place names, the legitimate place names that match the difficult-to-alternate place names, and location information (addresses, latitude and longitude) of the legitimate place names in association with each other. By storing the difficult-to-alternate place names and the legitimate place names that match the difficult-to-alternate place names, the location database can manage location information (latitude and longitude) more accurately, and can further reduce the time required to search for difficult-to-alternate place names.

[0039] The present invention can also be the voice response system having an initial response unit that queries the user for user information and place names; a map search unit that searches the difficult-to-choose place name database for place names and location information of the place names obtained by the initial response unit; a location identification inference unit that performs a search by the integrated control device by adding additional geographical information to the place names; a transfer unit that transfers processing to the location identification inference unit if the search by the map search unit fails to narrow down the number of detected place names that are easy to explain by voice, or if no place names are detected; a selection unit that, if the map search unit has detected a number of detected place names that are easy to explain by voice, assigns a voice code (such as a number from 1 to 3) to each legitimate place name and location information of the legitimate place name obtained by a multi-index search of the difficult-to-choose place name database and the location database, and outputs a question to the user requesting a selection; and a confirmation unit that outputs a question to confirm the correct answer to the one place name and location information selected by the user in response to the question in the selection unit.

[0040] The user makes a voice call via the voice response system of the present invention, and the user information can be information that identifies the user, such as the user's surname, given name, phone number, email address, address, etc. Furthermore, the user information can be linked and added with the name of the boarding location, the address of the boarding location, the store name, map coordinates, the boarding reservation time (now, in x minutes, x year x month x day x hour x minute), the destination name, the number of passengers, the order details such as the type of car, the usage history, etc.

[0041] The voice recognition engine, speech engine, map search unit, location identification inference unit, and inference selection unit (described later) can be included in an integrated control device (e.g., an AI commander) of the voice response system of the present invention. The integrated control device can ask the user only for the name of the boarding place through the voice recognition engine, without distinguishing between boarding places, disembarking places, destinations, and so on, and recognizes the place name as the boarding place name when it recognizes it in the user's speech, and repeats (as rebuttal evidence) that the recognized place name is the boarding place name to confirm to the user. The integrated control device can also ask the user for the name of the boarding place through the voice recognition engine and speech engine, and when it recognizes a place name in the user's speech, it recognizes it as the boarding place name, and repeats (refuting evidence) that the recognized place name is the boarding place name to confirm to the user; and further, it can ask the user for the name of the disembarking place through the voice recognition engine and speech engine, and when it recognizes a place name in the user's speech, it recognizes it as the disembarking place name, and repeats (refuting evidence) to confirm that the recognized place name is the disembarking place name (destination).

[0042] The difficult-to-choose place name database may have a user information management unit such as a user information recording unit, a user history recording unit, and a user remarks recording unit, and the integrated control device may respond to the user while referring to the user information management unit. The user information management unit may more efficiently respond to users who make roughly the same calls each time. When requesting personal information from the user, the integrated control device may obtain consent in advance to handling the personal information, and if consent is not obtained, may not store the personal information.

[0043] The initial response unit may execute an initial response operation in which the speech engine asks for user information and a place name when the integrated control device starts communication with the voice interface. When the user responds to the initial response operation by voice with a place name, the integrated control device outputs the acquired voice information to the voice recognition engine, and the voice recognition engine converts the place name into text information. The integrated control device may be replaced by an AI (artificial intelligence) that can hold a conversation.

[0044] The map search unit searches the difficult-to-identify place name in the difficult-to-alternate place name database based on the place name answered by the user in response to the question from the initial response unit. The map search unit can also search the mapping system for a legitimate place name based on the place name answered by the user, simultaneously with or before or after the search for the difficult-to-identify place name in the difficult-to-alternate place name database. The map search unit can search the difficult-to-alternate place name in the difficult-to-alternate place name database, and if the difficult-to-alternate place name is not found, attempt to search the mapping system for the difficult-to-alternate place name. The map search unit can also search the difficult-to-alternate place name database for a legitimate place name (including exact matches and partial matches) based on the place name answered by the user, for example. The map search unit is connected to the difficult-to-alternate place name database and a location database. Furthermore, the map search unit can be connected to the mapping system. The mapping system can be connected via, for example, a communication line, a communication circuit, a communication transmission circuit, or the like, and can be equipped with or incorporate the mapping system, a map engine, or the like.

[0045] The transfer unit causes the map search unit to search the difficult-to-identify place names in the difficult-to-alternate place name database based on the place names answered by the user in response to the question from the initial response unit, and determines whether the search results from the difficult-to-alternate place name database have been narrowed down to the number of detected place names that are easy to explain by voice. If the search results are not narrowed down sufficiently to the number of detected place names that are easy to explain by voice, the transfer unit can switch to a search by the location identification inference unit that targets the difficult-to-alternate place name database and the location database, adding additional geographic information to the place names. Furthermore, the transfer unit performs a search by the map search unit based on the place names answered by the user in response to the question from the initial response unit, targeting the difficult-to-alternate place name database, the location database, and the mapping system, and determines whether the search results are narrowed down sufficiently to the number of detected place names that are easy to explain by voice, and if the search results are not narrowed down sufficiently, can switch to a search by the location identification inference unit that targets the difficult-to-alternate place name database and the location database, adding additional geographic information to the place names.

[0046] The location database has common index information with the difficult-to-alternate place name database, using the identifier as a key, and can register the location information (address, latitude, longitude) of the difficult-to-identify place name linked to the difficult-to-alternate place name so that when the difficult-to-identify place name is output from the difficult-to-alternate place name database, a multi-index search can be performed on the location information (address, latitude, longitude) of the difficult-to-identify place name.

[0047] When the map search unit detects a number of detected place names that are easy to explain aloud, the selection unit speaks to the user through the speech engine to request one of the options. For example, when the number of detected place names that are easy to explain aloud is about 1 to 5, preferably about 1 to 3, the selection unit can assign phonetic codes to the regular place names and location information of the regular place names for each of the retrieved difficult-to-choose place names, and output a question requesting the selection of the phonetic codes. Furthermore, the selection unit can assign phonetic codes not to the regular place names but to the difficult-to-choose place names spoken by the user and location information of the regular place names, and output a question requesting the selection. Furthermore, when the number of detected place names that are easy to explain aloud is about 1 to 5, the selection unit can assign phonetic codes to the addresses of the difficult-to-choose place names or the names and location information of landmarks near the addresses (additional geographic information) for the difficult-to-choose place names, and output a question requesting the selection of the difficult-to-choose place names.

[0048] The selection section can output, via the speech engine, a question requesting a selection by adding a phonetic code, for example, numbers 1 to 3, A, B, C, or ii, ro, ha, etc. to the place names and additional geographic information of, for example, three places or less output by the map search section or the output section of the location identification inference unit described later. The user can answer the desired place name using the phonetic code, or the place name with the additional geographic information, etc.

[0049] The confirmation unit can output, via the speech engine, a question to confirm the correctness of one place name and location information or additional geographic information selected by the user in response to the multiple-choice question. The confirmation unit can output a question to confirm one legitimate place name and location information of the legitimate place name selected by the user in response to the multiple-choice question by attaching a voice code to the legitimate place name. The confirmation unit can also output a question to confirm one difficult-to-choose place name selected by the user in response to the multiple-choice question by attaching a voice code to the address of the difficult-to-choose place name or the name of a landmark close to the address.

[0050] The integrated control device may have an ordering unit that outputs the user information, a place name, and location information of the place name when the confirmation unit confirms the correct answer based on the user's response. The ordering unit may output the legitimate place name and location information of the legitimate place name, or output the difficult-to-alternate place name and the address of the difficult-to-alternate place name or the name of a landmark close to the address. The integrated control device may include the initial response unit, map search unit, transfer unit, selection unit, confirmation unit, and ordering unit.

[0051] The present invention can also be a voice response system having: a specific information addition unit that outputs a question for additional geographical information to the user when the location identification inference unit receives processing transfer from the map search unit; an inference selection unit that adds the additional geographical information to the place names and narrows down the difficult-to-choose place name database to the number of detected place names that are easy to explain aloud; and an output unit that sends the search results to the selection unit when the inference selection unit searches the difficult-to-choose place name database and the location database for the difficult-to-choose place names, the regular place names of the difficult-to-choose place names, and location information of the regular place names, and detects the number of detected place names that are easy to explain aloud.

[0052] The inference selection unit performs an index search for the difficult-to-choose place names based on the place names in the difficult-to-choose place name database and the location database, detects the number of detected place names that are easy to explain by voice, and outputs the legitimate place names of the difficult-to-choose place names and location information of the legitimate place names. The inference selection unit can output the difficult-to-choose place names and location information.

[0053] The inference selection unit assigns audio codes to each legitimate place name and location information of the legitimate place name searched by the inference selection unit, and outputs a question requesting a selection to the user. The inference selection unit can assign audio codes to the legitimate place names of the difficult-to-choose place names and location information of the legitimate place names, and output them. Furthermore, the inference selection unit can assign audio codes to the difficult-to-choose place names and location information, and output them.

[0054] The present invention relates to a response method using the voice response system, which includes an inference search process in which the location identification inference unit performs an index search of the difficult-to-select place names in the difficult-to-select place name database and the location database; a specific information addition process in which the inference selection unit outputs a question about additional geographical information if the inference search process fails to narrow down the number of detected place names that are easy to explain by voice, or if no place names are detected at all; an inference selection process in which the inference selection unit adds the additional geographical information obtained in the specific information addition process to the difficult-to-select place names, and searches the difficult-to-select place name database and the location database to narrow down the number of detected place names that are easy to explain by voice; and an output process in which the regular place names narrowed down in the inference selection process and the additional geographical information or addresses are sent to the selection unit.

[0055] In the inference search step, when the integrated control device recognizes a place name from the user's speech through the voice recognition engine, the integrated control device searches the difficult-to-choose place name database for the difficult-to-choose place name, and further performs an index search of the identified difficult-to-choose place name linked to the difficult-to-choose place name in the location database. By performing an index search based on the difficult-to-choose place name and the identifier in the difficult-to-choose place name database and the location database, searches can be performed more quickly and efficiently.

[0056] The specific information adding step outputs a question for additional geographical information if the inference search step fails to narrow down the number of detected place names to those that are easy to explain by voice, or if no place names are detected. By asking the user for additional geographical information, information that is advantageous for detecting place names can be acquired more efficiently. The inference selection step narrows down the number of detected place names to those that are easy to explain by voice, by performing a search that includes the additional geographical information obtained in the specific information adding step. The number of detected place names is preferably narrowed down to 1 to 5, preferably 1 to 3. The inference selection step assigns a voice code (such as a number 1 to 3) to the regular place names and additional geographical information or addresses narrowed down in the inference selection step, and outputs a question requesting selection. Assigning the voice code shortens the user's response time and enables more efficient and accurate calls.

[0057] The voice response system of the present invention may include a maintenance unit that maintains the difficult-to-choose place name database and the location database. For example, the maintenance unit may include an identifiable common place name registration unit that, when the integrated control device detects a difficult-to-choose place name, stores in the common place name dictionary an identifiable common place name as an identifiable difficult-to-choose place name, in which an identifier indicating that the common place name is a common place name is added to the common place name in kana characters as the difficult-to-choose place name, and the regular place name that matches the identifiable common place name, in association with each other.

[0058] The maintenance unit may have a discriminative homophonetic place name registration unit that stores, in the homophonetic place name dictionary, discriminative homophonetic place names as difficult-to-choose place names, in which place names in mispronounced kana characters and place names in homophonetic and different pronunciation kana characters are assigned an identifier indicating that they are homophonetic and different pronunciation place names, and the regular place names that match the discriminative homophonetic and different pronunciation place names.The maintenance unit may have a discriminative homophonetic place name registration unit that stores, in the homophonetic place name dictionary, discriminative homophonetic place names as difficult-to-choose place names, in which place names in homophonetic and different pronunciation kana characters are assigned an identifier indicating that they are homophonetic and different pronunciation place names, and the regular place names that match the discriminative homophonetic place names. The maintenance unit can generate maintenance information every time it finds a new difficult-to-choose place name and register it in the difficult-to-choose place name database and the location database for maintenance.

[0059] When the discriminative common place name registration unit detects a common place name that is a difficult-to-alternate place name, it links the discriminative common place name with the regular place name that matches the discriminative common place name to generate maintenance information and stores the maintenance information in the common place name dictionary. The common place name dictionary is provided, for example, in the difficult-to-alternate place name database, and the location database can store location information (address, latitude, longitude) linked with the discriminative common place name and the regular place name that matches the discriminative common place name. The common place name dictionary is provided, for example, in each of the difficult-to-alternate place name database and the location database, and the location database can store location information (address, latitude, longitude) linked with the common place name dictionary.

[0060] When the discriminative homophonic-phonetic place name registration unit detects a homophonic-phonetic place name that is a difficult-to-choose place name, it links the discriminative homophonic-phonetic place name with the regular place name that matches the discriminative homophonic-phonetic place name to generate maintenance information and stores the maintenance information in the homophonic-phonetic place name dictionary. The homophonic-phonetic place name dictionary is provided, for example, in the difficult-to-choose place name database, and the location database can store location information (address, latitude, longitude) that links the discriminative homophonic-phonetic place name with the regular place name that matches the discriminative homophonic-phonetic place name. The homophonic-phonetic place name dictionary is provided, for example, in both the difficult-to-choose place name database and the location database, and the location database can store location information (address, latitude, longitude) that is linked to the homophonic-phonetic place name dictionary.

[0061] When the discriminative homonymous place name registration unit detects a homonymous place name that is a difficult-to-choose place name, it generates maintenance information by linking the discriminative homonymous place name with the regular place name that matches the discriminative homonymous place name, and stores the maintenance information in the homonymous place name dictionary. The homonymous place name dictionary is provided, for example, in the difficult-to-choose place name database, and the location database can store location information (address, latitude, longitude) that links the discriminative homonymous place name with the regular place name that matches the discriminative homonymous place name. The homonymous place name dictionary is provided, for example, in both the difficult-to-choose place name database and the location database, and the location database can store location information (address, latitude, longitude) that is linked to the homonymous place name dictionary.

[0062] The maintenance unit maintains the difficult-to-alternate place name database and the location database. The maintenance unit registers multi-index information in each of the common place name dictionary, the homophonic and different pronunciation place name dictionary, and the homophonic and different spelling place name dictionary. The maintenance unit can be incorporated as software into the integrated control device. The maintenance unit can be incorporated as software into the speech recognition engine. The maintenance unit can be incorporated as software into the speech engine or the difficult-to-alternate place name database. The maintenance unit can be incorporated as software into the location database.

[0063] As an aspect related to the voice response system of the present invention, there is a maintenance method using the maintenance unit of the voice response system. The maintenance method for the voice response system can include, for example, an identifiable common place name registration step in which, when the maintenance unit detects a common place name that is a difficult-to-alternate place name and has not yet been stored in the difficult-to-alternate place name database, the identifiable common place name registration unit records in the common place name dictionary an identifiable common place name as an identifiable difficult-to-alternate place name, in which the common place name in kana characters as the difficult-to-alternate place name is given an identifier indicating that it is a common place name, and links it to the regular place name that matches the identifiable common place name, and further records in the location database the regular place name that matches the identifiable common place name and is linked to the identifiable common place name, and location information of the regular place name.

[0064] The maintenance method for the voice response system may be such that, when a recognition error occurs in the voice recognition engine of the voice response system, an operator or an administrator of the voice response system or the like adds a judgment based on human experience from the voice information and text information that caused the recognition error and registers difficult-to-choose place names. The registration work of difficult-to-choose place names that adds a judgment based on human experience can be performed by an accumulation and learning type AI engine instead of a human.

[0065] The AI ​​engine that performs maintenance on the voice response system may, for example, link the difficult-to-choose place name extracted from the user's recorded voice information with the location information (address, latitude, and longitude) where the operator dispatched the vehicle and record the link as unregistered information when the integrated control device is unable to respond to the user's voice information, an operator responds, completes the dispatch process, and the operator does not register the difficult-to-choose place name. Next, when the integrated control device recognizes the same difficult-to-choose place name as described above, is unable to respond to the user's voice information, an operator dispatches a vehicle to the same location information (address, latitude, and longitude), and the operator does not register the difficult-to-choose place name, count (record) this as a secondary record, and, based on the secondary record, determine that the accuracy of the difficult-to-choose place name and the location information (address, latitude, and longitude) has been confirmed, register the difficult-to-choose place name and the location information (address, latitude, and longitude) in the difficult-to-choose place name database and the location database.

[0066] Furthermore, as described above, when the AI ​​engine that performs maintenance on the voice response system recognizes the difficult-to-choose place name for the second time, it can confirm with the user by voice the location information (address, latitude, longitude) used by the operator to dispatch the vehicle the previous time, and if the information is valid, it can determine that the accuracy of the difficult-to-choose place name and location information (address, latitude, longitude) has been confirmed, and register the difficult-to-choose place name and location information (address, latitude, longitude) in the difficult-to-choose place name database and location database.

[0067] The maintenance method can include, for example, an identifying homophonic-variant pronunciation place name registration step in which, when the maintenance unit detects a mispronounced place name or a homophonic-variant pronunciation place name that is a difficult-to-alternate place name and has not yet been stored in the difficult-to-alternate place name database and the location database, the identifying homophonic-variant pronunciation place name registration unit links and records in the homophonic-variant pronunciation place name dictionary the mispronounced place name or the homophonic-variant pronunciation place name as the difficult-to-alternate place name, and the canonical place name that matches the identified homophonic-variant pronunciation place name, and further records in the location database the canonical place name that matches the identified homophonic-variant pronunciation place name that is linked to the identified homophonic-variant pronunciation place name, and location information for the canonical place name.

[0068] The maintenance method can include, for example, an identifying homophonetic place name registration step in which, when the maintenance unit detects a homophonetic place name and an insufficient kana character place name that are difficult to choose from and not yet stored in the difficult-to-choose place name database and the location database, the identifying homophonetic place name registration unit links and records in the homophonetic place name dictionary the identified homophonetic place name as an identified difficult-to-choose place name, with an identifier indicating that it is a homophonetic place name, and the regular place name that matches the identified homophonetic place name, and further records in the location database the regular place name that matches the identified homophonetic place name and is linked to the identified homophonetic place name.

[0069] The discriminative commonly used place name registration unit, the discriminative homophonic, different pronunciation place name registration unit, and the discriminative homophonic, different spelling place name registration unit can all register the difficult-to-alternate place names, each with an identifier that serves as index information for each dictionary, the regular place names linked to the difficult-to-alternate place names, location information (address, latitude and longitude), user information, etc., as multi-index information in the difficult-to-alternate place name database and the location database.

[0070] The distinguishable common place name registration step is executed when the maintenance unit detects the common place name that is a difficult-to-alternate place name that has not yet been stored in the difficult-to-alternate place name database. The distinguishable common place name registration step can be performed by the distinguishable common place name registration unit registering, in the common place name dictionary, a distinguishable common place name as a difficult-to-alternate place name, in which an identifier indicating that it is a common place name is added to a common place name in kana characters as the difficult-to-alternate place name, and linking the distinguishable common place name as a distinguishable difficult-to-alternate place name with the regular place name that matches the distinguishable common place name. Furthermore, the common place name registration step can register, in the location database, the regular place name that matches the distinguishable common place name and is linked to the distinguishable common place name, and location information of the regular place name.

[0071] The distinguishable homograph-different pronunciation place name registration step is executed when the maintenance unit detects a homograph-different pronunciation place name that is a difficult-to-choose place name and has not yet been stored in the difficult-to-choose place name database. The distinguishable homograph-different pronunciation place name registration step can be performed by the distinguishable homograph-different pronunciation place name registration unit to record in the homophone-different spelling place name dictionary, as the difficult-to-choose place name, a place name in mispronounced kana characters as the difficult-to-choose place name, and a distinguishable homograph-different pronunciation place name as a distinguishable difficult-to-choose place name in which an identifier indicating that the place name is a homograph-different pronunciation place name is added to the place name in homograph-different pronunciation kana characters, and the canonical place name that matches the distinguishable homograph-different pronunciation place name. Furthermore, the distinguishable homograph-different pronunciation place name registration step can register the canonical place name that matches the distinguishable homograph-different pronunciation place name and that is linked to the distinguishable homograph-different pronunciation place name, and location information for the canonical place name, in the location database.

[0072] The distinguishable homophonetic place name registration step is executed when the maintenance unit detects a homophonetic place name that is a difficult-to-choose place name and has not yet been stored in the difficult-to-choose place name database. The distinguishable homophonetic place name registration step is performed by the distinguishable homophonetic place name registration unit, which registers, in the homophonetic place name dictionary, a distinguishable homophonetic place name as a difficult-to-choose place name, in which an identifier indicating that the place name is a homophonetic place name is assigned to the distinguishable homophonetic place name, and the regular place name that matches the distinguishable homophonetic place name. Furthermore, the distinguishable homophonetic place name registration step can register, in the location database, the regular place name that matches the distinguishable homophonetic place name and is linked to the distinguishable homophonetic place name. The maintenance method of the present invention allows the voice response system and a response method using the voice response system to collect and store information necessary for the system, enabling more efficient operation.

[0073] The speech recognition engine can be, for example, Amivoice (registered trademark: manufactured by Advanced Media Corporation, hereafter abbreviated), which allows AI speech recognition and customization of server functions. The Amivoice server function can be customized to provide the difficult-to-alternate place name database. Furthermore, the difficult-to-alternate place name database can be provided with a dictionary of common place names (_e), a dictionary of homographs and different pronunciations of place names (_n), and a dictionary of homophones and different spellings of place names (_d). The integrated control device can add identifiers to the Amivoice speech recognition results (text information of difficult-to-alternate place names), such as (_e) for common place names, (_n) for homographs and different pronunciations of place names, and (_d) for homophones and different spellings of place names, to identify difficult-to-alternate place names and register them in the corresponding dictionary of common place names (_e), dictionary of homographs and different pronunciations of place names (_n), or dictionary of homophones and different spellings of place names (_d).

[0074] The location database can be, for example, Elasticsearch (registered trademark: Elasticsearch BV (Netherlands limited liability company, hereafter abbreviated)), a search and analysis engine that enables high-speed searches, in which the same dictionary of common place names (_e) as the difficult-to-choose place name database, a dictionary of homophonic but different pronunciation place names (_n) or a dictionary of homophonic but different spelling place names (_d), and location information for each place name are registered.

[0075] The Elasticsearch location database can have common index information with the Amivoice difficult-to-identify place name database, using the identifier as a key, and can be registered to enable a multi-index search when an output of an difficult-to-identify place name is received from Amivoice. The Elasticsearch can be registered in association with location information (e.g., address, latitude, and longitude) of the difficult-to-identify place name, and the output of the Amivoice difficult-to-identify place name can be linked to the Elasticsearch difficult-to-identify place name to detect the location information (e.g., address, latitude, and longitude) of the difficult-to-identify place name.

[0076] The voice recognition engine and database of difficult-to-choose place names provided in Amivoice and the location database provided in Elasticsearch are currently difficult to integrate due to storage capacity issues, but by increasing the storage capacity of the databases, it is possible to load them onto a single database.

[0077] The speech engine may be a cloud exchange, such as BIZTEL (registered trademark: Link Co., Ltd., hereinafter omitted). The speech engine may include the voice response system and telephone exchange system of the present invention. The location identification inference unit and the inference selection unit may be, for example, an AI commander. The AI ​​commander may have a map search unit in addition to the location identification inference unit and the inference selection unit. The AI ​​commander may have, for example, a database of difficult-to-choose place names, a location database, and a maintenance unit in addition to the location identification inference unit, the inference selection unit, and the map search unit. The map search unit may be a map engine, such as Increment P or Google Maps (registered trademark: Google LLC). The vehicle dispatch system may be, for example, a vehicle dispatch system provided by Cyber ​​Transportation Co., Ltd. (registered trademark: Cyber ​​Transportation Co., Ltd., hereinafter omitted). The AI ​​commander may operate in cooperation with Amivoice, BIZTEL, and Cyber ​​Transportation via an API (application programming interface, hereinafter omitted). [Effects of the Invention]

[0078] The voice response system and response method using the same of the present invention can, when receiving a voice utterance, search for location information for difficult-to-choose place names that are not found in a consumer mapping system in which legitimate place names are registered, or for which multiple correct place names are detected, and respond by voice.The voice response system and response method using the same of the present invention can achieve the excellent effect of efficiently narrowing down to one place name and location information by voice communication when multiple addresses, such as customary place names, place names with the same spelling but different pronunciations, and place names with the same phonetic spellings, are detected when receiving a voice utterance. [Brief explanation of the drawings]

[0079] [Figure 1] Conceptual diagram of the voice response system of the present invention. [Figure 2] Conceptual diagram of the voice response system of the present invention. [Figure 3]1 is a flowchart of a response method using the voice response system of the present invention. [Figure 4] 1 is a flowchart of a response method using the voice response system of the present invention. [Figure 5] FIG. 1 is a sequence diagram of a response method using the voice response system of the present invention. [Figure 6] FIG. 1 is a sequence diagram of a response method using the voice response system of the present invention. [Figure 7] FIG. 1 is a sequence diagram of a response method using the voice response system of the present invention. [Figure 8] 1 is a flowchart of a response method using the voice response system of the present invention. [Figure 9] 1 is a sequence diagram of the system functions of the voice response system of the present invention. [Figure 10] 1 is a sequence diagram of the system functions of the voice response system of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0080] The voice response system according to this embodiment and a response method using the same will be specifically described below with reference to the drawings. In particular, in this embodiment, Fig. 1 shows the basic configuration of a voice response system 1 of the present invention, and Fig. 2 shows a taxi dispatch system 11 incorporating the voice response system 1 of the present invention.

[0081] (Voice response system) 1 and 2, a voice response system 1 of the present invention is connected to a voice interface 9 having a microphone and a speaker. The voice response system 1 has a voice recognition engine 3 that converts voice information of a user 12 received via the voice interface 9 into text information, a speech engine 4 that converts text information into voice information, and an AI commander 1000 as an integrated control device that controls the voice recognition engine 3 and the speech engine 4 via a collaboration interface 2.

[0082] The voice recognition engine 3 and the speech engine 4 may be built into the voice response system 1 together with the collaboration interface 2. Alternatively, as shown in Figures 1 and 2, the voice recognition engine 3 and the speech engine 4 may be provided outside the voice response system 1 and connected to the voice response system 1 via an information and communication network such as the Internet or a LAN. For example, the collaboration interface 2 may be built into an API 1010, the voice recognition engine 3 into Amivoice 1020, and the speech engine 4 into BIZTEL (registered trademark: Link Corporation, hereinafter omitted) 1030 as a cloud exchange.

[0083] The AI ​​Commander 1000 connects to the voice recognition engine 3 and speech engine 4 via the collaboration interface 2, acquires voice information from the user 12 through the voice interface 9 and the speech engine 4 (cloud exchange: BIZTEL 1030), transmits the acquired voice information from the user 12 to the voice recognition engine 3 (Amivoice 1020), and acquires character information including kanji and hiragana information. The AI ​​Commander 1000 generates a reply message to the user 12 based on the acquired character information and speaks it to the user 12 via the collaboration interface 2, speech engine 4, and voice interface 9. The speech engine 4, under the control of the AI ​​Commander 1000, can speak a fixed phrase message pre-registered in the cloud exchange (BIZTEL) 1030 in accordance with a fixed phrase speech program installed in the speech engine 4. Furthermore, the speech engine 4 can be configured to have software, for example, in the cloud exchange (BIZTEL) 1030, that generates and speaks a sentence by adding information (place names, landmark names, address information, etc.) output by the AI ​​commander 1000 to the standard message.

[0084] The voice response system 1 is connected to a consumer-oriented mapping system 10 in which legitimate place names 13 are registered. The mapping system 10 can be multiple map databases (map engines), for example, two or more databases including a consumer-oriented map database 1 (101) and another consumer-oriented map database 2 (102). The mapping system 10 can be, for example, Increment p1040, Google Maps (registered trademark: Google LLC, hereinafter omitted), Yahoo! Maps (LINE Yahoo! Corporation, hereinafter omitted), etc. The voice response system 1 is connected to a difficult-to-choose place name database 5 that stores difficult-to-choose place names 14 that are difficult to search in the mapping system 10 or that are difficult to narrow down to the number of place names that are easy to explain by voice.

[0085] The difficult-to-choose place name database 5 can be set using a customization function of the Amivoice 1020. The voice response system 1 is also connected to a location database 6 that stores location information 15 of the difficult-to-choose place names 14 in the difficult-to-choose place name database 5. The location database 6 can be provided, for example, in a distributed search and analysis engine capable of high-speed searches, and more specifically, in Elasticsearch 1050. The AI ​​commander 1000 of the voice response system 1 has a map search unit 1004 and a location identification inference unit 7 that are connected to the difficult-to-choose place name database 5 and the location database 6, and search the difficult-to-choose place name database 5 and the location database 6 for information including the difficult-to-choose place names 14 and location information (e.g., address, latitude, longitude) 15 of the difficult-to-choose place names 14, and narrow down the number of place names that are easy to explain by voice (e.g., 1 to 5, preferably 1 to 3).

[0086] The AI ​​commander 1000 is a user interface (UI) that enables customization of a dialogue flow according to the purpose of the administrator of the voice response system 1. The dialogue flow can be executed by the cloud exchange, such as BIZTEL 1030, controlled by the AI ​​Commander 1000. The dialogue flow can be customized to include an initial response unit 1003 that queries the user 12 for user information and a place name, for example, when a voice response is initiated or a call line is connected. The dialogue flow of the AI ​​Commander 1000, for example, involves the AI ​​Commander 1000, the voice recognition engine 3, and the speech engine 4 cooperating with each other via the cooperation interface 2. The AI ​​Commander 1000 can be customized to reference the mapping system 10, the difficult-to-choose place name database 5, and the location database 6. The AI ​​Commander 1000 can be customized to include a map search unit 1004 that searches the mapping system 10, the difficult-to-choose place name database 5, and the location database 6.

[0087] The AI ​​Commander 1000 may have an initial response unit 1003 that queries the user 12 for user information and place names, and a map search unit 1004 that can search the difficult-to-alternate place name database 5, the location database 6, and the mapping system 10 for place names and location information for the place names obtained by the initial response unit 1003. The AI ​​Commander 1000 may have a transfer unit 1005 that transfers processing to the location identification inference unit 7 when the search by the map search unit 1004 fails to narrow down the number of detected place names to 1 to 5, preferably 1 to 3, that are easy to explain by voice, or when no place names are detected. The AI ​​commander 1000 has a location identification inference unit 7 connected to the difficult-to-alternate place name database 5, the location database 6, and the mapping system 10, and in particular performs a multi-index search on the multi-index data described below in the difficult-to-alternate place name database 5 and the location database 6, thereby enabling faster search for location information 15 (address, latitude and longitude) of the difficult-to-alternate place name 14.

[0088] The AI ​​commander 1000 may have a selection section 1006 that, when the map search section 1004 detects 1 to 5, preferably 1 to 3, detected place names that are easy to explain by voice, assigns a voice code (such as a number from 1 to 5 or 1 to 3) to each of the searched legitimate place names 13 and the location information 15 of the legitimate place names 13, and outputs a question to the user 12 requesting a selection, and a confirmation section 1007 that outputs a question (refuting evidence) to confirm the correct answer to the one place name and location information selected by the user 12 in response to the question in the selection section 1006.

[0089] As shown in Figures 1, 2 and Tables 1 to 3, the difficult-to-choose place name database 5 has the commonly used place name dictionary 50, the homophonic, different pronunciation place name dictionary 51 and the homophone, different spelling place name dictionary 52, and the commonly used place name dictionary 50, the homophonic, different pronunciation place name dictionary 51 and the homophone, different spelling place name dictionary 52 store identified difficult-to-choose place names 141 in which an identifier 140 is assigned to the difficult-to-choose place names 14 in character information, for example, kana character information.

[0090] The common place name dictionary 50 stores, in association with the common place names of the character information as the difficult-to-alternate place names 14, identifiable common place names as difficult-to-alternate place names 141 with an identifier 140 (e.g., "_e") indicating that it is a common place name, and the regular place names 13 that match the distinguishable common place names.

[0091] The homograph-but-different-pronunciation place name dictionary 51 stores, in association with the difficult-to-choose place names 14, place names with mispronunciation character information, and identifiable homograph-but-different-pronunciation place names as identifiable difficult-to-choose place names 141, which are homograph-but-different-pronunciation place names with an identifier 140 (e.g., "_n") added to the homograph-but-different-pronunciation place names, and the regular place names 13 that match the identifiable homograph-but-different-pronunciation place names.

[0092] The homophone place name dictionary 52 stores, in association with the difficult-to-choose place names 14, place names with character information of homophones, and place names with insufficient character information, as distinguishable difficult-to-choose place names 141, with an identifier 140 (for example, "_d") added to indicate that the place name is a homophone place name, and the regular place names 13 that match the distinguishable homophone place names.

[0093] The difficult-to-choose place name database 5 can further store, in addition to the information related to the place names, user information 120 linked to it. For example, the user information 120 can be linked to and registered with a difficult-to-choose place name 14 that is due to the pronunciation habits or misreading of kanji characters that are unique to the user 12, or the user information 120 of a user 12 who gets on or off at the same place (difficult-to-choose place name 14) every time can be linked to the difficult-to-choose place name 14 and registered.

[0094] The AI ​​Commander 1000 can be customized to allow registration of spoken text (hiragana information) and recognition result text (character information including kanji) in each of the dictionaries 50, 51, and 52. The AI ​​Commander 1000 can record and store in one of its storage units the portions of the user's 12 speech data that have been recognized as place names and the portions that have not been recognized as place names, and can refer to and compare the recorded data when a place name cannot be identified, thereby improving the accuracy and speed of place name recognition. The customization of registration in the difficult-to-choose place name database 5 and the location database 6 can be performed by registering the spoken item (spoken) and the output item (written) in separate fields, as shown in FIG. 8. Furthermore, the customization of registration can also include a text item (text). The output of the speech recognition engine 3 can be customized so that the difficult-to-identify place name 141 with the identifier 140 (any one of _e, _d, _n) added thereto is added to the output item (written). The identifier 140 is not limited to "_e", "_d", or "_n", and can be any identifier that enables multi-index search, and can be, for example, a character, character string, or symbol separated by "_", ".", ";", ":", or other symbols to the right of the character string of the difficult-to-identify place name 14.

[0095] The index information of the difficult-to-identify alternative place name 141 is the identifier 140 (for example, any one of _e, _d, _n). If there are no problems with the storage capacity of the difficult-to-identify alternative place name database 5 or the CPU processing speed, information such as the official place name and location information (address, latitude and longitude) can also be recorded in the difficult-to-identify alternative place name database 5 in a lump sum.

[0096] As shown in FIG. 2, the position identification inference unit 7 can have an inference selection unit 70 and be connected to a position database 6 that is searched through the inference selection unit 70. The position database 6 can be installed inside the voice response system 1 or can be located outside the voice response system 1 and connected through a communication network.

[0097] For the official place name 13 and the difficult-to-identify alternative place name 14, for example, in Kyoto City, the intersection name of "Sanjo Street" and "Kiyamachi Street" is called "Sanjo Kiyamachi" or "Kiyamachi Sanjo" differently by people. Not only the names (place names) of the place names are different, but the Chinese character "七条" may also be read as "しちじょう" or "ななじょう". Such derived words (place names) can be registered as multi-index information in the difficult-to-identify alternative place name database 5 (the common place name dictionary 50, the same notation different pronunciation place name dictionary 51, and the same pronunciation different notation place name dictionary 52) provided by customizing the voice recognition engine 3.

[0098] Table 1 is an excerpt of a list of intersection names (place names) with "七" registered in the difficult-to-identify alternative place name database 5. Table 2 illustrates differences in the names of places familiar to people around the area, the pronunciation of Chinese characters, using the names of east-west and north-south streets and intersection names as examples. Table 3 shows registration examples of the common place name dictionary 50, the same notation different pronunciation place name dictionary 51, the same pronunciation different notation place name dictionary 52, and additional geographic information 142 such as landmarks in the difficult-to-identify alternative place name database 5, and registration examples of the position information 15 (address, latitude and longitude) corresponding to the common place name dictionary 50, the same notation different pronunciation place name dictionary 51, and the same pronunciation different notation place name dictionary 52 and additional geographic information 142 such as landmarks in the position database 6.

[0099] [Table 1]

[0100] [Table 2]

[0101] [Table 3]

[0102] 2, the voice response system 1 may include a maintenance unit 8 that updates various information including the difficult-to-choose place names 14 and location information (address, latitude and longitude) 15 of the difficult-to-choose place names 14 registered in the difficult-to-choose place name database 5 and location database 6. The AI ​​commander 1000 may be connected to an external guidance system 11, such as an automobile navigation system installed in a vehicle, an automatic driving system for a vehicle, or a taxi dispatch system (11).

[0103] As shown in FIG. 2, the voice response system 1 is connected to, for example, a taxi dispatch system 11 via a BIZTEL 1030, which serves as a telephone switching system 91. The voice interface 9 can be, for example, a telephone network 90, including the Internet, optical fiber lines, and wireless communication networks, through which calls can be made with a user 12. The taxi dispatch system 11 can be installed outside the voice response system 1. The taxi dispatch system 11 is connected to an operator's PC 16 so as to be able to send and receive text information (such as a legitimate place name 13, location information 15, and user information 120). Furthermore, the operator's PC 16 can be connected to the telephone switching system 91 so that an operator 17 can directly talk to a user 12.

[0104] The location database 6 can store the difficult-to-alternate place names 14 that match the registered information in the difficult-to-alternate place name database 5, the difficult-to-identify place names 141, the correct place names 13 that match the difficult-to-identify place names 141, location information 15 such as the addresses of the correct place names 13 and the latitude and longitude of the correct place names, and additional geographic information 142 linked to the correct place names 13 and the location information 15. Furthermore, user information 120 can be associated with information related to these place names and stored.

[0105] The location identification inference unit 7 can include a specific information addition unit 71 that outputs a question about additional geographic information 142 to the user 12 when processing is transferred from the transfer unit 1005, an inference selection unit 70 that adds the additional geographic information 142 to the place names to narrow down the number of detected place names that are easy to explain by voice, and an output unit 72 that sends the search results to the selection unit 1006 when the inference selection unit 70 searches for the difficult-to-choose place names 14, the legitimate place names 13 of the difficult-to-choose place names 14, and the location information 15 of the legitimate place names 13, and detects a number of detected place names that are easy to explain by voice (for example, 1 to 5, preferably 1 to 3).

[0106] 2, the voice response system 1 has an AI commander (integrated control device) 1000 consisting of an AI commander 1001 in a broad sense including an AI commander 1002 in a narrow sense, and for example, the speech recognition engine 3 can be Amivoice 1020, and the difficult-to-choose place name database 5 can be provided using the customization function of Amivoice 1020. The speech engine 4 can be BIZTEL 1030, the collaboration interface 2 can be API 1010, the map engine 1040 can be Increment p, the location database 6 can be Elasticsearch 1050, and the vehicle dispatch system 11 can be Cyber ​​Transportation 1060.

[0107] (Response method using voice response system 1) As shown in FIGS. 2 and 3, the response method using the voice response system 1 of the present invention includes a reasoning search step 202 in which the map search unit 1004 under the control of the AI ​​commander 1000 performs a multi-index search for the difficult-to-identify alternative place name 141 in the difficult-to-choose place name database 5 and the location database 6, and a location determination reasoning unit 7 performs a search for additional geographic information 142 in the difficult-to-choose place name database 5 and the location database 6 if the number of detected place names that are easy to explain by voice cannot be narrowed down to a certain number or if no place names are detected. the specific information adding step 205 for outputting the question, an inference selecting step 207 for adding the additional geographical information 142 obtained in the specific information adding step 205 to the difficult-to-identify-choice place name 141, and for the location identifying inference unit 7 to search the difficult-to-choose place name database 5 and the location database 6 to narrow down the number of detected place names to those that are easy to explain by voice, and an output step 209 for transmitting the regular place names 13 and the additional geographical information 142 or addresses 15 narrowed down in the inference selecting step 207 to the selection section 1006.

[0108] As shown in Figures 2 and 3, when the voice response system 1 of the present invention is connected to a taxi dispatch system 11, it can operate as shown in Figures 3 and 9. In the initial response process 200, when a user 12 calls the voice response system 1, the AI ​​commander 1000 detects the input and executes an initial response. The speech engine 4 (e.g., BIZTEL 1030) can be configured to have a customized initial response unit 1003. When a user 12 calls the voice response system 1, the initial response unit 1003 executes an initial response and transmits voice (recorded) information of the user 12 to the AI ​​commander 1000.

[0109] The AI ​​commander 1000 transmits the received voice information to the voice recognition engine 3 (e.g., Amivoice 1020), and the voice recognition engine 3 outputs both character information including kanji and hiragana information from the voice information to the AI ​​commander 1000. If the AI ​​commander 1000 correctly recognizes the voice from the user 12, it outputs an initial response sentence in kana characters. The speech engine 4 that receives the initial response sentence outputs aloud, "Please tell us the name of the person receiving the guest, such as the address, facility name, building name, and building number, excluding the entrance name and building number." If the speech engine 4 can read out information including kanji and other characters, the output signal to the speech engine 4 can be various character information, not limited to kana character information, such as a mixture of kanji and English characters.

[0110] In the reply confirmation step 201, the AI ​​commander 1000 determines whether the destination place name is included in the response of the user 12 to the initial response step 200. If the place name cannot be recognized because the voice volume is too low or there is too much noise, the AI ​​commander 1000 returns to the initial response step 200. The AI ​​commander 1000 may have a place name identification unit 1008 that identifies place names from character information including kanji. When the AI ​​commander 1000 (map search unit 1004) recognizes the character information (for example, hiragana) as "Takeda Hospital," it executes the inference search step 202. In the reply confirmation step 201, when the voice recognition engine 3 recognizes the place name, it can repeat the place name to confirm with the user 12 before executing the inference search step 202.

[0111] In the inference search step 202, the map search unit 1004 searches the difficult-to-identify place name 14 "Takeda Hospital" in the common place name dictionary 50 "_e", the homophonic and different pronunciation place name dictionary 51 "_n", and the homophone and different spelling place name dictionary 52 "_d" in the difficult-to-identify place name database 5, and detects the difficult-to-identify place name 141 "Takeda Hospital_d". The "_d" added to the difficult-to-identify place name 141 is registered in the homophonic and different spelling place name dictionary 52, and as a result of the search, three incomplete place names (difficult-to-identify place names 141) "Takeda Hospital_d" can be detected.

[0112] Furthermore, in the inference search process 202, the map search unit 1004 does not determine whether the place name is a common place name (_e), a homophonic but differently pronounced place name (_n), or a homophoneically but differently spelled place name (_d), but searches the difficult-to-distinguish place name database 5 (common place name dictionary 50, homophonic but differently pronounced place name dictionary 51, and homophoneically but differently spelled place name dictionary 52) for each of the difficult-to-distinguish place names 141, which are the difficult-to-distinguish place names 14 with the identifiers 140 "_e", "_n", and "_d" attached, and can determine from the detection results that the difficult-to-distinguish place name 141 "Takeda Hospital _d" is a homophoneically but differently spelled place name (_d).

[0113] Furthermore, in the inference search process 202, the map search unit 1004 determines whether the place name answered by the user 12 is a common place name (_e), a place name with the same spelling but different pronunciation (_n), or a place name with the same spelling but different spelling (_d), and then, for example, can search only the homonymous place name dictionary 52 for the difficult-to-identify place name 141 (the difficult-to-identify place name 14 + identifier 140 "_d").

[0114] In the inference search step 202, the map search unit 1004 or the location identification inference unit 7 can perform a multi-index search of the location database 6 for location information 15 (address, latitude and longitude) registered as information matching the difficult-to-identify place name 141 "Takeda Hospital_d" based on the difficult-to-identify place name 141 "Takeda Hospital_d" detected from the difficult-to-identify place name database 5. For the difficult-to-identify place name 141 with the identifier 140 "_e", only the commonly used place name dictionary 50 is searched, for the difficult-to-identify place name 141 with the identifier 140 "_n", only the homonymous place name dictionary 51 is searched, and for the difficult-to-identify place name 141 with the identifier 140 "_d", only the homonymous place name dictionary 52 is searched. This increases the search speed and enables a response to the user 12 in a shorter time.

[0115] In the inference search process 202, the map search unit 1004 can search for related information such as place names and location information from at least one of the mapping system 10 including the consumer map database 1 (101) and another map database 2 (102), the difficult-to-alternate place name database 5, and the location database 6. The inference search process 202 can search for related information such as place names and location information from all of the mapping system 10, the difficult-to-alternate place name database 5, and the location database 6 simultaneously or sequentially.

[0116] The inference search process 202 searches for related information such as place names and location information from the difficult-to-choose place name database 5 and the location database 6, and if 1 to 5, preferably 1 to 3, place names that are easy to explain by voice are not detected, it can search for related information such as place names and location information from the mapping system 10. Furthermore, the inference search process 202 searches for related information such as place names and location information from the mapping system 10, and if 1 to 5, preferably 1 to 3, place names that are easy to explain by voice are not detected, it can search for related information such as place names and location information from the difficult-to-choose place name database 5 and the location database 6.

[0117] The map search unit 1004 searches the location database 6 for the incomplete place name (difficult-to-identify-and-choose place name 141) "Takeda Hospital_d" and detects each regular place name 13 and location information (address, latitude and longitude). The detection number determination step 203 detects three occurrences of "Takeda Hospital_d", and therefore determines that the number of place names that are easy to explain by voice is 1 to 5, preferably 1 to 3.

[0118] In the selection step 210, the selection part 1006 assigns a phonetic code (such as a number from 1 to 3) to each legitimate place name 13 based on the three detection results, and outputs (kana character information) a question requesting the user 12 to make an alternative selection to the speech engine 4. The speech engine 4 speaks, for example, "Sorry to keep you waiting. There are three options. First, Kyoto Totakeda Hospital, second, third, third, third, injury, stab, stab, one of the three? Please answer the number you are selecting. If you want to start the search again, say 'back,' 'again,' or 'no,' and if you want to ask again, say 'again.'" The alternative selection candidates can be presented either as phonetic codes (such as numbers 1 to 3) attached to each legitimate place name 13 and location information 15 such as the address, building name, latitude and longitude of the legitimate place name 13 that can be used as a reference for making a decision, or as additional geographic information 142 such as landmarks near the legitimate place name 13.

[0119] When the user 12 responds to the question in the multiple choice section 1006, the multiple choice section 1006 executes a selected answer determination step 211. The selected answer determination step 211 determines whether the user 12 has responded by selecting any one of the options 1 to 3, and if the user 12 has not responded correctly, the process returns to the multiple choice step 210. If the user 12 responds correctly, for example, "number one," the confirmation section 1007 executes a confirmation step 212.

[0120] In the confirmation step 212, the confirmation unit 1007 confirms the information in No. 1 by reading it out loud. The confirmation unit 1007 outputs text information to the speech engine 4, such as "Kyoto Takeda Hospital, right? Sakura Yakkyoku and Kyoto Nishi-shi Hospital are nearby. If it's OK, please answer 'Yes'. If it's not, please answer 'No'."

[0121] When the user 12 answers the question in the confirmation step 212, the place name identification unit 1008 performs a correct answer confirmation step 213. If the user 12 answers with a word other than "yes" that includes "no", the correct answer confirmation step 213 returns to the selection step 210. When the user 12 answers "yes", the correct answer confirmation step 213 informs the user 12 that the request has been received and ends the response. When the user 12 answers "yes", the correct answer confirmation step 213 executes a dispatch request to the taxi dispatch system (guidance system) 11. In the unlikely event that the taxi dispatch system (guidance system) 11 does not operate normally, the operator 17 can respond to the user 12 and the taxi dispatch system (guidance system) 11 via the operator PC 16 and the telephone switching system 91.

[0122] FIG. 10 also shows the responses from the initial response step 200 to the correct answer confirmation step 213 to the user 12 when the difficult-to-identify alternative place name 141 is a homophoneically spelled place name, for example, "saiganji_d." There is a custom of using the same pronunciation of place names or abbreviated facility names to represent multiple buildings or facilities that are not the same. For example, in Kyoto, similar to the aforementioned "Takeda Hospital," there are multiple facilities such as "Saiganji" and "Saiganji" that are uttered in the same utterance for the temple "saiganji." To solve this problem, the identifier 140 "_d" is added to the difficult-to-identify alternative place name 14, which defines the difficult-to-identify alternative place name 141 as having multiple candidates.

[0123] In the response of the correct answer confirmation step 213, it is possible to confirm the differences in kanji in place names and the official names by voice, but since it is difficult to explain using the radicals and structures of kanji in conversation and it is unclear whether the message will be conveyed, as mentioned above, the distinction is made by speaking up to the block level of the address. In the response of the correct answer confirmation step 213, for example, category information such as station names and line names, station names and city / ward / town / village names, or store names and branch names can be added to the place names to distinguish them.

[0124] 2, 3, and 5, the transfer unit 1005 executes a transfer step 204 when the number of easily voice-explained place names of 1 to 5, preferably 1 to 3, is not detected in the detection number determination step 203. The transfer step 204 transfers the search narrowing down to the number of easily voice-explained place names of 1 to 5, preferably 1 to 3, to the location identification inference unit 7. In the location identification inference unit 7 that has received the search transfer in the transfer step 204, the specific information addition unit 71 executes a specific information addition step 205. The specific information addition step 205 outputs a query for additional geographic information 142 to the user 12 via the speech engine 4.

[0125] Upon receiving a response from the user 12, the location identification reasoning unit 7 executes an information confirmation step 206 to confirm whether the user 12 has responded with additional geographic information 142. If the user 12 has not responded with additional geographic information 142, the location identification reasoning unit 7 returns to the specific information addition step 205. If the user 12 has responded with additional geographic information 142 in the information confirmation step 206, the location identification reasoning unit 7 executes an inference selection step 207. The inference selection step 207 searches the difficult-to-choose place name database 5 based on the difficult-to-identify place name 141 and the additional geographic information 142.

[0126] The common place name dictionary 50, the homophonic place name dictionary 51, and the homophonic place name dictionary 52 for the difficult-to-identify place names 141 store the difficult-to-identify place names 14, the difficult-to-identify place names 141 with identifiers 140, the correct place names 13, and additional geographic information 142 such as nearby landmarks, all linked together, so that a multi-index search can be performed based on each piece of information.When the inference selection step 207 has completed the search, the location identification inference unit 7 executes the detection number determination step 208.

[0127] The detection number determination step 208 determines whether the inference selection step 207 has detected 1 to 5, preferably 1 to 3, place names that are easy to explain by voice, and if the inference selection step 207 has not detected the number of place names that are easy to explain by voice, the process returns to the specific information addition step 205. If the inference selection step 207 has detected the number of place names that are easy to explain by voice, the process transfers the subsequent processing to the selection step 210.

[0128] As shown in Figures 1 to 4, the common place name dictionary 50, the homophonic and different pronunciation place name dictionary 51, and the homophonic and different spelling place name dictionary 52 of the difficult-to-choose place name database 5 register the difficult-to-choose place names 14, identifiers 140, difficult-to-identify place names 141, additional geographic information 142, and user information 120, as well as legitimate place names 13. Furthermore, the location database 6 also registers location information (address, latitude and longitude) 15 linked to the difficult-to-choose place names 14, identifiers 140, difficult-to-identify place names 141, additional geographic information 142, user information 120, and legitimate place names 13.

[0129] If the difficult-to-choose place name database 5 has a difficult-to-choose place name 14 registered, it can detect the registered difficult-to-choose place name 14 even when the user 12 utters the correct place name 13 of the difficult-to-choose place name 14. In this case, the AI ​​commander 1000 can detect location information (address, latitude and longitude) 15 from the location database 6 in a shorter time without searching the mapping system 10.

[0130] The voice recognition engine 3 converts the voice of the user 12 into kana character information, and when the AI ​​commander 1000 recognizes a place name in the converted kana character information, it outputs the place name to the location identification and inference unit 7 (map search unit 1004). The location identification and inference unit 7 can search place names independently, or can have the map search unit 1004 perform the search on its behalf. Furthermore, the location identification and inference unit 7 can simultaneously search different databases, map engines, etc. together with the map search unit 1004. The location identification inference unit 7 (map search section 1004) that receives the place name in the kana character information searches the difficult-to-choose place name database 5 to determine whether the place name in the kana character information is a registered difficult-to-choose place name 14 or an unregistered legitimate place name 13, and if it determines that the place name is a legitimate place name 13 that is not registered in the difficult-to-choose place name database 5, the location identification inference unit 7 (map search section 1004) can search the map database 1 (101) and another map database 2 (102) of the mapping system 10 for the legitimate place name 13 and location information 15 (address, latitude and longitude).

[0131] Furthermore, when the location identification inference unit 7 (map search unit 1004) determines that the place name of the kana character information is a difficult-to-choose place name 14, the location identification inference unit 7 (map search unit 1004) performs a multi-index search of the difficult-to-choose place name database 5 and the location database 6 to find a difficult-to-distinguish place name 141 obtained by adding an identifier 140 to the difficult-to-choose place name 14, and can search for the difficult-to-distinguish place name 14, the difficult-to-distinguish place name 141, the correct place name 13, location information 15 (address, latitude and longitude), and additional geographic information 142 such as landmarks.

[0132] Furthermore, the location identification inference unit 7 (map search unit 1004) that receives the place name in the kana character information can search for the difficult-to-alternate place name 14, the difficult-to-identify place name 141, the correct place name 13, location information 15 (address, latitude and longitude) and additional geographic information 142 such as landmarks simultaneously or sequentially, without determining whether the place name in the kana character information is a regular place name 13 or a difficult-to-alternate place name 14.

[0133] As shown in Figure 6, each of the dictionaries 50 to 52 of the difficult-to-choose place name database 5 can register, as the difficult-to-choose place names 14, misspelled place names, for example, place names that are slang expressions such as "Dai-nisseki" (Second Red Cross Hospital) to "Kyoto Second Red Cross Hospital," abbreviated place names such as "Takeda Byouin," place names with different sounds (which are only apparent because the voice response system 1 only supports hiragana), place names in the additional geographic information 142 such as landmarks not included in the consumer-oriented mapping system 10, etc.

[0134] The dictionaries 50 to 52 of the difficult-to-choose place name database 5 store, in association with each other, the difficult-to-choose place names 14, the difficult-to-distinguish place names 141 obtained by adding an identifier 140 to the difficult-to-distinguish place names 14, and the regular place names 13 registered in the mapping system 10. The identifier 140 of the difficult-to-distinguish place names 141 registered in the common place name dictionary 50 can be, for example, "_e." The identifier 140 in the homophonic, different pronunciation place name dictionary 51 can be, for example, "_n," and the identifier 140 in the homophonic, different spelling place name dictionary 52 can be, for example, "_d." The identifiers 140 are not limited to these, and can be letters, symbols, signs, etc. that are easy to distinguish so that mispronunciations are unlikely to occur, or digital classifications on the system (for example, flags, classifications, codes, etc.). In particular, using flags for the identifiers 140 can make searches more efficient.

[0135] The difficult-to-identify-and-alternate place names 141 in the difficult-to-alternate place name database 5 and the difficult-to-identify and alternative place names 141 and location information 15 (address, latitude and longitude) in the location database 6 are stored so that the identifiers 140 match for each place name. As shown in FIGS. 1 and 2, the voice response system 1 can have a maintenance unit 8 that maintains and manages the information stored in the difficult-to-alternate place name database 5 and the location database 6. The maintenance unit 8 allows an operator 17 to update the registration of the stored information, for example, approximately once every six months to one year, via an operator PC 16, while referring to the recording information and location information stored by the AI ​​commander.

[0136] The maintenance unit 8 may have, for example, a distinguishable commonly used place name registration unit 80, a distinguishable homophonic, differently pronounced place name registration unit 81, and a distinguishable homophonic, differently spelled place name registration unit 82, all of which are made up of software, and when each of the registration units 80-82 detects a difficult-to-alternate place name 14 that is not stored in each of the dictionaries 50-52 of the difficult-to-alternate place name database 5, it may update the stored information of each of the dictionaries 50-52 of the difficult-to-alternate place name database 5 and the location database 6 without the intervention of the operator 17. Each of the registration units 80-82 may have a periodic execution unit 83 that can set a period for periodic updates, and the periodic execution unit 83 may update the stored information at intervals (periods) input and set by the operator 17 or the administrator of the voice response system 1. The function of the periodic execution unit 83 to set the periodic update period allows data updates to be performed frequently when the voice response system 1 is first introduced, allowing more difficult-to-choose place names 14 to be registered more quickly, and when the frequency of new information to be registered decreases due to long-term use, the update period can be set longer, allowing the system to be used more efficiently according to the usage situation.

[0137] 7, at least one of the difficult-to-choose place name database 5 and the location database 6 may have a learning function unit 84 made up of software. The learning function unit 84 may, for example, record the names of places and times of boarding and disembarking for each user 12 when going out for a regular medical visit, shopping, etc., in association with information such as the name and telephone number of each user 12, and periodically update the information, thereby enabling smoother individual responses from each user 12. In addition, the learning function unit 84 may store the difficult-to-choose place name 141 and the desired arrival time of each user 12 heading to the same difficult-to-choose place name 141, thereby enabling, for example, dispatching a vehicle to arrive at a time that matches the stop time of a station, bus stop, etc.

[0138] 3 and 8, when the user 12 speaks the place name "Motomochi Animal Hospital" in the initial response step 200 of the voice response system 1, the AI ​​commander 1000 performs the inference search step 202. In the inference search step 202, the AI ​​commander 1000 searches the difficult-to-alternate place name database 5 for the place name "Motomochi Animal Hospital" and detects the difficult-to-distinguish alternative place name 141 "Motomochi Animal Hospital_e." Furthermore, the AI ​​commander 1000 performs a multi-index search of the difficult-to-distinguish alternative place name 141 "Motomochi Animal Hospital_e" in the location database 6 and queries the regular place name 13 and location information 15 (address, latitude and longitude) of the place name "Motomochi Animal Hospital."

[0139] Furthermore, in the inference search step 202, the AI ​​commander 1000 detects the search result of the difficult-to-alternate place name database 5, the legitimate place name 13 "Motomochi Animal Hospital," and the map search unit 1004 searches the map database 1 (101) and another map database 2 (102) of the mapping system 10 based on the legitimate place name 13 "Motomochi Animal Hospital," and can inquire about the location information 15 of "Motomochi Animal Hospital." Since only the legitimate place names 13 and not the difficult-to-alternate place names 14 are registered in the mapping system 10, in most cases when searching for place names based on the response of the user 12, the legitimate place name 13 and location information 15 cannot be inquired about without searching the difficult-to-alternate place name database 5 and the location database 6.

[0140] Furthermore, unlike the mapping system 10 for consumers, the difficult-to-choose place name database 5 and the location database 6 are dedicated to the voice response system 1, so there is no time lag due to concentrated access, and responses to the user 12 can be made more quickly and stably.

[0141] As shown in Figure 8, information obtained in the process of the voice response system 1 responding to the user 12 is recorded as user information 120 in at least one of the difficult-to-choose place name database 5 and the location database 6 in the form of data on customer utterances, etc., and various information such as destination landmarks, destinations, and destinations can also be registered as additional geographic information 142.

[0142] Kyoto has approximately 280 streets running north-south, east-west, and like a grid, and it is customary to write addresses using the names of their intersections. An easy-to-understand example of this is Shijo and Shichijo, which are pronounced "shijo" and "shichijo," respectively, and are virtually indistinguishable in speech. Some users 12 mistakenly pronounce "shichijo" as "shijo," intentionally pronouncing it as "yonjo" or "nanajo" to prevent mispronunciation. Another example is "Nakagyo (Chukyo) Ward," but some users 12, although not necessarily Kyoto residents, prefer to pronounce it as "chukyo-ku." When such an utterance is made, the voice response system 1 distinguishes between them by adding the "_n" identifier 140 and registering it in the place name dictionary 51.

[0143] As shown in Figures 2 and 3, for the difficult-to-identify place name 141, which is the difficult-to-identify place name 14 with an underscore (identifier 140) "_n," the selection process 210 assumes there are multiple candidates and, based on this assumption, adds phonetic symbols (e.g., numbers 1 through 3) to distinguish between them, vocally uttering them to prompt the user 12 to make a selection. In the confirmation process 211 and correct answer confirmation process 212, the first step is to provide the address down to the block. If this is difficult, landmarks (additional geographic information 142) near the latitude and longitude of the place name are searched and presented. This is spoken to the user 12 as a form of judgment using refutation evidence, and the user 12 is asked to respond and confirm the address to determine whether it is correct. To improve the accuracy of this process, the area can be limited by entering a telephone number, and the target area name (e.g., "Kyoto") can be added to the search term for confirmation. Furthermore, homophoneous place names can be confirmed by adding address evidence of the correct place name 13. [Industrial Applicability]

[0144] The voice response system of the present invention and the response method using the same can be used in order taking technology for goods or services via a voice interface, a guidance system via a voice interface, a navigation system operated via a voice interface, and the control of an autonomous vehicle operated via a voice interface. [Explanation of symbols]

[0145] 1. Voice response system 2 Collaboration Interface 3. Speech Recognition Engine 4. Speech Engine 5. Database of difficult-to-choose place names 50 Dictionary of Common Place Names (_e) 51 Dictionary of Place Names with the Same Spelling and Different Pronunciations (_n) 52 Dictionary of Homophones and Place Names (_d) 6. Location Database 7 Location Reasoning Unit 70 The same reasoning selection department 71 Specific information addition section 72 Output section 8 Maintenance Unit 80 Identifiable and Common Place Names Registration Department 81 Identification of Place Names with the Same Spelling and Different Pronunciations 82 Homophone Place Name Registration Department 83 Periodic execution section that can be set for the same period 84 Learning Functions Department 9. Voice Interface 90 Telephone network 91 Telephone Switching System 10. Mapping System 101 Map Database 1 102 Map Database 2 11 Taxi dispatch system (guidance system) 12 users 120 User Information 13 Official place names 14 Difficult place names to choose from 140 Same identifier 141 Difficult to distinguish between place names 142 Additional geographical information 15 Location information (address, latitude and longitude) 16 Operator PC 17 Operator 1000 AI Commander (Integrated Control Device) 1001 The same broad AI commander 1002 AI Commander in the Narrow Sense 1003 Initial Response Unit 1004 Map Search Department 1005 Transfer Department 1006 Same Choice Part 1007 Confirmation Department 1008 Place Name Identification Department 1010 API (application programming interface) 1020 Amivoice (registered trademark: Advanced Media Co., Ltd.) 1030 BIZTEL (registered trademark: Link Co., Ltd.) 1040 Increment P (Registered Trademark: Geo Technologies Co., Ltd.) 1050 Elasticsearch (registered trademark: Elasticsearch BV (Netherlands limited liability company) 1060 Dennou Kotsu (Registered Trademark: Dennou Kotsu Co., Ltd.) (Response method using voice response system 1) 200 Initial response process 201 Response confirmation process 202 Inference search process 303 Detection number determination process 204 Transfer Process 205 Specific information addition process 206 Information confirmation process 207 Reasoning Selection Process 208 Detection number determination process 209 Output Process 210 Selection Process 211 Selection decision process 212 Confirmation process 213 Correct answer confirmation process

Claims

1. a speech recognition engine that converts speech information received via a speech interface into text information; A speech engine that converts text information into audio information; A database of difficult-to-choose place names that is difficult to search using a mapping system where legitimate place names are registered, or that is difficult to narrow down to a number of place names that are easy to explain by voice; a location database that stores location information of difficult-to-choose place names in the difficult-to-choose place name database; an integrated control device that searches the difficult-to-choose place name database and the location database for the difficult-to-choose place names and location information of the difficult-to-choose place names, and narrows down the number of place names that can be easily explained by voice; A voice response system having:

2. The database of difficult-to-choose place names includes a dictionary of commonly used place names, a dictionary of place names with the same spelling but different pronunciations, and a dictionary of place names with the same spelling but different pronunciations, The common place name dictionary, the homophonic place name dictionary, and the homophonic place name dictionary store identified difficult-to-choose place names in which identifiers are assigned to the difficult-to-choose place names in character information. The common place name dictionary stores, in association with each other, discriminative common place names as distinguishable difficult-to-alternate place names, each of which is a common place name of character information as the difficult-to-alternate place name, and an identifier indicating that the common place name is a common place name, and the regular place name that matches the discriminative common place name; The homograph-different pronunciation place name dictionary stores, in association with the place names of misreading character information as the difficult-to-choose place names, and discriminative homograph-different pronunciation place names as discriminative difficult-to-choose place names in which an identifier indicating that the place names are homograph-different pronunciation place names is attached to the place names of homograph-different pronunciation character information, and the legitimate place names that match the discriminative homograph-different pronunciation place names; The homonym place name dictionary stores, in association with the place names with character information of homonyms as the difficult-to-choose place names and the regular place names that match the identified homonym place names, the homonym place names as the difficult-to-choose place names, and the regular place names that match the identified homonym place names. The voice response system according to claim 1.

3. The location database stores the difficult-to-identify-and-alternate place names, the regular place names that match the difficult-to-identify and alternative place names, and location information of the regular place names in association with each other. The voice response system according to claim 1.

4. The integrated control device an initial response unit that queries the user for user information and place names; a map search unit that searches the mapping system for the place names and location information of the place names obtained by the initial response unit; a location inference unit for searching the place name by adding additional geographic information; a transfer unit that transfers processing to the location specification inference unit when the search by the map search unit fails to narrow down the number of detected place names that are easy to explain by voice or when no place names are detected; a selection section that, when the map search section detects a number of detected place names that are easy to explain by voice, assigns a voice code (e.g., a number from 1 to 3) to each legitimate place name and location information of the legitimate place name obtained by a multi-index search of the difficult-to-choose place name database and the location database, and outputs a question to the user requesting a selection; a confirmation unit that outputs a question to confirm the correct answer of one place name and location information selected by the user in response to the question in the multiple-choice section; 2. The voice response system according to claim 1, comprising:

5. The location inference unit: a specific information adding unit that outputs a query for additional geographic information to a user when the process is transferred from the map search unit; an inference selection unit that adds the additional geographical information to the place names and narrows down the number of detected place names that can be easily explained by voice from the difficult-to-choose place name database; an output unit that transmits a search result to the selection unit when the inference selection unit performs a multi-index search for the difficult-to-choose place names, the correct place names of the difficult-to-choose place names, and the location information of the correct place names in the difficult-to-choose place name database and the location database, and detects a number of detected place names that are easy to explain by voice; 5. The voice response system according to claim 4, further comprising:

6. a deduction search step in which the location identification inference unit performs a multi-index search for the difficult-to-identify place name in the difficult-to-alternate place name database and the location database; a specific information adding step in which the inference selection unit outputs a question for additional geographic information when the number of detected place names that can be easily explained by voice cannot be narrowed down to a number that can be easily explained by voice in the inference search step, or when no place names are detected; an inference selection step in which the inference selection unit adds the additional geographical information obtained in the specific information addition step to the difficult-to-identify-choice place names, performs a multi-index search on the difficult-to-choose place name database and the location database, and narrows down the number of detected place names to those that are easy to explain by voice; an output step of transmitting the regular place names and additional geographical information or addresses narrowed down in the inference selection step to the selected part; 6. A response method using a voice response system according to claim 5, comprising:

Citation Information

Patent Citations

  • Method for automatic call recognition of arbitrarily spoken word

    JP1996320696A

  • Automated Call System

    JP2023002650A