Voice commerce system
The voice commerce system addresses the issue of non-personalized voice responses by using customer information to select appropriate automated voices, enhancing user satisfaction through personalized interactions.
Patent Information
- Application Number
- JP2024090882
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2025-12-16
AI Technical Summary
Existing automated voice response systems fail to adapt the played voice to match individual customer preferences, leading to suboptimal user experience.
A voice commerce system that connects to a customer terminal using a specified telephone number, stores customer information, and determines the voice playback based on this information, including preferences and demographics, using a determination unit to select an appropriate automated response voice.
Enhances customer satisfaction by providing a personalized voice response that aligns with individual preferences, improving the overall user experience.
Smart Images

Figure 2025183026000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a voice commerce system that is connected to a customer terminal using a predetermined telephone number and plays back audio on the customer terminal. [Background technology]
[0002] Conventionally, an automated voice response server that selectively plays back a guidance voice in response to an incoming call signal from a customer terminal has been known. The automated voice server "has a storage unit that stores a plurality of guidance voice data files, a voice playback unit that transmits a guidance voice signal requesting the transmission of a PB response signal to the PB customer terminal and stores the transmission time of the guidance voice signal as a request time, a PB signal receiving unit that stores the reception time of the PB response signal from the PB customer terminal, and a response management unit that stores the time difference between the request time and the reception time as a response time, and the voice playback unit selects and plays back a guidance voice data file from the storage unit in accordance with the response time" (see, for example, Patent Document 1).
[0003] The automated voice response server disclosed in Patent Document 1 uses the request time, which is the time at which the guidance voice signal is sent, and the time at which the PB response signal is received from the PB customer terminal, to select a guidance voice data file according to the response time, which is the time difference between the request time and the reception time, thereby selecting the optimal guidance voice for each user of the automated voice response system. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2010-034868 Summary of the Invention [Problem to be solved by the invention]
[0005] However, the automatic voice response server disclosed in Patent Document 1 changes the guidance voice only based on the response time when a customer (user) receives a call from the customer terminal, and is not capable of changing the played voice to suit the customer's preferences.
[0006] The present disclosure is intended to solve the above-mentioned problems and to provide a voice commerce system that can reproduce an automated response voice that matches the customer's preferences. [Means for solving the problem]
[0007] The voice commerce system of the present disclosure is a voice commerce system that is connected to a customer terminal used by a user using a specified telephone number and plays voice on the customer terminal, and is equipped with a memory unit that stores customer information including telephone numbers, an acquisition unit that acquires an incoming telephone number of the customer terminal, a reference unit that references customer information associated with the incoming telephone number from the memory unit, a storage unit that stores multiple automatic response voices, and a determination unit that determines the voice to be played from the multiple automatic response voices based on the referenced customer information. [Effects of the Invention]
[0008] According to the above, when a customer connects using a telephone number and orders a product or the like by voice, a response is made in a voice that matches the customer's preferences, which makes it possible to improve customer satisfaction. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a schematic configuration diagram showing a voice commerce system 100 according to a first embodiment. [Figure 2] 2 is an example of a hardware configuration of a server 10 according to the first embodiment. [Figure 3] 2 is an example of a functional block of the voice commerce system 100 according to the first embodiment. [Figure 4] 2 is an example of a configuration of a database used in the voice commerce system 100 according to the first embodiment. [Figure 5] 3 is a flowchart showing an example of information processing in the voice commerce system 100 according to the first embodiment. [Figure 6] 1 shows an example of a detailed configuration of a playback voice determination unit 14 in the voice commerce system 100 according to the first embodiment. [Figure 7] 10 is a flowchart showing an example of information processing in the voice commerce system 100 according to the second embodiment. [Figure 8] 11 is a flowchart showing an example of information processing in the voice commerce system 100 according to the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] Preferred embodiments of the voice commerce system of the present disclosure are described in detail below with reference to the drawings. Note that the embodiments described below are preferred specific examples and are subject to various technically preferable limitations, but the scope of the present disclosure is not limited to these aspects unless otherwise specified in the following description to the effect that the present disclosure is limited.
[0011] Embodiment 1 <Configuration of Voice Commerce System 100> FIG. 1 is a schematic diagram illustrating a voice commerce system 100 according to a first embodiment. The voice commerce system 100 illustrated in FIG. 1 is a system that responds to calls from a user using a predetermined telephone number on a customer terminal 30, and responds with an automated voice response. A telephone number for connecting to the voice commerce system 100 is set for each product, for example, and is displayed in various advertisements, television announcements, websites, etc., along with the product or service the user wishes to purchase. This telephone number is referred to as a connection telephone number. For example, a user checks a connection telephone number displayed in a newspaper advertisement or the like and connects to the voice commerce system 100 from the customer terminal 30 using the connection telephone number. Then, following the guidance of the automated voice response, the user responds using the customer terminal 30 and completes an order for a product or service. However, the voice commerce system 100 does not have to be used for selling products or services. For example, the voice commerce system 100 can be applied to various services, such as telephone-based donations, surveys, and advertising, which are primarily conducted via voice calls over a telephone communication network. In the first embodiment, an example will be described in which the voice commerce system 100 is configured such that a server 10 and a user's customer terminal 30 are connected via a network 90, as illustrated in FIG. 1.
[0012] The network 90 serves to connect the server 10 and the customer terminal 30. The network 90 is a communication network that connects the server 10 and the customer terminal 30 and establishes a connection path to enable data transmission and reception. At least a portion of the network 90 may be a wired network or a wireless network. The network 90 may also be an IP (Internet Protocol) network. Network 90 may include, for example, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), an ad hoc network, a metropolitan area network (MAN), a portion of the Internet, a portion of the Public Switched Telephone Network (PSTN), a cellular network, integrated service digital networks (ISDN), wireless LAN, long term evolution (LTE), code division multiple access (CDMA), Bluetooth, satellite communications, or a combination of two or more of these, enabling communication between terminals.
[0013] The server 10 has at least an automated voice response function. The server 10 may be any information processing device capable of implementing the functions described in each embodiment. The server 10 may include, for example, a server computer, a personal computer (desktop, laptop, tablet, etc.), a media computer platform (cable, satellite set-top box, digital video recorder), a handheld computer device (PDA, email client, etc.), or other types of computers. The server 10 may also be referred to as an information processing device.
[0014] The customer terminal 30 is a terminal used by a user of the voice commerce system 100. It can be any type of terminal, such as a landline, mobile phone, smartphone, tablet, or PC, as long as it can at least make a call to a connected telephone number. When ordering products or the like using the voice commerce system 100, the customer terminal 30 is sufficient as long as it can at least make a telephone call. In this case, it may be referred to as a call terminal 31 (see FIG. 3). The telephone line may be an analog line, a digital line, or an optical fiber line, and the type and carrier of the line may be any. However, the customer terminal 30 must have a telephone number, and the server 10 must be able to recognize the telephone number of the customer terminal 30 (the incoming telephone number). The customer terminal 30 may also be a smartphone-type mobile terminal or a personal computer configured to enable telephone communication. It is desirable that the customer terminal 30 is not a terminal that can be used by an unspecified number of people, such as a public telephone.
[0015] (Hardware configuration of voice commerce system 100) Fig. 2 shows an example of the hardware configuration of each of the server 10 and the customer terminal 30 according to the first embodiment. The server 10 and the customer terminal 30 have the hardware configuration shown in Fig. 2(a) or 2(b), for example. However, the customer terminal 30 may also be a terminal (call terminal 31) dedicated to calls via telephone lines such as a landline phone installed in an ordinary home.
[0016] The server 10 and the customer terminal 30, which are information processing devices, each include at least a control unit 51, a memory unit 52, and a communication unit 53. The control unit 51 is, for example, a central processing unit (CPU). The control unit 51 may also include a read-only memory (ROM) and a random access memory (RAM). The CPU is also referred to as a central processing unit, processor, microprocessor, microcomputer, or digital signal processor (DSP). In information processing devices, the CPU reads programs and data stored in the ROM and uses the RAM as a work area to control each information processing device. The CPU controls communication between information processing devices via the communication unit 53 and content displayed on an output unit 55 such as a display device or printer, based on information stored in the memory unit 52, such as the ROM and RAM. FIG. 2(b) shows a hardware configuration of an information processing device that includes an input unit 54 and an output unit 55 in addition to the control unit 51, the memory unit 52, and the communication unit 53. 2(b), the control unit 51 controls the content displayed on the output unit 55, such as a display device and a printer, based on the content input by the user via the input unit 54, and also controls various functions such as recognizing the content input by the user via the input unit 54 and storing the information in the memory unit 52. The input unit 54 of the server 10 also includes a voice input device, such as a mouthpiece of a telephone receiver, that inputs the user's voice at the customer terminal 30 connected via the network 90. In other words, the input unit 54 includes not only a device provided in the server 10, but also a device provided in a connected device, whether wireless or wired.
[0017] In addition, the control unit 51 may realize each process not only by a CPU having a control circuit, but also by a logic circuit (hardware) formed in an integrated circuit (IC (Integrated Circuit) chip, LSI (Large Scale Integration)) or a dedicated circuit.
[0018] The storage unit 52 is, for example, a non-volatile semiconductor memory such as a flash memory, an EPROM (Erasable Programmable ROM), an EEPROM (Electrically Erasable Programmable ROM), or an HDD (Hard Disk Drive), and serves as a so-called secondary storage device. In the first embodiment, the storage unit 52 may be installed within the terminal or may be a server connected via the network 90. Note that the storage unit 52 is not necessarily installed in the terminal, but may also be installed in a communicable manner on the network 90, and the network 90 may also be referred to as the storage unit 52. In other words, the control unit 51 can retrieve information from a communicable storage device, including the storage unit 52 installed in the terminal or the storage unit 52 connected to the network 90. The control unit 51 also retrieves information necessary for display on the display device from the storage unit 52 as appropriate. The storage unit 52 can also store user response information via the input unit 54.
[0019] The communication unit 53 is a part that transmits and receives various data via the network 90, and any communication protocol may be used as long as communication between the information processing devices can be performed, regardless of whether it is wired or wireless. The communication unit 53 has a function of performing communication between the information processing devices via the network 90. The communication unit 53 transmits various data to other information processing devices based on instructions from the control unit 51. The communication unit 53 also receives various data from other information processing devices and sends it to the control unit 51.
[0020] (Server 10) The control unit 51, memory unit 52, and communication unit 53 of the server 10 in the voice commerce system 100 may be configured in any form. For example, if the server 10 is configured as a network computer system consisting of multiple server computers, each function may be distributed among the multiple server computers. Furthermore, these server computers may be distributed across computer systems owned by multiple vendors or server administrators (e.g., computer systems owned or managed by other information providers, providers, hosting system providers, etc.). The server computers may be configured as a so-called cloud computer system. Furthermore, at least a portion of the components of the control unit 51, memory unit 52, and communication unit 53 may be provided in a client terminal constituting the voice commerce system 100. Note that the server 10 may further include an input unit 54, an output unit 55, and an imaging unit 56, as shown in FIG. 2(b), and each function may be distributed among multiple server computers.
[0021] (Customer terminal 30) The customer terminal 30 may be a terminal capable of making calls only via a telephone line, such as a landline phone for a general household, or may be configured to include a control unit 51 as shown in FIG. 2. The customer terminal 30 must be capable of making calls to a telephone number notified to the user by the server 10, have a telephone number, and allow the server 10 to recognize the telephone number of the customer terminal 30 when making a call. For example, the customer terminal 30 may be a landline phone installed in a general household or a mobile phone used by an individual. The customer terminal 30 is preferably capable of notifying the user of information on the server 10 by voice and of communicating with the user by voice. However, the customer terminal 30 does not necessarily have to be capable of communicating by voice. The customer terminal 30 may be, for example, a terminal that communicates only via text information, such as facsimile or SMS, or may be a terminal that allows the server 10 to provide voice guidance and then accept user operations via the keypad of the customer terminal 30.
[0022] The server 10 and the customer terminal 30, which are information processing devices, can realize the functions of the multiple functional units shown in the embodiment by reading and executing the programs stored in the storage media.
[0023] The program of the present disclosure may be provided to the server 10 and the customer terminal 30 via any transmission medium capable of transmitting the program (such as a communication network or broadcast waves). The server 10 and the customer terminal 30 execute a program downloaded via the Internet or the like, for example, to realize the functions of the multiple functional units shown in each embodiment.
[0024] Furthermore, the embodiments of the present disclosure may be realized in the form of a data signal embedded in a carrier wave in which a program is embodied by electronic transmission. At least a part of the processing in the server 10 and the customer terminal 30 may or may not be realized by cloud computing consisting of one or more computers.
[0025] At least a part of the processing in the customer terminal 30 may be configured to be performed by the server 10. In this case, at least a part of the processing of each functional unit of the control unit 51 of the customer terminal 30 may be configured to be performed by the server 10.
[0026] Unless explicitly stated otherwise, the judgment configuration in the embodiments of the present disclosure is not required, and a predetermined process may be performed when the judgment condition is met, or a predetermined process may be performed when the judgment condition is not met.
[0027] The programs disclosed herein are implemented using, for example, scripting languages such as JavaScript (registered trademark) and Python (registered trademark), object-oriented programming languages such as Java (registered trademark), markup languages such as HTML5, and functional programming languages such as Elixir.
[0028] An example of information processing according to Embodiment 1 will be described below. In the following description, the components of the voice commerce system 100 will be referred to as appropriate in FIG.
[0029] <Outline of Information Processing in Voice Commerce System 100> FIG. 3 shows an example of functional blocks of the voice commerce system 100 according to the first embodiment. The server 10 is a device that selectively plays an automated response voice in response to an incoming call signal from any of the customer terminals 30a to 30c that arrives via a network 90 such as a telephone communication network. The server 10 may have a function of transferring the incoming call signal to a call center or the like as necessary. The following describes a case where the server 10 is used for a user to order a product. Note that in practice, there may be many more customer terminals 30a to 30c connected to the server 10, and the number is not limited.
[0030] The server 10 includes an acquisition unit 11 that acquires an incoming telephone number, which is the telephone number of the customer terminal 30, a reference unit 12 that references customer information based on the incoming telephone number, a storage unit 13 that stores a plurality of automatic response voices to be played to the customer terminal 30, and a determination unit 14 that determines a voice to be played from the plurality of automatic response voices. The server 10 may also include a registration unit 15 that acquires customer information through the customer's operation of the registration terminal 32. This will be described later.
[0031] When an incoming call is received from the customer terminal 30, the acquisition unit 11 recognizes the telephone number of the customer terminal 30. The voice commerce system 100 identifies the user based on the incoming telephone number and plays an automatic response voice appropriate for the user.
[0032] The reference unit 12 references the customer database 20 based on the incoming telephone number recognized by the acquisition unit 11. The customer database 20 stores customer information associated with telephone numbers. The reference unit 12 extracts necessary information from the customer information and makes it possible for the playback audio in the determination unit 14 to be used for determination. Specifically, the reference unit 12 stores the information in a temporary storage unit 52 such as a memory.
[0033] Fig. 4 shows an example of the configuration of a database used in the voice commerce system 100 according to the first embodiment. As shown in Fig. 4(a), the customer database 20 stores a plurality of parameters such as the customer's name, telephone number, date of birth, gender, products purchased to date (purchase product history), favorite people (celebrities, actors, idols, voice actors, etc.) and characters (characters appearing in anime, manga, dramas, etc.) in association with each other. The table shown in Fig. 4(a) is an example of customer information, and is not limited to this, and may include other information related to the customer.
[0034] The determination unit 14 determines a playback voice suitable for a user from among the multiple automatic response voices stored in the storage unit 13 based on the customer information extracted by the reference unit 12. For example, if a specific user is a fan of actor X, the customer database stores actor X as a favorite person in association with a telephone number, and the determination unit 14 selects a voice by actor X from among the multiple automatic response voices stored in the storage unit 13 based on that information, and plays it on the customer terminal 30.
[0035] The determination unit 14 may not only select a voice based on information about a favorite person and character in the customer information, but may also extract, for example, a voice by actor Y who is popular in a particular generation. For example, the determination unit 14 may select a voice by actor Y from among multiple automatic response voices based on the customer's age information extracted by the reference unit 12, and play the voice on the customer terminal 30.
[0036] The determination unit 14 can select a voice to be played from among multiple automatic response voices based on at least one of age, gender, address, purchase history, and favorite person / character. For example, the storage unit 13 stores voices of popular (well-known) people and characters for each gender and age group, and the determination unit 14 first selects a voice to be played based on information about the favorite people and characters in the customer information extracted by the reference unit 12. If the voice corresponding to the favorite person / character in the customer information is not present in the storage unit 13, the determination unit 14 may refer to the gender and age of the customer information extracted by the reference unit 12 and select a voice of a person or character popular with the gender and age group to which the customer information belongs. The determination unit 14 may select a voice to be played by prioritizing each parameter, such as age, gender, address, purchase history, and favorite person / character in the extracted customer information, or may select a voice to be played by comprehensively evaluating each parameter.
[0037] The determination unit 14 may determine the audio to be reproduced using AI that uses a learning model, as will be described later.
[0038] The storage unit 13 stores multiple automated response voice data files (hereinafter simply referred to as voice data files). The voice data files are, for example, recordings of the voice of a specific person or character. The voice data files are created corresponding to the connection telephone number used by a user of the voice commerce system when connecting to the server 10. A connection telephone number is set, for example, for each product sold, and customers use this to connect to the server 10. When a user wants to purchase a product, they call the system, connect to the voice commerce system 100, and complete their order by following the automated response voice. For example, if a user sees a connection telephone number displayed with a product advertisement and intends to purchase the product, they call the system using the connection telephone number and connect to the server 10. In response to the incoming call signal, the server 10 uses the automated response voice to explain the product, enter the desired purchase quantity, and provide final confirmation of the order, thereby completing the order. In other words, the voice data files are created corresponding to the products sold in the voice commerce system 100 and are basically created corresponding to the connection telephone number.
[0039] The voice data files are stored in the voice database 21, and are associated with connection telephone numbers as shown in FIG. 4(b). Each voice data file has a target parameter 61 set to correspond to one of the parameters of the customer information. For example, the target parameter 61 has an item called "target gender" set to correspond to the gender of the customer information. The determination unit 14 selects, as the voice to be played, a voice data file whose target gender includes the gender of the customer information. However, since this alone may result in multiple voice data files being targeted, for example, a voice data file whose target age includes the age of the customer information may be selected as the voice to be played. In this way, the determination unit 14 may narrow down the voice data files to be played and ultimately determine the voice to be played.
[0040] The voice database 21 may also store multiple voice data files associated with multiple connection telephone numbers so that the server 10 can handle multiple connection telephone numbers. This allows the voice commerce system 100 to handle multiple products, services, etc. using a single server 10.
[0041] Furthermore, the voice data files stored in the storage unit 13 may be created by voice synthesis, which artificially generates voice based on the voice and speaking style of a specific person, for example. In this case, multiple voice models for each person / character are created by machine learning and stored in the storage unit 13 or the voice database 21. Text corresponding to multiple connection telephone numbers (i.e., sentences for voice guidance created corresponding to products, services, etc.) is also separately stored in the storage unit 13 or the voice database 21. The determination unit 14 may select text based on the connection telephone number, select one voice model from the multiple voice models for each person / character based on customer information, and determine the voice to be reproduced.
[0042] The voice selected by the decision unit 14 is played back on the customer terminal 30 via the communication unit 53, and the user operates or speaks in accordance with the voice to proceed with the procedure for the product or service.
[0043] The analysis unit 16 receives voice uttered by the user at the customer terminal 30 and analyzes the voice. Users of the voice commerce system 100 are not necessarily limited to those whose customer information associated with their telephone number is registered in the customer database. Therefore, when a call is received from a telephone number not registered in the customer database, the server 10 prompts the user to speak, and the analysis unit 16 analyzes the user's voice and estimates the user's gender, age, etc. The determination unit 14 determines the playback voice of a person or character that the user is likely to like based on the estimated gender, age, etc., and can guide the user through the customer information registration procedure using the playback voice.
[0044] The analysis unit 16 may also use the content of the response to the user via the automated response voice to evaluate whether the voice determined by the determination unit 14 as the playback voice is appropriate for the user. For example, the analysis unit 16 may analyze the tone of voice or the content of what the user said when responding to the user via the playback voice to evaluate whether the playback voice was appropriate. The analysis unit 16 may also ask the user a question about the playback voice and prompt the user to answer by speaking or dialing, etc., to obtain the user's direct evaluation of the playback voice and reflect this in the voice commerce system 100. The evaluation result may be reflected in the customer information in the customer database 20, or may adjust the content of the target parameter 61 associated with the voice data file in the voice database 21.
[0045] The receiving unit 17 receives input from a user's voice or operation using the customer terminal 30. The input information from the voice or operation is sent to the analyzing unit 16 and analyzed. In addition, the information from the user's voice or operation input to the receiving unit 17 may be stored in the storage unit 52.
[0046] (About the flow of information processing by the voice commerce system 100) FIG. 5 is a flowchart showing an example of information processing of the voice commerce system 100 according to the first embodiment. FIG. 5 illustrates a process flow from when a user connects to the server 10 to when the voice commerce system 100 completes procedures related to a product or service. Following the process flow shown in FIG. 5, the voice commerce system 100 recognizes and stores the order details from the user, identifies the shipping destination (recipient) of the product (service) based on the customer information stored in the customer database 20, and ships (provides) the product (service). The voice commerce system 100 identifies the product or service based on the connection telephone number used by the user to connect the customer terminal 30 to the server, and identifies the shipping destination of the product or the recipient of the service from the customer information based on the incoming telephone number. The voice commerce system 100 identifies the order details (such as the order quantity) based on the response of the user via the server 10 and the customer terminal 30, and confirms the order. Furthermore, if the voice commerce system 100 does not have customer information about the user, it may acquire the customer information about the user through a response.
[0047] The user uses their own customer terminal 30 to make a call to the notified connection telephone number (step S1). When the call is made, the server 10 recognizes the telephone number of the customer terminal 30 used by the user. The telephone number recognized by the server 10 is called the incoming telephone number. The server 10 connects to the customer terminal 30 and makes it possible to make a call with the user.
[0048] The server 10 determines whether the recognized incoming telephone number exists in the customer database 20 (step S2), and if a telephone number corresponding to the incoming telephone number exists in the customer database 20 (Yes in step S2), it obtains customer information from the customer database (step S3).
[0049] On the other hand, if the customer database 20 does not contain a telephone number corresponding to the incoming telephone number (No in step S2), the server 10 analyzes the user's call voice (step S4). For example, the server 10 acquires the user's call voice by prompting the user to speak using an automated response voice. For example, the server 10 plays an automated response voice saying, "Please tell us your name," on the customer terminal 30, to which the user responds by speaking their name. The server 10 then plays an automated response voice saying, "Please tell us your address," on the customer terminal 30, to which the user responds by speaking their address. The server 10 acquires and analyzes the user's speech. For example, the analysis is performed to convert the speech into text data to be stored by the server 10, which is known as speech recognition. When performing speech recognition, the server 10 may further ask questions about the user's age, gender, and preferences using the automated response voice, prompt the user to respond, convert the text into text data through speech recognition, and store the text data in the server 10. The analysis content may also include estimating at least one of the user's age and gender based on voice data uttered by the user.
[0050] If the customer database 20 does not contain a telephone number corresponding to the incoming telephone number (No in step S2), the server 10 may store the incoming telephone number together with acquiring the customer information.
[0051] The server 10 acquires customer information from the customer database 20 (step S3) or acquires or estimates customer information from the user (step S4), and then selects a voice data file that is deemed appropriate for the user from among multiple automatic response voices based on the customer information, and determines it as the voice to be played back (step S5). The voice to be played back is played back on the customer terminal 30 and recognized by the user (step S6).
[0052] The reproduced voice is stored as a history in the server 10 (step S7). For example, the history of the reproduced voice may be stored in the customer database 20 as part of the customer information about the user, or may be stored in the voice database 21, or may be stored in a temporary storage area and deleted later as necessary.
[0053] When the playback voice is played back on the customer terminal 30, the user responds according to the content of the voice (step S8). The user's response may be not only a vocal response, but also, for example, a push operation on the telephone terminal. If the user's response is a vocal response, the voice content is converted into content that can be stored in the server 10 using voice recognition technology, as in step S4, and is stored in the server 10.
[0054] The processes from step S6 to step S8 may be repeated until the necessary processes are completed in the voice commerce system 100. At that time, the data input by the user through speech or operation may be converted into data that can be stored in the server 10 as appropriate, and stored in the server 10. The stored data may be, for example, the order details (order quantity, etc.) of a product or service.
[0055] Next, the server 10 analyzes whether the response from the voice commerce system 100 was satisfactory to the user (step S9). In step S9, the server 10 may evaluate whether the user is satisfied with the reproduced voice based on the voice of the response between the server 10 and the user in steps S6 and S8. For example, the server 10 determines the user's emotion from the voice data of the user's utterance in step S8 and stores the result as an evaluation result (step S10).
[0056] Specifically, in step S10, for example, if the voice data uttered by the user indicates a "negative" emotion, the server 10 evaluates that the reproduced voice is not suitable for the user, and stores, for example, the fact that the reproduced voice is not to the user's liking as customer information in the customer database 20. Alternatively, the server 10 may adjust target parameters related to the reproduced voice stored in the voice database 21.
[0057] For example, if the voice data produced by the user's speech indicates a "positive" emotion, the server 10 evaluates the reproduced voice as being suitable for the user, and stores, for example, the fact that the reproduced voice is the user's preference as customer information in the customer database 20. Alternatively, the server 10 may adjust target parameters related to the reproduced voice stored in the voice database 21.
[0058] Furthermore, in step S9, the server 10 may prompt the user to respond to an evaluation of the reproduced voice. For example, the server 10 may play an automated response voice to the customer terminal 30 saying, "Please tell us your evaluation of the automated response voice," and allow the user to respond freely. The server 10 recognizes the user's response using voice recognition, and if the user responds with negative content such as "I don't like it," "I don't like it," or "I don't know," the server 10 evaluates the reproduced voice as not being suitable for the user. On the other hand, if the user responds with positive content such as "It was good," "I like it," or "I know it well," the server 10 evaluates the reproduced voice as being suitable for the user.
[0059] The server 10 may also prompt the user to respond by operating the system to evaluate the reproduced voice. For example, the server 10 may prompt the user by playing a voice message such as, "Please tell us your level of satisfaction with the automated response voice by pressing a number from 1 to 5. If you are most satisfied, press 5. If you are least satisfied, press 1." The server 10 evaluates the reproduced voice based on the user's operation.
[0060] When the necessary processing for ordering a product or service is completed in the voice commerce system 100, the server 10 ends the call and completes the processing (step S11).
[0061] The process of feeding back the evaluation result of the automatic response voice from the user's answer to the system in step S10 by the server 10 may be performed after the call ends. The step of evaluating the playback voice may be omitted.
[0062] (Variations of the Selection of Voice Playback in the Voice Commerce System 100) FIG. 6 shows an example of a detailed configuration of the playback speech determination unit 14 in the voice commerce system 100 according to the first embodiment. In the above description, the determination unit 14 determines the playback speech based on preset or stored playback speech determination criteria. However, this is not limited to this, and the determination unit 14 may determine the playback speech using a machine-learned model that estimates a playback speech suitable for a user. The playback speech estimation model 70 receives the user's customer information as input information and outputs a playback speech suitable (or estimated to be suitable) for the user as output information.
[0063] The reproduced speech estimation model 70 may be a model 71 trained using parameters such as age, gender, address, purchase history, and favorite people and characters contained in customer information. For example, the model 71 is a model trained by machine learning using parameters of customer information and favorite people and characters obtained by conducting a survey of customers or potential customers as training data. This model may be incorporated as the reproduced speech estimation model 70 in the determination unit 14. The model 71 may be generated by reinforcement learning or may be based on a neural network.
[0064] Furthermore, the reproduced speech estimation model 70 may be trained by evaluating a reproduced speech actually output by the voice commerce system 100. For example, when a certain reproduced speech is played back for a certain user, the reproduced speech and each parameter of the customer information may be used as input information, and the evaluation obtained from the user by the voice commerce system 100 may be used as a correct answer label, and the reproduced speech estimation model 70 may be trained.
[0065] Furthermore, the training data is not limited to user evaluations obtained through questionnaires or actual voice commerce system 100, but can also be data obtained from outside (such as data on age, gender, and popular celebrities).
[0066] When the above-described playback speech estimation model 70 is used in the determination unit 14, the determination unit 14 refers to the customer information of the user who made the call, uses each parameter of the customer information as input information, and estimates a person or character that the user is likely to like. Furthermore, the determination unit 14 extracts speech by the estimated person or character from a speech data file in the speech database 21.
[0067] (Collection of customer information) 3, the voice commerce system 100 may collect customer information from customers in advance. If the customer registers customer information in advance in the voice commerce system 100, when the customer connects as a user using a connection phone number, the voice commerce system 100 can refer to the customer information and provide an automated response voice appropriate for that customer.
[0068] A customer uses a customer terminal 30 to register customer information. The customer terminal 30 may be a registration terminal 32 integrated with a call terminal 31 used when using the voice commerce system 100 as a user, or may be a registration terminal 32 separate from the call terminal 31. Specifically, customer information can be registered in the voice commerce system 100 using an information terminal such as a PC, smartphone, or tablet that the customer owns.
[0069] A customer connects to the server 10, and the server 10 prompts the customer to input information by displaying a question about the customer information on a display device such as a display of the registration terminal 32 along with a text box for inputting the customer information. The server 10 may also prompt the customer to make a selection by displaying options along with the question about the customer information. In particular, when acquiring the customer's name, age, telephone number of the call terminal 31 used when using the voice commerce system 100, and address, the server 10 may be configured to display a text box on the display device of the registration terminal 32 to accept the customer's information input. Furthermore, with regard to the customer's favorite person or character, the server 10 may display options on the registration terminal 32 so that the customer can select one, or may input the option into the text box.
[0070] The server 10 processes the customer information registered by the customer in the registration unit 15 and stores it in the customer database 20.
[0071] (Effects of the First Embodiment) The voice commerce system 100 according to the first embodiment is connected to a customer terminal 30 using a predetermined telephone number, and plays back audio on the customer terminal 30. The voice commerce system 100 includes a memory unit 52 that stores customer information including telephone numbers, an acquisition unit 11 that acquires the incoming telephone number of the customer terminal 30, a reference unit 12 that references the customer information associated with the incoming telephone number from the memory unit 52, a storage unit 13 that stores a plurality of automated response voices, and a determination unit 14 that determines the voice to be played back from the plurality of automated response voices based on the extracted customer information.
[0072] With this configuration, the voice commerce system 100 can select a voice to be played back that is appropriate for the user using the voice commerce system 100 based on the incoming telephone number. The voice commerce system 100 has a plurality of pre-stored automatic response voices, and can provide the user with a voice that is appropriate for the user as the voice to be played back. The storage unit 52 may also include the customer database 20.
[0073] In the above-described voice commerce system 100, the multiple automated response voices include multiple voice data by different people or characters. The multiple voice data are associated with playback conditions set in correspondence with multiple parameters related to the user among the customer information, and the determination unit 14 selects, as the playback voice, a voice among the multiple voice data for which at least one of the multiple parameters satisfies the playback condition.
[0074] This allows the voice commerce system 100 to determine the voice to be played back using customer information stored in the customer database 20, for example.
[0075] In the above voice commerce system 100, the multiple parameters include at least one of the customer's name, address, gender, age, purchase history, and survey information.
[0076] As a result, the voice commerce system 100 can select the voice of a person or character that the user likes based on the customer information as the playback voice and provide it to the user who uses the voice commerce system 100. The survey information is information including the person or character that the customer likes. The survey information may also include the user's preferences regarding, for example, male or female, speaking style, tone of voice, etc.
[0077] The voice commerce system 100 includes a receiving unit 17 that receives input from a user's voice or operation using the customer terminal 30, and an analyzing unit 16 that analyzes the input content from the user's voice or operation. The analyzing unit 16 acquires the user's evaluation of the reproduced voice based on the user's voice or operation, and the storage unit 52 stores the evaluation as part of the customer information.
[0078] This allows the voice commerce system 100 to obtain evaluations of the voices played back by the users, thereby improving the accuracy of the selection of voices to be played back so that they are more suitable for the users.
[0079] In the above voice commerce system 100, the determination unit 14 inputs customer information about the user into a learning model trained using training data in which customer information is used as input information and one of multiple automatic response voices is used as output information, and outputs a playback voice.
[0080] This allows the voice commerce system 100 to determine the voice to be reproduced by using, for example, customer information stored in the customer database 20. In particular, even if there is no clear customer information about the voice that the user is likely to prefer, the voice that the user is likely to prefer can be estimated and provided to the user.
[0081] The voice commerce system 100 further includes a registration unit 15 into which the telephone number of the customer terminal 30 and customer information are input by the customer using the registration terminal 32. The registration unit 15 displays options related to preference information on the display device of the registration terminal 32 so that the customer can select from them, and stores the preference information selected by the customer in the storage unit 52 as customer information.
[0082] This allows the voice commerce system 100 to obtain customer information in advance before the customer connects as a user to the server 10. This can improve the user's satisfaction with the automated response voice.
[0083] By having the above-described configuration, the voice commerce system 100 can provide a voice that matches the user's preferences when the user is performing procedures such as ordering products and services, allowing the user to proceed with the procedures comfortably and encouraging the user to use the system again.
[0084] Embodiment 2 In the second embodiment, a different method of determining the voice to be reproduced will be described. The voice commerce system 100 according to the second embodiment determines the voice to be reproduced in accordance with the language used by the user.
[0085] Fig. 7 is a flowchart showing an example of information processing of the voice commerce system 100 according to embodiment 2. In embodiment 2, as in embodiment 1, a user connects to the server 10, and the voice commerce system 100 responds to the user via voice to complete a procedure related to a product or service, and Fig. 7 shows the flow of this processing.
[0086] The user uses their own customer terminal 30 to call the notified connection telephone number (step S201). When the call is made, the server 10 connects to the customer terminal 30 and makes it possible to talk to the user. Next, the server 10 prompts the user making a call using the customer terminal 30 to speak (step S202). This is done by playing back audio in English, for example, and notifying the user of content that prompts them to speak a predetermined message. Alternatively, audio in multiple languages, such as English and Japanese, may be played back in succession.
[0087] The user recognizes the voice from the server 10 prompting the user to speak and responds (step S203). The server 10 recognizes the user's voice when responding and performs voice analysis (step S204). Alternatively, the server 10 analyzes the voice input from the customer terminal 30 at a predetermined time after the voice prompting the user to speak in step 202 is played.
[0088] If the server 10 can identify the language from the analyzed voice (Yes in step S205), it determines the voice to be played that corresponds to the identified language from the multiple automatic response voices stored in the storage unit 13 (step S206) and plays it (step S207).
[0089] On the other hand, if the language cannot be identified from the analyzed voice (No in step S205), the server 10 prompts the user to speak again (return to step S202). At this time, the server 10 may cause the customer terminal 30 to play back a playback voice whose content is different from the voice that was originally played back in step S202. For example, the voice may be played in a language different from the original.
[0090] The subsequent steps S208 to S212 are similar to those in the first embodiment (steps S7 to S11 shown in FIG. 5), in which the played voice is stored as history (step S208), user evaluation is obtained (steps S209 and S210), feedback is provided to the voice commerce system 100 (step S211), and once the necessary procedures are completed, the response with the user is terminated (step S212).
[0091] The voice commerce system 100 according to the second embodiment may obtain a direct evaluation of the reproduced voice from the user, as in the case of the first embodiment. In addition, the voice commerce system 100 may obtain an evaluation from the user as to whether the language of the reproduced voice was appropriate or the ease of listening to the reproduced voice.
[0092] Furthermore, steps S208 to S211 may not be included in the information processing of the voice commerce system 100. Furthermore, the analysis and feedback processing of steps S210 and S211 may be performed after the call ends.
[0093] (Effects of the second embodiment) The voice commerce system 100 according to the second embodiment includes an analysis unit 16 that analyzes the user's speech. The multiple automated response voices include multiple voice data in multiple different languages. The analysis unit 16 determines which of the multiple different languages the user's speech at the customer terminal 30 corresponds to. The determination unit 14 determines the voice in the corresponding language from the multiple voice data as the voice to be reproduced.
[0094] With this configuration, the voice commerce system 100 can provide a user with a plurality of pre-stored automated response voices in a language that is suitable for the user, thereby enabling the voice commerce system 100 to smoothly complete procedures such as purchasing products and services with an automated response voice that is suitable for users whose native language is different.
[0095] Embodiment 3 In the second embodiment, a different method of determining the voice to be played back will be described. The voice commerce system 100 according to the third embodiment determines one of a plurality of automatic response voices as the voice to be played back by lottery every time a user connects to the server 10.
[0096] Fig. 8 is a flowchart showing an example of information processing of the voice commerce system 100 according to embodiment 3. In embodiment 3, as in embodiment 1, a user connects to the server 10, and the voice commerce system 100 responds to the user via voice to complete a procedure related to a product or service, and Fig. 8 shows the flow of this processing.
[0097] The user makes a call to the notified connection telephone number using the customer terminal 30 owned by the user (step S301). When the call is made, the server 10 connects to the customer terminal 30 and makes it possible to talk to the user.
[0098] The server 10 determines whether the recognized incoming telephone number exists in the customer database 20 (step S302), and if the telephone number corresponding to the incoming telephone number does not exist in the customer database 20 (No in step S302), the server 10 determines the voice to be played by lottery from multiple automatic response voices (step S303).
[0099] On the other hand, if the customer database 20 contains a telephone number corresponding to the called telephone number (Yes in step S302), the server 10 acquires customer information from the customer database 20 (step S304). For example, the server 10 checks, based on the called telephone number, whether the calling user has already connected using the connecting telephone number (whether the user has used the service before). If the user has already used the service before, the server 10 may acquire information about the voice data that has already been played for the user from the customer information, and may perform the voice lottery in step S303 by excluding the voice data that has already been played.
[0100] Once the voice to be played is determined, the server 10 plays the voice on the customer terminal 30 and the user recognizes it (step S305). The server 10 stores, in the voice database 21, a plurality of automatic response voices, such as voice guidance in the voices of people or characters, such as celebrities, and the voice to be played is determined by lottery from among the plurality of automatic response voices. The plurality of automatic response voices are made up of the voices of, for example, members of an idol group, and the voice of one of the members is determined by lottery to be the voice to be played.
[0101] The multiple automated response voices may be set to have different winning probabilities. For example, the multiple automated response voices associated with a certain connected telephone number stored in the voice database 21 may be composed of voice data files A through E, with the probability of voice data files A through D being played back being 24% and the probability of voice data file E being played back being 4%. In this case, voice data file E may be set to have special content as a specific voice (premium). For example, the specific voice may be a voice by a rare person or character, or may be the same person as the other voice data files A through E but with special content. In addition, the specific voice may be set to receive a special benefit for a product or service that was desired to be purchased using the connected telephone number in conjunction with the playback of the specific voice.
[0102] The subsequent steps S306 to S310 are similar to those in the first embodiment (similar to steps S7 to S11 shown in FIG. 5), in which the played voice is stored as history (step S306), user evaluation is obtained (steps S307 and S308), and feedback is provided to the voice commerce system 100 (step S309), and once the necessary procedures are completed, the response with the user is terminated (step S310).
[0103] If the history of the reproduced voices is stored as customer information in step S306, the stored history may be acquired in step S304 and used in the voice lottery.
[0104] The voice commerce system 100 according to the third embodiment may obtain a direct evaluation of the reproduced voice from the user in the same manner as in the first embodiment, and may not include steps S306 to S309 as information processing.
[0105] (Effects of the Third Embodiment) According to the voice commerce system 100 of the third embodiment, the determination unit 14 determines the voice to be reproduced by lottery from at least some of the multiple automatic response voices.
[0106] Furthermore, according to the voice commerce system 100 of embodiment 3, the memory unit 52 stores the voice data that has already been played back from among the multiple automated response voices in association with the incoming telephone number of the customer terminal 30, and the determination unit 14 determines the voice to be played back by lottery from among the multiple automated response voices other than the voice data that has already been played back.
[0107] With this configuration, the voice commerce system 100 can have multiple automated response voices prepared in advance and can provide a playback voice to a user by lottery from among them. This allows the voice commerce system 100 to provide a different automated response voice each time a user connects to process a product, service, or the like, thereby increasing the frequency of connections for the same user.
[0108] In the voice commerce system 100 according to the third embodiment, the plurality of automated response voices include at least one specific voice data. The specific voice data is less likely to be played than voice data other than the specific voice data among the plurality of automated response voices.
[0109] With this configuration, the voice commerce system 100 can create a premium feel for specific voice data, thereby increasing the frequency of user connections.
[0110] Although the present disclosure has been described above based on the embodiments, the present disclosure is not limited to the configurations of the above-described embodiments. In the above-described embodiments, the voice commerce system 100 is realized using a system such as a client-server system of a network computer system. However, the same functions as the voice commerce system 100 can also be realized on various computers, such as personal computers, that do not constitute a client-server system, or various communication terminals and mobile information terminals, such as mobile terminals and tablets. Furthermore, at least a portion of the functions of the voice commerce system 100 can also be realized by installing a computer program on various computers, communication terminals, and mobile information terminals. In other words, the present disclosure also includes programs for causing various computers to function as at least a portion of the voice commerce system 100. The above-described embodiments and variations may be implemented in appropriate combinations. It should be noted that the gist (technical scope) of the present disclosure also includes various modifications, applications, and uses that may be made by those skilled in the art as needed.
[0111] The voice commerce system 100 described above may also include combinations of the features shown in Supplementary Notes 1 to 10 below. These combinations are described below.
[0112] [Appendix 1] A voice commerce system that is connected to a customer terminal used by a user using a predetermined telephone number and plays back voice on the customer terminal, a storage unit for storing customer information including telephone numbers; an acquisition unit that acquires an incoming telephone number of the customer terminal; a reference unit that references customer information associated with the incoming telephone number from the storage unit; a storage unit for storing a plurality of automatic response voices; a determination unit that determines a playback voice from among the plurality of automatic response voices based on the extracted customer information, Voice commerce system. [Appendix 2] 3. The voice commerce system of claim 2, The plurality of automated response voices include: Contains multiple voice data by different people or characters, The plurality of audio data playback conditions set in correspondence with a plurality of parameters relating to the user among the customer information; The determination unit Among the plurality of pieces of audio data, audio in which at least one of the plurality of parameters satisfies the playback condition is determined as the playback audio. Voice commerce system. [Appendix 3] 3. The voice commerce system of claim 2, The plurality of parameters are: The information includes at least one of the user's name, address, gender, age, purchase history, and survey information. Voice commerce system. [Appendix 4] 4. The voice commerce system according to claim 2 or 3, a receiving unit that receives input by speech or operation of the user using the customer terminal; an analysis unit that analyzes input content by the user's voice or operation, The analysis unit acquiring an evaluation of the user regarding the reproduced voice based on an utterance or operation of the user; The storage unit storing the rating as part of the customer information; Voice commerce system. [Appendix 5] A voice commerce system according to any one of Supplementary Notes 1 to 3, The determination unit inputting customer information about the user into a learning model trained using training data in which the customer information is used as input information and any one of the plurality of automatic response voices is used as output information, and outputting the reproduced voice; Voice commerce system. [Appendix 6] 2. The voice commerce system of claim 1, an analysis unit that analyzes the user's speech, The plurality of automated response voices include: Contains multiple audio data in multiple different languages, The analysis unit determining which of the different languages the user's utterance corresponds to; The determination unit determining a voice in a corresponding language from the plurality of voice data as the voice to be reproduced; Voice commerce system. [Appendix 7] 2. The voice commerce system of claim 1, The determination unit determining the playback voice by lottery from at least some of the plurality of automatic response voices; Voice commerce system. [Appendix 8] 8. The voice commerce system of claim 7, The storage unit storing the voice data that has already been reproduced from the plurality of automatic response voices in association with the incoming telephone number of the customer terminal; The determination unit determining the voice to be reproduced by lottery from among the plurality of automatic response voices other than the voice data that has already been reproduced; Voice commerce system. [Appendix 9] 9. The voice commerce system according to claim 7 or 8, The plurality of automated response voices include: Contains at least one specific audio data; The specific audio data is the specific voice data is less likely to be played back than the voice data other than the specific voice data among the plurality of automatic response voices; Voice commerce system. [Appendix 10] A voice commerce system according to any one of claims 1 to 9, The system further includes a registration unit into which the telephone number of the customer terminal and the customer information input by the customer using the registration terminal are input, The registration unit displaying options relating to preference information on a display device of the registration terminal so that the customer can select from the options; storing the preference information selected by the customer in the storage unit as the customer information; Voice commerce system. [Explanation of symbols]
[0113] 10: Server 11: Acquisition part 12:Reference part 13: Storage area 14: Decision section 15: Registration Department 16:Analysis Department 17: Receiving unit 20: Customer database 21: Voice database 30: Customer terminal 30a: Customer terminal 30b: Customer terminal 30c: Customer terminal 31: Call terminal 32: Registered terminal 51: Control unit 52: Storage section 53: Communications Department 54: Input section 55: Output section 56: Imaging unit 61: Target parameter 70: Reproduced speech estimation model 71: Model 90: Network 100: Voice commerce system
Claims
1. A voice commerce system that is connected to a customer terminal used by a user using a predetermined telephone number and plays back voice on the customer terminal, a storage unit for storing customer information including telephone numbers; an acquisition unit that acquires an incoming telephone number of the customer terminal; a reference unit that references customer information associated with the incoming telephone number from the storage unit; a storage unit for storing a plurality of automatic response voices; a determination unit that determines a playback voice from the plurality of automatic response voices based on the referenced customer information, Voice commerce system.
2. 2. The voice commerce system of claim 1, The plurality of automated response voices include: Contains multiple voice data by different people or characters, The plurality of audio data a playback condition set in correspondence with a plurality of parameters relating to the user among the customer information is associated with the playback condition; The determination unit Among the plurality of pieces of audio data, audio in which at least one of the plurality of parameters satisfies the playback condition is determined as the playback audio. Voice commerce system.
3. 3. The voice commerce system of claim 2, The plurality of parameters are: The information includes at least one of the user's name, address, gender, age, purchase history, and survey information. Voice commerce system.
4. 4. The voice commerce system according to claim 2 or 3, a receiving unit that receives input by speech or operation of the user using the customer terminal; an analysis unit that analyzes input content by the user's voice or operation, The analysis unit acquiring an evaluation of the user regarding the reproduced voice based on an utterance or operation of the user; The storage unit storing the rating as part of customer information; Voice commerce system.
5. The voice commerce system according to any one of claims 1 to 3, The determination unit inputting customer information about the user into a learning model trained using training data in which customer information is used as input information and any one of the plurality of automatic response voices is used as output information, and outputting any one of the plurality of automatic response voices as the playback voice; Voice commerce system.
6. 2. The voice commerce system of claim 1, an analysis unit that analyzes the user's speech, The plurality of automated response voices include: Contains multiple audio data in multiple different languages, The analysis unit determining which of the different languages the user's utterance corresponds to; The determination unit determining a voice in a corresponding language from the plurality of voice data as the voice to be reproduced; Voice commerce system.
7. 2. The voice commerce system of claim 1, The determination unit determining the playback voice by lottery from at least some of the plurality of automatic response voices; Voice commerce system.
8. 8. The voice commerce system of claim 7, The storage unit storing voice data that has already been reproduced from among the plurality of automatic response voices in association with the incoming telephone number of the customer terminal; The determination unit determining the voice to be reproduced by lottery from among the plurality of automatic response voices other than the voice data that has already been reproduced; Voice commerce system.
9. 9. The voice commerce system according to claim 7 or 8, The plurality of automated response voices include: Contains at least one specific audio data; The specific audio data is the specific voice data is less likely to be played back than the voice data other than the specific voice data among the plurality of automatic response voices; Voice commerce system.
10. The voice commerce system according to any one of claims 1 to 3 and 6 to 8, The system further includes a registration unit into which the telephone number of the customer terminal and customer information input by the customer using the registration terminal are input, The registration unit displaying options relating to preference information on a display device of the registration terminal so that the customer can select from the options; storing the preference information selected by the customer in the storage unit as customer information; Voice commerce system.
Citation Information
Patent Citations
Automatic voice response server
JP2010034868A