Program, information processing device, and information processing method
The information processing system enhances voice quality in calls between users and customers by analyzing attributes and performing voice conversion, addressing the issue of suboptimal communication quality in existing technologies.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-19
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies fail to provide suitable voice quality in calls between multiple users, such as users and customers, leading to suboptimal communication.
An information processing system that includes a server, user terminals, a CRM system, and customer terminals, utilizing machine learning models to analyze user and call attributes, and perform voice conversion to enhance voice quality in calls.
Enables customers to communicate with users using more suitable voice quality, improving the effectiveness of calls between multiple parties.
Smart Images

Figure 0007836555000001 
Figure 0007836555000002 
Figure 0007836555000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a program, an information processing apparatus, and an information processing method.
Background Art
[0002] Conventionally, in acoustic devices such as earphones and headphones that are mainly worn on the head by users, acoustic devices capable of suppressing environmental sound (so-called noise) from the external environment and enhancing the sound insulation effect are known. Patent Document 1 discloses a technique that enables listening to sound in a more suitable manner without complicated operations even in a situation where the state and situation of the user change sequentially. Patent Document 2 discloses a technique for automatically determining an appropriate noise cancellation filter. Patent Document 3 discloses a voice recognition device that can cope with environments such as sudden environmental changes and alternating appearances of multiple environments.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in a call made between a plurality of users such as a user and a customer, a call more suitable for the user has not been realized.
[0005] Therefore, this disclosure has been made to solve the above-mentioned problems, and its purpose is to provide a technology that enables customers to communicate with users using more suitable voice quality in calls between multiple users. [Means for solving the problem]
[0006] A program comprising a processor and a memory unit, which causes a computer to perform a telephone call between a first user and a second user, wherein the program causes the processor to perform a voice acquisition step of acquiring telephone audio from the first user, a conversion step of converting the telephone audio acquired in the voice acquisition step, an output step of outputting the converted telephone audio to the second user, and an attribute acquisition step of acquiring telephone attributes related to the call, the conversion step including a step of converting the telephone audio acquired in the voice acquisition step based on the telephone attributes acquired in the attribute acquisition step. [Effects of the Invention]
[0007] According to this disclosure, in calls between multiple users, customers can communicate with users using more suitable voice. [Brief explanation of the drawing]
[0008] [Figure 1] This is a diagram showing the overall configuration of information processing system 1. [Figure 2] This block diagram shows the functional configuration of Server 10. [Figure 3] This is a block diagram showing the functional configuration of user terminal 20. [Figure 4] This is a diagram showing the functional configuration of CRM system 30. [Figure 5] This is a block diagram showing the functional configuration of customer terminal 50. [Figure 6] This diagram shows the data structure of user table 1012. [Figure 7] This diagram shows the data structure of organization table 1013. [Figure 8] It is a diagram showing the data structure of the call table 1014. [Figure 9] It is a diagram showing the data structure of the voice processing table 1015. [Figure 10] It is a diagram showing the data structure of the learning dataset 1031. [Figure 11] It is a diagram showing the data structure of the customer table 3012. [Figure 12] It is a flowchart showing the operation of the voice conversion process (first embodiment). [Figure 13] It is a flowchart showing the operation of the voice conversion process (second embodiment). [Figure 14] It is a flowchart showing the operation of the voice conversion process (third embodiment). [Figure 15] It is a diagram showing an example of the display screen of the user terminal 20 in the voice conversion process (third embodiment). [Figure 16] It is a block diagram showing the basic hardware configuration of the computer 90.
Mode for Carrying Out the Invention
[0009] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In all the drawings for describing the embodiments, common components are denoted by the same reference numerals, and repeated descriptions are omitted. Note that the following embodiments do not unduly limit the content of the present disclosure described in the claims. Also, not all of the components shown in the embodiments are essential components of the present disclosure. Further, each figure is a schematic diagram and is not necessarily drawn precisely.
[0010] <Overview of the Information Processing System 1> The information processing system 1 in the present disclosure is an information processing system that provides a call service according to the present disclosure. The information processing system 1 is an information processing system that provides a service related to a call made between a user and a customer, and stores and manages data related to the call.
[0011] <Basic Configuration of Information Processing System 1> The information processing system 1 is configured to include a server 10, a plurality of user terminals 20A, 20B, 20C, a CRM system 30, a voice server (PBX) 40 connected via a network N, and customer terminals 50A, 50B, 50C connected to the voice server (PBX) 40 via a telephone network T.
[0012] FIG. 1 is a diagram showing the overall configuration of the information processing system 1. FIG. 2 is a block diagram showing the functional configuration of the server 10. FIG. 3 is a block diagram showing the functional configuration of the user terminal 20. FIG. 4 is a block diagram showing the functional configuration of the CRM system 30. FIG. 5 is a block diagram showing the functional configuration of the customer terminal 50.
[0013] The server 10 is an information processing device that provides a service for storing and managing data (call data) related to calls made between users and customers.
[0014] The user terminal 20 is an information processing device operated by a user who uses the service. The user terminal 20 may be, for example, a desktop PC (Personal Computer), a laptop PC, or a mobile terminal such as a smartphone or a tablet. It may also be a wearable terminal such as an HMD (Head Mount Display) or a wristwatch-type terminal.
[0015] The CRM system 30 is an information processing device managed and operated by a business operator (CRM operator) that provides a CRM (Customer Relationship Management) service. Examples of CRM services include SalesForce, HubSpot, Zoho CRM, and kintone.
[0016] The voice server (PBX) 40 is an information processing device that functions as a switchboard, enabling calls between the user terminal 20 and the customer terminal 50 by connecting the network N and the telephone network T to each other.
[0017] The customer terminal 50 is an information processing device that the customer operates when making a call with the user. The customer terminal 50 may be, for example, a mobile device such as a smartphone or tablet, or a stationary PC (Personal Computer) or laptop PC. It may also be a wearable device such as an HMD (Head Mount Display) or a smartwatch.
[0018] Each information processing device consists of a computer equipped with an arithmetic unit and a memory device. The basic hardware configuration of the computer and the basic functional configuration of the computer realized by said hardware configuration will be described later. For each of the server 10, user terminal 20, CRM system 30, voice server (PBX) 40, and customer terminal 50, explanations that overlap with the basic hardware configuration and basic functional configuration of the computer described later will be omitted.
[0019] The configuration and operation of each device are described below.
[0020] <Server 10 Functional Configuration> Figure 2 shows the functional configuration realized by the hardware configuration of server 10. Server 10 includes a storage unit 101 and a control unit 104.
[0021] <Configuration of the storage unit of Server 10> The storage unit 101 of the server 10 includes an application program 1011, a user table 1012, an organization table 1013, a call table 1014, a speech processing table 1015, an evaluation model 1021, a generation model 1022, a speech processing model 1023, and a training dataset 1031. Figure 6 shows the data structure of user table 1012. Figure 7 shows the data structure of organization table 1013. Figure 8 shows the data structure of the call table 1014. Figure 9 shows the data structure of the speech processing table 1015. Figure 10 shows the data structure of the training dataset 1031.
[0022] User Table 1012 is a table that stores and manages information about member users (hereinafter referred to as "users") who use the service. When a user registers to use the service, their information is stored in a new record in User Table 1012. This allows the user to use the service related to this disclosure. User Table 1012 has User ID as its primary key and contains columns for User ID, CRM ID, Organization ID, User Name, and User Attributes.
[0023] The User ID is an item that stores user identification information to identify a user. The CRMID is an item in the CRM system 30 that stores identification information for identifying a user. Users can access CRM services by logging into the CRM system 30 using their CRMID. In other words, the user ID on the server 10 and the CRMID on the CRM system 30 are linked. The Organization ID is an item that stores the Organization ID of the organization to which the user belongs. The username field is where the user's name is stored. User attributes are fields that store information about the user's attributes, such as age, gender, place of origin, dialect, and occupation (sales, customer support, etc.).
[0024] Organization Table 1013 is a table that defines information about the organizations to which a user belongs. Organizations include any organization or group, such as companies, corporations, corporate groups, clubs, and various other organizations. Organizations may also be defined for more detailed subgroups, such as departments within a company (sales department, general affairs department, customer support department). Organization Table 1013 is a table with Organization ID as the primary key and has columns for Organization ID, Organization Name, and Organization Attributes.
[0025] The Organization ID is an item that stores organizational identification information used to identify an organization. The "Organization Name" field is used to remember the name of an organization. This field includes any organization or group name, such as company names, legal entity names, corporate group names, club names, or various other group names. Organizational attributes are items that store information about the attributes of an organization, such as the type of organization (company, corporate group, other organization, etc.) and industry (real estate, finance, etc.).
[0026] The call table 1014 is a table that stores and manages call data related to calls made between a user and a customer. The call table 1014 has a call ID as its primary key and has columns for call ID, user ID, customer ID, call category, call type (incoming / outgoing), and voice data.
[0027] The call ID is an item that stores call data identification information used to identify call data. The User ID is an item used to store the user's User ID (user identification information) during phone calls between the user and the customer. The Customer ID is an item used to store the customer's ID (customer identification information) during a call between the user and the customer. The call category is an item that stores the type (category) of call made between the user and the customer. Call data is classified by call category. Depending on the purpose of the call between the user and the customer, the call category may store values such as telephone operator, telemarketing, customer support, or technical support. The "Call Type" field stores information to distinguish whether a call between the user and a customer was initiated by the user (outbound) or received by the user (inbound). The audio data field stores audio data from conversations between the user and the customer. Various audio data formats, such as mp4 and wav, can be used. It may also store reference information (paths) to audio data files located elsewhere. The audio data may be in a format in which identifiers are set that allow the user's voice and the customer's voice to be independently identified. In this case, the control unit 104 of the server 10 can perform independent analysis processing on the user's voice and the customer's voice. In this disclosure, video data containing audio information may be used instead of audio data. Furthermore, in this disclosure, "audio data" is a concept that also includes audio data contained within video data.
[0028] The audio processing table 1015 is a table that stores information related to audio processing, such as effects and filters applied to audio data (audio processing information). The voice processing table 1015 is a table whose primary key is the voice processing ID, and which has columns for voice processing ID and voice processing content.
[0029] The voice processing ID is an item that stores voice processing identification information used to identify the content of voice processing. The audio processing details section is an item that stores the audio processing details to be applied to the audio data. It may also store references to functions, methods, programs, etc. that perform audio processing located elsewhere. The audio processing includes voice conversion processing, such as changing the voice of the audio data. Voice conversion includes converting audio data into a male voice, a female voice, the voice of a specific person, or a specific character. Voice conversion includes converting audio data into voices expressing specific emotions (joy, sadness, anger, surprise, fear, disgust). Voice conversion includes conversion processing that changes the shape of the intensity of each frequency contained in the audio data (the spectral structure of the voice, the frequency distribution). Voice conversion includes conversion processing that changes the fundamental frequency, the strength of the intonation, the speaking speed, and changes the intonation (making it louder, making it softer). Voice conversion includes processing to remove fillers contained in the voice (for example, um, uh, etc.). The speech conversion process may include a process for converting the voice components of a person contained in the audio data, but may not include a process for converting other speech components such as background noise, static, and background sounds.
[0030] Evaluation model 1021 is a learning model that takes user attributes, customer attributes, call attributes such as call category and call type, and voice data as input data to output (infer) evaluation metric values. The evaluation model 1021 does not need to be a single learning model; it may be implemented by switching between multiple independent learning models depending on the type of evaluation metric to be output (evaluation type). For example, the evaluation model 1021 may include multiple different independent learning models depending on the type of evaluation metric to be output (first metric, second metric, etc.). One example of an evaluation model is the SIIB (Speech Intelligibility in Bits) model (hereinafter referred to as the first model). By applying speech data to the first model, a SIIB score (hereinafter referred to as the first index) can be obtained. The SIIB score is described in arXiv:2104.08499, etc., and is a quantitative evaluation index related to the ease of listening (ease of perception) of speech data, using speech data as input data. The evaluation model may be prepared according to evaluation metrics such as HASPI (Hearing-aid speech index), ESTOI (Extended short-time subjective intelligibility), PESQ (Perceptual evaluation of speech quality), and ViSQOL (Virtual speech quality objective listener). Furthermore, any machine learning, deep learning, or artificial intelligence model can be constructed using call attributes and audio data as input data, and particularly using evaluation results from listener surveys as training data. For example, the evaluation results from surveys may include items such as reliability, credibility, pleasantness, comfort, preference, stress level, intimidation level, and interest. In other words, it can also be a learning model that uses call attributes and audio data as input data and outputs evaluation indicators such as reliability, credibility, pleasantness, comfort, preference, stress level, intimidation level, and interest from the listener's perspective.
[0031] Generative model 1022 is a learning model that takes call attributes and voice data as input data and outputs (infers) converted voice data. The training process for generative model 1022 will be described later. The generative model 1022 does not need to be a single learning model; it may be implemented by switching between multiple independent learning models for each call attribute and evaluation type. Specifically, the generative model 1022 may include multiple independent learning models for each input call attribute and output evaluation type. For example, the generative model 1022 may be implemented by selectively switching between multiple generative models that output suitable converted speech data depending on each combination of call attributes. For example, the generative model 1022 may be implemented by selectively switching between a first generative model, a second generative model, and so on, which output suitable converted speech data for each of multiple evaluation metrics such as the first metric and the second metric.
[0032] The speech processing model 1023 is a learning model that takes call attributes and voice data as input data and outputs (infers) a speech processing ID. The training process for the speech processing model 1023 will be described later. The speech processing model 1023 does not need to be a single learning model; it may be implemented by switching between multiple independent learning models for each call attribute and evaluation type. Specifically, the speech processing model 1023 may include multiple independent learning models for each input call attribute and output evaluation type. For example, the speech processing model 1023 may be implemented by selectively switching between multiple speech processing models that output suitable converted speech data according to each combination of call attributes. For example, the speech processing model 1023 may be implemented by selectively switching between a first speech processing model, a second speech processing model, etc., which output suitable converted speech data for each of multiple evaluation metrics such as the first metric and the second metric.
[0033] The evaluation model 1021, the generative model 1022, and the speech processing model 1023 are, for example, types of machine learning, artificial intelligence, and deep learning models. As an example of evaluation model 1021, generative model 1022, and speech processing model 1023, we will explain a deep learning model using a deep neural network in deep learning. The deep learning model can be any learning model that takes arbitrary time-series data as input, such as RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), and GRU (Gated Recurrent Unit). The learning model can include any deep learning model that includes, for example, Attention and Transformer. The evaluation model 1021, the generative model 1022, and the speech processing model 1023 do not necessarily have to be deep learning models; any machine learning or artificial intelligence model is acceptable.
[0034] The training dataset 1031 is a table that stores the datasets used in the training process of the generative model 1022. The training dataset 1031 is a dataset used in machine learning, deep learning, and other training processes, in which user attribute information, customer attribute information, and call attribute information in call data are stored in association with each other. The training dataset 1031 may also be created by combining a call table 1014, which stores past call data between users and customers, a user table 1012, an organization table 1013, a customer table 3012, and so on. The training dataset 1031 is a table with columns for user attribute information, customer attribute information, call attribute information, voice data, first metric, second metric, and third metric.
[0035] User attribute information is an item in call data that stores information about the user's user attributes, the name of the organization to which the user belongs, or the organization's attributes. User attribute information may also include information about the user's emotions (joy, sadness, anger, surprise, fear, disgust) in the call data (emotional information). Customer attribute information is an item in call data that stores information about the customer's attributes, the name of the organization to which the customer belongs, or organizational attributes. Customer attribute information may also include customer sentiment information in the call data. Call attribute information is an item in the call data that stores information such as the call category and the type of caller / recipient. Call attribute information may also include user and customer sentiment information in the call data. The voice data field stores the audio data of calls between the user and the customer. Since it is similar to the voice data in call table 1014, a detailed explanation is omitted.
[0036] <Configuration of the control unit of server 10> The control unit 104 of the server 10 includes a user registration control unit 1041, a voice conversion unit 1042, and a learning unit 1051. The control unit 104 realizes each functional unit by executing the application program 1011 stored in the storage unit 101.
[0037] The user registration control unit 1041 processes information of users who wish to use the services related to this disclosure and stores it in the user table 1012. Information stored in the user table 1012 is obtained when a user opens a web page operated by the service provider from any information processing terminal, enters the information into a designated input form, and sends it to the server 10. The user registration control unit 1041 of the server 10 stores the received information in a new record in the user table 1012, and user registration is completed. As a result, users stored in the user table 1012 can use the service. Prior to the registration of user information in the user table 1012 by the user registration control unit 1041, the service provider may perform a prescribed review and restrict whether or not the user can use the service. The user ID can be any string or number that can identify the user, and may be any string or number desired by the user, or the user registration control unit 1041 of the server 10 may automatically set any string or number.
[0038] The voice conversion unit 1042 performs voice conversion processing (first embodiment), voice conversion processing (second embodiment), voice conversion processing (third embodiment), and voice conversion processing (fourth embodiment). Details will be described later. The learning unit 1051 executes the learning process. Details will be described later.
[0039] <Functional configuration of user terminal 20> Figure 3 shows the functional configuration realized by the hardware configuration of the user terminal 20. The user terminal 20 comprises a storage unit 201, a control unit 204, an input device 206 connected to the user terminal 20, and an output device 208. The input device 206 includes a camera 2061, a microphone 2062, a position information sensor 2063, a motion sensor 2064, a keyboard 2065, and a mouse 2066. The output device 208 includes a display 2081 and a speaker 2082.
[0040] <Configuration of the storage unit of user terminal 20> The storage unit 201 of the user terminal 20 stores a user ID 2011, an application program 2012, and a CRM ID 2013 for identifying the user using the user terminal 20. The User ID is the user's account ID for Server 10. The user sends User ID 2011 from User Terminal 20 to Server 10. Server 10 identifies the user based on User ID 2011 and provides the services related to this disclosure to the user. The User ID includes information such as a session ID that is temporarily assigned by Server 10 to identify the user using User Terminal 20. CRMID is the user's account ID for the CRM system 30. The user sends CRMID2013 from the user terminal 20 to the CRM system 30. The CRM system 30 identifies the user based on CRMID2013 and provides CRM services to the user. CRMID2013 also includes information such as a session ID that is temporarily assigned by the CRM system 30 to identify the user using the user terminal 20. The application program 2012 may be pre-stored in the memory unit 201, or it may be configured to be downloaded from a web server operated by the service provider via a communication interface. The application program 2012 includes an interpreted programming language such as JavaScript (registered trademark) that is executed on a web browser application stored in the user terminal 20.
[0041] <Configuration of the Control Unit of User Terminal 20> The control unit 204 of the user terminal 20 includes an input control unit 2041 and an output control unit 2042. By executing the application program 2012 stored in the storage unit 201, the control unit 204 realizes the functional units of the input control unit 2041 and the output control unit 2042. The input control unit 2041 of the user terminal 20 acquires information output from input devices such as a camera 2061, a microphone 2062, a position information sensor 2063, a motion sensor 2064, a keyboard 2065, and a mouse 2066 connected to the user terminal 20 and executes various processes. The input control unit 2041 of the user terminal 20 executes a process of transmitting the information acquired from the input device 206 to the server 10 together with the user ID 2011. Similarly, the input control unit 2041 of the user terminal 20 executes a process of transmitting the information acquired from the input device 206 to the CRM system 30 together with the CRM ID 2013. The output control unit 2042 of the user terminal 20 receives operations by the user on the input device 206 and information from the server 10 and the CRM system 30, and executes control processes for the display content of the display 2081 and the audio output content of the speaker 2082 connected to the user terminal 20.
[0042] <Functional Configuration of CRM System 30> The functional configuration realized by the hardware configuration of the CRM system 30 is shown in FIG. 4. The CRM system 30 includes a storage unit 301 and a control unit 304. The user has separately concluded a contract with a CRM operator and can receive CRM services by accessing (logging in) a website operated by the CRM operator via a web browser or the like using the CRM ID 2013 set for each user.
[0043] <Configuration of the Storage Unit of CRM System 30> The storage unit 301 of the CRM system 30 includes a customer table 3012. FIG. 11 is a diagram showing the data structure of the customer table 3012.
[0044] Customer table 3012 is a table for storing and managing customer information. Customer table 3012 has Customer ID as its primary key and contains columns for Customer ID, User ID, Name, Telephone Number, Customer Attributes, Customer Organization Name, and Customer Organization Attributes.
[0045] The customer ID is an item that stores customer identification information to identify a customer. The User ID is an item that stores the User ID (user identification information) of a user associated with a customer. Users can view a list of customers associated with their User ID and make calls to those customers. In this disclosure, customers are linked to users, but they may also be linked to organizations (organization IDs in organization table 1013). In that case, users belonging to an organization can view a list of customers linked to their own organization ID and can send messages to those customers. The "Name" field is used to store the customer's name. The phone number field is used to store the customer's phone number. The user can access the website provided by the CRM system, select the customer they wish to call, and perform a predetermined operation such as "Call" to initiate a call to the customer's phone number from the user terminal 20. Customer attributes are items that store information about customer attributes such as age, gender, place of origin, dialect, and occupation (sales, customer support, etc.). The customer organization name is an item that stores the name of the organization to which the customer belongs. The organization name can include any organization name or group name, such as company name, corporate name, corporate group name, club name, or various group name. Customer organizational attributes are items that store information about the attributes of the customer's organization, such as the type of organization (company, corporate group, other organization, etc.) and industry (real estate, finance, etc.). Customer attributes, customer organization name, and customer organization attributes may be stored by the user through input, or they may be entered by the customer when they access a designated website.
[0046] <Configuration of the Control Unit of the CRM System 30> The control unit 304 of the CRM system 30 includes a user registration control unit 3041. The control unit 304 realizes each functional unit by executing the application program 3011 stored in the storage unit 301.
[0047] The CRM system 30 provides functions called API (Application Programming Interface), SDK (Software Development Kit), and code snippets (hereinafter referred to as "beacons"). The user can perform linkage settings such as account information for the server 10 and the CRM system 30 according to the present disclosure in advance, so that the control unit 104 of the server 10 and the control unit 304 of the CRM system 30 can communicate with each other and realize any information processing.
[0048] <Overview of the Voice Server (PBX) 40> When there is an outgoing call from the user to the customer, the voice server (PBX) 40 makes an outgoing call (rings) to the customer terminal 50. When there is an incoming call from the customer to the user, the voice server (PBX) 40 sends a message indicating that (hereinafter referred to as "incoming call notification message") to the user terminal 20. In addition, the voice server (PBX) 40 can send an incoming call notification message to the beacons, SDK, API, etc. provided by the server 10.
[0049] <Functional Configuration of the Customer Terminal 50> The functional configuration realized by the hardware configuration of the customer terminal 50 is shown in FIG. 5. The customer terminal 50 includes a storage unit 501, a control unit 504, a touch panel 506, a touch-sensitive device 5061, a display 5062, a microphone 5081, a speaker 5082, a position information sensor 5083, a camera 5084, and a motion sensor 5085.
[0050] <Configuration of the Storage Unit of the Customer Terminal 50> The memory unit 501 of the customer terminal 50 stores the customer's telephone number 5011 and application program 5012. The application program 5012 may be pre-stored in the memory unit 501, or it may be configured to be downloaded from a web server operated by the service provider via a communication interface. The application program 5012 includes an interpreted programming language such as JavaScript (registered trademark) that is executed on a web browser application stored in the customer terminal 50.
[0051] <Configuration of the control unit of customer terminal 50> The control unit 504 of the customer terminal 50 comprises an input control unit 5041 and an output control unit 5042. The control unit 504 realizes the functional units of the input control unit 5041 and the output control unit 5042 by executing an application program 5012 stored in the storage unit 501. The input control unit 5041 of the customer terminal 50 acquires information from the user's operations on the touch-sensitive device 5061 of the touch panel 506, voice input to the microphone 5081, and information output from input devices such as the position information sensor 5083, camera 5084, and motion sensor 5085, and performs various processes. The output control unit 5042 of the customer terminal 50 receives information from the user's operation on the input device and from the server 10, and performs control processing such as the display content of the display 5062 and the audio output content of the speaker 5082.
[0052] <Operation of Information Processing System 1> The following describes each process of Information Processing System 1. Figure 12 is a flowchart showing the operation of the speech conversion process (first embodiment). Figure 13 is a flowchart showing the operation of the speech conversion process (second embodiment). Figure 14 is a flowchart showing the operation of the speech conversion process (third embodiment). Figure 15 shows an example of the display screen of the user terminal 20 in the voice conversion process (third embodiment).
[0053] <Term definition> In explaining each process of Information Processing System 1, the following terms are defined: Call data is data relating to calls made between a user and a customer, and includes data stored in each item of the call table 1014. Call attributes are data relating to the attributes of a call between a user and a customer, and include user attributes, the name or attributes of the organization to which the user belongs, information about the user's emotions during the call (user attribute information), customer attributes, the name or attributes of the organization to which the customer belongs, information about the customer's emotions during the call (customer attribute information), call category, caller / recipient type, and information about emotions during the call (call attribute information). In other words, call data is characterized by call attributes such as user attribute information, customer attribute information, and call attribute information. In this disclosure, call attributes include attribute information about individual users and customers, but do not include information about the user's and customer's surrounding environment or call environment. For example, it does not include information about noise or ambient noise conditions around the user and customer.
[0054] <Outgoing call processing> Outgoing call processing is the process of a user making an outgoing call to a customer.
[0055] <Overview of outgoing call processing> The outgoing call process is a series of operations in which the user selects a customer they wish to call from among multiple customers displayed on the screen of the user terminal 20, and then makes a call to that customer by performing the call operation.
[0056] <Details of the outgoing call process> This section describes the transmission process of Information Processing System 1 when a user sends a message to a customer.
[0057] When a user makes a call to a customer, the following processes are executed in Information Processing System 1.
[0058] The user operates the user terminal 20 to launch a web browser and access the website of the CRM service provided by the CRM system 30. By opening the customer management screen provided by the CRM service, the user can view a list of their customers on the display 2081 of the user terminal 20. Specifically, the user terminal 20 sends a request to the CRM system 30 to display a list of CRMID2013 and customers. Upon receiving the request, the CRM system 30 searches the customer table 3012 and sends information about the user's customers, such as customer ID, name, phone number, customer attributes, customer organization name, and customer organization attributes, to the user terminal 20. The user terminal 20 displays the received customer information on its display 2081.
[0059] The user selects the customer they wish to call from the list of customers displayed on the user terminal 20's display 2081. With the customer selected, the user sends a request including the phone number to the CRM system 30 by pressing the "Call" button or the phone number button displayed on the user terminal 20's display 2081. The CRM system 30, upon receiving the request, sends the request including the phone number to the server 10. The server 10, upon receiving the request, sends a call request to the voice server (PBX) 40. Upon receiving the call request, the voice server (PBX) 40 makes a call to the customer terminal 50 based on the received phone number.
[0060] Accordingly, the user terminal 20 controls the speaker 2082 and other components to make a sound indicating that a call is being made by the voice server (PBX) 40. In addition, the display 2081 of the user terminal 20 displays information indicating that a call is being made to the customer by the voice server (PBX) 40. For example, the display 2081 of the user terminal 20 may display the words "Calling".
[0061] The customer can make the customer terminal 50 ready for a call by lifting the handset (not shown) on the customer terminal 50 or by pressing the "Receive" button displayed on the customer terminal 50's touch panel 506 when an incoming call is received. In response, the voice server (PBX) 40 sends information indicating that the customer terminal 50 has responded (hereinafter referred to as a "response event") to the user terminal 20 via the server 10, CRM system 30, etc. As a result, the user and the customer become able to communicate using the user terminal 20 and the customer terminal 50, respectively, and can communicate with each other. Specifically, the user's voice, picked up by the microphone 2062 of the user terminal 20, is output from the speaker 5082 of the customer terminal 50. Similarly, the customer's voice, picked up by the microphone 5081 of the customer terminal 50, is output from the speaker 2082 of the user terminal 20.
[0062] When the user terminal 20 becomes ready to make a call, the display 2081 receives an answer event and displays information indicating that a call is in progress. For example, the display 2081 of the user terminal 20 may display the text "Answering".
[0063] <Incoming Call Processing> Incoming call processing is the process by which a user receives an incoming call from a customer.
[0064] <Overview of incoming call processing> Incoming call processing is a series of processes that occur when a customer makes a call to a user while the user has launched an application on the user terminal 20, and the user receives the call.
[0065] <Details of incoming call processing> This section describes the incoming call processing of Information Processing System 1 when a user receives an incoming call from a customer.
[0066] When a user receives a call from a customer, the following processes are executed in Information Processing System 1.
[0067] The user operates the user terminal 20 to launch a web browser and access the website for the CRM service provided by the CRM system 30. At this time, the user is assumed to be logged into the CRM system 30 with their account in the web browser and waiting. The user only needs to be logged into the CRM system 30 and may be performing other tasks related to the CRM service.
[0068] The customer operates the customer terminal 50, enters a predetermined telephone number assigned to the voice server (PBX) 40, and makes a call to the voice server (PBX) 40. The voice server (PBX) 40 receives the call from the customer terminal 50 as an incoming event.
[0069] The voice server (PBX) 40 sends an incoming call event to the server 10. Specifically, the voice server (PBX) 40 sends an incoming call request to the server 10, including the customer's telephone number 5011. The server 10 then sends the incoming call request to the user terminal 20 via the CRM system 30. Accordingly, the user terminal 20 controls the speaker 2082 and other components to make a sound indicating that an incoming call is being received by the voice server (PBX) 40. The display 2081 of the user terminal 20 displays information indicating that an incoming call is being received from a customer by the voice server (PBX) 40. For example, the display 2081 of the user terminal 20 may display the words "Incoming Call".
[0070] The user terminal 20 accepts responses from the user. Responses are performed, for example, by the user lifting the handset (not shown) on the user terminal 20, or by the user using the mouse 2066 to press a button on the user terminal 20's display 2081 that displays "Answer the call". When the user terminal 20 receives a response request, it sends a response request to the voice server (PBX) 40 via the CRM system 30 and server 10. The voice server (PBX) 40 receives the received response request and establishes voice communication. As a result, the user terminal 20 becomes ready to communicate with the customer terminal 50. The display 2081 of the user terminal 20 displays information indicating that a call is in progress. For example, the display 2081 of the user terminal 20 may display the words "Call in Progress".
[0071] When a call becomes possible, the following voice conversion processes are executed: (First Embodiment), (Second Embodiment), (Third Embodiment), and (Fourth Embodiment). The system may be configured to perform one of the following speech conversion processes for each specific speaker-listener pair: speech conversion process (first embodiment), speech conversion process (second embodiment), or speech conversion process (third embodiment). The speech conversion process (fourth embodiment) may be executed simultaneously with the speech conversion process (first embodiment), speech conversion process (second embodiment), and speech conversion process (third embodiment). If a call is taking place involving three or more people, one of the following speech conversion processes may be performed for each pair of speakers and listeners among the three: speech conversion process (first embodiment), speech conversion process (second embodiment), or speech conversion process (third embodiment). In other words, a different speech conversion process may be performed for each pair of two different speakers and listeners among the three.
[0072] <Variation> Furthermore, the method by which a user can become available to communicate with a customer is not limited to outgoing or incoming call processing; any method is permitted to enable communication between the user and the customer. For example, a virtual communication space called a "room" for communication between the user and the customer could be created on the server 10, and the user and the customer could become available to communicate by accessing this room via a web browser or application program stored on the user terminal 20 and the customer terminal 50. In this case, the voice server (PBX) 40 would not be necessary. Specifically, the user who is the host of the call operates the input device 206 of the user terminal 20 and sends a request to the server 10 to initiate the call. Upon receiving the request, the control unit 104 of the server 10 issues room identification information, such as a unique room ID, and sends a response to the user terminal 20. The user sends the received room identification information to the customer, the other party to the call, via email or any other means of communication. The user can enter the room by operating the input device 206 of the user terminal 20, accessing the URL that provides room-related services on the server 10 using a web browser, and entering the room identification information. Similarly, the customer can enter the room by operating the touch panel 506 of the customer terminal 50, accessing the URL that provides room-related services on the server 10 using a web browser, and entering the room identification information. As a result, the user and the customer can communicate via the user terminal 20 and customer terminal 50, respectively, within a virtual call space called a room, which is associated with the room identification information. By entering room identification information, multiple users and multiple customers can enter a single room. This allows multiple users and multiple customers to communicate via user terminals 20 and customer terminals 50 within a virtual communication space called a room, which is associated with the room identification information.
[0073] <Call memory processing> Call memory processing is the process of storing data related to calls made between a user and a customer.
[0074] <Overview of Call Memory Processing> The call memory processing is a series of processes that store data related to a call in the call table 1014 when a call is initiated between a user and a customer.
[0075] <Details of call memory processing> When a call is initiated between a user and a customer, the voice server (PBX) 40 records the voice data related to the call between the user and the customer and sends it to the server 10. Upon receiving the voice data, the control unit 104 of the server 10 creates a new record in the call table 1014 and stores the data related to the call between the user and the customer. Specifically, the control unit 104 of the server 10 stores the user ID, customer ID, call category, call type, and the content of the voice data in the call table 1014.
[0076] The control unit 104 of the server 10 obtains the user's user ID 2011 from the user terminal 20 during outgoing or incoming call processing and stores it in the user ID field of a new record. The control unit 104 of server 10 queries the CRM system 30 based on the telephone number during outgoing or incoming call processing. The CRM system 30 retrieves the customer ID by searching the customer table 3012 using the telephone number and sends it to server 10. The control unit 104 of server 10 stores the retrieved customer ID in the customer ID field of a new record. The control unit 104 of the server 10 stores the call category value, which has been set in advance for each user or customer, in the call category field of the new record. Alternatively, the call category may be stored by the user selecting or entering a value for each call. The control unit 104 of the server 10 identifies whether the call being made was initiated by the user or by the customer, and stores either an outbound (initiated by the user) or inbound (initiated by the customer) value in the call type field of the new record. The control unit 104 of server 10 stores the voice data received from the voice server (PBX) 40 in the voice data field of a new record. Alternatively, the voice data may be stored as a voice data file in another location, and reference information (path) to the voice data file may be stored after the call ends. Furthermore, the control unit 104 of server 10 may be configured to store data after the call ends.
[0077] <Voice Conversion Processing (First Embodiment)> The speech conversion process (first embodiment) is a process that outputs converted speech data to the customer by applying a generative model 1022 selected based on call attributes to the speech data spoken by the user. This allows customers to communicate with users using a more suitable voice.
[0078] <Overview of the speech conversion process (first embodiment)> The voice conversion process (first embodiment) is initiated when the user and the customer are in a state where they can communicate. The voice conversion process (first embodiment) is a series of processes that acquire call attributes, input the call attributes and the voice data spoken by the user as input data into the generation model 1022, and output the converted voice data to the customer.
[0079] <Details of the speech conversion process (first embodiment)> In step S101, when the user and the customer are able to communicate, the voice conversion process (first embodiment) is started.
[0080] In step S102, the voice conversion unit 1042 of the server 10 acquires call attributes related to the call. Specifically, the voice conversion unit 1042 of server 10 searches the user ID field in user table 1012 based on the user's user ID and obtains the organization ID and user attribute fields. Based on the obtained organization ID, the voice conversion unit 1042 of server 10 searches the organization ID field in organization table 1013 and obtains the organization name and organization attribute fields. In other words, the voice conversion unit 1042 of server 10 obtains attribute information about the user. The voice conversion unit 1042 of server 10 may also estimate the user's emotional state from the user's spoken voice and obtain the user's emotional information as attribute information about the user. The voice conversion unit 1042 of server 10 sends a query request containing the customer's customer ID to the CRM system 30. Based on the customer ID included in the received request, the CRM system 30 searches the customer ID field in the customer table 3012, retrieves the customer attribute, customer organization name, and customer organization attribute fields, and sends them to server 10. The voice conversion unit 1042 of server 10 retrieves the customer attribute, customer organization name, and customer organization attribute fields from the CRM system 30. In other words, the voice conversion unit 1042 of server 10 retrieves attribute information about the customer. The voice conversion unit 1042 of server 10 may also estimate the customer's emotional state from the customer's spoken voice and retrieve the customer's emotional information as attribute information about the customer. The voice conversion unit 1042 of server 10 refers to the call table 1014 and obtains information on the call category and call type included in the call data stored by the call storage process. Based on the call ID related to the call, the voice conversion unit 1042 of server 10 searches the call ID item in the call table 1014 and obtains the call category and call type items. In other words, the voice conversion unit 1042 of server 10 obtains attribute information related to the call. The voice conversion unit 1042 of server 10 may also estimate the emotional state of the user and customer from the user's and customer's spoken voice and obtain user and customer emotional information as attribute information related to the call.
[0081] The voice conversion unit 1042 of the server 10 may acquire at least one of the following as call attributes: user attribute information, customer attribute information, or call attribute information. For example, it may acquire only user attribute information as call attributes. Furthermore, the voice conversion unit 1042 of the server 10 may acquire one of the following as a call attribute: user attributes, the name or attributes of the organization to which the user belongs, user sentiment information, customer attributes, the name or attributes of the organization to which the customer belongs, customer sentiment information, call category, caller / recipient type, or sentiment information of the call. For example, only the call category may be acquired as a call attribute.
[0082] In step S104, the voice conversion unit 1042 of the server 10 acquires call audio from the user and converts the acquired call audio. At this time, the voice conversion unit 1042 of the server 10 converts the acquired call audio based on the call attributes acquired in step S102. The voice conversion unit 1042 of the server 10 converts the acquired call audio by applying the generation model 1022 to the call attributes and call audio acquired in step S102. Specifically, the voice conversion unit 1042 of server 10 sequentially acquires voice data spoken by the user from the voice server (PBX) 40. It is desirable for the voice conversion unit 1042 of server 10 to acquire the spoken voice data with as little delay as possible after the user has spoken. The voice conversion unit 1042 of server 10 inputs the acquired call attributes and voice data as input data to the generation model 1022 and obtains the output converted voice data.
[0083] The voice conversion unit 1042 of the server 10 may selectively switch and apply one of several generation models 1022 depending on the type of evaluation metric for the converted voice data output to the customer, and convert the voice data. For example, the audio data may be transformed using a generative model 1022 that is more trustworthy to the customer. For example, the audio data may be transformed using a generative model 1022 that is easier for the customer to listen to (easier to understand). For example, the audio data may be transformed using a generative model 1022 that is suitable for the customer in terms of evaluation metrics such as SIIB, HASPI, ESTOI, PESQ, and ViSQOL, as well as evaluation metrics such as reliability, trustworthiness, pleasantness, comfort, preference, stress level, intimidation level, and interest.
[0084] The voice conversion unit 1042 of the server 10 may convert the call audio using at least one of the following as call attributes: user attribute information, customer attribute information, or call attribute information. For example, it may convert the call audio using only user attribute information as call attributes. Furthermore, the voice conversion unit 1042 of the server 10 may convert the call audio using one of the following as a call attribute: user attribute, name or attribute of the organization to which the user belongs, user sentiment information, customer attribute, name or attribute of the organization to which the customer belongs, customer sentiment information, call category, caller / recipient type, or sentiment information of the call. For example, the call audio may be converted using only the call category as a call attribute.
[0085] In step S105, the voice conversion unit 1042 of the server 10 outputs the call audio converted in step S104 to the customer. The voice conversion unit 1042 of the server 10 transmits the converted voice data to the voice server (PBX) 40. The voice server (PBX) 40 outputs the received converted voice data to the customer terminal 50. The speaker 5082 of the customer terminal 50 outputs the received converted voice data as the user's call audio. In other words, the audio data related to the user's voice, collected by the microphone 2062 of the user terminal 20, is converted into converted audio data by the audio conversion unit 1042 of the server 10 and output from the speaker 5082 of the customer terminal 50.
[0086] <Voice Conversion Processing (Second Embodiment)> The speech conversion process (second embodiment) is a process that outputs converted speech data to the customer by applying the speech processing model 1023 to the speech data spoken by the user and applying the identified speech processing content. This allows users to select speech processing that can convert the speech into a format more suitable for the customer.
[0087] <Overview of the speech conversion process (second embodiment)> The voice conversion process (second embodiment) is initiated when the user and the customer are in a state where they can communicate. The voice conversion process (second embodiment) is a series of processes that acquire call attributes, input the call attributes as input data to the voice processing model 1023, select the voice processing content based on the output voice processing ID, apply the selected voice processing content to the voice data spoken by the user, and output the converted voice data to the customer.
[0088] <Details of the speech conversion process (second embodiment)> In step S301, when the user and the customer are able to communicate, the voice conversion process (second embodiment) is started.
[0089] In step S302, the voice conversion unit 1042 of the server 10 acquires call attributes related to the call. Step S302 is the same as step S102 in the voice conversion process (first embodiment), so its explanation is omitted.
[0090] In step S303, the voice conversion unit 1042 of the server 10 selects a predetermined voice processing from among multiple voice processing methods based on the acquired call attributes. The voice conversion unit 1042 of the server 10 selects a predetermined voice processing method by applying the voice processing model 1023 to the call attributes acquired in step S302. Specifically, the voice conversion unit 1042 of server 10 inputs the acquired call attributes as input data into the voice processing model 1023 and obtains the output voice processing ID. Based on the acquired voice processing ID, the voice conversion unit 1042 of server 10 searches the voice processing table 1015 for the corresponding voice processing ID item and obtains the voice processing content. In other words, the voice conversion unit 1042 of server 10 identifies and selects the voice processing content based on the call attributes.
[0091] In step S304, the voice conversion unit 1042 of the server 10 acquires call audio from the user and converts the acquired call audio. At this time, the voice conversion unit 1042 of the server 10 converts the acquired call audio based on the call attributes acquired in step S302. The voice conversion unit 1042 of the server 10 converts the call audio by applying a predetermined voice processing selected in step S302 to the call audio. Specifically, the voice conversion unit 1042 of server 10 sequentially acquires voice data spoken by the user from the voice server (PBX) 40. It is desirable for the voice conversion unit 1042 of server 10 to acquire the spoken voice data with as little delay as possible after the user has spoken. The voice conversion unit 1042 of the server 10 applies the voice processing content selected in step S303 to the acquired call attributes and voice data, and obtains the output converted voice data.
[0092] The voice conversion unit 1042 of the server 10 may selectively switch and apply one of several voice processing models 1023 according to the type of evaluation indicator for the converted voice data output to the customer, and convert the voice data. For example, the audio data may be transformed using an audio processing model 1023 that is more trustworthy to the customer. For example, the audio data may be transformed using an audio processing model 1023 that is easier for the customer to listen to (easier to understand). For example, the audio data may be transformed using an audio processing model 1023 that is suitable for the customer in terms of evaluation metrics such as SIIB, HASPI, ESTOI, PESQ, and ViSQOL, as well as evaluation metrics such as reliability, trustworthiness, pleasantness, comfort, preference, stress level, intimidation level, and interest.
[0093] The voice conversion unit 1042 of the server 10 may convert the call audio using at least one of the following as call attributes: user attribute information, customer attribute information, or call attribute information. For example, it may convert the call audio using only user attribute information as call attributes. Furthermore, the voice conversion unit 1042 of the server 10 may convert the call audio using one of the following as a call attribute: user attribute, name or attribute of the organization to which the user belongs, user sentiment information, customer attribute, name or attribute of the organization to which the customer belongs, customer sentiment information, call category, caller / recipient type, or sentiment information of the call. For example, the call audio may be converted using only the call category as a call attribute.
[0094] In step S305, the voice conversion unit 1042 of the server 10 outputs the call audio converted in step S304 to the customer. Step S305 is the same as step S105 in the voice conversion process (first embodiment), so its explanation is omitted.
[0095] <Voice Conversion Processing (Third Embodiment)> The speech conversion process (third embodiment) is a process that outputs converted speech data to the customer by applying the speech processing model 1023 selected by the user to the speech data spoken by the user. This allows customers to have more comfortable voice conversations with users, depending on the user's selection and instructions.
[0096] <Overview of the speech conversion process (third embodiment)> The voice conversion process (third embodiment) is initiated when the user and the customer are in a state where they can communicate. The voice conversion process (third embodiment) is a series of processes in which the user selects the voice processing content, applies the selected voice processing content to the voice data spoken by the user, and outputs converted voice data to the customer.
[0097] <Details of the speech conversion process (third embodiment)> In step S501, when the user and the customer are able to communicate, the voice conversion process (third embodiment) is started.
[0098] In step S503, based on the selection instruction received from the user, a predetermined voice processing method is selected from among several voice processing methods. For example, selecting a voice processing method that reduces intonation can reduce the intimidating impression on the customer. If the customer is female, selecting a voice processing method that reduces the volume of a loud voice can reduce the intimidating impression on the customer. Users may choose to perform voice processing at any time during a call between them and a customer. Alternatively, users may pre-select voice processing before the call begins. Specifically, the user operates the input device 206 of the user terminal 20 to send a request to the server 10 that includes a voice processing ID related to the voice processing content they wish to apply. The voice conversion unit 1042 of the server 10 searches the voice processing table 1015 for the voice processing ID included in the received request and retrieves the voice processing content. In other words, the voice conversion unit 1042 of the server 10 identifies and selects the voice processing content based on the selection instruction received from the user.
[0099] Figure 15 illustrates an example of the display screen of the user terminal 20 in the voice conversion process (third embodiment). The user terminal 20's display 2081 shows the call screen 80. The call screen 80 shows the customer information 801 currently on the call and a user interface 802 for selecting the voice processing content. The customer information 801 may include attribute information about the customer, such as customer attributes, the name or organizational attributes of the organization to which the customer belongs, and customer sentiment information. The user selects a predetermined voice processing from among several voice processing options by operating the input device 206 of the user terminal 20 and pressing a switch 803 associated with the voice processing content. Figure 15 shows that the voice processing content with voice processing ID M002 has been selected.
[0100] <Variation> Furthermore, the user may operate the input device 206 of the user terminal 20 to select the type of evaluation metric they wish to optimize. Specifically, the user may operate the input device 206 of the user terminal 20 to select evaluation metrics such as SIIB, HASPI, ESTOI, PESQ, ViSQOL, or evaluation metrics such as reliability, trustworthiness, comfort, pleasantness, preference, stress level, intimidation level, and interest. For example, the user may operate the input device 206 of the user terminal 20 to select options such as making the audio easier for customers to hear (easier to understand) or making the audio more trustworthy to customers. The display 2081 of the user terminal 20 may also be configured to display a list of selectable evaluation metrics to the user. The user may also select the type of evaluation metric they wish to optimize by operating the input device 206 of the user terminal 20 to select the item they wish to optimize from the displayed list of evaluation metrics.
[0101] The user terminal 20 sends a request to the server 10 that includes the type of evaluation metric selected.
[0102] The speech conversion unit 1042 of server 10 selects a speech processing model 1023 based on the type of evaluation metric included in the received request. Specifically, the speech conversion unit 1042 of server 10 selects a speech processing model 1023 that has been optimized (learned) to obtain a higher evaluation metric, depending on the type of evaluation metric received. At this time, the voice conversion unit 1042 of the server 10 acquires call attributes related to the call, similar to steps S302 and S303 of the voice conversion process (second embodiment), and selects a predetermined voice process from among multiple voice processes based on the acquired call attributes.
[0103] Furthermore, the speech conversion unit 1042 of the server 10 may select a generative model 1022 based on the type of evaluation metric included in the received request. Specifically, the speech conversion unit 1042 of the server 10 may select a generative model 1022 that has been optimized (learned) to obtain a larger evaluation metric, depending on the type of evaluation metric received.
[0104] In other words, in step S503, the voice conversion unit 1042 of the server 10 may directly receive the voice processing ID and identify and select the voice processing content based on the selection instruction received from the user, or it may indirectly identify the voice processing ID and identify and select the voice processing content using the voice processing model 1023. Alternatively, the voice conversion unit 1042 of the server 10 may identify and select the generation model 1022 based on the selection instruction received from the user.
[0105] In step S504, the voice conversion unit 1042 of the server 10 acquires the call audio from the user and converts the acquired call audio. At this time, the voice conversion unit 1042 of the server 10 converts the acquired call audio based on the voice processing content selected in step S503. Alternatively, the voice conversion unit 1042 of the server 10 may convert the acquired call audio based on the generation model 1022 selected in step S503. Specifically, the voice conversion unit 1042 of server 10 sequentially acquires voice data spoken by the user from the voice server (PBX) 40. It is desirable for the voice conversion unit 1042 of server 10 to acquire the spoken voice data with as little delay as possible after the user has spoken. The voice conversion unit 1042 of server 10 applies the voice processing content selected in step S503 to the acquired call attributes and voice data, and obtains the output converted voice data. Alternatively, the voice conversion unit 1042 of server 10 may apply the generation model 1022 selected in step S503 to the acquired call attributes and voice data, and obtain the output converted voice data.
[0106] In step S505, the voice conversion unit 1042 of the server 10 outputs the call audio converted in step S504 to the customer. Step S505 is the same as step S105 in the voice conversion process (first embodiment), so its explanation is omitted.
[0107] <Variation> In the first, second, and third embodiment of the speech conversion process, only the user's spoken voice can be converted, and the system may be configured in such a way that customer spoken voices cannot be converted. Specifically, the customer may not be able to make selection instructions regarding voice processing, and the voice conversion unit 1042 of the server 10 may be configured not to accept selection instructions regarding voice processing from the customer. In this case, the voice conversion unit 1042 of the server 10 may be configured to output the customer's spoken voice to the user without converting it. Specifically, the voice data of the customer's voice collected by the microphone 5081 of the customer terminal 50 is output from the speaker 2082 of the user terminal 20 without being converted by the voice conversion unit 1042 of the server 10. This allows users to conduct calls with customers while simultaneously listening to their voices without altering them.
[0108] <Voice Conversion Processing (Fourth Embodiment)> The speech conversion process (fourth embodiment) is a process that outputs converted speech data to the user by applying the speech processing model 1023 selected by the user to the speech data spoken by the customer.
[0109] <Overview of the speech conversion process (fourth embodiment)> The voice conversion process (fourth embodiment) is initiated when the user and the customer are in a state where they can communicate. The voice conversion process (fourth embodiment) is a series of processes in which the user selects the voice processing content, and the selected voice processing content is applied to the voice data spoken by the customer, and the converted voice data is output to the user.
[0110] <Details of the speech conversion process (fourth embodiment)> When the user and the customer are able to communicate, the voice conversion process (fourth embodiment) is initiated. Based on the selection instructions received from the user, a predetermined voice processing method is selected from among several voice processing methods. For example, a predetermined voice processing method may be selected from among several voice processing methods that results in a voice that is easier for the user to listen to. For example, a predetermined voice processing method may be selected from among several voice processing methods that provides a greater sense of comfort to the user. For example, a predetermined voice processing method may be selected from among several voice processing methods that matches the user's preference, reduces stress levels, or increases interest. For example, if a customer is angry, the user can reduce the psychological stress associated with interacting with the customer by selecting a voice processing method from among several voice processing methods that reduces intonation or lowers the volume. Users may choose to perform voice processing at any time during a call between them and a customer. Alternatively, users may pre-select voice processing before the call begins. Specifically, the user operates the input device 206 of the user terminal 20 to send a request to the server 10 that includes a voice processing ID related to the voice processing content they wish to apply. The voice conversion unit 1042 of the server 10 searches the voice processing table 1015 for the voice processing ID included in the received request and retrieves the voice processing content. In other words, the voice conversion unit 1042 of the server 10 identifies and selects the voice processing content based on the selection instruction received from the user.
[0111] <Variation> Furthermore, the user may operate the input device 206 of the user terminal 20 to select the type of evaluation metric they wish to optimize. Specifically, the user may operate the input device 206 of the user terminal 20 to select evaluation metrics such as SIIB, HASPI, ESTOI, PESQ, ViSQOL, or evaluation metrics such as reliability, trustworthiness, comfort, pleasantness, preference, stress level, intimidation level, and interest. For example, the user may operate the input device 206 of the user terminal 20 to select options that make the audio easier to hear (easier to understand) or options that provide the user with greater comfort. The display 2081 of the user terminal 20 may also be configured to display a list of selectable evaluation metrics to the user. The user may also select the type of evaluation metric they wish to optimize by operating the input device 206 of the user terminal 20 to select the item they wish to optimize from the displayed list of evaluation metrics. The user terminal 20 sends a request to the server 10 that includes the type of evaluation metric selected. The speech conversion unit 1042 of server 10 selects a speech processing model 1023 based on the type of evaluation metric included in the received request. Specifically, the speech conversion unit 1042 of server 10 selects a speech processing model 1023 that has been optimized (learned) to obtain a higher evaluation metric, depending on the type of evaluation metric received. At this time, the voice conversion unit 1042 of the server 10 acquires call attributes related to the call, similar to steps S302 and S303 of the voice conversion process (second embodiment), and selects a predetermined voice process from among multiple voice processes based on the acquired call attributes. However, unlike steps S302 and S303 of the voice conversion process (second embodiment), the attribute information related to the user and the attribute information related to the customer are swapped and applied as call attributes. This is because, in the voice conversion process (fourth embodiment), the customer becomes the speaker of the voice data, and the user becomes the listener of the converted voice data.
[0112] In other words, the voice conversion unit 1042 of the server 10 may directly receive a voice processing ID and identify and select the voice processing content based on the selection instruction received from the user, or it may indirectly identify the voice processing ID and identify and select the voice processing content using the voice processing model 1023.
[0113] The voice conversion unit 1042 of server 10 acquires call audio from the customer and converts the acquired call audio. At this time, the voice conversion unit 1042 of server 10 converts the acquired call audio based on the selected voice processing content. Specifically, the voice conversion unit 1042 of server 10 sequentially acquires voice data spoken by the customer from the voice server (PBX) 40. It is desirable for the voice conversion unit 1042 of server 10 to acquire the spoken voice data with as little delay as possible after the customer has spoken. The voice conversion unit 1042 of server 10 applies the selected voice processing content to the acquired call attributes and voice data, and obtains the output converted voice data.
[0114] The voice conversion unit 1042 of the server 10 outputs the converted call audio from step S504 to the user. Step S505 is the same as step S105 in the voice conversion process (first embodiment), except that the user and customer are swapped, so its explanation is omitted.
[0115] <Variation> In steps S104, S304, and S504 of the voice conversion process (first embodiment), voice conversion process (second embodiment), and voice conversion process (third embodiment), respectively, if the user and customer make a call in a virtual call space called a room, the voice conversion unit 1042 of the server 10 may be configured to sequentially acquire the voice data spoken by the user received by the server 10 without going through the voice server (PBX) 40. Similarly, the voice conversion unit 1042 of the server 10 may be configured to sequentially acquire the voice data spoken by the customer received by the server 10 without going through the voice server (PBX) 40.
[0116] Similarly, in steps S105, S305, and S505 of the voice conversion process (first embodiment), voice conversion process (second embodiment), and voice conversion process (third embodiment), respectively, if the user and customer make a call within a virtual call space called a room, the voice conversion unit 1042 of the server 10 may be configured to output the converted voice data to the customer terminal 50 without going through the voice server (PBX) 40. In other words, the voice server (PBX) 40 is not an essential configuration requirement in the voice conversion process (first embodiment), voice conversion process (second embodiment), and voice conversion process (third embodiment).
[0117] In the voice conversion process (fourth embodiment), when the user and customer make a call in a virtual call space called a room, the voice conversion unit 1042 of the server 10 may be configured to sequentially acquire voice data spoken by the customer received by the server 10 without going through the voice server (PBX) 40. Similarly, the voice conversion unit 1042 of the server 10 may be configured to sequentially acquire voice data spoken by the user received by the server 10 without going through the voice server (PBX) 40.
[0118] Similarly, in the voice conversion process (fourth embodiment), if the user and customer make a call within a virtual call space called a room, the voice conversion unit 1042 of the server 10 may be configured to output the converted voice data to the user terminal 20 without going through the voice server (PBX) 40. In other words, the voice server (PBX) 40 is not an essential configuration requirement in the voice conversion process (fourth embodiment).
[0119] <Learning Process> The learning process for the generative model 1022 and the speech processing model 1023 is described below. Note that the learning process described below pertains to a specific evaluation metric (for example, the first metric). If multiple evaluation metrics are used, the learning process will be performed for each of the multiple generative models 1022 and speech processing models 1023 prepared for each evaluation type, such as the first metric and the second metric.
[0120] <Training process for generative model 1022> The training process for generative model 1022 involves training the training parameters of the deep neural network included in generative model 1022 using deep learning.
[0121] <Overview of the training process for generative model 1022> The learning process for generative model 1022 involves using deep learning to train the learning parameters of the deep neural network included in generative model 1022, taking user attribute information, customer attribute information, call attribute information, and voice data as input data (input vectors), so that it outputs transformed voice data that yields a higher evaluation metric.
[0122] <Details of the training process for generative model 1022> The learning unit 1051 of server 10 obtains call attributes and audio data items from the training dataset 1031. Using the call attributes and audio data as input data, the learning unit 1051 of server 10 applies the learning parameters of the deep neural network included in the generative model 1022 while changing them, and generates multiple converted audio data. In this case, the learning unit 1051 of the server 10 may include at least one of the following as call attributes in the input data: user attribute information, customer attribute information, and call attribute information, and exclude the rest before performing the learning process. Alternatively, the learning unit 1051 of the server 10 may include at least one of the following as call attributes in the input data: user attribute, organization name or organizational attribute of the organization to which the user belongs, user sentiment information, customer attribute, organization name or organizational attribute of the organization to which the customer belongs, customer sentiment information, call category, caller / recipient type, and call sentiment information, and exclude the rest before performing the learning process. The learning unit 1051 of server 10 applies call attributes and voice data as input data to the evaluation model 1021, thereby obtaining listener-side evaluation metrics for each of the multiple converted voice data. The learning unit 1051 of server 10 optimizes the learning parameters of the deep neural network included in the generative model 1022 to obtain higher evaluation metrics. This allows us to obtain a generative model 1022 that takes call attributes and voice data as input data and outputs converted voice data that yields a larger evaluation metric.
[0123] The learning unit 1051 of server 10 may configure the generative model 1022 as any learning model such as a GAN (Generative Adversarial Network).
[0124] <Training process of speech processing model 1023> The training process for speech processing model 1023 involves training the training parameters of the deep neural network included in speech processing model 1023 using deep learning.
[0125] <Overview of the learning process for speech processing model 1023> The training process for the speech processing model 1023 involves using call attributes such as user attributes, customer attributes, call category, and call type as input data (input vectors) to train the training parameters of the deep neural network included in the speech processing model 1023 using deep learning so that it outputs speech processing content that yields a higher evaluation metric. Specifically, in the training process of the speech processing model 1023, the speech processing model 1023 is a learning model that takes call attributes as input data (input vector) and outputs a speech processing ID in the speech processing table 1015, which provides a larger evaluation metric.
[0126] <Details of the learning process for speech processing model 1023> The learning unit 1051 of server 10 generates multiple audio processing data by applying the audio processing content stored in the audio processing table 1015 to the audio data associated with call attributes included in the learning dataset 1031. The learning unit 1051 of server 10 takes call attributes as input data and applies multiple speech processing data associated with those call attributes to the evaluation model 1021, thereby obtaining an evaluation index for each of the multiple speech processing data. The learning unit 1051 of server 10 optimizes the learning parameters of the deep neural network included in the speech processing model 1023 so that a speech processing ID that yields a higher evaluation index can be obtained. This makes it possible to obtain a speech processing model 1023 that takes call attributes as input data and outputs a speech processing ID that provides a larger evaluation metric.
[0127] In addition, during the training process of the speech processing model 1023, the input data (input vector) may include speech data in addition to call attributes.
[0128] <Basic Computer Hardware Configuration> Figure 16 is a block diagram showing the basic hardware configuration of computer 90. Computer 90 comprises at least a processor 901, main memory 902, auxiliary storage 903, and a communication interface IF991. These are electrically connected to each other by a communication bus 921.
[0129] The processor 901 is hardware for executing the instruction set written in a program. The processor 901 consists of an arithmetic unit, registers, peripheral circuits, etc.
[0130] Main memory 902 is used to temporarily store programs and data processed by programs, etc. For example, it is a volatile memory such as DRAM (Dynamic Random Access Memory).
[0131] Auxiliary storage device 903 refers to a storage device for saving data and programs. Examples include flash memory, HDD (Hard Disc Drive), magneto-optical disk, CD-ROM, DVD-ROM, and semiconductor memory.
[0132] The IF991 communication interface is an interface for inputting and outputting signals for communication with other computers via a network using wired or wireless communication standards. A network consists of various mobile communication systems, such as the internet, LANs, and wireless base stations. For example, a network includes 3G, 4G, and 5G mobile communication systems, LTE (Long Term Evolution), and wireless networks that can connect to the internet via designated access points (e.g., Wi-Fi®). When connecting wirelessly, communication protocols include, for example, Z-Wave®, ZigBee®, and Bluetooth®. When connecting via a wired connection, the network also includes connections made directly via USB (Universal Serial Bus) cables, etc.
[0133] Furthermore, by distributing all or part of each hardware configuration across multiple computers 90 and connecting them to each other via a network, a computer 90 can be virtually realized. Thus, the concept of computer 90 includes not only a computer 90 housed in a single enclosure or case, but also a virtualized computer system.
[0134] <Basic Functional Configuration of Computer 90> The functional configuration of the computer realized by the basic hardware configuration of computer 90 (Figure 16) is described below. The computer comprises at least one functional unit: a control unit, a memory unit, and a communication unit.
[0135] Furthermore, the functional units of computer 90 can also be realized by distributing all or part of each functional unit across multiple computers 90 interconnected via a network. The concept of computer 90 includes not only a single computer 90 but also a virtualized computer system.
[0136] The control unit is realized when the processor 901 reads various programs stored in the auxiliary storage device 903, loads them into the main memory device 902, and executes processing according to those programs. The control unit can realize various functional units that perform information processing depending on the type of program. In this way, the computer is realized as an information processing device that performs information processing.
[0137] The memory unit is implemented by the main memory 902 and the auxiliary memory 903. The memory unit stores data, various programs, and various databases. The processor 901 can also reserve memory areas corresponding to the memory unit in the main memory 902 or the auxiliary memory 903 according to the program. The control unit can also cause the processor 901 to perform operations such as adding, updating, and deleting data stored in the memory unit according to the various programs.
[0138] A database, specifically a relational database, is used to manage and link together tabular data sets called masters, which are structurally defined by rows and columns. In a database, tables are called tables, masters are called masters, the columns of tables are called columns, and the rows of tables are called records. In a relational database, relationships can be established and linked between tables and masters. Typically, each table and master has a primary key column to uniquely identify records, but setting a primary key column is not mandatory. The control unit can instruct the processor 901 to add, delete, or update records in specific tables and masters stored in the memory unit, according to various programs.
[0139] Furthermore, the databases and masters in this disclosure may include any data structures (lists, dictionaries, associative arrays, objects, etc.) in which information is structurally defined. Data structures also include data that can be considered as data structures by combining data with functions, classes, methods, etc., written in any programming language.
[0140] The communication unit is implemented by the communication IF991. The communication unit provides the functionality to communicate with other computers 90 via the network. The communication unit can receive information transmitted from other computers 90 and input it to the control unit. The control unit can cause the processor 901 to perform information processing on the received information according to various programs. The communication unit can also transmit information output from the control unit to other computers 90.
[0141] <Note> The details described in each of the above embodiments are noted below.
[0142] (Note 1) A program comprising a processor and a memory unit, which causes a computer to perform a telephone call between a first user and a second user, wherein the program causes the processor to perform a voice acquisition step (S104, S304) for acquiring telephone audio from the first user, a conversion step (S104, S304) for converting the telephone audio acquired in the voice acquisition step, an output step (S105, S305) for outputting the converted telephone audio to the second user, and an attribute acquisition step (S102, S302) for acquiring telephone attributes related to the call, wherein the conversion step includes a step of converting the telephone audio acquired in the voice acquisition step based on the telephone attributes acquired in the attribute acquisition step. This allows customers to communicate with users using a more appropriate voice based on the call attributes during calls between multiple users.
[0143] (Note 2) The program described in Appendix 1 is a program in which the conversion step is to convert the call audio by applying a generation model to the call attributes obtained in the attribute acquisition step and the call audio obtained in the audio acquisition step. This allows customers to communicate with users using a more appropriate voice based on the call attributes during calls between multiple users.
[0144] (Note 3) The program, as described in Appendix 1, causes the processor to perform a selection step (S303) in which it selects a predetermined voice processing from among a plurality of voice processing based on the call attributes acquired in the attribute acquisition step, and the conversion step is a step in which the program converts the call audio by applying the predetermined voice processing selected in the selection step to the call audio acquired in the voice acquisition step. This allows customers to communicate with users using a more appropriate voice based on the call attributes during calls between multiple users.
[0145] (Note 4) The selection step is the step of selecting a predetermined audio processing method from among several audio processing methods that results in audio that is easier for a second user to hear, as described in Appendix 3 of the program. This allows customers to communicate with other users in a clearer voice, depending on the call attributes, during calls between multiple users.
[0146] (Note 5) The selection step is a step of selecting a predetermined voice processing method from among several voice processing methods that is more reliable for a second user, as described in Appendix 3 or 4 of the program. This allows users to project a trustworthy image to customers through calls, depending on the call attributes, during the call.
[0147] (Note 6) The selection step is a step in which a predetermined speech processing is selected (S304, S504) by applying a speech processing model to the call attributes acquired in the attribute acquisition step, as described in any of the programs described in Appendix 3 to 5. This allows customers to communicate with users using a more appropriate voice based on the call attributes during calls between multiple users.
[0148] (Note 7) A program comprising a processor and a memory unit, which enables a computer to perform a call between a first user and a second user, wherein the program causes the processor to perform a voice acquisition step (S504) to acquire call audio from the first user, a conversion step (S504) to convert the call audio acquired in the voice acquisition step, an output step (S505) to output the call audio converted in the conversion step to the second user, and a selection step (S503) to select a predetermined voice processing from among a plurality of voice processing based on a selection instruction received from the first user, wherein the conversion step is a step of converting the call audio by applying the predetermined voice processing selected in the selection step to the call audio acquired in the voice acquisition step. This allows customers to communicate with users using a more appropriate voice based on user selection instructions during calls between multiple users.
[0149] (Note 8) The program, as described in Appendix 7, causes the processor to perform an attribute acquisition step to acquire call attributes related to a call, and the selection step includes receiving an instruction from a first user to select a type of evaluation metric, and selecting a predetermined voice processing from among a plurality of voice processing based on the call attributes acquired in the attribute acquisition step and the received evaluation metric. This allows customers to communicate with users using a more appropriate voice in calls between multiple users, based on the type of evaluation metric selected by the user and the call attributes.
[0150] (Note 9) The program as described in Appendix 8 includes a selection step of receiving an instruction from a first user to select a type of evaluation metric to be optimized, and a step of selecting a predetermined voice processing method from among multiple voice processing methods that optimizes the evaluation metric, based on the call attributes obtained in the attribute acquisition step and the received evaluation metric. This allows customers to communicate with users using a more appropriate voice in calls between multiple users, based on the type of evaluation metric selected by the user and the call attributes.
[0151] (Note 10) The program as described in Appendix 8 or 9, wherein the selection step includes receiving an instruction from a first user to select a type of evaluation metric, and selecting a predetermined voice processing from among multiple voice processing methods by applying a voice processing model to the call attributes acquired in the attribute acquisition step and the received evaluation metric. This allows customers to communicate with users using a more appropriate voice in calls between multiple users, based on the type of evaluation metric selected by the user and the call attributes.
[0152] (Note 11) A program as described in any of Appendix 8 to 10, wherein the voice acquisition step includes a step of acquiring second call audio from a second user, the selection step does not allow selection instructions to be received from the second user, and the conversion step does not convert the second call audio acquired in the voice acquisition step. This allows the first user to communicate with the second user while listening to the second user's voice, without converting the second user's voice.
[0153] (Note 12) The program described in any of the appendices 1 to 6 or 8 to 11, wherein the attribute acquisition step includes a step of acquiring attribute information about a second user, and the conversion step includes a step of converting the call audio acquired in the voice acquisition step based on the attribute information about the second user acquired in the attribute acquisition step. This allows customers to communicate with other users using a more appropriate voice in calls between multiple users, based on their attribute information.
[0154] (Note 13) The program is one of the programs described in any of Appendix 1 to 6 or 8 to 12, wherein the attribute acquisition step includes a step of acquiring attribute information about a first user, and the conversion step includes a step of converting the call audio acquired in the voice acquisition step based on the attribute information about the first user acquired in the attribute acquisition step. This allows customers to communicate with users using a more appropriate voice based on user attribute information during calls between multiple users.
[0155] (Note 14) The program is one of the programs described in any of Appendix 1 to 6 or 8 to 13, wherein the attribute acquisition step includes a step of acquiring attribute information relating to a call, and the conversion step includes a step of converting the call audio acquired in the audio acquisition step based on the attribute information relating to the call acquired in the attribute acquisition step. This allows customers to communicate with other users using a more appropriate voice in calls between multiple users, based on attribute information related to the call.
[0156] (Note 15) The program described in any of the appendices 1 to 6 or 8 to 14, wherein the attribute acquisition step includes acquiring attribute information relating to a call, and the conversion step includes converting the call audio acquired in the voice acquisition step based on information relating to the emotions of the user or customer in the call acquired in the attribute acquisition step. This allows customers to communicate with users using a more appropriate voice based on the emotional information of the user or customer during calls between multiple users. For example, a customer can communicate with a user using a more appropriate voice based on the emotional state of the user or customer.
[0157] (Note 16) The call attributes acquired in the attribute acquisition step do not include information about the user's and customer's surrounding environment or the call environment, and are provided by any of the programs described in Appendix 1 to 15. This allows customers to communicate with users using a more appropriate voice based on the call attributes during calls between multiple users.
[0158] (Note 17) The program described in any of Appendix 1 to 16, wherein the conversion step includes a step of converting the voice component of a person from the call audio acquired in the voice acquisition step, but does not include a step of converting other voice components such as background noise, noise, and background sounds from the call audio acquired in the voice acquisition step. This allows customers to communicate with users using a more appropriate voice based on the call attributes during calls between multiple users.
[0159] (Note 18) The program is one of the programs described in any of Appendix 1 to 17, wherein the program causes the processor to perform a second selection step of selecting a second voice processing from among multiple voice processing methods based on a second selection instruction received from a first user; the voice acquisition step includes a step of acquiring second call audio from a second user; the conversion step includes a step of converting the second call audio acquired in the acquisition step by applying the second voice processing selected in the second selection step to the second call audio acquired in the acquisition step; and the output step includes a step of outputting the second call audio converted in the conversion step to the first user. This allows, for example, to conduct conversations with customers using a voice that is easier for the user to listen to from among multiple voice processing options. For example, it allows to conduct conversations with customers using a predetermined voice that provides greater comfort to the user from among multiple voice processing options. For example, it allows to conduct conversations with customers using a predetermined voice that the user prefers, reduces stress levels, or increases interest from among multiple voice processing options. For example, if a customer is angry, the user can reduce the psychological stress associated with interacting with the customer by selecting a voice processing option that reduces intonation or lowers the volume from among multiple voice processing options.
[0160] (Note 19) An information processing device comprising a processor and a memory unit, wherein the processor is made to perform a voice acquisition step (S104, S304) for acquiring call audio from a first user, a conversion step (S104, S304) for converting the call audio acquired in the voice acquisition step, an output step (S105, S305) for outputting the call audio converted in the conversion step to a second user, and an attribute acquisition step (S102, S302) for acquiring call attributes related to the call, wherein the conversion step includes a step of converting the call audio acquired in the voice acquisition step based on the call attributes acquired in the attribute acquisition step. This allows customers to communicate with users using a more appropriate voice based on the call attributes during calls between multiple users.
[0161] (Note 20) An information processing method performed by a computer comprising a processor and a memory unit, wherein the processor is instructed to perform a voice acquisition step (S104, S304) for acquiring call audio from a first user, a conversion step (S104, S304) for converting the call audio acquired in the voice acquisition step, an output step (S105, S305) for outputting the call audio converted in the conversion step to a second user, and an attribute acquisition step (S102, S302) for acquiring call attributes related to the call, wherein the conversion step includes a step of converting the call audio acquired in the voice acquisition step based on the call attributes acquired in the attribute acquisition step. This allows customers to communicate with users using a more appropriate voice based on the call attributes during calls between multiple users. [Explanation of Symbols]
[0162] 1 Information processing system, 10 Server, 101 Storage unit, 103 Control unit, 20A, 20B, 20C User terminals, 201 Storage unit, 204 Control unit, 30 CRM system, 301 Storage unit, 304 Control unit, 50A, 50B, 50C Customer terminals, 501 Storage unit, 504 Control unit
Claims
1. A program comprising a processor and a memory unit, which enables a computer to perform a call between a first user and a second user, The program is provided to the processor: A voice acquisition step to obtain call audio from the first user, A voice processing step for processing the call audio acquired in the voice acquisition step, An output step which outputs the voice-processed call audio from the voice processing step to a second user, A presentation step in which information indicating multiple voice processing methods is presented to the first user, A selection step that allows the first user to select one or more pieces of information indicating predetermined audio processing from among the pieces of information indicating the plurality of audio processing presented in the presentation step, Make it run, The aforementioned voice processing step is a step of processing the call audio by applying the call audio acquired in the voice acquisition step to a predetermined generation model identified based on information indicating the predetermined voice processing, The aforementioned presentation step is a step of presenting information indicating the types of quantitative evaluation indicators related to auditory perception of multiple listeners, The selection step is a step that can accept the selection of information indicating a predetermined type of evaluation indicator from among the information indicating the types of the multiple evaluation indicators, The voice processing step is a step of processing the call audio acquired in the voice acquisition step based on the predetermined generation model that optimizes the predetermined evaluation index. program.
2. The aforementioned presentation step is to present information indicating the type of at least one of the following evaluation indicators: SIIB, HASPI, ESTOI, PESQ, ViSQOL, reliability, credibility, comfort, pleasantness, preference, stress level, intimidation level, and interest level. The program according to claim 1.
3. A program comprising a processor and a memory unit, which enables a computer to perform a call between a first user and a second user, The program is provided to the processor: A voice acquisition step to obtain call audio from the first user, A voice processing step for processing the call audio acquired in the voice acquisition step, An output step which outputs the voice-processed call audio from the voice processing step to a second user, A presentation step in which information indicating multiple voice processing methods is presented to the first user, A selection step that allows the first user to select one or more pieces of information indicating predetermined audio processing from among the pieces of information indicating the plurality of audio processing presented in the presentation step, An attribute acquisition step to acquire call attributes relating to the aforementioned call, including user attributes, customer attributes, call category, and call type, Make it run, The aforementioned voice processing step is a step of processing the call audio by applying the call audio acquired in the voice acquisition step to a predetermined generation model identified based on information indicating the predetermined voice processing, The aforementioned selection step is, A step in which the first user provides instructions for selecting the type of quantitative evaluation index regarding the listener's auditory perception, A step of selecting a predetermined generation model from among the plurality of generation models based on the call attributes obtained in the attribute acquisition step and the received evaluation index, including, program.
4. The aforementioned selection step is, A step in which the first user provides instructions on the type of evaluation metric to be optimized, Based on the call attributes obtained in the attribute acquisition step and the received evaluation metrics, a step is to select a predetermined generation model from among a plurality of generation models that optimizes the evaluation metrics. including, The program according to claim 3.
5. The aforementioned selection step is, A step in which the first user provides instructions for selecting the type of evaluation metric, The steps include: selecting a predetermined generation model from among the plurality of generation models by applying the call attributes acquired in the attribute acquisition step and the received evaluation index to the voice processing model; including, The program according to claim 3 or 4.
6. The aforementioned voice acquisition step includes the step of acquiring second call audio from a second user, The aforementioned selection step cannot accept selection instructions from the second user. The aforementioned voice processing step is a step in which the second call audio acquired in the aforementioned voice acquisition step is not subjected to voice processing. The program according to any one of claims 3 to 5.
7. The attribute acquisition step includes the step of acquiring attribute information relating to a second user, The voice processing step includes processing the call audio acquired in the voice acquisition step based on the attribute information relating to the second user acquired in the attribute acquisition step. The program according to any one of claims 3 to 5.
8. The attribute acquisition step includes the step of acquiring attribute information relating to the first user, The voice processing step includes processing the call audio acquired in the voice acquisition step based on the attribute information relating to the first user acquired in the attribute acquisition step. The program according to any one of claims 3 to 6.
9. The attribute acquisition step includes the step of acquiring attribute information relating to the call, The voice processing step includes processing the call audio acquired in the voice acquisition step based on the attribute information relating to the call acquired in the attribute acquisition step. The program according to any one of claims 3 to 8.
10. The attribute acquisition step includes the step of acquiring attribute information relating to the call, The voice processing step includes processing the call audio acquired in the voice acquisition step based on information about the emotions of the user or customer in the call acquired in the attribute acquisition step. The program according to any one of claims 3 to 9.
11. The call attributes obtained in the attribute acquisition step do not include information about the user's and customer's surrounding environment or the call environment. The program according to any one of claims 3 to 10.
12. The aforementioned audio processing step is: The step includes processing the voice component of the person's voice from the call audio acquired in the voice acquisition step, The voice acquisition step does not include a step of processing audio components other than the voice components of the person, such as background noise, noise, and other sounds, from the acquired call audio. A program according to any one of claims 1 to 11.
13. The program is provided to the processor: A second selection step involves selecting a second generative model from among multiple speech processing methods based on a second selection instruction received from the first user, Make it run, The aforementioned voice acquisition step includes a second acquisition step of acquiring second call audio from a second user, The aforementioned audio processing step includes processing the second call audio by applying the second call audio acquired in the second acquisition step to the second generation model selected in the second selection step, The output step includes outputting the second call audio processed in the voice processing step to the first user. The program according to any one of claims 1 to 12.
14. A method to be performed on an information processing apparatus comprising a processor and a storage unit, wherein the processor performs all steps performed in the invention according to any one of claims 1 to 13.
15. An information processing apparatus comprising a processor and a storage unit, wherein the processor performs all steps performed in the invention according to any one of claims 1 to 13.
Citation Information
Patent Citations
Method and device for recognizing speech
JP2000330587A
Voice conversion apparatus
JP2010103704A
Relay system, relay method, and program
JP2014164241A
Voice conversion model learning device, voice conversion device, method, and program
JP2018136430A
Voice processing condition setting device, radio communication device and voice processing condition setting method
JP2020086099A