Information processing system, information processing method and program

The information processing system enhances user convenience by converting a user's voice into another person's voice through impression-based voice selection and processing adjustments, ensuring accurate and clear output.

JP2025078119APending Publication Date: 2025-05-20YAMAHA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023190460
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-08
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

Existing voice conversion technologies lack user convenience in terms of flexibility and accuracy in converting a user's voice into the voice of another person.

Method used

An information processing system that includes a processor to receive impression information, set a registered voice corresponding to the desired impression, and utilize a voice conversion model to output the user's voice as the voice of another person, with options for superimposing and signal processing to adjust delay and pronunciation clarity.

Benefits of technology

Improves user convenience by allowing accurate and flexible conversion of voices based on user preferences, reducing delays, and clarifying phonemes in the output voice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025078119000001_ABST
    Figure 2025078119000001_ABST
Patent Text Reader

Abstract

To provide an information processing system, method and program for converting a user's voice to that of another person different from the user, which can enhance user convenience.SOLUTION: An information processing system includes the steps of: receiving impression information of another person's voice to be converted from a user (user terminal) with an information processing device; setting a registered voice based on the input impression information; receiving input of the user's voice; inputting the user's voice into a voice conversion model corresponding to the set registered voice, and transmitting the voice of another person output from the voice conversion model; and outputting the voice of another person with the user terminal.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an information processing system, an information processing method, and a program. [Background technology]

[0002] As disclosed in the following document, a technique is disclosed for converting an inputted user's voice (audio data) into another person's voice. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-033260 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the above technology still has room for improvement in terms of user convenience.

[0005] In view of the above circumstances, the present invention provides a technique that can improve user convenience. [Means for solving the problem]

[0006] According to one aspect of the present invention, there is provided an information processing system for converting a user's voice into the voice of another person different from the user. The information processing system includes a processor capable of executing a program to perform the following steps: In the receiving step, impression information indicating an impression of the voice of the other person desired by the user is received. In the setting step, a registered voice corresponding to the impression information is set as the voice of the other person from among a plurality of registered voices prepared in advance.

[0007] According to the present disclosure, user convenience can be improved. [Brief description of the drawings]

[0008] [Figure 1] 1 is a configuration diagram illustrating an information processing system 1. [Diagram 2] 2 is a block diagram showing a hardware configuration of an information processing device 2. FIG. [Diagram 3] FIG. 2 is a block diagram showing a hardware configuration of a user terminal 3. [Figure 4] 2 is a block diagram showing functions realized by an information processing device 2 (processor 23). FIG. [Diagram 5] FIG. 2 is a block diagram showing an example of the configuration of a voice conversion model M1. [Figure 6] 11 is a diagram illustrating an example of the relationship between the allowable delay time (allowable delay amount) and the superimposition ratio. FIG. [Figure 7] FIG. 11 is a diagram illustrating an example of a relationship between pronunciation clarity and a superimposition ratio. [Figure 8] 11 is a diagram showing an example of a pronunciation start portion SP in a user's voice UV. FIG. [Figure 9] FIG. 2 is an activity diagram showing the flow of a first information processing method. [Figure 10] 13 is a diagram showing an example of an impression information input screen IS displayed on a user terminal 3. FIG. [Figure 11] FIG. 11 is an activity diagram showing the flow of a second information processing method. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, embodiments of the present invention will be described with reference to the drawings. Various characteristic features shown in the following embodiments can be combined with each other.

[0010] Incidentally, the program for realizing the software appearing in this embodiment may be provided as a non-transitory computer-readable recording medium, or may be provided so as to be downloadable from an external server, or may be provided so that the program is started on an external computer and its functions are realized on a client terminal (so-called cloud computing).

[0011] In addition, in this embodiment, the term "unit" may include, for example, a combination of hardware resources implemented by a circuit in the broad sense and software information processing that can be specifically realized by these hardware resources. In addition, in this embodiment, various information is handled, and this information is represented, for example, by physical values ​​of signal values ​​representing voltage and current, high and low signal values ​​as a binary bit collection consisting of 0 or 1, or quantum superposition (so-called quantum bits), and communication and calculation can be performed on the circuit in the broad sense.

[0012] In addition, a circuit in the broad sense is a circuit realized by at least appropriately combining a circuit, circuitry, a processor, a memory, etc. In other words, it includes an application specific integrated circuit (ASIC), a programmable logic device (e.g., a simple programmable logic device (SPLD), a complex programmable logic device (CPLD), and a field programmable gate array (FPGA)), etc.

[0013] 1. Hardware Configuration This section describes the hardware configuration.

[0014] <Information Processing System 1> FIG. 1 is a configuration diagram showing an information processing system 1. The information processing system 1 includes an information processing device 2 and a user terminal 3. The information processing device 2 and the user terminal 3 are configured to be able to communicate with each other via a telecommunications line. In one embodiment, the information processing system 1 is made up of one or more devices or components. For example, if the information processing system 1 is made up of only the information processing device 2, the information processing system 1 can be the information processing device 2. These components will be described below.

[0015] <Information processing device 2> 2 is a block diagram showing a hardware configuration of the information processing device 2. The information processing device 2 includes a communication bus 20, a communication unit 21, a storage unit 22, and a processor 23. The communication unit 21, the storage unit 22, and the processor 23 are electrically connected via the communication bus 20 inside the information processing device 2.

[0016] <Communications Division 21> The communication unit 21 is preferably a wired communication means such as USB, IEEE1394, Thunderbolt (registered trademark), wired LAN network communication, etc., but may also include wireless LAN network communication, mobile communication such as 3G / LTE / 5G, BLUETOOTH (registered trademark) communication, etc. as necessary. In other words, it is more preferable to implement it as a collection of multiple communication means. In other words, the information processing device 2 may communicate various information from the outside via the communication unit 21 and the network.

[0017] <Storage section 22> The storage unit 22 stores various information defined by the above description. This can be implemented, for example, as a storage device such as a solid state drive (SSD) that stores various programs and the like related to the information processing device 2 executed by the processor 23, or as a memory such as a random access memory (RAM) that stores temporarily required information (arguments, arrays, etc.) related to the program calculations. The storage unit 22 stores various programs, variables, etc. related to the information processing device 2 executed by the processor 23.

[0018] <Processor 23> The processor 23 performs processing and control of the overall operation related to the information processing device 2. The processor 23 is, for example, a central processing unit (CPU). The processor 23 realizes various functions related to the information processing device 2 by reading out a predetermined program stored in the storage unit 22. That is, information processing by software stored in the storage unit 22 can be specifically realized by the processor 23, which is an example of hardware, and executed as each functional unit included in the processor 23. That is, the processor 23 can execute a program so that each functional unit is executed. These will be described in more detail in the next section. The processor 23 is not limited to being single, and may be implemented to have multiple processors 23 for each function. Also, a combination of these may be used.

[0019] <User terminal 3> 3 is a block diagram showing a hardware configuration of the user terminal 3. The user terminal 3 includes a communication bus 30, a communication unit 31, a storage unit 32, a processor 33, a display unit 34, and an input unit 35. The communication unit 31, the storage unit 32, the processor 33, the display unit 34, and the input unit 35 are electrically connected via the communication bus 30 inside the user terminal 3. Descriptions of the communication unit 31, the storage unit 32, and the processor 33 are omitted because they are similar to the descriptions of the respective units in the information processing device 2.

[0020] <Display section 34> The display unit 34 displays a screen of a graphical user interface (GUI) that can be operated by the user. The display unit 34 may be included in the housing of the user terminal 3, or may be attached externally. Specifically, the display unit 34 may be implemented as a display device such as a CRT display, a liquid crystal display, an organic EL display, or a plasma display. It is preferable that these display devices are implemented by selectively using them according to the type of the user terminal 3.

[0021] <Input section 35> The input unit 35 accepts an operation input made by a user. The operation input is transferred as a command signal to the processor 33 via the communication bus 30. The processor 33 may execute a predetermined control or calculation based on the transferred command signal as necessary. The input unit 35 may be included in the housing of the user terminal 3, or may be externally attached. For example, the input unit 35 may be implemented as a touch panel integrated with the display unit 34. When the input unit 35 is implemented as a touch panel, the user can input a tap operation, a swipe operation, or the like to the input unit 35. As the input unit 35, a switch button, a mouse, a QWERTY keyboard, or the like can be adopted instead of a touch panel.

[0022] The information processing device 2 may be an on-premise type or a cloud type. As the information processing device 2 in the cloud type, the above-mentioned functions and processes may be provided in the form of, for example, SaaS (Software as a Service) or cloud computing.

[0023] The user terminal 3 may be a general-purpose computer as shown in Fig. 1A, or may be configured as a combination of an audio interface 3A and a computer 3B connected to the audio interface 3A as shown in Fig. 1B. That is, the information processing system 1 may include at least one of the computer 3B and the audio interface 3A. With such a configuration, the computer 3B or the audio interface 3A can be provided as a component of the information processing system 1. The audio interface 3A may have a function of transmitting audio data to the computer 3B and receiving audio data from the computer 3B, as well as a function of distributing audio via a network.

[0024] 2. Functional configuration In this section, the functional configuration of this embodiment will be described. Information processing by software stored in the storage unit 22 is specifically realized by the processor 23, which is an example of hardware, and can be executed as each functional unit included in the processor 23. The information processing system 1 functions as a system that converts the voice of a user into the voice of a person different from the user.

[0025] 4 is a block diagram showing functions realized by the information processing device 2 (processor 23). Specifically, the information processing device 2 (processor 23) includes a reception unit 231, a setting unit 232, an output control unit 233, a superimposition unit 234, and a signal processing unit 235.

[0026] The information processing device 2 is configured to realize a first conversion function, a second conversion function, and a function that combines these functions. The first conversion function is a function that provides an interface for a user to select another person's voice. The second conversion function is a function that adjusts the output of the converted voice. Below, each unit of the information processing device 2 will be described separately for the first conversion function and the second conversion function.

[0027] <First conversion function> In the first conversion function, mainly the reception unit 231, the setting unit 232, and the output control unit 233 in FIG. 4 are used.

[0028] <Reception Department 231> The receiving unit 231 is configured to acquire various information. Specifically, the receiving unit 231 receives an input of impression information indicating an impression of the voice of another person desired by the user and an input of the user's own voice from the user terminal 3. The impression information is input from the input unit 35 of the user terminal 3, for example, via a UI displayed on the user terminal 3. The user's own voice is input from a microphone included in the user terminal 3 (or connected to the user terminal 3), for example.

[0029] "Impression information" is not quantitative information represented by parameters such as numerical values, but qualitative information represented by words or sentences (i.e., natural language) that are close to the impression of the voice that the user imagines to convert. Impression information includes words that modify the impression of the voice, such as "kind," "cute," "handsome," "mature," "clear voice," and "strong voice." Impression information may be composed of a combination of multiple words, such as "kind" + "handsome," or may be composed of a sentence containing multiple words, such as "a kind and tall person."

[0030] When the impression information is composed of a combination of multiple words or a sentence including multiple words, the reception unit 231 may present the first option (multiple selectable words) to the user and have the user select a word, and then sequentially present the next option to the user and have the user select a word, repeating this process to acquire the selected multiple words as elements of the impression information. In this case, the reception unit 231 may set a word included in the next option based on the previously selected word. For example, when "kind" is selected as the first option, the reception unit 231 may display a word expressing the degree of "kindness" as the next option, or display a word that can be combined with (co-occurs with) "kind" as the next option.

[0031] Furthermore, the impression information may be composed of a combination of a word expressing an impression and the degree of the impression. Examples of such impression information include "cute: 5", "cute: 5, handsome: 3", etc. The numerical value attached to the word indicates the magnitude of the impression expressed by the word. The magnitude of the impression can be input by the user on the user terminal 3, for example, using a UI such as a slide bar.

[0032] Furthermore, the receiving unit 231 may display a UI including a free entry field on the user terminal 3 and receive input of an arbitrary sentence (for example, "a kind and tall person") from the user. The receiving unit 231 receives the sentence itself input by the user as impression information. In this case, it is preferable to set an upper limit on the number of characters that can be input in order to avoid receiving an excessive number of words.

[0033] The receiving unit 231 may receive designation information that designates the voice of another person other than the impression information. For example, the receiving unit 231 may receive input of speaker information (corresponding to designation information) instead of impression information. The speaker information is information that directly designates a registered voice, which will be described later. In other words, the speaker information includes a numerical value (or a combination of symbols such as alphabets that can be converted into a numerical value) for the user to designate a desired registered voice. When receiving input of speaker information, the receiving unit 231 allows the user to select the speaker information using an exclusive selection type UI (e.g., radio buttons).

[0034] The receiving unit 231 may further receive the pitch of the user's voice. For example, the receiving unit 231 receives an input of the user's own voice from the user terminal 3, and acquires the pitch (fundamental frequency) of the user's voice from the user's own voice using an F0 (fundamental frequency) estimator. When acquiring the pitch from the user's own voice in this way, the receiving unit 231 presents a prepared sentence for pitch acquisition to the user, has the user read the sentence for pitch acquisition aloud, and records the pronunciation. The receiving unit 231 may also receive an input (designation) of a numerical value representing the pitch from the user.

[0035] <Settings section> The setting unit 232 is configured to set a registered voice corresponding to the impression information as the voice of another person from among a plurality of registered voices prepared in advance. With such a configuration, the user's voice can be converted to a voice close to the user's image (impression). In other words, the user can select a registered voice based on their intuition.

[0036] Here, the "registered voice" refers to a voice for which reference information (for example, a voice conversion model which is a trained model described later) capable of converting the user's voice into another person's voice in real time is recorded in the information processing system 1 (specifically, the storage unit 22) in the output control unit 233 described later. Each of the multiple registered voices has an ID (identification label). Each registered voice may be assigned multiple IDs (for example, an ID for impression information and an ID for speaker information) so that the registered voice can be set based on both impression information and information other than impression information (speaker information). In addition, each of the multiple registered voices may be set to a pitch (average pitch).

[0037] The setting unit 232 sets a registered voice with an ID that is the same as or similar to the numerical information derived from the impression information as the voice of another person. Such a configuration makes it easy to register and manage registered voices using IDs. If there is a registered voice with the same ID as the numerical information derived from the impression information, the setting unit 232 sets this registered voice as the voice of another person. On the other hand, if there is no registered voice with the same ID as the numerical information derived from the impression information, the setting unit 232 sets, from among the registered voice IDs, a registered voice with an ID closest to the numerical information as the voice of another person. Note that the ID of the registered voice referred to here is an ID for impression information.

[0038] As a method of deriving the numerical information to be matched with the ID of the registered voice, for example, the setting unit 232 obtains the numerical information by vectorizing the impression information by natural language processing (i.e., embedding in a vector space). With such a configuration, it is possible to select a registered voice close to the image of the user based on the input of qualitative impression information. Specifically, the setting unit 232 inputs words or sentences included in the impression information to a vector converter (typically a classifier that is a trained model) using an algorithm such as the TF-IDF method, the LSI method, the LDA method, Word2vec, or BERT, and outputs a corresponding vector.

[0039] When the impression information is a combination of words, the setting unit 232 inputs each word into a vector converter, and calculates (for example adds) the vectors of each word to obtain the numerical information derived from the impression information. When the impression information is a sentence, the setting unit 232 inputs the sentence into a vector converter configured to vectorize the sentence, and obtains the vectorized sentence to obtain the numerical information derived from the impression information. When vectorizing each word, the vector into which one word is converted is constant as long as the same vector converter is used. On the other hand, when vectorizing a sentence, the vector into which one word is converted changes depending on the context, even if the same vector converter is used.

[0040] Furthermore, when the impression information is a sentence, the setting unit 232 may use a morphological analysis engine such as MeCab to decompose the sentence into one or more words. At this time, characters that cannot be extracted as words are excluded from the vector conversion. The setting unit 232 inputs each decomposed word into a vector converter, and calculates the vector of each word to obtain numerical information derived from the impression information.

[0041] As another method of deriving the numerical information to be matched with the ID of the registered voice, for example, the setting unit 232 may derive the numerical information by assigning numerical values ​​to words input by the user in advance. Specifically, the setting unit 232 assigns vectors such as {1,0,0,0} to "cute" in the impression information and {0,1,0,0} to "mature" in advance. In this case, the setting unit 232 acquires the numerical information by referring to a table in which the relationship between the words and vectors included in the impression information is described (words and vectors are linked one-to-one) without natural language processing the impression information (i.e., without using a vector converter). Note that when multiple words are included in the impression information, the vectors of each word are calculated (for example, added) as in the case of natural language processing, and the numerical information derived from the impression information is the result.

[0042] The setting unit 232 sets a registered voice with an ID that is the same as or similar to the designated information (speaker information) other than the impression information as the voice of another person. That is, the setting unit 232 sets a registered voice based on the speaker information in addition to setting a registered voice based on the impression information. With this configuration, it is possible to provide the user with a variety of methods for designating a voice to be converted. That is, the user can set a registered voice by inputting impression information, or can directly designate a registered voice by inputting speaker information. Note that the ID of the registered voice referred to here is an ID for speaker information.

[0043] The setting unit 232 may convert the impression information into speaker information by referring to the reference information stored in the storage unit 22, and set the registered voice based on the speaker information. In this case, the speaker information converted from the impression information corresponds to "numerical information derived from the impression information", and the registered voice is set by matching the speaker information with an ID for the speaker information.

[0044] As the reference information for converting the impression information into the speaker information, for example, a first information conversion model that has been machine-learned to learn the correspondence between the impression information and the speaker information is used. The first information conversion model is a trained model that has been machine-learned to receive the impression information and the acoustic data (voice for input) as input and to output the speaker information. The first information conversion model is trained by teacher data in which the impression information and the acoustic data are labeled with the speaker information corresponding to the combination of these. The setting unit 232 inputs the impression information and the user's own voice received by the receiving unit 231 into the first information conversion model, and sets the registered voice with an ID that is the same as or similar to the speaker information output from the first information conversion model as the voice of another person. In addition, the reference information for converting the impression information into the speaker information may be a table that describes the relationship between the impression information and the speaker information, a conversion formula that converts numerical information derived from the impression information into the speaker information, or the like.

[0045] Furthermore, the setting unit 232 may convert the speaker information into impression information by referring to the reference information. For example, a second information conversion model that has learned the correspondence between the speaker information and the impression information through machine learning is used as the reference information for converting the speaker information into the impression information. The second information conversion model is a trained model that has been machine-learned to receive the speaker information and the acoustic data (voice for input) as input and to output the impression information. The second information conversion model is trained by teacher data in which the speaker information and the acoustic data are labeled with the impression information corresponding to the combination of these. The setting unit 232 inputs the speaker information and the user's own voice received by the receiving unit 231 into the second information conversion model, and sets the registered voice using the impression information output from the second information conversion model.

[0046] The setting unit 232 may set, from among a plurality of registered voices, a registered voice corresponding to the impression information and the pitch of the user's voice as the voice of another person. With this configuration, the user's voice can be converted into a voice of another person suitable for the pitch of the user's voice. For example, the setting unit 232 may correct the fundamental frequency of the registered voice corresponding to the impression information to match the pitch of the user's voice. Specifically, as will be described later, the setting unit 232 sets the registered voice with the adjusted pitch by instructing the output control unit 233 to input the amount of pitch correction into the voice conversion model as a variable.

[0047] <Output control unit 233> The output control unit 233 is configured to output various sounds including the user's voice to an output device such as a speaker. Specifically, the output control unit 233 inputs the user's voice and the ID of the registered voice set as the voice of another person into a voice conversion model, thereby outputting the user's voice converted into the registered voice. The voice conversion model is a model that machine-learns the relationship between the input voice and the vocalization by the registered voice for each ID. With this configuration, the input voice can be accurately converted into the voice of another person based on the user's impression information. The output control unit 233 outputs the sound to, for example, a sound output device such as a speaker that the user terminal 3 has (or is connected to the user terminal 3), a sound output device connected to the user terminal 3 through a network, etc.

[0048] <Voice conversion model> The voice conversion model is stored in, for example, the storage unit 22. The voice conversion model is a trained model that has been trained in advance by machine learning. FIG. 5 is a block diagram showing an example of the configuration of the voice conversion model M1. The voice conversion model M1 receives acoustic data D1 as input and outputs synthetic acoustic data D3. The voice conversion model M1 includes an analysis unit M11, an encoder M12, a decoder M16, and a vocoder M17.

[0049] The acoustic data D1 includes waveform data of voice (including singing), waveform data of musical instrument sounds, etc. The user's voice input to the user terminal 3 is input to the voice conversion model M1 as the acoustic data D1. The synthetic acoustic data D3 is waveform data of a voice obtained by converting the acoustic data D1 (user's voice) into a specified registered voice (a different person's voice).

[0050] The analysis unit M11 converts the user's voice (sound data D1) into sound feature data AF by frequency analysis of the user's voice and acquires the fundamental frequency (pitch) of the user's voice. The sound feature data AF is a frequency spectrum of the sound represented by the sound data D1, for example, a Mel-Scale Log-Spectrum (MSLS).

[0051] The encoder M12 is a trained model configured to receive user voice data (acoustic feature data AF) and output first intermediate features MF1. The first intermediate features MF1 are intermediate data for causing the decoder M16 to output synthetic acoustic feature data AFS. A method for training the encoder M12 will be described later.

[0052] The decoder M16 is a trained model configured to receive the first intermediate feature MF1 and output the data of the registered voice (synthetic acoustic feature data AFS) into which the user's voice is converted. The synthetic acoustic feature data AFS is a frequency spectrum generated based on the intermediate feature, and is, for example, a Mel-frequency logarithmic spectrum.

[0053] The decoder M16 may receive the second intermediate feature MF2 as input. That is, the decoder M16 may be configured to receive the second intermediate feature MF2 instead of the first intermediate feature MF1 or together with the first intermediate feature MF1, and output the synthetic acoustic feature data AFS. The second intermediate feature MF2 is any information, and may be, for example, a feature generated by an encoder other than the encoder M12 for causing the decoder M16 to output the synthetic acoustic feature data AFS, or data in which language information is converted into numerical information by vectorization.

[0054] The vocoder M17 generates synthetic acoustic data D3 based on the synthetic acoustic feature data AFS output by the decoder M16. Specifically, the vocoder M17 generates synthetic acoustic data D3 by restoring the synthetic acoustic feature data AFS to an acoustic signal.

[0055] <Training method for voice conversion model M1> In training the voice conversion model M1, the encoder M12 and the decoder M16 are each trained using a machine learning device. The machine learning device may be incorporated in the information processing system 1 (e.g., may be a part of the information processing device 2), or may be an external device to the information processing system 1. For example, a convolutional neural network (CNN), a recurrent neural network (RNN), or a combination thereof may be used as a training method. An autoregressive model, an attention model, or the like may be used as a training model.

[0056] Training is performed for each registered voice ID. First, multiple teacher datasets in which training acoustic data is linked to training synthetic acoustic feature data that is the correct answer for each registered voice ID are stored in the machine learning device. The training acoustic data included in the teacher dataset includes the voice (including singing) of the owner of the registered voice (the person who provides his / her voice as the registered voice). Furthermore, in one training step, training acoustic data equivalent to a certain number of frames in frequency analysis is used. The training step is repeated until all frames of the teacher dataset have been referenced.

[0057] After preparing the teacher data set, the machine learning device causes the analysis unit M11 to generate acoustic feature data AF from the training acoustic data. Next, the machine learning device inputs the acoustic feature data AF to the training target encoder M12, which generates first intermediate features MF1.

[0058] The machine learning device inputs the first intermediate feature MF1 generated by the encoder M12 together with the ID of the registered voice to the decoder M16 to be trained, and generates synthetic acoustic feature data AFS. The machine learning device trains the encoder M12 and the decoder M16 so that the synthetic acoustic feature data AFS output by the decoder M16 approaches the training synthetic acoustic feature data.

[0059] Specifically, the machine learning device performs back propagation and updates the variables constituting the encoder M12 and the variables constituting the decoder M16 so as to reduce the difference between the synthetic acoustic feature data AFS generated by the decoder M16 and the training synthetic acoustic feature data included in the teacher dataset. By repeating this training step, a voice conversion model M1 corresponding to the ID of a registered voice (i.e., one voice provider) is obtained. In other words, a voice conversion model M1 capable of synthesizing a voice with the sound quality of the registered voice corresponding to the ID is obtained in response to the specification of the ID of the registered voice.

[0060] Furthermore, the machine learning device sets the pitch of the voice conversion model from the training acoustic data used in training. Specifically, the machine learning device stores the average value of the frequency of the data excluding outliers from the training acoustic data as the pitch of the voice conversion model together with the ID, etc.

[0061] After the above-mentioned basic training, the machine learning device may further perform auxiliary training on the decoder M16. In the auxiliary training, only the decoder M16 of the voice conversion model M1 is trained. The auxiliary training is performed for each registered voice ID. In the auxiliary training, first, a plurality of teacher data sets in which training acoustic data is linked to training synthetic acoustic feature data that is the correct answer are stored in the machine learning device. In one training step, training acoustic data equivalent to a certain number of frames in the frequency analysis is used. The training step is repeated until all frames of the teacher data set are referenced.

[0062] After preparing the teacher data set, the machine learning device causes the analysis unit M11 to generate acoustic feature data AF from the training acoustic data. Next, the machine learning device inputs the acoustic feature data AF to a trained encoder M12 to generate first intermediate features MF1. Furthermore, the machine learning device inputs the first intermediate features MF1 generated by the encoder M12 to a decoder M16 to be trained to generate synthetic acoustic feature data AFS.

[0063] The machine learning device trains the decoder M16 so that the synthetic acoustic feature data AFS output by the decoder M16 approaches the training synthetic acoustic feature data. In this way, in the auxiliary training, the decoder M16 can be further trained using only the acoustic feature data as input.

[0064] The trained voice conversion model M1 is assigned an ID of the registered voice (e.g., an ID for impression information, an ID for speaker information, etc.). The voice conversion model M1 includes a decoder M16 labeled with at least the ID of the registered voice (i.e., trained in association with the ID of the registered voice).

[0065] The encoder M12 included in the voice conversion model M1 may be trained in association with the ID of the enrolled voice, or may be trained without being associated with the ID of the enrolled voice. When the encoder M12 is associated with the ID of the enrolled voice, the encoder M12 becomes a dedicated encoder for the associated ID and is used only in combination with the decoder M16 having the same ID. On the other hand, when the encoder M12 is not associated with the ID of the enrolled voice, the trained encoder M12 is shared among multiple enrolled voices (multiple decoders M16) with different IDs.

[0066] As described above, when training the voice conversion model M1, the ID of the registered voice is input together with the teacher data set. At this time, by inputting two or more types of IDs (for example, an ID for impression information and an ID for speaker information) that are assigned to the same registered voice while randomly switching between them to the voice conversion model M1 and training it, it is possible to obtain a voice conversion model M1 to which two or more types of IDs are assigned.

[0067] When the ID of the registered voice is labeled in the encoder M12, the output control unit 233 selects the encoder M12 to input the user's voice and the decoder M16 to input the intermediate features based on the ID of the registered voice. With this configuration, both the encoder M12 and the decoder M16 prepared for each registered voice are selected based on the impression information input by the user, thereby improving the accuracy of conversion to a different person's voice.

[0068] The machine learning device may perform encoder assisted training on the encoder M12. In the encoder assisted training, only the encoder M12 of the voice conversion model M1 is trained. The encoder assisted training is also performed for each registered voice ID. The teacher data set used in the encoder assisted training is the same as that used in the assisted training of the decoder M16, and is obtained by linking the training acoustic data with the training synthetic acoustic feature data that is the correct answer.

[0069] After preparing the teacher data set, the machine learning device causes the analysis unit M11 to generate acoustic feature data AF from the training acoustic data. Next, the machine learning device inputs this acoustic feature data AF to the encoder M12 to be trained, and generates a first intermediate feature MF1. Furthermore, the machine learning device inputs the first intermediate feature MF1 generated by the encoder M12 to a trained decoder M16, and generates synthetic acoustic feature data AFS. The machine learning device trains the encoder M12 so that the synthetic acoustic feature data AFS output by the decoder M16 approaches the training synthetic acoustic feature data.

[0070] The output control unit 233 may input the difference (e.g., 120 Hz) between the pitch of the user accepted by the accepting unit 231 (e.g., 440 Hz) and the pitch of the registered voice of the specified ID (e.g., 680 Hz) as the pitch correction amount to the voice conversion model M1. The voice conversion model M1 outputs synthetic acoustic data D3 whose fundamental frequency is the frequency obtained by adding the pitch correction amount to the fundamental frequency of the input acoustic data D1 (user's voice).

[0071] <Second conversion function> In the second conversion function, mainly, in addition to the reception unit 231, the setting unit 232, and the output control unit 233 in Fig. 4, the superimposition unit 234 and the signal processing unit 235 are further used. The second conversion function is based on the premise that a voice of another person (registered voice) to convert the user's voice is set. The registered voice may be set by the above-mentioned first conversion function, or may be set by a method or means other than the first conversion function.

[0072] <Reception Department 231> The receiving unit 231 receives designation information of the voice of another person desired by the user and the pitch of the user's voice, similar to the first conversion function. The "designation information" in the second conversion function may be impression information or information other than impression information (e.g., speaker information). The method of receiving the pitch of the user's voice is similar to the first conversion function.

[0073] In the second conversion function, the reception unit 231 receives an input from the user terminal 3 indicating an allowable delay time for the voice (voice converted into another person's voice) to be output by the output control unit 233 in response to the input of the user's voice. The "delay time" refers to the time lag from when the user inputs their voice into the information processing system 1 to when the voice of another person into which the user's voice has been converted is output from the audio output device via the output control unit 233.

[0074] The receiving unit 231 may accept the allowable delay time as an arbitrary numerical value (e.g., any time between 20 ms and 100 ms), or may accept the allowable delay time as a label (option) indicating the degree of the allowable delay time, such as "large," "medium," and "small."

[0075] In the second conversion function, the receiving unit 231 receives the pronunciation clarity of the other person's voice as an input from the user terminal 3. The "pronunciation clarity" is the degree of improvement of the accent at the beginning of pronunciation in the other person's voice.

[0076] The receiving unit 231 may receive the pronunciation clarity at any position on the slider, or may receive the pronunciation clarity as a label (option) indicating the degree of pronunciation clarity, such as "good," "medium," and "poor."

[0077] The delay time tolerance and pronunciation clarity are parameters for determining the superimposition ratio, which will be described later. The receiving unit 231 may also receive an input of the superimposition ratio itself. In other words, the receiving unit 231 may directly receive the superimposition ratio without using the delay time tolerance.

[0078] <Setting section 232> As in the first conversion function, the setting unit 232 sets a registered voice corresponding to the designated information and the pitch of the user's voice as the voice of another person from among a plurality of registered voices prepared in advance. With this configuration, the second conversion function can also convert the user's voice into a voice of another person that is suitable for the user's pitch.

[0079] <Superimposition unit 234> The superimposing unit 234 is configured to determine the superimposing ratio of the user's voice that is not input to the voice conversion model M1 to the voice of another person that is converted from the user's voice by the voice conversion model M1. That is, the superimposing unit 234 determines the ratio (superimposing ratio) when simultaneously outputting the first voice (the converted voice of another person) that has a certain large output delay due to the conversion process of the voice conversion model M1 and the second voice (the user's own voice, or the voice obtained by signal processing the user's own voice) that has not been subjected to the conversion process of the voice conversion model M1 and has a small output delay. For example, when the superimposing ratio is 50%, the first voice and the second voice are output at a ratio of 1:1. When the superimposing ratio is 0%, only the first voice is output. When the superimposing ratio is 100%, only the second voice is output, and in this case, the first voice is replaced with the second voice.

[0080] The superimposing unit 234 has a first superimposing function for suppressing delay in the output voice caused by the processing time of the voice conversion model, and a second superimposing function for clarifying the phonemes of the output voice.

[0081] <First overlay function> In the first superimposition function, the superimposition unit 234 determines the superimposition ratio based on the allowable delay time input by the user (accepted by the accepting unit 231). With this configuration, the superimposition ratio can be set according to the allowable delay amount preferred by the user. Specifically, the superimposition unit 234 increases the superimposition ratio as the allowable delay time input by the user is smaller. In other words, when the allowable delay time is small (i.e., the allowable delay time is short), the superimposition unit 234 increases the superimposition ratio so that the ratio of the second voice increases, and when the allowable delay time is large (i.e., the allowable delay time is long), the superimposition unit 234 decreases the superimposition ratio so that the ratio of the first voice increases. With this configuration, the ratio of the voice converted by the voice conversion model can be adjusted according to the allowable amount.

[0082] The superimposing unit 234 converts the allowable amount of delay time into a superimposing ratio based on a previously prepared judgment formula. The judgment formula may be a function that converts the allowable amount into a superimposing ratio linearly or nonlinearly, or may be a conditional formula that determines the superimposing ratio for each range of the allowable amount (for example, when the allowable amount is 20 ms or more and less than 30 ms, the superimposing ratio is 100%, and when the allowable amount is 30 ms or more and 70 ms or less, the superimposing ratio is 50%). In addition, when the accepting unit 231 accepts the allowable amount by a label (for example, "large", "medium", "small", etc.), the superimposing unit 234 may determine the superimposing ratio by using a table in which the superimposing ratio for each label is described.

[0083] The second voice, which has not been subjected to the conversion process of the voice conversion model M1, includes the user's own voice (hereinafter also referred to as "raw voice") and a voice obtained by signal processing the raw voice by the signal processing unit 235 (hereinafter also referred to as "signal processed voice"). When the raw voice is superimposed on the first voice, the superimposition unit 234 determines the superimposition ratio of the user's own voice (raw voice) to the voice of another person (first voice) as the superimposition ratio. When the signal processed voice is superimposed on the first voice, the superimposition unit 234 determines the superimposition ratio of the user's voice converted by signal processing (signal processed voice) to the voice of another person (first voice) as the superimposition ratio.

[0084] The reception unit 231 receives, for example, an input from the user terminal 3 as to whether to superimpose the raw voice or the signal-processed voice as the second voice. The relationship between the allowable delay time and the superimposition ratio when the signal-processed voice is used as the second voice may be different from the relationship between the allowable delay time and the superimposition ratio when the raw voice is used as the second voice.

[0085] FIG. 6 is a diagram showing an example of the relationship between the delay time tolerance (delay tolerance) and the superimposition ratio. In the example of the superimposition ratio when a signal-processed voice is used as the second voice in FIG. 6A, the superimposition ratio is set in three stages according to the delay tolerance, in order from the largest delay tolerance: a first stage (only the first voice) with a superimposition ratio of 0%, a second stage (cross-fade between the first voice and the signal-processed voice) with a superimposition ratio of 0-100%, and a third stage (only the signal-processed voice) with a superimposition ratio of 100%. In the second stage, for example, the superimposition ratio increases linearly as the delay tolerance becomes smaller. Note that the maximum value of the delay tolerance is set to the maximum delay time of the first voice (for example, 100 ms), and the minimum value is set to the minimum delay time of the second voice (signal-processed voice) (for example, 20 ms). Also, the minimum delay time of the first voice is, for example, 50 ms, and the corresponding delay tolerance is included in the second stage.

[0086] In the example of the superimposition ratio when a natural voice is used as the second voice in Fig. 6B, the superimposition ratio is set in three stages, as in the case of using a signal-processed voice. However, when a natural voice is used, the range of the third stage (only natural voice) is narrower (i.e., the upper limit is smaller) than when a signal-processed voice is used. In other words, when a natural voice is used, the range of the second stage (cross-fade) is wider (i.e., the lower limit is smaller) than when a signal-processed voice is used.

[0087] The superimposing unit 234 may superimpose both the live voice and the signal-processed voice on the first voice. In this case, the superimposing unit 234 may determine the respective ratios of the live voice and the signal-processed voice in the second voice, in addition to the superimposing ratio of the second voice on the first voice. That is, the superimposing unit 234 may determine the superimposing ratio of the live voice on the first voice (live voice superimposing ratio) and the superimposing ratio of the signal-processed voice on the first voice (signal-processed voice superimposing ratio). The maximum value of the sum of the live voice superimposing ratio and the signal-processed voice superimposing ratio is 100%. The live voice superimposing ratio and the signal-processed voice superimposing ratio may be input from the user terminal 3. The superimposing unit 234 may also determine the live voice superimposing ratio and the signal-processed voice superimposing ratio according to the allowable delay time.

[0088] <Second overlay function> In the second superimposition function, the superimposition unit 234 determines a superimposition ratio such that the voice of another person (first voice) is superimposed on the voice of a user not input to the voice conversion model (second voice) at the pronunciation start portion of the input user's voice. With this configuration, it is possible to clarify the phonemes at the pronunciation start portion of the converted voice. Specifically, the superimposition unit 234 determines a superimposition ratio such that the voice of another person (first voice) is superimposed on the voice of the user converted by signal processing (signal processing voice) at the pronunciation start portion of the user's voice. The superimposition unit 234 may also determine a superimposition ratio such that the first voice is superimposed on the live voice at the pronunciation start portion of the user's voice. Furthermore, the superimposition unit 234 may superimpose both the live voice and the signal processing voice on the first voice at the pronunciation start portion of the user's voice.

[0089] The superimposing unit 234 determines the superimposing ratio based on the pronunciation clarity input by the user (accepted by the accepting unit 231). With this configuration, it is possible to set the superimposing ratio according to the pronunciation clarity preferred by the user. Specifically, the higher (better) the pronunciation clarity input by the user, the larger the superimposing ratio is. That is, when the pronunciation clarity is high (i.e., the pronunciation is desired to be clear), the superimposing unit 234 increases the superimposing ratio so that the ratio of the second voice increases, and when the pronunciation clarity is low (i.e., the pronunciation does not need to be clear), the superimposing unit 234 decreases the superimposing ratio so that the ratio of the first voice increases. With this configuration, it is possible to adjust the ratio of the voice converted by the voice conversion model according to the pronunciation clarity.

[0090] Fig. 7 is a diagram showing an example of the relationship between pronunciation clarity and the overlapping ratio. In the example of Fig. 7, the overlapping ratio is set in three stages according to the pronunciation clarity, in order from the worst pronunciation clarity: a first stage (first voice only) with an overlapping ratio of 0%, a second stage (cross-fade between the first voice and the second voice) with an overlapping ratio of 0-100%, and a third stage (second voice only) with an overlapping ratio of 100%. In the second stage, for example, as the pronunciation clarity improves, the overlapping ratio increases linearly.

[0091] The superimposing unit 234 determines the moment when the speech content of the input user's voice switches from unvoiced to voiced as the start point of the "pronunciation start portion." The superimposing unit 234 determines the range from this start point to a preset unit time number as the "pronunciation start portion," and sets this as the target portion for superimposing processing in the first voice.

[0092] FIG. 8 is a diagram showing an example of the pronunciation start portion SP in the user's voice UV. The superimposing unit 234 determines the pronunciation start portion SP by, for example, VAD (Voice Activity Detection) for the user's voice UV. As an algorithm for the VAD, one that can determine the pronunciation start portion SP within the conversion processing time of the user's voice UV by the voice conversion model is preferable, and for example, a volume switch type VAD is suitable. The superimposing unit 234 may extract an unvoiced part and a voiced part from the result of the F0 estimator to which the user's voice UV is input, and determine the pronunciation start portion SP. In parallel with the output control unit 233 converting the user's voice UV to the voice conversion model, the superimposing unit 234 extracts the pronunciation start portion SP to be superimposed by the second voice (signal processing voice or raw voice).

[0093] The second overlay function (overlay onto the pronunciation start portion SP) may be executed by default at a predetermined overlay ratio when converting the user's voice, or the second overlay function may be enabled or disabled by the user's selection (setting of the overlay ratio).

[0094] <Signal processing unit 235> The signal processing unit 235 converts the input user's voice into a signal-processed voice by signal processing. The signal processing performed by the signal processing unit 235 is, for example, a well-known waveform processing technique (algorithm) for converting voice quality, such as pitch conversion, formant conversion, etc. The signal-processed voice superimposed on the first voice (or substituted for the first voice) may be one that has been subjected to only pitch conversion, or one that has been subjected to both pitch and formant conversion.

[0095] The signal processor 235 performs signal processing on the user's voice so as to obtain a signal processed voice as close as possible to the first voice. For example, the signal processor 235 performs signal processing to obtain a voice close to a registered voice selected by the user using a table in which registered voices and signal processing parameters are previously associated with each other.

[0096] <Output control unit 233> In the second conversion function, the output control unit 233 outputs to the user terminal 3 or the like a voice in which the voice of another person (first voice) and the voice of the user (second voice) are superimposed based on the superimposition ratio determined by the superimposition unit 234. With such a configuration, it is possible to suppress delays in the output voice caused by the processing time of the voice conversion model, and to clarify the phonemes of the output voice.

[0097] The output control unit 233 may receive a composite waveform in which the first voice and the second voice are superimposed by the superimposing unit 234 based on the superimposing ratio from the superimposing unit 234, and output this composite waveform to the user terminal 3, etc. Also, the output control unit 233 may receive the superimposing ratio determined by the superimposing unit 234 from the superimposing unit 234, generate a composite waveform in which the first voice and the second voice are superimposed based on this superimposing ratio, and output this composite waveform to the user terminal 3, etc.

[0098] The specific procedure for superimposing the first voice and the second voice by the superimposing unit 234 and the output control unit 233 is as follows. When the superimposing ratio is 0% (i.e., when superimposing is not performed on the first voice), the output control unit 233 outputs only the first voice. At this time, the first voice is output with a delay of the total time (e.g., 50 ms to 100 ms) of the time required for data conversion and transmission / reception in the user terminal 3, the time required for conversion processing of the voice conversion model, and the communication time to the audio output device. Also, when the superimposing ratio is 100% (i.e., when the first voice is replaced with the second voice), the output control unit 233 outputs only the second voice. At this time, the second voice is output with a delay of the time obtained by subtracting the time required for conversion processing of the voice conversion model from the delay time of the first voice (e.g., 20 ms).

[0099] When the overlap ratio is greater than 0% and less than 100%, the output control unit 233 overlaps the first voice and the second voice at the overlap ratio with a delay of the allowable delay time input by the user, and outputs the two. When the allowable amount is smaller than the actual delay time of the first voice (for example, the allowable amount is 40 ms and the actual delay time of the first voice is 60 ms), the output control unit 233 outputs the first voice with a delay of the allowable amount (40 ms), and outputs the second voice with the actual delay time (60 ms). Therefore, in this case, the second voice is output with a delay from the first voice.

[0100] The output control unit 233 causes the first superimposition function of the superimposition unit 234 to output a voice in which the voice of another person (first voice) and the voice of the user converted by signal processing (signal processed voice) are superimposed based on the superimposition ratio. Also, the output control unit 233 causes a voice in which the voice of another person (first voice) and the user's own voice (live voice) are superimposed based on the superimposition ratio. With such a configuration, it is possible to reduce a delay in feedback of the output voice to the user while converting the user's voice.

[0101] Furthermore, the output control unit 233 outputs the superimposed voice to the user terminal 3 or the like so that the user can recognize it. With this configuration, it is possible to provide the user with voice feedback with reduced delay. Therefore, in distribution, concerts, etc., the sense of incongruity of the output voice of the user who is the speaker or singer is reduced. The superimposed voice may be output so that only the user who is the voice input can hear it and the viewer or audience cannot hear it. In other words, the output control unit 233 may output the superimposed voice only to the user terminal 3 used by the user who is the input, and not to the audio output device of the distribution destination or venue. Furthermore, the output control unit 233 may output the superimposed voice to both the user terminal 3 used by the user who is the input, and the audio output device of the distribution destination or venue. The user can arbitrarily set the output destination of the superimposed voice from the user terminal 3.

[0102] For the second superimposing function of the superimposing unit 234, the output control unit 233 outputs a voice in which the user's voice converted by signal processing (signal processed voice) is superimposed on the pronunciation start part of the voice of another person (first voice). The output control unit 233 may also output a voice in which the user's own voice (live voice) is superimposed on the pronunciation start part of the voice of another person (first voice). With this configuration, it is possible to clarify the phonemes at the pronunciation start part in the converted voice while reducing the sense of incongruity in the superimposed part. That is, in the converted voice, accents (obscuration) are likely to occur at the pronunciation start part by prioritizing the processing speed, but this accent can be reduced by superimposing the second voice.

[0103] For the second superimposition function of the superimposition unit 234, the output control unit 233 may superimpose the second voice onto the pronunciation start portion in the following procedure. First, the output control unit 233 separates each of the first voice and the second voice into a sine wave and a noise model, such as SMS (Spectral Modeling Synthesis). Next, the output control unit 233 adds or replaces the envelope of the sine wave (harmonic envelope) of the pronunciation start portion of the first voice with the envelope of the sine wave of the pronunciation start portion of the second voice.

[0104] 3. Information processing method This section describes an information processing method of the information processing device 2. This information processing method is executed by a computer, with each unit of the information processing device 2 acting as each step.

[0105] The first information processing method for executing the first conversion function includes a receiving step, a setting step, and an output control step. In the receiving step, impression information indicating an impression of the voice of another person desired by the user is received. In the setting step, a registered voice corresponding to the impression information is set as the voice of the other person from among a plurality of registered voices prepared in advance. In the output control step, the user's voice and the ID of the registered voice set as the voice of the other person are input to a voice conversion model, thereby outputting the user's voice converted into the registered voice.

[0106] 9 is an activity diagram showing the flow of the first information processing method. In the following, the first information processing method will be described along with each activity in this activity diagram.

[0107] First, the user inputs impression information of the voice of another person to be converted in the user terminal 3 (activity A110). FIG. 10 is a diagram showing an example of an input screen IS of impression information displayed in the user terminal 3. For example, the information processing device 2 displays an input screen IS for allowing a user to select a word constituting the impression information alternatively using radio buttons, as shown in FIG. 10A. In this input screen IS, the selected word is accepted as impression information by inputting an input button SB. The information processing device 2 accepts input of a plurality of words constituting the impression information by repeating input of impression information through this input screen IS and display of the next option. In addition, the information processing device 2 may accept input of the degree of a word using a slider SD arranged on the input screen IS, as shown in FIG. 10B. Furthermore, the information processing device 2 may display an input screen IS including a free entry field IF on the user terminal 3, as shown in FIG. 10C, and accept input of a sentence constituting the impression information.

[0108] The user terminal 3 transmits the input impression information to the information processing device 2 (activity A120). The information processing device 2 receives the impression information from the user terminal 3 (activity A130). Next, the information processing device 2 sets a registered voice based on the input impression information (activity A140).

[0109] After the registered voice is set, the information processing device 2 accepts input of the user's voice from the user terminal 3 (activity A150). When the information processing device 2 is in a state in which it can accept voice input, the user's voice is input to the user terminal 3 (activity A160). The user terminal 3 transmits the input voice to the information processing device 2 at any time (activity A170). The information processing device 2 inputs the user's voice transmitted from the user terminal 3 into a voice conversion model corresponding to the set registered voice (activity A180). Furthermore, the information processing device 2 transmits the voice of another person output from the voice conversion model to the user terminal 3 (activity A190). The user terminal 3 receives the voice of another person from the information processing device 2 and outputs it as a voice (activity A200).

[0110] While the user continues to input voice, the information processing device 2 repeats the activities from voice input reception (A150) to transmission of another person's voice (A190).

[0111] Furthermore, for example, when the voice input from the user terminal 3 is singing, the information processing device 2 may accept input of impression information or speaker information while the user is singing. When impression information or speaker information is input while the user is singing, the registered voice is immediately changed by the activity of the registered voice setting (A140), and the voice of another person output from the user terminal 3 is changed in real time. For example, at the beginning of singing, a "clear voice" is input as impression information, and when the user sings one phrase, a voice of another person corresponding to the "clear voice" is output for this one phrase. After the user sings one frame, when a "tense voice" is input as impression information, the voice of another person is changed to one corresponding to the "tense voice" from the next phrase. In other words, the voice of another person output almost in real time changes depending on the input of impression information or speaker information while singing.

[0112] A second information processing method for executing the second conversion function includes a receiving step, a superimposing step, and an output control step. In the receiving step, a delay time tolerance or pronunciation clarity of a voice to be output in the output control step in response to an input of a user's voice is received. In the superimposing step, a superimposing ratio of the user's voice that has not been input to the voice conversion model to a voice of another person into which the user's voice has been converted by the voice conversion model is determined. In the output control step, a voice in which the voice of the other person and the user's voice are superimposed is output based on the superimposing ratio.

[0113] 11 is an activity diagram showing the flow of the second information processing method. In the following, the second information processing method will be described along with each activity in this activity diagram.

[0114] First, the user inputs a delay time tolerance (tolerable delay amount) or pronunciation clarity from the user to the user terminal 3 (activity A310). The user terminal 3 transmits the input delay tolerance amount or pronunciation clarity to the information processing device 2 (activity A320). The information processing device 2 receives the delay tolerance amount or pronunciation clarity from the user terminal 3 (activity A330). Next, the information processing device 2 sets a superimposition ratio based on the input delay tolerance amount or pronunciation clarity (activity A340).

[0115] After the superimposition ratio is determined, the information processing device 2 accepts the input of the user's voice from the user terminal 3 (activity A350). When the information processing device 2 is in a state in which it can accept voice input, the user's voice is input to the user terminal 3 (activity A360). The user terminal 3 transmits the input voice to the information processing device 2 at any time (activity A370). The information processing device 2 inputs the user's voice transmitted from the user terminal 3 to a voice conversion model corresponding to the set registered voice (activity A380). Furthermore, the information processing device 2 transmits to the user terminal 3 a voice (superimposed voice) in which a signal-processed voice or a raw voice (second voice) is superimposed at a superimposition ratio on the voice of another person (first voice) output from the voice conversion model (activity A390). The user terminal 3 receives the superimposed voice from the information processing device 2 and outputs it as a voice (activity A400).

[0116] The information processing device 2 repeats the activities from voice input reception (A350) to transmission of another person's voice (A390) while the user continues to input voice. The information processing device 2 may also receive input of the allowable delay amount or pronunciation clarity while the user is inputting voice. If the allowable delay amount or pronunciation clarity is changed while the voice is being input, the superimposition ratio is immediately changed by the activity of determining the superimposition ratio (A340), and the superimposition ratio in the superimposed voice output from the user terminal 3 is changed in real time.

[0117] 4. Effect The effects of this embodiment can be summarized as follows: That is, the convenience for the user when converting the input voice of the user into the voice of another person is improved.

[0118] Although the embodiment of the present invention has been described above, the present invention is not limited to this, and can be modified as appropriate without departing from the technical concept of the invention.

[0119] 5.Other In the above embodiment, the information processing device 2 performs various storage and control, but multiple external devices may be used instead of the information processing device 2. That is, various information and programs may be distributed and stored in multiple external devices using block chain technology or the like.

[0120] The aspect of the present embodiment is not limited to the information processing system 1, and may be an information processing method or a program. The information processing method includes each step of the information processing device 2. The program causes a computer to function as the information processing device 2.

[0121] The information processing system 1 may be an integration of an information processing device 2 and a user terminal 3. That is, the information processing device 2 may input impression information, delay tolerance, voice, and the like from a user, and the user terminal 3 may perform voice conversion (synthesis) processing.

[0122] The information processing system 1 may have the following configuration. That is, the information processing system 1 includes a processor capable of executing a program so that each of the following units is operated. The reception unit receives designation information of a voice of another person desired by the user and the pitch of the user's voice. The setting unit sets a registered voice corresponding to the designation information and the pitch of the user's voice as the voice of another person from among a plurality of registered voices prepared in advance. With this configuration, the user's voice can be converted into a voice of another person suitable for the pitch of the user. Note that in the information processing system 1, it is not necessarily required to set the registered voice based on impression information. Also, in the information processing system 1, it is not necessarily required to superimpose the first voice and the second voice.

[0123] It may be provided in any of the following ways:

[0124] (1) An information processing system that converts a user's voice into the voice of another person different from the user, comprising a processor capable of executing a program to perform each of the following steps: in a reception step, impression information indicating the impression of the voice of the other person desired by the user is received; and in a setting step, the registered voice corresponding to the impression information is set as the voice of the other person from among a plurality of registered voices prepared in advance.

[0125] According to such a configuration, the user's voice can be converted into a voice that is closer to the image (impression) of the user.

[0126] (2) In the information processing system described in (1) above, in the setting step, the registered voice assigned an ID identical or similar to the numerical information derived from the impression information is set as the voice of the other person.

[0127] According to such a configuration, it becomes easy to register and manage the registration voice using the ID.

[0128] (3) In the information processing system described in (2) above, in the setting step, the numerical information is obtained by vectorizing the impression information through natural language processing.

[0129] According to this configuration, it is possible to select a registered voice that is closest to the user's image based on the input of qualitative impression information.

[0130] (4) In the information processing system described in (2) or (3) above, in the output control step, the user's voice and the ID of the registered voice set as the voice of the other person are input into a voice conversion model, thereby outputting the user's voice converted into the registered voice, wherein the voice conversion model is a model that has been machine-learned to determine the relationship between the input voice and the pronunciation of the registered voice for each ID.

[0131] According to such a configuration, an input voice can be accurately converted into a voice of another person based on the user's impression information.

[0132] (5) In the information processing system described in (4) above, the voice conversion model includes a plurality of encoders corresponding to the plurality of enrollment voices, respectively, and a plurality of decoders corresponding to the plurality of enrollment voices, the encoders being configured to receive data of the user's voice as an input and to output intermediate features, and the decoders being configured to receive data of the enrollment voice into which the user's voice has been converted, and the output control step selects the encoder that inputs the user's voice and the decoder that inputs the intermediate features based on the ID.

[0133] According to such a configuration, both the encoder and the decoder prepared for each registered voice are selected based on the impression information input by the user, so that the accuracy of conversion into a voice of another person is improved.

[0134] (6) An information processing system according to any one of (2) to (5) above, wherein the receiving step receives designation information other than the impression information that designates the voice of the other person, and the setting step sets the registered voice having an ID that is the same as or similar to the designation information as the voice of the other person.

[0135] According to this configuration, it is possible to provide the user with a variety of methods for specifying the voice to be converted.

[0136] (7) In the information processing system described in any one of (1) to (6) above, the receiving step further receives the pitch of the user's voice, and the setting step sets, from among the multiple registered voices, the registered voice that corresponds to the impression information and the pitch of the user's voice as the voice of the other person.

[0137] According to this configuration, the user's voice can be converted into another person's voice that matches the pitch of the user.

[0138] (8) The information processing system according to any one of (1) to (7) above, further comprising at least one of a computer and an audio interface.

[0139] With this configuration, it is possible to provide a component device of an information processing system as a computer or audio interface.

[0140] (9) An information processing system that converts a user's voice into the voice of another person different from the user, comprising a processor capable of executing a program to perform each of the following steps: in a superimposition step, a superimposition ratio of the user's voice that has not been input to the voice conversion model relative to the voice of the other person into which the user's voice has been converted by a voice conversion model is determined; and in an output control step, a voice in which the voice of the other person and the user's voice are superimposed based on the superimposition ratio is output.

[0141] According to such a configuration, it is possible to suppress delays in the output voice caused by the processing time of the voice conversion model, and to clarify the phonemes of the output voice.

[0142] (10) An information processing system as described in (9) above, further comprising: in the signal processing step, converting the voice of the user by signal processing; in the superimposition step, determining as the superimposition ratio a superimposition ratio of the voice of the user converted by the signal processing relative to the voice of the other person; and in the output control step, outputting a voice in which the voice of the other person and the voice of the user converted by the signal processing are superimposed based on the superimposition ratio.

[0143] According to such a configuration, it is possible to convert the user's voice while suppressing the delay in feedback of the output voice to the user.

[0144] (11) An information processing system as described in (9) or (10) above, wherein in the superimposition step, a superimposition ratio of the user's own voice relative to the voice of the other person is determined as the superimposition ratio, and in the output control step, a voice in which the voice of the other person and the user's own voice are superimposed based on the superimposition ratio is output.

[0145] According to such a configuration, it is possible to convert the user's voice while suppressing the delay in feedback of the output voice to the user.

[0146] (12) An information processing system according to any one of (10) to (11) above, further comprising, in the reception step, receiving an allowable delay time for the voice to be output in the output control step in response to input of the user's voice, and, in the superimposition step, determining the superimposition ratio based on the allowable delay time.

[0147] According to this configuration, the superimposition ratio can be set according to the tolerable amount of delay preferred by the user.

[0148] (13) The information processing system according to (12) above, wherein in the superimposing step, the superimposing ratio is increased as the allowable amount is smaller.

[0149] According to this configuration, the proportion of the voice converted by the voice conversion model can be adjusted according to the allowable amount.

[0150] (14) In the information processing system according to any one of (9) to (13) above, in the output control step, the superimposed voice is output so as to be recognizable by the user.

[0151] With this configuration, it is possible to provide the user with voice feedback with reduced delay.

[0152] (15) In the information processing system described in (9) above, in the superimposition step, the superimposition ratio is determined so that the voice of the other person is superimposed on the voice of the user that has not been input to the voice conversion model at the pronunciation start part of the user's voice.

[0153] According to this configuration, it is possible to clarify the phoneme at the beginning of pronunciation in the converted voice.

[0154] (16) In the information processing system described in (15) above, further, in the signal processing step, the user's voice is converted by signal processing, and in the superimposition step, the superimposition ratio is determined so that the voice of the other person is superimposed on the voice of the user converted by the signal processing at the pronunciation start part, and in the output control step, a voice in which the voice of the user converted by the signal processing is superimposed at the pronunciation start part of the voice of the other person is output.

[0155] According to this configuration, it is possible to clarify the phonemes at the beginning of pronunciation in the converted voice while reducing the sense of incongruity in the superimposed portion.

[0156] (17) In the information processing system described in any one of (9) to (16) above, in a receiving step, designation information of the voice of the other person desired by the user and the pitch of the user's voice are received, and in a setting step, the registered voice corresponding to the designation information and the pitch of the user's voice is set as the voice of the other person from among a plurality of registered voices prepared in advance.

[0157] According to this configuration, the user's voice can be converted into another person's voice that matches the pitch of the user.

[0158] (18) The information processing system according to any one of (9) to (17) above, further comprising at least one of a computer and an audio interface.

[0159] With this configuration, it is possible to provide a component device of an information processing system as a computer or audio interface.

[0160] (19) An information processing method comprising each step of the information processing system according to any one of (1) to (18) above.

[0161] (20) A program that causes a computer to execute each step of the information processing system described in any one of (1) to (18) above. Of course, this is not the case.

[0162] Finally, although various embodiments according to the present disclosure have been described, these are presented as examples and are not intended to limit the scope of the invention. The novel embodiment can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. The embodiments and their modifications are included within the scope and spirit of the invention, and are included in the scope of the invention and its equivalents described in the claims. [Explanation of symbols]

[0163] 1: Information processing system 2: Information processing equipment 3: User terminal 3A: Audio interface 3B: Computer 20: Communication bus 21: Communications Department 22: Storage section 23: Processor 30: Communication bus 31: Communications Department 32: Storage section 33: Processor 34:Display section 35: Input section 231: Reception 232: Setting section 233: Output control section 234: Overlapped section 235: Signal processing section AF: Acoustic feature data AFS: Synthetic acoustic feature data D1: Acoustic data D3: Synthetic acoustic data IS: Input screen M1: Voice conversion model M11:Analysis Department M12: Encoder M16: Decoder M17: Vocoder MF1: ​​First intermediate feature MF2: Second intermediate feature SB: Selection button SD: Slider SP: Start of pronunciation UV: User feedback

Claims

1. An information processing system for converting a user's voice into a voice of a person different from the user, comprising: A processor capable of executing a program to perform the following steps: In the receiving step, impression information indicating an impression of the voice of the other person desired by the user is received; In the setting step, the information processing system sets a registered voice corresponding to the impression information from among a plurality of registered voices prepared in advance as the voice of the different person.

2. 2. The information processing system according to claim 1, In the setting step, the registered voice assigned an ID that is the same as or similar to numerical information derived from the impression information is set as the voice of the different person.

3. 3. The information processing system according to claim 2, An information processing system, wherein in the setting step, the numerical information is obtained by vectorizing the impression information through natural language processing.

4. 3. The information processing system according to claim 2, Furthermore, in an output control step, the user's voice and the ID of the registered voice, which is set as the voice of the other person, are input into a voice conversion model, thereby outputting the user's voice converted into the registered voice, wherein the voice conversion model is a machine-learned model of the relationship between the input voice and the pronunciation of the registered voice for each ID.

5. 5. The information processing system according to claim 4, The voice conversion model is a plurality of encoders corresponding to the plurality of enrollment voices, respectively; a plurality of decoders corresponding to the plurality of enrollment voices, respectively; Including, The encoder is configured to receive voice data of the user and output intermediate features; the decoder is configured to receive the intermediate features as an input and output the enrollment voice data obtained by converting the voice of the user; In the output control step, the information processing system selects the encoder to which the user's voice is input and the decoder to which the intermediate feature is input, based on the ID.

6. 3. The information processing system according to claim 2, In the receiving step, designation information for designating a voice of the other person other than the impression information is received, In the setting step, the registered voice having an ID identical or similar to the specified information is set as the voice of the different person.

7. 2. The information processing system according to claim 1, The receiving step further includes receiving a pitch of the user's voice; In the setting step, a registered voice corresponding to the impression information and the pitch of the user's voice is set as the voice of the other person from among the multiple registered voices.

8. 2. The information processing system according to claim 1, An information processing system comprising at least one of a computer and an audio interface.

9. An information processing system for converting a user's voice into a voice of a person different from the user, comprising: A processor capable of executing a program to perform the following steps: In the superimposing step, a superimposing ratio of the voice of the user that has not been input to the voice conversion model to the voice of the other person obtained by converting the voice of the user by the voice conversion model is determined; In the output control step, a voice in which the voice of the other person and the voice of the user are superimposed based on the superimposition ratio is output.

10. 10. The information processing system according to claim 9, Furthermore, in the signal processing step, the voice of the user is converted by signal processing, In the superimposing step, a superimposing ratio of the voice of the user converted by the signal processing to the voice of the other person is determined as the superimposing ratio; The information processing system, in the output control step, outputs a voice in which the voice of the other person and the voice of the user converted by the signal processing are superimposed based on the superimposition ratio.

11. 10. The information processing system according to claim 9, In the superimposing step, a superimposing ratio of the user's own voice to the voice of the other person is determined as the superimposing ratio; In the output control step, a voice in which the voice of the other person and the user's own voice are superimposed based on the superimposition ratio is output.

12. 11. The information processing system according to claim 10, Furthermore, in the receiving step, a permissible delay time of the voice to be output in the output control step with respect to the input of the user's voice is received, In the superimposing step, the superimposing ratio is determined based on the allowable amount.

13. 13. The information processing system according to claim 12, In the superimposing step, the superimposing ratio is increased as the allowable amount decreases.

14. 10. The information processing system according to claim 9, In the output control step, the superimposed voice is output so as to be recognizable by the user.

15. 10. The information processing system according to claim 9, An information processing system in which, in the superimposition step, the superimposition ratio is determined so that the voice of the other person is superimposed on the voice of the user that has not been input to the voice conversion model at the pronunciation start part of the user's voice.

16. 16. The information processing system according to claim 15, Furthermore, in the signal processing step, the voice of the user is converted by signal processing, In the superimposing step, the superimposing ratio is determined so that the voice of the other person is superimposed on the voice of the user converted by the signal processing at the pronunciation start portion; In the output control step, the information processing system outputs a voice in which the voice of the user converted by the signal processing is superimposed on the pronunciation start portion of the voice of the other person.

17. 10. The information processing system according to claim 9, Furthermore, in the receiving step, the designation information of the voice of the other person desired by the user and the pitch of the user's voice are received, Furthermore, in the setting step, the information processing system sets, from among a plurality of registered voices prepared in advance, the registered voice corresponding to the specified information and the pitch of the user's voice as the voice of the different person.

18. 10. The information processing system according to claim 9, An information processing system comprising at least one of a computer and an audio interface.

19. 1. An information processing method, comprising: An information processing method comprising the steps of the information processing system according to any one of claims 1 to 18.

20. A program, A program causing a computer to execute each step of the information processing system according to any one of claims 1 to 18.

Citation Information

Patent Citations

  • Training method, speaker identification method, and recording medium

    JP2021033260A