Speech processing system, speech processing method, and speech processing program

The voice processing system addresses the challenge of adjusting audio for multiple users by using processors to manage individual volume requests, ensuring each user receives optimized audio output.

JP2025173123APending Publication Date: 2025-11-27SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024078531
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-14
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Conventional audio systems struggle to output appropriate audio for each user individually, as adjusting the volume of one user affects the volumes of other users, leading to difficulty in hearing specific voices.

Method used

A voice processing system that includes an acquisition processor, reception processor, and output processor to manage multiple audio devices, allowing for individual volume adjustments based on user requests, ensuring each user receives optimized audio output.

Benefits of technology

The system effectively outputs appropriate voice for each user by performing voice adjustment processing on specific audio data, maintaining consistent volume levels for all users involved in a conversation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025173123000001_ABST
    Figure 2025173123000001_ABST
Patent Text Reader

Abstract

To provide a speech processing system, a speech processing method, and a speech processing program that can output appropriate speech for each user of an audio device.SOLUTION: In a speech processing device 1, an acquisition processing unit 111 acquires a plurality of speech data corresponding to the user's speech from each of a plurality of user terminals 2. A reception processing unit 113 accepts an adjustment request for first speech data from the plurality of speech data acquired by the acquisition processing unit 111. An adjustment processing unit 114 performs speech adjustment processing on the first speech data in response to the adjustment request. An output processing unit 115 causes the first user terminal 2, which is the source of the adjustment request, to output the adjusted speech data that has been subjected to the speech adjustment processing and second speech data that is the plurality of speech data excluding the first speech data, and also causes the plurality of speech data to be output from a second user terminal 2 from the plurality of speech devices excluding the first user terminal 2.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a technology for controlling audio when multiple users individually use audio devices to have a conversation. [Background technology]

[0002] Conventionally, there is known a system in which multiple users can have a conversation using audio devices each equipped with a microphone and a speaker. Also, there is known a technology for automatically adjusting the volume level of the audio for each speaker during the conversation (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-069136 Summary of the Invention [Problem to be solved by the invention]

[0004] However, with conventional technology, if a specific user's voice is difficult to hear, adjusting the volume of that user will also change the volume of other users. As such, with conventional technology, it is difficult to output appropriate audio for each user.

[0005] An object of the present disclosure is to provide a voice processing system, a voice processing method, and a voice processing program that are capable of outputting appropriate voice for each user of an audio device. [Means for solving the problem]

[0006] A voice processing system according to one aspect of the present disclosure includes an acquisition processor, a reception processor, an adjustment processor, and an output processor. The acquisition processor acquires a plurality of voice data corresponding to a user's speech from a plurality of voice devices. The reception processor accepts an adjustment request for first voice data from the plurality of voice data acquired by the acquisition processor. The adjustment processor performs voice adjustment processing on the first voice data in response to the adjustment request. The output processor causes the first voice device, which is the source of the adjustment request, to output adjusted voice data resulting from the voice adjustment processing and second voice data, which is the plurality of voice data excluding the first voice data, and causes the plurality of voice data to be output from a second voice device, which is the plurality of voice devices excluding the first voice device.

[0007] Another aspect of the present disclosure is an audio processing method executed by one or more processors, which includes acquiring multiple audio data corresponding to a user's speech from each of multiple audio devices, accepting an adjustment request for first audio data from the multiple audio data, performing an audio adjustment process on the first audio data in response to the adjustment request, outputting the adjusted audio data resulting from the audio adjustment process and second audio data from the multiple audio data excluding the first audio data from the first audio device that requested the adjustment, and outputting the multiple audio data from a second audio device from the multiple audio devices excluding the first audio device.

[0008] Another aspect of the present disclosure is a voice processing program for causing one or more processors to execute the following: acquiring multiple voice data corresponding to a user's spoken voice from each of multiple voice devices; accepting an adjustment request for first voice data from the multiple voice data; performing voice adjustment processing on the first voice data in response to the adjustment request; outputting the adjusted voice data after the voice adjustment processing and second voice data from the multiple voice data excluding the first voice data from the first voice device that requested the adjustment request; and outputting the multiple voice data from a second voice device from the multiple voice devices excluding the first voice device. [Effects of the Invention]

[0009] According to the present disclosure, it is possible to provide a voice processing system, a voice processing method, and a voice processing program that are capable of outputting appropriate voice for each user of an audio device. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating an application example of a voice processing system according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a block diagram showing a configuration of a voice processing system according to an embodiment of the present disclosure. [Figure 3] FIG. 3 is a diagram illustrating a specific example of audio processing by the audio processing system according to an embodiment of the present disclosure. [Figure 4] FIG. 4 is a diagram illustrating a specific example of audio processing by the audio processing system according to an embodiment of the present disclosure. [Figure 5] FIG. 5 is a diagram illustrating a specific example of audio processing by the audio processing system according to an embodiment of the present disclosure. [Figure 6] FIG. 6 is a diagram illustrating a specific example of audio processing by the audio processing system according to an embodiment of the present disclosure. [Figure 7]FIG. 7 is a flowchart illustrating an example of a procedure of a voice control process executed in a voice processing device according to an embodiment of the present disclosure. [Figure 8] FIG. 8 is a diagram illustrating a specific example of voice processing by the voice processing system according to the first modification of the present disclosure. [Figure 9] FIG. 9 is a flowchart illustrating an example of a procedure of a voice control process executed in a voice processing device according to the first modification of the present disclosure. [Figure 10] FIG. 10 is a diagram illustrating a specific example of voice processing by the voice processing system according to the second modification of the present disclosure. [Figure 11] FIG. 11 is a flowchart illustrating an example of the procedure of a voice control process executed in a voice processing device according to the second modification of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that the following embodiments are examples that embody the present disclosure and do not limit the technical scope of the present disclosure.

[0012] The voice processing system according to the present disclosure can be applied to, for example, a web meeting (online meeting) in which multiple users, each in different locations (such as an office conference room or their own homes), hold a conversation (conference) using user terminals (examples of the voice device according to the present disclosure) such as laptops and smartphones. The voice processing system can also execute the online meeting using a conversation application, which is general-purpose software for executing the online meeting.

[0013] FIG. 1 shows an application example of a voice processing system 10 according to this embodiment. As shown in FIG. 1, user A participates in a conference room Ra, user B participates in a conference room Rb, and user C participates in a conference room Rc. Users A, B, and C converse using user terminals 2a, 2b, and 2c, respectively. In another embodiment, each user may converse using a microphone / speaker device. For example, each user may use a neckband-type microphone / speaker device that can be worn around the neck, or a stationary microphone / speaker device installed in the conference room. The user terminal 2 and the microphone / speaker device are examples of voice devices of the present disclosure.

[0014] The voice processing system 10 enables multiple users to hold online meetings in remote locations by executing a conversation application installed on each user terminal 2. The conversation application is general-purpose software, and multiple users participating in the same meeting select a common conversation application.

[0015] For example, users A, B, and C start the conversation application on their own user terminals 2a, 2b, and 2c, respectively.

[0016] The voice processing system 10 may be configured such that a camera connectable to the user terminal 2 is connected to each location (such as a conference room or a home) and camera images can be transmitted in two directions. The camera may be built into the user terminal 2.

[0017] The voice processing device 1 and the user terminal 2 are connected to each other via a network N1. The network N1 is a communication network such as the Internet, a LAN, a WAN, or a public telephone line.

[0018] [User device 2] 2, the user terminal 2 includes a control unit 21, a storage unit 22, an operation display unit 23, a microphone 24, a speaker 25, and a communication unit 26. The user terminal 2 is an information processing device such as a laptop computer, a smartphone, or a tablet terminal. Each user terminal 2 may have the same configuration.

[0019] The communication unit 26 is a communication unit that connects the user terminal 2 to the network N1 by wire or wirelessly and executes data communication with other devices (for example, the voice processing device 1) via the network N1 in accordance with a predetermined communication protocol.

[0020] The microphone 24 collects the user's speech. The speaker 25 reproduces the speech output from the speech processing device 1. The microphone 24 and the speaker 25 may be configured as an integrated unit. Furthermore, the microphone 24 and the speaker 25 may be built into the user terminal 2, or may be arranged externally and connected to the user terminal 2 by wire or wirelessly.

[0021] The operation display unit 23 is a user interface that includes a display unit such as a liquid crystal display or an organic EL display that displays various information, and an operation unit such as a mouse, keyboard, or touch panel that accepts operations from the user.

[0022] The storage unit 22 is a non-volatile storage unit such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), or a flash memory that stores various types of information. The storage unit 22 stores a control program for causing the control unit 21 to execute various processes. For example, the control program is non-temporarily recorded on a computer-readable recording medium such as a CD or a DVD, and is read by a reading device (not shown) such as a CD drive or a DVD drive provided in the user terminal 2 and stored in the storage unit 22. The control program may be distributed from a cloud server and stored in the storage unit 22.

[0023] The storage unit 22 also has installed therein one or more conversation applications for providing online meeting services.

[0024] The control unit 21 has control devices such as a CPU, a ROM, and a RAM. The CPU is a processor that executes various types of arithmetic processing. The ROM stores in advance control programs such as a BIOS and an OS that cause the CPU to execute various types of processing. The RAM stores various types of information and is used as a temporary storage memory (work area) for the various types of processing executed by the CPU. The control unit 21 controls the user terminal 2 by having the CPU execute various control programs that are pre-stored in the ROM or the storage unit 22. The control unit 21 also functions as a processing unit that executes the conversation application.

[0025] The control unit 21 includes various processing units. The control unit 21 functions as the various processing units by executing various processes in accordance with the control program using the CPU. Some or all of the processing units included in the control unit 21 may be configured with electronic circuits. The control program may be a program for causing multiple processors to function as the various processing units.

[0026] For example, the control unit 21 executes various processes related to the online meeting in accordance with the conversation application. Specifically, when the control unit 21 receives an operation (login operation) by the user to start the conversation application, it transmits a start request to the voice processing device 1. When the voice processing device 1 authenticates the start request, the control unit 21 displays a conversation screen on the user terminal 2 to start the online meeting.

[0027] When the online meeting starts, for example, the control unit 21 of the user terminal 2a of user A outputs the voice data of the user A's speech input to the microphone 24 of the user terminal 2a to the voice processing device 1. Furthermore, when the control unit 21 of the user terminal 2a acquires voice data obtained by synthesizing the voices of each user output from the voice processing device 1, it causes the speaker 25 to play back the voice.

[0028] The control unit 21 also accepts various operations from the user. For example, when a user feels that the voice of the person he or she is having difficulty hearing and requests an adjustment of that voice, the control unit 21 accepts the operation for the adjustment request. For example, when users A, B, and C are having a conversation and user B feels that only user C's voice is too quiet and difficult to hear, user B inputs a request for adjusting user C's voice on the user terminal 2b. For example, user B selects the target user (user C in this case) on the conversation screen and makes an adjustment request (a request to increase the volume) by pressing an adjustment request button. This causes the control unit 21 of the user terminal 2b to accept user C's voice adjustment request. Upon accepting the adjustment request, the control unit 21 outputs the adjustment request to the voice processing device 1. The voice processing device 1 performs a voice adjustment process on the voice data based on the adjustment request. The control unit 21 plays back the voice of the adjusted voice data after the voice adjustment process. A specific example of the voice adjustment process will be described later.

[0029] When the control unit 21 receives an operation (termination operation) by the user to terminate the conversation application, the control unit 21 transmits a termination request to the voice processing device 1. When the voice processing device 1 authenticates the termination request, the control unit 21 terminates the online meeting in the user terminal 2.

[0030] Each user participating in the online meeting starts the online meeting by launching a conversation application on his / her own user terminal 2. Furthermore, each user ends the online meeting by closing the conversation application on his / her own user terminal 2.

[0031] [Speech processing device 1] 2, the voice processing device 1 is an information processing device including a control unit 11, a storage unit 12, a communication unit 13, etc. The voice processing device 1 may be configured, for example, with one or more servers (for example, cloud servers).

[0032] The communication unit 13 is a communication unit that connects the voice processing device 1 to the network N1 and executes data communication with an external device such as the user terminal 2 via the network N1 in accordance with a predetermined communication protocol.

[0033] The storage unit 12 is a non-volatile storage unit such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory that stores various types of information. Specifically, the storage unit 12 may store data such as information that can identify the user terminal 2 (such as a device number or device ID).

[0034] The storage unit 12 also stores control programs such as a voice control program (an example of a voice processing program of the present disclosure) for causing the control unit 11 to execute voice control processing (see FIGS. 7, 9, and 11) described below. For example, the voice control program may be non-temporarily recorded on a computer-readable recording medium such as a CD or a DVD, read by a reading device (not shown) such as a CD drive or a DVD drive provided in the voice processing device 1, and stored in the storage unit 12.

[0035] The control unit 11 has control devices such as a CPU, a ROM, and a RAM. The CPU is a processor that executes various arithmetic processes. The ROM is a non-volatile storage unit that pre-stores control programs such as a BIOS and an OS that cause the CPU to execute various arithmetic processes. The RAM is a volatile or non-volatile storage unit that stores various information and is used as a temporary storage memory (work area) for various processes executed by the CPU. The control unit 11 controls the audio processing device 1 by having the CPU execute various control programs pre-stored in the ROM or the storage unit 12.

[0036] Specifically, as shown in Fig. 2, the control unit 11 includes various processing units such as an acquisition processing unit 111, a voice processing unit 112, a reception processing unit 113, an adjustment processing unit 114, and an output processing unit 115. The control unit 11 functions as the various processing units by executing various processes in accordance with the control program using the CPU. Some or all of the processing units may be configured with electronic circuits. The control program may be a program for causing multiple processors to function as the processing units.

[0037] The acquisition processing unit 111 acquires a plurality of pieces of voice data corresponding to the user's voice from each of the plurality of user terminals 2. For example, when the online meeting starts and each user speaks, the acquisition processing unit 111 acquires the voice data output from each user terminal 2. The acquisition processing unit 111 stores the acquired voice data in the storage unit 12 with the identification information (device number) of the user terminal 2 or the identification information (user ID) of the user terminal 2.

[0038] The audio processing unit 112 performs predetermined audio processing on the audio data. Specifically, the audio processing unit 112 performs well-known audio processing (pre-processing) on ​​each piece of audio data, such as gain adjustment, noise removal, and echo cancellation.

[0039] Furthermore, the audio processing unit 112 executes a synthesis process to synthesize the audio data that has been subjected to the audio processing. Specifically, the audio processing unit 112 executes synthesis and encoding processes to generate audio data (synthesized audio data) to be output (distributed) to the user terminal 2.

[0040] The reception processing unit 113 accepts a request to adjust specific voice data (first voice data of the present disclosure) from among the multiple voice data acquired by the acquisition processing unit 111. For example, when users A, B, and C are having a conversation, if user B feels that only user C's voice is too quiet and difficult to hear and inputs a request to adjust user C's voice into user terminal 2b, the reception processing unit 113 accepts the request to adjust user C's voice from user terminal 2b. In this way, the reception processing unit 113 accepts the request to adjust the voice data in response to an instruction to request adjustment of the voice data from the user of user terminal 2.

[0041] In response to the adjustment request, the adjustment processing unit 114 performs audio adjustment processing on the audio data to be adjusted. Specifically, the adjustment processing unit 114 performs adjustment processing on the audio data to be adjusted in accordance with the adjustment request, such as volume adjustment, frequency (pitch) adjustment, speed adjustment, etc. For example, when user B makes an adjustment request to increase the volume of user C's voice, the adjustment processing unit 114 increases the volume of user C's voice.

[0042] The output processing unit 115 outputs the processed audio data (synthesized audio data) to the user terminal 2. The output processing unit 115 also outputs the adjusted audio data to the user terminal 2 that has requested the adjustment. The output processing unit 115 transmits the audio data to each user terminal 2, causing it to be output (played) from each user terminal 2.

[0043] Below, we will explain specific examples of the voice processing (preprocessing), voice synthesis processing, voice adjustment processing, and output processing executed by the control unit 11. For example, as shown in Fig. 3, the control unit 11 acquires voice data Va of the voice uttered by user A from the user terminal 2a, voice data Vb of the voice uttered by user B from the user terminal 2b, and voice data Vc of the voice uttered by user C from the user terminal 2c.

[0044] Upon acquiring the voice data Va, Vb, and Vc (see FIG. 3), the control unit 11 performs well-known voice processing (preprocessing) such as gain adjustment, noise removal, and echo cancellation on each of the voice data Va', Vb', and Vc' to generate processed voice data Va', Vb', and Vc'. Next, the control unit 11 performs a synthesis process to synthesize and encode the voice data Va', Vb', and Vc' to generate synthesized voice data Vm1 (Va'+Vb'+Vc') (see FIG. 4). The control unit 11 then outputs (distributes) the synthesized voice data Vm1 to each of the user terminals 2a, 2b, and 2c using channel Ch1 (an example of the second channel of the present disclosure) of the voice processing device 1 (see FIG. 4).

[0045] For ease of explanation, it is assumed here that the same synthetic voice data Vm1(Va'+Vb'+Vc') is output to each user terminal 2, but in reality, the control unit 11 mutes the user's own voice before outputting the data. For example, when transmitting to user terminal 2a, the control unit 11 mutes the voice of user A and transmits synthetic voice data Vm1(Vb'+Vc'), and when transmitting to user terminal 2b, the control unit 11 mutes the voice of user B and transmits synthetic voice data Vm1(Va'+Vc').

[0046] Upon receiving the synthetic voice data Vm1, each of the user terminals 2a, 2b, and 2c reproduces the voice corresponding to the synthetic voice data Vm1 from the speaker 25.

[0047] For example, if user B feels that only user C's voice is difficult to hear (e.g., the volume is low) among the voices played back from user terminal 2b, user B makes an adjustment request to increase the volume of user C's voice (see FIG. 4). In this case, upon receiving the adjustment request from user terminal 2b, control unit 11 executes a voice adjustment process on user C's voice data Vc' in accordance with the adjustment request. Here, control unit 11 generates voice data Vc'' (adjusted voice data) by increasing the volume of voice data Vc'.

[0048] When the control unit 11 executes the audio adjustment process, it generates synthetic audio data including the audio data that has undergone the audio adjustment process. For example, the control unit 11 executes a synthesis process to synthesize audio data Va', Vb', and Vc" to generate synthetic audio data Vm2 (Va'+Vb'+Vc') (see FIG. 5). Then, the control unit 11 outputs (distributes) the synthetic audio data Vm2 to the user terminal 2b that has requested the adjustment, using channel Ch2 (an example of the first channel of the present disclosure) of the audio processing device 1 (see FIG. 5).

[0049] In this way, when user B makes the adjustment request, the control unit 11 outputs the synthetic voice data Vm1 (Va'+Vb'+Vc') to the user terminals 2a and 2c using channel Ch1 of the voice processing device 1, and outputs the synthetic voice data Vm2 (Va'+Vb'+Vc') to the user terminal 2b using channel Ch2 of the voice processing device 1 (see FIG. 5). This makes it easier for user B to hear user C's voice. Furthermore, since the quality of user A's voice does not change, user B can hear the voices of users A and C with the same quality. Furthermore, since user terminal 2a plays user C's voice (voice data Vc') that has not been subjected to voice adjustment processing, the quality of user C's voice does not change. In this way, appropriate voice is output for each user, making it easier for all users to hear the voices.

[0050] 6 shows an example of the audio adjustment process. The control unit 11 generates synthetic audio data Vm1 (Va'+Vb'+Vc') based on the audio data Va, Vb, and Vc acquired from each user terminal 2, and also performs audio adjustment processing in response to an adjustment request (here, an adjustment request for volume, speed, or frequency) to generate synthetic audio data Vm2 (Va'+Vb'+Vc'). The audio adjustment processing includes processing such as increasing / decreasing the volume, increasing / decreasing the speed (higher pitch) / decreasing the speed (lower pitch), and increasing / decreasing the pitch.

[0051] [Voice control processing] FIG. 7 shows an example of the procedure of the voice control process executed by the control unit 11 of the voice processing device 1.

[0052] The present disclosure can be understood as a voice control method (voice processing method of the present disclosure) that executes one or more steps included in the voice control process. Furthermore, one or more steps included in the voice control process described here may be omitted as appropriate. Furthermore, the steps in the voice control process may be executed in a different order as long as the same operational effect is achieved. Furthermore, while the description here takes as an example a case where the control unit 11 executes each step in the voice control process, in other embodiments, one or more processors may execute each step in the voice control process in a distributed manner.

[0053] Here, as shown in FIG. 1, an example will be described in which users A, B, and C hold a conference using user terminals 2a, 2b, and 2c.

[0054] <Step S11> First, in step S11, the control unit 11 acquires voice data from the user terminal 2. Here, the control unit 11 acquires voice data Va of the speech of user A from the user terminal 2a, voice data Vb of the speech of user B from the user terminal 2b, and voice data Vc of the speech of user C from the user terminal 2c (see FIG. 3).

[0055] <Step S12> Next, in step S12, the control unit 11 performs predetermined audio processing (pre-processing) on ​​the acquired audio data. For example, the control unit 11 performs audio processing such as gain adjustment, noise removal, and echo cancellation on the audio data Va, Vb, and Vc. The audio data after audio processing is represented as audio data Va', Vb', and Vc'.

[0056] <Step S13> Next, in step S13, the control unit 11 generates synthetic voice data. Specifically, the control unit 11 generates synthetic voice data Vm1 by synthesizing the voice data Va', Vb', and Vc' after the voice processing (see FIG. 4).

[0057] <Step S14> In step S14, the control unit 11 outputs the synthetic voice data Vm1 (Va'+Vb'+Vc') to the user terminals 2a, 2b, and 2c using channel Ch1 of the voice processing device 1 (see FIG. 4). The user terminals 2a, 2b, and 2c each play back the voice of the synthetic voice data Vm1 (Va'+Vb'+Vc').

[0058] <Step S15> Next, in step S15, the control unit 11 determines whether or not an audio adjustment request has been received from the user terminal 2. If the control unit 11 has received an audio adjustment request from the user terminal 2 (S15: Yes), the control unit 11 proceeds to step S16. On the other hand, if the control unit 11 has not received an audio adjustment request from the user terminal 2 (S15: No), the control unit 11 proceeds to step S19.

[0059] For example, if user B feels that only user C's voice is difficult to hear (for example, the volume is low) among the voices played from user terminal 2b, user B makes an adjustment request to increase the volume of user C's voice (see FIG. 4). In this case, control unit 11 accepts the adjustment request for user C's voice from user terminal 2b and causes the process to proceed to step S16.

[0060] <Step S16> In step S16, the control unit 11 executes a sound adjustment process. Here, the control unit 11 executes a sound adjustment process on the sound data Vc' of the user C in accordance with the adjustment request received from the user terminal 2b. For example, the control unit 11 generates sound data Vc'' (adjusted sound data) by increasing the volume of the sound data Vc' (see FIG. 5).

[0061] <Step S17> In step S17, the control unit 11 generates synthetic voice data. Specifically, the control unit 11 generates synthetic voice data Vm2 (Va'+Vb'+Vc') by combining the voice data Va' and Vb' after the voice processing and the voice data Vc'' after the voice adjustment processing (see FIG. 5).

[0062] <Step S18> In step S18, the control unit 11 outputs the generated synthetic speech data to the user terminal 2. Specifically, the control unit 11 outputs synthetic speech data Vm1 (Va'+Vb'+Vc') to the user terminals 2a and 2c using channel Ch1 of the speech processing device 1, and outputs synthetic speech data Vm2 (Va'+Vb'+Vc') to the user terminal 2b using channel Ch2 of the speech processing device 1 (see FIG. 5).

[0063] The user terminals 2a and 2c play back the voice of the synthetic voice data Vm1 (Va'+Vb'+Vc'), and the user terminal 2b plays back the voice of the synthetic voice data Vm2 (Va'+Vb'+Vc').

[0064] <Step S19> In step S19, the control unit 11 determines whether the conference has ended. For example, if the user performs an operation to end the conference on the user terminal 2, the control unit 11 determines that the conference has ended (S19: Yes) and ends the voice control process. If the control unit 11 determines that the conference has not ended (S19: No), the control unit 11 returns the process to step S11. The control unit 11 repeatedly executes the above process until the conference ends.

[0065] As described above, when a user makes an adjustment request, the control unit 11 outputs the synthetic voice data Vm1 (Va'+Vb'+Vc') to the user terminal 2 of a user other than the user who made the adjustment request, using channel Ch1 of the voice processing device 1, and outputs the synthetic voice data Vm2 (Va'+Vb'+Vc') to the user terminal 2 of the user who made the adjustment request, using channel Ch2 of the voice processing device 1.

[0066] [Variation 1] A first modification of this embodiment will be described with reference to Figs. 8 and 9. In the first modification, for example, when user B makes a request to adjust user C's voice (see Fig. 4), the control unit 11 generates synthetic voice data Vm2 (Va' + Vb') by synthesizing voice data Va and Vb after voice processing (preprocessing) based on voices other than user C's voice (voices of users A and B) (see Fig. 8). Then, the control unit 11 outputs (distributes) the synthetic voice data Vm2 to the user terminal 2b that has made the adjustment request, using channel Ch2(1) of the voice processing device 1, and also outputs (distributes) user C's voice data Vc as the original sound (voice data before voice processing) using channel Ch2(2) of the voice processing device 1 (see Fig. 8).

[0067] In this case, when the control unit 21 of the user terminal 2b acquires the voice data Vc, it performs a voice adjustment process on the voice data Vc. For example, the control unit 21 performs a process to adjust the volume, speed, pitch, etc. of the voice data Vc in response to an adjustment request from user B, thereby generating voice data Vc1. The control unit 21 then simultaneously plays back the voice of the synthetic voice data Vm2 (Va'+Vb') and the voice of the voice data Vc1 after the voice adjustment. Note that the control unit 21 may also re-synthesize the voice of the synthetic voice data Vm2 (Va'+Vb') and the voice of the voice data Vc1 and play them back.

[0068] The user terminals 2a and 2c acquire the synthesized voice data Vm1 (Va'+Vb'+Vc') distributed from the channel Ch1 of the voice processing device 1 and play back the voice.

[0069] Fig. 9 shows a flowchart of the audio control process according to Modification 1. Steps S21 to S25 are the same as steps S11 to S15 shown in Fig. 7, and therefore description thereof will be omitted. Furthermore, in Modification 1, the audio adjustment process (step S16) shown in Fig. 7 is omitted.

[0070] <Step S26> In step S26, the control unit 11 synthesizes the voice data Va and Vb of the voices (voices of users A and B) excluding the voice to be adjusted (voice of user C) to generate synthetic voice data Vm2 (Va'+Vb') (see FIG. 8).

[0071] <Step S27> In step S27, the control unit 11 outputs the synthetic voice data Vm1 (Va'+Vb'+Vc') to the user terminals 2a and 2c using channel Ch1 of the voice processing device 1, outputs the synthetic voice data Vm2 (Va'+Vb') to the user terminal 2b using channel Ch2(1) of the voice processing device 1, and outputs the voice data Vc (original sound) to be adjusted to the user terminal 2b using channel Ch2(2) of the voice processing device 1 (see Figure 8).

[0072] User terminals 2a and 2c play back the voice of synthetic voice data Vm1 (Va'+Vb'+Vc'). User terminal 2b performs voice adjustment processing on the voice data Vc to be adjusted in response to the adjustment request to generate adjusted voice data Vc1. User terminal 2b then plays back the voice of synthetic voice data Vm2 (Va'+Vb') and the voice of voice data Vc1 simultaneously.

[0073] As described above, in the first modification, the control unit 11 of the voice processing device 1 outputs the voice data (original voice) of the voice to be adjusted and the remaining voice data (synthetic voice data) of the multiple voice data excluding the voice data to be adjusted to the user terminal 2 that requested the adjustment, using channels Ch2(1) and Ch(2), and outputs the synthesized voice data obtained by synthesizing the multiple voice data to the other user terminals 2, using channel Ch1. That is, in the first modification, the voice processing device 1 outputs the voice to be adjusted to the user terminal 2 without performing the voice adjustment process, and the user terminal 2 performs the voice adjustment process. For example, the control unit 21 of the user terminal 2 performs the voice adjustment process on the voice data before performing the voice processing including gain adjustment and noise removal. In the first modification, the control unit 21 of the user terminal 2 may function as the adjustment processing unit and output processing unit of the present disclosure.

[0074] [Variation 2] A second modification of this embodiment will be described with reference to Fig. 10 and Fig. 11. In the second modification, for example, when user B makes an adjustment request (see Fig. 4), the control unit 11 generates synthesized voice data Vm2 (Va'+Vb') by combining voice data Va and Vb after voice processing (preprocessing), and voice data Vc2 by performing voice processing (preprocessing) and voice adjustment processing on the voice data Vc to be adjusted (see Fig. 10). For example, the control unit 11 performs processing to adjust the volume, speed, pitch, etc. of the voice data Vc in response to the adjustment request from user B, and generates the voice data Vc2.

[0075] Then, the control unit 11 outputs (distributes) the synthetic voice data Vm2 to the user terminal 2b that requested the adjustment using channel Ch2(1) of the voice processing device 1, and outputs (distributes) the voice data Vc2 after the voice adjustment using channel Ch2(2) of the voice processing device 1 (see Figure 10).

[0076] User terminal 2b acquires synthetic voice data Vm2(Va'+Vb') distributed from channel Ch2(1) of voice processing device 1 and voice data Vc2 distributed from channel Ch2(2) of voice processing device 1, and simultaneously plays back the voice of each voice data. User terminals 2a and 2c also acquire synthetic voice data Vm1(Va'+Vb'+Vc') distributed from channel Ch1 of voice processing device 1, and play back the voice.

[0077] Fig. 11 shows a flowchart of the audio control process according to Modification 2. Steps S31 to S36 are the same as steps S11 to S16 shown in Fig. 7, and therefore a description thereof will be omitted. In step S36, the control unit 11 executes audio adjustment processing on the audio data Vc' of the user C in accordance with the adjustment request received from the user terminal 2b. For example, the control unit 11 generates audio data Vc2 (adjusted audio data) by increasing the volume of the audio data Vc' (see Fig. 10).

[0078] <Step S37> In step S37, the control unit 11 synthesizes the voice data Va and Vb of the voices (voices of users A and B) excluding the voice to be adjusted (voice of user C) to generate synthetic voice data Vm2 (Va'+Vb') (see FIG. 10).

[0079] <Step S38> In step S38, the control unit 11 outputs the synthetic voice data Vm1 (Va' + Vb' + Vc') to the user terminals 2a and 2c using channel Ch1 of the voice processing device 1, outputs the synthetic voice data Vm2 (Va' + Vb') to the user terminal 2b using channel Ch2(1) of the voice processing device 1, and outputs the voice data Vc2 that has been subjected to voice adjustment processing to the user terminal 2b using channel Ch2(2) of the voice processing device 1 (see Figure 10).

[0080] User terminals 2a and 2c play back the voice of synthetic voice data Vm1 (Va'+Vb'+Vc'). User terminal 2b plays back the voice of synthetic voice data Vm2 (Va'+Vb') and the voice of voice data Vc2 simultaneously.

[0081] As described above, in Modification 2, the control unit 11 of the voice processing device 1 outputs voice data of the voice to be adjusted (adjusted voice data) and the remaining voice data (synthetic voice data) of the multiple voice data excluding the voice data to be adjusted to the user terminal 2 that requested the adjustment, using channels Ch2(1) and Ch(2), and outputs synthetic voice data obtained by synthesizing the multiple voice data to the other user terminals 2, using channel Ch1. That is, in Modification 2, the voice processing device 1 performs voice adjustment processing on the voice to be adjusted, and outputs the voice data after the adjustment processing and the synthetic voice data of the other voices separately to the user terminal 2 that requested the adjustment. In Modification 2, the control unit 21 of the user terminal 2 may function as the output processing unit of the present disclosure.

[0082] As described in the above embodiments, the voice processing system 10 according to the present disclosure acquires, from each of a plurality of user terminals 2 (voice devices), a plurality of voice data corresponding to a user's speech, and accepts an adjustment request for specific voice data (first voice data) from the acquired plurality of voice data. The voice processing system 10 performs a voice adjustment process on the first voice data in response to the adjustment request, and outputs (plays) the adjusted voice data obtained by the voice adjustment process and second voice data, which is the plurality of voice data excluding the first voice data, from the first user terminal 2 (first voice device) that issued the adjustment request, and also outputs the plurality of voice data from a second user terminal 2, which is the plurality of user terminals 2 excluding the first user terminal 2.

[0083] For example, the voice processing system 10 outputs synthetic voice data Vm2 obtained by synthesizing the adjusted voice data and the second voice data to the first user terminal 2, and outputs synthetic voice data Vm1 obtained by synthesizing the plurality of voice data to the second user terminal 2 (see FIG. 5). Also, for example, the voice processing system 10 outputs synthetic voice data Vm1 to the first user terminal 2 on channel Ch1, and outputs synthetic voice data Vm2 to the second user terminal 2 on channel Ch2.

[0084] Also, for example, the voice processing system 10 outputs the first voice data to be adjusted (e.g., the original sound of voice data Vc) and the second voice data (e.g., synthesized voice data Vm2 of voice data Va and Vb) to the first user terminal 2 on channel Ch2 (e.g., channels Ch2(1) and Ch2(2)), and outputs synthesized voice data Vm1 obtained by synthesizing the multiple voice data to the second user terminal 2 on channel Ch1 (see Figure 8).

[0085] For example, the voice processing system 10 outputs adjusted voice data that has undergone voice adjustment processing (e.g., voice data Vc2 obtained by subjecting voice data Vc to voice adjustment processing) and second voice data (e.g., synthesized voice data Vm2 of voice data Va and Vb) to the first user terminal 2 on channel Ch2 (e.g., channels Ch2(1) and Ch2(2)), and outputs synthesized voice data Vm1 obtained by synthesizing the multiple voice data to the second user terminal 2 on channel Ch1 (see FIG. 10).

[0086] The speech processing system 10 according to this embodiment may have a translation function for translating speech in a first language received from a user terminal 2 into a second language. For example, the speech processing device 1 acquires Japanese speech data Va and Vb from the user terminals 2a and 2b, and acquires English speech data Vc from the user terminal 2c. When a user B makes an adjustment request (translation request), the speech processing device 1 outputs (distributes) synthesized speech data Vm2 (Va' + Vb') obtained by synthesizing the speech data Va and Vb to the user terminal 2b that issued the adjustment request, via channel Ch2(1), and also outputs (distributes) speech data Vc (original speech) (see FIG. 8) or speech data Vc2 after speech adjustment (see FIG. 10) via channel Ch2(2). The user terminal 2b translates the speech data Vc or Vc2 from English to Japanese. In this way, when different languages ​​are included, the speech processing device 1 can improve translation accuracy by distributing speech data to the user terminal 2 using different channels for each language.

[0087] The translation process may be performed by the speech processing device 1. For example, when the speech processing device 1 receives Japanese speech data Va and Vb from user terminals 2a and 2b and receives English speech data Vc from user terminal 2c, it translates the speech data Vc from English to Japanese. Then, the speech processing device 1 outputs (distributes) synthesized speech data Vm2 (Va'+Vb') to user terminal 2b via channel Ch2(1), and outputs (distributes) translated speech data after translation via channel Ch2(2).

[0088] The speech processing system 10 according to this embodiment may include a text conversion function (transcription function) for converting speech data into text. For example, when the speech processing device 1 acquires speech data Va, Vb, and Vc from user terminals 2a, 2b, and 2c and receives a request from user B to adjust the speech of user C, the speech processing device 1 outputs (distributes) synthesized speech data Vm2 (Va' + Vb') obtained by synthesizing the speech data Va and Vb to user terminal 2b via channel Ch2(1), and also converts speech data Vc (original speech) (see FIG. 8) or speech data Vc2 after speech adjustment (see FIG. 10) into text separately from the synthesized speech data Vm2 and the speech data Vc or speech data Vc2 after speech adjustment. This improves the accuracy of the text conversion.

[0089] The text conversion process may be performed by the speech processing device 1. For example, when the speech processing device 1 acquires speech data Va, Vb, and Vc from the user terminals 2a, 2b, and 2c, it converts the synthetic speech data Vm2 (Va'+Vb') into text, and performs a speech adjustment process on the speech data Vc to be adjusted, and then performs the text conversion process. Then, the speech processing device 1 outputs text information corresponding to the synthetic speech data Vm2 and the speech data Vc to the user terminal 2b. That is, the adjustment processing unit 114 of the present disclosure may perform at least one of volume adjustment, frequency adjustment, speed adjustment, translation, and text conversion.

[0090] In each of the above-described embodiments, the user terminal 2 outputs an adjustment request to the voice processing device 1 when receiving an adjustment request from a user. In another embodiment, the user terminal 2 may analyze voice data acquired from the voice processing device 1 to determine whether or not it is necessary to perform voice adjustment processing, and output an adjustment request instruction to the voice processing device 1 when it determines that it is necessary to perform voice adjustment processing. The control unit 11 of the voice processing device 1 may accept a voice data adjustment request when the user terminal 2 analyzes multiple voice data and outputs an instruction to request voice data adjustment. For example, the user terminal 2 compares the volume, frequency, speed, etc. of the multiple voice data to determine whether or not it is necessary to perform voice adjustment processing. In another embodiment, the voice processing device 1 may determine whether or not it is necessary to perform voice adjustment processing based on the voice data acquired from the user terminal 2.

[0091] [Disclosure Note] The following is a summary of the disclosure extracted from the above-described embodiment. Note that the configurations and processing functions described in the following supplementary notes can be selected and combined as desired.

[0092] <Appendix 1> an acquisition processing unit that acquires a plurality of pieces of voice data corresponding to a user's speech from each of a plurality of voice devices; a reception processing unit that receives a request to adjust first audio data among the plurality of audio data acquired by the acquisition processing unit; an adjustment processing unit that performs an audio adjustment process on the first audio data in response to the adjustment request; an output processing unit that outputs the adjusted audio data obtained by the audio adjustment process and second audio data, which is the audio data excluding the first audio data, from a first audio device that is a requester of the adjustment request, and outputs the audio data from a second audio device, which is the audio device excluding the first audio device, from the audio devices; A voice processing system comprising:

[0093] <Appendix 2> the reception processing unit receives a request to adjust the first voice data in response to an instruction to request adjustment of the first voice data from a user of the first voice device; 10. The speech processing system of claim 1.

[0094] <Appendix 3> the reception processing unit receives a request to adjust the first voice data when the first voice device analyzes the plurality of voice data and outputs an instruction to request adjustment of the first voice data; 10. The speech processing system of claim 1.

[0095] <Appendix 4> the output processing unit outputs first synthetic voice data obtained by synthesizing the adjusted voice data and the second voice data to the first voice device, and outputs second synthetic voice data obtained by synthesizing the plurality of voice data to the second voice device. 4. A speech processing system according to any one of Supplementary notes 1 to 3.

[0096] <Appendix 5> the output processing unit outputs the first synthetic voice data to the first voice device through a first channel, and outputs the second synthetic voice data to the second voice device through a second channel; 5. The speech processing system of claim 4.

[0097] <Appendix 6> the adjustment processing unit included in the first audio device performs the audio adjustment process on the first audio data; the output processing unit included in the first audio device causes the adjusted audio data and the second audio data to be output from the first audio device; 6. A speech processing system according to any one of Supplementary notes 1 to 5.

[0098] <Appendix 7> the adjustment processing unit included in the first audio device performs the audio adjustment processing on the first audio data before audio processing including gain adjustment and noise removal is performed; 7. The speech processing system of claim 6.

[0099] <Appendix 8> the output processing unit outputs the first audio data or the adjusted audio data and the second audio data to the first audio device via a first channel, and outputs synthesized audio data obtained by synthesizing the plurality of audio data to the second audio device via a second channel. 6. A speech processing system according to any one of Supplementary notes 1 to 5.

[0100] <Appendix 9> the adjustment processing unit performs at least one of volume adjustment, frequency adjustment, speed adjustment, translation, and text conversion; 9. A speech processing system according to any one of Supplementary notes 1 to 8. [Explanation of symbols]

[0101] 1: Audio processing device 2: User device 11: Control section 12: Storage section 13: Communications Department 21: Control unit 22: Storage section 23: Operation display section 24:Mike 25: Speaker 26: Communications Department 100: Audio processing system 111: Acquisition processing unit 112: Audio processing unit 113: Reception processing unit 114: Adjustment processing unit 115: Output processing section Ch1: Channel Ch2: Channel

Claims

1. an acquisition processing unit that acquires a plurality of pieces of voice data corresponding to a user's speech from each of a plurality of voice devices; a reception processing unit that receives a request to adjust first audio data among the plurality of audio data acquired by the acquisition processing unit; an adjustment processing unit that performs an audio adjustment process on the first audio data in response to the adjustment request; an output processing unit that outputs the adjusted audio data obtained by the audio adjustment process and second audio data, which is the audio data excluding the first audio data, from a first audio device that is a requester of the adjustment request, and outputs the audio data from a second audio device, which is the audio device excluding the first audio device, from the audio devices; A voice processing system comprising:

2. the reception processing unit receives a request to adjust the first voice data in response to an instruction to request adjustment of the first voice data from a user of the first voice device; The audio processing system of claim 1 .

3. the reception processing unit receives a request to adjust the first voice data when the first voice device analyzes the plurality of voice data and outputs an instruction to request adjustment of the first voice data; The audio processing system of claim 1 .

4. the output processing unit outputs first synthetic voice data obtained by synthesizing the adjusted voice data and the second voice data to the first voice device, and outputs second synthetic voice data obtained by synthesizing the plurality of voice data to the second voice device. The audio processing system of claim 1 .

5. the output processing unit outputs the first synthetic speech data to the first audio device through a first channel, and outputs the second synthetic speech data to the second audio device through a second channel; 5. The audio processing system of claim 4.

6. the adjustment processing unit included in the first audio device performs the audio adjustment process on the first audio data; the output processing unit included in the first audio device causes the adjusted audio data and the second audio data to be output from the first audio device; The audio processing system of claim 1 .

7. the adjustment processing unit included in the first audio device performs the audio adjustment processing on the first audio data before audio processing including gain adjustment and noise removal is performed; 7. The audio processing system of claim 6.

8. the output processing unit outputs the first audio data or the adjusted audio data and the second audio data to the first audio device via a first channel, and outputs synthesized audio data obtained by synthesizing the plurality of audio data to the second audio device via a second channel. The audio processing system of claim 1 .

9. the adjustment processing unit performs at least one of volume adjustment, frequency adjustment, speed adjustment, translation, and text conversion; The voice processing system according to any one of claims 1 to 8.

10. acquiring a plurality of pieces of voice data corresponding to a user's speech from each of a plurality of audio devices; receiving a request to adjust first audio data among the plurality of audio data; performing a sound adjustment process on the first sound data in response to the adjustment request; outputting the adjusted audio data obtained by the audio adjustment process and second audio data obtained by excluding the first audio data from the plurality of audio data from a first audio device that is a requester of the adjustment request, and outputting the plurality of audio data from a second audio device obtained by excluding the first audio device from the plurality of audio devices; An audio processing method executed by one or more processors.

11. acquiring a plurality of pieces of voice data corresponding to a user's speech from each of a plurality of audio devices; receiving a request to adjust first audio data among the plurality of audio data; performing a sound adjustment process on the first sound data in response to the adjustment request; outputting the adjusted audio data obtained by the audio adjustment process and second audio data obtained by excluding the first audio data from the plurality of audio data from a first audio device that is a requester of the adjustment request, and outputting the plurality of audio data from a second audio device obtained by excluding the first audio device from the plurality of audio devices; An audio processing program for execution by one or more processors.

Citation Information

Patent Citations

  • Communication conference device having sound volume adjustment function for each speaker

    JP2015069136A