Information Processing Method, Information Processing Apparatus, and Program

The information processing method allows users to see the effect of voice clarification by displaying spectrum changes, addressing the lack of feedback in existing voice processing systems.

JP7710637B1Active Publication Date: 2025-07-18RADIUS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025502631
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-07-18
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

The existing voice processing techniques, as described in Patent Document 1, do not allow the speaker or listener to confirm the effect of voice clarification, leading to a need for a method to present the user with the impact of voice processing.

Method used

An information processing method involving a processor that acquires voice data, generates clarified voice data through processing, and displays a visualization image showing the spectrum changes before and after clarification, allowing users to see the effect of voice processing.

Benefits of technology

Enables users to visually confirm the effect of voice clarification, enhancing user experience by making the impact of voice processing transparent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710637000001
    Figure 0007710637000001
  • Figure 0007710637000002
    Figure 0007710637000002
  • Figure 0007710637000003
    Figure 0007710637000003
Patent Text Reader

Abstract

In order to present the user with the effect of clarification by voice processing for clarifying the voice, at least one processor executes a first acquisition process (S104) for acquiring voice data indicating the voice of the speaker, a second acquisition process (S105) for acquiring clarified voice data generated by voice processing for clarifying the voice with respect to the voice data, and a visualization control process (S108) for displaying a visualization image obtained by visualizing the voice processing based on the spectrum of the voice data and the spectrum of the clarified voice data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technique for displaying information related to voice processing.

Background Art

[0002] Patent Document 1 discloses a technique for performing voice processing to clarify the voice of the speaker in a call. According to this technique, a voice signal of the speaker subjected to voice processing so that the user on the receiving side feels clear is output to the receiving side.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Here, in the technique described in Patent Document 1, there is a problem that the speaker on the transmitting side cannot confirm the effect of clarification on the receiving side due to voice processing on their own voice. In addition, there is a problem that the user on the receiving side cannot confirm the effect of clarification in the audible voice. Therefore, a technique for presenting the user with the effect of clarification by voice processing for clarifying the voice is required.

[0005] One aspect of the present disclosure aims to realize a technique for presenting the user with the effect of clarification by voice processing for clarifying the voice.

Means for Solving the Problems

[0006] In order to solve the above problems, an information processing method according to an aspect of the present disclosure includes at least one processor performing a first acquisition process of acquiring voice data indicating the voice of a speaker, a second acquisition process of acquiring clarified voice data generated by voice processing for clarifying the voice with respect to the voice data, and a visualization control process for displaying a visualization image obtained by visualizing the voice processing based on the spectrum of the voice data and the spectrum of the clarified voice data.

[0007] In order to solve the above problems, an information processing apparatus according to an aspect of the present disclosure includes the at least one processor, and the at least one processor executes each process included in the above-described information processing method.

[0008] In order to solve the above problems, a program according to an aspect of the present disclosure causes the at least one processor to execute each process included in the above-described information processing method. Note that a non-transitory computer-readable recording medium recording the program also falls within the scope of the present disclosure.

Effect of the Invention

[0009] According to an aspect of the present disclosure, it is possible to present to the user the effect of clarification by voice processing for clarifying the voice.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

MODE FOR CARRYING OUT THE INVENTION

[0011] 〔Embodiment 1〕 Hereinafter, an information processing system 100 according to Embodiment 1 of the present disclosure will be described in detail with reference to the drawings.

[0012] (Configuration of Information Processing System 100) FIG. 1 is a block diagram showing the configuration of the information processing system 100. As shown in FIG. 1, the information processing system 100 includes a user terminal 1 used by user U1 and a user terminal 2 used by user U2. The user terminal 1 and the user terminal 2 can be connected via a network N for user U1 and user U2 to make a call. Although FIG. 1 shows one user terminal 1 and one user terminal 2 respectively, the number of each may be plural.

[0013] Here, a call refers to at least real-time exchange of voice between multiple users via the network N, which may be a voice call that exchanges only voice, or a video call that exchanges voice and video. Also, the call method may be a circuit switching method via a public switched telephone network or the like, or a packet switching method via an IP (Internet Protocol) network or the like, but is not limited thereto.

[0014] The network N may be, for example, a public switched telephone network, a private line network, a mobile communication network, an IP network, a wireless or wired LAN (Local Area Network), or a combination of some or all of these. However, the network N only needs to be a network that can connect the user terminal 1 and the user terminal 2 to make a call in a predetermined call method, and is not limited to the above examples.

[0015] The user terminal 1 is a terminal used by user U1. The user terminal 1 has a function of connecting to the user terminal 2 via the network N by a predetermined call method (for example, the circuit switching method or the packet switching method described above, etc.). That is, the user U1 can make a call to the user U2 using the user terminal 1. The user terminal 1 is an aspect of the information processing apparatus described in the claims, and has a clarifying function of clarifying its own or the other party's voice in a call, and a visualization function of visualizing the clarification of the voice. The user terminal 1 is a computer including at least one processor and a memory. The user terminal 1 may be, for example, a mobile phone, a smartphone, a tablet, a smartwatch, a notebook computer (laptop personal computer), a desktop computer (desktop personal computer), etc., but is not limited thereto.

[0016] The user terminal 2 is a terminal used by user U2. The user terminal 2 has a function of connecting to the user terminal 1 via the network N by the same call method as the user terminal 1. That is, the user U2 can make a call to the user U1 using the user terminal 2. Hereinafter, the user terminal 2 will be described as not having the above-described clarifying function and clarifying visualization function, but this is not limited, and it may have these functions. The user terminal 2 is a computer including at least one processor and a memory. The user terminal 2 may be, for example, a mobile phone, a smartphone, a tablet, a smartwatch, a notebook computer, a desktop computer, etc., but is not limited thereto.

[0017] (Configuration of User Terminal 1) FIG. 2 is a block diagram showing the functional configuration of the user terminal 1. As shown in FIG. 2, the user terminal 1 includes a control unit 110, a storage unit 120, a display unit 130, an input unit 140, a voice input unit 150, a voice output unit 160, and a communication unit 170. The control unit 110 is realized, for example, by a processor executing a program stored in a memory, and comprehensively controls each unit of the user terminal 1. Details of each functional block included in the control unit 110 will be described later. The storage unit 120 is constituted by, for example, a memory, and stores various data and programs used by the control unit 110.

[0018] The display unit 130 displays an image generated by the control unit 110. The display unit 130 may include, for example, a liquid crystal display, an organic EL (Electro luminescence) display, or the like, but is not limited thereto. The input unit 140 receives an operation of the user U1, acquires input information indicated by the operation, and outputs the input information to the control unit 110. The input unit 140 may include, for example, a mouse, a keyboard, a touch pad, or a combination of some or all of these, but is not limited thereto. Further, the display unit 130 and the input unit 140 may be integrally formed as a touch panel.

[0019] The voice input unit 150 converts an externally input voice signal into voice data which is a digital signal, and outputs the voice data to the control unit 110. The voice input unit 150 may include, for example, a microphone and an AD (Analog to Digital) converter, but is not limited thereto. The voice output unit 160 converts voice data input from the control unit 110 into a voice signal which is an analog signal, and outputs the voice signal to the outside as voice. The voice output unit 160 may include, for example, a DA converter and a speaker, but is not limited thereto.

[0020] The communication unit 170 is connected to the network N to communicate with the outside. The communication unit 170 transmits the information input from the control unit 110 via the network N. Also, the communication unit 170 outputs the information received via the network N to the control unit 110.

[0021] Note that part or all of the storage unit 120, the display unit 130, the input unit 140, the voice input unit 150, the voice output unit 160, and the communication unit 170 may be connected as peripheral devices instead of being built into the user terminal 1.

[0022] (Functional blocks of the control unit 110) As shown in FIG. 2, the control unit 110 includes a call application unit 111, a first acquisition unit 112, a second acquisition unit 113, a voice output control unit 114, a clarification UI (User Interface) unit 115, and a visualization control unit 116. The first acquisition unit 112, the second acquisition unit 113, the voice output control unit 114, the clarification UI unit 115, and the visualization control unit 116 may constitute a clarification application that extends the call function by the call application unit 111.

[0023] The call application unit 111 provides a function for making a call according to a predetermined call method. For example, the call application unit 111 connects to the call application unit 211 of the user terminal 2 via the network N by a predetermined call method in response to an operation of the user U1 for starting a call. The operation for starting the call may include, for example, an operation for designating a call destination and an operation for instructing the transmission of a call request to the call destination. Also, the operation for starting the call may include, for example, an operation for responding to a received call request. Also, the call application unit 111 transmits voice data indicating the voice of the user U1 input from the voice input unit 150 to the connected user terminal 2. Also, the call application unit 111 receives voice data indicating the voice of the user U2 from the user terminal 2 and outputs the voice data from the voice output unit 160. Also, the call application unit 111 ends the connection with the user terminal 2 in response to an operation of the user U1 for instructing the end of the call.

[0024] The first acquisition unit 112 executes a first acquisition process. The first acquisition process is a process of acquiring voice data indicating the voice of the speaker. Here, the voice data to be acquired may be the voice of the speaker in a call made by connecting a plurality of user terminals 1 and 2. Hereinafter, "voice data indicating the voice of the speaker" will also be referred to as "the speaker's voice data". Note that the voice data is acquired in real time according to the progress of the speech by the speaker.

[0025] For example, when user U1 speaks in a call between user U1 and user U2, user U1 is the speaker and user U2 is the listener. In this case, the first acquisition unit 112 acquires the voice data of user U1 (the speaker) using the user terminal 1 via the voice input unit 150.

[0026] Also, for example, when user U2 speaks in a call between user U1 and user U2, user U2 is the speaker and user U1 is the listener. In this case, the first acquisition unit 112 acquires the voice data of user U2 (the speaker) received by the call application unit 111.

[0027] The second acquisition unit 113 executes a second acquisition process. The second acquisition process is a process of acquiring clarified voice data generated by voice processing for clarifying the voice with respect to the speaker's voice data. Here, the clarified voice data may be acquired in real time. That the clarified voice data is acquired in real time means that the clarified voice data generated from the voice data is acquired in parallel with the process of acquiring the voice data according to the progress of the speech.

[0028] For example, the second acquisition unit 113 may generate clarified voice data by executing voice processing for clarifying the voice on the voice data. For example, when the second acquisition unit 113 executes the voice processing in real time according to the progress of the speech by the speaker, the clarified voice data is acquired in real time.

[0029] Further, for example, the second acquisition unit 113 may acquire the clarified voice data from a voice processing device that performs the voice processing. Such a voice processing device may be configured by a voice processing circuit that performs voice processing on a voice signal that is an analog signal, or may be configured by a computer that executes voice processing on voice data that is a digital signal. When the voice processing device executes the voice processing in real time according to the progress of the speech by the speaker, the clarified voice data is acquired in real time.

[0030] In each of the embodiments described below, an example in which the second acquisition unit 113 (or the second acquisition unit 313 described later) executes voice processing to acquire the clarified voice data will be mainly described. However, the description of each embodiment is similarly applicable when the second acquisition unit 113 (or 313) acquires the clarified voice data from an external voice processing device.

[0031] For example, the voice processing for clarifying the voice may be a process of amplifying higher-order formant components including at least second-order formant components in the voice data. Here, in the spectral analysis of the voice, a plurality of peak frequencies appear at integer multiples of points. These plurality of peak frequencies are called the first formant, second formant, third formant, fourth formant, etc. in ascending order of frequency. Although each peak frequency varies depending on the skeleton of the speaker, etc., generally, it is known that in order to understand the language, it is necessary to surely hear the components of the first to fourth formants. Also, users with insufficient hearing often have a reduced hearing in the band including higher-order formant components such as the second to fourth formants. Therefore, by amplifying the higher-order formant components including the second formant, the voice is clarified so that users with insufficient hearing can feel it is clear.

[0032] For example, the second acquisition unit 113 may perform a process of amplifying the level of a band equal to or higher than the lower frequency in the audio data. In this case, the lower frequency is determined such that a higher-order harmonic component including at least a second-order harmonic component is included in the band equal to or higher than the lower frequency. The second acquisition unit 113 may further determine an upper frequency and perform a process of amplifying the level of a band equal to or higher than the lower frequency and equal to or lower than the upper frequency in the audio data. In this case, the lower frequency and the upper frequency are determined such that a higher-order harmonic component including at least a second-order harmonic component is included in the band. As an example, the lower frequency may be 400 Hz and the upper frequency may be 5 KHz. There is a high possibility that the second to fourth harmonic components in general human speech are included in the band from 400 Hz to 5 KHz. However, the lower frequency and the upper frequency are not limited to the above example. Further, for example, as the lower frequency, the frequency of the first-order harmonic component detected in real time in the audio data may be applied. Note that the audio processing for clarifying the audio is not limited to the above-described processing, and known processing can be applied.

[0033] In addition, the degree of clarification of the audio in the audio processing can be set according to the first operation. The first operation is an operation for setting the degree of clarification, and in the present embodiment, it is performed by the user U1. The degree of clarification may be, for example, a gain for amplifying the level of a band equal to or higher than the lower frequency and equal to or lower than the upper frequency described above. The first operation is received by a clarification UI unit 115 described later. For example, the degree of clarification of the audio may be discretely set as a plurality of levels. Also, the degree of clarification of the audio may be continuously set.

[0034] The voice output control unit 114 executes voice output control processing. The voice output control processing is processing for controlling the voice indicated by the clarified voice data to be output from the user terminal of the recipient among a plurality of user terminals (for example, a plurality of user terminals including user terminal 1 and user terminal 2). For example, when user U1 speaks, user U1 is the speaker and user U2 is the recipient. Therefore, when the clarified voice data of user U1 is generated, the voice output control unit 114 transmits the clarified voice data to user terminal 2 in order to output the clarified voice data from user terminal 2. Also, for example, when user U2 speaks during a call between user U1 and user U2, user U2 is the speaker and user U1 is the recipient. Therefore, when the clarified voice data of user U2 is generated, the voice output control unit 114 controls the voice output unit 160 to output the clarified voice data.

[0035] The clarification UI unit 115 receives a first operation for setting the degree of clarification of the voice in voice processing. For example, the first operation may include an operation for setting the degree of clarification for the transmitted voice, or may include an operation for setting the degree of clarification for the received voice. Here, the transmitted voice refers to the voice of user U1 when user U1 is the speaker. Also, the received voice refers to the voice that user U1 hears when user U2 is the speaker.

[0036] For example, the first operation may include an operation of selecting any one of a plurality of levels as the degree of clarification. The plurality of levels may be, for example, three levels such as weak, medium, and strong. Also, the degree of clarification may include not performing clarification. In this case, the plurality of levels may be, for example, four levels such as off, weak, medium, and strong, or may be two levels of on and off. However, the names of each level and the number of levels are not limited to this.

[0037] Further, for example, the first operation may include an operation of specifying any degree of clarification included in a continuous range. For example, any degree of clarification may be represented by any numerical value from 1 to m, where the degree of clarification "off" is set to 1 times and the degree of clarification "strong" is set to m times (m is a number greater than 1). Further, the first operation may be an operation on an operation object (e.g., a slider object, a volume object, etc.) continuously corresponding to values from 1 to m. Information indicating the degree of clarification for the transmitted voice or received voice received by the first operation is stored in the storage unit 120 as setting information.

[0038] The visualization control unit 116 executes visualization control processing. The visualization control processing is processing for displaying a visualization image that visualizes voice processing based on the spectrum of voice data and the spectrum of the clarified voice data. Thereby, since the change in the spectrum before and after the voice is clarified is visualized, the user U1 can confirm the effect of voice clarification.

[0039] For example, assume that the speaker in a call is the user U1 and voice processing is performed on the transmitted voice, which is the voice of the user U1. In this case, the visualization control unit 116 may display the visualization image on the display unit 130 of the user terminal 1 of the user U1 (the speaker) among a plurality of user terminals (e.g., a plurality of user terminals including the user terminal 1 and the user terminal 2). Thereby, the user U1 (the speaker) can confirm the effect on the user U2 (the listener), who is the call partner, due to the clarification of the transmitted voice, which is his / her own voice.

[0040] Further, for example, assume that the speaker in a call is the user U2 and voice processing is performed on the received voice, which is the voice of the user U2. In this case, the visualization control unit 116 may display the visualization image on the display unit 130 of the user terminal 1 of the user U1 (the listener) among a plurality of user terminals (e.g., a plurality of user terminals including the user terminal 1 and the user terminal 2). Thereby, the user U1 (the listener) can confirm the effect of the clarification of the received voice.

[0041] Further, the visualization control unit 116 may update the visualization image in real time according to the progress of the speech by the speaker. Here, since the clarified speech data is generated in real time from the speech data that changes according to the progress of the speech, the spectrum of the speech data and the spectrum of the clarified speech data change in real time. Therefore, by updating the visualization image in real time based on these two spectra, the visualization image is displayed as a moving image.

[0042] Also, for example, the visualization image is displayed in a display mode for comparing these two spectra. For example, the visualization image may include an image in which an image showing the spectrum of the speech data and an image showing the spectrum of the clarified speech data are superimposed. Also, for example, the visualization image may include an image showing the difference between the spectrum of the speech data and the spectrum of the clarified speech data. Also, for example, the visualization image may include an image in which an image showing the spectrum of the speech data and an image showing the spectrum of the clarified speech data are arranged side by side. However, the display mode of the visualization image is not limited to the above-described examples. Thereby, since it is visualized which band in the spectrum of the speech is amplified by how much before and after clarification, the user U1 can confirm the effect of the clarification of the speech.

[0043] Also, the visualization image may further include a label image indicating each formant component. Thereby, since it is visualized how much each higher-order formant component in the spectrum of the speech is amplified, the user U1 can confirm the effect of the clarification of the speech.

[0044] In addition, the visualization control unit 116 displays a visualization image that reflects the degree of clarification set according to the first operation. Here, since the spectrum of the clarification voice data changes when the degree of clarification changes, the visualization image also changes. That is, the visualization image is updated while reflecting the set degree of clarification. Thereby, when the user U1 changes the setting of the degree of clarification, the user U1 can confirm how the effect of voice clarification changes.

[0045] (Configuration of User Terminal 2) FIG. 3 is a block diagram showing the functional configuration of the user terminal 2. As shown in FIG. 3, the user terminal 2 includes a control unit 210, a storage unit 220, a display unit 230, an input unit 240, a voice input unit 250, a voice output unit 260, and a communication unit 270. Each of these units will be described in the same manner as the corresponding functional blocks provided in the user terminal 1.

[0046] (Functional Blocks of Control Unit 210) As shown in FIG. 3, the control unit 210 includes a call application unit 211. The call application unit 211 provides a function for making a call according to the same call method as the call application unit 111. For details of the call application unit 211, in the description of the call application unit 111, the user U1 and the user U2, the user terminal 1 and the user terminal 2, and the corresponding functional blocks with the same name in the user terminal 1 and the user terminal 2 are read mutually, and the description is made in the same manner.

[0047] (Flow of Information Processing Method S1) The information processing system 100 configured as described above executes an information processing method S1. In the information processing method S1, voice processing for clarifying the transmitted voice of the user U1 is performed when the user U1 is the speaker. The information processing method S1 is a method of presenting a visualization image obtained by visualizing the voice processing for the transmitted voice of the user U1 to the user U1. FIG. 4 is a flowchart for explaining the flow of the information processing method S1. As shown in FIG. 4, the information processing method S1 includes steps S101 to S111.

[0048] In step S101, the call application unit 111 of the user terminal 1 connects to the call application unit 211 of the user terminal 2 based on the operation of the user U1 to start a call with the user U2. Also, in step S102, the call application unit 211 of the user terminal 2 connects to the call application unit 111 of the user terminal 1 based on the operation of the user U2. Thereby, a call between the user U1 and the user U2 is started. Note that steps S101 and S102 are not limited to being executed in this order, and the reverse order, or a part or all of them may be executed in parallel. Also, steps S101 and S102 do not limit which of the user U1 and the user U2 requests a call to the other party.

[0049] In step S103, the clarification UI unit 115 of the user terminal 1 receives a first operation from the user U1 to set the clarification level of the transmitted voice. The clarification level is an example of the degree of clarification, and refers to the stage when the degree of clarification is set discretely. Hereinafter, an example where the degree of clarification can be set discretely will be mainly described, but it is not limited to this. The setting information including the clarification level of the transmitted voice set by the first operation is stored in the storage unit 120. Note that the timing of executing step S103 is not limited to after the start of the call (after step S101), and it may be executed before the start of the call (before step S101). Also, the timing is not limited to immediately after the start of the call, and it may be executed at any point during the call (any point after step S104 described below). Also, if the first operation is not received, the pre-stored setting information may be applied.

[0050] Step S104 is an example of the first acquisition process. In step S104, the first acquisition unit 112 acquires the voice data of the user U1 via the voice input unit 150. That is, it is assumed that the user U1 spoke during the call.

[0051] Step S105 is an example of the second acquisition process. In step S105, the second acquisition unit 113 executes voice processing on the voice data of user U1 according to the setting information. As a result, clarified voice data clarified according to the clarification level indicated by the setting information is acquired. When the setting information indicates the clarification level "off" of the transmitted voice, the voice processing is not executed and the clarified voice data is not acquired.

[0052] In step S106, the voice output control unit 114 transmits the clarified voice data to the user terminal 2. When the clarified voice data is not generated, the voice data of user U1 is transmitted instead.

[0053] In step S107, the call application unit 211 of the user terminal 2 receives the clarified voice data. Also, the call application unit 211 outputs the voice indicated by the received clarified voice data via the voice output unit 260. As a result, the clarified voice of user U1 can be heard by user U2. That is, user U1 can have user U2, who is the call partner, hear the clarified voice of himself / herself even if the user terminal 2 used by user U2 does not have a clarification function. When voice data is received instead of the clarified voice data, the unclarified voice indicated by the voice data is output.

[0054] Step S108 is an example of the visualization control process. In step S108, the visualization control unit 116 of the user terminal 1 generates a visualization image based on the spectrum of the voice data of user U1 and the spectrum of the clarified voice data. Also, the visualization control unit 116 displays the visualization image on the display unit 130. Here, steps S104 to S108 are executed in real time according to the progress of the speech by user U1. Therefore, the visualization image displayed in step S108 is a moving image that changes according to the progress of the speech. User U1 can confirm the effect on the user U2 side due to the clarification of his / her own voice by visually recognizing the visualization image.

[0055] Thereafter, when user U2 speaks, user terminal 1 receives the voice data of user U2 from user terminal 2 and outputs the received voice data via the voice output unit 160. As a result, user U1 can hear the voice of user U2 that has not been clarified. Note that it is also possible to configure such that steps S204 to S208 of information processing method S2 described later are executed when user U2 speaks. In this case, user U1 can hear the clarified voice of user U2. Further, when user U1 speaks, steps S104 to S108 are repeated.

[0056] Also, when step S103 is executed again during a call and the clarification level of the transmitted voice is changed, the changed clarification level is reflected in the clarified voice data acquired in step S105 and the voice audible to user U2 in step S107. Further, the changed clarification level is reflected in the visualization image presented to user U1 in step S108.

[0057] In step S110, the call application unit 111 of user terminal 1 terminates the connection with the call application unit 211 of user terminal 2 based on the operation of user U1. Also, in step S111, the call application unit 211 of user terminal 2 terminates the connection with the call application unit 111 of user terminal 1. As a result, the call between user U1 and user U2 ends. Note that steps S110 and S111 are not limited to being executed in this order, and the reverse order, or part or all of them may be executed in parallel. Also, the operation for ending the call is not limited to user U1, and user U2 may perform it, or both may perform it.

[0058] (Screen example) FIG. 5 is a diagram schematically showing a screen example G1 displayed on the display unit 130 of user terminal 1 in information processing method S1. Screen example G1 provides a user interface related to the clarification function and visualization function of one's own voice (transmitted voice) in a call. As shown in FIG. 5, screen example G1 includes information G11 to G13, operation objects G14-1 to G14-4, G15, and a visualization image G16.

[0059] The information G11 (for example, "on a call with Mr. U2") indicates that a call is currently in progress and the call partner. The information G12 (for example, "1:02") indicates the elapsed time of the call. The information G13 (for example, "Your voice is being clarified and can be heard by Mr. U2") indicates that the voice of the user U1 using the user terminal 1 is being clarified.

[0060] The operation objects G14-1 "weak", G14-2 "medium", G14-3 "strong", and G14-4 "off" accept an operation (an example of the first operation) for selecting the clarification level regarding the transmitted voice. In this example, any one of the four levels of clarification (off, weak, medium, strong) can be selected. Also, G14-2 "medium" which is highlighted among the operation objects G14-1 to G14-4 indicates that the currently selected clarification level is "medium". Note that in FIG. 5, as a highlighting mode, a thick frame and gray filling are shown, but it is not limited to this. When an operation on any of the operation objects G14-1 to G14-4 is accepted, the setting information including the corresponding clarification level is stored in the storage unit 120. Also, the operation objects G14-1 "weak", G14-2 "medium", G14-3 "strong", and G14-4 "off" can accept operations at any time before and during the call. Also, multiple operations may be accepted, and in that case, the setting information based on the most recent operation is stored.

[0061] The operation object G15 accepts an operation of the user U1 instructing the end of the call. When an operation on the operation object G15 is accepted, the call application unit 111 of the user terminal 1 ends the connection with the call application unit 211 of the user terminal 2.

[0062] The visualization image G16 includes an image in which the spectrum G16-1 before the voice of the user U1, who is the speaker, is clarified and the spectrum G16-2 after clarification are superimposed. The visualization image G16 is a moving image that changes according to the progress of the speech. With the visualization image G16, the user U1 can superimpose and compare the spectra G16-1 and G16-2 before and after their own voice is clarified, and thus can confirm the effect that their own voice is clearly heard by the user U2, who is the other party in the call.

[0063] Here, the screen example G1 may include the visualization images G16a to G16d shown in FIG. 6 instead of the visualization image G16. FIG. 6 is a diagram schematically showing variations of the visualization image.

[0064] As shown in FIG. 6, the visualization image G16a is an image in which label images G16a-1 to G16a-4 indicating each formant component are superimposed on the visualization image G16. With the visualization image G16a, the user U1 can more clearly confirm that the first formant component, which is highly likely to be heard by the user U2, the other party in the call, is not amplified, and the higher-order formant components, which may be difficult to hear, are amplified.

[0065] Also, as shown in FIG. 6, the visualization image G16b is an image showing the difference spectrum G16b-1 between the spectrum G16-1 before clarification and the spectrum G16-2 after clarification. With the visualization image G16b, the user U1 can more clearly confirm the difference in the degree to which the higher-order formant components in their own voice are amplified.

[0066] Also, as shown in FIG. 6, the visualization image G16c is an image in which the visualization image G16 and the visualization image G16b are arranged side by side. That is, the visualization image G16c includes an image in which the spectra G16-1 and G16-2 before and after clarification are superimposed, and an image showing the difference between the spectra G16-1 and G16-2. With the visualization image G16b, the user U1 can more clearly confirm the effect on the user U2 side due to the clarification of their own voice.

[0067] Also, as shown in FIG. 6, the visualization image G16d is an image in which the visualization image G16d-1 and the visualization image G16d-2 are arranged side by side. The visualization image G16d-1 shows the spectrum G16-1 before being clarified. The visualization image G16d-2 shows the spectrum G16-2 after being clarified. With the visualization image G16d, the user U1 can compare the spectra G16-1 and G16-2 before and after being clarified side by side, so that the user U1 can clearly confirm the effect on the user U2 side due to the clarification of his own voice.

[0068] Note that the screen example G1 may include not only the visualization images G16, G16a to G16d, but also visualization images in other forms, or combinations of some or all of these.

[0069] (Flow of the information processing method S2) Also, the information processing system 100 executes the information processing method S2. In the information processing method S2, voice processing for clarifying the received voice that can be heard by the user U1 when the user U2 is the speaker (in other words, the voice of the user U2) is performed. The information processing method S2 is a method of presenting a visualization image obtained by visualizing the voice processing for the received voice to the user U1. FIG. 7 is a flowchart for explaining the flow of the information processing method S2. As shown in FIG. 7, the information processing method S2 includes steps S201 to S211.

[0070] Steps S201 and S202 are steps for starting a call between the user U1 and the user U2. Since steps S201 and S202 are explained in the same way as steps S101 and S102, detailed explanations will not be repeated.

[0071] In step S203, the clarification UI unit 115 of the user terminal 1 receives a first operation from the user U1 to set the clarification level of the received voice. The setting information including the clarification level of the received voice set by the first operation is stored in the storage unit 120. Details of the clarification level of the received voice are the same as those of the clarification level of the transmitted voice. Note that the timing of executing step S203 is not limited to after the start of the call (after step S201), and it may be executed before the start of the call (before step S201). The timing is not limited to immediately after the start of the call, and it may be executed at any point during the call (any point after step S204 described below). Also, when the first operation is not received, the pre-stored setting information may be applied.

[0072] In step S204, the call application unit 211 of the user terminal 2 acquires the voice data of the user U2 via the voice input unit 250. That is, it is assumed that the user U2 spoke during the call. Also, the call application unit 211 transmits the acquired voice data of the user U2 to the user terminal 1.

[0073] Step S205 is an example of the first acquisition process. In step S205, the first acquisition unit 112 acquires by receiving the voice data of the user U2.

[0074] Step S206 is an example of the second acquisition process. In step S206, the second acquisition unit 113 executes voice processing on the voice data of the user U2 according to the setting information. Thereby, the clarified voice data clarified according to the clarification level indicated by the setting information is acquired. Note that when the setting information indicates the received voice clarification level "off", the voice processing is not executed and the clarified voice data is not acquired.

[0075] In step S207, the voice output control unit 114 outputs the voice indicated by the clarified voice data via the voice output unit 160. As a result, the clarified voice of the user U2 can be heard by the user U1. That is, the user U1 can clarify and hear the voice of the call partner, the user U2, by means of the clarification function of the user terminal 1 that the user U1 is using. If the clarified voice data has not been generated, the voice indicated by the voice data of the user U2 is output.

[0076] Step S208 is an example of the visualization control process. In step S208, the visualization control unit 116 of the user terminal 1 generates a visualization image based on the spectrum of the voice data of the user U2 and the spectrum of the clarified voice data. Further, the visualization control unit 116 displays the visualization image on the display unit 130. Here, steps S204 to S208 are executed in real time according to the progress of the speech by the user U2. Therefore, the visualization image displayed in step S208 is a moving image that changes according to the progress of the speech. The user U1 can confirm the effect that the voice of the call partner, the user U2, is clarified by visually recognizing the visualization image.

[0077] Thereafter, when the user U1 speaks, the user terminal 1 transmits the voice data of the user U1 acquired via the voice input unit 150 to the user terminal 2. Further, the user terminal 2 outputs the received voice data via the voice output unit 260. As a result, the unclarified voice of the user U1 can be heard by the user U2. It is also possible to configure such that steps S104 to S108 of the information processing method S1 described above are executed when the user U1 speaks. In this case, the clarified voice of the user U1 can be heard by the user U2. Further, when the user U2 speaks, steps S204 to S208 are repeated.

[0078] Also, when step S203 is executed again during a call and the clarification level of the received voice is changed, the changed clarification level is reflected in the clarified voice data obtained in step S206 and the voice audible to user U1 in step S207. Further, the changed clarification level is reflected in the visualization image presented to user U1 in step S208.

[0079] Steps S210 and S211 are steps for ending the call between user U1 and user U2. Since steps S210 and S211 are explained in the same way as steps S110 and S111, a detailed explanation will not be repeated.

[0080] (Screen example) FIG. 8 is a diagram schematically showing a screen example G2 displayed on the display unit 130 of the user terminal 1 in the information processing method S2. In addition to providing a user interface regarding its own voice (transmitted voice) similar to the screen example G1 shown in FIG. 5, the screen example G2 provides a user interface regarding the clarification function and visualization function of the voice of the call partner (received voice). The screen example G2 shown in FIG. 8 includes information G21 to G22, regions G23 and G25, and an operation object G27.

[0081] The information G21 and G22 indicate that a call is in progress and are explained in the same way as the information G11 and G12 in the screen example G1. The operation object G27 "End Call" is explained in the same way as the operation object G15 in the screen example G1.

[0082] Area G23 is an area that provides a user interface related to clarifying one's own voice. Information G231, operation objects G232-1 to G232-4, and visualization image G233 included in area G23 are explained in the same way as information G13, operation objects G14-1 to G14-4, and visualization image G16 in screen example G1. Note that in screen example G2, the operation object G232-4 "Off" is highlighted, that is, an example of turning off the clarification of one's own voice is shown. Therefore, the visualization image G233 includes the spectrum G23-3 of the voice of the user U1 (the speaker) before clarification and does not include the spectrum of the clarified voice.

[0083] Area G25 is an area that provides a user interface related to clarifying the voice of the call partner. Area G25 includes information G251, operation objects G252-1 to G252-4, and visualization image G253. Information G251 (as an example, "The voice of Mr. U2 is clarified") indicates that the voice of the user U2, who is the call partner, is clarified.

[0084] The operation objects G252-1 "Weak", G252-2 "Medium", G252-3 "Strong", and G252-4 "Off" accept an operation (an example of the first operation) for selecting a clarification level regarding the received voice. For the operation objects G252-1 to G252-2, they are similarly explained by replacing the transmitted voice with the received voice in the description of the operation objects G14-1 to G14-4 in the screen example G1. However, the number of levels of the clarification level regarding the received voice does not necessarily have to be the same as the number of levels of the clarification level regarding the transmitted voice. For example, the number of levels that can be set regarding the received voice may be more than the number of levels that can be set regarding the transmitted voice. Thereby, while clarifying one's own voice to the call partner, the degree of clarification can be set more finely for the voice that one should hear. Also, for example, the number of levels that can be set regarding the received voice may be less than the number of levels that can be set regarding the transmitted voice. Thereby, while finely clarifying one's own voice so that the call partner can hear it, the setting of the degree of clarification for the voice that one should hear can be easily performed.

[0085] Also, among the operation objects G252-1 to G252-4, the emphasized G252-1 "Weak" indicates that the currently selected clarification level is "Weak".

[0086] The visualization image G253 includes an image in which the spectrum G253-1 before the voice of the user U2 who is the call partner is clarified and the spectrum G253-2 after clarification are superimposed. The visualization image G253 is a moving image that changes according to the progress of the utterance. With the visualization image G253, the user U1 can superimpose and compare the spectra G253-1 and G253-2 before and after the voice of the call partner U2 is clarified, so that the user U1 can confirm the effect that the call partner U2 is clarified.

[0087] Here, the screen example G2 may include a visualization image in a display mode corresponding to the visualization images G16a to G16d shown in FIG. 6 instead of the visualization images G233 and G253. Note that the display mode of the visualization image included in the screen example G2 is not limited to the above-described example. For example, the visualization image is not limited to showing the spectrum as a continuous waveform, and may be shown in a mode of a histogram for each unit band (for example, a stacked graph like a graphic equalizer).

[0088] (Modification Example 1 of Embodiment 1) In the information processing system 100 according to Embodiment 1, an example in which the user terminal 2 does not have the clarification function has been mainly described. However, when the user terminal 2 has the clarification function, the second acquisition unit 113 in the user terminal 1 is modified as follows.

[0089] For example, it is assumed that voice processing for clarifying the uttered voice is executed in the user terminal 2 based on the operation of the user U2 (speaker). In this case, the second acquisition unit 113 of the user terminal 1 cannot accept the operation of the user U1 (listener) for executing the voice processing on the received voice, or does not execute the voice processing even if the operation of the user U1 (listener) is accepted.

[0090] For example, the second acquisition unit 113 may determine whether voice processing for clarifying the uttered voice has been executed in the user terminal 2 based on the spectrum of the received voice. Also, for example, the second acquisition unit 113 may determine whether voice processing for clarifying the uttered voice has been executed in the user terminal 2 based on the received uttered-side clarification flag. The uttered-side clarification flag is, for example, a flag transmitted to the receiving side together with voice data indicating the uttered voice or clarified voice data. For example, the user terminal 2 may transmit an uttered-side clarification flag indicating clarification on to the user terminal 1 together with the clarified voice data, and may transmit an uttered-side clarification flag indicating clarification off to the user terminal 1 together with the unclarified voice data. Note that the user terminal 1 may also be provided with a function of transmitting the uttered-side clarification flag. Thereby, the user terminal 2 can also operate based on the uttered-side clarification flag in the same manner as the user terminal 1.

[0091] Also, for example, assume that voice processing for clarifying the received voice is executed in the user terminal 2 based on the operation of the user U2 (the receiver). In this case, the second acquisition unit 113 of the user terminal 1 either cannot accept the operation of the user U1 (the speaker) for executing the voice processing on the uttered voice, or does not execute the voice processing even if the operation of the user U1 (the speaker) is accepted.

[0092] For example, the second acquisition unit 113 may determine whether voice processing for clarifying the received voice has been executed in the user terminal 2 based on the received-side clarification flag. The received-side clarification flag is, for example, a flag transmitted to the speaking side when voice processing for clarifying the received voice is performed. For example, the user terminal 2 may transmit a received-side clarification flag indicating "clarification on" to the user terminal 1 when performing voice processing for clarifying the received voice. In this case, when the second acquisition unit 113 of the user terminal 1 receives the received-side clarification flag indicating "clarification on", it determines that voice processing for clarifying the received voice has been executed in the user terminal 2. Also, when the second acquisition unit 113 of the user terminal 1 has not received the received-side clarification flag indicating "clarification on", it may determine that voice processing for clarifying the received voice has not been executed in the user terminal 2.

[0093] In this modification example, for example, the screen example G2 shown in FIG. 8 is modified as follows. For example, it is assumed that voice processing for clarifying the speaking voice has been determined to be executed in the user terminal 2. In this case, in the screen example G2 displayed on the user terminal 1, the operation objects G252-1 to G252-4 related to the clarification of the received voice may be hidden or may be in a state where operations are not accepted (for example, grayed out). Note that even in this case, the visualization image G253 may be in a displayed state. Also, for example, it is assumed that voice processing for clarifying the received voice has been determined to be executed in the user terminal 2. In this case, in the screen example G2 displayed on the user terminal 1, the operation objects G232-1 to G232-4 related to the clarification of the speaking voice may be hidden or may be in a state where operations are not accepted (for example, grayed out). Note that even in this case, the visualization image G253 may be in a displayed state.

[0094] In this modification example, since voice processing for doubly clarifying the voice on both the speaking side and the receiving side is not executed for the same speaking voice, there is an effect of suppressing the unnatural speaking voice from being heard by the listener.

[0095] (Modification Example 2 of Embodiment 1) In the information processing system 100 according to Embodiment 1, the second acquisition unit 113 is modified as follows. The second acquisition unit 113 executes voice processing for clarifying the voice of the speaker based on any setting information selected from the preset setting information according to the characteristics of each of the plurality of types of voices. The characteristics of the voice may be, for example, but not limited to, the characteristics of the language, the characteristics of the dialect, the characteristics of the gender, or the characteristics of the individual, etc. Here, according to the characteristics of the voice, the frequencies of the higher-order formant components including the second-order formant component to be amplified for clarification may be different.

[0096] Therefore, the setting information may include information indicating the frequency band of the higher-order formant component including the second-order formant component according to the characteristics of the voice. For example, the first setting information may include information indicating the frequency band according to the Japanese language characteristics, and the second setting information may include information indicating the frequency band according to the English language characteristics. Note that the number of types of voice characteristics for which the setting information is preset is not limited to two, and may be three or more. Also, the process of selecting any one from the plurality of setting information may be performed by the user's operation or by a computer. The selection of the setting information by the computer can be executed based on the characteristics obtained by analyzing the uttered voice in real time.

[0097] In this modification example, for example, the screen example G1 shown in FIG. 5 and the screen example G2 shown in FIG. 8 are modified as follows. For example, the screen examples G1 and G2 may include operation objects for selecting the setting information to be applied to the clarification of their own voice from a plurality of types of setting information (for example, setting information according to Japanese, English,... etc.). Also, for example, the screen example G2 may include an operation object for selecting the setting information to be applied to the clarification of the voice of the call partner from the plurality of types of setting information.

[0098] In this modification example, there is an effect that the voice of the speaker can be accurately clarified according to the characteristics of the voice.

[0099] 〔Embodiment 2〕 The user terminal 1A according to Embodiment 2 of the present disclosure will be described below. For convenience of explanation, members having the same functions as those described in the above embodiments are given the same reference numerals, and their descriptions will not be repeated.

[0100] The user terminal 1A is an aspect of the information processing apparatus described in the claims and has a function of clarifying and visualizing the uttered voice. The user terminal 1A can be used, for example, as a pre-test before a call, but is not limited thereto.

[0101] FIG. 9 is a block diagram showing a functional configuration of the user terminal 1A. As shown in FIG. 9, in addition to the same configuration as that of the user terminal 1, the user terminal 1A includes a reproduction control unit 117 in the control unit 110. The reproduction control unit 117 may be included in a clarification application that extends the call function by the call application unit 111. Further, the visualization control unit 116 in the present embodiment is a modified aspect of the visualization control unit 116 in Embodiment 1. Since the other configurations are the same as those in Embodiment 1, detailed descriptions will not be repeated.

[0102] The visualization control unit 116 displays a visualization image in response to a second operation after the utterance by the speaker in the visualization control process. The second operation is, for example, an operation in which the user U1 of the user terminal 1A instructs to clarify the uttered voice after the utterance. However, it is sufficient that at least the display of the visualization image is executed in response to the second operation, and the voice processing for generating the clarified voice data and the process for generating the visualization image itself may be performed in response to the second operation or may be started before the second operation.

[0103] The reproduction control unit 117 executes reproduction control processing. The reproduction control processing is processing for reproducing the selected voice in response to a third operation of selecting either the voice indicated by the voice data or the voice indicated by the clarified voice data after the speaker's speech. For example, after the user U1's speech, the user U1 may sequentially perform both a third operation of selecting the voice indicated by the voice data and a third operation of selecting the voice indicated by the clarified voice data in order to compare the voices before and after clarification. Note that, in this case, the order of performing each third operation is not limited to the order described above.

[0104] (Flow of information processing method S3) The user terminal 1A configured as described above executes the information processing method S3. In the information processing method S3, it is assumed that the user U1 performs a preliminary test of clarification for his / her own voice before a call. Hereinafter, "performing a preliminary test of clarification before a call" is also referred to as "test call". The information processing method S3 is a method of presenting a visualization image that visualizes the voice processing for the voice of the user U1 to the user U1. FIG. 10 is a flowchart showing the flow of the information processing method S3. As shown in FIG. 10, the information processing method S3 includes steps S301 to S306.

[0105] Step S301 is an example of the first acquisition processing. In step S301, the first acquisition unit 112 acquires the voice data of the user U1 via the voice input unit 150. For example, assume that the user U1 spoke during a test call. The acquired voice data is stored in the storage unit 120.

[0106] In step S302, the clarification UI unit 115 receives a second operation for instructing clarification of the voice from the user U1. Also, in step S302, the clarification UI unit 115 may further receive a first operation for setting the clarification level. The setting information including the clarification level set by the first operation is stored in the storage unit 120. Also, when the first operation is not received, the previously stored setting information may be applied.

[0107] Step S303 is an example of the second acquisition process. In step S303, the second acquisition unit 113 performs voice processing on the voice data of user U1 according to the setting information. As a result, clarified voice data clarified according to the clarification level indicated by the setting information is acquired. The acquired clarified voice data is stored in the storage unit 120.

[0108] Step S304 is an example of the visualization control process. In step S304, the visualization control unit 116 generates a visualization image based on the spectrum of the voice data and the spectrum of the clarified voice data by referring to the voice data and the clarified voice data stored in the storage unit 120. The visualization image is a moving image corresponding to the temporal length of the voice data (or the clarified voice data). Further, the visualization control unit 116 displays the visualization image on the display unit 130. Here, the moving image that is the visualization image may be reproduced, or at least one still image constituting the moving image may be displayed. By visually recognizing the visualization image, user U1 can confirm the clarification effect on his own voice. For example, user U1 can confirm the clarification effect of how his own voice is clarified and heard by the call partner during an actual call before the call.

[0109] In step 305, the playback control unit 117 receives a third operation of selecting the voice before clarification or the voice after clarification.

[0110] In step S306, the playback control unit 117 outputs the voice selected by the third operation received in step S305 via the voice output unit 160. As a result, the user U1 can listen to his own voice before clarification or his own voice after clarification. Further, the user U1 can compare his own voice before and after clarification by repeating steps S305 to S306. In this way, the user U1 can confirm the clarification effect on his own voice by listening to his own voice before or after clarification. For example, the user U1 can confirm the clarification effect of how his own voice is clarified and heard by the call partner during a call before the actual call.

[0111] (Screen example) FIG. 11 is a diagram schematically showing a screen example G3 displayed on the display unit 130 of the user terminal 1A in the information processing method S3. The screen example G3 provides a user interface related to the clarification function and visualization function of the user's own voice in a test call. As shown in FIG. 11, the screen example G3 includes operation objects G31, G32-1 to G32-3, G34 to G36, and a visualization image G33.

[0112] The operation object G31 receives an operation to start speaking. When an operation on the operation object G31 is received, the first acquisition unit 112 acquires voice data indicating the voice of the user U1 input via the voice input unit 150. For example, the acquired voice data may indicate the voice input until a predetermined speaking time has elapsed after the operation on the operation object G31 is received. Further, for example, the acquired voice data may indicate the voice input until the end of speaking is detected after the operation on the operation object G31 is received. Further, for example, the acquired voice data may indicate the voice input until the operation is received again after the operation on the operation object G31 is received.

[0113] The operation objects G32-1 "weak", G32-2 "medium", and G32-3 "strong" accept an operation (an example of the first operation) to select a clarification level in a test call. Also, in this example, when an operation on any of the operation objects G32-1 to G32-3 is accepted, voice processing is executed according to the clarification level. That is, the operation on the operation objects G32-1 to G32-3 is an example of the first operation and also an example of the second operation that instructs voice clarification. Also, in the screen example G3, G32-2 "medium", which is highlighted among the operation objects G32-1 to G32-3, indicates that an instruction to execute voice processing at the clarification level "medium" has been accepted. Note that the number of levels, names, and highlighting modes of the clarification level in a test call are not limited to the examples shown in FIG. 11.

[0114] The visualization image G33 is a moving image generated based on the spectra of the voice data and the clarified voice data. Note that the visualization image G33 may be played as a moving image, or at least one still image constituting the moving image may be displayed. Also, when a still image is displayed, for example, it may be played as a moving image when operations on the operation objects G34 and G35 described later are performed. With the visualization image G33, the user U1 can arrange and compare the spectra G33-1 and G33-2 before and after clarification, so that the effect of how his / her voice is clarified and heard by the call partner during a call can be confirmed before an actual call.

[0115] Here, the screen example G3 may include a visualization image in a display mode corresponding to the visualization images G16a to G16d shown in FIG. 6 instead of the visualization image G33. Note that the display mode of the visualization image included in the screen example G3 is not limited to the examples described above.

[0116] The operation object G34, "Listen to the voice before clarification", accepts an operation (an example of the third operation) of selecting the voice before clarification as the voice to be played. The operation object G35, "Listen to the clarified voice", accepts an operation (an example of the third operation) of selecting the voice after clarification as the voice to be played. For example, along with the playback of any of the voices, the above-described visualization image G33 may be played in synchronization with the voice.

[0117] The operation object G36 accepts an operation for instructing the end of a test call. When this operation is accepted, for example, the screen example G3 may transition to a screen for starting a call.

[0118] In the screen example G3, before any of the operation objects G32-1 to G32-3 for instructing voice clarification is operated and the visualization image G33 is displayed, the operation objects G34 and G35 may be unable to accept operations or may be in a non-displayed state. In this case, when any of the operation objects G32-1 to G32-3 is operated and the visualization image G33 is displayed, the operation objects G34 and G35 may become able to accept operations or may change from non-display to display. In other words, the third operation may be in a mode where it can be accepted after the second operation is accepted.

[0119] Also, in the screen example G3, the operation objects G34 and G35 may not be included. For example, when any of the operation objects G32-1 to G32-3 is operated, voice processing is performed, the visualization image G33 is displayed, and the voice before clarification and the voice after clarification are sequentially played. In this case, the operation on the operation objects G32-1 to G32-3 is an example of the second operation for instructing voice clarification and also an example of the third operation for selecting the voice to be played. That is, the same operation may be applied as the second operation and the third operation.

[0120] (Modification of Embodiment 2) In the user terminal 1A according to the second embodiment, the second acquisition unit 113 can be deformed in the same manner as the second acquisition unit 113 in the second modification of the first embodiment. As a result, even when a pretest is performed, the same effects as those of the second modification of the first embodiment can be obtained.

[0121] [Embodiment 3] The information processing system 100B according to the third embodiment of the present disclosure will be described below. For the sake of convenience of explanation, members having the same functions as those described in the above embodiments are denoted by the same reference numerals, and the description thereof will not be repeated.

[0122] (Configuration of Information Processing System 100B) FIG. 12 is a block diagram showing the configuration of the information processing system 100B. As shown in FIG. 12, the information processing system 100B includes a user terminal 1B used by the user U1, a user terminal 2 used by the user U2, and a server 3. The user terminal 1B and the user terminal 2 can be connected via the server 3 for the users U1 and U2 to make a call. Note that the user terminal 1B, the user terminal 2, and the server 3 can be connected via the network N. In FIG. 12, one user terminal 1B, one user terminal 2, and one server 3 are shown, but the number of each may be plural. The server 3 is a computer having a function of relaying between the user terminal 1B and the user terminal 2 by a predetermined call method. The server 3 is an aspect of the information processing apparatus described in the claims, and provides at least an audio clarification function and a visualization function to the user terminal 1B.

[0123] The user terminal 1B is a terminal used by the user U1 and is a modified form of the user terminal 1. The user terminal 1B has a user interface that utilizes the clarification function and the visualization function provided by the server 3. Regarding the user terminal 2, as described above, it will be described as not having the clarification function and the visualization function, but it is not limited thereto, and it may have such functions.

[0124] (Configuration of Server 3) FIG. 13 is a block diagram showing a functional configuration of each device constituting the information processing system 100B. As shown in FIG. 13, the server 3 includes a control unit 310, a storage unit 320, and a communication unit 370. The control unit 310 is realized, for example, by a processor executing a program stored in a memory, and comprehensively controls each part of the server 3. Details of each functional block included in the control unit 310 will be described later. The storage unit 320 is constituted by, for example, a memory, and stores various data and programs used by the control unit 310. The communication unit 370 is connected to the network N and communicates with the outside. The communication unit 370 transmits the information input from the control unit 310 via the network N. Further, the communication unit 370 outputs the information received via the network N to the control unit 310. Note that the storage unit 320 and the communication unit 370 may be connected as peripheral devices instead of being built in the server 3.

[0125] (Functional Blocks of Control Unit 310) As shown in FIG. 13, the control unit 310 includes a call relay unit 311, a first acquisition unit 312, a second acquisition unit 313, an audio output control unit 314, and a visualization control unit 316.

[0126] The call relay unit 311 has a function of relaying calls between a plurality of user terminals (for example, user terminal 1B and user terminal 2) according to a predetermined call method. For example, the call relay unit 311 generates a call session in response to a request from any of the plurality of user terminals. Further, the call relay unit 311 causes the user terminal to participate in the generated call session in response to a request from any of the plurality of user terminals. Also, for example, the call relay unit 311 relays audio data between the plurality of participating user terminals during the period in which the call session is generated. Also, for example, the call relay unit 311 causes the user terminal to leave the call session or ends the call session itself in response to a request from any of the plurality of user terminals.

[0127] For example, the call relay unit 311 may be at least part of a server application that provides a voice call function between two parties or a voice call function in a group. Also, for example, the call relay unit 311 may be at least part of a server application that provides a video call function between two parties or a video call function in a group. In the following, an example in which the call relay unit 311 relays between two user terminals, the user terminal 1B and the user terminal 2, will be mainly described, but the number of user terminals to be relayed may be three or more.

[0128] The first acquisition unit 312 executes a first acquisition process by receiving voice data of the speaker (for example, user U1 or user U2) from the user terminal 1B or the user terminal 2.

[0129] The second acquisition unit 313 acquires clear voice data generated by voice processing for clarifying the voice of the voice data of the speaker (for example, user U1 or user U2). Since the second acquisition unit 313 will be described in the same manner as the second acquisition unit 113, a detailed description will not be repeated.

[0130] The voice output control unit 314 controls to output the voice indicated by the clear voice data from the user terminal by transmitting the clear voice data to the user terminal of the recipient (for example, the user terminal 1B or the user terminal 2). For example, when user U1 is the speaker and the clear voice data of the user U1 is generated, the voice output control unit 314 transmits the clear voice data to the user terminal 2. Thereby, the voice indicated by the clear voice data is output from the user terminal 2. Also, for example, when user U2 is the speaker and the clear voice data of the user U2 is generated, the voice output control unit 314 transmits the clear voice data to the user terminal 1B. Thereby, the voice indicated by the clear voice data is output from the user terminal 1B.

[0131] The visualization control unit 316 executes visualization control processing. For example, the visualization control unit 316 transmits a visualization image to the user terminal 1B having a visualization function. Note that the visualization image may not be transmitted to the user terminal 2 that does not have the visualization function. For example, when the user U1 is the speaker and the clarified voice data of the user U1 is generated, a visualization image that visualizes the clarification of the transmitted voice is transmitted to the user terminal 1B. Also, for example, when the user U2 is the speaker and the clarified voice data of the user U2 is generated, a visualization image that visualizes the clarification of the received voice is transmitted to the user terminal 1B.

[0132] (Configuration of User Terminal 1B) As shown in FIG. 13, the user terminal 1B includes a control unit 110, a storage unit 120, a display unit 130, an input unit 140, a voice input unit 150, a voice output unit 160, and a communication unit 170, similar to the user terminal 1. However, the functional blocks included in the control unit 110 are different. Since the storage unit 120, the display unit 130, the input unit 140, the voice input unit 150, the voice output unit 160, and the communication unit 170 are as described above, detailed descriptions will not be repeated.

[0133] (Functional Blocks of Control Unit 110) As shown in FIG. 13, the control unit 110 includes a call application unit 111 and a clarification UI unit 115. The clarification UI unit 115 may constitute a clarification application that extends the call function by the call application unit 111.

[0134] The call application unit 111 provides a function for making calls according to a predetermined call method. The call application unit 111 is a modified form of the call application unit 111 in Embodiment 1. Hereinafter, the description will focus on the points modified in the call application unit 111, and the same points will not be repeated. The call application unit 111 makes a connection with the call destination via the server 3. Also, the call application unit 111 transmits and receives voice data with the call destination via the server 3. Also, the call application unit 111 can make calls with a plurality of call destinations, not limited to one call destination. For example, the call application unit 111 transmits a request for generating a call session with one or a plurality of call destinations to the server 3. Also, the call application unit 111 participates in the call session by transmitting a request for participating in the call session to the server 3. Also, the call application unit 111 transmits and receives voice data with other user terminals participating in the same call session via the server 3. Also, the call application unit 111 exits the call session by transmitting a request for exiting the call session to the server 3. Also, the call application unit 111 ends the call session by transmitting a request for ending the call session to the server 3.

[0135] The clarification UI unit 115 receives a first operation and generates setting information. Since the details of the first operation and the setting information are as described above, a detailed description will not be repeated. The generated setting information is transmitted to the server 3 and stored in the storage unit 320 of the server 3. Also, the clarification UI unit 115 displays the visualization image transmitted from the server 3 on the display unit 130.

[0136] (Configuration of User Terminal 2) Since the configuration of the user terminal 2 is as described above, a detailed description will not be repeated. However, the call application unit 211 is a modified form of the call application unit 211 in Embodiment 1. The points modified in the call application unit 211 are described in the same way as the points modified in the call application unit 111.

[0137] (Flow of Information Processing Method S4) The information processing system 100B configured as described above executes the information processing method S4. In the information processing method S4, when the user U1 is the speaker, voice processing for clarifying the transmitted voice of the user U1 is performed in the server 3 that relays the call. Further, a visualization image that visualizes the voice processing for the transmitted voice of the user U1 is generated by the server 3. In the following, an example of a two-party call in the information processing method S4 will be described. However, the information processing method S4 can also be applied to group calls of three or more parties by considering that at least one of the user terminal 1B and the user terminal 2 exists in plural. FIG. 14 is a flowchart for explaining the flow of the information processing method S4. As shown in FIG. 14, the information processing method S4 includes steps S401 to S422.

[0138] In step S401, the call application unit 111 of the user terminal 1B transmits a call session generation request and a participation request to the server 3 based on an operation of the user U1 for starting a call with the user U2. Further, in step S402, the call relay unit 311 of the server 3 generates a call session and allows the user terminal 1B to participate. Further, in step S403, the call application unit 211 of the user terminal 2 transmits a participation request to the call session to the server 3 based on an operation of the user U2 for responding to a call with the user U1. The call relay unit 311 of the server 3 allows the user terminal 2 to participate in the call session. Thereby, the call between the user U1 and the user U2 is started. Note that steps S401 to S403 are not limited to being executed in this order, and can be executed in the reverse order, or partially or entirely in parallel. Further, steps S401 to S403 do not limit which of the user U1 and the user U2 requests a call to the other.

[0139] In step S404, the clarification UI unit 115 of the user terminal 1 receives a first operation for setting the clarification level of the transmitted voice from the user U1. Details of step S404 are explained in the same manner as step S103. Further, the clarification UI unit 115 transmits the setting information to the server 3.

[0140] In step S405, the second acquisition unit 313 of the server 3 stores the setting information including the clarification level of the transmitted voice in the storage unit 320.

[0141] In step S406, the call application unit 111 of the user terminal 1B acquires the voice data of the user U1 via the voice input unit 150. That is, it is assumed that the user U1 spoke during the call. Also, the call application unit 111 transmits the voice data of the user U1 to the server 3.

[0142] Step S407 is an example of the first acquisition process. In step S407, the first acquisition unit 312 of the server 3 acquires the voice data of the user U1 by receiving it from the user terminal 1B.

[0143] Step S408 is an example of the second acquisition process. In step S408, the second acquisition unit 313 performs voice processing on the voice data of the user U1 according to the setting information. As a result, the clarified voice data clarified according to the clarification level indicated by the setting information is acquired. Note that when the setting information indicates the clarification level of the transmitted voice as "off", the voice processing is not performed and the clarified voice data is not acquired.

[0144] In step S409, the voice output control unit 314 transmits the clarified voice data to the user terminal 2. When the clarified voice data is not generated, instead, the voice data of the user U1 is transmitted.

[0145] In step S410, the call application unit 211 of the user terminal 2 receives the clarified voice data. Also, the call application unit 211 outputs the voice indicated by the received clarified voice data via the voice output unit 260. As a result, the user U2 can hear the clarified voice of the user U1. Note that when voice data is received instead of the clarified voice data, the unclarified voice indicated by the voice data is output.

[0146] Step S411 is an example of visualization control processing. In step S411, the visualization control unit 316 of the server 3 generates a visualization image based on the spectrum of the voice data of the user U1 and the spectrum of the clarified voice data. Further, the visualization control unit 316 transmits the visualization image to the user terminal 1B.

[0147] In step S412, the clarification UI unit 115 of the user terminal 1B displays the visualization image received from the server 3 on the display unit 130. Here, steps S406 to S412 are executed in real time according to the progress of the speech by the user U1. Therefore, the visualization image displayed in step S412 is a moving image that changes according to the progress of the speech. The user U1 can confirm the effect on the user U2 side due to the clarification of his own voice by visually recognizing the visualization image.

[0148] Thereafter, when the user U2 speaks, the user terminal 1B receives the voice data of the user U2 transmitted from the user terminal 2 from the server 3, and outputs the received voice data via the voice output unit 160. As a result, the user U1 can hear the voice of the user U2 that has not been clarified. Note that it is also possible to configure such that steps S506 to S512 of the information processing method S5 described later are executed when the user U2 speaks. In this case, the user U1 can hear the clarified voice of the user U2. Further, when the user U1 speaks again, steps S406 to S412 are repeated.

[0149] Also, when step S404 is executed again during the call and the clarification level of the transmitted voice is changed, the changed clarification level is reflected in the clarified voice data obtained in step S408 and the voice that can be heard by the user U2 in step S410. Further, the changed clarification level is reflected in the visualization image presented to the user U1 in step S412.

[0150] In step S420, the call application unit 111 of the user terminal 1 transmits a request to end the call session to the server 3 based on the operation of the user U1. Further, in step S421, the call relay unit 311 of the server 3 ends the call session in response to the end request. Further, in step S422, the call application unit 211 of the user terminal 2 performs a process to complete the exit from the ended call session. As a result, the call between the user U1 and the user U2 ends. Note that steps S420 to S422 are not limited to being executed in this order, and the reverse order, or a part or all of them may be executed in parallel. Further, the operation for ending the call session is not limited to the user U1, and the user U2 may perform it, or both may perform it. Further, at least one of the user U1 and the user U2 may perform an operation for exiting instead of the operation for ending the call session.

[0151] According to the information processing method S4, for example, on the display unit 130 of the user terminal 1B, a screen example G1 is displayed in the same manner as in the information processing method S1. Details of the screen example G1 are as described above. By visually recognizing the screen example G1, the user U1 can confirm the effect that his / her voice is clearly heard by the user U2 who is the call partner.

[0152] (Flow of the information processing method S5) Further, the information processing system 100B executes the information processing method S5. In the information processing method S5, voice processing for clarifying the received voice that can be heard by the user U1 when the user U2 is the speaker is performed in the server 3 that relays the call. Further, a visualization image that visualizes the voice processing for the received voice that can be heard by the user U1 is generated by the server 3. Hereinafter, similar to the information processing method S4, an example of a two-party call in the information processing method S5 will be described. However, the information processing method S5 is applicable to group calls of three or more parties by assuming that at least one of the user terminal 1B and the user terminal 2 exists in plural. FIG. 15 is a flowchart for explaining the flow of the information processing method S5. As shown in FIG. 15, the information processing method S5 includes steps S501 to S522.

[0153] Steps S501 to S503 are steps for starting a conversation between user U1 and user U2. Since steps S501 to S503 are explained in the same way as steps S401 to S403, detailed explanations will not be repeated.

[0154] In step S504, the clarification UI unit 115 of the user terminal 1B receives a first operation from the user U1 to set the clarification level of the received voice. Details of step S504 are explained in the same way as step S203. Also, the clarification UI unit 115 transmits the setting information to the server 3.

[0155] In step S505, the second acquisition unit 313 of the server 3 stores the setting information including the clarification level of the received voice in the storage unit 320.

[0156] In step S506, the call application unit 211 of the user terminal 2 acquires the voice data of the user U2 via the voice input unit 250. That is, it is assumed that the user U2 spoke during the call. Also, the call application unit 211 transmits the voice data of the user U2 to the server 3.

[0157] Step S507 is an example of the first acquisition process. In step S507, the first acquisition unit 312 of the server 3 acquires the voice data of the user U2 by receiving it from the user terminal 2.

[0158] Step S508 is an example of the second acquisition process. In step S508, the second acquisition unit 313 performs voice processing on the voice data of the user U2 according to the setting information. As a result, clarified voice data clarified according to the clarification level indicated by the setting information is acquired. When the setting information indicates that the clarification level of the received voice is "off", voice processing is not performed and clarified voice data is not acquired.

[0159] In step S509, the voice output control unit 314 transmits the clarified voice data to the user terminal 1B. If the clarified voice data has not been generated, the voice data of user U2 is transmitted instead.

[0160] In step S510, the call application unit 111 of the user terminal 1B receives the clarified voice data. Further, the call application unit 111 outputs the voice indicated by the received clarified voice data via the voice output unit 160. Thereby, the clarified voice of user U1 can be heard by user U1. Note that if voice data is received instead of the clarified voice data, the unclarified voice indicated by the voice data is output.

[0161] Step S511 is an example of visualization control processing. In step S511, the visualization control unit 316 of the server 3 generates a visualization image based on the spectrum of the voice data of user U2 and the spectrum of the clarified voice data. Further, the visualization control unit 316 transmits the visualization image to the user terminal 1B.

[0162] In step S512, the clarified UI unit 115 of the user terminal 1B displays the visualization image received from the server 3 on the display unit 130. Here, steps S506 to S512 are executed in real time according to the progress of the speech by user U2. Therefore, the visualization image displayed in step S512 is a moving image that changes according to the progress of the speech. User U1 can confirm the effect that the voice of user U2, who is the call partner, is clarified by visually recognizing the visualization image.

[0163] Hereafter, when the user U1 speaks, the user terminal 2 receives the voice data of the user U1 transmitted from the user terminal 1B from the server 3, and outputs the received voice data via the voice output unit 260. As a result, the unclear voice of the user U1 can be heard by the user U2. Note that when the user U1 speaks, it is also possible to configure the steps S406 to S412 of the above-described information processing method S4 to be executed. In this case, the clarified voice of the user U1 can be heard by the user U2. Further, when the user U2 speaks, steps S506 to S512 are repeated.

[0164] Also, when step S504 is executed again during the call and the clarification level of the received voice is changed, the changed clarification level is reflected in the clarified voice data obtained in step S508 and the voice audible to the user U1 in step S510. Further, the changed clarification level is reflected in the visualization image presented to the user U1 in step S512.

[0165] Steps S520 to S522 are steps for ending the call between the user U1 and the user U2. Since steps S520 to S522 are described in the same manner as steps S420 to S422, a detailed description will not be repeated.

[0166] By the information processing method S5, for example, a screen example G2 is displayed on the display unit 130 of the user terminal 1B in the same manner as the information processing method S2. The details of the screen example G2 are as described above. By visually recognizing the screen example G2, the user U1 can confirm the effect that the voice of the user U2, who is the call partner, is clarified.

[0167] (Modification Example 1 of Embodiment 3) In Embodiment 3, the visualization control unit 316 of the server 3 may be modified to display a visualization image in response to a second operation after the user U1's utterance. Further, the control unit 110 of the server 3 may be modified to include a playback control unit that executes playback control processing. As described in Embodiment 2, the playback control processing is processing for playing back the selected voice in response to a third operation of selecting either the voice indicated by the voice data or the voice indicated by the clarified voice data after the user U1's utterance. Further, the clarification UI unit 115 of the user terminal 1B may be modified to receive a second operation or a third operation after the user U1's utterance. Details of the second operation and the third operation are as described in Embodiment 2.

[0168] According to Modification Example 1, the user U1 can confirm the effect of clarification on his or her own voice by reproducing the visualization image or the voice before and after clarification regardless of being in a call. For example, a screen example G3 similar to that in the embodiment may be displayed on the display unit 130 of the user terminal 1B. Thereby, the user U1 can confirm the effect of clarification of how his or her own voice is clarified and heard by the call partner during the call before the actual call.

[0169] (Modification Example 2 of Embodiment 3) In Embodiment 3, instead of the server 3 including the visualization control unit 316, the user terminal 1B may be modified to include a visualization control unit 116. For example, in the information processing method S4 for visualizing the clarification of the transmitted voice, the visualization control processing in step S411 may be executed by the user terminal 1B instead of the server 3. In this case, for example, in step S409, the server 3 not only transmits the clarified voice data of the user U1 to the user terminal 2 but also to the user terminal 1B. As a result, it becomes possible for the user terminal 1B to execute the visualization control processing in step S411. As a result, instead of transmitting a moving image as a visualization image from the server 3 to the user terminal 1B, it is sufficient to transmit clarified voice data having a generally smaller capacity than the moving image, so that the communication cost is reduced.

[0170] Also, for example, in the information processing method S5 for clarifying received voice, the visualization control process in step S511 may be executed by the user terminal 1B instead of the server 3. For example, in step S509, the server 3 transmits not only the clarified voice data of the user U2 but also the voice data before voice processing to the user terminal 1B. As a result, the user terminal 1B can execute the visualization control process in step S511. Consequently, instead of transmitting a moving image as a visualization image from the server 3 to the user terminal 1B, voice data with a generally smaller capacity than the visualization image may be transmitted, thus reducing the communication cost.

[0171] (Modification Example 3 of Embodiment 3) In the information processing system 100B according to Embodiment 3, the second acquisition unit 313 of the server 3 can be modified in the same manner as the second acquisition unit 113 in Modification Example 1 or 2 of Embodiment 1. Thereby, even when the server 3 relays a call, the same effects as those in Modification Example 1 or 2 of Embodiment 1 can be achieved.

[0172] [Embodiment 4] The telephone system 100C according to Embodiment 4 of the present disclosure will be described below. The telephone system 100C includes a voice processing device 10C that executes voice processing for clarifying transmitted voice or received voice. Further, the voice processing device 10C incorporates a microcomputer 1C that visualizes the voice processing. The microcomputer 1C is an aspect of the information processing device described in the claims. Note that the microcomputer 1C is not limited to what is called a "microcomputer" and may be a computer equipped with a processor and a memory and capable of being incorporated in the voice processing device 10C. Hereinafter, for convenience of explanation, members having the same functions as the members described in the above embodiments are denoted by the same reference numerals, and their descriptions will not be repeated.

[0173] (Configuration of Telephone System 100C) FIG. 16 is a block diagram showing the configuration of the telephone system 100C. As shown in FIG. 16, the telephone system 100C includes a voice processing device 10C, a handset 20, a main body 30, and cables 41, 42, and 43. The telephone system 100C is used for making a call by the user U1. Also, the call partner who makes a call via the telephone system 100C is described as the user U2.

[0174] (Handset 20) The handset 20 includes a microphone 21, a speaker 22, and a port P21. Each of the microphone 21 and the speaker 22 is connected to the port P21. The microphone 21 is an example of a transmitter, and the speaker 22 is an example of a receiver.

[0175] The microphone 21 converts the voice uttered by the user U1 into a voice signal S V and supplies the voice signal S V to the port P21. In the present embodiment, the voice signal S V is an analog signal and an electrical signal. However, the microphone 21 may be configured to convert the voice uttered by the user U1 into a voice signal S V which is a digital signal.

[0176] As the microphone 21, for example, an electric condenser type microphone can be adopted. A DC voltage for driving is applied to the electric condenser type microphone. Also, the electric condenser type microphone is often connected using different types of connectors depending on the manufacturer or model.

[0177] Speaker 22 converts an audio signal supplied from port P21, which is an audio signal obtained by converting the voice uttered by user U2, into sound and outputs the sound. In the present embodiment, the audio signal supplied from port P21 is an analog signal and also an electrical signal. However, the audio signal supplied from port P21 may be a digital signal, and speaker 22 may be configured to convert an audio signal that is a digital signal into sound.

[0178] Port P21 is an example of a port of the transceiver 20. In the present embodiment, a modular jack compliant with the RJ-9 standard is employed. However, port P21 is not limited to a modular jack compliant with the RJ-9 standard. Any connector that can input and output an audio signal may be employed as port P21.

[0179] (Main body 30) Main body 30 is provided with ports P31 and P32. Port P31 is an example of a port on the transceiver 20 side of main body 30. In the present embodiment, a modular jack compliant with the RJ-9 standard is employed in the same manner as port P21.

[0180] Port P32 is an example of a port on the line side of main body 30. In the present embodiment, a modular jack compliant with the RJ-11 standard is employed. However, port P32 is not limited to a modular jack compliant with the RJ-11 standard. As port P32, for example, a modular jack compliant with the RJ-12 standard or the RJ-14 standard may be employed.

[0181] Main body 30 is a main body of a telephone using a two-wire telephone line. Main body 30 is configured to be able to perform two-way communication of an audio signal with the main body of the receiving-side telephone by connecting the telephone line to port P32 and connecting transceiver 20 to port P31. Main body 30 may be configured in the same manner as the main body of an existing telephone. Therefore, in the present embodiment, the description of main body 30 is omitted.

[0182] The main body 30 outputs the synthesized voice signal S supplied from the addition synthesizer 16 of the voice processing device 10C described later to the port P31 to a cable 43 described later. SV

[0183] (Cables 41, 42, 43) Each of the cables 41, 42, and 43 includes a plurality of signal lines. Each of the plurality of signal lines is composed of a conductor so as to be able to transmit a voice signal which is an electrical signal.

[0184] The cable 41 connects the port P11 of the voice processing device 10C described later and the port P21 of the handset 20, the cable 42 connects the port P12 of the voice processing device 10C and the port P31 of the main body 30, and the cable 43 constitutes an end of the telephone line and is connected to the port P32 of the main body 30.

[0185] In the present embodiment, each of the cables 41 and 42 is a cable provided with modular plugs compliant with the RJ-9 standard at both ends. Also, in the present embodiment, the cable 43 is a cable provided with modular plugs compliant with the RJ-11 standard at both ends. Note that in FIG. 16, the black-painted squares shown at both ends of each of the cables 41 and 42 indicate modular plugs compliant with the RJ-9 standard, and the white squares shown at the ends of the cable 43 indicate modular plugs compliant with the RJ-11 standard.

[0186] (Voice processing device 10C) FIG. 17 is a block diagram showing a detailed configuration of the voice processing device 10C. As shown in FIG. 17, the voice processing device 10C includes a microcomputer 1C, a display unit 130, an AD (Analog-Digital) converter 180, and a voice processing circuit 200. The voice processing device 10C is provided so as to be interposed between the handset 20 and the main body 30.

[0187] ​Also, in this embodiment, the housing of the voice processing device 10C is made of aluminum. However, the material constituting the housing is not limited to aluminum, and may be a metal such as copper or stainless steel, for example. Further, the material constituting the housing may be mainly resin with a metal layer provided on its surface. By covering at least the surface of the housing with metal, electromagnetic waves entering from the outside can be shielded from entering the inside.

[0188] (Voice processing circuit 200) The voice processing circuit 200 performs voice processing on the voice signal S input from the transceiver 20 V to generate a synthesized voice signal S on which voice processing has been performed so that the voice is perceived as clear SV and outputs it to the main body 30. Also, the voice signal S input to the voice processing circuit 200 V is also supplied to the AD converter 180 and input to the microcomputer 1C as voice data converted into a digital signal by the AD converter 180. Also, the synthesized voice signal S output from the voice processing circuit 200 SV is also supplied to the AD converter 180 and input to the microcomputer 1C as clarified voice data converted into a digital signal by the AD converter 180.

[0189] The voice processing circuit 200 includes a port P11, a port P12, a branch unit 11, a phase corrector 12, a filter 13, an amplifier 14, an adjuster 15, and an addition synthesizer 16. The voice processing circuit 200 further includes a power supply unit, a detection unit, and a control unit (not shown in FIG. 17).

[0190] The port P11 is a port connected to the transceiver 20. The port P11 and the port P21 of the transceiver 20 are connected using a cable 41. Among the terminals of the port P11, the terminal connected to the microphone 21 receives the voice signal S V converted from the voice uttered by the user U1 by the microphone 21 and supplied from the microphone 21.

[0191] The voice signal SV includes a primary formant component, a secondary formant component, ···, an n-th formant component. Here, n depends on the user U1 who speaks and is at least a positive integer of 4 or more, although there are some individual differences.

[0192] The branch portion 11 branches the audio signal S V into a first audio signal S V1 and a second audio signal S V2 . In the present embodiment, the branch portion 11 is configured such that the intensity ratio between the first audio signal S V1 and the second audio signal S V2 is 1:1, that is, the distribution ratio is 1:1. However, the distribution ratio of the branch portion 11 is not limited to 1:1 and can be set as appropriate.

[0193] Note that the spectrum of each of the first audio signal S V1 and the second audio signal S V2 is the same as the spectrum of the audio signal S V . That is, each of the first audio signal S V1 and the second audio signal S V2 includes a primary formant component, a secondary formant component, ···, an n-th formant component.

[0194] The filter 13 is a high-pass filter that removes low-frequency components that are components with a frequency of less than 400 Hz from the first audio signal S V1 that has passed through the branch portion 11. The filter 13 outputs the first audio signal S V1’ that does not include the low-frequency components. Therefore, the first audio signal S V1’ is an audio signal that does not include a primary formant component and includes a secondary formant component, a tertiary formant component, ···, an n-th formant component.

[0195] The amplifier 14 outputs a first audio signal S V1’ in which the intensities of the secondary formant component, the tertiary formant component, ···, the n-th formant component are increased by amplifying the first audio signal S V1” that has passed through the filter 13.

[0196] The regulator 15 adjusts the gain of the amplifier 14. In this embodiment, the regulator 15 is configured using a switch capable of adjusting the gain in three steps. When the regulator 15 selects "0" by the switch, the gain becomes 1 times; when it selects "1" by the switch, the gain becomes 3 times; and when it selects "2" by the switch, the gain becomes 5 times.

[0197] However, the number of steps of the switch constituting the regulator 15 and the gain selected by the switch are not limited to the above-described example and can be appropriately selected. Also, instead of a switch that discretely changes the gain, the regulator 15 can employ a volume that continuously changes the gain. Further, when the receiving-side user U2 can be identified to some extent, such as when the telephone system 100C is placed at home, the gain can be preset and the regulator 15 can be omitted.

[0198] The phase corrector 12 corrects the phase of the second audio signal S V1” to match the phase of the first audio signal S V2 that has passed through the amplifier 14. That is, the phase corrector 12 outputs the second audio signal S V1” whose phase matches that of the first audio signal S V2’ .

[0199] The addition synthesizer 16 generates a synthesized audio signal S V1” by adding and synthesizing the first audio signal S V2’ that has passed through the amplifier 14 and the second audio signal S SV that has passed through the phase corrector 12, and supplies the synthesized audio signal S SV to the port P12.

[0200] Port P12 is a port connected to the main body 30 and is an example of the second port. Port P12 and port P31 of the main body 30 are connected using cable 42. To the terminal of port P12 to which the adder synthesizer 16 is connected, the synthesized audio signal S SV is supplied from the adder synthesizer 16.

[0201] In the present embodiment, the microphone 21 supplies the audio signal S V , which is an analog signal, to port P11 via port P21. Port P12 supplies the synthesized audio signal S SV , which is an analog signal, to the main body 30 via port P31. However, the microphone 21 may be configured to supply the audio signal S V , which is a digital signal, to port P11, and port P12 may be configured to supply the synthesized audio signal S SV , which is a digital signal, to the main body 30.

[0202] Also, in the present embodiment, each of the phase corrector 12, the filter 13, the amplifier 14, and the adder synthesizer 16 is constituted by an analog circuit. However, each of the phase corrector 12, the filter 13, the amplifier 14, and the adder synthesizer 16 may be constituted by a digital circuit.

[0203] Also, in the present embodiment, the audio processing circuit 200 includes a filter 13, which is a high-pass filter, as a filter that performs filtering processing on the first audio signal S V1 . However, in the audio processing circuit 200, the filter 13 is the first audio signal S that has passed through the branch unit 11 V1It may be configured to remove (1) a low-frequency component that is a component with a frequency of less than 400 Hz and (2) a high-frequency component that is a component with a frequency exceeding 7 kHz. That is, the filter 13 may be a band-pass filter that passes components with a frequency of 400 Hz or more and 7 kHz or less, or may be a band-pass filter that passes components with a frequency of 400 Hz or more and 5 kHz or less. In this case, the filter 13 may be realized as a single band-pass filter, or may be realized by connecting a high-pass filter and a low-pass filter in series.

[0204] Note that in the audio processing circuit 200, the terminal of the port P11 that is connected to the speaker 22 and the predetermined terminal of the port P12 are directly connected. Therefore, the audio processing circuit 200 outputs to the port P11 without performing any audio processing on the audio signal supplied from the main body 30 to the port P12 and converted from the voice uttered by the user U2. Therefore, the telephone system 100C outputs from the speaker 22 without performing any audio processing on the voice uttered by the user U2.

[0205] (AD converter 180) The audio signal S supplied to the port P11 V branches and is input to the AD converter 180 and the branch unit 11. The AD converter 180 V converts the audio signal S into audio data that is a digital signal and outputs it to the microcomputer 1C.

[0206] The synthesized audio signal S output from the addition synthesizer 16 SV branches and is supplied to the AD converter 180 and the port P12. The AD converter 180 SV converts the synthesized audio signal S into clarified audio data that is a digital signal and outputs it to the microcomputer 1C.

[0207] (Display unit 130) The display unit 130 displays the image generated by the microcomputer 1C. The display unit 130 may be configured to include, for example, a liquid crystal display, an organic EL (Electro Luminescence) display, or the like, but is not limited thereto.

[0208] (Microcomputer 1C) The microcomputer 1C includes a control unit 110 and a storage unit 120. The control unit 110 is realized, for example, by a processor executing a program stored in a memory, and comprehensively controls each part of the voice processing apparatus 10C. The storage unit 120 is constituted by, for example, a memory, and stores various data and programs used by the control unit 110.

[0209] The control unit 110 includes a first acquisition unit 112, a second acquisition unit 113, and a visualization control unit 116.

[0210] The first acquisition unit 112 executes a first acquisition process. In the present embodiment, the first acquisition unit 112 acquires the voice data output from the AD converter 180. The voice data indicates the voice of the user U1 who speaks toward the handset 20 using the telephone system 100C.

[0211] The second acquisition unit 113 executes a second acquisition process. In the present embodiment, the second acquisition unit 113 acquires the clarified voice data output from the AD converter 180.

[0212] The visualization control unit 116 executes a visualization control process. Details of the visualization control unit 116 will be described in the same manner as in Embodiment 1. As a result, a visualization image based on the spectrum of the voice data and the spectrum of the clarified voice data is displayed on the display unit 130.

[0213] (Flow of the information processing method S6) In the telephone system 100C configured as described above, the microcomputer 1C executes an information processing method S6. The information processing method S6 is a method executed by the microcomputer 1C when the user U1 makes a call with a call partner using the telephone system 100C. Incidentally, in parallel with the execution of the information processing method S6, an audio signal S V is synthesized into an audio signal S SV in real time by the audio processing circuit 200. FIG. 18 is a flowchart showing the flow of the information processing method S6. As shown in FIG. 18, the information processing method S6 includes steps S601 to S603.

[0214] In step S601, the first acquisition unit 112 acquires the audio data output from the AD converter 180. In step S602, the second acquisition unit 113 acquires the clarified audio data output from the AD converter 180. In step S603, the visualization control unit 116 displays a visualization image on the display unit 130 based on the spectrum of the audio data and the spectrum of the clarified audio data. By visually recognizing the visualization image, the user U1 can confirm the clarification effect of how his / her voice is clarified and heard by the call partner.

[0215] (Modification Example 1 of Embodiment 4) The audio processing device 10C according to Embodiment 4 can be modified to include another audio processing circuit (hereinafter, an audio processing circuit for received voice) for clarifying the received voice in addition to the audio processing circuit 200 for clarifying the transmitted voice. In this case, the audio processing circuit for received voice is configured in the same manner as the audio processing circuit 200, but is arranged between the port P11 and the port P12 so as to be in the reverse direction to the audio processing circuit 200. That is, the adder synthesizer 16 included in the audio processing circuit for received voice is connected to the terminal corresponding to the speaker 22 of the port P11. Further, the branching unit 11 included in the audio processing circuit for received voice is connected to a predetermined terminal of the port P12.

[0216] In the voice processing circuit for received voice, a voice signal indicating the received voice is input from the main body 30, and a synthesized voice signal obtained by performing voice processing on the voice signal so that the voice is felt to be clear is output to the speaker 22. Further, the voice signal indicating the received voice input to the voice processing circuit for received voice is branched and also supplied to the AD converter 180, where it is converted into voice data which is a digital signal. The voice data is output to the microcomputer 1C. Further, the synthesized voice signal branched from the voice processing circuit for received voice is branched and also supplied to the AD converter 180, where it is converted into clarified voice data which is a digital signal. The clarified voice data is output to the microcomputer 1C.

[0217] The first acquisition unit 112 of the microcomputer 1C acquires voice data indicating the received voice. The second acquisition unit 113 acquires clarified voice data corresponding to the voice data. The visualization control unit 116 displays a visualization image on the display unit 130 based on the spectrum of the voice data indicating the received voice and the spectrum of the clarified voice data.

[0218] According to this modification, the user U1 can confirm the clarification effect that the voice of the call partner is clarified and audible by visually recognizing the visualization image.

[0219] (Modification 2 of Embodiment 4) In the voice processing apparatus 10C according to Embodiment 4, the voice data input to the microcomputer 1C is not limited to the voice signal S converted from the port P11 into a digital signal. For example, it can be said that the second voice signal S output from the phase corrector 12 indicates the voice before the gain of the higher-order harmonic component is amplified (i.e., before being clarified). Therefore, instead of the voice signal S, the second voice signal S is configured to be input to the AD converter 180, and the voice data obtained by converting the second voice signal S into a digital signal may be input to the microcomputer 1C. V For example, the second voice signal S output from the phase corrector 12 V2’ can be said to indicate the voice before the gain of the higher-order harmonic component is amplified (i.e., before being clarified). Therefore, instead of the voice signal S V , the second voice signal S V2’ is configured to be input to the AD converter 180, and the voice data obtained by converting the second voice signal S V2’ into a digital signal may be input to the microcomputer 1C.

[0220] Also, it is not always necessary for the microcomputer 1C to receive voice data and clarified voice data. It is sufficient to input data necessary for generating a visualized image. For example, for the voice signal S V and the synthesized voice signal S SV instead, the first voice signal S V1” output from the amplifier 14 may be input to the AD converter 180. It can be said that the first voice signal S V1” indicates the difference between the voice signals before and after clarification. In this case, the differential voice data obtained by converting the first voice signal S V1” into a digital signal is input to the microcomputer 1C. The visualization control unit 116 can generate a visualization image (for example, the visualization image G16b in FIG. 6) indicating the difference between the spectrum of the voice data and the spectrum of the clarified voice data based on the differential voice data.

[0221] (Modification Example 3 of Embodiment 4) Instead of including the voice processing circuit 200, the voice processing apparatus 10C according to Embodiment 4 may be configured such that the control unit 110 of the microcomputer 1C executes voice processing. In this case, the voice processing apparatus 10C further includes a DA (Digital - Analog) converter that converts the clarified voice data generated by the microcomputer 1C into an analog signal. The analog signal is output to the handset 20 or the main body 30.

[0222] (Another Modification Example of Embodiment 4) The voice processing apparatus C according to Embodiment 4 may be wirelessly connected, not limited to a wired connection, to one or both of the handset 20 and the main body 30. Also, the voice processing apparatus 10C is not necessarily configured as a separate body from the handset 20 and the main body 30. For example, the voice processing apparatus 10C may be housed in the housing of the main body 30 so as to be integrated with the main body 30. Also, the voice processing apparatus 10C may be housed in the housing of the handset 20 so as to be integrated with the handset 20.

[0223] In addition, the telephone system 100C according to Embodiment 4 may include a computer having a call function instead of the main body 30. The computer may be, for example, a mobile phone, a smartphone, a tablet, a smartwatch, a notebook computer, a desktop computer, etc., but is not limited thereto. Further, the handset 20 is not limited to the mode shown in FIG. 16. For example, the handset 20 may be a headset or the like.

[0224] 〔Example of implementation by software〕 The functions of the user terminals 1, 1A, 1B, the microcomputer 1C, and the server 3 (hereinafter referred to as "devices") are programs for causing a computer to function as the device, and can be realized by programs for causing a computer to function as each control block (particularly each part included in the control units 110 and 310) of the device.

[0225] In this case, the above device includes a computer having at least one control device (for example, a processor) and at least one storage device (for example, a memory) as hardware for executing the above program. By executing the above program with this control device and storage device, each function described in each of the above embodiments is realized.

[0226] The above program may be recorded on one or more computer-readable recording media, not temporarily. This recording medium may or may not be provided in the above device. In the latter case, the above program may be supplied to the above device via any wired or wireless transmission medium.

[0227] In addition, part or all of the functions of the above control blocks can also be realized by a logic circuit. For example, an integrated circuit in which a logic circuit functioning as each of the above control blocks is formed is also included in the scope of the present disclosure. In addition to this, for example, it is also possible to realize the functions of the above control blocks by a quantum computer.

[0228] In addition, each process described in each of the above embodiments may be executed by AI (Artificial Intelligence). In this case, the AI may operate in the above control device, or may operate in another device (for example, an edge computer or a cloud server, etc.).

[0229] 〔Summary〕 The information processing method according to Aspect 1 includes at least one processor executing a first acquisition process of acquiring voice data indicating the voice of the speaker, a second acquisition process of acquiring clarified voice data generated by voice processing for clarifying the voice with respect to the voice data, and a visualization control process for displaying a visualization image obtained by visualizing the voice processing based on the spectrum of the voice data and the spectrum of the clarified voice data. With the above configuration, there is an effect that the effect of clarification by voice processing for clarifying the voice can be presented to the user.

[0230] The information processing method according to Aspect 2 is, in Aspect 1, the visualization image includes an image obtained by superimposing an image showing the spectrum of the voice data and an image showing the spectrum of the clarified voice data. With the above configuration, the user can visually compare the spectra before and after clarification, so that the effect of clarification by the comparison can be confirmed.

[0231] The information processing method according to Aspect 3 is, in Aspect 1 or Aspect 2, the visualization image includes an image showing the difference between the spectrum of the voice data and the spectrum of the clarified voice data. With the above configuration, the user can visually recognize the difference in the spectra before and after clarification, so that the effect of clarification by the difference can be confirmed.

[0232] The information processing method according to Aspect 4, in any one of Aspects 1 to 3, the degree of clarification in the voice processing can be set according to a first operation, and in the visualization control processing, the at least one processor displays the visualization image reflecting the degree of clarification set according to the first operation. With the above configuration, the user can confirm how the clarification effect changes due to the change in the degree of clarification.

[0233] The information processing method according to Aspect 5, in any one of Aspects 1 to 4, the voice data indicates the voice of the speaker in a call made by connecting a plurality of user terminals, and in the visualization control processing, the at least one processor displays the visualization image on the user terminal of the speaker among the plurality of user terminals. With the above configuration, the user can confirm the clarification effect on the receiving side by voice processing for their own voice in a call.

[0234] The information processing method according to Aspect 6, in any one of Aspects 1 to 5, the voice data indicates the voice of the speaker in a call made by connecting a plurality of user terminals, and in the visualization control processing, the at least one processor displays the visualization image on the user terminal of the recipient among the plurality of user terminals. With the above configuration, the user can confirm the clarification effect by voice processing for the voice of the other party in a call.

[0235] The information processing method according to Aspect 7, in any one of Aspects 1 to 6, in the visualization control processing, the at least one processor updates the visualization image in real time according to the progress of the speech by the speaker. With the above configuration, the user can confirm the clarification effect in real time according to the progress of the speech by the speaker.

[0236] The information processing method according to Aspect 8 is, in any one of Aspects 1 to 7, in the visualization control process, the at least one processor displays the visualization image in response to a second operation after speech by the speaker. With the above configuration, the user can confirm the effect of clarification by voice processing on their own voice by visually recognizing the visualization image after their own speech.

[0237] The information processing method according to Aspect 9 is, in any one of Aspects 1 to 8, the at least one processor further executes reproduction control processing for reproducing the selected voice in response to a third operation of selecting either the voice indicated by the voice data or the voice indicated by the clarified voice data after speech by the speaker. With the above configuration, the user can confirm the effect of clarification on their own voice by listening to the voice before and after clarification after their own speech.

[0238] The information processing method according to Aspect 10 is, in any one of Aspects 1 to 9, the voice processing is a process of amplifying higher-order formant components including at least second-order formant components in the voice data. With the above configuration, the voice can be clarified.

[0239] The information processing apparatus according to Aspect 11 includes the at least one processor, and the at least one processor executes each process included in the information processing method according to any one of Aspects 1 to 10. With the above configuration, the same effect as any one of Aspects 1 to 10 is achieved.

[0240] The program according to Aspect 12 causes the at least one processor to execute each process included in the information processing method according to any one of Aspects 1 to 10. With the above configuration, the same effect as any one of Aspects 1 to 10 is achieved.

[0241] The non-transitory computer-readable recording medium according to Embodiment 13 stores the program according to Embodiment 12. With the above configuration, the same effect as any one of Embodiments 1 to 10 is achieved.

[0242] Each embodiment according to the present disclosure has an effect of being able to present to the user the effect of clarification by voice processing for clarifying the voice. Such an effect also contributes to the achievement of, for example, Goal 3, "Good health and well-being for all", of the Sustainable Development Goals (SDGs) proposed by the United Nations.

[0243] The present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope shown in the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present disclosure.

Explanation of Reference Numerals

[0244] 1, 1A, 1B, 2 User terminals 1C Microcomputer 3 Server 10C Voice processing device 11 Branch unit 12 Phase corrector 13 Filter 14 Amplifier 15 Regulator 16 Addition synthesizer 20 Receiver 20 Handset 21 Microphone 22 Speaker 30 Main body 41, 42, 43 Cable 100, 100B Information processing system 100C Telephone system 110, 210, 310 Control unit 111, 211 Call application unit 112, 312 First acquisition unit 113, 313 Second acquisition unit 114, 314 Voice output control unit 115 Clarification UI Unit 116, 316 Visualization Control Unit 117 Reproduction Control Unit 120, 220, 320 Memory Unit 130, 230 Display Unit 140, 240 Input Unit 150, 250 Voice Input Unit 160, 260 Voice Output Unit 170, 270, 370 Communication Unit 180 AD Converter 200 Voice Processing Circuit 311 Call Relay Unit 370 Communication

Claims

1. at least one processor performs: a first acquisition process of acquiring voice data indicating the voice of a speaker; a second acquisition process of acquiring clarified voice data generated by voice processing for clarifying the voice with respect to the voice data; a visualization control process for displaying a visualization image obtained by visualizing the voice processing based on the spectrum of the voice data and the spectrum of the clarified voice data; wherein the voice data indicates the voice of a speaker in a call performed by connecting a plurality of user terminals, and in the visualization control process, the at least one processor updates the visualization image in real time according to the progress of the speech by the speaker during the call. An information processing method.

2. The visualization image includes an image obtained by superimposing an image showing the spectrum of the voice data and an image showing the spectrum of the clarified voice data. The information processing method according to Claim 1.

3. The visualization image includes an image showing the difference between the spectrum of the voice data and the spectrum of the clarified voice data. The information processing method according to Claim 1.

4. The degree of clarification in the voice processing can be set according to a first operation, and in the visualization control process, the at least one processor displays the visualization image reflecting the degree of clarification set according to the first operation. The information processing method according to Claim 1.

5. In the visualization control process, the at least one processor displays the visualization image on the user terminal of the speaker among the plurality of user terminals. The information processing method according to Claim 1.

6. In the visualization control process, the at least one processor displays the visualization image on the user terminal of the listener among the plurality of user terminals. The information processing method according to Claim 1.

7. As the voice data, voice data indicating the voice of the speaker before the call is further acquired, and in the visualization control process, the at least one processor displays the visualization image obtained by visualizing the voice processing for the voice data before the call according to a second operation after the speech by the speaker before the call. The information processing method according to Claim 1.

8. The at least one processor ​ In response to a third operation of selecting either the voice indicated by the voice data or the voice indicated by the clarified voice data after the speaker's speech before the call, further execute a playback control process for playing back the selected voice. The information processing method according to claim 7.

9. The voice processing is a process of amplifying a higher-order formant component including at least a second-order formant component in the voice data. The information processing method according to claim 1.

10. The visualization image includes a label image indicating each formant component. The information processing method according to claim 9.

11. At least one processor performs a first acquisition process of acquiring voice data indicating the voice of a speaker, a second acquisition process of acquiring clarified voice data generated by voice processing for clarifying the voice with respect to the voice data, and a visualization control process for displaying a visualization image obtained by visualizing the voice processing based on the spectrum of the voice data and the spectrum of the clarified voice data, including executing wherein the voice data indicates the voice of a speaker in a call made by connecting a plurality of user terminals, in the visualization control process, the at least one processor displays the visualization image on the user terminal of the recipient among the plurality of user terminals. Information processing method.

12. comprising the at least one processor, An information processing apparatus, wherein the at least one processor executes each process included in the information processing method according to any one of claims 1 to 11.

13. A program for causing the at least one processor to execute each process included in the information processing method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Communication device

    JP1993075988A

  • Telephone apparatus

    JP2008271481A

  • Remote conference apparatus and remote conference method

    JP2011193374A

  • Voice clarification device and voice clarifying method

    JP2021117359A

  • Telephone and voice processor

    JP2021108429A