Information processing method, information processing device, and program
The information processing method and device enhance communication by clarifying voice data through real-time amplification of higher-order formant components, addressing distance-dependent clarity and feedback issues in hearing aids, enabling smooth interactions for users with hearing impairments.
Patent Information
- Application Number
- JP2025502630
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2026-03-02
- Estimated Expiration
- 2044-12-27
AI Technical Summary
Existing hearing aids struggle with clear communication, especially when users with poor hearing interact with others, due to distance-dependent sound clarity and potential feedback issues.
An information processing method and device that clarifies voice data in real-time by amplifying higher-order formant components, allowing users to adjust clarity levels for both outgoing and incoming voices during calls, enhancing communication for users with hearing impairments.
Supports smooth communication between users with and without hearing impairments by clarifying voices in real-time, addressing distance-dependent clarity and feedback issues.
Smart Images

Figure 0007822516000001 
Figure 0007822516000002 
Figure 0007822516000003
Abstract
Description
[Technical Field]
[0001] FIELD OF THE DISCLOSURE This disclosure relates to audio processing for speech intelligibility. [Background technology]
[0002] Patent document 1 discloses a hearing aid that performs frequency analysis on acoustic signals that change over time to determine the signal in the frequency band with the highest level in real time, and then increases the gain of signals in a higher frequency range compared to the signals in that frequency band, thereby making sound clearer. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent No. 3731179 Summary of the Invention [Problem to be solved by the invention]
[0004] Here, when talking to someone with poor hearing, it may be difficult to communicate smoothly with the hearing aid described in Patent Document 1. For example, it is difficult for a user talking to someone with poor hearing to force the other person to wear a hearing aid. Furthermore, even if the other person is wearing a hearing aid, the receiver must be placed close to the hearing aid. For example, if the hearing aid is an ear-hook type, the received voice may not be sufficiently clear depending on the distance between the hearing aid and the earpiece. Furthermore, if the earpiece is placed too close to the hearing aid, feedback may occur.
[0005] An aspect of the present disclosure aims to realize a technology that supports smooth communication during calls between multiple users, including users with hearing impairments. [Means for solving the problem]
[0006] In order to solve the above problem, an information processing method according to one embodiment of the present disclosure includes at least one processor executing an acquisition process to acquire voice data indicating the voice of a speaker input to a user terminal of the speaker among a plurality of user terminals connected to make a call, a voice process to generate clarified voice data by performing voice processing on the voice data to clarify the voice, and a voice output control process to control the voice indicated by the clarified voice data to be output from a user terminal of a listener among the plurality of user terminals.
[0007] In order to solve the above problem, an information processing device according to one aspect of the present disclosure includes the at least one processor, and the at least one processor executes each process included in the above information processing method.
[0008] To solve the above problems, a program according to one embodiment of the present disclosure causes the at least one processor to execute each process included in the information processing method described above. Note that a non-transitory computer-readable recording medium on which the program is recorded also falls within the scope of the present disclosure. [Effects of the Invention]
[0009] According to one aspect of the present disclosure, it is possible to support smooth communication during calls between multiple users, including users with poor hearing. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a block diagram showing a configuration of an information processing system according to a first embodiment of the present disclosure. [Figure 2] FIG. 2 is a block diagram showing a functional configuration of a user terminal according to the first embodiment of the present disclosure. [Figure 3] FIG. 10 is a block diagram showing a functional configuration of another user terminal according to the first embodiment of the present disclosure. [Figure 4] 1 is a flow diagram illustrating a flow of an information processing method according to a first embodiment of the present disclosure. [Figure 5] FIG. 2 is a diagram schematically illustrating an example of a screen displayed in the first embodiment of the present disclosure. [Figure 6] FIG. 10 is a flow diagram illustrating the flow of another information processing method according to the first embodiment of the present disclosure. [Figure 7] FIG. 10 is a diagram schematically illustrating another example of a screen displayed in the first embodiment of the present disclosure. [Figure 8] FIG. 10 is a diagram schematically illustrating yet another example of a screen displayed in the first embodiment of the present disclosure. [Figure 9] FIG. 10 is a block diagram showing a configuration of an information processing system according to a second embodiment of the present disclosure. [Figure 10] FIG. 10 is a block diagram showing the functional configuration of each device constituting an information processing system according to a second embodiment of the present disclosure. [Figure 11] FIG. 10 is a flow diagram illustrating the flow of an information processing method according to a second embodiment of the present disclosure. [Figure 12] FIG. 10 is a flow diagram illustrating the flow of another information processing method according to the second embodiment of the present disclosure. [Figure 13] FIG. 10 is a block diagram showing the configuration of a telephone system according to a third embodiment of the present disclosure. [Figure 14] FIG. 10 is a block diagram showing a configuration of a voice processing device according to a third embodiment of the present disclosure. [Figure 15] FIG. 10 is a flowchart showing the flow of an information processing method according to a third embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0011] [Embodiment 1] Hereinafter, the information processing system 100 according to the first embodiment of the present disclosure will be described in detail with reference to the drawings.
[0012] (Configuration of information processing system 100) Fig. 1 is a block diagram showing the configuration of an information processing system 100. As shown in Fig. 1, the information processing system 100 includes a user terminal 1 used by a user U1 and a user terminal 2 used by a user U2. The user terminal 1 and the user terminal 2 can be connected via a network N so that the users U1 and U2 can make calls. Note that, although Fig. 1 shows one user terminal 1 and one user terminal 2, there may be more than one of each.
[0013] Here, a call refers to the exchange of at least voice in real time between multiple users via a network N, and may be a voice call in which only voice is exchanged, or a video call in which voice and video are exchanged. The call method may be a circuit switching method via a public line network or the like, or a packet switching method via an IP (Internet Protocol) network or the like, but is not limited to these.
[0014] The network N may be, for example, a public line network, a dedicated line network, a mobile communication network, an IP network, a wireless or wired LAN (Local Area Network), or a combination of some or all of these. However, the network N is not limited to the above examples as long as it is a network that connects the user terminal 1 and the user terminal 2 so that they can make calls using a predetermined calling method.
[0015] The user terminal 1 is a terminal used by the user U1. The user terminal 1 has a function of connecting to the user terminal 2 via the network N by a predetermined communication method (for example, the circuit switching method or packet switching method described above). That is, the user U1 can communicate with the user U2 using the user terminal 1. The user terminal 1 is one aspect of the information processing device described in the claims, and has a voice clarity function that clarifies the voice of the user or the other party during a call. The user terminal 1 is a computer including at least one processor and memory. The user terminal 1 may be, for example, a mobile phone, a smartphone, a tablet, a smartwatch, a notebook computer (notebook personal computer), a desktop computer (stationary personal computer), etc., but is not limited to these.
[0016] The user terminal 2 is a terminal used by the user U2. The user terminal 2 has a function of connecting to the user terminal 1 via the network N using the same call method as the user terminal 1. In other words, the user U2 can make a call with the user U1 using the user terminal 2. In the following, the user terminal 2 is described as not having the above-mentioned clarification function, but this is not limited thereto and the user terminal 2 may have the clarification function. The user terminal 1 is a computer including at least one processor and memory. The user terminal 2 may be, for example, but is not limited to, a mobile phone, a smartphone, a tablet, a smartwatch, a laptop, a desktop computer, etc.
[0017] (Configuration of user terminal 1) Fig. 2 is a block diagram showing the functional configuration of the user terminal 1. As shown in Fig. 2, the user terminal 1 includes a control unit 110, a storage unit 120, a display unit 130, an input unit 140, an audio input unit 150, an audio output unit 160, and a communication unit 170. The control unit 110 is realized, for example, by a processor executing a program stored in a memory, and controls each unit of the user terminal 1 in an integrated manner. Details of each functional block of the control unit 110 will be described later. The storage unit 120 is, for example, configured by a memory, and stores various data and programs used by the control unit 110.
[0018] The display unit 130 displays an image generated by the control unit 110. The display unit 130 may be configured to include, for example, a liquid crystal display or an organic EL (Electro Luminescence) display, but is not limited to these. The input unit 140 accepts an operation by the user U1, acquires input information indicated by the operation, and outputs the input information to the control unit 110. The input unit 140 may be configured to include, for example, a mouse, a keyboard, a touchpad, or a combination of some or all of these, but is not limited to these. Furthermore, the display unit 130 and the input unit 140 may be integrally formed as a touch panel.
[0019] The audio input unit 150 converts an audio signal input from the outside into audio data, which is a digital signal, and outputs the audio data to the control unit 110. The audio input unit 150 may be configured to include, for example, a microphone and an AD (Analog to Digital) converter, but is not limited to this. The audio output unit 160 converts the audio data input from the control unit 110 into an audio signal, which is an analog signal, and outputs the audio signal as sound to the outside. The audio output unit 160 may be configured to include, for example, a DA converter and a speaker, but is not limited to this.
[0020] The communication unit 170 connects to the network N and communicates with the outside. The communication unit 170 transmits information input from the control unit 110 via the network N. The communication unit 170 also outputs information received via the network N to the control unit 110.
[0021] Note that some or all of the memory unit 120, display unit 130, input unit 140, audio input unit 150, audio output unit 160, and communication unit 170 may be connected as peripheral devices instead of being built into the user terminal 1.
[0022] (Functional blocks of the control unit 110) 2, control unit 110 includes a call application unit 111, an acquisition unit 112, an audio processing unit 113, an audio output control unit 114, and an clarification UI (User Interface) unit 115. Acquisition unit 112, audio processing unit 113, audio output control unit 114, and clarification UI unit 115 may constitute an clarification application that expands the call function of call application unit 111.
[0023] The call application unit 111 provides a function for making a call according to a predetermined call method. For example, in response to an operation by the user U1 to start a call, the call application unit 111 connects to the call application unit 211 of the user terminal 2 via the network N using a predetermined call method. The operation for starting the call may include, for example, an operation for specifying the call destination and an operation for instructing transmission of a call request to the call destination. The operation for starting the call may also include, for example, an operation for responding to a received call request. The call application unit 111 also transmits audio data representing the voice of the user U1 input from the audio input unit 150 to the connected user terminal 2. The call application unit 111 also receives audio data representing the voice of the user U2 from the user terminal 2 and outputs the audio data from the audio output unit 160. The call application unit 111 also terminates the connection with the user terminal 2 in response to an operation by the user U1 to instruct termination of the call.
[0024] The acquisition unit 112 executes an acquisition process. The acquisition process is a process of acquiring voice data indicating the voice of a speaker input to a user terminal of the speaker among multiple user terminals (for example, multiple user terminals including user terminal 1 and user terminal 2) connected to make a call. Hereinafter, "voice data indicating the voice of the speaker" will also be referred to as "voice data of the speaker." Note that the voice data is acquired in real time according to the progress of the speaker's speech.
[0025] For example, when user U1 speaks in a call between user U1 and user U2, user U1 is the speaker and user U2 is the listener. In this case, the acquisition unit 112 acquires voice data of user U1 (speaker) who is using the user terminal 1 via the voice input unit 150.
[0026] For example, when user U2 speaks in a call between user U1 and user U2, user U2 is the speaker and user U1 is the receiver. In this case, the acquisition unit 112 acquires the voice data of user U2 (speaker) received by the call application unit 111.
[0027] The voice processing unit 113 executes voice processing. The voice processing is processing for generating clarified voice data by executing voice processing for clarifying the voice on the voice data. Here, the clarified voice data may be generated in real time. Generating the clarified voice data in real time means that the clarified voice data is generated from the voice data in parallel with the processing for acquiring the voice data in accordance with the progress of the speech.
[0028] For example, speech processing for clarifying speech may involve amplifying higher-order formant components, including at least the second formant component, in speech data. Here, in speech spectrum analysis, multiple peak frequencies appear at integer multiples. These multiple peak frequencies are referred to as the first formant, second formant, third formant, fourth formant, etc., in ascending order of frequency. Although each peak frequency varies depending on the speaker's skeletal structure, it is generally known that in order to understand language, it is necessary to reliably hear the first to fourth formant components. Furthermore, users with insufficient hearing often have reduced hearing in bands including higher-order formant components, such as the second to fourth formant. Therefore, by amplifying the higher-order formant components, including the second formant, speech is made clearer so that users with insufficient hearing can perceive it as clear.
[0029] For example, the audio processing unit 113 may perform processing to amplify the level of a band above a lower frequency limit in the audio data. In this case, the lower frequency limit is determined so that the band above the lower frequency limit includes higher-order formant components, including at least a second formant component. The audio processing unit 113 may further determine an upper frequency limit and perform processing to amplify the level of a band above the lower frequency limit and below the upper frequency limit in the audio data. In this case, the lower and upper frequency limits are determined so that the band includes higher-order formant components, including at least a second formant component. As an example, the lower frequency limit may be 400 Hz and the upper frequency limit may be 5 kHz. The band from 400 Hz to 5 kHz is likely to include second to fourth formant components in typical human voices. However, the lower and upper frequency limits are not limited to the above example. Furthermore, for example, the frequency of a first formant component detected in real time from the audio data may be applied as the lower frequency limit. Note that the audio processing for clarifying audio is not limited to the above processing, and known processing can be applied.
[0030] Furthermore, the voice processing unit 113 may perform voice processing on the voice data of a speaker based on an operation of the speaker. For example, when user U1 speaks in a call between users U1 and U2, user U1 is the speaker and user U2 is the listener. In this case, the voice processing unit 113 performs voice processing on the voice data of user U1 based on an operation of user U1 (speaker). As a result, for example, when user U1, who is the speaker, wants user U2, who is the other party in the call, to hear his / her voice more clearly, he / she can instruct the user terminal 1 to clarify his / her voice.
[0031] Furthermore, the voice processing unit 113 may set the degree of clarity in voice processing for the voice data of a speaker based on an operation of the speaker. For example, the voice processing unit 113 sets the degree of clarity in voice processing for the voice data of user U1 (speaker) based on an operation of the user U1. This allows the speaker, user U1, to set the degree to which his or her voice should be cleared for the other party, user U2, to hear.
[0032] Furthermore, the voice processing unit 113 may perform voice processing on the voice data of the speaker based on an operation of the receiver. For example, when user U2 speaks in a call between user U1 and user U2, user U2 is the speaker and user U1 is the receiver. In this case, the voice processing unit 113 performs voice processing on the voice data of user U2 (speaker) received by the call application unit 111 based on an operation of user U1 (receiver) who is using the user terminal 1. As a result, for example, when user U1 who is the receiver finds it difficult to hear the voice of user U2, who is the other party in the call, he or she can instruct the user terminal 1 to make the voice of user U2 clearer.
[0033] Furthermore, the voice processing unit 113 may set the degree of clarity in voice processing for the voice data of the speaker based on an operation of the listener. For example, the voice processing unit 113 sets the degree of clarity in voice processing for the voice data of user U2 (speaker) based on an operation of user U1 (listener). This allows user U1, who is the listener, to set the degree to which the voice of user U2, who is the other party in the call, should be made clearer.
[0034] The degree of clarity may be, for example, a gain that amplifies the level of the band above the lower limit frequency and below the upper limit frequency. An operation for setting the degree of clarity is accepted by the clarity UI unit 115, which will be described later. For example, the degree of clarity of the voice may be set discretely in multiple stages. Alternatively, the degree of clarity of the voice may be set continuously.
[0035] The audio output control unit 114 executes audio output control processing. The audio output control processing is processing for controlling the output of audio represented by the cleared audio data from a user terminal that is a listener among multiple user terminals (e.g., multiple user terminals including user terminal 1 and user terminal 2). For example, when user U1 speaks, user U1 is the speaker and user U2 is the listener. Therefore, when cleared audio data for user U1 is generated, the audio output control unit 114 transmits the cleared audio data to user terminal 2 to output the cleared audio data from user terminal 2. Furthermore, for example, when user U2 speaks in a call between users U1 and U2, user U2 is the speaker and user U1 is the listener. Therefore, when cleared audio data for user U2 is generated, the audio output control unit 114 controls the audio output unit 160 to output the cleared audio data.
[0036] The clarity UI unit 115 accepts an operation for setting the degree of clarity of voice in voice processing. For example, the setting operation may include an operation for setting the degree of clarity for the outgoing voice, or may include an operation for setting the degree of clarity for the received voice. Here, the outgoing voice refers to the voice of user U1 when user U1 is the speaker. Also, the received voice refers to the voice heard by user U1 when user U2 is the speaker.
[0037] For example, the setting operation may include an operation of selecting one of multiple levels as the degree of clarity. The multiple levels may be, for example, three levels such as weak, medium, and strong. The degree of clarity may also include no clarity. In this case, the multiple levels may be, for example, four levels such as off, weak, medium, and strong, or two levels such as on and off. However, the names of the levels and the number of levels are not limited to these.
[0038] Furthermore, for example, the setting operation may include an operation to specify an arbitrary degree of clarity included in a continuous range. For example, the arbitrary degree of clarity may be expressed by an arbitrary numerical value from 1 to m, where clarity "off" is 1 and clarity "strong" is m times (m is a number greater than 1). Furthermore, the setting operation may be an operation on an operation object (e.g., a slider object, a volume object, etc.) that corresponds continuously to values from 1 to m. Information indicating the degree of clarity for the transmitted voice or the received voice, received by the setting operation, is stored in the storage unit 120 as setting information.
[0039] (Configuration of user terminal 2) Fig. 3 is a block diagram showing the functional configuration of the user terminal 2. As shown in Fig. 3, the user terminal 2 includes a control unit 210, a storage unit 220, a display unit 230, an input unit 240, an audio input unit 250, an audio output unit 260, and a communication unit 270. Each of these units will be described in the same manner as the functional blocks of the same name that are included in the user terminal 1.
[0040] (Functional blocks of the control unit 210) 3, the control unit 210 includes a call application unit 211. The call application unit 211 provides a function for making calls according to the same call method as the call application unit 111. Details of the call application unit 211 can be similarly explained by replacing the functional blocks with the same names in the user U1 and the user U2, the user terminal 1 and the user terminal 2, and the user terminal 1 and the user terminal 2 in the explanation of the call application unit 111.
[0041] (Flow of information processing method S1) The information processing system 100 configured as described above executes an information processing method S1. The information processing method S1 is a method in which, when user U1 is the speaker, voice processing for clarifying the voice transmitted by user U1 is executed in the terminal of user U1 who is the speaker. In other words, a processor provided in the user terminal 1 of user U1 who is the speaker executes at least the voice processing. FIG. 4 is a flow diagram illustrating the flow of information processing method S1. As shown in FIG. 4, the information processing method S1 includes steps S101 to S111.
[0042] In step S101, the call application unit 111 of the user terminal 1 connects to the call application unit 211 of the user terminal 2 based on an operation by the user U1 to start a call with the user U2. Also, in step S102, the call application unit 211 of the user terminal 2 connects to the call application unit 111 of the user terminal 1 based on an operation by the user U2. This starts a call between the user U1 and the user U2. Note that steps S101 and S102 do not necessarily have to be executed in this order, but may be executed in the reverse order, or some or all of them may be executed in parallel. Also, steps S101 and S102 do not limit which of the user U1 and the user U2 requests a call to the other.
[0043] In step S103, the clarity UI unit 115 of the user terminal 1 receives from the user U1 an operation for setting the clarity level of the transmitted voice (hereinafter also referred to as a level setting operation). The clarity level is an example of the degree of clarity and refers to a stage when the degree of clarity is set discretely. Hereinafter, an example in which the degree of clarity can be set discretely will be mainly described, but is not limited to this. Setting information including the clarity level of the transmitted voice set by the level setting operation is stored in the storage unit 120. Note that the timing of executing step S103 is not limited to after the start of the call (after step S101), but may also be executed before the start of the call (before step S101). Furthermore, the timing is not limited to immediately after the start of the call, but may also be executed at any time during the call (any time after step S104 described below). Furthermore, if the level setting operation is not received, pre-stored setting information may be applied.
[0044] Step S104 is an example of an acquisition process. In step S104, the acquisition unit 112 acquires voice data of the user U1 via the voice input unit 150. That is, it is assumed that the user U1 has spoken during a call.
[0045] Step S105 is an example of voice processing. In step S105, the voice processor 113 performs voice processing on the voice data of the user U1 in accordance with the setting information. As a result, clear voice data that has been cleared according to the clarity level indicated by the setting information is generated. Note that if the setting information indicates that the clarity level of the transmitted voice is "off," no voice processing is performed and no clear voice data is generated.
[0046] Step S106 is an example of a voice output control process. In step S106, the voice output control unit 114 transmits the cleared voice data to the user terminal 2. If the cleared voice data has not been generated, the voice data of the user U1 is transmitted instead.
[0047] In step S107, the call application unit 211 of the user terminal 2 receives the clarified voice data. The call application unit 211 also outputs the voice represented by the received clarified voice data via the voice output unit 260. As a result, user U2 hears the clarified voice of user U1. In other words, even if the user terminal 2 used by user U2, who is the other party in the call, does not have a clarification function, user U1 can have user U2 hear his or her clarified voice. Note that when voice data is received instead of the clarified voice data, the unamplified voice represented by the voice data is output.
[0048] Note that steps S104 to S107 are executed in real time in accordance with the progress of the speech by user U1. That is, user U2 hears the clarified voice of user U1 in real time in accordance with the progress of the speech by user U1. Thereafter, when user U2 speaks, user terminal 1 receives voice data of user U2 from user terminal 2 and outputs the received voice data via the audio output unit 160. As a result, user U1 hears the unclarified voice of user U2. Note that it is also possible to configure so that steps S204 to S207 of information processing method S2, which will be described later, are executed when user U2 speaks. In this case, user U1 hears the clarified voice of user U2. Furthermore, when user U1 speaks again, steps S104 to S107 are repeated.
[0049] Furthermore, if step S103 is executed again during a call to change the clarity level of the transmitted voice, the changed clarity level is referenced in the voice processing of step S105, and the changed clarity level is reflected in the voice heard by user U2 in step S107.
[0050] In step S110, the call application unit 111 of the user terminal 1 terminates the connection with the call application unit 211 of the user terminal 2 based on an operation by the user U1. Also, in step S111, the call application unit 211 of the user terminal 2 terminates the connection with the call application unit 111 of the user terminal 1. This ends the call between the user U1 and the user U2. Note that steps S110 and S111 do not necessarily have to be performed in this order, but may be performed in the reverse order, or some or all of them may be performed in parallel. Also, the operation to end the call may be performed not only by the user U1 but also by the user U2, or may be performed by both users.
[0051] (Screen example) 5 is a diagram schematically illustrating a screen example G1 displayed on the display unit 130 of the user terminal 1 in the information processing method S1. The screen example G1 provides a user interface related to a function for clarifying the user's own voice (outgoing voice) during a call. As shown in FIG. 5, the screen example G1 includes information G11 to G13 and operation objects G14-1 to G14-4 and G15.
[0052] Information G11 (for example, "Currently on a call with U2") indicates that a call is currently in progress and the other party. Information G12 (for example, "1:02") indicates the elapsed time of the call. Information G13 (for example, "Your voice is being amplified and can be heard by U2") indicates that the voice of user U1 using the user terminal 1 is being amplified.
[0053] The operation objects G14-1 "Weak", G14-2 "Medium", G14-3 "Strong", and G14-4 "Off" accept a level setting operation for selecting a clarity level for the transmitted voice. In this example, any one of four clarity levels (off, weak, medium, strong) can be selected. Furthermore, the highlighted G14-2 "Medium" among the operation objects G14-1 to G14-4 indicates that the currently selected clarity level is "Medium". Note that, although a bold frame and gray fill are shown as the highlighting mode in FIG. 5, this is not limiting. When a level setting operation for any of the operation objects G14-1 to G14-4 is accepted, setting information including the corresponding clarity level is stored in the storage unit 120. Furthermore, the operation objects G14-1 "Weak", G14-2 "Medium", G14-3 "Strong", and G14-4 "Off" can accept level setting operations at any time before the start of a call or during a call. Furthermore, level setting operations may be accepted multiple times, and in that case, setting information based on the most recent operation is stored.
[0054] The operation object G15 receives an operation from the user U1 to instruct the user U1 to end the call. When the operation on the operation object G15 is received, the call application unit 111 of the user terminal 1 ends the connection with the call application unit 211 of the user terminal 2.
[0055] By operating the example screen G1, the user U1 can clarify his / her own voice (outgoing voice) and set the degree of clarity, which allows the user U1 to communicate smoothly with the other user U2 even if the other user U2 has poor hearing.
[0056] (Flow of information processing method S2) The information processing system 100 also executes an information processing method S2. The information processing method S2 is a method in which, when a user U2 is the speaker, voice processing for clarifying the received voice heard by the user U1 is executed in the user terminal 1 of the user U1 who is the listener. In other words, a processor included in the user terminal 1 of the listener executes at least the voice processing. Fig. 6 is a flow diagram illustrating the flow of the information processing method S2. As shown in Fig. 6, the information processing method S2 includes steps S201 to S211.
[0057] Steps S201 and S202 are steps for starting a call between user U1 and user U2. Steps S201 and S202 are described in the same manner as steps S101 and S102, and therefore detailed description thereof will not be repeated.
[0058] In step S203, the clarity UI unit 115 of the user terminal 1 receives from the user U1 a level setting operation for setting the clarity level of the received voice. Setting information including the clarity level of the received voice set by the level setting operation is stored in the storage unit 120. Details of the clarity level of the received voice are explained in the same manner as for the clarity level of the transmitted voice. Note that the timing for executing step S203 is not limited to after the start of the call (after step S201), but may be executed before the start of the call (before step S201). The timing is not limited to immediately after the start of the call, but may be executed at any time during the call (any time after step S204 described below). Furthermore, if the level setting operation is not received, pre-stored setting information may be applied.
[0059] In step S204, the call application unit 211 of the user terminal 2 acquires voice data of the user U2 via the voice input unit 250. That is, it is assumed that the user U2 has spoken during the call. The call application unit 211 also transmits the acquired voice data of the user U2 to the user terminal 1.
[0060] Step S205 is an example of an acquisition process. In step S205, the acquisition unit 112 receives and acquires the voice data of the user U2.
[0061] Step S206 is an example of voice processing. In step S206, the voice processing unit 113 performs voice processing on the voice data of the user U2 in accordance with the setting information. As a result, clear voice data that has been cleared according to the clarity level indicated by the setting information is generated. Note that if the setting information indicates that the clarity level of the received voice is "off," no voice processing is performed and no clear voice data is generated.
[0062] In step S207, the audio output control unit 114 outputs the audio represented by the clarified audio data via the audio output unit 160. As a result, user U1 hears the clarified audio of user U2. That is, user U1 can hear the clarified audio of user U2, who is the other party in the call, by using the clarification function of the user terminal 1 that user U1 is using. Note that if clarified audio data has not been generated, the audio represented by the audio data of user U2 is output.
[0063] Note that steps S204 to S207 are executed in real time in accordance with the progress of the speech by user U2. That is, user U1 hears the clarified voice of user U2 in real time in accordance with the progress of the speech by user U2. Thereafter, when user U1 speaks, user terminal 1 transmits the voice data of user U1 acquired via the voice input unit 150 to user terminal 2. Furthermore, user terminal 2 outputs the received voice data via the voice output unit 260. As a result, user U2 hears the unclarified voice of user U1. Note that it is also possible to configure so that steps S104 to S107 of the information processing method S1 described above are executed when user U1 speaks. In this case, user U2 hears the clarified voice of user U1. Furthermore, when user U2 speaks again, steps S204 to S207 are repeated.
[0064] Furthermore, if step S203 is executed again during a call to change the clarity level of the received voice, the changed clarity level is referenced in the voice processing of step S206, and the changed clarity level is reflected in the voice heard by user U1 in step S207.
[0065] Steps S210 and S211 are steps for terminating the call between user U1 and user U2. Steps S210 and S211 are described in the same manner as steps S110 and S111, and therefore detailed description thereof will not be repeated.
[0066] (Screen example) Fig. 7 is a diagram schematically showing a screen example G2 displayed on the display unit 130 of the user terminal 1 in the information processing method S2. Similar to the screen example G1 shown in Fig. 5, the screen example G2 provides a user interface related to the user's own voice (outgoing voice), and also provides a user interface related to a function for clarifying the voice of the other party (received voice). The screen example G2 shown in Fig. 7 includes information G21 to G22, areas G23 and G25, and an operation object G27.
[0067] Information G21 and G22 indicate that a call is in progress and are explained in the same way as information G11 and G12 in the example screen G1. Operation object G27 "end call" is explained in the same way as operation object G15 in the example screen G1.
[0068] Area G23 is an area that provides a user interface for clarifying one's own voice. Information G231 and operation objects G232-1 to G232-4 included in area G23 are explained in the same manner as information G13 and operation objects G14-1 to G14-4 in screen example G1. Note that in screen example G2, operation object G232-4 "OFF" is highlighted, which shows an example of turning off the clarity of one's own voice.
[0069] Area G25 is an area that provides a user interface related to the clarity of the voice of the other party. Area G25 includes information G251 and operation objects G252-1 to G252-4. Information G251 (for example, "U2's voice has been cleared") indicates that the voice of user U2, who is the other party, has been cleared.
[0070] The operation objects G252-1 "Weak," G252-2 "Medium," G252-3 "Strong," and G252-4 "Off" accept a level setting operation for selecting a clarity level for the received voice. The operation objects G252-1 to G252-2 can be similarly described in the description of the operation objects G14-1 to G14-4 in the example screen G1, by substituting "received voice" for "transmitted voice." However, the number of clarity levels for the received voice does not necessarily have to be the same as the number of clarity levels for the transmitted voice. For example, the number of levels that can be set for the received voice may be greater than the number of levels that can be set for the transmitted voice. This allows the user to make their own voice clear to the other party while setting a more detailed clarity level for the voice that they themselves need to hear. Also, for example, the number of levels that can be set for the received voice may be fewer than the number of levels that can be set for the transmitted voice. This allows the user to easily set the level of clarity for the voice that the user wants to hear, while finely clarifying the user's own voice so that the other party can hear it.
[0071] Moreover, the highlighted G252-1 "Weak" among the operation objects G252-1 to G252-4 indicates that the currently selected clarity level is "Weak."
[0072] By operating the example screen G2, the user U1 can clarify one or both of his / her own voice (outgoing voice) and the voice (received voice) of the other party, user U2, and set the degree of clarity. Therefore, the user U1 can communicate smoothly with the other party even if the user U1 and / or the other party has poor hearing.
[0073] (Other screen examples) 8 is a diagram schematically illustrating a screen example G3 displayed on the display unit 130 of the user terminal 1 when setting the clarity level before starting a call. That is, the screen example G3 is displayed in step S103 or S203 when the information processing method S1 or S2 is executed before steps S101 to S102 or S201 to S202. As shown in FIG. 8, the screen example G3 includes information G31, areas G32 and G33, and an operation object G34.
[0074] Information G31 indicates that the other party attempting to start a call is user U2. Operation object G34 accepts an operation for starting a call with the other party. When an operation on operation object G34 is accepted, step S101 or S201 for starting the call is executed.
[0075] Area G32 is an area that provides a user interface for clarifying one's own voice. Operation objects G32-1 to G32-4 included in area G32 are explained in the same manner as operation objects G14-1 to G14-4 in screen example G1. However, operation objects G32-1 to G32-4 accept settings for the clarity level of the voice to be transmitted in a call that is about to start, rather than settings for the clarity level of the voice to be transmitted during a call.
[0076] Area G33 is an area that provides a user interface for clarifying the voice of the other party. Operation objects G33-1 to G33-4 included in area G33 are explained in the same manner as operation objects G252-1 to G252-4 in screen example G2. However, operation objects G33-1 to G33-4 accept settings for the clarity level of the received voice for a call that is about to start, rather than settings for the clarity level of the received voice during a call.
[0077] By operating the example screen G3, user U1 can set in advance whether to clarify one or both of the outgoing voice and the incoming voice, and the degree of clarity to be used, before starting a call. Therefore, user U1 can prepare in advance for smooth communication in a call with user U2, for example, when user U1 knows in advance that one or both of the user U1 and the callee have poor hearing.
[0078] (Modification 1 of Embodiment 1) In the information processing system 100 according to the first embodiment, the example in which the user terminal 2 does not have a clarification function has been mainly described. However, if the user terminal 2 has a clarification function, the voice processing unit 113 in the user terminal 1 is modified as follows.
[0079] For example, suppose that voice processing for clarifying a spoken voice is executed based on an operation by a user U2 (speaker) in the user terminal 2. In this case, the voice processing unit 113 of the user terminal 1 either does not accept an operation by the user U1 (listener) to execute the voice processing on the received voice, or does not execute the voice processing even if an operation by the user U1 (listener) is accepted.
[0080] For example, the audio processor 113 may determine whether audio processing for clarifying the speech has been performed in the user terminal 2 based on the spectrum of the received speech. Alternatively, for example, the audio processor 113 may determine whether audio processing for clarifying the speech has been performed in the user terminal 2 based on a received speaker clarity flag. The speaker clarity flag is a flag that is transmitted to the receiver together with, for example, audio data indicating the speech or cleared audio data. For example, the user terminal 2 may transmit to the user terminal 1 a speaker clarity flag indicating clarity on along with the cleared audio data, and may transmit to the user terminal 1 a speaker clarity flag indicating clarity off along with the non-clarified audio data. The user terminal 1 may also have a function for transmitting the speaker clarity flag. This allows the user terminal 2 to operate based on the speaker clarity flag, just like the user terminal 1.
[0081] Also, for example, suppose that voice processing for clarifying the received voice is executed based on an operation by user U2 (listener) in user terminal 2. In this case, the voice processing unit 113 of user terminal 1 either cannot accept an operation by user U1 (speaker) to execute the voice processing on the uttered voice, or does not execute the voice processing even if an operation by user U1 (speaker) is accepted.
[0082] For example, the voice processing unit 113 may determine whether voice processing for clarifying the received voice has been performed in the user terminal 2 based on a receiver-side clarification flag. The receiver-side clarification flag is, for example, a flag that is transmitted to the speaker when voice processing for clarifying the received voice has been performed. For example, the user terminal 2 may transmit a receiver-side clarification flag indicating clarification on to the user terminal 1 when voice processing for clarifying the received voice has been performed. In this case, the voice processing unit 113 of the user terminal 1 determines that voice processing for clarifying the received voice has been performed in the user terminal 2 when it receives a receiver-side clarification flag indicating clarification on. Furthermore, the voice processing unit 113 of the user terminal 1 may determine that voice processing for clarifying the received voice has not been performed in the user terminal 2 when it does not receive a receiver-side clarification flag indicating clarification on.
[0083] In this modification, for example, the screen example G2 shown in FIG. 7 is modified as follows. For example, it is assumed that it is determined that audio processing for clarifying the speech is being executed in the user terminal 2. In this case, in the screen example G2 displayed on the user terminal 1, the operation objects G252-1 to G252-4 related to the clarity of the received speech may be hidden or may be in a mode in which no operation is accepted (for example, grayed out). Furthermore, it is assumed that it is determined that audio processing for clarifying the received speech is being executed in the user terminal 2. In this case, in the screen example G2 displayed on the user terminal 1, the operation objects G232-1 to G232-4 related to the clarity of the spoken speech may be hidden or may be in a mode in which no operation is accepted (for example, grayed out).
[0084] In this modified example, voice processing to clarify the same speech is not performed twice on both the speaking and receiving sides, which has the effect of preventing the receiver from hearing unnatural speech.
[0085] (Modification 2 of Embodiment 1) In the information processing system 100 according to the first embodiment, the voice processing unit 113 is modified as follows. The voice processing unit 113 executes voice processing for clarifying the voice of a speaker based on any setting information selected from setting information previously determined according to each of a plurality of types of voice characteristics. The voice characteristics may be, for example, but are not limited to, language characteristics, dialect characteristics, gender characteristics, or personal characteristics. Here, the frequencies of higher-order formant components, including second-order formant components, that should be amplified for clarity may differ depending on the voice characteristics.
[0086] Therefore, the setting information may include information indicating the frequency band of higher-order formant components including second-order formant components according to the characteristics of the voice. For example, the first setting information may include information indicating the frequency band according to the characteristics of Japanese, and the second setting information may include information indicating the frequency band according to the characteristics of English. Note that the number of types of voice characteristics for which setting information is predefined is not limited to two, and may be three or more. Furthermore, the process of selecting one of the multiple setting information may be performed by a user operation or by a computer. The selection of setting information by a computer can be performed based on characteristics obtained by analyzing the speech voice in real time.
[0087] In this modification, for example, the example screen G1 shown in Fig. 5, the example screen G2 shown in Fig. 7, and the example screen G3 shown in Fig. 8 may be modified as follows. For example, the example screens G1, G2, and G3 may include an operation object for selecting setting information to be applied to clarifying one's own voice from a plurality of types of setting information (for example, setting information corresponding to Japanese, English, etc.). Furthermore, for example, the example screens G2 and G3 may include an operation object for selecting setting information to be applied to clarifying the voice of the other party from the plurality of types of setting information.
[0088] This modification has the effect of enabling accurate clarification according to the characteristics of the speaker's voice.
[0089] [Embodiment 2] An information processing system 100B according to a second embodiment of the present disclosure will be described below. For ease of explanation, components having the same functions as those described in the above embodiments will be denoted by the same reference numerals, and their description will not be repeated.
[0090] (Configuration of information processing system 100B) FIG. 9 is a block diagram showing the configuration of an information processing system 100B. As shown in FIG. 9, the information processing system 100B includes a user terminal 1B used by a user U1, a user terminal 2 used by a user U2, and a server 3. The user terminal 1B and the user terminal 2 can be connected via the server 3 so that the users U1 and U2 can make calls. The user terminal 1B, the user terminal 2, and the server 3 can be connected via a network N. Although FIG. 9 shows one user terminal 1B, one user terminal 2, and one server 3, there may be more than one of each. The server 3 is a computer having a function of relaying calls between the user terminal 1B and the user terminal 2 using a predetermined call method. The server 3 is one aspect of the information processing device described in the claims, and provides a voice clarity function to at least the user terminal 1B.
[0091] User terminal 1B is a terminal used by user U1, and is a modified version of user terminal 1. User terminal 1B has a user interface that uses the clarification function provided by server 3. User terminal 2 is as described above and will be described as not having the clarification function, but is not limited to this and may have the function.
[0092] (Server 3 configuration) FIG. 10 is a block diagram showing the functional configuration of each device constituting the information processing system 100B. As shown in FIG. 10, the server 3 includes a control unit 310, a storage unit 320, and a communication unit 370. The control unit 310 is realized, for example, by a processor executing a program stored in a memory, and controls each unit of the server 3. Details of each functional block of the control unit 310 will be described later. The storage unit 320 is, for example, configured by a memory, and stores various data and programs used by the control unit 310. The communication unit 370 is connected to a network N to communicate with the outside. The communication unit 370 transmits information input from the control unit 310 via the network N. The communication unit 370 also outputs information received via the network N to the control unit 310. Note that the storage unit 320 and the communication unit 370 may be connected as peripheral devices instead of being built into the server 3.
[0093] (Functional blocks of the control unit 310) As shown in FIG. 10, the control unit 310 includes a call relay unit 311, an acquisition unit 312, an audio processing unit 313, and an audio output control unit 314.
[0094] The call relay unit 311 has a function of relaying calls between multiple user terminals (e.g., user terminal 1B and user terminal 2) according to a predetermined call method. For example, the call relay unit 311 generates a call session in response to a request from one of the multiple user terminals. Furthermore, in response to a request from one of the multiple user terminals, the call relay unit 311 causes the user terminal to participate in the generated call session. Furthermore, for example, the call relay unit 311 relays voice data between the multiple participating user terminals while the call session is being generated. Furthermore, for example, in response to a request from one of the multiple user terminals, the call relay unit 311 causes the user terminal to exit the call session or terminates the call session itself.
[0095] For example, the call relay unit 311 may be at least a part of a server application that provides a two-party voice call function or a group voice call function. Furthermore, for example, the call relay unit 311 may be at least a part of a server application that provides a two-party video call function or a group video call function. The following description focuses on an example in which the call relay unit 311 relays between two user terminals, user terminal 1B and user terminal 2, but the number of user terminals that the call relay unit 311 relays may be three or more.
[0096] The acquisition unit 312 receives voice data of a speaker (for example, user U1 or user U2) from the user terminal 1B or user terminal 2, and thereby executes an acquisition process.
[0097] The voice processing unit 313 generates clear voice data by performing voice processing for clarifying the voice of the speaker (e.g., user U1 or user U2). The voice processing unit 313 will be described in the same manner as the voice processing unit 113, and therefore detailed description thereof will not be repeated.
[0098] The audio output control unit 314 transmits the clarified audio data to the user terminal of the listener (for example, user terminal 1B or user terminal 2), thereby controlling the output of audio represented by the clarified audio data from the user terminal. For example, if user U1 is the speaker and clarified audio data for user U1 has been generated, the audio output control unit 314 transmits the clarified audio data to user terminal 2. As a result, audio represented by the clarified audio data is output from user terminal 2. Furthermore, for example, if user U2 is the speaker and clarified audio data for user U2 has been generated, the audio output control unit 314 transmits the clarified audio data to user terminal 1B. As a result, audio represented by the clarified audio data is output from user terminal 1B.
[0099] (Configuration of user terminal 1B) 10, user terminal 1B includes control unit 110, storage unit 120, display unit 130, input unit 140, audio input unit 150, audio output unit 160, and communication unit 170, similar to user terminal 1, but differs in the functional blocks included in control unit 110. As storage unit 120, display unit 130, input unit 140, audio input unit 150, audio output unit 160, and communication unit 170 have been described above, detailed description thereof will not be repeated.
[0100] (Functional blocks of the control unit 110) 10, control unit 110 includes a call application unit 111 and a clarification UI unit 115. Clarification UI unit 115 may constitute a clarification application that extends the call function of call application unit 111.
[0101] The call application unit 111 provides a function for making a call in accordance with a predetermined call method. The call application unit 111 is a modified version of the call application unit 111 in the first embodiment. The following description will focus on the modified aspects of the call application unit 111, and the same aspects will not be described repeatedly. The call application unit 111 connects to a call destination via the server 3. The call application unit 111 also transmits and receives voice data to and from the call destination via the server 3. The call application unit 111 is capable of making calls with multiple call destinations, not just one call destination. For example, the call application unit 111 transmits to the server 3 a request to create a call session with one or multiple call destinations. The call application unit 111 also participates in the call session by transmitting a request to join the call session to the server 3. The call application unit 111 also transmits and receives voice data to and from other user terminals participating in the same call session via the server 3. The call application unit 111 also exits from the call session by transmitting a request to exit from the call session to the server 3. Furthermore, the call application unit 111 transmits a call session termination request to the server 3, thereby terminating the call session.
[0102] The clarification UI unit 115 receives the level setting operation and generates setting information. The level setting operation and the setting information have been described above in detail, so detailed description will not be repeated. The generated setting information is transmitted to the server 3 and stored in the storage unit 320 of the server 3.
[0103] (Configuration of user terminal 2) The configuration of user terminal 2 is as described above, and therefore detailed description will not be repeated. However, call application unit 211 is a modified version of call application unit 211 in embodiment 1. The modifications in call application unit 211 will be described in the same manner as the modifications in call application unit 111.
[0104] (Flow of information processing method S4) The information processing system 100B configured as described above executes an information processing method S4. The information processing method S4 is a method in which, when the user U1 is the speaker, the server 3 that relays the call executes voice processing to clarify the voice transmitted by the user U1. In other words, a processor provided in the server 3 that relays the call executes at least the voice processing. Note that, below, an example of a two-party call will be described using the information processing method S4. However, the information processing method S4 can also be applied to a group call involving three or more parties by regarding at least one of the user terminal 1B and the user terminal 2 as being present in plural. FIG. 11 is a flow diagram illustrating the flow of the information processing method S4. As shown in FIG. 11, the information processing method S4 includes steps S401 to S422.
[0105] In step S401, the call application unit 111 of the user terminal 1B transmits a call session creation request and a call session participation request to the server 3 based on an operation by user U1 to start a call with user U2. Also, in step S402, the call relay unit 311 of the server 3 generates a call session and allows user terminal 1B to participate. Also, in step S403, the call application unit 211 of the user terminal 2 transmits a call session participation request to the server 3 based on an operation by user U2 to answer the call with user U1. The call relay unit 311 of the server 3 allows user terminal 2 to participate in the call session. This initiates a call between user U1 and user U2. Note that steps S401 to S403 are not limited to being executed in this order, and may be executed in the reverse order, or some or all of them may be executed in parallel. Also, steps S401 to S403 do not limit which of user U1 and user U2 requests a call from the other.
[0106] In step S404, the clarity UI unit 115 of the user terminal 1B accepts a level setting operation from the user U1 to set the clarity level of the transmitted voice. The details of step S404 are the same as those of step S103. The clarity UI unit 115 also transmits the setting information to the server 3.
[0107] In step S405, the voice processing unit 313 of the server 3 stores the setting information including the clarity level of the transmitted voice in the storage unit 320.
[0108] In step S406, the call application unit 111 of the user terminal 1B acquires the voice data of the user U1 via the voice input unit 150. That is, it is assumed that the user U1 has spoken during the call. The call application unit 111 also transmits the voice data of the user U1 to the server 3.
[0109] Step S407 is an example of the acquisition process. In step S407, the acquisition unit 312 of the server 3 acquires the voice data of the user U1 by receiving it from the user terminal 1B.
[0110] Step S408 is an example of voice processing. In step S408, the voice processing unit 313 performs voice processing on the voice data of the user U1 in accordance with the setting information. As a result, clear voice data that has been cleared according to the clarity level indicated by the setting information is generated. Note that if the setting information indicates that the clarity level of the transmitted voice is "off," no voice processing is performed and no clear voice data is generated.
[0111] In step S409, the audio output control unit 314 transmits the cleared audio data to the user terminal 2. If the cleared audio data has not been generated, the audio data of the user U1 is transmitted instead.
[0112] In step S410, the call application unit 211 of the user terminal 2 receives the clarified voice data. The call application unit 211 also outputs the voice represented by the received clarified voice data via the voice output unit 260. As a result, the user U2 hears the clarified voice of the user U1. Note that if voice data is received instead of the clarified voice data, the unamplified voice represented by the voice data is output.
[0113] Note that steps S406 to S410 are executed in real time in accordance with the progress of user U1's speech. That is, user U2 hears the clarified voice of user U1 in real time in accordance with the progress of user U1's speech. Thereafter, when user U2 speaks, user terminal 1B receives from server 3 the voice data of user U2 transmitted from user terminal 2 and outputs the received voice data via audio output unit 160. As a result, user U1 hears the unclarified voice of user U2. Note that it is also possible to configure the system so that steps S506 to S510 of information processing method S5, which will be described later, are executed when user U2 speaks. In this case, user U1 hears the clarified voice of user U2. Furthermore, when user U1 speaks again, steps S406 to S410 are repeated.
[0114] Furthermore, if step S404 is executed again during a call to change the clarity level of the transmitted voice, the changed clarity level is referenced in the voice processing of step S408, and the changed clarity level is reflected in the voice heard by user U2 in step S410.
[0115] In step S420, the call application unit 111 of the user terminal 1B transmits a call session termination request to the server 3 based on the operation of user U1. In step S421, the call relay unit 311 of the server 3 terminates the call session in response to the termination request. In step S422, the call application unit 211 of the user terminal 2 performs processing to complete the withdrawal from the terminated call session. This terminates the call between user U1 and user U2. Note that steps S420 to S422 do not necessarily have to be performed in this order, and they may be performed in the reverse order, or some or all of them may be performed in parallel. In addition, the operation to terminate the call session may be performed not only by user U1 but also by user U2, or may be performed by both users. In addition, at least one of user U1 and user U2 may perform an operation to withdraw instead of an operation to terminate the call session.
[0116] According to the information processing method S4, for example, the display unit 130 of the user terminal 1B displays the screen example G1 in the same manner as the information processing method S1. The details of the screen example G1 are as described above. By operating the screen example G1, the user U1 can make his / her voice clearer so that the other party, the user U2, can hear it, even if the other party, the user U2, has insufficient hearing. As a result, this contributes to smooth communication between the users U1 and U2.
[0117] (Flow of information processing method S5) The information processing system 100B also executes information processing method S5. In information processing method S5, when user U2 is the speaker, voice processing is executed in server 3 that relays the call to clarify the received voice heard by user U1. In other words, a processor included in server 3 that relays the call executes at least the voice processing. Note that, as with information processing method S4, an example of a two-party call in information processing method S5 will be described below. However, information processing method S5 can also be applied to a group call involving three or more parties by regarding at least one of user terminal 1B and user terminal 2 as being present multiple times. FIG. 12 is a flow diagram illustrating the flow of information processing method S5. As shown in FIG. 12, information processing method S5 includes steps S501 to S522.
[0118] Steps S501 to S503 are steps for starting a conversation between user U1 and user U2. Steps S501 to S503 will be described in the same manner as steps S401 to S403, and therefore detailed description thereof will not be repeated.
[0119] In step S504, the clarity UI unit 115 of the user terminal 1B accepts a level setting operation from the user U1 to set the clarity level of the received voice. The details of step S504 are the same as those of step S203. The clarity UI unit 115 also transmits the setting information to the server 3.
[0120] In step S505, the voice processing unit 313 of the server 3 stores the setting information including the clarity level of the received voice in the storage unit 320.
[0121] In step S506, the call application unit 211 of the user terminal 2 acquires the voice data of the user U2 via the voice input unit 250. That is, it is assumed that the user U2 has spoken during the call. The call application unit 211 also transmits the voice data of the user U2 to the server 3.
[0122] Step S507 is an example of the acquisition process. In step S507, the acquisition unit 312 of the server 3 receives the voice data of the user U2 from the user terminal 2 and acquires it.
[0123] Step S508 is an example of voice processing. In step S508, the voice processing unit 313 performs voice processing on the voice data of the user U2 in accordance with the setting information. As a result, clear voice data that has been cleared according to the clarity level indicated by the setting information is generated. Note that if the setting information indicates that the clarity level of the received voice is "off," no voice processing is performed and no clear voice data is generated.
[0124] In step S509, the audio output control unit 314 transmits the cleared audio data to the user terminal 1 B. If the cleared audio data has not been generated, the audio data of the user U2 is transmitted instead.
[0125] In step S510, the call application unit 111 of the user terminal 1B receives the clarified voice data. The call application unit 111 also outputs the voice represented by the received clarified voice data via the voice output unit 160. As a result, the user U1 hears the clarified voice of the user U1. Note that if voice data is received instead of the clarified voice data, the unamplified voice represented by the voice data is output.
[0126] Note that steps S506 to S510 are executed in real time in accordance with the progress of the speech by user U2. That is, user U1 hears the clarified voice of user U2 in real time in accordance with the progress of the speech by user U2. Thereafter, when user U1 speaks, user terminal 2 receives from server 3 the voice data of user U1 transmitted from user terminal 1B and outputs the received voice data via the voice output unit 260. As a result, user U2 hears the unclarified voice of user U1. Note that it is also possible to configure steps S406 to S410 of the information processing method S4 described above to be executed when user U1 speaks. In this case, user U2 hears the clarified voice of user U1. Furthermore, when user U2 speaks again, steps S506 to S510 are repeated.
[0127] Furthermore, if step S504 is executed again during a call to change the clarity level of the received voice, the changed clarity level is referenced in the voice processing of step S508, and the changed clarity level is reflected in the voice heard by user U1 in step S510.
[0128] Steps S520 to S522 are steps for terminating the call between user U1 and user U2. Steps S520 to S522 will be described in the same manner as steps S420 to S422, and therefore detailed description thereof will not be repeated.
[0129] According to the information processing method S5, for example, the display unit 130 of the user terminal 1B displays the example screen G2 in the same manner as in the information processing method S2. The details of the example screen G2 are as described above. By operating the example screen G2, the user U1 can clearly hear the voice of the other party, user U2, even if the user U1's own hearing is insufficient. This contributes to smooth communication between the users U1 and U2. (Modification of the second embodiment) In the information processing system 100B according to the second embodiment, the audio processing unit 313 of the server 3 can be modified in the same manner as the audio processing unit 113 in the first or second modification of the first embodiment. As a result, even when the server 3 relays a call, the same effects as those in the first or second modification of the first embodiment can be achieved.
[0130] [Embodiment 3] A telephone system 100C according to a third embodiment of the present disclosure will be described below. The telephone system 100C includes a voice processing device 10C that performs voice processing to clarify transmitted or received voice. The voice processing device 10C also includes a microcomputer 1C that performs the voice processing. The microcomputer 1C is one aspect of an information processing device described in the claims. Note that the microcomputer 1C is not limited to what is called a "microcomputer," but may be any computer that includes a processor and memory and can be built into the voice processing device 10C. For ease of explanation, the following components having the same functions as those described in the above embodiments will be denoted by the same reference numerals, and their descriptions will not be repeated.
[0131] (Configuration of telephone system 100C) Fig. 13 is a block diagram showing the configuration of a telephone system 100C. As shown in Fig. 13, the telephone system 100C includes a voice processing device 10C, a handset 20, a main unit 30, and cables 41, 42, and 43. The telephone system 100C is used by a user U1 to make a call. The other party making a call via the telephone system 100C is referred to as a user U2.
[0132] (Handset 20) The handset 20 includes a microphone 21, a speaker 22, and a port P21. The microphone 21 and the speaker 22 are each connected to the port P21. The microphone 21 is an example of a transmitter, and the speaker 22 is an example of a receiver.
[0133] The microphone 21 converts the voice uttered by the user U1 into an audio signal and supplies the audio signal to the port P21. In this embodiment, the audio signal is an analog electrical signal. However, the microphone 21 may be configured to convert the voice uttered by the user U1 into a digital audio signal.
[0134] For example, an electric condenser microphone can be used as the microphone 21. A DC voltage is applied to an electric condenser microphone for driving it. Furthermore, electric condenser microphones are often connected using connectors of different types depending on the manufacturer or model.
[0135] The speaker 22 converts the audio signal supplied from the audio processing device 10C based on the audio uttered by the user U2, which is supplied from the port P21, into audio and outputs the audio. In this embodiment, the audio signal supplied from the port P21 is an analog signal and an electrical signal. However, the audio signal supplied from the port P21 may be a digital signal, and the speaker 22 may be configured to convert the digital audio signal into audio.
[0136] Port P21 is an example of a port of handset 20, and in this embodiment, a modular jack conforming to the RJ-9 standard is used. However, port P21 is not limited to a modular jack conforming to the RJ-9 standard. Any connector capable of inputting and outputting audio signals may be used as port P21.
[0137] (Main body 30) The main body 30 has ports P31 and P32. Port P31 is an example of a port on the handset 20 side of the main body 30, and in this embodiment, like port P21, it uses a modular jack that complies with the RJ-9 standard.
[0138] Port P32 is an example of a line-side port of main unit 30, and in this embodiment, a modular jack conforming to the RJ-11 standard is used. However, port P32 is not limited to a modular jack conforming to the RJ-11 standard. Port P32 may also be a modular jack conforming to the RJ-12 standard or the RJ-14 standard, for example.
[0139] The main body 30 is a main body of a telephone that uses a two-wire telephone line. The main body 30 is configured to enable two-way communication of voice signals with the main body of a telephone on the receiving side by connecting the telephone line to port P32 and connecting handset 20 to port P31. The main body 30 may be configured in the same way as the main body of an existing telephone. Therefore, in this embodiment, a description of the main body 30 will be omitted.
[0140] The main body 30 outputs an audio signal supplied to a port P31 from an audio processing device 10C (described later) to a cable 43 (described later).
[0141] (Cables 41, 42, 43) Each of the cables 41, 42, and 43 includes a plurality of signal lines, each of which is made of a conductor so as to be capable of transmitting an audio signal, which is an electrical signal.
[0142] Cable 41 connects port P11 of the voice processing device 10C described later to port P21 of the handset 20, cable 42 connects port P12 of the voice processing device 10C to port P31 of the main body 30, and cable 43 forms the end of the telephone line and is connected to port P32 of the main body 30.
[0143] In this embodiment, each of cables 41 and 42 is a cable provided with modular plugs conforming to the RJ-9 standard at both ends. Also, in this embodiment, cable 43 is a cable provided with modular plugs conforming to the RJ-11 standard at both ends. Note that in Fig. 13, the black squares shown at both ends of each of cables 41 and 42 indicate modular plugs conforming to the RJ-9 standard, and the white square shown at the end of cable 43 indicates a modular plug conforming to the RJ-11 standard.
[0144] (Sound processing device 10C) The voice processing device 10C is a device that executes voice processing to clarify the transmitted voice. FIG. 14 is a block diagram showing a detailed configuration of the voice processing device 10C. As shown in FIG. 14, the voice processing device 10C includes a microcomputer 1C, an AD (Analog-Digital) converter 180, a DA converter 190, and an adjuster 15. The voice processing device 10C is provided so as to be interposed between the handset 20 and the main body 30.
[0145] In this embodiment, the housing of the audio processing device 10C is made of aluminum. However, the material constituting the housing is not limited to aluminum, and may be metal such as copper or stainless steel. The material constituting the housing may also be primarily resin with a metal layer provided on its surface. By covering at least the surface of the housing with metal, electromagnetic waves penetrating from the outside into the inside can be blocked.
[0146] (AD converter 180, DA converter 190) The AD converter 180 converts the audio signal (corresponding to the transmitted voice) supplied to the port P11 into audio data, which is a digital signal, and outputs it to the microcomputer 1C. The DA converter 190 also converts the clarified audio data output from the microcomputer 1C into an analog audio signal and supplies it to the port P12.
[0147] (Adjustor 15) The adjuster 15 is a user interface for setting the clarity level. In this embodiment, the adjuster 15 is configured using a switch that can select one of three clarity levels (0, 1, 2). For example, 0, 1, and 2 that can be selected by the switch correspond to clarity levels "off," "weak," and "strong." For example, the adjuster 15 outputs a signal indicating one of 0, 1, and 2 selected by the switch to the microcomputer 1C.
[0148] The number of switch stages constituting adjuster 15 is not limited to the above example and can be selected as appropriate. Also, adjuster 15 can employ a volume that continuously changes the degree of clarity instead of a switch that discretely changes the degree of clarity. Also, when telephone system 100C is installed in a home, and the receiving user U2 can be identified to some extent, the degree of clarity can be set in advance, and adjuster 15 can be omitted.
[0149] (Microcomputer 1C) The microcomputer 1C includes a control unit 110 and a storage unit 120. The control unit 110 is realized, for example, by a processor executing a program stored in a memory, and controls each unit of the audio processing device 10C in an integrated manner. The storage unit 120 is, for example, configured by a memory, and stores various data and programs used by the control unit 110. The control unit 110 includes an acquisition unit 112, an audio processing unit 113, and an audio output control unit 114.
[0150] The acquisition unit 112 executes an acquisition process. In this embodiment, the acquisition unit 112 acquires voice data output from the AD converter 180. The voice data indicates the voice (transmitted voice) of the user U1 speaking into the handset 20.
[0151] The audio processing unit 113 executes audio processing, thereby generating clear audio data corresponding to the transmitted audio. Details of the audio processing unit 113 will be described in the same manner as in the first embodiment.
[0152] The audio output control unit 114 executes audio output control processing by outputting, to the DA converter 190, cleared audio data corresponding to the transmitted audio.
[0153] (Flow of information processing method S7) In the telephone system 100C configured as above, the microcomputer 1C executes an information processing method S7. The information processing method S7 is a method executed by the microcomputer 1C when the user U1 makes a call with a call partner using the telephone system 100C. Fig. 15 is a flow diagram showing the flow of the information processing method S7. As shown in Fig. 15, the information processing method S7 includes steps S701 to S704.
[0154] In step S701, the audio processing unit 113 acquires the clarity level from the adjuster 15. In step S702, the acquisition unit 112 acquires audio data (corresponding to the transmitted audio) output from the AD converter 180. In step S703, the audio processing unit 113 performs audio processing on the audio data in accordance with the clarity level to generate clear audio data corresponding to the transmitted audio. In step S703, the audio output control unit 114 outputs the clear audio data corresponding to the transmitted audio to the DA converter 190.
[0155] As a result, the cleared voice data corresponding to the transmitted voice is converted into an analog signal, and the resulting voice signal is transmitted to the user terminal of the call partner, user U2, via the main unit 30. As a result, the user terminal of user U2 is controlled to output the cleared voice of user U1. Even if the call partner, user U2, has poor hearing, user U1 can use the voice processing device 10C to clear his or her own voice so that the call partner, user U2, can hear it. As a result, this contributes to smooth communication between users U1 and U2.
[0156] (Modification of the third embodiment) The voice processing device 10C according to the third embodiment can be modified to perform voice processing for clarifying a received voice in addition to voice processing for clarifying a transmitted voice. In this case, the voice processing device 10C further includes an AD converter that converts a voice signal (corresponding to the received voice) supplied from the port P12 into a digital signal and outputs the digital signal to the microcomputer 1C. The voice processing device 10C also includes a DA converter that converts the cleared voice data (corresponding to the received voice) output from the microcomputer 1C into an analog signal and supplies the analog signal to the port P11. The voice processing device 10C also includes an adjuster that sets a clarity level for the received voice.
[0157] In this modification, the acquisition unit 112 acquires voice data supplied from the port P12 via an AD converter. The voice data indicates the voice (received voice) of the user U2, who is the other party in the call, received by the main unit 30.
[0158] The voice processing unit 113 acquires a clarity level for the received voice from the adjuster. The voice processing unit 113 then performs voice processing on the voice data of the user U2 according to the clarity level. This generates clear voice data corresponding to the received voice. Details of the voice processing unit 113 will be described in the same manner as in the first embodiment.
[0159] The audio output control unit 114 executes audio output control processing. In this embodiment, the audio output control processing is executed by supplying cleared audio data corresponding to the received voice to the port P11 via a DA converter. As a result, the cleared voice of the user U2 is output from the speaker 22.
[0160] According to this modification, even if the user U1 has poor hearing, the user U1 can hear the voice of the other party, user U2, clearly using the voice processing device 10C used by the user U1, which contributes to smooth communication between the users U1 and U2.
[0161] (Another Modification 1 of Embodiment 3) The voice processing device C according to the third embodiment may be connected to one or both of the handset 20 and the main body 30 not only by wire but also by wireless. The voice processing device 10C is not limited to being configured as a separate entity from the handset 20 and the main body 30. For example, the voice processing device 10C may be housed in the housing of the main body 30 so as to be integrated with the handset 20. The voice processing device 10C may be housed in the housing of the handset 20 so as to be integrated with the handset 20.
[0162] The telephone system 100C according to the third embodiment may include a computer having a call function instead of the main unit 30. The computer may be, for example, a mobile phone, a smartphone, a tablet, a smartwatch, a notebook computer, a desktop computer, or the like, but is not limited to these. The handset 20 is not limited to the embodiment shown in FIG. 13. For example, the handset 20 may be a headset or the like.
[0163] (Modification 2 of Embodiment 3) In the telephone system 100C according to the third embodiment, the voice processing unit 113 of the microcomputer 1C can be modified in the same manner as the voice processing unit 113 in the first or second modification of the first embodiment. As a result, the same effects as those in the first or second modification of the first embodiment can be achieved in a call using the voice processing device 10C incorporating the microcomputer 1C.
[0164] [Software implementation example] The functions of the user terminals 1, 1B, microcomputer 1C, and server 3 (hereinafter referred to as "devices") can be realized by a program that causes a computer to function as the device, and a program that causes a computer to function as each control block of the device (particularly each part included in the control units 110, 310).
[0165] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., a memory) as hardware for executing the program. The control device and storage device execute the program, thereby realizing the functions described in each of the above embodiments.
[0166] The program may be non-transitory and may be recorded on one or more computer-readable recording media. The recording media may or may not be included in the device. In the latter case, the program may be supplied to the device via any wired or wireless transmission medium.
[0167] In addition, some or all of the functions of each of the control blocks can be realized by logic circuits. For example, integrated circuits in which logic circuits that function as each of the control blocks are formed are also included in the scope of the present disclosure. In addition, the functions of each of the control blocks can also be realized by, for example, a quantum computer.
[0168] Furthermore, each process described in each of the above embodiments may be executed by AI (Artificial Intelligence). In this case, the AI may run on the control device or on another device (for example, an edge computer or a cloud server).
[0169] 〔summary〕 An information processing method according to a first aspect includes, by at least one processor, executing an acquisition process for acquiring voice data representing a voice of a speaker input to a user terminal of the speaker among a plurality of user terminals connected to perform a call, a voice process for generating clarified voice data by executing a voice process for clarifying the voice on the voice data, and a voice output control process for controlling a user terminal of a listener among the plurality of user terminals to output the voice represented by the clarified voice data. With the above configuration, the voice of the speaker is cleared in a call between multiple users, thereby achieving the effect of supporting smooth communication in a call between multiple users, including a user with insufficient hearing.
[0170] In the information processing method according to aspect 2, in accordance with aspect 1, the at least one processor executes the voice processing based on an operation by the speaker. With the above configuration, when the speaker's hearing is insufficient, the speaker can perform an operation to make his / her voice clearer, thereby allowing the speaker's voice to be heard by the other party.
[0171] In the information processing method according to aspect 3, in accordance with aspect 2, the at least one processor sets a degree of clarity in the voice processing based on an operation by the speaker. With this configuration, the speaker can set a degree of clarity for their own voice in accordance with the hearing ability of the other party.
[0172] In the information processing method according to aspect 4, in aspect 2 or aspect 3, when the voice processing is executed based on the operation of the speaker, the at least one processor either does not accept the operation of the receiver to execute the voice processing, or does not execute the voice processing even if the operation of the receiver is accepted. With the above configuration, voice processing to clarify the voice of the speaker is not performed twice on the speaker side and the receiver side, thereby making it possible to prevent unnatural received voice from being heard.
[0173] An information processing method according to aspect 5 is any one of aspects 1 to 4, wherein the at least one processor executes the voice processing based on an operation by the listener. With the above configuration, when the listener's hearing is insufficient, the listener can perform an operation to clarify the voice of the other party, thereby being able to hear the clearer voice of the other party.
[0174] In the information processing method according to aspect 6, in accordance with aspect 5, the at least one processor sets a degree of clarity in the voice processing based on an operation by the listener. With this configuration, the speaker can set a degree of clarity for the voice of the other party in accordance with his or her own hearing ability.
[0175] In the information processing method according to aspect 7, in aspect 5 or aspect 6, when the voice processing is executed based on the operation of the receiver, the at least one processor either does not accept the operation of the speaker to execute the voice processing, or does not execute the voice processing even if the operation of the speaker is accepted. With the above configuration, voice processing to clarify the voice of the speaker is not performed twice on the speaker side and the receiver side, thereby preventing unnatural received voice from being heard.
[0176] An information processing method according to aspect 8 is any one of aspects 1 to 7, wherein the at least one processor performs the voice processing based on any setting information selected from setting information predetermined according to each of a plurality of types of voice characteristics. With the above configuration, the voice of the speaker can be made clearer with greater accuracy according to the characteristics of the speaker's voice.
[0177] An information processing method according to aspect 9 is any one of aspects 1 to 8, wherein a processor included in a server that relays the call performs at least the voice processing. With the above configuration, the speaker and the recipient can clarify the speaker's voice even if their own user terminals used in the call do not have a function to perform voice processing to clarify the voice.
[0178] An information processing method according to aspect 11 is any one of aspects 1 to 8, wherein a processor included in the user terminal of the speaker executes at least the voice processing. With the above configuration, the speaker can use his / her own user terminal to clarify his / her own voice, regardless of whether the user terminal of the other party has a function to execute voice processing for voice clarity.
[0179] An information processing method according to aspect 12 is any one of aspects 1 to 8, wherein a processor included in the user terminal of the recipient executes at least the voice processing. With the above configuration, the recipient can use his / her own user terminal to make the voice of the other party clearer, regardless of whether the user terminal of the other party has a function to execute voice processing for making the voice clearer.
[0180] An information processing method according to aspect 13 is any one of aspects 1 to 12, wherein the audio processing is processing for amplifying higher-order formant components including at least second-order formant components in the audio data. With the above configuration, audio can be made clearer.
[0181] The information processing device according to aspect 14 includes the at least one processor, and the at least one processor executes each process included in the information processing method according to any one of aspects 1 to 12. With the above configuration, the same effect as any one of aspects 1 to 9 is achieved.
[0182] A program according to aspect 11 causes the at least one processor to execute each process included in the information processing method according to any one of aspects 1 to 9. With the above configuration, the same effect as any one of aspects 1 to 9 is achieved.
[0183] A non-transitory computer-readable recording medium according to aspect 12 stores a program according to aspect 12. The above configuration provides the same effect as any one of aspects 1 to 9.
[0184] The embodiments of the present disclosure provide an advantage of being able to present to a user the clarity of speech produced by speech processing for speech clarity, which contributes to achieving, for example, Goal 3 of the United Nations' Sustainable Development Goals (SDGs), "Ensure good health and well-being for all."
[0185] The present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present disclosure. [Explanation of symbols]
[0186] 1, 1B, 2 User terminal 1C microcomputer 3 Server 10C Audio Processing Device 15 Regulator 20 Handset 21 Microphone 22 speakers 30 Main Unit 41, 42, 43 Cables 100, 100B Information Processing System 100C Telephone System 110, 210, 310 Control unit 111, 211 Call App 112, 312 Acquisition Department 113, 313 Audio processing unit 114, 314 Audio output control unit 115 Clarification UI part 120, 220, 320 storage section 130, 230 display section 140, 240 input section 150, 250 Audio input section 160, 260 Audio output section 170, 270, 370 Communications Department 180 AD converter 190 DA converter 311 Call Relay Department
Claims
1. an acquisition process in which a user terminal of a speaker among a plurality of user terminals connected to make a call acquires voice data indicating the voice of the speaker input to the user terminal of the speaker; a user terminal of the speaker or a user terminal of a listener among the plurality of user terminals, performing a voice process on the voice data to clarify the voice, thereby generating clear voice data; a voice output control process in which the user terminal of the listener controls the user terminal of the listener to output the voice represented by the cleared voice data; performing When the user terminal of the speaker executes the voice processing based on the operation of the speaker, the user terminal of the listener is unable to accept an operation by the listener on the user terminal of the listener to execute the voice processing, or does not further execute the voice processing on the clarified voice data even if the user terminal of the listener accepts an operation by the listener on the user terminal of the listener. Information processing methods.
2. A server that relays calls between multiple user terminals an acquisition process of acquiring voice data indicating a voice of a speaker input to a user terminal of the speaker among the plurality of user terminals; a voice processing unit for generating cleared voice data by performing a voice processing on the voice data to clarify the voice; a voice output control process for controlling the voice represented by the cleared voice data to be output from a user terminal of a listener among the plurality of user terminals; performing The server When the voice processing is performed based on an operation of the speaker on the user terminal of the speaker, the operation of the listener on the user terminal of the listener to perform the voice processing is not accepted, or even if the operation of the listener on the user terminal of the listener is accepted, the voice processing is not further performed on the clarified voice data. Information processing methods.
3. an acquisition process in which a user terminal of a speaker among a plurality of user terminals connected to make a call acquires voice data indicating the voice of the speaker input to the user terminal of the speaker; a user terminal of the speaker or a user terminal of a listener among the plurality of user terminals, performing a voice process on the voice data to clarify the voice, thereby generating clear voice data; a voice output control process in which the user terminal of the listener controls the user terminal of the listener to output the voice represented by the cleared voice data; performing When the user terminal of the listener executes the voice processing based on the operation of the listener, the user terminal of the speaker is unable to accept an operation by the speaker on the user terminal of the speaker for executing the voice processing, or does not execute the voice processing even if the operation by the speaker on the user terminal of the speaker is accepted; Information processing methods.
4. A server that relays calls between multiple user terminals an acquisition process of acquiring voice data indicating a voice of a speaker input to a user terminal of the speaker among the plurality of user terminals; a voice processing unit for generating cleared voice data by performing a voice processing on the voice data to clarify the voice; a voice output control process for controlling the voice represented by the cleared voice data to be output from a user terminal of a listener among the plurality of user terminals; performing The server When the voice processing is executed based on an operation of the listener on the user terminal of the listener, the operation of the speaker on the user terminal of the speaker to execute the voice processing is made unacceptable, or the voice processing is not executed even when the operation of the speaker on the user terminal of the speaker is accepted. Information processing methods.
Citation Information
Patent Citations
Voice processing device and voice processing method
JP2019035894A
Hands-free speech device and method for controlling hands-free speech device
JP2020077933A
Telephone and voice processor
JP2021108429A
hearing aid
JP3731179B2