Signal Processing Method, Signal Processing Apparatus, and Signal Processing Program

By measuring network delay and selecting signal processing with the longest acceptable delay, the method enhances accuracy and user comfort in signal processing.

JP7700455B2Active Publication Date: 2025-07-01YAMAHA CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021002750
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-01-12
Publication Date
2025-07-01
Estimated Expiration
2041-01-12

AI Technical Summary

Technical Problem

Existing signal processing methods that minimize delay time compromise accuracy, leading to user discomfort when network delays exceed a certain threshold.

Method used

A method that measures network delay time, calculates an upper limit for acceptable delay, and selects signal processing with the longest delay within this limit to ensure high accuracy without causing user discomfort.

Benefits of technology

Improves signal processing accuracy while maintaining user comfort by optimizing processing based on network conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700455000001
    Figure 0007700455000001
  • Figure 0007700455000002
    Figure 0007700455000002
  • Figure 0007700455000003
    Figure 0007700455000003
Patent Text Reader

Abstract

To provide a signal processing method, a signal processing device, and a signal processing program that improve the accuracy of signal processing without giving discomfort to a user.SOLUTION: A signal processing method includes measuring a network delay time with other devices connected via a network, obtaining an input signal, calculating an allowable upper limit value of the delay time generated in the output signal with respect to the input signal by performing signal processing on the basis of the measured network delay time and the allowable total delay time, selecting the signal processing with the longest delay time that is less than or equal to the upper limit value, processing the input signal with the selected signal processing, and transmitting the input signal after signal processing to the other device as the output signal.SELECTED DRAWING: Figure 2A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One embodiment of the present invention relates to a signal processing method, a signal processing apparatus, and a signal processing program for processing an audio signal or a video signal.

Background Art

[0002] Patent Document 1 discloses a configuration for measuring a delay time in wireless communication and setting an encoder parameter having the smallest delay time among a plurality of encoder parameters.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The delay in communication with a remote other device includes a delay time due to signal processing and a network delay time. When the sum of these delay times exceeds a predetermined time, the user feels discomfort.

[0005] Since the configuration of Patent Document 1 sets the encoder parameter with the minimum delay time, the user feels less discomfort. However, since the configuration of Patent Document 1 sets the encoder parameter with the minimum delay time, the accuracy of signal processing may decrease.

[0006] Therefore, one of the objects of one embodiment of the present invention is to provide a signal processing method, a signal processing apparatus, and a signal processing program that do not give discomfort to the user and improve the accuracy of signal processing.

Means for Solving the Problems

[0007] The signal processing method according to an embodiment of the present invention measures the network delay time with other devices connected via a network, acquires an input signal, and based on the measured network delay time and an acceptable total delay time, calculates an upper limit value of an acceptable delay time among the delay times generated in the output signal with respect to the input signal by performing signal processing, selects signal processing having the longest delay time within the upper limit value, processes the input signal with the selected signal processing, and transmits the input signal after signal processing to the other device as the output signal.

Advantages of the Invention

[0008] According to an embodiment of the present invention, it is possible to improve the accuracy of signal processing without giving discomfort to the user.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2A

Figure 2B

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Embodiments for Carrying Out the Invention

[0010] FIG. 1 is a block diagram showing the configuration of the signal processing apparatus 1. The signal processing apparatus 1 includes a communication unit 11, a processor 12, a RAM 13, a flash memory 14, a microphone 15, an amplifier 16, and a speaker 17.

[0011] The signal processing apparatus 1 constitutes, for example, a remote conversation apparatus that connects to another apparatus at a remote location and transmits and receives voice data. The signal processing apparatus 1 performs predetermined signal processing on the sound signal acquired by the microphone 15. The signal processing apparatus 1 transmits the sound signal subjected to signal processing as voice data to the remote side. Further, the signal processing apparatus 1 outputs sound from the speaker 17 based on the sound signal of the voice data received from the remote side.

[0012] The communication unit 11 connects to a remote conversation apparatus on the remote side via a network and transmits and receives voice data to and from the remote conversation apparatus on the remote side.

[0013] The processor 12 reads a program from the flash memory 14, which is a storage medium, and temporarily stores it in the RAM 13 to perform various operations. The program includes a signal processing program 141. The flash memory 14 stores, among other things, operation programs for the processor 12 such as firmware.

[0014] The microphone 15 is an example of an input signal acquisition unit, and acquires various sounds such as the voice of a speaker and noise as a sound signal. The microphone 15 digitally converts the acquired sound signal. The microphone 15 outputs the digitally converted sound signal to the processor 12.

[0015] The processor 12 performs predetermined signal processing on the sound signal acquired by the microphone 15. For example, the processor 12 performs noise removal processing on the sound signal acquired by the microphone 15. Further, the processor 12 performs echo cancellation processing on the sound signal acquired by the microphone 15. The processor 12 transmits the sound signal after signal processing as voice data to the remote side via the communication unit 11. Further, the processor 12 outputs the voice data received via the communication unit 11 as a sound signal to the amplifier 16.

[0016] Amplifier 16 analog-converts and amplifies the sound signal received from Processor 12. Amplifier 16 outputs the amplified sound signal to Speaker 17. Speaker 17 outputs sound based on the sound signal output from Amplifier 16.

[0017] Processor 12 implements the sound signal processing method of the present invention. FIG. 2A is a block diagram showing the functional configuration of Processor 12. Functionally, Processor 12 includes Buffer 121, Noise Removal Unit 122, Transmission Unit 123, Reception Unit 124, Measurement Unit 125, and Delay Time Calculation Unit 126. These configurations are realized by Signal Processing Program 141.

[0018] Buffer 121 temporarily holds the sound signal acquired by Microphone 15 for a predetermined time. Noise Removal Unit 122 is an example of a signal processing unit, and performs noise removal processing using the sound signal held in Buffer 121. Transmission Unit 123 transmits the sound signal from which noise has been removed by Noise Removal Unit 122 as voice data to the connected device. Reception Unit 124 receives voice data from the connected device and outputs it as a sound signal to Amplifier 16. Measurement Unit 125 measures the network delay time. Delay Time Calculation Unit 126 calculates the upper limit value of the allowable delay time among the delay times generated in the output signal with respect to the input signal by performing signal processing in Signal Processing Program 141 based on the network delay time. Further, Delay Time Calculation Unit 126 selects the signal processing with the longest delay time within the upper limit value.

[0019] FIG. 3 is a flowchart showing the operation of Signal Processing Program 141. Measurement Unit 125 measures the network delay time (S11). FIG. 4 is a flowchart showing the detailed operation of measuring the network delay time. Measurement Unit 125 first transmits a first DTMF (Dual-Tone Multi-Frequency) signal as a test signal to the connected device via Transmission Unit 123 and records the transmission time (S101). The first DTMF signal is embedded, for example, in the payload of VoIP (Voice over Internet Protocol).

[0020] The destination device receives the first DTMF signal (S201). The destination device returns a second DTMF signal as a response to the first DTMF signal (S202). The second DTMF signal is also embedded, for example, in the payload of VoIP. The measurement unit 125 receives the second DTMF signal via the reception unit 124 and records the reception time (S102). The measurement unit 125 measures the network delay time from the difference between the recorded transmission time and reception time (S103).

[0021] The network delay time corresponds to the time difference from when a certain data is transmitted until it is received by the destination device. The difference between the transmission time and reception time recorded by the measurement unit 125 is the time difference from when a certain data is transmitted until a response is received. Therefore, the measurement unit 125 sets half of the difference between the recorded transmission time and reception time as the network delay time.

[0022] Note that the measurement of the network delay time may be performed during a conversation, but it is preferably performed immediately after establishing the connection between the devices. Thereby, the measurement unit 125 does not interfere with the conversation by the sound of the DTMF signal.

[0023] When measuring the network delay time during a conversation, the measurement unit 125 preferably does not affect the user's conversation by embedding a test signal in a high frequency band (for example, a band of about 20 kHz).

[0024] Also, the measurement unit 125 may measure the network delay time by imparting specific frequency characteristics or phase characteristics to the sound signal of the conversation sound. The measurement unit 125, for example, imparts a dip to a specific frequency (for example, 1 kHz) of the sound signal. When the destination device detects a dip at the said frequency, it makes a return. The return may be the above-mentioned second DTMF signal, or specific frequency characteristics or phase characteristics may be imparted to the sound signal of the conversation sound.

[0025] Note that the measurement unit 125 may embed special information corresponding to the first DTMF signal, for example, not in the payload within VoIP but in the header of an RTP (Real-time Transport Protocol) packet. When the destination device extracts the special information from the header of the RTP packet, it performs a reply. The reply may be the second DTMF signal described above, or information for reply may be embedded in the header of the RTP packet.

[0026] Also, the measurement unit 125 may acquire the transmission time of the packet data received from the destination device from a remote conversation program (a program for transmitting and receiving voice data). FIG. 5 is a flowchart showing the measurement operation of the network delay time according to a modified example. In this modified example, a remote conversation program transmits voice data with a transmission time.

[0027] The destination device transmits voice data with a transmission time (S301). The measurement unit 125 receives the voice data via the reception unit 124 and records the reception time (S401). The measurement unit 125 extracts the transmission time from the received voice data (S402). Then, the measurement unit 125 calculates the network delay time from the difference between the extracted transmission time and the recorded reception time (S403).

[0028] In this example, since the remote conversation program transmits voice data with a transmission time, it is not necessary to transmit and receive test signals such as DTMF signals. Also, in this example, since the measurement unit 125 uses the time information attached to the voice data of the conversation sound, even if the measurement is performed during the conversation, it does not affect the user's conversation.

[0029] Returning to FIG. 3, the delay time calculation unit 126 calculates an upper limit value based on the network delay time measured by the measurement unit 125 (S12). For example, the upper limit value corresponds to the difference between an acceptable total delay time (for example, 200 msec) that does not give discomfort to the user and the network delay time. When the network delay time is large, the upper limit value becomes short, and when the network delay time is small, the upper limit value becomes long.

[0030] Then, the delay time calculation unit 126 selects the signal processing with the longest delay time among those with calculated upper limit values or less (S13). In the example of FIG. 2A, the delay time calculation unit 126 changes the buffer amount of the buffer 121 without changing the processing content of the noise removal unit 122. That is, the delay time calculation unit 126 sets the buffer amount to the longest among those with upper limit values or less. The noise removal unit 122 performs noise removal processing using the sound signal temporarily held with the set longest buffer amount (S14). The transmission unit 123 transmits the sound signal after the noise removal processing to the destination device (S15).

[0031] The noise removal processing is an example of processing that determines whether it is a target signal and passes the target signal. The noise removal processing passes the target sound (voice) and removes other sounds as noise. For example, the noise removal processing is a filter process that converts a certain input signal into a certain output signal using a predetermined algorithm such as a learned neural network (especially, a convolutional neural network (CNN), a recurrent neural network (RNN), or a long short-term memory (LSTM)). The algorithm of the filter process is constructed by machine learning. The noise removal unit 122 repeatedly performs processing and learning for converting a certain input sound signal into a sound signal with noise removed in advance to construct a learned model. The noise removal unit 122 performs noise removal processing using the learned model.

[0032] The accuracy of the noise removal processing using such a learned neural network depends on the information amount of the input signal. The higher the information amount of the input signal, the higher the accuracy of the noise removal processing. The delay time calculation unit 126 of the present embodiment sets the buffer amount to the longest among those with upper limit values or less. Therefore, the accuracy of the noise removal unit 122 is set to the highest accuracy among those with upper limit values or less.

[0033] As described above, when the network delay time is large, the upper limit value becomes shorter, and when the network delay time is small, the upper limit value becomes longer. That is, the signal processing apparatus 1 of the present embodiment performs high-precision noise removal processing under good communication environment conditions, and also performs noise removal processing without delay to such an extent that the user does not feel discomfort even under bad communication environment conditions. Therefore, the signal processing apparatus 1 can perform optimal noise removal processing according to the communication environment.

[0034] In the above embodiment, as an example of selecting signal processing with the longest delay time within the upper limit value, an example was shown in which the buffer amount of the buffer 121 was set to the longest without changing the processing content of the noise removal unit 122. However, the delay time calculation unit 126 may change the content of the signal processing of the noise removal unit 122. For example, the delay time calculation unit 126 may change the algorithm according to the upper limit value.

[0035] For example, as shown in FIG. 2B, the processor 12 may not include the buffer 121 and directly input the sound signal acquired by the microphone 15 to the noise removal unit 122. In this case, the delay time calculation unit 126 may change the content of the signal processing of the noise removal unit 122. For example, the delay time calculation unit 126 may select signal processing such as a recurrent neural network or LSTM that has the longest delay time within the upper limit value. Since a recurrent neural network or LSTM holds internal variables, a configuration that does not explicitly have a buffer for storing the sound signal acquired by the microphone 15 is also possible.

[0036] In the above embodiment, noise removal processing was shown as an example of signal processing. However, the signal processing is not limited to noise removal processing. For example, echo cancellation processing may be performed as the signal processing. Also in echo cancellation processing, the delay time calculation unit 126 sets the buffer amount to the longest within the upper limit value.

[0037] Also, the signal processing may be a process of performing speech recognition processing and converting it into text data. Further, the signal processing may perform determination (speech recognition) as to whether the voice is that of a specific speaker, and perform processing to emphasize the voice of the specific speaker or remove the voice of the specific speaker.

[0038] Also, the signal processing is not limited to the processing of audio signals. FIG. 6 is a block diagram showing the configuration of the signal processing apparatus 1A according to Modification 1. The configurations common to those in FIG. 1 are denoted by the same reference numerals, and the description thereof is omitted. The signal processing apparatus 1A further includes a display 18 and a camera 19 with respect to the signal processing apparatus 1.

[0039] FIG. 7 is a block diagram showing the functional configuration of the processor 12 in the signal processing apparatus 1A. The configurations common to those in FIG. 2A are denoted by the same reference numerals, and the description thereof is omitted. The processor 12 of the signal processing apparatus 1A includes an autoframing processing unit 152 instead of the noise removal unit 122. Other configurations are the same as those of the processor 12 in the signal processing apparatus 1.

[0040] The buffer 121 holds the video signal captured by the camera 19 for a predetermined time. The autoframing processing unit 152 performs autoframing processing to cut out and enlarge the face of the speaker from the video signal held in the buffer 121. The autoframing processing is also an example of a process of determining whether it is a target signal and passing the target signal.

[0041] More specifically, the autoframing processing is a process of performing face recognition (image recognition) and cutting out the recognized face portion. The autoframing processing may be a process of cutting out the face image of a specific speaker. Further, the autoframing processing may be a process of cutting out only the face image of the speaker during conversation.

[0042] Similar to the noise removal processing, the autoframing processing is a filtering process of converting a certain input signal into a certain output signal using a predetermined algorithm such as a neural network. The algorithm of the autoframing processing is also constructed by machine learning.

[0043] The accuracy of the autoframing process using such a neural network also depends on the amount of information in the input signal. The delay time calculation unit 126 sets the buffer amount to the longest within the upper limit value. Therefore, the accuracy of the autoframing processing unit 152 is set to the highest accuracy within the upper limit value. Also, the delay time calculation unit 126 may change the autoframing process algorithm according to the upper limit value. Similarly to the above, the processor 12 may directly input the video signal acquired by the camera 19 to the autoframing processing unit 152 without providing a buffer. In this case, the delay time calculation unit 126 may select signal processing such as a regression neural network or LSTM that has the longest delay time within the upper limit value.

[0044] The signal processing device 1A performs high-precision autoframing processing under good communication environment conditions, and also performs autoframing processing without causing the user to feel discomfort even under poor communication environment conditions. Therefore, the signal processing device 1A can perform optimal autoframing processing according to the communication environment.

[0045] The description of this embodiment should be considered illustrative in all respects and not restrictive. The scope of the present invention is indicated by the claims rather than the above-described embodiments. Furthermore, the scope of the present invention includes the scope equivalent to the claims.

Explanation of Reference Numerals

[0046] 1, 1A... Signal processing device 11... Communication unit 12... Processor 13... RAM 14... Flash memory 15... Microphone 16... Amplifier 17... Speaker 18... Display 19... Camera 121... Buffer 122... Noise removal unit 123... Transmission unit 124… Receiver unit 125… Measurement unit 126… Delay time calculation unit 141… Signal processing program 152… Auto-framing processing unit

Claims

1. A signal processing method for a remote conversation device that connects to another device at a remote location via a network and transmits and receives audio data, comprising: measuring a network delay time with the other device connected via the network; acquiring, as an input signal, an audio signal acquired by a microphone; calculating an upper limit value of an allowable delay time of an output signal generated with respect to the input signal by performing signal processing based on a difference between an allowable total delay time that does not give discomfort to a user and the measured network delay time; temporarily holding the input signal for the longest time within the upper limit value; performing signal processing on the input signal, the accuracy of which depends on the amount of information of the input signal, based on the temporarily held input signal; transmitting the input signal after signal processing to the other device as the output signal; a signal processing method.

2. The signal processing includes determining whether it is a target signal based on the input signal and includes a process of passing the target signal. The signal processing method according to Claim 1.

3. The determination is performed by a neural network that has been machine-learned. The signal processing method according to Claim 2.

4. The determination includes determining whether it is audio or noise. The signal processing method according to Claim 2 or Claim 3.

5. The signal processing includes a process of removing the noise. The signal processing method according to Claim 4.

6. The determination includes face recognition. The input signal includes a video signal. Performing auto-framing processing to cut out a face image recognized by the face recognition from the video signal. The signal processing method according to any one of Claims 2 to 5.

7. The network delay time is measured based on information included in a protocol used for communication with the other device. The signal processing method according to any one of Claims 1 to 6.

8. The measurement is performed at the start of connection with the other device. The signal processing method according to any one of Claims 1 to 7.

9. A signal processing device corresponding to a remote conversation device that connects to another device at a remote location via a network and transmits and receives audio data, comprising: a measurement unit that measures a network delay time with another device connected via the network; an input signal acquisition unit that acquires, as an input signal, an audio signal acquired by a microphone; A time calculation unit that calculates an allowable upper limit value of the delay time generated in the output signal with respect to the input signal by performing signal processing based on the difference between an allowable total delay time that does not give discomfort to the user and the network delay time measured by the measurement unit; temporarily holds the input signal for the longest time equal to or less than the upper limit value calculated by the time calculation unit; a signal processing unit that performs signal processing on the input signal whose accuracy depends on the amount of information of the input signal based on the input signal held temporarily; a transmission unit that transmits the input signal after signal processing to the other device as the output signal; A signal processing device comprising:

10. The signal processing includes determining whether it is a target signal based on the input signal and passing the target signal. The signal processing device according to claim 9.

11. The determination is performed by a neural network that has been trained by machine learning. The signal processing device according to claim 10.

12. The determination includes determining whether it is voice or noise. The signal processing device according to claim 10 or claim 11.

13. The signal processing includes processing for removing the noise. The signal processing device according to claim 12.

14. The determination includes face recognition. The input signal includes a video signal. The signal processing performs autoframing processing to cut out a face image recognized by the face recognition from the video signal. The signal processing device according to any one of claims 10 to 13.

15. The measurement unit measures the network delay time based on information included in a protocol used for communication with the other device. The signal processing device according to any one of claims 9 to 14.

16. The measurement unit performs the measurement at the start of connection with the other device. The signal processing device according to any one of claims 9 to 15.

17. For a signal processing device corresponding to a remote conversation device that connects to another device at a remote location via a network and performs transmission and reception of voice data, measures the network delay time with the other device connected via the network; acquires a sound signal acquired by a microphone as an input signal; calculates an allowable upper limit value of the delay time generated in the output signal with respect to the input signal by performing signal processing based on the difference between an allowable total delay time that does not give discomfort to the user and the measured network delay time; acquires an input signal temporarily hold the input signal for the longest time below the upper limit value, perform signal processing on the input signal, where the accuracy depends on the information amount of the input signal, based on the temporarily held input signal, transmit the input signal after the signal processing to the other device as the output signal, A signal processing program for executing the processing.

Citation Information

Patent Citations

  • Video monitoring device

    JP2007336260A

  • Information processing device and method

    JP2010141659A

  • Broadcasting communication cooperation system

    JP2013009343A

  • Transmitter, transmission method, receiver, reception method, synchronous transmission system, synchronous transmission method, and program

    JP2013134119A

  • Information processing device and control method of the same

    JP2014120830A