Signal processing method, signal processing device and storage medium
By measuring the network delay time and calculating the upper limit of the allowable delay time, and selecting the signal processing method with the longest delay time, the contradiction between signal processing accuracy and discomfort in the prior art is solved, and high-precision signal processing is achieved.
Patent Information
- Application Number
- CN202210020817.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-01-12
- Filing Date
- 2022-01-10
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2042-01-10
AI Technical Summary
The prior art reduces discomfort when setting encoder parameters with a minimum delay time, but may lead to a decrease in signal processing accuracy.
By measuring the network delay time, a tolerable upper limit value of delay time is calculated, and a signal processing method with the longest delay time below the upper limit value is selected to perform signal processing.
It realizes the optimal signal processing accuracy without causing discomfort to users and adapts to different communication environments.
Smart Images

Figure CN114765796B_ABST
Abstract
Description
Technical Field
[0001] One embodiment of the present invention relates to a signal processing method, a signal processing device, and a signal processing program for processing an audio signal or an image signal. Background Art
[0002] Patent Document 1 discloses a configuration in which the delay time of wireless communication is measured and an encoder parameter with the shortest delay time is set among a plurality of encoder parameters.
[0003] Patent Document 1: Japanese Patent Application Laid-Open No. 2014-120830
[0004] The delay in communicating with other remote devices includes network delay and delay caused by signal processing. If the sum of these delays exceeds a specified time, users will feel uncomfortable.
[0005] The configuration of Patent Document 1 sets encoder parameters for the minimum delay time, so there is less chance of discomfort. However, the configuration of Patent Document 1 sets encoder parameters for the minimum delay time, so there is a possibility of reduced accuracy in signal processing. Summary of the Invention
[0006] Therefore, one object of one embodiment of the present invention is to provide a signal processing method, a signal processing device, and a signal processing program that do not cause discomfort to users and improve the accuracy of signal processing.
[0007] A signal processing method according to one embodiment of the present invention measures a network delay time with other devices connected via a network, obtains an input signal, calculates an allowable upper limit of the delay time generated in an output signal relative to the input signal due to signal processing based on the measured network delay time and the allowable total delay time, selects a signal processing with the longest delay time below the upper limit, processes the input signal using the selected signal processing, and sends the processed input signal as the output signal to the other device.
[0008] Effects of the Invention
[0009] According to one embodiment of the present invention, the accuracy of signal processing can be improved without causing discomfort to the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 It is a block diagram showing the structure of the signal processing device 1.
[0011] Figure 2A It is a block diagram showing the functional structure of the processor 12.
[0012] Figure 2B It is a block diagram showing the functional structure of the processor 12.
[0013] Figure 3 This is a flowchart showing the operation of the signal processing program 141.
[0014] Figure 4 This is a flowchart showing the detailed operation of measuring network delay time.
[0015] Figure 5 This is a flowchart showing the network delay time measurement operation according to the modification.
[0016] Figure 6 This is a block diagram showing the configuration of a signal processing device 1A according to Modification 1.
[0017] Figure 7 1A is a block diagram showing the functional configuration of the processor 12 of the signal processing device 1A. DETAILED DESCRIPTION
[0018] Figure 1 1 is a block diagram showing the configuration of the signal processing device 1. The signal processing device 1 includes a communication unit 11, a processor 12, a RAM 13, a flash memory 14, a microphone 15, an amplifier 16, and a speaker 17.
[0019] The signal processing device 1, for example, constitutes a remote conversation device that connects to another device at a remote location to transmit and receive voice data. The signal processing device 1 performs predetermined signal processing on the voice signal received by the microphone 15. The signal processing device 1 transmits the processed voice signal as voice data to the remote end. Furthermore, the signal processing device 1 outputs sound from the speaker 17 based on the voice signal of the voice data received from the remote end.
[0020] The communication unit 11 is connected to a remote dialogue device on a remote side via a network, and transmits and receives voice data to and from the remote dialogue device on the remote side.
[0021] The processor 12 performs various operations by reading programs from the flash memory 14 as a storage medium and temporarily storing them in the RAM 13. The programs include a signal processing program 141. In addition to the above programs, the flash memory 14 also stores programs for operating the processor 12, such as firmware.
[0022] The microphone 15 is an example of an input signal acquisition unit, and acquires various sounds such as a speaker's voice and noise as audio signals. The microphone 15 digitally converts the acquired audio signals and outputs the digitally converted audio signals to the processor 12.
[0023] The processor 12 performs predetermined signal processing on the sound signal received by the microphone 15. For example, the processor 12 performs noise reduction processing on the sound signal received by the microphone 15. Furthermore, the processor 12 performs echo reduction processing on the sound signal received by the microphone 15. The processor 12 transmits the processed sound signal as voice data to the remote end via the communication unit 11. Furthermore, the processor 12 outputs the voice data received via the communication unit 11 as a voice signal to the amplifier 16.
[0024] The amplifier 16 performs analog conversion on the sound signal received from the processor 12 and amplifies the sound signal. The amplifier 16 outputs the amplified sound signal to the speaker 17. The speaker 17 outputs sound based on the sound signal output from the amplifier 16.
[0025] The processor 12 implements the sound signal processing method of the present invention. Figure 2A This is a block diagram showing the functional configuration of the processor 12. The processor 12 functionally comprises a buffer 121, a noise removal unit 122, a transmitter 123, a receiver 124, a measurer 125, and a delay time calculator 126. These configurations are implemented by a signal processing program 141.
[0026] The buffer 121 temporarily stores the sound signal obtained by the microphone 15 for a predetermined period of time. The noise removal unit 122 is an example of a signal processing unit, and performs noise removal processing using the sound signal stored in the buffer 121. The sending unit 123 sends the sound signal after noise removal by the noise removal unit 122 as voice data to the device of the connection target. The receiving unit 124 receives the voice data from the device of the connection target and outputs it to the amplifier 16 as a sound signal. The measuring unit 125 measures the network delay time. The delay time calculation unit 126 calculates the allowable upper limit value of the delay time generated by the input signal in the output signal through the signal processing performed by the signal processing program 141 based on the network delay time. In addition, the delay time calculation unit 126 selects the signal processing with the longest delay time below the upper limit value.
[0027] Figure 3 1 is a flowchart showing the operation of the signal processing program 141. The measuring unit 125 measures the network delay time (S11). Figure 4 This is a flowchart showing the detailed operation of measuring network delay time. The measuring unit 125 first transmits a first DTMF (Dual-Tone Multi-Frequency) signal as a test signal to the target device via the transmitting unit 123 and records the transmission time (S101). The first DTMF signal is embedded in the payload of a VoIP (Voice over Internet Protocol) signal, for example.
[0028] The target device receives a first DTMF signal (S201). The target device sends back a second DTMF signal as a response to the first DTMF signal (S202). The second DTMF signal is also embedded in the payload of the VoIP service, for example. The measuring unit 125 receives the second DTMF signal via the receiving unit 124 and records the time of reception (S102). The measuring unit 125 measures the network delay time based on the difference between the recorded sending time and receiving time (S103).
[0029] Network latency corresponds to the time difference between the transmission of certain data and the reception of that data by the destination device. The difference between the transmission and reception times recorded by measurement unit 125 is the time difference between the transmission of certain data and the reception of a response. Therefore, measurement unit 125 sets half of the difference between the transmission and reception times as the network latency.
[0030] The network delay time can be measured during a conversation, but it is preferably measured immediately after the connection between the devices is established. This prevents the measurement unit 125 from interrupting the conversation due to the sound generated by the DTMF signal.
[0031] When measuring the network delay time during a conversation, the measuring unit 125 preferably embeds a test signal in a high frequency band (eg, a frequency band of approximately 20 kHz) to avoid affecting the user's conversation.
[0032] Alternatively, measurement unit 125 can measure network latency by adding specific frequency or phase characteristics to the conversational audio signal. For example, measurement unit 125 assigns a dip to a specific frequency (e.g., 1kHz) of the audio signal. The destination device, upon detecting the dip at that frequency, responds with a feedback signal. This feedback signal can be the aforementioned second DTMF signal, or it can be a signal that assigns specific frequency or phase characteristics to the conversational audio signal.
[0033] Furthermore, the measurement unit 125 may embed special information corresponding to the first DTMF signal into the header of an RTP (Real-time Transport Protocol) packet, rather than the payload within VoIP. The destination device extracts the special information from the header of the RTP packet and returns it. The return may be the second DTMF signal described above, or information for return may be embedded in the header of the RTP packet.
[0034] Alternatively, the measurement unit 125 may obtain the transmission time of the packet data received from the connection destination device from a remote communication program (a program for transmitting and receiving voice data). Figure 5 This is a flowchart showing the network delay time measurement operation according to a modification example. In this modification example, a remote communication program transmits voice data with a transmission time.
[0035] The destination device transmits voice data with a transmission time (S301). The measurement unit 125 receives the voice data via the receiving unit 124 and records the reception time (S401). The measurement unit 125 extracts the transmission time from the received voice data (S402). The measurement unit 125 then calculates the network delay time based on the difference between the extracted transmission time and the recorded reception time (S403).
[0036] In this example, the remote conversation program transmits voice data with the transmission time, eliminating the need for sending and receiving test signals such as DTMF signals. Furthermore, in this example, the measurement unit 125 utilizes the time information assigned to the conversation voice data, so even if measurement is performed during a conversation, it does not affect the user's conversation.
[0037] Back to Figure 3 The delay time calculation unit 126 calculates an upper limit based on the network delay time measured by the measurement unit 125 (S12). For example, the upper limit corresponds to the difference between the total delay time (e.g., 200 msec) that is permissible to the user without causing discomfort and the network delay time. When the network delay time is large, the upper limit is shortened, and when the network delay time is small, the upper limit is lengthened.
[0038] Then, the delay time calculation unit 126 selects the signal processing having the longest delay time below the calculated upper limit value (S13). Figure 2A In this example, delay time calculation unit 126 changes the buffer size of buffer 121 without changing the processing content of noise suppression unit 122. Specifically, delay time calculation unit 126 sets the buffer size to the maximum value within the upper limit. Noise suppression unit 122 performs noise suppression processing using the audio signal temporarily stored in the set maximum buffer size (S14). Transmission unit 123 transmits the noise-suppressed audio signal to the destination device (S15).
[0039] Noise removal processing is an example of a process that determines whether it is a target signal and allows the target signal to pass. The noise removal process allows the target sound (voice) to pass and removes other sounds as noise. For example, the noise removal process is a filtering process that converts a certain input signal into a certain output signal using a prescribed algorithm such as a trained neural network (especially, a convolutional neural network (CNN (Convolutional Neural Network)), a recurrent neural network (RNN (Recurrent Neural Network)) or an LSTM (Long-Short Term Model)). The algorithm for the filtering process is constructed through machine learning. The noise removal unit 122 repeatedly performs the process of converting a certain input sound signal into a sound signal from which noise has been removed and learns in advance to construct a trained model. The noise removal unit 122 performs noise removal processing using the trained model.
[0040] The accuracy of noise removal using this trained neural network depends on the amount of information in the input signal. The greater the amount of information in the input signal, the higher the accuracy of the noise removal process. In this embodiment, the delay time calculation unit 126 sets the buffer size to its maximum value when the buffer size is below the upper limit. Therefore, the accuracy of the noise removal unit 122 is set to the highest value when the buffer size is below the upper limit.
[0041] As described above, when network delay is high, the upper limit is shortened, while when network delay is low, the upper limit is lengthened. In other words, the signal processing device 1 of this embodiment performs high-precision noise reduction processing in a good communication environment, and even in a poor communication environment, performs noise reduction processing with minimal delay to a level that does not cause user discomfort. Therefore, the signal processing device 1 can perform optimal noise reduction processing tailored to the communication environment.
[0042] In the above embodiment, as an example of selecting the signal processing with the longest delay time below the upper limit, the buffer size of buffer 121 is set to the maximum without changing the processing content of noise reduction unit 122. However, delay time calculation unit 126 may also change the content of the signal processing of noise reduction unit 122. For example, delay time calculation unit 126 may also change the algorithm based on the upper limit.
[0043] For example, Figure 2BAs shown, the processor 12 may not include the buffer 121 and may directly input the sound signal obtained by the microphone 15 to the noise removal unit 122. In this case, the delay time calculation unit 126 may also change the content of the signal processing of the noise removal unit 122. For example, the delay time calculation unit 126 may select a signal processing method such as a recurrent neural network or LSTM that has the longest delay time below the upper limit value. Recurrent neural networks and LSTMs store internal variables, so they can also be configured to explicitly not include a buffer for storing the sound signal obtained by the microphone 15.
[0044] In the above embodiment, noise reduction processing is described as an example of signal processing. However, signal processing is not limited to noise reduction processing. For example, echo reduction processing may also be performed as signal processing. In echo reduction processing, the delay time calculation unit 126 also sets the buffer size to the maximum value if the buffer size is below the upper limit.
[0045] In addition, signal processing may also include voice recognition processing and conversion into text data. In addition, signal processing may also include determining whether the voice is that of a specific speaker (voice recognition) and emphasizing or removing the voice of a specific speaker.
[0046] Furthermore, signal processing is not limited to processing of sound signals. Figure 6 1A is a block diagram showing the structure of a signal processing device 1A according to Modification 1. Figure 1 The same reference numerals are used to designate common configurations, and their descriptions are omitted. The signal processing device 1A further includes a display 18 and a camera 19 in comparison with the signal processing device 1 .
[0047] Figure 7 This is a block diagram showing the functional structure of the processor 12 of the signal processing device 1A. Figure 2A The processor 12 of the signal processing device 1A includes an auto-framing processing unit 152 in place of the noise removal unit 122. The other configurations are the same as those of the processor 12 of the signal processing device 1.
[0048] Buffer 121 stores the video signal captured by camera 19 for a predetermined period of time. Automatic frame capture processing unit 152 performs automatic frame capture processing to cut out and magnify the face of the speaker in the video signal stored in buffer 121. Automatic frame capture processing is also an example of processing that determines whether the signal is a target signal and allows the target signal to pass.
[0049] More specifically, automatic framing is the process of performing facial recognition (image recognition) and cropping the recognized face portion. Automatic framing can also crop the facial image of a specific speaker. Furthermore, automatic framing can crop only the facial image of the speaker in the conversation.
[0050] Similar to noise removal, automatic framing is a filtering process that converts an input signal into an output signal using a predetermined algorithm such as a neural network. The algorithm for automatic framing is also constructed through machine learning.
[0051] The accuracy of the automatic framing process using such a neural network also depends on the amount of information in the input signal. The delay time calculation unit 126 sets the buffer amount to the maximum value below the upper limit. Therefore, the accuracy of the automatic framing processing unit 152 is set to the highest accuracy below the upper limit. In addition, the delay time calculation unit 126 may also change the algorithm of the automatic framing process according to the upper limit. Similar to the above, the processor 12 may not have a buffer and directly input the image signal obtained by the camera 19 to the automatic framing processing unit 152. In this case, the delay time calculation unit 126 selects a signal processing method such as a regression neural network or LSTM that has the longest delay time below the upper limit.
[0052] The signal processing device 1A performs high-precision automatic framing in good communication environments, and even in poor communication environments, performs automatic framing without delay, to a degree that does not cause discomfort to the user. Therefore, the signal processing device 1A can perform the optimal automatic framing according to the communication environment.
[0053] The description of the present embodiment is illustrative in all aspects and is not intended to be restrictive. The scope of the present invention is not indicated by the above-described embodiment but by the claims. Furthermore, the scope of the present invention includes all modifications within the meaning and scope equivalent to the claims.
[0054] Description of the label
[0055] 1. 1A…Signal processing device, 11…Communication unit, 12…Processor, 13…RAM, 14…Flash memory, 15…Microphone, 16…Amplifier, 17…Speaker, 18…Display, 19…Camera, 121…Buffer, 122…Noise reduction unit, 123…Transmitter, 124…Receiver, 125…Measurement unit, 126…Delay time calculation unit, 141…Signal processing program, 152…Automatic frame processing unit.
Claims
1. A signal processing method, It measures the network latency with other devices connected via the network. Get the input signal, calculating an upper limit value of a delay time allowed in an output signal relative to the input signal due to signal processing based on the measured network delay time and the allowable total delay time, Select the signal processing with the longest delay time below the upper limit value, processing the input signal by the selected signal processing, The input signal after signal processing is transmitted to the other device as the output signal.
2. The signal processing method according to claim 1, wherein: The signal processing includes temporarily storing the input signal. The selection includes a process of temporarily storing the input signal for a maximum time period below the upper limit value.
3. The signal processing method according to claim 1 or 2, wherein: The signal processing includes determining whether the input signal is a target signal based on the input signal and allowing the target signal to pass. The signal processing method according to claim 3 , wherein: The determination is performed by a machine-trained neural network. The signal processing method according to claim 3 , wherein: The determination includes determining whether it is speech or noise. The signal processing method according to claim 5 , wherein: The signal processing includes processing for removing the noise.
7. The signal processing method according to claim 3, wherein: The determination includes facial recognition, The input signal includes an image signal, An automatic frame extraction process is performed to cut out the facial image recognized by the facial recognition in the video signal. The signal processing method according to claim 1 , wherein: The network delay time is measured based on information included in a protocol used in communication with the other device.
9. The signal processing method according to claim 1, wherein: The measurement is performed at the beginning of the connection with the other device.
10. A signal processing device comprising: a measuring unit that measures a network delay time with other devices connected via the network; an input signal acquiring unit that acquires an input signal; a time calculation unit that calculates an allowable upper limit value of a delay time generated in an output signal with respect to the input signal due to signal processing, based on the network delay time measured by the measurement unit and the allowable total delay time; a signal processing unit that selects a signal process having the longest delay time below the upper limit value calculated by the time calculation unit, and processes the input signal by the selected signal process; and A transmitting unit transmits the input signal after signal processing as the output signal to the other device. The signal processing device according to claim 10 , wherein: The signal processing includes temporarily storing the input signal. The selection includes a process of temporarily storing the input signal for a maximum time period below the upper limit value.
12. The signal processing device according to claim 10 or 11, wherein: The signal processing includes determining whether the input signal is a target signal based on the input signal and allowing the target signal to pass.
13. The signal processing device according to claim 12, wherein: The determination is performed by a machine-trained neural network.
14. The signal processing device according to claim 12, wherein: The determination includes determining whether it is speech or noise.
15. The signal processing device according to claim 14, wherein: The signal processing includes processing for removing the noise.
16. The signal processing device according to claim 12, wherein: The determination includes facial recognition, The input signal includes an image signal, The signal processing performs automatic frame extraction processing for cutting out a facial image recognized by the facial recognition in the video signal.
17. The signal processing device according to claim 10, wherein: The measuring unit measures the network delay time based on information included in a protocol used for communication with the other device.
18. The signal processing device according to claim 10, wherein: The measuring unit performs the measurement when the connection with the other device is started.
19. A storage medium storing a signal processing program, wherein the signal processing program causes a signal processing device to perform the following processing: Measure network latency with other devices connected via the network. Get the input signal, calculating an upper limit value of a delay time allowed in an output signal relative to the input signal due to signal processing based on the measured network delay time and the allowable total delay time, Select the signal processing with the longest delay time below the upper limit value, processing the input signal by the selected signal processing, The input signal after signal processing is transmitted to the other device as the output signal.
Citation Information
Patent Citations
Information processing device and control method of the same
JP2014120830A
Video transmission device, video transmission method, video receiving device, and video receiving method
CN103210656A
Optimizing packetization for minimal end-to-end delay in VoIP networks
US20050094628A1