Audio signal complementation method, audio signal complementation device, and program

The audio signal interpolation method and device synchronize and generate complementary audio signals to address packet loss issues, ensuring continuous and synchronized audio playback in electrical communication systems.

WO2026100076A1PCT designated stage Publication Date: 2026-05-15NT T INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NT T INC
Filing Date
2024-11-11
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing audio systems face challenges in maintaining clear and synchronized audio playback when using electrical communication methods due to packet loss and transmission delays, causing listener discomfort and audio gaps.

Method used

An audio signal interpolation method and device that generates a complementary audio signal based on synchronized audio signals from different positions to fill gaps caused by packet loss, ensuring synchronized and continuous audio playback.

Benefits of technology

Compensates for missing audio due to packet loss, maintaining clear and synchronized audio playback without significant listener discomfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024040027_15052026_PF_FP_ABST
    Figure JP2024040027_15052026_PF_FP_ABST
Patent Text Reader

Abstract

This audio signal complementation method includes: a signal acquisition step for acquiring a first audio signal and a second audio signal collected at a position different from that of the first audio signal; a complementary audio signal generation step for, when the first audio signal is missing, generating a complementary audio signal for complementing the missing first audio signal, on the basis of the second audio signal; and an output step for outputting the complementary audio signal when the first audio signal is missing, and outputting the first audio signal when the first audio signal is not missing.
Need to check novelty before this filing date? Find Prior Art

Description

Audio signal augmentation method, audio signal augmentation device, and program

[0001] This invention relates to a method for audio signal augmentation, an audio signal augmentation device, and a program.

[0002] Open-ear headphones are becoming increasingly popular. Compared to conventional headphones, open-ear headphones allow users to hear ambient sounds while simultaneously listening to the sound played through the headphones. This offers the advantage of not blocking the ears, allowing users to naturally hear ambient sounds while listening to the played sound. By utilizing the advantage of naturally hearing ambient sounds and amplifying the sound picked up near the ear before playing it back through the headphones, they are expected to be used in hearing aids and sound pickups to improve hearing.

[0003] Taking this application further, by transmitting acoustic signals picked up near the sound source using electrical communication methods such as wired or wireless connections, and then playing them back through headphones in sync with the acoustic signals that have propagated through space and arrived at the listener, it becomes possible to hear the desired sounds with far greater clarity than conventional hearing aids or sound pickups that pick up and play back sounds at the listener's ear. Such an acoustic system can be used for purposes such as hearing the voice of a specific singer, performer, or musician on stage at a concert, or hearing the voice of a speaker at a distance in a noisy environment.

[0004] To realize the above-described acoustic system, the acoustic signal picked up near the sound source must be transmitted simultaneously with the acoustic signal propagating through space using electrical communication means such as wired or wireless connections. Furthermore, considering the delays required for sound pickup processing, transmission processing, and playback processing within headphones, the acoustic signal must be transmitted electrically with a smaller delay than the acoustic signal propagating through space. Assuming that the acoustic signal propagates through the air at a speed of 300 meters per second, it would take only 20 milliseconds for the acoustic signal from a sound source 6 meters away to travel through the air and reach the listener. To realize the above-described acoustic system, the acoustic signal must be transmitted by electrical communication means with a delay far less than this.

[0005] When transmitting acoustic signals using electrical communication methods, transmission errors can occur due to interference from radio waves or congestion in the communication path. When transmitting acoustic signals as digital signals using currently prevalent packet transmission methods, transmission errors are detected as packet loss. Furthermore, even if packets are not lost, if congestion in the transmission path delays the arrival of packets, and the arriving packets do not arrive in time for playback at the receiving end, those packets cannot be used for playback and must be discarded, which can be considered an event equivalent to packet loss.

[0006] When packet loss occurs, methods to recover from signal loss generally include: 1. retransmitting the packet, 2. correcting the error using a forward error correction code (FEC), and 3. reducing auditory discomfort by supplementing the acoustic signal. However, as mentioned above, the delay allowed for transmission by electrical communication means in the above-mentioned acoustic system is extremely small. In a typical IP network, end-to-end packet transmission usually takes tens of milliseconds, so if retransmission is performed, it will not be possible to reproduce it on the receiving end in time, making it impractical to implement 1. packet retransmission in the above-mentioned acoustic system. Furthermore, 2. error correction using a forward error correction code (FEC) is also difficult to implement because the calculation and correction process of the error correction code requires calculation in blocks that store information from multiple packets, resulting in a delay due to packet buffering.

[0007] 3. Reducing auditory discomfort through acoustic signal interpolation is a method that reduces the listener's discomfort by not restoring the missing acoustic signal, but by artificially creating and reproducing acoustic information equivalent to the missing packet on the receiving side, based on the audio information contained in the packets that have arrived so far. For example, common methods include repeating the previously played sound, or, when repeating the previously played sound, attenuating the sound pressure and adding reverberation to make it sound like an echo to the listener. The method of repeating the previously played sound is disclosed, for example, in Non-Patent Document 4. However, in the case of the above acoustic system, the listener also hears the acoustic signal that has traveled through the air, so the listener feels discomfort because they perceive the difference between the artificially created reproduced sound and the acoustic signal that has traveled through the air. To eliminate this discomfort, since the listener hears the acoustic signal that has traveled through the air, one could consider a method of not reproducing the missing acoustic signal and canceling the playback. However, in this case, the sound pressure that the listener was perceiving from the sound played through the headphones suddenly drops to zero, and then just as quickly returns to normal once the gap is removed, causing the listener to experience a sense of unease regarding sound pressure.

[0008] "Transmission Control Protocol", [online], Wikipedia, [Accessed September 4, 2024, 11:07 AM], Internet <URL: https: / / ja.wikipedia.org / wiki / Transmission_Control_Protocol> "World's first technology to extract the 'voice of the person you want to hear' based on voice characteristics realized - New deep learning technology enables extraction of only specific voices in noisy environments -" [online], Nippon Telegraph and Telephone Corporation, [Accessed October 25, 2024], Internet <URL: https: / / group.ntt / jp / newsrelease / 2018 / 05 / 28 / 180528c.html> "Devised a new audio signal processing method to extract only the conversation of interest - ConceptBeam, a filter that separates and extracts voices by meaning -" [online], Nippon Telegraph and Telephone Corporation, [Accessed October 25, 2024], Internet <URL: https: / / group.ntt / jp / newsrelease / 2023 / 05 / 30 / 230530d.html〉“JT-G711 PCM Encoding Method for Voice Frequency Band Signals” [online], Information and Communications Technology Committee, [Accessed November 5, 2024], Internet <URL: https: / / www.ttc.or.jp / application / files / 9315 / 5425 / 2192 / JT-G711v6.pdf>

[0009] In view of the above circumstances, the present invention aims to provide a technology that can compensate for missing audio when audio is lost.

[0010] One aspect of the present invention is an audio signal interpolation method comprising: a signal acquisition step of acquiring a first audio signal and a second audio signal picked up at a position different from the first audio signal; a complementary audio signal generation step of generating a complementary audio signal that complements the missing first audio signal based on the second audio signal when there is a gap in the first audio signal; and an output step of outputting the complementary audio signal when there is a gap in the first audio signal, and outputting the first audio signal when there is no gap in the first audio signal.

[0011] One aspect of the present invention is an audio signal enhancement device comprising: a signal acquisition unit that acquires a first audio signal and a second audio signal picked up at a position different from the first audio signal; a supplemental audio signal generation unit that generates a supplemental audio signal to fill in the missing first audio signal based on the second audio signal when there is a gap in the first audio signal; and an output unit that outputs the supplemental audio signal when there is a gap in the first audio signal and outputs the first audio signal when there is no gap in the first audio signal.

[0012] This invention makes it possible to compensate for missing audio when audio is lost.

[0013] This is a diagram showing the configuration of the audio system 1 according to this embodiment. This is a diagram showing an example of the configuration of the audio signal receiving device 10 according to this embodiment. This is a diagram showing an example of the configuration of the audio signal supplementation device 11 according to the first embodiment. This is a flowchart showing the operation of the audio signal receiving device 10 and the audio signal supplementation device 11 according to the first embodiment. This is a diagram showing an example of the configuration of the audio signal supplementation device 11 according to the second embodiment.

[0014] Embodiments of the present invention will be described in detail below with reference to the drawings.

[0015] Figure 1 shows an example of the configuration of the audio system 1 according to this embodiment. The audio system 1 comprises an audio signal receiving device 10, an audio signal enhancement device 11, and a sound emission device 12. In the audio system 1, the audio signal receiving device 10 receives sound from a target sound source S, the audio signal enhancement device 11 enhances any missing sound, and the sound emission device 12 emits the sound with the missing parts enhanced. Details of the audio signal receiving device 10 and the audio signal enhancement device 11 will be described later.

[0016] The target sound source S is, for example, an artist on stage at a concert. The target sound source S is, for example, a speaker located at a distance from the listener U. The sound emission device 12 is, for example, headphones, which are attached to the listener U's ears. The listener U hears both the sound emitted from the target sound source S and the sound output from the sound emission device 12 via the audio signal receiving device 10 and the audio signal enhancement device 11.

[0017] Figure 2 shows an example of the configuration of the audio signal receiving device 10 according to this embodiment. The audio signal receiving device 10 includes an audio signal acquisition unit 101, a signal processing unit 102, an output unit 103, and a storage unit 109. The audio signal receiving device 10 also includes a microphone 100.

[0018] The microphone 100 picks up sound from the target sound source S and converts it into an audio signal. The microphone 100 is installed, for example, in the vicinity of the target sound source S.

[0019] The audio signal acquisition unit 101 acquires an audio signal from the microphone 100. The audio signal acquired from the microphone 100 is an analog signal. The signal processing unit 102 converts the analog audio signal into a digital signal. The signal processing unit 102 then packets the digital signal. The output unit 103 outputs the packetized audio signal to the audio signal enhancement device 11.

[0020] Figure 3 shows an example of the configuration of the audio signal enhancement device 11 according to the first embodiment. The audio signal enhancement device 11 includes a packet acquisition unit 111, an audio loss determination unit 112, an audio signal acquisition unit 113, a signal processing unit 114, an enhanced audio signal generation unit 115, an output unit 116, and a storage unit 119. The audio signal enhancement device 11 also includes a microphone 110.

[0021] The packet acquisition unit 111 acquires audio signal packets from the audio signal receiving device 10. The audio loss determination unit 112 determines whether or not there are any missing audio signals in the audio signals acquired from the audio signal receiving device 10. Missing audio signals occur due to errors in packet transmission. For example, if the sequence numbers assigned to the packets acquired from the audio signal receiving device 10 are discontinuous, it is determined that there are missing audio signals; if they are continuous, it is determined that there are no missing audio signals.

[0022] The microphone 110 picks up sound from the sound source and converts it into an audio signal. It is preferable that the microphone 110 be installed near the listener U. The microphone 110 is, for example, built into the sound emission device 12.

[0023] The voice signal acquisition unit 113 acquires a voice signal from the microphone 110. The signal processing unit 114 converts the voice signal, which is an analog signal acquired from the microphone 110, into a digital signal and packetizes it. Hereinafter, the voice signal acquired from the voice signal receiving device 10 is referred to as the "first voice signal", and the voice signal acquired from the microphone 110 is referred to as the "second voice signal".

[0024] The complementary voice signal generation unit 115 generates a complementary voice signal. Hereinafter, the method for generating the complementary voice signal will be described. The complementary voice signal is generated based on the second voice signal that is temporally synchronized with the missing first voice signal. The temporal synchronization relationship between the first voice signal and the second voice signal is determined in consideration of the difference between the time until the sound emitted from the target sound source S is input to the microphone 100 and the time until the sound emitted from the target sound source S is input to the microphone 110. Also, the temporal synchronization relationship between the first voice signal and the second voice signal is determined in consideration of the difference between the time from when the sound is input to the microphone 100 until it is transmitted to the voice signal complementing device 11 and the time from when the sound is input to the microphone 110 until it is transmitted to the voice signal complementing device 11. The time T until the sound emitted from the target sound source S is input to the microphone 100 0 and the time T from when the sound is input to the microphone 100 until it is transmitted to the voice signal complementing device 11 1 The sum T 0 + T 1 and the time T until the sound emitted from the target sound source S is input to the microphone 110 2 and the time T from when the sound is input to the microphone 110 until it is transmitted to the voice signal complementing device 11 3 The sum T 2 + T 3 Among them, by matching the longer one, the first voice signal and the second voice signal are temporally synchronized. T 0 + T 1 and T 2 + T 3 The adjustment of T and T is performed by a buffer or the like provided in the voice signal receiving device 10 or the voice signal complementing device 11.

[0025] The supplementary audio signal is generated by adjusting the sound pressure of the second audio signal. The sound pressure of the supplementary audio signal is the sound pressure obtained by adding a predetermined sound pressure to the sound pressure of the first audio signal. Furthermore, the sound pressure of the supplementary audio signal may be adjusted to be equivalent to the sound pressure of the first audio signal at a time prior to the missing first audio signal. Ideally, the sound pressure of the supplementary audio signal should be adjusted to be equivalent to the sound pressure of the first audio signal immediately before the missing first audio signal.

[0026] The output unit 116 outputs the first audio signal to the sound emission device 12. If the first audio signal is missing, the output unit 116 outputs the supplemental audio signal generated by the supplemental audio signal generation unit 115 to the sound emission device 12 as the missing first audio signal.

[0027] The sound emission device 12 decodes the audio signal acquired from the audio signal enhancement device 11 and outputs the audio. As a result, the listener U hears both the audio directly from the target sound source S and the audio output from the sound emission device 12. The two audios heard by the listener U are synchronized in time. Here, the two sounds heard by the listener U are synchronized in time. The time T from when the sound emitted from the target sound source S is heard by the listener U. 4 The time T is the time it takes for the sound emitted from the target sound source S to be input to the microphone 100. 0 The time T from when the sound is input to the microphone 100 until it is output to the sound emission device 12. 5 Japanese T 0 +T 5 To ensure that the durations are the same, the audio signal receiving device 10 and the audio signal supplementing device 11 perform a waiting period using buffers, etc. 0 +T 5 ga T 4 By adjusting to match this, the two sounds heard by listener U are synchronized in time.

[0028] Figure 4 is a flowchart showing the operation of the audio signal receiving device 10 and the audio signal enhancement device 11 according to the first embodiment. In the audio signal receiving device 10, the audio signal acquisition unit 101 acquires a first audio signal from the microphone 100 (step S11). Subsequently, the signal processing unit 102 processes the first audio signal, converts it into a digital signal, and packets it (step S12). The output unit 103 outputs the first audio signal to the audio signal enhancement device 11 (step S13).

[0029] In the audio signal enhancement device 11, the audio signal acquisition unit 113 acquires a second audio signal from the microphone 110 (step S21). Subsequently, the signal processing unit 114 processes the second audio signal, converts it into a digital signal, and packets it (step S22). Also, the packet acquisition unit 111 acquires the packetized first audio signal output from the audio signal receiving device 10 (step S23). The audio loss determination unit 112 determines whether or not there is a loss in the first audio signal (step S24).

[0030] If there is a void in the first audio signal (step S25: YES), the supplemental audio signal generation unit 115 generates a supplemental audio signal (step S26), and the output unit 116 outputs the supplemental audio signal as the missing first audio signal to the sound emission device 12 (step S27). If there is no void in the first audio signal (step S25: NO), the output unit 116 outputs the first audio signal to the sound emission device 12 (step S28).

[0031] As a result, the audio lost during transmission is compensated for, and the listener U can hear the compensated audio from the sound emission device 12. The sound pressure of the compensated audio signal does not need to be such that the listener U does not perceive any discomfort in terms of the change in sound pressure from the preceding first audio signal. In other words, the sound pressure of the compensated audio signal does not need to be the same as the sound pressure of the first audio signal acquired immediately before the lost first audio signal, and a difference in sound pressure of a predetermined magnitude (e.g., 3 decibels) or less is acceptable.

[0032] Next, a second embodiment will be described. FIG. 5 is a diagram showing a configuration example of the voice signal completion device 11 according to the second embodiment. The voice signal completion device 11 according to the second embodiment includes a voice signal extraction unit 117 in addition to the voice signal completion device 11 according to the first embodiment. The voice signal extraction unit 117 extracts a desired voice signal from the second voice signal acquired by the voice signal acquisition unit 113. The voice signal extraction unit 117 removes, for example, a frequency band with a large amount of noise components from the second voice signal. The voice signal extraction unit 117 extracts, for example, a signal of voice uttered by a desired person from the second voice signal. The voice signal extraction unit 117 extracts, for example, a signal of conversation voice for a specific purpose. The extraction of the voice signal can be performed, for example, by using the method described in Non-Patent Document 2 or 3.

[0033] In the second embodiment, the complementary voice signal generation unit 115 generates a complementary voice signal based on the voice signal extracted by the voice signal extraction unit 117 instead of the second voice signal.

[0034] In the second embodiment, by generating the complementary voice signal based on the voice signal extracted by the voice signal extraction unit 117, a complementary voice signal with better quality can be generated.

[0035] As described above, an embodiment of the present invention has been described in detail with reference to the drawings. However, the specific configuration is not limited to the above, and various design changes and the like can be made without departing from the gist of the present invention.

[0036] The time T until the sound emitted from the target sound source S enters the microphone 110 2 may be calculated according to the position of the microphone 110. Also, the time T until the listener U hears the sound emitted from the target sound source S 4 may be calculated according to the position of the sound playback device 12. The position of the sound playback device 12 may be the position of the listener U. When the microphone 110 is built into the sound playback device 12, the position of the microphone 110 may be the position of the sound playback device 12 or the position of the listener U.

[0037] For example, when it is known in advance that the listener U is seated in a specific seat in a concert hall, a theater, or the like, the distance from the sound source S to the listener U can be calculated in advance. Therefore, the time until the sound emitted from the target sound source S is input to the microphone 110 and the time until the listener U hears the sound emitted from the target sound source S can be calculated based on the position of the seat and the position of the target sound source S.

[0038] For example, when the position of the listener U changes moment by moment, such as when the listener U can walk freely in a live house or the like, by acquiring the position information from the position information transmission device (such as a beacon) possessed by the listener U or the position information transmission device attached to the microphone 110 or the sound playback device 12, the time until the sound emitted from the target sound source S is input to the microphone 110 and the time until the listener U hears the sound emitted from the target sound source S can be calculated.

[0039] The processing of the audio signal receiving device 10 and the audio signal supplementation device 11 in the above-described embodiment may be implemented by a computer using software. In that case, the program for implementing this function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be loaded into a computer system and executed. Here, "computer system" includes hardware such as the OS and peripheral devices. Furthermore, "computer-readable recording medium" refers to portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and storage devices such as hard disks built into a computer system. Moreover, "computer-readable recording medium" may also include those that dynamically hold programs for a short period of time, such as communication lines used when transmitting programs via networks such as the Internet or communication lines such as telephone lines, and those that hold programs for a certain period of time, such as volatile memory inside a computer system that acts as a server or client in such cases. Furthermore, the above-mentioned program may be for implementing a part of the above-mentioned function, or it may be a program that can implement the above-mentioned function in combination with a program already recorded in the computer system, or it may be implemented using a programmable logic device such as an FPGA (Field Programmable Gate Array).

[0040] 1 Audio system, 10 Audio signal receiving device, 100 Microphone, 101 Audio signal acquisition unit, 102 Signal processing unit, 103 Output unit, 109 Storage unit, 11 Audio signal supplementation device, 111 Packet acquisition unit, 112 Audio loss detection unit, 113 Audio signal acquisition unit, 114 Signal processing unit, 115 Supplementary audio signal generation unit, 116 Output unit, 117 Audio signal extraction unit, 119 Storage unit

Claims

1. A method for supplementing an audio signal, comprising: a signal acquisition step of acquiring a first audio signal and a second audio signal picked up at a position different from the first audio signal; a supplemental audio signal generation step of generating a supplemental audio signal that supplements the missing first audio signal based on the second audio signal when there is a gap in the first audio signal; and an output step of outputting the supplemental audio signal when there is a gap in the first audio signal, and outputting the first audio signal when there is no gap in the first audio signal.

2. The method for supplementing an audio signal according to claim 1, wherein in the step of generating a supplemented audio signal, the sound pressure of the second audio signal is adjusted to be equivalent to the sound pressure of the first audio signal prior to the missing first audio signal, thereby generating the supplemented audio signal.

3. The audio signal interpolation method according to claim 1 or 2, wherein the second audio signal is picked up near the listener of the audio signal output by the output step.

4. Whether or not there is a loss in the first audio signal is determined by whether or not the sequence number assigned to the audio signal packet is discontinuous, the audio signal interpolation method according to claim 1 or 2.

5. A method for audio signal complementation according to claim 1 or 2, further comprising: an audio signal extraction step of extracting a desired audio signal from the second audio signal, wherein in the complementary audio signal generation step, a complementary audio signal is generated based on the extracted audio signal instead of the second audio signal.

6. An audio signal enhancement device comprising: a signal acquisition unit that acquires a first audio signal and a second audio signal picked up at a position different from the first audio signal; a supplemental audio signal generation unit that generates a supplemental audio signal to fill in the missing first audio signal based on the second audio signal when there is a gap in the first audio signal; and an output unit that outputs the supplemental audio signal when there is a gap in the first audio signal, and outputs the first audio signal when there is no gap in the first audio signal.

7. A program that causes a computer to perform the audio signal interpolation method described in claim 1.