Sound signal receiving and decoding method, sound signal receiving side device, communication method, telephone system, computer program product, and recording medium
By combining the code strings of the first communication line and the second communication line in the terminal device and using the frame number matching of the mono code and the extended code, the contradiction between the high-quality decoded sound signal and the delay time in the existing technology is solved, and a high-quality decoded sound signal is achieved.
Patent Information
- Application Number
- CN201980097331.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-13
- Filing Date
- 2019-12-27
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2039-12-27
AI Technical Summary
The prior art has difficulty in simultaneously achieving high-quality decoded audio signals in a plurality of frames without generating delays, or has difficulty in achieving high-quality decoded audio signals without generating delays.
By utilizing the code strings of the first and second communication lines in the receiving and decoding steps in the terminal device, combining the mono code and the extension code, and outputting a high-quality decoded digital sound signal based on the principle of frame number matching or near matching.
Without significantly increasing the delay time, high-quality decoding of sound signals is achieved, thereby improving the quality of the sound signals.
Smart Images

Figure CN114144832B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to at least one of a decoding technology for a voice signal and a corresponding encoding technology for a voice signal in a terminal device connected to at least two communication networks having different information transmission priorities. Background Art
[0002] Patent Document 1 is a prior art technology for encoding and decoding audio signals between terminal devices connected to two communication networks with different information transmission priorities. The encoding device of Patent Document 1 performs scalable coding on the input audio signal for each specific time interval, or frame, to generate low-band code 1, a base layer code, and low-band code 2 and high-band code, extension layer codes. Low-band code 1 is included in a high-priority packet and transmitted at least to network B, which has a guaranteed bandwidth. Low-band code 2 and high-band code are included in a low-priority packet and transmitted to network A, which does not have a guaranteed bandwidth. The decoding device of Patent Document 1 monitors the passage of a time limit upon receiving a high-priority packet. Once the time limit has expired, decoding is performed using the packet that has already been received. That is, if it is assumed that the delay of network A is larger than that of network B, then the decoding device of patent document 1 essentially uses low-band code 2 and high-band code to perform decoding processing after the above-mentioned limited time has passed since the code of the base layer arrived, if low-band code 2 and high-band code also arrived, to obtain a decoded sound signal with high sound quality. If low-band code 2 and high-band code have not arrived, only low-band code 1 is used to perform decoding processing to obtain a decoded sound signal with the required minimum sound quality.
[0003] Prior art literature
[0004] Patent Literature
[0005] Patent Document 1: Japanese Patent Application Laid-Open No. 2005-117132 Summary of the Invention
[0006] Problems to be solved by the invention
[0007] In order to obtain high-quality decoded audio signals in most frames, the technique of Patent Document 1 requires setting the aforementioned time limit to a delay significantly longer than would be required if only a minimum-quality decoded audio signal were to be obtained. Consequently, the technique of Patent Document 1 presents a problem: to obtain high-quality decoded audio signals in most frames, the time limit must be set to a delay time long enough to create a sense of discomfort during double-talk. Furthermore, if the time limit is set close to zero to avoid discomfort during double-talk, the proportion of frames in which high-priority packets arrive within the time limit becomes extremely small. Consequently, the technique of Patent Document 1 presents a problem: if the time limit is set to avoid discomfort during double-talk, high-quality decoded audio signals cannot be obtained in almost all frames.
[0008] Therefore, an object of the present invention is to provide a technique for obtaining a decoded audio signal of high sound quality without significantly increasing the delay time compared to a configuration in which a decoded audio signal of only the minimum required sound quality is obtained.
[0009] Means for solving problems
[0010] One embodiment of the present invention is a sound signal receiving and decoding method, which is performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, comprising: a receiving step, in which, for each frame, when the extended code included in the second code string input from the second communication line includes an extended code with a frame number identical to that of the mono code included in the first code string input from the first communication line, the mono code included in the first code string input from the first communication line and the extended code with a frame number identical to that of the mono code are output; and when the extended code included in the second code string input from the second communication line does not include an extended code with a frame number identical to that of the mono code included in the first code string input from the first communication line, the extended code with a frame number closest to that of the mono code included in the first code string input from the first communication line and the extended code included in the second code string input from the second communication line is output; and a decoding step, in which, for each frame, based on the monaural code output in the receiving step and the extended code output in the receiving step, C channels of decoded digital sound signals are obtained and output.
[0011] One embodiment of the present invention is a sound signal decoding method, which is performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, and includes: a decoding step, for each frame, when the extended code included in the second code string input from the second communication line includes an extended code with a frame number identical to that of the mono code included in the first code string input from the first communication line, based on the mono code included in the first code string input from the first communication line and the extended code with a frame number identical to the mono code, C (C is an integer greater than 2) channels of decoded digital sound signals are obtained and output; when the extended code included in the second code string input from the second communication line does not include an extended code with a frame number identical to that of the mono code included in the first code string input from the first communication line, based on the mono code included in the first code string input from the first communication line and the extended code with a frame number closest to the mono code included in the second code string input from the second communication line, C channels of decoded digital sound signals are obtained and output.
[0012] One embodiment of the present invention is a method for encoding and transmitting a sound signal, which is performed by a terminal device connected to a first communication line and a second communication line having a lower priority than the first communication line, and includes: an encoding step, in which, for each frame, a mono code representing a signal formed by mixing digital sound signals of C channels (C is an integer greater than 2) to be input, and an extended code representing a characteristic parameter, wherein the characteristic parameter is a parameter representing the difference between the channels of the digital sound signals of the C channels to be input and representing information that depends on the relative position of the sound source and the microphone in space; and a transmitting step, in which, for each frame, a first code string including the mono code obtained in the encoding step is output to the first communication line, and a second code string including the extended code obtained in the encoding step is output to the second communication line.
[0013] One embodiment of the present invention is a method for encoding and transmitting a sound signal, which is performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, and includes: an encoding step, for each frame, obtaining a mono code representing a signal mixed with digital sound signals of C channels (C is an integer greater than 2) to be input, and for a predetermined frame among multiple frames, obtaining an extended code representing a characteristic parameter, the characteristic parameter being a parameter representing a characteristic of the difference between the channels of the digital sound signals of the C channels to be input and representing information that depends on the relative position of the sound source and the microphone in space; and a transmitting step, for each frame, outputting a first code string including the mono code obtained in the encoding step to the first communication line, and for the predetermined frame, outputting a second code string including the extended code obtained in the encoding step to the second communication line.
[0014] One embodiment of the present invention is a method for encoding and transmitting a sound signal, which is performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, and includes: an encoding step, for each frame, obtaining a mono code representing a signal formed by mixing digital sound signals of C channels (C is an integer greater than 2) to be input, obtaining a characteristic parameter for each frame, the characteristic parameter being a parameter representing a characteristic of the difference between the channels of the digital sound signals of the C channels to be input and representing information that depends on the relative position of the sound source and the microphone in space, and obtaining an extended code representing an average or weighted average of the characteristic parameters for a predetermined frame among multiple frames; and a transmitting step, for each frame, outputting a first code string including the mono code obtained in the encoding step to the first communication line, and outputting a second code string including the extended code obtained in the encoding step to the second communication line for the predetermined frame.
[0015] One embodiment of the present invention is a sound signal encoding method, which is performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, and includes: an encoding step, obtaining a mono code and an extended code for each frame and outputting them, the mono code is a signal that represents a mixture of digital sound signals of C channels (C is an integer greater than 2) to be input and is included in a first code string and output to the first communication line, the extended code is a code that represents characteristic parameters and is included in a second code string and output to the second communication line, the characteristic parameters are parameters that represent the characteristics of the difference between the channels of the digital sound signals of the C channels to be input and represent information that depends on the relative position of the sound source and the microphone in space.
[0016] One embodiment of the present invention is a sound signal encoding method, which is performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, and includes: an encoding step, obtaining and outputting a mono code for each frame, the mono code being a signal representing a mixture of digital sound signals of C channels (C is an integer greater than 2) to be input and a code output to the first communication line in a first code string; obtaining and outputting an extended code for a predetermined frame among multiple frames, the extended code being a code output to the second communication line in a second code string representing a characteristic parameter, the characteristic parameter being a parameter representing a characteristic of the difference between the channels of the digital sound signals of the C channels to be input and representing information dependent on the relative position of the sound source and the microphone in space.
[0017] One embodiment of the present invention is a sound signal encoding method, which is performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, and includes: an encoding step, obtaining and outputting a mono code for each frame, the mono code being a signal representing a mixture of digital sound signals of C channels (C is an integer greater than 2) to be input and a code output to the first communication line in a first code string; obtaining a characteristic parameter for each frame, the characteristic parameter being a parameter representing a characteristic of the difference between the channels of the digital sound signals of the C channels to be input and representing information dependent on the relative position of the sound source and the microphone in space; obtaining and outputting an extended code for a predetermined frame among multiple frames, the extended code being a code representing an average or weighted average of the characteristic parameters and output to the second communication line in a second code string.
[0018] Effects of the Invention
[0019] According to the present invention, a decoded audio signal of high sound quality can be obtained without significantly increasing the delay time compared to a configuration in which a decoded audio signal of only the minimum required sound quality is obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a block diagram showing an example of a telephone system.
[0021] Figure 2 This is a block diagram showing an example of a multi-line supporting terminal device.
[0022] Figure 3 This is a flowchart showing an example of processing by a voice signal transmitting-side device of a multi-line supporting terminal device.
[0023] Figure 4 This is a flowchart showing an example of processing by a voice signal receiving device of a multi-line supporting terminal device.
[0024] Figure 5 This is a diagram schematically showing the temporal relationship between an input code and an output signal in a voice signal receiving-side device of a multi-channel supporting terminal device.
[0025] Figure 6 This is a diagram schematically showing the temporal relationship between an input code and an output signal in a conventional audio signal receiving device.
[0026] Figure 7 This is a block diagram showing an example of a multi-site control device.
[0027] Figure 8 This is a flowchart showing a processing example of the multi-site control device.
[0028] Figure 9This is a block diagram showing an example of a multi-site control device.
[0029] Figure 10 This is a flowchart showing a processing example of the multi-site control device.
[0030] Figure 11 This is a block diagram showing an example of a dedicated telephone line terminal device.
[0031] Figure 12 This is a flowchart showing an example of processing by the voice signal transmitting-side device of the telephone line dedicated terminal device.
[0032] Figure 13 This is a flowchart showing an example of processing by a voice signal receiving-side device of a telephone line dedicated terminal.
[0033] Figure 14 This is a diagram showing an example of a functional configuration of a computer for realizing each device in the embodiment of the present invention. DETAILED DESCRIPTION
[0034] <Telephone System 100>
[0035] Telephone system 100 Figure 1 As shown, the telephone system 100 includes a multi-line supporting terminal device 200-m (m is an integer greater than or equal to 1 and less than or equal to M, and M is an integer greater than or equal to 2), a first communication network 400, and a second communication network 500. Figure 1 The dashed line in the middle shows dedicated telephone line terminal devices 300-n (n is an integer greater than or equal to 1 and less than or equal to N, where N is an integer greater than or equal to 1). Each multi-line supporting terminal device 200-m can connect to another terminal device via a first communication line 410-m, which is each communication line of the first communication network 400. Furthermore, each multi-line supporting terminal device 200-m can connect to another multi-line supporting terminal device via a second communication line 510-m, which is each communication line of the second communication network 500. Each dedicated telephone line terminal device 300-n can connect to another terminal device via a first communication line 420-n, which is each communication line of the first communication network 400.
[0036] <First Communication Network 400, Second Communication Network 500>
[0037] The first communication network 400 and the second communication network 500 have different information transmission priorities. The first communication network 400 has a higher information transmission priority than the second communication network 500 and is capable of transmitting a code string at a specific bit rate from one terminal device to another with a short delay. The first communication network 400 is, for example, a communication network used for two-way communication between a terminal device such as a conventional mobile phone or smartphone and another terminal device such as a conventional mobile phone or smartphone, and is a communication network having communication lines generally referred to as telephone lines. The second communication network 500 has a lower information transmission priority than the first communication network 400 and is capable of transmitting a code string from one terminal device to another without imposing delay constraints. The second communication network 500 is, for example, a communication network used when transmitting data such as an image or a character string from a terminal device such as a smartphone to another terminal device such as a smartphone, and is a communication network having communication lines generally referred to as Internet lines.
[0038] exist Figure 1 Although described as a first communication network 400 and a second communication network 500, the first communication network 400 and the second communication network 500 do not need to be physically separated; they only need to be logically separated. Similarly, when a terminal device is connected to both the first communication line 410-m and the second communication line 510-m, the first communication line 410-m and the second communication line 510-m do not need to be physically separated; they only need to be logically separated. In other words, each terminal device can be connected to a single IP communication network via a single IP communication line. Through packet priority control, etc., a logical structure is established: a first communication network 400 and a first communication line 410-m, which are communication networks and communication lines with a higher priority for information transmission; and a second communication network 500 and a second communication line 510-m, which are communication networks and communication lines with a lower priority for information transmission than the first communication network 400 and the first communication line 410-m. For example, the multi-line supporting terminal device 200-m may be a smartphone supporting VoLTE (Voice over LTE), examples of the first communication network 400 and the first communication line 410-m may be an LTE communication network and a VoLTE communication network and a VoLTE line in the LTE line, and examples of the second communication network 500 and the second communication line 510-m may be an LTE communication network and an Internet communication network and an Internet line in the LTE line.
[0039] In addition, the above-mentioned examples of communication networks, communication lines, and terminal devices are all examples of mobile communications, but there is no limitation on whether each communication network is a communication network for fixed communication or a communication network for mobile communication, whether each communication line is wired or wireless, and whether each terminal device is a fixed telephone or a portable telephone.
[0040] <First embodiment>
[0041] The multi-line supporting terminal device according to the first embodiment will be described.
[0042] <Multi-line Support Terminal Device 200-m>
[0043] The multi-line supporting terminal device 200-m is, for example, a smartphone supporting VoLTE, such as Figure 2 As shown, the system includes a sound signal transmitting device 210-m and a sound signal receiving device 220-m. The sound signal transmitting device 210-m includes a sound receiving unit 211-m, an encoding device 212-m, and a transmitting unit 213-m. The sound signal receiving device 220-m includes a receiving unit 221-m, a decoding device 222-m, and a playback unit 223-m. The encoding device 212-m includes a signal analysis unit 2121-m and a mono encoding unit 2122-m. The decoding device 222-m includes a mono decoding unit 2221-m and an extended decoding unit 2222-m. Furthermore, as indicated by the dotted lines in the figure, the signal analysis unit 2121-m and the mono encoding unit 2122-m are collectively referred to as the encoding unit 2129-m, and the mono decoding unit 2221-m and the extended decoding unit 2222-m are collectively referred to as the decoding unit 2229-m. In addition, the encoding device 212-m and the decoding device 222-m are sometimes referred to as the sound signal encoding device 212-m and the sound signal decoding device 222-m, respectively. Figure 3 The multi-line support terminal device 200-m performs the following processing of steps S211 to S213, and the multi-line support terminal device 200-m performs the processing of the voice signal receiving side device 220-m. Figure 4 And the processing of steps S221 to S223 illustrated below.
[0044] [Sound signal transmitting device 210-m]
[0045] The sound signal sending side device 210-m obtains a code string containing a monaural code corresponding to the digital sound signal of the two channels, i.e., a first code string, and outputs it to the first communication line 410-m, for example, at each specific time interval of 20 ms, i.e., each frame, and obtains a code string containing an extended code corresponding to the digital sound signal of the two channels, i.e., a second code string, and outputs it to the second communication line 510-m.
[0046] [[Sound receiving unit 211-m]]
[0047] The sound pickup unit 211-m includes two microphones and two AD conversion units. Each microphone is associated with each AD conversion unit on a one-to-one basis. The microphone picks up the sound generated in the spatial domain around the microphone, converts it into an analog electrical signal, and outputs it to the AD conversion unit. The AD conversion unit converts the input analog electrical signal into a digital sound signal, such as a PCM signal with a sampling frequency of 8kHz, and outputs it. That is, the sound pickup unit 211-m outputs two-channel digital sound signals corresponding to the sounds picked up by the two microphones, such as a two-channel stereo digital sound signal of the left and right channels, to the encoding device 212-m (step S211).
[0048] Furthermore, all or part of the sound pickup unit 211-m may not be located within the sound signal transmitting device 210-m, but may be connected to the sound signal transmitting device 210-m. For example, the sound pickup unit 211-m of the sound signal transmitting device 210-m may not include a microphone, but may instead input two analog electrical signals from a microphone connected to the sound signal transmitting device 210-m to the AD converter of the sound pickup unit 211-m of the sound signal transmitting device 210-m. Alternatively, the sound signal transmitting device 210-m may not include the sound pickup unit 211-m, but may instead input two channels of digital sound signals from a radio device such as an AD converter connected to the sound signal transmitting device 210-m to the encoding device 212-m of the sound signal transmitting device 210-m.
[0049] [[Encoding device 212-m]]
[0050] Two-channel digital audio signals are input to encoding device 212-m from sound receiving unit 211-m or a radio connected to audio signal transmitting device 210-m. Encoding device 212-m generates a monaural code and an extension code corresponding to the input two-channel digital audio signals for each frame, and outputs them to transmitting unit 213-m (step S212).
[0051] [[[Signal Analysis Unit 2121-m]]]
[0052] Signal analysis unit 2121-m generates, for each frame, a monaural signal (a signal obtained by mixing the two-channel digital audio signals) and a spreading code representing characteristic parameters that exhibit minimal temporal variation and represent the differential characteristics of the two-channel digital audio signals. Signal analysis unit 2121-m outputs the resulting monaural signal to monaural encoding unit 2122-m and the resulting spreading code to transmission unit 213-m. Parameters exhibiting minimal temporal variation have low temporal dependence and low temporal resolution.
[0053] [First Example of Signal Analysis Unit 2121-m]
[0054] As a first example, the operation of the signal analysis unit 2121-m for each frame is described when information representing the time difference between the two input channels of digital audio signals is used as a feature parameter. The signal analysis unit 2121-m first obtains a feature parameter representing the time difference between the two input channels of digital audio signals (step S2121-11). The time difference between the two input channels of digital audio signals can be obtained using any well-known method. For example, the signal analysis unit 2121-m calculates the correlation value between the sample sequence of the digital audio signal of one channel (first channel) and the sample sequence obtained by advancing the sample sequence of the digital audio signal of the other channel (second channel) by the candidate sample number for each candidate sample number of time differences within a predetermined range, and obtains the time difference sample number of the candidate sample number with the largest correlation value as the feature parameter.
[0055] The signal analysis unit 2121-m then obtains one of the following sequences: a sequence of summed samples, a sequence of average values of the corresponding samples, or a sequence obtained by modifying the above-mentioned summed or averaged sequences, for the sample sequence of the digital audio signal of the first channel and the sample sequence of the digital audio signal of the second channel, as a signal obtained by mixing the digital audio signals of the two channels, namely, a monaural signal (step S2121-12). The sample sequence obtained by assigning the time difference indicated by the characteristic parameter to the sample sequence of the digital audio signal of the second channel is, for example, a sample sequence obtained by advancing the sample sequence of the digital audio signal of the second channel by the number of samples indicated by the time difference indicated by the characteristic parameter.
[0056] Signal analysis unit 2121-m then obtains a spreading code representing the characteristic parameter (step S2121-13). The spreading code representing the characteristic parameter can be obtained using a well-known method. For example, signal analysis unit 2121-m performs scalar quantization on the number of time difference samples of the input two-channel digital audio signal to obtain a code, and outputs the obtained code as the spreading code. Alternatively, signal analysis unit 2121-m may output the binary number representing the number of time difference samples of the input two-channel digital audio signal as the spreading code.
[0057] [Second Example of Signal Analysis Unit 2121-m]
[0058] As a second example, the operation of signal analysis section 2121-m for each frame is described, using information representing the intensity difference in each frequency band of the input two-channel digital audio signal as a feature parameter. While a specific example using the complex DFT (Discrete Fourier Transformation) is described below, well-known frequency domain transformation methods other than the complex DFT may also be used.
[0059] Signal analysis unit 2121-m first performs a complex DFT on each of the two input channels of digital audio signals to obtain a complex DFT coefficient sequence (step S2121-21). The complex DFT coefficient sequence can also be obtained using well-known methods, such as applying a window that overlaps between frames or taking into account the symmetry of the complex numbers obtained by the complex DFT. For example, if a frame consists of 128 samples, a complex DFT is performed on a continuous 256-point digital audio signal sample sequence, including the last 64 samples of the immediately preceding frame and the first 64 samples of the immediately following frame. The first half of the 128 complex number sequences of the resulting 256 complex number sequences are then used as the complex DFT coefficient sequence. Subsequently, f is set to integers between 1 and 128, the complex DFT coefficients of the complex DFT coefficient sequence for the first channel are set to V1(f), and the complex DFT coefficients of the complex DFT coefficient sequence for the second channel are set to V2(f). The signal analysis unit 2121-m then obtains a sequence of radius values for each complex DFT coefficient on the complex plane based on the complex DFT coefficient sequences for the two channels (step S2121-22). The radius values for each complex DFT coefficient on the complex plane for each channel correspond to the intensity of each frequency bin of the digital audio signal for each channel. Subsequently, the radius value for the complex DFT coefficient V1(f) of the first channel on the complex plane is set to V1r(f), and the radius value for the complex DFT coefficient V2(f) of the second channel on the complex plane is set to V2r(f). The signal analysis unit 2121-m then obtains the average value of the ratio of the radius value of one channel to the radius value of the other channel for each frequency band, and obtains the sequence of average values as a feature parameter (step S2121-23). This sequence of average values is a feature parameter that represents information about the intensity difference for each frequency band of the input two-channel digital audio signal. For example, if there are four bands, for the four bands f from 1 to 32, from 33 to 64, from 65 to 96, and from 97 to 128, the value of the radius of the first channel V1r(f) is divided by the value of the radius of the second channel V2r(f) to obtain the average values of 32 values Mr(1), Mr(2), Mr(3), and Mr(4), and the sequence of average values {Mr(1), Mr(2), Mr(3), Mr(4)} is obtained as the feature parameter.
[0060] Furthermore, the number of bands can be any value less than the number of frequency bins, and can be the same as the number of frequency bins or 1. When the number of bands is the same as the number of frequency bins, the signal analysis unit 2121-m obtains the ratio of the radius of the vocal tract of one frequency bin to the radius of the other frequency bin, and uses the sequence of these ratios as a characteristic parameter. When the number of bands is 1, the signal analysis unit 2121-m obtains the ratio of the radius of the vocal tract of one frequency bin to the radius of the other frequency bin, and uses the average value of these ratios across the entire band as a characteristic parameter. Furthermore, when the number of bands is set to multiple, the number of frequency bins included in each frequency band is arbitrary; for example, a low-frequency band may include fewer frequency bins than a high-frequency band.
[0061] Furthermore, instead of using the ratio of the radius of one vocal channel to the radius of the other vocal channel, the signal analysis unit 2121-m may use the difference between the radius of one vocal channel and the radius of the other vocal channel. That is, in the above example, instead of using the value obtained by dividing the radius of the first vocal channel V1r(f) by the radius of the second vocal channel V2r(f), the value obtained by subtracting the radius of the second vocal channel V2r(f) from the radius of the first vocal channel V1r(f) may be used.
[0062] The signal analysis unit 2121-m also obtains, for the sample sequence of the digital audio signal of the first channel and the sample sequence of the digital audio signal of the second channel, one of a sequence of the sum of the corresponding samples, a sequence of the average values of the corresponding samples, or a sequence obtained by modifying the above-mentioned summed or averaged sequences, as a signal obtained by mixing the digital audio signals of the two channels, namely, a monophonic signal (step S2121-24). Furthermore, the signal analysis unit 2121-m may obtain the average value VMr(f) of the radius and the average value VMθ(f) of the complex DFT coefficients V1(f) of the first channel and V2(f) of the second channel obtained in step S2121-21, and perform an inverse complex DFT on the sequence of complex numbers VM(f) with a radius VMr(f) and an angle VMθ(f) on the complex plane to obtain a signal obtained by mixing the digital audio signals of the two channels, namely, a monophonic signal (step S2121-24').
[0063] The signal analysis unit 2121-m then obtains a spreading code representing the characteristic parameter (step S2121-25). The spreading code representing the characteristic parameter can be obtained using a well-known method. For example, the signal analysis unit 2121-m performs vector quantization on the sequence of values obtained in step S2121-23 to obtain a code, and outputs the obtained code as the spreading code. Alternatively, for example, the signal analysis unit 2121-m performs scalar quantization on each of the values included in the sequence of values obtained in step S2121-23 to obtain a code, and combines the obtained codes to output as the spreading code. Furthermore, if the signal analysis unit 2121-m obtains a single value in step S2121-23, the code obtained by scalar quantization of the single value can be output as the spreading code.
[0064] The time difference between the two-channel digital audio signals input as described in the first example of signal analysis unit 2121-m, or the intensity difference between each frequency band of the two-channel digital audio signals input as described in the second example of signal analysis unit 2121-m, depends on the location of the sound source. For typical sound sources such as people or musical instruments, the location of the sound source rarely changes over time. Even if the location of the sound source changes over time, the time difference between the two-channel digital audio signals input as well as the intensity difference between each frequency band will not change much unless the sound source moves abruptly.
[0065] Therefore, signal analysis unit 2121-m may also calculate an average or weighted average of characteristic parameters obtained from the input two-channel digital audio signals for a plurality of consecutive frames including the target frame, and output a spreading code representing the obtained characteristic parameters. The weight used for weighted averaging may be set to the maximum value for the target frame and to a smaller value for frames further away from the target frame. Furthermore, if characteristic parameters for frames further in the future than the target frame are used, look-ahead is required, increasing the delay. Therefore, signal analysis unit 2121-m preferably uses a plurality of consecutive frames in the past, including the target frame. Furthermore, when a characteristic parameter includes multiple elements, such as information representing the intensity differences of multiple frequency bands, the average or weighted average of the characteristic parameters is a numerical sequence whose elements are the average or weighted average of each element of the characteristic parameter.
[0066] Furthermore, for example, the waveform difference of the two-channel digital audio signal input, that is, the sample sequence representing the difference between corresponding samples of the two-channel digital audio signal input, will be completely different from the waveform difference of the two-channel digital audio signal input even if the timing of each sample is shifted by only one sample. Therefore, this information is highly time-dependent, has high temporal resolution, and exhibits significant temporal fluctuations. Similarly, the phase difference between the two-channel digital audio signal input, for example, the difference in angle on the complex plane between each complex DFT coefficient V1(f) of the complex DFT coefficient sequence for the first channel obtained in step S2121-21 and each complex DFT coefficient V2(f) of the complex DFT coefficient sequence for the second channel, is highly time-dependent, has high temporal resolution, and exhibits significant temporal fluctuations.
[0067] That is, the characteristic parameters represented by the extended code obtained by the signal analysis unit 2121-m are not parameters that express information that depends on the waveform of the sound signal emitted by the sound source in the difference between the two channels of digital sound signals input, such as the difference in waveform of the two channels of digital sound signals input or the phase difference between the two channels of digital sound signals input as exemplified immediately above, but are parameters that express information that depends on the spatial relative position of the sound source and the microphone in the difference between the two channels of digital sound signals input, such as the time difference between the two channels of digital sound signals input as shown in the first example of the signal analysis unit 2121-m or the intensity difference in each frequency band of the two channels of digital sound signals input as shown in the second example of the signal analysis unit 2121-m. In short, the characteristic parameter represented by the extended code obtained by the signal analysis unit 2121-m can be said to be a parameter that expresses the difference characteristics of the two channels of digital sound signals input and has low time resolution, or a parameter that expresses the difference characteristics of the two channels of digital sound signals input and has small temporal variation, or a parameter that expresses the difference characteristics of the two channels of digital sound signals input and has low dependence on time, or a parameter that expresses the difference characteristics between the channels of the two channels of digital sound signals input and depends on information about the relative spatial position of the sound source and the microphone.
[0068] [[[Mono encoding unit 2122-m]]]
[0069] Mono encoding section 2122-m encodes the input mono signal frame by frame using a specific encoding method, generates a mono code, and outputs it to transmitting section 213-m. The encoding method should be one that provides a mono code bit rate that is less than the communication capacity of first communication line 410-m. For example, a method for encoding voice in the telephone band used in mobile phones, such as the 13.2 kbps mode of the 3GPP EVS standard (3GPP TS26.442), can be used.
[0070] Specifically, encoding device 212-m generates, for each frame, a monaural code representing a signal obtained by mixing the two-channel digital audio signals being input, and an extended code representing a characteristic parameter having low temporal resolution and representing the inter-channel difference characteristics of the two-channel digital audio signals being input. As will be described later, the monaural code generated by encoding device 212-m is the code included in the first code string and output to the first communication channel, while the extended code generated by encoding device 212-m is the code included in the second code string and output to the second communication channel.
[0071] In addition, the encoding device 212-m can also use, as an extension code, a code that represents the average or weighted average of characteristic parameters obtained based on the digital sound signals of the two channels of the frame being processed, i.e., the current frame, and characteristic parameters obtained based on the digital sound signals of the two channels of a frame that is closer to the current frame being processed.
[0072] [[Sending unit 213-m]]
[0073] The sending unit 213-m outputs the code string containing the mono code input from the encoding device 221-m, i.e., the first code string, to the first communication line 410-m per frame, and outputs the code string containing the extended code input from the encoding device 221-m, i.e., the second code string, to the second communication line 510-m (step S213).
[0074] Transmitting unit 213-m outputs a mono code in a manner that allows identification of the frame in the first code string. For example, transmitting unit 213-m includes information that can identify the frame, such as the frame number or the time corresponding to the frame, as auxiliary information in the first code string and outputs it. Similarly, transmitting unit 213-m outputs a spread code in a manner that allows identification of the frame in the second code string. For example, transmitting unit 213-m includes information that can identify the frame, such as the frame number or the time corresponding to the frame, as auxiliary information in the second code string and outputs it. Furthermore, in the sound signal receiving device 220-m of this first embodiment and subsequent embodiments and variations, an example is described in which the frame number is included as auxiliary information in both the first and second code strings.
[0075] [Sound Signal Receiving Device 220-m]
[0076] The sound signal receiving side device 220-m outputs sound based on the monaural code contained in the first code string input from the first communication line 410-m and the extended code contained in the second code string input from the second communication line 510-m, for example, at a specific time interval of each 20ms, that is, each frame.
[0077] [[Receiving unit 221-m]]
[0078] The receiving unit 221-m outputs the mono code contained in the first code string input from the first communication line 410-m and the extended code contained in the second code string input from the second communication line 510-m, whose frame number is closest to the mono code, to the decoding device 222-m in each frame (step S221).
[0079] First communication line 410-m is a high-priority communication network for two-way communication. Therefore, a first code string containing a monaural code is input from first communication line 410-m to receiving unit 221-m, so that the monaural code outputted by encoding unit 212-m' of voice signal transmitting device 210-m' of multi-line supporting terminal device 200-m' (m' is an integer different from m and greater than 1 and less than M) on the other end of the call can be outputted in frame number order at intervals of the frame length (i.e., at specific intervals of 20 ms, for example). Furthermore, since telephone system 100 aims to smoothly facilitate two-way communication, receiving unit 221-m preferably outputs the code outputted by encoding unit 212-m' of voice signal transmitting device 210-m' on the other end of the call to decoding unit 222-m with as little delay as possible. Therefore, the receiving unit 221-m outputs the mono code contained in the first code string output by the sound signal sending side device 210-m' of the other party of the call to the decoding device 222-m in the order of the frame numbers output by the sound signal sending side device 210-m' of the other party of the call, at a time interval of the frame length, regardless of whether the second code string containing the extended code with the same frame number as each mono code is input to the receiving unit 221-m.
[0080] Second communication line 510-m is a low-priority communication network. Therefore, the second code string of a frame output by the other party's voice signal transmitting device 210-m' is typically input to receiving unit 221-m via second communication line 510-m after the first code string of that frame is input to receiving unit 221-m via first communication line 410-m. In other words, at the time receiving unit 221-m outputs a monaural code to decoding unit 222-m, the second code string containing the same extension code with the same frame number as the monaural code has typically not been input to receiving unit 221-m, preventing the extension code with the same frame number from being output to decoding unit 222-m. Furthermore, because second communication line 510-m is a low-priority communication network, the second code strings of each frame output by the other party's voice signal transmitting device 210-m' are not necessarily input from second communication line 510-m in frame number order. Of course, depending on the status of the second communication network 500, for example, if the second communication network 500 is idle, the second code string of a certain frame output by the voice signal transmitting device 210-m' of the other party may be input to the receiving unit 221-m from the second communication line 510-m at the same time as, or before, the first code string of the frame is input to the receiving unit 221-m from the first communication line 410-m. In other words, there is a possibility that, at the time when the receiving unit 221-m outputs the monaural code to the decoding device 222-m, the second code string including the extension code with the same frame number as the monaural code has already been input to the receiving unit 221-m, and the extension code with the same frame number as the monaural code may be output to the decoding device 222-m. Therefore, receiving unit 221-m outputs, for each frame, to decoding device 222-m, the spreading code included in the second code string input from second communication line 510-m, whose frame number is closest to the monaural code output to decoding device 222-m, instead of the spreading code included in the second code string input from second communication line 510-m, whose frame number is the same as the monaural code output to decoding device 222-m. In other words, receiving unit 221-m outputs, for each frame, the spreading code included in the second code string input from second communication line 510-m, whose frame number is closest to the first code string included in the monaural code output to decoding device 222-m.
[0081] Here, as the spreading code included in the second code string input from the second communication line 510-m, the spreading code having a frame number closest to the monaural code output to the decoding device 222-m is the spreading code included in the second code string input from the second communication line 510-m, and when the spreading code included in the second code string input from the second communication line 510-m includes a spreading code having a frame number identical to the monaural code output to the decoding device 222-m, the spreading code included in the second code string input from the second communication line 510-m, and the monaural code having a frame number identical to the monaural code output to the decoding device 222-m is the spreading code included in the second code string input from the second communication line 510-m. When the second code string input from the second communication channel 510-m does not include an extension code with the same frame number as the monaural code output to the decoding device 222-m, the extension code is the extension code with the frame number closest to the monaural code output to the decoding device 222-m (that is, the extension code included in the second code string input from the second communication channel 510-m, which has a frame number that is different from the monaural code output to the decoding device 222-m but is closest to the monaural code output to the decoding device 222-m). This also applies to the embodiments and modifications described below.
[0082] Specifically, receiving unit 221-m outputs, for each frame, the monaural code included in the first code string input from first communication line 410-m and the extended code included in the second code string input from second communication line 510-m, whose frame number is closest to the monaural code. Receiving unit 221-m outputs the monaural codes in frame number order. More specifically, the receiving unit 221-m accepts input of a first code string from the first communication line 410-m and input of a second code string from the second communication line 510-m, and outputs the mono code contained in the first code string input from the first communication line 410-m (i.e., the mono code in frame number order) for each frame. If the extended code contained in the second code string input from the second communication line 510-m includes an extended code with the same frame number as the mono code, the receiving unit 221-m outputs the extended code with the same frame number as the mono code. If the extended code contained in the second code string input from the second communication line 510-m does not include an extended code with the same frame number as the mono code, the receiving unit 221-m outputs the extended code with the frame number closest to the mono code among the extended codes contained in the second code string input from the second communication line (i.e., the extended code with the frame number closest to the mono code, although different from the frame number of the mono code, among the extended codes contained in the second code string input from the second communication line).
[0083] Although not described in detail due to well-known technology, receiving unit 221-m includes a storage unit (not shown) that accumulates multiple frames of code strings received asynchronously from various communication lines due to communication involving fluctuations or retransmission control. While code strings are not necessarily input to receiving unit 221-m from various communication lines at specific time intervals or in frame number order, receiving unit 221-m can output any code included in the code strings accumulated in the storage unit. Specifically, receiving unit 221-m receives and stores a first code string from first communication line 410-m, stores the inputted first code string, and can output any code string that is part of the stored first code string. Furthermore, receiving unit 221-m receives and stores a second code string from second communication line 510-m, stores the inputted second code string, and can output any code string that is part of the stored second code string. Therefore, the receiving unit 221 - m can extract the mono code in the order of frame numbers in each specific time interval, that is, each frame, or extract the spreading code with the frame number closest to the mono code.
[0084] [[Decoding device 222-m]]
[0085] The monaural code and extended code outputted by the receiving unit 221-m are inputted to the decoding unit 222-m on a frame-by-frame basis. The decoding unit 222-m obtains a two-channel decoded digital audio signal corresponding to the input monaural code and extended code on a frame-by-frame basis and outputs it to the playback unit 223-m (step S222).
[0086] Input to decoding device 222-m are the monaural codes in frame number order contained in the first code string input from first communication line 410-m in frame number order, and the spreading codes with frame numbers closest to the monaural codes contained in the second code string input from second communication line 510-m. In other words, decoding device 222-m obtains and outputs a two-channel decoded digital audio signal for each frame based on the monaural code contained in the first code string input from first communication line 410-m and the spreading code with frame numbers closest to the monaural code contained in the second code string input from second communication line 510-m. The monaural codes used by decoding device 222-m are naturally in frame number order.
[0087] In other words, the input to decoding device 222-m is the monaural code in the order of frame numbers output by encoding device 212-m' of the other party's audio signal transmitting device 210-m', along with the extension code with the frame number closest to the monaural code. Specifically, decoding device 222-m uses the monaural code in the order of frame numbers output by encoding device 212-m' of the other party's audio signal transmitting device 210-m', along with the extension code with the frame number closest to the monaural code, for each frame, and outputs it to playback unit 223-m.
[0088] Here, the spreading code input to decoding device 222-m is, in the case where the spreading code included in the second code string input from second communication channel 510-m includes a frame number identical to the monaural code included in the first code string input from first communication channel 410-m, the spreading code included in the second code string input from second communication channel 510-m is the spreading code included in the second code string input from second communication channel 510-m. In the case where the spreading code included in the second code string input from second communication channel 510-m does not include a frame number identical to the monaural code included in the first code string input from first communication channel 410-m, the spreading code included in the second code string input from second communication channel 510-m is the spreading code with the frame number closest to the monaural code of the frame (i.e., the spreading code with the frame number different from the monaural code of the frame but closest to the monaural code of the frame). This also applies to the embodiments and variations described below.
[0089] Thus, the decoding device 222-m obtains and outputs a decoded digital sound signal of two channels based on the mono code (i.e., the mono code in the frame number sequence) contained in the first code string input from the first communication line 410-m and the extension code with the same frame number as the mono code, in each frame, when the extension code contained in the second code string input from the second communication line 510-m contains an extension code having the same frame number as the mono code (i.e., the mono code in the frame number sequence) contained in the first code string input from the first communication line 410-m, and outputs the decoded digital sound signal. In a case where the included extended codes do not include an extended code having the same frame number as the mono code included in the first code string input from the first communication line 410-m (i.e., the mono code in frame number sequence), based on the mono code included in the first code string input from the first communication line 410-m (i.e., the mono code in frame number sequence) and the extended code having a frame number closest to the mono code included in the second code string input from the second communication line 510-m (i.e., the extended code having a frame number that is different from the mono code but is closest to the mono code), decoded digital sound signals of two channels are obtained and output.
[0090] [[[Mono decoding unit 2221-m]]]
[0091] The monaural code input to the decoding device 222-m is input to the monaural decoding unit 2221-m on a frame-by-frame basis. The monaural decoding unit 2221-m decodes the input monaural code using a specific decoding method for each frame, obtaining a decoded monaural digital audio signal and outputting it to the extended decoding unit 2222-m. The specific decoding method used is the same as the one used by the monaural encoding unit 2122-m' of the encoding device 212-m' of the voice signal transmitting device 210-m' of the other party in the call.
[0092] The input to monaural decoding section 2221-m is the monaural code in the frame number order output by encoding section 212-m' of audio signal transmitting device 210-m' of the other party. Specifically, monaural decoding section 2221-m obtains, for each frame, the monaural decoded digital audio signal encoded by encoding section 212-m' of audio signal transmitting device 210-m' of the other party and outputs it to extended decoding section 2222-m.
[0093] [[[Extended decoding unit 2222-m]]]
[0094] The mono decoded digital audio signal output by the monaural decoding unit 2221-m and the extension code input to the decoding device 222-m are input to the extension decoding unit 2222-m on a frame-by-frame basis. The extension decoding unit 2222-m obtains a two-channel decoded digital audio signal based on the input mono decoded digital audio signal and the extension code on a frame-by-frame basis, and outputs the signal to the playback unit 223-m.
[0095] The mono decoded digital audio signal input to extension decoding unit 2222-m follows the frame number sequence encoded by encoding unit 212-m' of the other party's audio signal transmitting device 210-m'. The extension code input to decoding unit 222-m is the extension code with the frame number closest to the mono decoded digital audio signal. In other words, extension decoding unit 2222-m generates two-channel decoded digital audio signals for each frame based on the mono decoded digital audio signal in the frame number sequence output by encoding unit 212-m' of the other party's audio signal transmitting device 210-m' and the extension code with the frame number closest to the mono decoded digital audio signal. These signals are then output to playback unit 223-m. Furthermore, the extension code represents a characteristic parameter obtained by encoding unit 212-m' of the other party's audio signal transmitting device 210-m' of the multi-channel terminal 200-m', representing a parameter representing the differential characteristics of the two-channel digital audio signals. That is, the extended decoding unit 2222-m regards the input mono decoded digital sound signal as a signal obtained by mixing the decoded digital sound signals of the two channels for each frame, regards the characteristic parameters obtained based on the extended code as information representing the differential characteristics of the digital sound signals of the two channels, obtains the decoded digital sound signals of the two channels and outputs them to the playback unit 223-m.
[0096] [First Example of Extension Decoding Unit 2222-m]
[0097] As a first example, the operation of extension decoding unit 2222-m per frame will be described, where the characteristic parameter is information representing the time difference between two channels of digital audio signals. Extension decoding unit 2222-m first obtains information representing the time difference, which serves as the characteristic parameter represented by the input extension code, based on the input extension code (step S2222-11). Extension decoding unit 2222-m obtains the characteristic parameter based on the extension code in a manner similar to how signal analysis unit 2121-m' of encoding unit 212-m' of the other party's audio signal transmitting device 210-m' obtains the extension code based on the characteristic parameter. The information representing the time difference, which serves as the characteristic parameter, may be, for example, the number of time difference samples. For example, extension decoding unit 2222-m performs scalar decoding on the input extension code, obtaining a scalar value corresponding to the input extension code as the number of time difference samples. Alternatively, extension decoding unit 2222-m may interpret the input extension code as a binary number and obtain a decimal number corresponding to the binary number as the number of time difference samples.
[0098] The extended decoding unit 2222-m then treats the input mono decoded digital audio signal and the characteristic parameters obtained in step S2222-11 as a mixture of two decoded digital audio signals, and uses the characteristic parameters as information representing the time difference between the two decoded digital audio signals. It then obtains and outputs two decoded digital audio signals (step S2222-12). More specifically, the extended decoding unit 2222-m obtains and outputs one of the following: the sample sequence of the input mono digital audio signal itself, a sequence of values obtained by dividing the values of each sample in the input mono digital audio signal sample sequence by 2, or a sequence obtained by transforming one of the above sample sequences, and outputs the result as the digital audio signal of the first channel (step S2222-121). The extended decoding unit 2222-m further obtains and outputs the sample sequence of the first channel digital audio signal delayed by the number of samples represented by the time difference characteristic parameters as the sample sequence of the second channel digital audio signal (step S2222-122).
[0099] [Second Example of Extension Decoding Unit 2222-m]
[0100] As a second example, the operation of extension decoding unit 2222-m per frame will be described, where the characteristic parameter is information representing the intensity difference for each frequency band of a two-channel digital audio signal. Extension decoding unit 2222-m first decodes the input extension code to obtain information representing the intensity difference for each frequency band (step S2222-21). Extension decoding unit 2222-m obtains the characteristic parameter from the extension code in a manner similar to how signal analysis unit 2121-m' of encoding unit 212-m' of the other party's audio signal transmitting device 210-m' obtains the extension code from information representing the intensity difference for each frequency band. For example, extension decoding unit 2222-m performs vector decoding on the input extension code to obtain the values of each element of the vector corresponding to the input extension code as information representing the intensity difference for each of the multiple frequency bands. Alternatively, extension decoding unit 2222-m performs scalar decoding on each code contained in the input extension code to obtain information representing the intensity difference for each frequency band. When the number of bands is 1, the extension decoding unit 2222-m performs scalar decoding on the input extension code to obtain information indicating the intensity difference of one frequency band, that is, the entire band.
[0101] The extended decoding unit 2222-m then treats the input mono decoded digital audio signal and the characteristic parameters obtained in step S2222-21 as a signal obtained by mixing two decoded digital audio signals, and uses the characteristic parameters as information representing the difference in intensity per frequency band between the two decoded digital audio signals, thereby obtaining and outputting two decoded digital audio signals (step S2222-22). If the signal analysis unit 2121-m' of the encoding device 212-m' of the voice signal transmitting device 210-m' of the other party in the call performs the above-described specific example operation using the complex DFT, the extended decoding unit 2222-m performs the following operation.
[0102] Extension decoding unit 2222-m first performs a complex DFT on the input monaural decoded digital audio signal to obtain a complex DFT coefficient sequence (step S2222-221). Each complex DFT coefficient in the monaural complex DFT coefficient sequence obtained by extension decoding unit 2222-m is then set to MQ(f). Extension decoding unit 2222-m then obtains, from the monaural complex DFT coefficient sequence, the radius MQr(f) of each complex DFT coefficient on the complex plane and the angle MQθ(f) of each complex DFT coefficient on the complex plane (step S2222-222). The extended decoding unit 2222-m then obtains the value obtained by multiplying the value MQr(f) of each radius by the square root of the corresponding value in the characteristic parameter as the value VLQr(f) of each radius for the first channel, and obtains the value obtained by dividing the value MQr(f) of each radius by the square root of the corresponding value in the characteristic parameter as the value VRQr(f) of each radius for the second channel (steps S2222-223). In the example of the four bands described above, the corresponding value in the characteristic parameter for each frequency bin is Mr(1) for f from 1 to 32, Mr(2) for f from 33 to 64, Mr(3) for f from 65 to 96, and Mr(4) for f from 97 to 128. In addition, when the signal analysis unit 2121-m' of the encoding device 212-m' of the sound signal sending side device 210-m' of the other party in the call uses the difference between the radius value of the first channel and the radius value of the second channel instead of the ratio between the radius value of the first channel and the radius value of the second channel, the extended decoding unit 2222-m obtains the value MQr(f) of each radius plus the value obtained by dividing the corresponding value of the characteristic parameter by 2 as the value VLQr(f) of each radius of the first channel, and obtains the value VRQr(f) of each radius of the second channel obtained by subtracting the value MQr(f) of each radius from the value MQr(f) of each radius from the value of the corresponding value of the characteristic parameter divided by 2. The extended decoding unit 2222-m then performs an inverse complex DFT on a sequence of complex numbers with a radius of VLQr(f) and an angle of MQθ(f) on the complex plane to obtain and output a decoded digital sound signal of the first channel, and performs an inverse complex DFT on a sequence of complex numbers with a radius of VRQr(f) and an angle of MQθ(f) on the complex plane to obtain and output a decoded digital sound signal of the second channel (steps S2222-224).
[0103] [[Playback unit 223-m]]
[0104] The playback unit 223 - m outputs audio corresponding to the input two-channel decoded digital audio signal (step S223 ).
[0105] The playback unit 223-m includes, for example, two DA conversion units and two speakers. The DA conversion unit converts the input decoded digital sound signal into an analog electrical signal and outputs it. The speakers produce sound corresponding to the analog electrical signal input from the DA conversion unit. The speakers can also be configured as stereo headphones or stereo earphones. In this case, for example, the playback unit 223-m associates the DA conversion unit with the speakers on a one-to-one basis, and produces sounds (decoded sound signals) corresponding to the two decoded digital sound signals, respectively, from the two speakers.
[0106] Furthermore, all or part of the playback unit 223-m may not be located within the audio signal receiving device 220-m, but may be connected to the audio signal receiving device 220-m. For example, the playback unit 223-m of the audio signal receiving device 220-m may not include a speaker, but may output two analog electrical signals generated by the DA converter of the playback unit 223-m of the audio signal receiving device 220-m to a speaker connected to the audio signal receiving device 220-m. Alternatively, the audio signal receiving device 220-m may not include the playback unit 223-m, but may instead have the decoder 222-m of the audio signal receiving device 220-m output two channels of decoded digital audio signals to a playback device such as a DA converter connected to the audio signal receiving device 220-m.
[0107] [Operation Example of the Sound Signal Receiving Device 220-m]
[0108] Figure 5 This diagram schematically shows the temporal relationship between the monaural code included in the first code string input from the first communication line 410-m to the sound signal receiving side device 220-m, the extended code included in the second code string input from the second communication line 510-m to the sound signal receiving side device 220-m, and the decoded sound signal output by the sound signal receiving side device 220-m, excluding processing delays that depend on the processing capabilities of the device. Figure 5 The horizontal axis is the time axis. The sequence number i in parentheses is the frame number in the encoding device 212-m' of the voice signal transmitting device 210-m' of the multi-line supporting terminal device 200-m' of the other party. CM(i) is the monaural code included in the first code string input from the first communication line 410-m to the voice signal receiving device 220-m. CE(i) is the extension code included in the second code string input from the second communication line 510-m to the voice signal receiving device 220-m. YS'(i) is the decoded voice signal output by the voice signal receiving device 220-m. Figure 5This is an example in which, although the second code string is input from the second communication line 510-m, which is a communication network with a low priority, to the sound signal receiving side device 220-m in frame number order, the second code string is input 5 frames later than the first code string in frame number order from the first communication line 410-m, which is a communication network with a high priority.
[0109] When receiving the first code string including the monaural code CM(6) of frame number 6 from the first communication line 410-m, the receiving unit 221-m outputs the monaural code CM(6) included in the first code string input from the first communication line 410-m and the extension code CE(1) included in the second code string having the frame number closest to the monaural code CM(6) among the second code strings input from the second communication line 510-m to the decoding unit 222-m. When receiving the monaural code CM(6) and the extension code CE(1), the decoding unit 222-m obtains a decoded digital audio signal of two channels corresponding to the input monaural code CM(6) and the extension code CE(1) and outputs it to the playback unit 223-m. The playback unit 223-m starts outputting the two-channel decoded audio signals YS'(6) corresponding to the two input decoded digital audio signals from the moment the two-channel decoded digital audio signals corresponding to the monaural code CM(6) and the extended code CE(1) are input. Thus, the audio signal receiving side device 220-m can obtain the two-channel decoded audio signals YS'(6) from the monaural code CM(6) of frame number 6 and the extended code CE(1) included in the second code string having the closest frame number to the monaural code CM(6) and start outputting them when the receiving unit 221-m receives the first code string including the monaural code CM(6) of frame number 6 from the first communication line 410-m.
[0110] Similarly, the sound signal receiving side device 220-m and thereafter, when the receiving unit 221-m receives the first code string including the mono code CM(7) of the frame number 7 from the first communication line 410-m, it obtains the decoded sound signal YS'(7) of the two channels based on the mono code CM(7) of the frame number 7 and the extension code CE(2) included in the second code string with the frame number closest thereto, and starts outputting it. When the receiving unit 221-m receives the first code string including the mono code CM(8) of the frame number 8 from the first communication line 410-m, it obtains the decoded sound signal YS'(8) of the two channels based on the mono code CM(8) of the frame number 8 and the extension code CE(3) included in the second code string with the frame number closest thereto, and starts outputting it. ..., and the operation is carried out in this manner.
[0111] Figure 6This diagram schematically shows the temporal relationship between the monaural code included in the first code string input from the first communication line 410-m to the sound signal receiving side device, the extended code included in the second code string input from the second communication line 510-m to the sound signal receiving side device 220-m, and the decoded sound signal output by the sound signal receiving side device when the technology of patent document 1 is used, excluding the processing delay that depends on the processing capability of the device. Figure 6 The horizontal axis, the serial number i in the brackets, CM(i), CE(i) and Figure 5 YS(i) is a decoded audio signal output by the audio signal receiving device using the technology of Patent Document 1. Figure 6 also with Figure 5 Similarly, although the second code string is input from the second communication line 510-m, which is a communication network with a low priority, to the sound signal receiving side device in frame number order, the second code string is input 5 frames later than the first code string in frame number order from the first communication line 410-m, which is a communication network with a high priority. Figure 6 This is an example in which the aforementioned limited time in the audio signal receiving device using the technology of Patent Document 1 is a time corresponding to five frames.
[0112] The sound signal receiving side device using the technology of Patent Document 1 obtains a decoded sound signal YS(6) of two channels corresponding to the monaural code CM(6) input from the first communication line 410-m and the extended code CE(6) input from the second communication line 510-m exactly after a limit time of 5 frames has passed since the monaural code CM(6) was input, and starts outputting it. The sound signal receiving side device using the technology of Patent Document 1 also obtains and starts outputting the decoded sound signal YS(7) of two channels based on the mono code CM(7) of frame number 7 and the extension code CE(7) of frame number 7 input from the second communication line 510-m at the time when five frames have passed since the mono code CM(7) was received from the first communication line 410-m, and obtains and starts outputting the decoded sound signal YS(8) of two channels based on the mono code CM(8) of frame number 8 and the extension code CE(8) of frame number 8 input from the second communication line 510-m at the time when five frames have passed since the mono code CM(8) was received from the first communication line 410-m, and operates in this way.
[0113] 〔Effect〕
[0114] according to Figure 6 and Figure 5It can also be confirmed that, in the technology of Patent Document 1, a delay of five frames is required to obtain a high-quality decoded audio signal, compared to obtaining a decoded audio signal of minimum quality. However, in the technology of the first embodiment, a high-quality decoded audio signal can be obtained without significantly increasing the delay time compared to the case of obtaining a decoded audio signal of minimum quality, that is, with a delay time that does not cause a sense of discomfort during a two-way conversation.
[0115] <Second embodiment>
[0116] In the first embodiment, a spreading code is obtained and output for each frame, but a spreading code may be obtained and output only once for a plurality of frames. This method will be described as a second embodiment.
[0117] The second embodiment differs from the first embodiment in the operations of the signal analysis unit 2121-m and the transmission unit 213-m of the encoding device 212-m of the audio signal transmitting device 210-m.
[0118] [[[Signal Analysis Unit 2121-m]]]
[0119] The signal analysis unit 2121-m is similar to the signal analysis unit 2121-m of the first embodiment. For each frame, the signal analysis unit 2121-m obtains and outputs a monaural signal that is a signal obtained by mixing the digital sound signals of the two channels to be input based on the digital sound signals of the two channels to be input. However, unlike the signal analysis unit 2121-m of the first embodiment, the signal analysis unit 2121-m obtains and outputs an extended code that represents a characteristic parameter only for a predetermined frame among multiple frames. The characteristic parameter is a parameter that represents a characteristic of the difference between the digital sound signals of the two channels to be input and has little temporal variation.
[0120] For example, for frames with odd frame numbers, signal analysis unit 2121-m obtains characteristic parameters from the input two-channel digital audio signals and obtains and outputs a spreading code representing the characteristic parameters. However, for frames with even frame numbers, no characteristic parameters are obtained, and no spreading code representing the characteristic parameters are obtained, and no such parameters are output. Furthermore, if signal analysis unit 2121-m employs a configuration that uses characteristic parameters when obtaining a monaural signal, signal analysis unit 2121-m obtains a monaural signal for frames for which no characteristic parameters are obtained, using the input two-channel digital audio signals for that frame and the characteristic parameters corresponding to the most recent spreading code among the spreading codes that have been output.
[0121] Alternatively, for example, the signal analysis unit 2121-m obtains characteristic parameters from the input two-channel digital sound signals for frames with odd frame numbers, but does not obtain a spreading code representing the characteristic parameters and does not output them. For frames with even frame numbers, the signal analysis unit 2121-m obtains characteristic parameters from the input two-channel digital sound signals and obtains a spreading code representing the average or weighted average of the characteristic parameters of the immediately preceding frame (for which no spreading code representing the characteristic parameters was obtained and not output) and the characteristic parameters of the current frame, and outputs the result. The weight used for the weighted average may be set to a value greater than the weight of the immediately preceding frame.
[0122] The above two examples are configured to obtain and output a spreading code once for two frames, but a spreading code may be obtained and output once for three or more frames, or a spreading code may be obtained and output for a predetermined frame among a plurality of frames.
[0123] That is, the encoding device 212-m of the second embodiment obtains a monaural code representing a signal formed by mixing the digital sound signals of the two channels to be input for each frame, and obtains an extended code representing a characteristic parameter for a predetermined frame among multiple frames, wherein the characteristic parameter is a parameter having a low time resolution and representing the difference between the channels of the digital sound signals of the two channels to be input.
[0124] Alternatively, the encoding device 212-m of the second embodiment obtains a monaural code representing a signal obtained by mixing two channels of input digital audio signals for each frame, obtains a feature parameter representing a characteristic difference between the channels of the input two channels of digital audio signals and having a low temporal resolution for each frame, and obtains a spreading code representing an average or weighted average of the feature parameters obtained in frames subsequent to the immediately preceding predetermined frame for a predetermined frame among a plurality of frames. The weight used for the weighted average may be set to the maximum value for the frame in question and to a smaller value for frames further away from the predetermined frame.
[0125] As described later, the monaural code obtained by the encoding device 212-m is included in the first code string and output to the first communication channel, and the spread code obtained by the encoding device 212-m is included in the second code string and output to the second communication channel.
[0126] [[Sending unit 213-m]]
[0127] The sending unit 213-m outputs a code string including an input mono code, i.e., a first code string, to the first communication line 410-m for each frame, similarly to the sending unit 213-m of the first embodiment. However, unlike the sending unit 213 of the first embodiment, the sending unit 213-m outputs a code string including an input extended code, i.e., a second code string, to the second communication line 510-m only for frames to which an extended code is input, i.e., only for predetermined frames among a plurality of frames.
[0128] 〔Effect〕
[0129] As described in the first embodiment, the spreading code used by the sound signal receiving device 220-m is the spreading code with the closest frame number to the monaural code. Therefore, a spreading code with the same frame number as the monaural code is not necessarily input to the sound signal receiving device 220-m. Furthermore, characteristic parameters inherently have little temporal variation. Therefore, according to this embodiment, by adopting a configuration that obtains and outputs a spreading code only once for multiple frames, the processing load of the signal analysis unit 2121-m can be reduced without significantly degrading the quality of the decoded sound signal compared to the first embodiment. Furthermore, the amount of code used to transmit the characteristic parameters can be reduced compared to the first embodiment.
[0130] <Third embodiment>
[0131] In the first embodiment, the audio signal receiving device 220-m obtains a spreading code for decoding for each frame. However, the audio signal receiving device 220-m may obtain a spreading code for decoding only once for a plurality of frames. This method will be described as a third embodiment.
[0132] The sound signal receiving device 220-m of the third embodiment differs from the sound signal receiving device 220-m of the first embodiment in the operations of the receiving unit 221-m and the extension decoding unit 2222-m of the decoding device 222-m. The differences between the third embodiment and the first embodiment are described below.
[0133] [[Receiving unit 221-m]]
[0134] Similar to the receiving unit 221-m of the first embodiment, the receiving unit 221-m outputs the monaural code included in the first code string input from the first communication line 410-m to the decoding device 222-m for each frame. However, unlike the receiving unit 221-m of the first embodiment, the receiving unit 221-m obtains and outputs the spreading code with the frame number closest to the monaural code from among the spreading codes included in the input second code string, only for predetermined frames among the multiple frames. More specifically, the receiving unit 221-m obtains and outputs the spreading code with the frame number closest to the monaural code from a storage unit (not shown) within the receiving unit 221-m, only for predetermined frames among the multiple frames.
[0135] [[[Extended decoding unit 2222-m]]]
[0136] Similar to the extension decoding unit 2222-m of the first embodiment, the monaural decoded digital audio signal output by the monaural decoding unit 2221-m is input to the extension decoding unit 2222-m for each frame. However, unlike the extension decoding unit 2222-m of the first embodiment, the extension code is input only for predetermined frames among the multiple frames. Similar to the extension decoding unit 2222-m of the first embodiment, the extension decoding unit 2222-m obtains and outputs two-channel decoded digital audio signals based on the input monaural decoded digital audio signal and the extension code for the predetermined frames among the multiple frames, i.e., frames to which the extension code has been input. However, different from the extension decoding unit 2222-m of the first embodiment, the extension decoding unit 2222-m obtains and outputs two-channel decoded digital audio signals based on the input monaural decoded digital audio signal and the latest extension code among the input extension codes.
[0137] That is, the decoding device 222-m obtains and outputs a decoded digital sound signal of two channels for a predetermined frame among a plurality of frames based on the mono code contained in the first code string input from the first communication line 410-m and the extended code whose frame number is closest to the mono code contained in the second code string input from the second communication line 510-m. For frames other than the predetermined frame, the decoding device 222-m obtains and outputs a decoded digital sound signal of two channels based on the mono code contained in the first code string input from the first communication line 410-m and the latest extended code used in the predetermined frame. Specifically, the decoding device 222-m obtains and outputs a two-channel decoded digital sound signal based on the mono code (i.e., the mono code in the frame number sequence) contained in the first code string input from the first communication line 410-m and the extension code with the same frame number as the mono code, when the extension code included in the second code string input from the second communication line 510-m includes an extension code having the same frame number as the monaural code (i.e., the monaural code in the frame number sequence) contained in the first code string input from the first communication line 410-m, and outputs the decoded digital sound signal, and when the extension code included in the second code string input from the second communication line 510-m does not include the same frame number as the monaural code (i.e., the monaural code in the frame number sequence) contained in the first code string input from the first communication line 410-m. In the case where the same extension code is used for the frames (i.e., the mono code in the order of the frame numbers), decoded digital sound signals of two channels are obtained and output based on the mono code contained in the first code string input from the first communication line 410-m (i.e., the mono code in the order of the frame numbers) and the extension code with a frame number closest to the mono code (i.e., the extension code with a frame number that is different from the mono code but closest to the mono code) contained in the second code string input from the second communication line 510-m. For frames other than the predetermined frames, decoded digital sound signals of two channels are obtained and output based on the mono code contained in the first code string input from the first communication line 410-m (i.e., the mono code in the order of the frame numbers) and the latest extension code used in the predetermined frame.
[0138] More specifically, monaural decoding section 2221-m of decoding device 222-m decodes the monaural code included in the first code string input from first communication line 410-m for each frame to obtain a monaural decoded digital audio signal. Extension decoding section 2222-m of decoding device 222-m treats the monaural decoded digital audio signal as a mixture of two-channel decoded digital audio signals for a predetermined frame among a plurality of frames. It uses a feature parameter obtained based on the extension code whose frame number, included in the second code string input from second communication line 510-m, is closest to the monaural code included in the first code string input from first communication line 410-m as information representing the characteristics of the inter-channel difference in the two-channel decoded digital audio signals, thereby obtaining and outputting the two-channel decoded digital audio signals. Furthermore, extension decoding section 2222-m uses the feature parameter obtained based on the extension code in the predetermined frame, allowing it to store the feature parameter in advance and use it in frames other than the predetermined frame. That is, in frames other than predetermined frames, the extended decoding unit 2222-m regards the mono decoded digital sound signal as a signal obtained by mixing the decoded digital sound signals of the two channels, regards the latest feature parameters obtained in the predetermined frame as information representing the characteristics of the difference between the channels in the decoded digital sound signals of the two channels, obtains the decoded digital sound signals of the two channels and outputs them.
[0139] That is, the mono decoding unit 2221-m of the decoding device 222-m decodes the mono code (i.e., the mono code in the order of the frame numbers) contained in the first code string input from the first communication line 410-m for each frame to obtain a mono decoded digital sound signal, and the extended decoding unit 2222-m of the decoding device 222-m decodes the extended code contained in the second code string input from the second communication line 510-m for a predetermined frame among the plurality of frames, including the frame number and the number of the extended code contained in the first code string input from the first communication line 410-m. In the case where the mono code (i.e., the mono code in the order of the frame numbers) is the same as the extended code, the mono decoded digital sound signal is regarded as a signal obtained by mixing the decoded digital sound signals of the two channels, and the characteristic parameters obtained based on the extended code with the same frame number as the mono code are regarded as information representing the characteristics of the difference between the channels in the decoded digital sound signals of the two channels, and the decoded digital sound signals of the two channels are obtained and output, and the extended code included in the second code string input from the second communication line 510-m does not include the frame number and the extended code from the first channel. In the case where the mono code (i.e., the mono code in the order of the frame numbers) included in the first code string input from the second communication line 410-m is the same as the spreading code, the decoded mono digital audio signal is regarded as a signal obtained by mixing the decoded digital audio signals of the two channels, and the spreading code (i.e., the spreading code having a frame number that is closest to the mono code included in the first code string input from the first communication line 410-m, although not identical to the mono code) based on the frame number included in the second code string input from the second communication line 510-m is used. The characteristic parameters obtained are regarded as information expressing the characteristics of the difference between the channels in the decoded digital sound signals of the two channels, and the decoded digital sound signals of the two channels are obtained and output. In frames other than predetermined frames, the decoded digital sound signal of the mono channel is regarded as a signal obtained by mixing the decoded digital sound signals of the two channels, and the latest characteristic parameters obtained in the predetermined frames are regarded as information expressing the characteristics of the difference between the channels in the decoded digital sound signals of the two channels, and the decoded digital sound signals of the two channels are obtained and output.
[0140] <Modification of the Third Embodiment>
[0141] Alternatively, instead of the third embodiment, the extended decoding unit 2222-m may perform the same operation as the first embodiment, and the receiving unit 221-m may output, for a predetermined frame among a plurality of frames, the mono code included in the first code string input from the first communication line 410-m and the extended code whose frame number is closest to the mono code among the extended codes included in the second code string input from the second communication line 510-m, and output, for frames other than the predetermined frame among the plurality of frames, the mono code included in the first code string input from the first communication line 410-m and the latest extended code among the extended codes already output.
[0142] More specifically, the receiving unit 221-m may also output, with respect to a predetermined frame among a plurality of frames, the mono code and the extended code with the same frame number as the mono code when the extended code included in the second code string input from the second communication line 510-m includes an extended code with the same frame number as the mono code included in the first code string input from the first communication line 410-m (i.e., a mono code in a frame number sequence), and the extended code with the same frame number as the mono code when the extended code included in the second code string input from the second communication line 510-m does not include an extended code with the same frame number as the mono code included in the first code string input from the first communication line 410-m (i.e., a mono code in a frame number sequence). The mono code contained in the first code string input from the communication line 410-m (i.e., the mono code in frame number order), and the extended code whose frame number is closest to the mono code among the extended codes contained in the second code string input from the second communication line 510-m (i.e., the extended code whose frame number is different from the mono code but closest to the mono code among the extended codes contained in the second code string input from the second communication line 510-m), with respect to frames other than predetermined frames among the multiple frames, the mono code contained in the first code string input from the first communication line 410-m (the mono code in frame number order), and the latest extended code among the extended codes already output are output.
[0143] 〔Effect〕
[0144] As described in the first embodiment, the spreading code used by audio signal receiving device 220-m is the spreading code with the closest frame number to the monaural code. Therefore, the spreading code with the same frame number as the monaural code is not necessarily input to spreading decoding section 2222-m. Furthermore, characteristic parameters inherently exhibit minimal temporal fluctuations. Therefore, according to this embodiment and this variation, by employing a configuration in which only a single spreading code is obtained for multiple frames, the amount of computational processing performed by receiving section 221-m and the amount of information output can be reduced without significantly degrading the quality of the decoded audio signal compared to the first embodiment.
[0145] <Fourth embodiment>
[0146] As the characteristic parameters used by the sound signal receiving device 220-m in the first embodiment when obtaining the two decoded digital sound signals, the average or weighted average of the characteristic parameters represented by the spreading code input in the processing target frame and the characteristic parameters of the previous frames may be used. This method will be described as the fourth embodiment.
[0147] The fourth embodiment differs from the first embodiment in the operation of the extended decoding unit 2222-m of the decoding device 222-m of the audio signal receiving device 220-m. The following describes the differences between the fourth embodiment and the first embodiment. Hereinafter, the frame being processed by the extended decoding unit 2222-m on a frame-by-frame basis is referred to as the current frame, and the previous frame is referred to as the past frame.
[0148] [[[Extended decoding unit 2222-m]]]
[0149] Similar to the extended decoding unit 2222-m of the first embodiment, the mono decoded digital audio signal output by the mono decoding unit 2221-m and the extension code input to the decoding device 222-m are input to the extended decoding unit 2222-m for each frame. The extended decoding unit 2222-m includes a storage unit (not shown). The storage unit stores feature parameters obtained by the extended decoding unit 2222-m in previous frames. For each frame, the extended decoding unit 2222-m obtains a two-channel decoded digital audio signal based on the input mono decoded digital audio signal, the input extension code, and the feature parameters of previous frames stored in the storage unit, and outputs it to the playback unit 223-m. Specifically, the extended decoding unit 2222-m performs the following steps S2222-31 to S2222-35 for each frame.
[0150] The extended decoding unit 2222-m first obtains the characteristic parameters represented by the input extended code (step S2222-31), and stores the obtained characteristic parameters in the storage unit (step S2222-32). The extended decoding unit 2222-m then reads out K (K is an integer greater than 1) of the characteristic parameters of the past frames stored in the storage unit (step S2222-33). For example, the characteristic parameters of the past K frames that are continuous with the current frame are read out. The extended decoding unit 2222-m then obtains the average or weighted average of the characteristic parameters of the K past frames read from the storage unit and the characteristic parameters of the current frame (step S2222-34). The weight used for the weighted average can be set to the maximum value for the characteristic parameters of the current frame and to a smaller value for frames farther away from the current frame. The extended decoding unit 2222-m then treats the input mono decoded digital audio signal and the average or weighted average of the feature parameters obtained in step S2222-34 as a mixture of two decoded digital audio signals. The average or weighted average of the feature parameters obtained in step S2222-34 is then used as information representing the difference between the two decoded digital audio signals. The two decoded digital audio signals are then output to the playback unit 223-m (step S2222-35). Alternatively, instead of storing the feature parameters represented by the extension code in the storage unit in step S2222-32, the extended decoding unit 2222-m may store the average or weighted average obtained in step S2222-34 as the feature parameters for the current frame. Furthermore, since the storage unit of the extended decoding unit 2222-m only needs to store K feature parameters for past frames, feature parameters for K+1 or more past frames can be deleted from the storage unit during processing of the frame following the current frame.
[0151] <Modification of the Fourth Embodiment>
[0152] Similar to the sound signal receiving device 220-m of the first embodiment, the sound signal receiving device 220-m of the third embodiment may use an average or weighted average of the characteristic parameters represented by the spreading code input in the target frame and characteristic parameters of previous frames as the characteristic parameters used when obtaining the two decoded digital sound signals. Specifically, the extension decoding unit 2222-m of the decoding device 222-m of the sound signal receiving device 220-m of the third embodiment may use an average or weighted average of the characteristic parameters represented by the spreading code input in the target frame and characteristic parameters of previous frames as the characteristic parameters used when obtaining the two decoded digital sound signals for a predetermined frame among a plurality of frames. This embodiment will be described as a modified example of the fourth embodiment.
[0153] The variation of the fourth embodiment differs from the third embodiment in the operation of the extended decoding unit 2222-m of the decoding device 222-m of the audio signal receiving device 220-m. The following describes the differences between the variation of the fourth embodiment and the third embodiment. Hereinafter, the frame being processed by the extended decoding unit 2222-m at that moment is referred to as the current frame, and the previous frame is referred to as the past frame.
[0154] [[[Extended decoding unit 2222-m]]]
[0155] Similar to the extension decoding unit 2222-m of the third embodiment, the mono decoded digital audio signal output by the monaural decoding unit 2221-m is input to the extension decoding unit 2222-m for each frame. The extension code is input to the extension decoding unit 2222-m only for predetermined frames among the plurality of frames. The extension decoding unit 2222-m includes a storage unit (not shown). The storage unit stores at least the average or weighted average of the feature parameters obtained by the extension decoding unit 2222-m for past frames, and may also store the feature parameters represented by the extension code of the past frame.
[0156] The extension decoding unit 2222-m performs the following steps S2222-41 to S2222-46 on a predetermined frame among a plurality of frames, that is, a frame to which the extension code is also input.
[0157] The extended decoding unit 2222-m first obtains the characteristic parameters represented by the input extended code based on the input extended code (step S2222-41) and stores the obtained characteristic parameters in the storage unit (step S2222-42). The extended decoding unit 2222-m then reads K (K is an integer greater than or equal to 1) characteristic parameters of past frames stored in the storage unit (step S2222-43). For example, the characteristic parameters of the K past frames closest to the current frame are read. Since characteristic parameters are only stored in the storage unit for frames that have also received the extended code, the characteristic parameters read are those of the K frames that are consecutive to the current frame among the frames that have also received the extended code. The extended decoding unit 2222-m then obtains the average or weighted average of the characteristic parameters of the K past frames read from the storage unit and the characteristic parameters of the current frame (step S2222-44), and stores the obtained average or weighted average of the characteristic parameters in the storage unit (step S2222-45). The weights used for weighted averaging can be set to the maximum value for the feature parameters of the current frame and to the smallest value for frames farther from the current frame. The extended decoding unit 2222-m then treats the input mono decoded digital audio signal and the average or weighted average of the feature parameters obtained in step S2222-44 as a mixture of two decoded digital audio signals, and treats the average or weighted average of the feature parameters obtained in step S2222-44 as information representing the difference between the two decoded digital audio signals. The two decoded digital audio signals are then output to the playback unit 223-m (step S2222-46). Alternatively, instead of storing the feature parameters represented by the extended code in the storage unit in step S2222-42, the extended decoding unit 2222-m may instead read the average or weighted average stored in the storage unit in step S2222-45 as the feature parameters of the previous frame in step S2222-43. Furthermore, the storage unit of the extended decoding unit 2222-m only needs to store K feature parameters of past frames. Therefore, when processing the next frame of the current frame, the feature parameters of past frames K+1 or more can be deleted from the storage unit. Furthermore, the storage unit of the extended decoding unit 2222-m only needs to store the most recent average or weighted average of the feature parameters obtained in step S2222-44. Therefore, the average or weighted average of the feature parameters stored in the storage unit at the time of performing step S2222-45 can be deleted from the storage unit.
[0158] The extension decoding unit 2222 - m of the modification of the fourth embodiment performs the following steps S2222 - 47 to S2222 - 48 on frames other than predetermined frames among a plurality of frames, that is, frames to which no extension code is input.
[0159] The extended decoding unit 2222-m first reads the average or weighted average of the most recent feature parameters stored in the storage unit from the storage unit (step S2222-47). The extended decoding unit 2222-m then treats the input mono decoded digital audio signal and the average or weighted average of the feature parameters obtained in step S2222-47 as a signal obtained by mixing two decoded digital audio signals, and treats the average or weighted average of the feature parameters obtained in step S2222-47 as information representing the difference between the two decoded digital audio signals, thereby obtaining two decoded digital audio signals and outputting them to the playback unit 223-m (step S2222-48).
[0160] 〔Effect〕
[0161] While characteristic parameters are statistically small in temporal variation, they reflect the characteristics of the audio signal in each frame and therefore rarely have exactly the same value across multiple frames. Furthermore, their values may vary significantly between frames. Therefore, in the audio signal receiving device 220-m, by averaging or weighted averaging characteristic parameters represented by multiple temporally close spreading codes, as in the fourth embodiment and its variations, compared to characteristic parameters represented by a spreading code different from the original spreading code of the frame, it is possible to suppress sudden fluctuations between channels in the decoded audio signal and the occurrence of unusual sounds.
[0162] <Fifth embodiment>
[0163] In the first embodiment, the sound signal receiving device 220-m uses a monaural code and the spreading code with the closest frame number for each frame to obtain a two-channel decoded digital sound signal. However, for frames without a spreading code within a specific limited time range with the monaural code, the monaural code may be decoded and the obtained decoded digital sound signal may be used as a two-channel decoded digital sound signal. This method will be described as the fifth embodiment.
[0164] The fifth embodiment differs from the first embodiment in the operations of the receiving unit 221-m and the decoding unit 222-m of the sound signal receiving device 220-m. Furthermore, within the decoding unit 222-m, the fifth embodiment performs operations that differ from those of the first embodiment in the extension decoding unit 2222-m. The differences between the fifth embodiment and the first embodiment are described below.
[0165] [[Receiving unit 221-m]]
[0166] Receiving unit 221-m outputs the monaural code included in the first code string input from first communication channel 410-m and the spreading code included in the second code string input from second communication channel 510-m, with the frame number closest to the monaural code, for frames in which the difference in frame number between the monaural code included in the first code string input from first communication channel 410-m and the spreading code included in the second code string input from second communication channel 510-m is less than a predetermined value. Receiving unit 221-m outputs the monaural code included in the first code string input from first communication channel 410-m, for frames in which the difference in frame number is not less than the predetermined value. Specifically, receiving unit 221-m performs the following steps S221-11 to S221-15 for each frame.
[0167] Receiving unit 221-m outputs the mono code contained in the first code string input from first communication line 410-m to decoding device 222-m (step S221-11). Receiving unit 221-m then obtains the frame number of the mono code output in step S221-11 (step S221-12). Receiving unit 221-m then obtains the extension code contained in the second code string input from second communication line 510-m, the second code string having a frame number closest to the frame number of the mono code obtained in step S221-12, and the frame number of the extension code (step S221-13). Receiving unit 221-m then determines whether the difference between the frame number of the mono code obtained in step S221-12 and the frame number of the extension code obtained in step S221-13 is less than a predetermined value (step S221-14). Next, if the difference between the monaural code frame number and the extended code frame number in step S221-14 is less than a predetermined value, the receiving unit 221-m outputs the extended code to the decoding device 222-m (step S221-15). If the difference between the monaural code frame number and the extended code frame number in step S221-14 is not less than a predetermined value, the receiving unit 221-m does not output the extended code. In other words, if the difference between the monaural code frame number and the extended code frame number in step S221-14 is not less than a predetermined value, the receiving unit 221-m only outputs the monaural code.
[0168] Here, the predetermined value is a value greater than or equal to 2. That is, the receiving unit 221-m outputs the mono code (i.e., the mono code in the order of the frame numbers) included in the first code string input from the first communication line 410-m and the extended code whose frame number is closest to the mono code among the extended codes included in the second code string input from the second communication line 510-m, with respect to the frame in which the difference in frame number between the mono code (i.e., the mono code in the order of the frame numbers) included in the first code string input from the first communication line 410-m and the extended code whose frame number is closest to the mono code among the extended codes included in the second code string input from the second communication line 510-m is 0. For the extended codes with the same mono code, for frames whose frame number difference is greater than 0 and less than a predetermined value, the mono code contained in the first code string input from the first communication line 410-m (i.e., the mono code in frame number order) and the extended code contained in the second code string input from the second communication line 510-m with a frame number closest to the mono code (i.e., the extended code contained in the second code string input from the second communication line 510-m with a frame number that is different from the mono code but closest to the mono code) are output. For frames whose frame number difference is not less than a predetermined value, only the mono code contained in the first code string input from the first communication line 410-m (i.e., the mono code in frame number order) is output.
[0169] [[Decoding device 222-m]]
[0170] Decoding device 222-m receives the mono code output by receiving unit 221-m, and sometimes receives the extended code output by receiving unit 221-m, for each frame. Decoding device 222-m obtains two-channel decoded digital audio signals corresponding to the input mono code and extended code, or the input mono code, for each frame, and outputs them to playback unit 223-m. Specifically, for frames where the difference in frame number is less than a predetermined value, decoding device 222-m obtains and outputs two-channel decoded digital audio signals based on the monaural code and extended code output by receiving unit 221-m. For frames where the difference in frame number is not less than a predetermined value, decoding device 222-m outputs the monaural digital signal based on the monaural code output by receiving unit 221-m as a two-channel decoded digital audio signal.
[0171] [[[Extended decoding unit 2222-m]]]
[0172] Extension decoding unit 2222-m receives the mono decoded digital audio signal output from monaural decoding unit 2221-m for each frame, and may also receive the extension code input to decoding device 222-m. For frames receiving the mono decoded digital audio signal and the extension code, extension decoding unit 2222-m performs the same operation as extension decoding unit 2222-m in the first embodiment, obtaining a two-channel decoded digital audio signal based on the input mono decoded digital audio signal and the extension code, and outputs the signal to playback unit 223-m. For frames receiving only the mono decoded digital audio signal, extension decoding unit 2222-m obtains the input mono decoded digital audio signal as a two-channel decoded digital audio signal, and outputs the signal to playback unit 223-m.
[0173] That is, for frames in which the difference in frame number between the mono code included in the first code string input from the first communication line 410-m and the extended code whose frame number is closest to the mono code included in the second code string input from the second communication line 510-m is less than a predetermined value, the decoding device 222-m obtains and outputs decoded digital sound signals of two channels based on the mono code and the extended code whose frame number is closest to the mono code. For frames in which the difference in frame number is not less than a predetermined value, the decoded digital sound signal based on the mono code included in the first code string input from the first communication line 410-m is output as a decoded digital sound signal of two channels as it is.
[0174] More specifically, the decoding device 222-m obtains and outputs decoded digital sound signals of two channels based on the mono code and the extended code with the same frame number as the mono code, with respect to a frame in which the difference in frame number between the mono code contained in the first code string input from the first communication line 410-m (i.e., the mono code in the frame number sequence) and the extended code with the frame number closest to the mono code contained in the second code string input from the second communication line 510-m is 0 (i.e., the frame in which the second code string input from the second communication line 510-m contains an extended code with the same frame number as the mono code contained in the first code string input from the first communication line 410-m), and the difference in frame number between the two channels is greater than 0 and less than a predetermined value. The decoded digital sound signals of two channels are obtained and outputted based on the mono code (i.e., the mono code in the order of the frame numbers) contained in the first code string input from the first communication line 410-m and the extended code whose frame number is closest to the mono code (i.e., the extended code whose frame number is different from the mono code but closest to the mono code among the extended codes contained in the second code string input from the second communication line 510-m). For frames whose difference in the above-mentioned frame numbers is not less than a predetermined value, the decoded digital sound signals based on the mono code (i.e., the mono code in the order of the frame numbers) contained in the first code string input from the first communication line 410-m are outputted as decoded digital sound signals of two channels.
[0175] <Modification of the Fifth Embodiment>
[0176] The above describes the sound signal receiving side device 220-m of the fifth embodiment and its operation based on the structure of the sound signal receiving side device 220-m of the first embodiment, but the sound signal receiving side device 220-m of the fifth embodiment based on the sound signal receiving side device 220-m of the third embodiment and the fourth embodiment and one of their variations can also be constructed and operated.
[0177] 〔Effect〕
[0178] The encoding device 212-m' of the voice signal transmitting device 210-m' of the multi-channel terminal 200-m' on the other end of the call encodes frames within each specific time interval. Therefore, the difference between the frame number of the monaural code and the frame number of the extended code corresponds to the time difference between the digital voice signals encoded by the encoding device 212-m' of the multi-channel terminal 200-m' on the other end of the call. For example, if the frame length is 20ms and the frame number difference is 150, there is a time difference of 3 seconds between the digital voice signal obtained by the monaural code and the digital voice signal obtained by the extended code. Even parameters with minimal temporal fluctuations can have significantly different values if the timing differs significantly. Therefore, if there is a time difference that significantly differs the characteristic parameters represented by the extended code, there is a risk of significant errors in the signal segmentation between the two channels in the decoded voice signals of the two channels, which reflect the differential characteristics of the two channels. According to the fifth embodiment, by not applying a difference to the decoded audio signals of the two channels for frames in which the frame number difference between the monaural code included in the first code string received from the first communication channel and the spreading code closest to the monaural code in the spreading code included in the second code string received from the second communication channel is large, it is possible to suppress significant errors in signal segmentation between the channels of the decoded audio signals. For example, assuming that the characteristic parameter differs significantly when the time difference is 400 ms or longer, and that the characteristic parameter differs significantly when the frame number difference is 20 or more when the frame length is 20 ms, the predetermined value can be set to, for example, 20.
[0179] <Sixth embodiment>
[0180] The sound signal receiving device 220-m may also treat the decoded digital sound signal obtained by decoding the monaural code as a two-channel decoded digital sound signal based on the average value of the time difference between the first code string input from the first communication line 410-m and the second code string having the same frame number as the first code string input from the second communication line 510-m, measured within a specific time range, if the average value of the time difference is not within a predetermined time limit. This method will be described as the sixth embodiment.
[0181] The sixth embodiment differs from the first embodiment in the operations of the receiving unit 221-m and the decoding unit 222-m of the sound signal receiving device 220-m. Furthermore, within the decoding unit 222-m, the sixth embodiment performs operations that differ from those of the first embodiment in the extension decoding unit 2222-m. The differences between the sixth embodiment and the first embodiment are described below.
[0182] [[Receiving unit 221-m]]
[0183] The first code string output by the voice signal transmitting device 210-m' of the other party is input to the receiving unit 221-m via the first communication line 410-m, and the second code string output by the voice signal transmitting device 210-m' of the other party is input to the receiving unit 221-m via the second communication line 510-m. The second communication line is a low-priority communication network. Therefore, the second code string of a frame output by the voice signal transmitting device 210-m' of the other party is typically input to the receiving unit 221-m via the second communication line 510-m after the first code string of the frame is input to the receiving unit 221-m via the first communication line 410-m.
[0184] Receiving unit 221-m first determines, for each pair consisting of a first code string received from first communication channel 410-m and a corresponding second code string received from second communication channel 510-m, whether the average value of the difference between the times at which the first and second code strings were received, across multiple pairs, is less than a predetermined time limit Tmax. Time limit Tmax is, for example, 400 ms.
[0185] For example, receiving unit 221-m performs the following steps S221-21 to S221-24. Receiving unit 221-m reads the frame number for a predetermined number of first code strings starting from the start of receiving the first code string, measures the time of reception, and stores the frame number in a storage unit (not shown) within receiving unit 221-m, associating it with the time of reception of the first code string (step S221-21). Receiving unit 221-m also reads the frame number for the received second code string, and if the read frame number matches one of the frame numbers stored in the storage unit, measures the time of reception, and stores the time of reception of the second code string in the storage unit, associating it with the frame number stored in the storage unit and the time of reception of the first code string (step S221-22). The receiving unit 221-m then uses the frame number, the time when the first code string was received, and the time when the second code string was received, which are stored in association with each other in the storage unit, to obtain the average value of the predetermined number of values obtained by subtracting the time when the first code string was received from the time when the second code string was received for each frame number (step S221-23). The receiving unit 221-m then determines whether the average value obtained in step S221-23 is less than the predetermined time limit Tmax (step S221-24).
[0186] Next, if the average value is less than the time limit Tmax in the above determination, receiving unit 221-m outputs the monaural code included in the first code string input from first communication channel 410-m and the extended code included in the second code string input from second communication channel 510-m, whose frame number is closest to the monaural code, to decoding device 222-m for subsequent frames. If the average value is not less than the time limit Tmax in the above determination, receiving unit 221-m outputs the monaural code included in the first code string input from first communication channel 410-m to decoding device 222-m for subsequent frames. If the average value is not less than the time limit Tmax in the above determination, receiving unit 221-m does not output the extended code for subsequent frames. In other words, receiving unit 221-m only outputs the monaural code if the average value is not less than the time limit Tmax in the above determination.
[0187] That is, the receiving unit 221-m, with respect to a group consisting of a first code string received from the first communication line 410-m and a second code string received from the second communication line 510-m corresponding to the first code string, when the average value of the difference between the time when the first code string and the second code string are received among multiple groups is less than a predetermined limit time Tmax, with respect to subsequent frames, when the extended code contained in the second code string input from the second communication line 510-m includes an extended code having the same frame number as the mono code (i.e., a mono code in the order of frame numbers) contained in the first code string input from the first communication line 410-m, the mono code and the extended code having the same frame number as the mono code are output to the decoding device 222-m, and when the extended code contained in the second code string input from the second communication line 510-m does not include the same frame number as the mono code input from the first communication line 410-m In the case where the mono code included in the first code string input from the first communication line 410-m (i.e., the mono code in sequence of the frame numbers) is the same as the extended code, the mono code included in the first code string input from the first communication line 410-m (i.e., the mono code in sequence of the frame numbers) and the extended code included in the second code string input from the second communication line 510-m, the frame number of which is closest to the mono code (i.e., the frame number of which is different from the mono code but closest to the mono code) are output to the decoding device 222-m. When the above-mentioned average value is not less than the limit time Tmax, for subsequent frames, only the mono code included in the first code string input from the first communication line 410-m (i.e., the mono code in sequence of the frame numbers) is output to the decoding device 222-m.
[0188] In addition, until the above-mentioned judgment is completed, the receiving unit 221-m may output nothing, or may output the mono code and the extended code to the decoding device 222-m in the same way as the first embodiment, or may output the mono code to the decoding device 222-m without outputting the extended code, or may always output the mono code to the decoding device 222-m in the same way as the fifth embodiment, and output the extended code to the decoding device 222-m only when the difference in frame number between the mono code and the extended code is small.
[0189] [[Decoding device 222-m]]
[0190] If the average value determined by the receiving unit 221-m is less than the predetermined time limit Tmax, the monaural code and the extended code are input to the decoding unit 222-m for each frame, similar to the decoding unit 222-m of the first embodiment. If the average value determined by the receiving unit 221-m is not less than the predetermined time limit Tmax, the monaural code output by the receiving unit 221-m is input to the decoding unit 222-m for each frame without the extended code.
[0191] Furthermore, until the above-described determination by receiving unit 221-m is completed, nothing is input to decoding unit 222-m, or a monaural code is input instead of an extension code, or both a monaural code and an extension code are input. Decoding unit 222-m obtains a two-channel decoded digital audio signal corresponding to the input monaural code and extension code, or the input monaural code, for each frame, and outputs it to playback unit 223-m.
[0192] [[[Extended decoding unit 2222-m]]]
[0193] When a monaural decoded digital audio signal and a spread code are input, that is, when the average value is less than the predetermined time limit Tmax in the above-mentioned determination, the extension decoding unit 2222-m performs the same operation as the extension decoding unit 2222-m of the first embodiment on a frame-by-frame basis based on the input monaural decoded digital audio signal and the spread code, thereby obtaining a two-channel decoded digital audio signal and outputting it to the playback unit 223-m. When a monaural decoded digital audio signal is input, that is, when the average value is not less than the predetermined time limit Tmax in the above-mentioned determination, the extension decoding unit 2222-m obtains the input monaural decoded digital audio signal as a two-channel decoded digital audio signal without modification and outputs it to the playback unit 223-m.
[0194] That is, the decoding device 222-m obtains and outputs a decoded digital sound signal for two channels based on the mono code contained in the first code string input from the first communication line 410-m and the extended code whose frame number is closest to the mono code contained in the second code string input from the second communication line 510-m, with respect to a group consisting of a first code string received from the first communication line 410-m and a second code string received from the second communication line 510-m corresponding to the first code string, when the average value of the difference between the times when the first code string and the second code string are received among multiple groups is less than a predetermined limit time Tmax. When the above-mentioned average value is not less than the limit time Tmax, the decoded digital sound signal for the mono channel based on the mono code contained in the first code string input from the first communication line 410-m is output as a decoded digital sound signal for two channels as it is.
[0195] More specifically, the decoding device 222-m, with respect to a group consisting of a first code string received from the first communication line 410-m and a second code string received from the second communication line 510-m corresponding to the first code string, when the average value of the difference between the time when the first code string and the second code string are received among multiple groups is less than a predetermined limit time Tmax, with respect to a frame in which the extended code contained in the second code string input from the second communication line 510-m contains an extended code having the same frame number as the mono code (i.e., a mono code in the order of frame numbers) contained in the first code string input from the first communication line 410-m, obtains and outputs a decoded digital sound signal of two channels based on the mono code and the extended code having the same frame number as the mono code, and with respect to a frame in which the extended code contained in the second code string input from the second communication line 510-m does not contain a frame number identical to the first code string input from the first communication line 410-m. Frames having the same extended code as the mono code contained in the code string (i.e., the mono code in sequence of frame numbers) are obtained and output as decoded digital sound signals for two channels based on the mono code contained in the first code string input from the first communication line 410-m (i.e., the mono code in sequence of frame numbers) and the extended code with a frame number closest to the mono code contained in the second code string input from the second communication line 510-m (i.e., the extended code with a frame number that is different from the mono code but closest to the mono code among the extended codes contained in the second code string input from the second communication line 510-m). When the above-mentioned average value is not less than the limit time Tmax, the decoded digital sound signal for the mono channel based on the mono code contained in the first code string input from the first communication line 410-m (i.e., the mono code in sequence of frame numbers) is output as is as decoded digital sound signals for two channels.
[0196] In addition, until the above-mentioned judgment performed by the receiving unit 221-m is completed, the extended decoding unit 2222-m obtains a two-channel decoded digital sound signal and outputs it to the playback unit 223-m based on the input mono decoded digital sound signal and the extended code, or obtains the input mono decoded digital sound signal as a two-channel decoded digital sound signal as it is and outputs it to the playback unit 223-m, or outputs nothing.
[0197] <Modification of Sixth Embodiment>
[0198] The above description describes the sound signal receiving device 220-m and its operation according to the sixth embodiment, which is based on the configuration of the sound signal receiving device 220-m according to the first embodiment. However, the sound signal receiving device 220-m according to the sixth embodiment may be configured and operated based on the sound signal receiving device 220-m according to any of the third to fifth embodiments and their variations. Furthermore, in the above example, the specific time range is the period from the start of receiving the first code string until a predetermined number of first code strings are received. However, the specific time range may also be defined as starting at any time. For example, the specific time range may be the period starting at a certain time after the start of receiving the first code string, or the specific time range may be defined as each period starting at a plurality of times after the start of receiving the first code string.
[0199] 〔Effect〕
[0200] As also explained in the fifth embodiment, even characteristic parameters with minimal temporal fluctuations can have significantly different values if the timing differs significantly. Therefore, if it is determined that there is a time difference between the first and second communication channels that significantly differs the characteristic parameter represented by the spreading code, significant errors may occur in the signal segmentation between the two channels in the decoded audio signals of the two channels, which reflect the characteristics of the difference between the two channels. According to this sixth embodiment, by not applying a difference to the decoded audio signals of the two channels when there is a significant difference between the time when the first code string is received from the first communication channel and the time when the second code string is received from the second communication channel for the same frame, significant errors in the signal segmentation between the channels of the decoded audio signals can be suppressed.
[0201] <Seventh embodiment>
[0202] The sound signal receiving device 220-m may also use a monaural code and a spreading code having the same frame number as the monaural code as the decoded digital sound signal for two channels based on the average value of the time difference between the first code string input from the first communication line 410-m and the second code string input from the second communication line 510-m having the same frame number as the first code string, measured within a specific time range, and if the average value of the time difference is within a predetermined time limit. This method will be described as the seventh embodiment.
[0203] The seventh embodiment differs from the first embodiment in the operation of the receiving unit 221 - m of the audio signal receiving device 220 - m .
[0204] [[Receiving unit 221-m]]
[0205] The first code string output by the other party's voice signal transmitting device 210-m' is input to the receiving unit 221-m via the first communication line 410-m, and the second code string output by the other party's voice signal transmitting device 210-m' is input to the receiving unit 221-m via the second communication line 510-m. Since the second communication line is a low-priority communication network, the second code string of a frame output by the other party's voice signal transmitting device 210-m' is typically input to the receiving unit 221-m via the second communication line 510-m after the first code string of the frame is input to the receiving unit 221-m via the first communication line 410-m.
[0206] Receiving unit 221-m first determines, for each pair consisting of a first code string received from first communication channel 410-m and a corresponding second code string received from second communication channel 510-m, whether the average value of the difference between the times at which the first and second code strings were received, across multiple pairs, is less than a predetermined time limit Tmin. Time limit Tmin is, for example, twice the frame length. For example, if the frame length is 20 ms, time limit Tmin is, for example, 40 ms.
[0207] For example, receiving unit 221-m performs the following steps S221-31 to S221-34. Receiving unit 221-m reads the frame number of a predetermined number of first code strings from the start of receiving the first code string, measures the time of reception, and stores the frame number in a storage unit (not shown) within receiving unit 221-m, in association with the time of reception of the first code string (step S221-31). Receiving unit 221-m also reads the frame number of the received second code string. If the read frame number matches one of the frame numbers stored in the storage unit, receiving unit 221-m measures the time of reception and stores the time of reception of the second code string in the storage unit, in association with the frame number stored in the storage unit and the time of reception of the first code string (step S221-32). The receiving unit 221-m then uses the frame number, the time when the first code string was received, and the time when the second code string was received, which are stored in association with each other in the storage unit, to obtain the average value of the predetermined number of values obtained by subtracting the time when the first code string was received from the time when the second code string was received for each frame number (step S221-33). The receiving unit 221-m then determines whether the average value obtained in step S221-33 is less than the predetermined time limit Tmin (step S221-34).
[0208] The receiving unit 221-m then outputs, for subsequent frames, the mono code contained in the first code string input from the first communication line 410-m and the extended code contained in the second code string input from the second communication line 510-m, whose frame number is the same as that of the mono code, to the decoding device 222-m if the average value is less than the limit time Tmin in the above judgment. If the average value is not less than the limit time Tmin in the above judgment, for subsequent frames, the mono code contained in the first code string input from the first communication line 410-m and the extended code contained in the second code string input from the second communication line 510-m, whose frame number is closest to that of the mono code, to the decoding device 222-m. Among them, the time from when the first code string is received on the first communication line 410-m to when the second code string of the frame is received from the second communication line 510-m is assumed to be the average time required to obtain the average value in step S221-33. Therefore, the receiving unit 221-m needs to operate so that the time from when the first code string is received on the first communication line 410-m to when it is output to the decoding device 222-m becomes the average value obtained in step S221-33 or a value greater than it.
[0209] That is, the receiving unit 221-m outputs to the decoding device 222-m, with respect to a group consisting of a first code string received from the first communication line 410-m and a second code string received from the second communication line 510-m corresponding to the first code string, the mono code (i.e., the mono code in the order of the frame numbers) contained in the first code string input from the first communication line 410-m and the extended code with the same frame number as the mono code among the extended codes contained in the second code string input from the second communication line 510-m, with respect to subsequent frames, and when the above-mentioned average value is not less than the limit time Tmin, the extended code contained in the second code string input from the second communication line 510-m contains the same frame number as the first code string input from the first communication line 410-m. In the case where the extended code contained in the second code string input from the second communication line 510-m does not contain an extended code having a frame number identical to the mono code contained in the first code string input from the first communication line 410-m (i.e., the mono code in the sequence of frame numbers), the mono code contained in the first code string (i.e., the mono code in the sequence of frame numbers) and the extended code with a frame number closest to the mono code among the extended codes contained in the second code string input from the second communication line 510-m (i.e., the extended code with a frame number closest to the mono code among the extended codes contained in the second code string input from the second communication line 510-m) are output to the decoding line 222-m.
[0210] The operation of the decoding device 222-m of the sound signal receiving device 220-m of the seventh embodiment is similar to that of the decoding device 222-m of the sound signal receiving device 220-m of the first embodiment. Based on the monaural code and the extension code output by the receiving unit 221-m, the decoding device 222-m obtains and outputs a two-channel decoded digital sound signal. However, the extension code output by the receiving unit 221-m of the seventh embodiment may differ from the extension code output by the receiving unit 221-m of the first embodiment, depending on the circumstances. Therefore, the decoding device 222-m specifically performs the following operation.
[0211] That is, the decoding device 222-m obtains and outputs a decoded digital sound signal for two channels based on the monaural code contained in the first code string input from the first communication line 410-m and the extended code having the same frame number as the monaural code contained in the second code string input from the second communication line 510-m, when the average value of the difference between the times at which the first code string and the second code string are received among multiple groups is less than a predetermined limit time Tmin. When the above-mentioned average value is not less than the limit time Tmin, the decoding device 222-m obtains and outputs a decoded digital sound signal for two channels based on the monaural code contained in the first code string input from the first communication line 410-m and the extended code having the same frame number as the monaural code contained in the second code string input from the second communication line 510-m.
[0212] More specifically, the decoding device 222-m obtains and outputs a decoded digital sound signal of two channels based on the monaural code (i.e., the monaural code in the order of frame numbers) contained in the first code string input from the first communication line 410-m and the extended code having the same frame number as the monaural code contained in the second code string input from the second communication line 510-m, when the average value of the difference between the times at which the first code string and the second code string are received among a plurality of groups is less than a predetermined limit time Tmin. When the average value is not less than the limit time Tmin, the decoding device 222-m obtains and outputs a decoded digital sound signal of two channels based on the monaural code (i.e., the monaural code in the order of frame numbers) contained in the first code string input from the first communication line 410-m and the extended code having the same frame number as the monaural code contained in the second code string input from the second communication line 510-m. For frames in which the extended code contained in the second code string input from the second communication line 510-m does not contain an extended code with the same frame number as the mono code contained in the first code string input from the first communication line 410-m (i.e., a mono code with a sequential frame number), decoded digital sound signals for two channels are obtained and output. For frames in which the extended code contained in the second code string input from the second communication line 510-m does not contain an extended code with the same frame number as the mono code contained in the first code string input from the first communication line 410-m (i.e., a mono code with a sequential frame number), decoded digital sound signals for two channels are obtained and output based on the mono code contained in the first code string input from the first communication line 410-m (i.e., a mono code with a sequential frame number) and the extended code with a frame number closest to the mono code contained in the second code string input from the second communication line 510-m (i.e., an extended code with a frame number that is different from the mono code but closest to the mono code among the extended codes contained in the second code string input from the second communication line 510-m).
[0213] In addition, until the above-mentioned judgment performed by the receiving unit 221-m is completed, for example, the receiving unit 221-m can output the mono code and the extended code to the decoding device 222-m in the same way as the first embodiment, and the decoding device 222-m can use the mono code and the extended code in the same way as the first embodiment to obtain the decoded digital sound signal of the two channels and output it to the playback unit 223-m.
[0214] <Modification of Seventh Embodiment>
[0215] The above description describes the sound signal receiving device 220-m and its operation according to the seventh embodiment, which is based on the configuration of the sound signal receiving device 220-m according to the first embodiment. However, the sound signal receiving device 220-m according to the seventh embodiment may be configured and operated based on the sound signal receiving device 220-m according to any of the third to fifth embodiments and their variations. Furthermore, in the above example, the specific time range is the time from the start of receiving the first code string until a predetermined number of first code strings are received. However, the specific time range may also be defined as starting at any time. For example, the specific time range may be the time period starting at a certain time after the start of receiving the first code string, or the specific time range may be defined as each time period starting at a plurality of times after the start of receiving the first code string.
[0216] 〔Effect〕
[0217] Even feature parameters with small temporal variations may have slightly different values depending on the time. Therefore, if decoding can be performed using the feature parameters of the same frame by only slightly increasing the delay, a high-quality decoded audio signal may be obtained. Therefore, in this seventh embodiment, a predetermined time limit is set for the average value of a specific time range between the time when the first code string of the same frame is received from the first communication line and the time when the second code string is received from the second communication line. When the time limit is less than the predetermined time limit, a monaural code and an extension code of the same frame as the monaural code are used as the two-channel decoded digital audio signal, with an intentionally slightly increased delay. This allows for a high-quality decoded audio signal.
[0218] <Eighth embodiment>
[0219] The sound signal receiving device 220-m may also use the average value of the time difference between a first code string input from the first communication line 410-m and a second code string input from the second communication line 510-m having the same frame number as the first code string, measured within a specific time range. If the average value of the time difference is less than a first limit time, the device may use a monaural code and a spreading code with the same frame number as the monaural code to obtain a decoded digital sound signal for the two channels. If the average value of the time difference is greater than a predetermined second limit time that is greater than the first limit time, the decoded digital sound signal obtained by decoding the monaural code is used as the decoded digital sound signal for the two channels. If the average value of the time difference is greater than the first limit time and less than the second limit time, the device may use a monaural code and a spreading code with a frame number closest to the monaural code to obtain a decoded digital sound signal for the two channels. In short, the sixth embodiment and the seventh embodiment may also be implemented in combination. This method will be described as the eighth embodiment.
[0220] The eighth embodiment differs from the first embodiment in the operations of the receiving unit 221-m and the decoding unit 222-m of the sound signal receiving device 220-m. The operations of the decoding unit 222-m of the sound signal receiving device 220-m are identical to those of the decoding unit 222-m of the sixth embodiment. The following describes the operations of the receiving unit 221-m, which differ from both the first and sixth embodiments in the eighth embodiment.
[0221] [[Receiving unit 221-m]]
[0222] The first code string output by the other party's voice signal transmitting device 210-m' is input to the receiving unit 221-m via the first communication line 410-m, and the second code string output by the other party's voice signal transmitting device 210-m' is input to the receiving unit 221-m via the second communication line 510-m. Since the second communication line is a low-priority communication network, the second code string of a frame output by the other party's voice signal transmitting device 210-m' is typically input to the receiving unit 221-m via the second communication line 510-m after the first code string of the frame is input to the receiving unit 221-m via the first communication line 410-m.
[0223] Receiving unit 221-m first determines, for each pair consisting of a first code string received from first communication channel 410-m and a corresponding second code string received from second communication channel 510-m, whether the average value of the difference between the times at which the first and second code strings were received, across multiple pairs, is less than a predetermined first time limit Tmin, is greater than or equal to a predetermined second time limit Tmax that is greater than the first time limit Tmin, or is greater than or equal to the first time limit Tmin and less than the second time limit Tmax. Furthermore, the first time limit Tmin is, for example, twice the frame length. That is, if the frame length is 20 ms, the first time limit Tmin is, for example, 40 ms. Furthermore, the second time limit Tmax is, for example, 400 ms.
[0224] For example, receiving unit 221-m performs the following steps S221-41 to S221-44. Receiving unit 221-m reads the frame number of a predetermined number of first code strings from the start of receiving the first code string, measures the time of reception, and stores the frame number in a storage unit (not shown) within receiving unit 221-m, in association with the time of reception of the first code string (step S221-41). Receiving unit 221-m also reads the frame number of the received second code string, and if the read frame number matches one of the frame numbers stored in the storage unit, measures the time of reception, and stores the time of reception of the second code string in the storage unit, in association with the frame number stored in the storage unit and the time of reception of the first code string (step S221-42). The receiving unit 221-m then uses the frame number, the time when the first code string was received, and the time when the second code string was received, which are stored in association with each other in the storage unit, to obtain an average value of the predetermined number of values obtained by subtracting the time when the first code string was received from the time when the second code string was received for each frame number (step S221-43). The receiving unit 221-m then determines whether the average value obtained in step S221-43 is less than the predetermined first time limit Tmin, is greater than or equal to the predetermined second time limit Tmax, which is greater than the first time limit Tmin, or is greater than or equal to the first time limit Tmin and less than or equal to the second time limit Tmax (step S221-44).
[0225] Then, if the average value is less than the first limit time Tmin in the above judgment, the receiving unit 221-m outputs the monaural code included in the first code string input from the first communication line 410-m and the extended code included in the second code string input from the second communication line 510-m, whose frame number is the same as that of the monaural code, to the decoding device 222-m for subsequent frames. If the average value is greater than the first limit time Tmin and less than the second limit time Tmax in the above judgment, the monaural code included in the first code string input from the first communication line 410-m and the extended code included in the second code string input from the second communication line 510-m, whose frame number is closest to that of the monaural code, to the decoding device 222-m for subsequent frames. If the average value is not less than the second limit time Tmax in the above judgment, the monaural code included in the first code string input from the first communication line 410-m is output to the decoding device 222-m for subsequent frames. If the average value in the above determination is not less than the second time limit Tmax, receiving unit 221-m does not output the extension code for subsequent frames. In other words, if the average value in the above determination is not less than the second time limit Tmax, receiving unit 221-m may simply output the monaural code. The time from the reception of the first code string on the first communication channel to the reception of the second code string on the second communication channel for the same frame is assumed to be the average value obtained in step S221-43. Therefore, receiving unit 221-m must operate so that the time from the reception of the first code string on the first communication channel to the output to decoding device 222-m is equal to or greater than the average value obtained in step S221-43.
[0226] That is, the receiving unit 221-m, with respect to a group consisting of a first code string received from the first communication line 410-m and a second code string received from the second communication line 510-m corresponding to the first code string, when the average value of the difference between the time when the first code string and the second code string are received among a plurality of groups is less than a predetermined limit time Tmin, the receiving unit 221-m, with respect to subsequent frames, replaces the mono code (i.e., the mono code in the order of the frame number) contained in the first code string input from the first communication line 410-m and the mono code input from the second communication line 510-m with the mono code (i.e., the mono code in the order of the frame number) contained in the first code string input from the first communication line 410-m. Among the extension codes included in the second code string, the extension code having the same frame number as the monaural code is output to the decoding device 222-m. When the above-mentioned average value is greater than the first limit time Tmin and less than the second limit time Tmax, with respect to the subsequent frames, when the extension code included in the second code string input from the second communication line 510-m includes an extension code having the same frame number as the monaural code (i.e., the monaural code in the order of the frame numbers) included in the first code string input from the first communication line 410-m, the monaural code and the frame The extended code having the same frame number as the mono code is output to the decoding device 222-m. In the case where the extended code included in the second code string input from the second communication line 510-m does not include an extended code having the same frame number as the mono code included in the first code string input from the first communication line 410-m (i.e., the mono code in the sequence of frame numbers), the extended code included in the first code string input from the first communication line 410-m (i.e., the mono code in the sequence of frame numbers) and the extended code included in the second code string input from the second communication line 510-m are output to the decoding device 222-m. The extended code whose frame number is closest to the mono code (that is, the extended code whose frame number is different from the mono code but closest to the mono code among the extended codes included in the second code string input from the second communication line 510-m) is output to the decoding device 222-m. When the above-mentioned average value is not less than the second limit time Tmax, for subsequent frames, only the mono code included in the first code string input from the first communication line 410-m (that is, the mono code in frame number order) is output to the decoding device 222-m.
[0227] In addition, until the above-mentioned judgment is completed, the receiving unit 221-m may output nothing, or may output the mono code and the extended code to the decoding device 222-m in the same way as the first embodiment, or may output the mono code to the decoding device 222-m without outputting the extended code, or may always output the mono code to the decoding device 222-m in the same way as the fifth embodiment, and output the extended code to the decoding device 222-m only when the difference in frame number between the mono code and the extended code is small.
[0228] The operation of the decoding device 222-m of the sound signal receiving device 220-m of the eighth embodiment is the same as that of the decoding device 222-m of the sound signal receiving device 220-m of the sixth embodiment. However, the spreading code output by the receiving unit 221-m of the eighth embodiment may differ from the spreading code output by the receiving unit 221-m of the sixth embodiment depending on the situation. Therefore, the decoding device 222-m specifically performs the following operation.
[0229] That is, when the average value in the above judgment is less than the first limit time Tmin, and when the average value in the above judgment is greater than the first limit time Tmin and less than the second limit time Tmax, the decoding device 222-m obtains and outputs a decoded digital sound signal of two channels based on the mono code output by the receiving unit 221-m and the extended code output by the receiving unit 221-m for subsequent frames. When the average value in the above judgment is greater than the second limit time Tmax, for subsequent frames, the mono decoded digital sound signal based on the mono code output by the receiving unit 221-m is output as a decoded digital sound signal of two channels as it is.
[0230] More specifically, the decoding device 222-m obtains and outputs decoded digital sound signals of two channels based on the monaural code included in the first code string input from the first communication line 410-m and the extended code with the same frame number as the monaural code included in the second code string input from the second communication line 510-m, when the average value of the difference between the time when the first code string and the second code string are received among a plurality of groups is less than a predetermined first limit time Tmin. When the decoding time Tmin is greater than the predetermined second limit time Tmax, the mono decoded digital sound signal based on the mono code included in the first code string input from the first communication line 410-m is output as a 2-channel decoded digital sound signal. When the above-mentioned average value is greater than the first limit time Tmin and less than the second limit time Tmax, the 2-channel decoded digital sound signal is obtained and output based on the mono code included in the first code string input from the first communication line 410-m and the extended code whose frame number is closest to the mono code included in the second code string input from the second communication line 510-m.
[0231] More specifically, the decoding device 222-m decodes a group consisting of a first code string received from the first communication line 410-m and a second code string received from the second communication line 510-m corresponding to the first code string, based on the mono code (i.e., the mono code in the order of the frame number) contained in the first code string input from the first communication line 410-m and the mono code (i.e., the mono code in the order of the frame number) contained in the second code string input from the second communication line 510-m, when the average value of the difference between the time when the first code string and the second code string are received among a plurality of groups is less than a predetermined first limit time Tmin. The frame number of the mono code is the same as the extended code, and the decoded digital sound signal of the two channels is obtained and output. When the above average value is longer than the predetermined second limit time Tmax which is larger than the first limit time Tmin, the decoded digital sound signal of the mono channel based on the mono code included in the first code string input from the first communication line 410-m (that is, the mono code in the order of the frame number) is output as a decoded digital sound signal of the two channels. When the above average value is longer than the first limit time Tmin and shorter than the second limit time Tmax, the decoded digital sound signal of the mono channel based on the mono code (that is, the mono code in the order of the frame number) is output as a decoded digital sound signal of the two channels. In this case, regarding a frame in which the extended code included in the second code string input from the second communication line 510-m includes an extended code having the same frame number as the mono code (i.e., a mono code in the order of frame numbers) included in the first code string input from the first communication line 410-m, a decoded digital sound signal of two channels is obtained and output based on the mono code and the extended code having the same frame number as the mono code, and regarding a frame in which the extended code included in the second code string input from the second communication line 510-m does not include a frame number that is the same as the mono code included in the first code string input from the first communication line 410-m. The decoded digital sound signals of two channels are obtained and output based on the mono code (i.e., the mono code in the order of frame numbers) contained in the first code string input from the first communication line 410-m and the extended code with a frame number closest to the mono code contained in the second code string input from the second communication line 510-m (i.e., the extended code with a frame number different from the mono code but closest to the mono code among the extended codes contained in the second code string input from the second communication line 510-m).
[0232] Furthermore, until the above-described determination by receiving unit 221-m is completed, nothing is input to decoding unit 222-m, or a monaural code is input instead of an extension code, or both a monaural code and an extension code are input. Decoding unit 222-m obtains a two-channel decoded digital audio signal corresponding to the input monaural code and extension code, or the input monaural code, for each frame, and outputs it to playback unit 223-m.
[0233] <Modification of Eighth Embodiment>
[0234] While the above description relates to the sound signal receiving device 220-m of the eighth embodiment, which is based on the configuration of the sound signal receiving device 220-m of the first embodiment, and its operation, the sound signal receiving device 220-m of the eighth embodiment may be configured and operated based on the sound signal receiving device 220-m of any of the third to fifth embodiments and their variations. Furthermore, in the above example, the specific time range is the time from the start of receiving the first code string until a predetermined number of first code strings are received. However, the specific time range may be set at any time. For example, the specific time range may be the time range starting at a certain time after the start of receiving the first code string, or the specific time range may be each of multiple time ranges starting at each of multiple time points after the start of receiving the first code string.
[0235] 〔Effect〕
[0236] According to the eighth embodiment, when the difference between the time when the first code string of the same frame is received from the first communication line and the time when the second code string is received from the second communication line is large, a large error in signal segmentation between channels of the decoded sound signal is suppressed, and when the above-mentioned difference is small, a high-quality decoded sound signal can be obtained.
[0237] <Ninth embodiment>
[0238] In a multi-site control device (MCU) for conducting a conference call at multiple locations, digital audio signals corresponding to audio signals from two different locations can be treated as two-channel digital audio signals, and the same operation as that of the audio signal transmitting device 210-m in the above-described embodiments can be performed. This method will be described as the ninth embodiment.
[0239] <Multi-location Control Device 600>
[0240] The multi-location control device 600 is as follows Figure 7 As shown, it includes a receiving unit 610, a mono decoding unit 620, a location selection unit 630, a signal analysis unit 640, a mono encoding unit 650, and a transmitting unit 660. In the following, it is described that the multi-location control device 600 is connected to terminal devices at P locations (P is an integer greater than or equal to 3), and transmits location m2 to the multi-line supporting terminal device 200-m1. P The multi-location control device 600 performs the sound signal of at most two locations among the P-1 locations, for example, in each frame of a specific time interval of 20ms. Figure 8 And the processing of steps S610 to S660 illustrated below.
[0241] [Receiving Unit 610]
[0242] The receiving unit 610 receives the data from the multi-line supporting terminal device 200-m. else (else is an integer greater than or equal to 2 and less than or equal to P) P-1 first code strings are output via the first communication channel. Receiving section 610 outputs the monaural codes contained in each of the input P-1 first code strings to monaural decoding section 620 (step S610).
[0243] [Mono decoding unit 620]
[0244] Mono decoding section 620 decodes each of the P-1 mono codes input from receiving section 610 using a specific decoding method, obtains a decoded mono signal as a mono decoded digital audio signal, and outputs it to location selection section 630 (step S620). The specific decoding method is as described in the first embodiment.
[0245] [Location Selection Unit 630]
[0246] Based on a predetermined selection criterion, location selection section 630 selects two decoded mono signals from the P-1 decoded mono signals input from monaural decoding section 620 and outputs them to signal analysis section 640 (step S630). The predetermined selection criterion can be any criterion that allows for selection of decoded mono signals at highly important locations, and location selection section 630 can perform the selection. For example, if the power of the audio signal is used as the selection criterion, location selection section 630 outputs the decoded mono signal with the highest power and the decoded mono signal with the second highest power among the P-1 decoded mono signals input to signal analysis section 640 for each frame.
[0247] [Signal Analysis Unit 640]
[0248] Signal analysis unit 640 obtains a monaural signal, a signal obtained by mixing the two decoded monaural signals, from the two inputted decoded monaural signals and outputs it to monaural encoding unit 650. It also obtains a spreading code representing a characteristic parameter representing a characteristic difference between the two inputted decoded monaural signals and having minimal temporal variation, and outputs it to transmitting unit 660 (step S640). Signal analysis unit 640 performs the same operation as signal analysis unit 2121-m of encoding device 212-m of audio signal transmitting device 210-m of multi-channel terminal device 200-m in the first embodiment. In this ninth embodiment, since the two inputted decoded monaural signals correspond to audio signals emitted from different locations, information representing the intensity difference for each frequency band as shown in the second example is more suitable as the characteristic parameter, rather than information representing the time difference as shown in the first example of signal analysis unit 2121-m. Alternatively, information representing the ratio or difference in power between the two inputted decoded monaural signals may be used as the characteristic parameter.
[0249] [Mono encoding unit 650]
[0250] The monaural encoding unit 650 encodes the input monaural signal using a specific encoding method to obtain a monaural code, and outputs the code to the transmitting unit 660 (step S650). The specific encoding method is as described in the first embodiment.
[0251] [Sending unit 660]
[0252] The sending unit 660 outputs the first code string, which is a code string including the mono code input from the mono encoding unit 650, to the multi-channel supporting terminal device 200-m1 via the first communication line for each frame, and outputs the second code string, which is a code string including the extension code input from the signal analysis unit 640, to the multi-channel supporting terminal device 200-m1 via the second communication line (step S660).
[0253] 〔Effect〕
[0254] By operating the multi-location control device 600 according to the ninth embodiment, the multi-line support terminal device 200-m1 can virtually distribute and play audio signals from two locations to the left and right, making it clear which location the speech was made at or whether the speech was made at different locations.
[0255] <Modification of Ninth Embodiment>
[0256] In the site selection unit 630 of the multi-site control device 600 of the ninth embodiment, since power is used to select two decoded monaural signals, the spreading code can be obtained by the site selection unit 630 rather than the signal analysis unit 640. This method is described as a modified example of the ninth embodiment, focusing on the differences from the ninth embodiment.
[0257] <Multi-location Control Device 600>
[0258] A multi-site control device 600 according to a modified example of the ninth embodiment is as follows: Figure 9 As shown in FIG. 1 , the multi-site control device 600 includes a signal mixing unit 670 instead of the signal analyzing unit 640 included in the ninth embodiment. The multi-site control device 600 performs the signal mixing operation for each frame. Figure 10 The processes of steps S610 to S630, step S670, and steps S650 to S660 are shown as examples. The essential differences from the ninth embodiment are step S630 performed by the location selection unit 630 and step S670 performed by the signal mixing unit 670. Step S660 performed by the transmission unit 660 is the same as the ninth embodiment, except that the spreading code is input from the location selection unit 630 instead of the signal analysis unit 640.
[0259] [Location Selection Unit 630]
[0260] The location selection unit 630 selects the decoded mono signal with the largest power and the decoded mono signal with the second largest power among the P-1 decoded mono signals input from the mono decoding unit 620, and outputs them to the signal analysis unit 640. Furthermore, the ratio or difference between the powers of the two selected decoded mono signals is obtained as a characteristic parameter, and an extended code as a code representing the obtained characteristic parameter is obtained and output to the sending unit 660 (step S630).
[0261] [Signal mixing unit 670]
[0262] Signal mixing section 670 obtains a monaural signal obtained by mixing the two input decoded monaural signals based on the two input decoded monaural signals, and outputs the signal to monaural encoding section 650 (step S670).
[0263] Furthermore, to emphasize the virtual left and right distribution of audio signals from two locations in the multi-channel terminal 200-m1, the location selection unit 630 may obtain information identifying the location with the higher power of the two selected decoded monaural signals as a characteristic parameter, obtain a spreading code representing the obtained characteristic parameter, and output it to the transmission unit 660. In this case, the spreading decoding unit 2222-m1 of the decoding device 222-m1 of the audio signal receiving device 220-m1 of the multi-channel terminal 200-m1 can obtain two-channel decoded digital audio signals so that the audio signals are localized to the predetermined left and right positions for each location. Furthermore, in this case, the signal mixing unit 670 may select the higher power of the two input decoded monaural signals and output it to the monaural encoding unit 650. Alternatively, the signal mixing unit 670 may not be included, and the location selection unit 630 may select and output only the decoded monaural signal with the highest power.
[0264] <Tenth embodiment>
[0265] In the above-described embodiments and variations, for simplicity of explanation, an example is described in which two channels of audio signals from a multi-channel terminal device 200-m are processed. However, the number of channels is not limited to this; it can be two or more. If the number of channels is C (where C is an integer greater than or equal to 2), the above-described embodiments and variations can be implemented with C channels (where C is an integer greater than or equal to 2) instead of two channels.
[0266] For example, the sound receiving unit 211-m of the audio signal transmitting device 210-m of the multi-channel supporting terminal device 200-m may include C microphones and C A / D converters. The encoding device 212-m of the audio signal transmitting device 210-m of the multi-channel supporting terminal device 200-m may generate a monaural code and an extension code based on the C channels of digital audio signals received. Specifically, the encoding device 212-m may encode a signal obtained by mixing the C channels of digital audio signals using a specific first encoding method to generate a monaural code and an extension code including a code representing information representing inter-channel differences between the C channels of digital audio signals received. This information representing inter-channel differences between the C channels of digital audio signals may be, for example, information representing the differences between the digital audio signals of the reference channel and the digital audio signals of the reference channel.
[0267] Furthermore, the decoding device 222-m of the audio signal receiving device 220-m of the multi-channel terminal 200-m can obtain and output C channels of decoded digital audio signals based on the input monaural code and extension code. Specifically, the monaural decoding unit 2221-m of the decoding device 222-m decodes the input monaural code to obtain a monaural decoded digital audio signal. The extension decoding unit 2222-m of the decoding device 222-m treats the monaural decoded digital audio signal as a signal obtained by mixing the C channels of decoded digital audio signals. The feature parameters obtained based on the input extension code are used as information representing the characteristics of the inter-channel differences in the C channels of decoded digital audio signals, thereby obtaining and outputting the C channels of decoded digital audio signals. Furthermore, in this case, the playback unit 223-m of the audio signal receiving device 220-m of the multi-channel terminal 200-m can also include a maximum of C DA conversion units and a maximum of C speakers.
[0268] <Other Implementation Methods>
[0269] <<System that also includes dedicated telephone line terminal equipment>>
[0270] When the telephone system 100 also includes the telephone line dedicated terminal device 300-n, the telephone line dedicated terminal device 300-n performs a well-known operation as follows.
[0271] <Telephone Line Dedicated Terminal Device 300-n>
[0272] The telephone line dedicated terminal device 300-n is, for example, a conventional mobile phone or a conventional smart phone, such as Figure 11 As shown, the device 310-n includes a sound signal transmitting side device and a sound signal receiving side device. The sound signal transmitting side device 310-n includes a sound receiving unit 311-n, an encoding unit 312-n, and a transmitting unit 313-n. The sound signal receiving side device 320-n includes a receiving unit 321-n, a decoding unit 322-n, and a playing unit 323-n. The sound signal transmitting side device 310-n of the telephone line dedicated terminal device 300-n performs Figure 12 The voice signal receiving side device 320-n of the telephone line dedicated terminal device 300-n performs the following processing from step S311 to step S313. Figure 13 And the processing of steps S321 to S323 illustrated below.
[0273] [Sound Signal Transmitting Device 310-n]
[0274] The audio signal transmitting device 310-n obtains a first code string including a monaural code corresponding to a digital audio signal of one channel at a specific time interval of, for example, 20 ms, i.e., at each frame, and outputs it to the first communication line 420-n.
[0275] [[Sound receiving unit 311-n]]
[0276] Sound pickup unit 311-n includes a microphone and an A / D converter. The microphone picks up sounds generated in the spatial domain surrounding the microphone, converts them into analog electrical signals, and outputs them to the A / D converter. The A / D converter converts the input analog electrical signals into digital audio signals, such as PCM signals with a sampling frequency of 8kHz, and outputs them. Specifically, sound pickup unit 311-n outputs a single-channel digital audio signal corresponding to the sound picked up by the microphone to encoding device 312-n (step S311).
[0277] [[Encoding device 312-n]]
[0278] The encoding device 312-n encodes the one-channel digital audio signal input from the sound pickup unit 311-n for each frame using the aforementioned specific encoding method to obtain a monaural code and outputs it to the transmitting unit 313-n (step S312).
[0279] [[Sending unit 313-n]]
[0280] The transmitting unit 313 - n outputs the first code string, which is a code string including the monaural code input from the encoding device 312 - n , to the first communication line 420 - n for each frame (step S313 ).
[0281] [Sound Signal Receiving Device 320-n]
[0282] The audio signal receiving device 320 - n outputs audio based on the monaural code included in the first code string input from the first communication line 420 - n at a specific time interval of, for example, every 20 ms, that is, at each frame.
[0283] [[Receiving unit 321-n]]
[0284] The receiving unit 321 - n outputs the monaural code included in the first code string input from the first communication line 420 - n to the decoding device 322 - n for each frame (step S321 ).
[0285] [[Decoding device 322-n]]
[0286] The mono code output by the receiving unit 321-n is input to the decoding device 322-n on a frame-by-frame basis. The decoding device 322-n decodes the input mono code using the aforementioned specific decoding method on a frame-by-frame basis, obtains a decoded digital audio signal, and outputs it to the playback unit 323-n (step S322).
[0287] [[Playback unit 323-n]]
[0288] The playback unit 323 - n outputs the sound corresponding to the input one decoded digital sound signal (step S323 ).
[0289] The playback unit 323-n includes, for example, one DA conversion unit and one speaker. The DA conversion unit converts the input decoded digital sound signal into an analog electrical signal and outputs it. The speaker produces sound corresponding to the analog electrical signal input from the DA conversion unit. The speaker can also be configured on a stereo headset or stereo headphones. When using the speakers provided by stereo headphones or stereo headphones, that is, two speakers, for example, the playback unit 323-n inputs the electrical signal output by the DA conversion unit to the two speakers, and the two speakers produce sound corresponding to the one decoded digital sound signal (decoded sound signal).
[0290] 〔Effect〕
[0291] The dedicated telephone line terminal device 300-n also uses the same encoding and decoding methods as the multi-line supporting terminal device 200-m. Therefore, while ensuring compatibility in the dedicated telephone line terminal device 300-n so that a decoded audio signal of minimum sound quality can be obtained, the multi-line supporting terminal device 200-m can obtain a high-quality decoded audio signal with a delay time that is substantially the same as that used to obtain a decoded audio signal of minimum sound quality, that is, a delay time that does not cause discomfort during a two-way conversation.
[0292] <<There is also a code method that is neither a monaural code nor an extended code>>
[0293] The audio signal transmitting device 210-m of the multi-channel supporting terminal device 200-m may also obtain and output a code (supplementary code) that is neither the monaural code nor the extended code described above. Specifically, the encoding device 212-m may also obtain and output the supplementary code to the transmitting unit 213-m, and the transmitting unit 213-m may output the supplementary code input from the encoding device 212-m to one of the first communication line 410-m and the second communication line 510-m. The supplementary code is, for example, a code that represents the characteristics of the high-band component of a signal obtained by mixing the input digital audio signals of C channels (C is an integer greater than or equal to 2).
[0294] Similarly, a code (supplementary code) that is neither the monaural code nor the extended code described above may be input to the audio signal receiving device 220-m of the multi-channel supporting terminal device 200-m, and the audio signal receiving device 220-m of the multi-channel supporting terminal device 200-m may further use the supplementary code to obtain and output a decoded audio signal. Specifically, the receiving unit 221-m may output the supplementary code input from one of the first communication line 410-m and the second communication line 510-m to the decoding device 222-m, and the decoding device 222-m may further use the supplementary code input from the receiving unit 221-m to obtain a decoded audio signal.
[0295] <Program and Recording Medium>
[0296] The processing of each unit of the multi-line terminal device 200-m can also be implemented by a computer. In other words, the processing of each step of the encoding method and the decoding method in the multi-line terminal device 200-m can also be executed by a computer. In this case, the processing of each step is described in a program. Then, by executing the program on the computer, the processing of each step is implemented on the computer. Figure 14 This diagram shows an example of a functional configuration of a computer for implementing the above-mentioned processing. This processing can be implemented by having the recording unit 2020 read a program for causing the computer to function as the above-mentioned device and operating the control unit 2010, input unit 2030, output unit 2040, etc.
[0297] Each program describing the processing contents can be recorded in advance on a computer-readable recording medium. The computer-readable recording medium may be any medium such as a magnetic recording device, an optical disc, a magneto-optical recording medium, or a semiconductor memory.
[0298] Furthermore, the processing of each unit may be configured by executing a specific program on a computer, or at least a part of the processing may be realized by hardware.
[0299] It is obvious that other changes can be appropriately made without departing from the scope of the present invention.
Claims
1. A method for receiving and decoding a sound signal, performed by a terminal device connected to a first communication line and a second communication line having a lower priority than the first communication line, comprising: a receiving step of outputting, for each frame, the monaural code included in the first code string input from the first communication line and the spreading code having the same frame number as the monaural code, when the spreading code included in the second code string input from the second communication line includes the same spreading code having a frame number as the monaural code included in the first code string input from the first communication line; outputting, when the spreading codes included in the second code string input from the second communication line do not include a spreading code having the same frame number as the monaural code included in the first code string input from the first communication line, a spreading code having a frame number closest to the monaural code among the spreading codes included in the first code string input from the first communication line and the second code string input from the second communication line; and a decoding step of obtaining and outputting C channels of decoded digital audio signals for each frame based on the monaural code outputted in the receiving step and the spread code outputted in the receiving step, where C is an integer greater than or equal to 2; The spread code is a code expressing a characteristic parameter that expresses a characteristic of a difference between channels of a digital audio signal of C channels and is a parameter that expresses information that depends on the relative positions of a sound source and a microphone in space.
2. The method for receiving and decoding a sound signal according to claim 1, wherein: The characteristic parameter is a parameter representing a time difference between the digital audio signals of the C channels, or a parameter representing an intensity difference in each frequency band of the digital audio signals of the C channels.
3. The sound signal receiving and decoding method according to claim 1 or 2, wherein: The decoding step comprises: a mono decoding step of decoding the mono code outputted in the receiving step to obtain a mono decoded digital sound signal; as well as In the extended decoding step, the mono decoded digital sound signal is regarded as a signal obtained by mixing the decoded digital sound signals of the C channels, and the characteristic parameters obtained based on the extended code output in the receiving step are regarded as information representing the characteristics of the difference between the channels of the decoded digital sound signals of the C channels, thereby obtaining and outputting the decoded digital sound signals of the C channels.
4. A communication method based on the sound signal receiving and decoding method according to any one of claims 1 to 3, and a sound signal encoding and transmitting method performed by a terminal device connected to the first communication line and the second communication line and different from the terminal device, The method for encoding and sending a sound signal comprises: an encoding step of obtaining, for each frame, a monaural code representing a signal obtained by mixing C channels of digital audio signals to be input, and an extended code representing a characteristic parameter representing a characteristic of a difference between channels of the C channels of digital audio signals to be input and representing information dependent on the relative positions of a sound source and a microphone in space, wherein C is an integer greater than or equal to 2; and The sending step outputs the first code string including the monaural code obtained in the encoding step to the first communication channel and outputs the second code string including the extended code obtained in the encoding step to the second communication channel for each frame.
5. The communication method according to claim 4, wherein: The spread code obtained in the encoding step is a code representing an average or weighted average of feature parameters obtained from the digital audio signals of the C channels of the current frame and feature parameters of past frames.
6. A communication method based on the sound signal receiving and decoding method according to any one of claims 1 to 3, and a sound signal encoding and transmitting method performed by a terminal device different from the terminal device and connected to the first communication line and the second communication line, The method for encoding and sending a sound signal comprises: The encoding step obtains a monaural code representing a signal obtained by mixing C channels of digital audio signals to be input, for each frame, where C is an integer greater than or equal to 2. obtaining, for a predetermined frame among a plurality of frames, a spreading code representing a characteristic parameter representing a characteristic of a difference between channels of the input C channels of the digital audio signal and representing information dependent on the relative positions of a sound source and a microphone in space; and a sending step of outputting the first code string including the monaural code obtained in the encoding step to the first communication line for each frame; Regarding the predetermined frame, a second code string including the spread code obtained in the encoding step is output to the second communication channel.
7. A communication method based on the sound signal receiving and decoding method according to any one of claims 1 to 3, and a sound signal encoding and transmitting method performed by a terminal device connected to the first communication line and the second communication line and different from the terminal device, The method for encoding and sending a sound signal comprises: The encoding step obtains a monaural code representing a signal obtained by mixing C channels of digital audio signals to be input, for each frame, where C is an integer greater than or equal to 2. For each frame, a feature parameter is obtained, which is a parameter that expresses the characteristics of the difference between the channels of the digital sound signal of the C channels input and is a parameter that expresses information that depends on the relative position of the sound source and the microphone in space, For a predetermined frame among a plurality of frames, obtaining a spreading code representing an average or weighted average of the characteristic parameters; and a sending step of outputting the first code string including the monaural code obtained in the encoding step to the first communication line for each frame; Regarding the predetermined frame, a second code string including the spread code obtained in the encoding step is output to the second communication channel.
8. A sound signal receiving device, included in a terminal device connected to a first communication line and a second communication line having a lower priority than the first communication line, comprising: The receiving unit outputs, for each frame, the monaural code included in the first code string input from the first communication line and the monaural code having the same frame number as the monaural code included in the first code string input from the first communication line, when the spreading code included in the second code string input from the second communication line includes a spreading code having the same frame number as the monaural code included in the first code string input from the first communication line. outputting, when the spreading codes included in the second code string input from the second communication line do not include a spreading code having the same frame number as the monaural code included in the first code string input from the first communication line, a spreading code having a frame number closest to the monaural code among the spreading codes included in the first code string input from the first communication line and the second code string input from the second communication line; and A decoding device, for each frame, obtains and outputs C channels of decoded digital audio signals based on the monaural code output by the receiving unit and the extended code output by the receiving unit, where C is an integer greater than or equal to 2. The spread code is a code expressing a characteristic parameter that expresses a characteristic of a difference between channels of a digital audio signal of C channels and is a parameter that expresses information that depends on the relative positions of a sound source and a microphone in space.
9. The sound signal receiving device according to claim 8, wherein: The characteristic parameter is a parameter representing a time difference between the digital audio signals of the C channels, or a parameter representing an intensity difference in each frequency band of the digital audio signals of the C channels.
10. The sound signal receiving device according to claim 8 or 9, wherein: The decoding device comprises: a mono decoding unit, which decodes the mono code output by the receiving unit to obtain a mono decoded digital sound signal; and An extended decoding unit regards the mono decoded digital sound signal as a signal obtained by mixing the decoded digital sound signals of the C channels, regards the characteristic parameters obtained based on the extended code output by the receiving unit as information representing the characteristics of the difference between the channels of the decoded digital sound signals of the C channels, obtains the decoded digital sound signals of the C channels, and outputs them.
11. A telephone system comprising the voice signal receiving device according to any one of claims 8 to 10, and a voice signal transmitting device included in a terminal device different from the terminal device and connected to a first communication line and the second communication line, The sound signal sending side device includes: an encoding unit that obtains, for each frame, a monaural code representing a signal obtained by mixing C channels of input digital audio signals, and an extended code representing a characteristic parameter representing a characteristic of a difference between channels of the C channels of input digital audio signals and representing information dependent on the relative positions of a sound source and a microphone in space, wherein C is an integer greater than or equal to 2; The transmitting unit outputs a first code string including the monaural code obtained by the encoding unit to the first communication channel and outputs a second code string including the spread code obtained by the encoding unit to the second communication channel for each frame.
12. The telephone system according to claim 11, wherein The spread code obtained by the encoding unit is a code representing an average or weighted average of feature parameters obtained from the digital audio signals of the C channels of the current frame and feature parameters of the previous frame.
13. A telephone system comprising the voice signal receiving device according to any one of claims 8 to 10, and a voice signal transmitting device included in a terminal device different from the terminal device and connected to a first communication line and the second communication line, The sound signal sending side device includes: The encoding unit obtains a monaural code representing a signal obtained by mixing C-channel digital audio signals inputted for each frame, where C is an integer greater than or equal to 2. obtaining, for a predetermined frame among a plurality of frames, a spreading code representing a characteristic parameter representing a characteristic of a difference between channels of the input C channels of the digital audio signal and representing information dependent on the relative positions of a sound source and a microphone in space; and a transmitting unit that outputs the first code string including the mono code obtained by the encoding unit to the first communication line for each frame, Regarding the predetermined frame, a second code string including the spread code obtained by the encoding unit is output to the second communication channel.
14. A telephone system comprising the voice signal receiving device according to any one of claims 8 to 10, and a voice signal transmitting device included in a terminal device different from the terminal device and connected to a first communication line and the second communication line, The sound signal sending side device includes: The encoding unit obtains a monaural code representing a signal obtained by mixing C-channel digital audio signals inputted for each frame, where C is an integer greater than or equal to 2. For each frame, a feature parameter is obtained, which is a parameter that expresses the characteristics of the difference between the channels of the digital sound signal of the C channels input and is a parameter that expresses information that depends on the relative position of the sound source and the microphone in space, For a predetermined frame among a plurality of frames, obtaining a spreading code representing an average or weighted average of the characteristic parameters; and a transmitting unit that outputs the first code string including the mono code obtained by the encoding unit to the first communication line for each frame, Regarding the predetermined frame, a second code string including the spread code obtained by the encoding unit is output to the second communication channel.
15. A computer program product comprising a program for causing a computer to execute the sound signal receiving and decoding method according to any one of claims 1 to 3.
16. A computer program product comprising a program for causing a computer to execute the communication method according to claim 6 or 7.
17. A computer-readable recording medium recording a program for causing a computer to execute the sound signal reception and decoding method according to any one of claims 1 to 3.
18. A computer-readable recording medium recording a program for causing a computer to execute the communication method according to claim 6 or 7.
Citation Information
Patent Citations
Audio signal packet communication method, audio signal packet transmission method and reception method, apparatus for the same, program thereof and recording medium
JP2005117132A