Sound signal receiving and decoding method, sound signal decoding method, sound signal receiving-side device, decoding device, computer program product, and recording medium
By receiving and decoding code strings from communication lines of different priority levels, using time and frequency band characteristic parameters, the contradiction between high-quality decoding sound signals and delay time in the prior art is solved, and high-quality decoding is achieved without adding significant delay.
Patent Information
- Application Number
- CN201980097320.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-13
- Filing Date
- 2019-12-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2039-12-27
Smart Images

Figure CN113966530B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to at least one of a decoding technique for a voice signal and a corresponding encoding technique for a voice signal in a terminal device connected to at least two communication networks having different priorities for information transmission. Background Art
[0002] As a prior art for encoding and decoding a voice signal between terminal devices connected to two communication networks having different priorities for information transmission, there is the technique of Patent Document 1. The encoding device of Patent Document 1 performs scalable encoding on the input voice signal for each specific time interval, i.e., each frame, to obtain a low-band code 1 as a code of a base layer, a low-band code 2 as a code of an extended layer, and a high-band code. The low-band code 1 is included in a packet with a high priority and sent to at least Network B where the bandwidth is guaranteed, and the low-band code 2 and the high-band code are included in a packet with a low priority and sent to Network A where the bandwidth is not guaranteed. The decoding device of Patent Document 1 starts monitoring the elapsed time of a limit when receiving a packet with a high priority, and if the limit time has elapsed, it decodes using the received packet at that moment. That is, if it is assumed that the delay of Network A is generally larger than that of Network B, the decoding device of Patent Document 1 substantially decodes using the low-band code 2 and the high-band code to obtain a high-quality decoded voice signal if the low-band code 2 and the high-band code also arrive after the above limit time has elapsed since the arrival of the code of the base layer, and decodes using only the low-band code 1 to obtain a decoded voice signal with the minimum required quality if the low-band code 2 and the high-band code do not arrive.
[0003] Prior Art Documents
[0004] Patent Documents
[0005] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2005-117132 Summary of the Invention
[0006] Problems to be Solved by the Invention
[0007] In the technology of Patent Document 1, in order to obtain a decoded audio signal with high sound quality in most frames, it is necessary to set a time much longer than the delay time generated in the structure for obtaining a decoded audio signal with only the minimum required sound quality as the above-mentioned limit time. Therefore, there is a problem in the technology of Patent Document 1: If you want to obtain a decoded audio signal with high sound quality in most frames, you have to set the above-mentioned limit time so long that it causes a sense of disharmony during a two-way call. In addition, in the technology of Patent Document 1, if the limit time is made close to 0 so as not to cause a sense of disharmony during a two-way call, the proportion of frames in which packets with high priority arrive within the limit time is very small. Therefore, there is a problem in the technology of Patent Document 1: If the limit time is set so as not to cause a sense of disharmony during a two-way call, it is almost impossible to obtain a decoded audio signal with high sound quality in all frames.
[0008] Therefore, in the present invention, an object is to provide a technology that can obtain a decoded audio signal with high sound quality without significantly increasing the delay time compared to a structure for obtaining a decoded audio signal with only the minimum required sound quality.
[0009] Means for Solving the Problem
[0010] One aspect of the present invention is an audio signal reception decoding method performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, including: a reception step of determining whether an average value of differences in the reception times of a first code string received from the first communication line and a second code string received from the second communication line corresponding to the first code string among a plurality of groups is less than a predetermined limit time Tmin. When the average value is less than the limit time Tmin in the determination, for frames after the determination, output a mono code included in the first code string input from the first communication line and an extended code included in the second code string input from the second communication line, where the frame number of the extended code is the same as that of the mono code. When the average value is not less than the limit time Tmin in the determination, for frames after the determination, output a mono code included in the first code string input from the first communication line and an extended code included in the second code string input from the second communication line, where the frame number of the extended code is the closest to that of the mono code; and a decoding step of obtaining and outputting a decoded digital audio signal of C (C is an integer of 2 or more) channels based on the mono code output in the reception step and the extended code output in the reception step for frames after the determination.
[0011] One aspect of the present invention is a method for receiving and decoding a sound signal, which is performed by a terminal device connected to a first communication line and a second communication line having a lower priority than the first communication line, and includes: a receiving step of determining, for a group formed by a first code string received from the first communication line and a second code string received from the second communication line corresponding to the first code string, whether an average value of a difference in reception times of the first code string and the second code string among a plurality of groups is less than a predetermined limit time Tmax, and in a case where the average value is less than the limit time Tmax in the determination, for frames after the determination, outputting an extended code in which a frame number is closest to a mono code included in the first code string input from the first communication line and a mono code included in the first code string input from the first communication line and an extended code included in the second code string input from the second communication line, and in a case where the average value is not less than the limit time Tmax in the determination, for frames after the determination, outputting the mono code included in the first code string input from the first communication line; and a decoding step of, in a case where the average value is less than the limit time Tmax in the determination, for frames after the determination, obtaining and outputting a decoded digital sound signal of C (C is an integer of 2 or more) channels based on the mono code output in the receiving step and the extended code output in the receiving step, and in a case where the average value is not less than the limit time Tmax in the determination, for frames after the determination, outputting the decoded digital sound signal based on the mono code output in the receiving step as the decoded digital sound signal of C channels.
[0012] One embodiment of the present invention is a method for receiving and decoding a voice signal, which is performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, and includes: a receiving step of determining, for a group consisting of a first code string received from the first communication line and a second code string received from the second communication line corresponding to the first code string, whether the difference in the reception times of the first code string and the second code string is less than a predetermined first limit time Tmin, greater than or equal to a predetermined second limit time Tmax that is greater than the first limit time Tmin, or greater than or equal to the first limit time Tmin and less than the second limit time Tmax. In the case where the average value is less than the first limit time Tmin in the determination, for the frames after the determination, an extended code with the same frame number as the mono code included in the first code string input from the first communication line and the extended code included in the second code string input from the second communication line is output. In the case where the average value is greater than or equal to the first limit time Tmin and less than the second limit time Tmax in the determination, for the frames after the determination, an extended code with the frame number closest to the mono code among the mono code included in the first code string input from the first communication line and the extended code included in the second code string input from the second communication line is output. In the case where the average value is greater than or equal to the second limit time Tmax in the determination, for the frames after the determination, the mono code included in the first code string input from the first communication line is output; and a decoding step of, in the case where the average value is less than the first limit time Tmin in the determination and in the case where the average value is greater than or equal to the first limit time Tmin and less than the second limit time Tmax in the determination, for the frames after the determination, obtaining a decoded digital voice signal of C (C is an integer greater than or equal to 2) channels based on the mono code output in the receiving step and the extended code output in the receiving step and outputting it. In the case where the average value is greater than or equal to the second limit time Tmax in the determination, for the frames after the determination, the decoded digital voice signal based on the mono code output in the receiving step is output as the decoded digital voice signal of C channels.
[0013] One aspect of the present invention is a method for receiving and decoding a sound signal, which is performed by a terminal device connected to a first communication line and a second communication line having a lower priority than the first communication line, and includes: a receiving step of outputting, for a frame in which a difference between a frame number of a mono code included in a first code string input from the first communication line and an extended code having a frame number closest to the mono code among extended codes included in a second code string input from the second communication line is less than a predetermined value, the mono code included in the first code string input from the first communication line and the extended code having a frame number closest to the mono code among the extended codes included in the second code string input from the second communication line, and outputting the mono code included in the first code string input from the first communication line for a frame in which the difference is not less than the predetermined value; and a decoding step of obtaining and outputting a decoded digital sound signal of C (C is an integer of 2 or more) channels based on the mono code output in the receiving step and the extended code output in the receiving step for a frame in which the difference is less than the predetermined value, and outputting a decoded digital sound signal based on the mono code output in the receiving step as a decoded digital sound signal of C channels for a frame in which the difference is not less than the predetermined value.
[0014] One aspect of the present invention is a method for decoding a sound signal, which is performed by a terminal device connected to a first communication line and a second communication line having a lower priority than the first communication line, and includes: a decoding step of obtaining and outputting a decoded digital sound signal of C (C is an integer of 2 or more) channels based on a mono code included in a first code string received from the first communication line and an extended code having the same frame number as the mono code included in a second code string received from the second communication line corresponding to the first code string when a difference in time at which the first code string and the second code string are received is less than an average value among a plurality of groups of a predetermined limit time Tmin, and obtaining and outputting a decoded digital sound signal of C channels based on the mono code included in the first code string received from the first communication line and an extended code having a frame number closest to the mono code among the extended codes included in the second code string received from the second communication line when the average value is not less than the limit time Tmin.
[0015] One aspect of the present invention is a method for decoding an audio signal, which is performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, and includes: a decoding step, for a group consisting of a first code string received from the first communication line and a second code string received from the second communication line corresponding to the first code string, when the average value of the time differences between the times when the first code string and the second code string are received among multiple groups is less than a predetermined limit time Tmax, based on the mono code included in the first code string input from the first communication line and the extended code with the frame number closest to the mono code included in the second code string input from the second communication line, obtain a decoded digital audio signal of C (C is an integer greater than or equal to 2) channels and output it. When the average value is not less than the limit time Tmax, output the decoded digital audio signal based on the mono code included in the first code string input from the first communication line as the decoded digital audio signal of C channels.
[0016] One aspect of the present invention is a method for decoding an audio signal, which is performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, and includes: a decoding step, for a group consisting of a first code string received from the first communication line and a second code string received from the second communication line corresponding to the first code string, when the average value of the time differences between the times when the first code string and the second code string are received among multiple groups is less than a predetermined first limit time Tmin, based on the mono code included in the first code string input from the first communication line and the extended code with the same frame number as the mono code included in the second code string input from the second communication line, obtain a decoded digital audio signal of C (C is an integer greater than or equal to 2) channels and output it. When the average value is greater than or equal to a predetermined second limit time Tmax greater than the first limit time Tmin, output the decoded digital audio signal based on the mono code included in the first code string input from the first communication line as the decoded digital audio signal of C channels. When the average value is greater than or equal to the first limit time Tmin and less than the second limit time Tmax, based on the mono code included in the first code string input from the first communication line and the extended code with the frame number closest to the mono code included in the second code string input from the second communication line, obtain a decoded digital audio signal of C channels and output it.
[0017] One aspect of the present invention is a method for decoding an audio signal, which is performed by a terminal device connected to a first communication line and a second communication line having a lower priority than the first communication line, and includes: a decoding step, for frames in which the difference in frame numbers between the mono code included in the first code string input from the first communication line and the extended code with the frame number closest to the mono code included in the second code string input from the second communication line is less than a predetermined value, obtaining a decoded digital audio signal for C (C is an integer of 2 or more) channels based on the mono code and the extended code and outputting it; for frames in which the difference is not less than the predetermined value, outputting the decoded digital audio signal based on the mono code as the decoded digital audio signal for C channels.
[0018] Effect of the Invention
[0019] According to the present invention, compared with a structure that only obtains a decoded audio signal with the minimum required sound quality, a decoded audio signal with high sound quality can be obtained without significantly increasing the delay time. Description of the Drawings
[0020] Figure 1 It is a block diagram showing an example of a telephone system.
[0021] Figure 2 It is a block diagram showing an example of a multi-line support terminal device.
[0022] Figure 3 It is a flowchart showing an example of the processing of the audio signal transmission side device of a multi-line support terminal device.
[0023] Figure 4 It is a flowchart showing an example of the processing of the audio signal reception side device of a multi-line support terminal device.
[0024] Figure 5 It is a diagram schematically showing the temporal relationship between the input code and the output signal in the audio signal reception side device of a multi-line support terminal device.
[0025] Figure 6 It is a diagram schematically showing the temporal relationship between the input code and the output signal in the audio signal reception side device using the prior art.
[0026] Figure 7 It is a block diagram showing an example of a multi-site control device.
[0027] Figure 8 It is a flowchart showing an example of the processing of a multi-site control device.
[0028] Figure 9 It is a block diagram showing an example of a multi-site control device.
[0029] Figure 10 It is a flowchart showing a processing example of a multi-site control device.
[0030] Figure 11 It is a block diagram showing an example of a dedicated telephone line terminal device.
[0031] Figure 12 It is a flowchart showing a processing example of the voice signal transmission side device of the dedicated telephone line terminal device.
[0032] Figure 13 It is a flowchart showing a processing example of the voice signal reception side device of the dedicated telephone line terminal device.
[0033] Figure 14 It is a diagram showing an example of the functional structure of a computer for implementing each device in the embodiment of the present invention. Detailed Embodiment
[0034] <Telephone System 100>
[0035] As shown in Figure 1 the telephone system 100 includes multi-line support terminal devices 200-m (m is each integer from 1 or more to M or less, and M is an integer of 2 or more), a first communication network 400, and a second communication network 500. The telephone system 100 may also include, as shown by the dotted line in Figure 1 dedicated telephone line terminal devices 300-n (n is each integer from 1 or more to N or less, and N is an integer of 1 or more). Each multi-line support terminal device 200-m can be connected to other terminal devices via a first communication line 410-m that is each communication line of the first communication network 400. Further, each multi-line support terminal device 200-m can be connected to other multi-line support terminal devices via a second communication line 510-m that is each communication line of the second communication network 500. Each dedicated telephone line terminal device 300-n can be connected to other terminal devices via a first communication line 420-n that is each communication line of the first communication network 400.
[0036] <First Communication Network 400, Second Communication Network 500>
[0037] The first communication network 400 and the second communication network 500 are communication networks with different information transmission priorities. The first communication network 400 is a communication network with a higher information transmission priority compared to the second communication network 500, and it is a communication network capable of transmitting a code string of a specific bit rate from a certain terminal device to another terminal device with a shorter delay time. The first communication network 400 is, for example, a communication network for two-way calls between a terminal device such as an existing mobile phone or a smart phone and another terminal device such as an existing mobile phone or a smart phone, and it is a communication network equipped with a communication line generally called a telephone line. The second communication network 500 is a communication network with a lower information transmission priority compared to the first communication network 400, and it is a communication network capable of transmitting a code string from a certain terminal device to another terminal device without setting a constraint on the delay time. The second communication network 500 is, for example, a communication network used when transmitting data such as images or character strings from a terminal device such as a smart phone to another terminal device such as a smart phone, and it is a communication network equipped with a communication line generally called an Internet line.
[0038] It is described in Figure 1 that there are the first communication network 400 and the second communication network 500, but the first communication network 400 and the second communication network 500 do not need to be physically separated, as long as they are logically separated. Similarly, when a terminal device is connected to both the first communication line 410-m and the second communication line 510-m, the first communication line 410-m and the second communication line 510-m do not need to be physically separated, as long as they are logically separated. That is, each terminal device can also be connected to an IP communication network through an IP communication line, and through packet priority control, etc., logically construct the first communication network 400 and the first communication line 410-m as a communication network and communication line with a high information transmission priority, and the second communication network 500 and the second communication line 510-m as a communication network and communication line with a lower information transmission priority compared to the first communication network 400 and the first communication line 410-m. For example, it can also be that the multi-line support terminal device 200-m is a smart phone that supports VoLTE (Voice over LTE), and examples of the first communication network 400 and the first communication line 410-m are the VoLTE communication network and the VoLTE line in the LTE communication network and the LTE line, and examples of the second communication network 500 and the second communication line 510-m are the Internet communication network and the Internet line in the LTE communication network and the LTE line.
[0039] In addition, all of the above examples of communication networks, communication lines, and terminal devices are examples of mobile communication. However, it is not limited whether each communication network is a communication network for fixed communication or mobile communication, whether each communication line is wired or wireless, and whether each terminal device is a landline phone or a mobile phone, etc.
[0040] <First Embodiment>
[0041] A multi-line support terminal device according to the first embodiment will be described.
[0042] <Multi-line support terminal device 200-m>
[0043] The multi-line support terminal device 200-m is, for example, a smartphone that supports VoLTE. As Figure 2 shown, it includes a voice signal transmission-side device 210-m and a voice signal reception-side device 220-m. The voice signal transmission-side device 210-m includes a sound collection unit 211-m, an encoding device 212-m, and a transmission unit 213-m. The voice signal reception-side device 220-m includes a reception unit 221-m, a decoding device 222-m, and a playback unit 223-m. The encoding device 212-m includes a signal analysis unit 2121-m and a mono encoding unit 2122-m. The decoding device 222-m includes a mono decoding unit 2221-m and an expansion decoding unit 2222-m. In addition, as shown by the dotted line in the figure, the signal analysis unit 2121-m and the mono encoding unit 2122-m are collectively referred to as an encoding unit 2129-m, and the mono decoding unit 2221-m and the expansion decoding unit 2222-m are collectively referred to as a decoding unit 2229-m. In addition, the encoding device 212-m and the decoding device 222-m may sometimes be referred to as a voice signal encoding device 212-m and a voice signal decoding device 222-m, respectively. The voice signal transmission-side device 210-m of the multi-line support terminal device 200-m performs Figure 3 and the processing of steps S211 to S213 exemplified below, and the voice signal reception-side device 220-m of the multi-line support terminal device 200-m performs Figure 4 and the processing of steps S221 to S223 exemplified below.
[0044] [Voice signal transmission-side device 210-m]
[0045] The voice signal transmission-side device 210-m, for example, for each specific time interval of 20 ms, that is, for each frame, obtains a code string, that is, a first code string, including a mono code corresponding to a digital voice signal of two channels and outputs it to the first communication line 410-m, and obtains a code string, that is, a second code string, including an expansion code corresponding to the digital voice signal of the two channels and outputs it to the second communication line 510-m.
[0046] [[Sound receiving unit 211-m]]
[0047] The sound receiving unit 211-m includes two microphones and two AD conversion units. Each microphone is associated with each AD conversion unit on a one-to-one basis. The microphone picks up the sound generated in the spatial domain around the microphone, converts it into an analog electrical signal, and outputs it to the AD conversion unit. The AD conversion unit converts the input analog electrical signal into a digital sound signal, such as a PCM signal with a sampling frequency of 8 kHz, and outputs it. That is, the sound receiving unit 211-m outputs two-channel digital sound signals corresponding to the sounds picked up by the two microphones, for example, two-channel stereo digital sound signals of the left channel and the right channel, to the encoding device 212-m (step S211).
[0048] In addition, all or part of the sound receiving unit 211-m may not be configured inside the sound signal transmitting side device 210-m, but may be connected to the sound signal transmitting side device 210-m. For example, the sound receiving unit 211-m of the sound signal transmitting side device 210-m may not have a microphone, and two analog electrical signals may be input from a microphone connected to the sound signal transmitting side device 210-m to the AD conversion unit of the sound receiving unit 211-m of the sound signal transmitting side device 210-m. Or, the sound signal transmitting side device 210-m may not have the sound receiving unit 211-m, and two-channel digital sound signals may be input from a sound receiving device such as an AD converter connected to the sound signal transmitting side device 210-m to the encoding device 212-m of the sound signal transmitting side device 210-m.
[0049] [[Encoding device 212-m]]
[0050] Two-channel digital sound signals are input to the encoding device 212-m from the sound receiving unit 211-m or a sound receiving device connected to the sound signal transmitting side device 210-m. The encoding device 212-m obtains a monaural code and an extended code corresponding to the input two-channel digital sound signals for each frame, and outputs them to the transmitting unit 213-m (step S212).
[0051] [[[Signal analysis unit 2121-m]]]
[0052] The signal analysis unit 2121-m obtains, for each frame, a monaural signal that is a signal obtained by mixing the two input digital audio signals of two channels, and an extended code representing characteristic parameters. The characteristic parameters are parameters that represent the difference between the two input digital audio signals of two channels and have little temporal variation. The signal analysis unit 2121-m outputs the obtained monaural signal to the monaural encoding unit 2122-m and outputs the obtained extended code to the transmission unit 213-m. The parameter with little temporal variation is a parameter with low dependence on time and low time resolution.
[0053] 〔First example of the signal analysis unit 2121-m〕
[0054] As a first example, the operation per frame of the signal analysis unit 2121-m in the case where the information representing the time difference between the two input digital audio signals of two channels is used as the characteristic parameter will be described. The signal analysis unit 2121-m first obtains the characteristic parameter that is the information representing the time difference between the two input digital audio signals of two channels (step S2121-11). The time difference between the two input digital audio signals of two channels can be obtained by any well-known method. For example, the signal analysis unit 2121-m calculates the correlation value between the sample sequence of the digital audio signal of one channel (the first channel) and the sample sequence obtained by advancing the sample sequence of the digital audio signal of the other channel (the second channel) by the number of candidate samples for each time difference within a predetermined range, and obtains the number of time difference samples that is the candidate sample with the largest correlation value as the characteristic parameter.
[0055] Next, the signal analysis unit 2121-m obtains, for the sample sequence of the digital audio signal of the first channel and the sample sequence obtained by giving the time difference represented by the characteristic parameter to the sample sequence of the digital audio signal of the second channel, one of the sequence obtained by adding the corresponding samples, the sequence of the average values of the corresponding samples, and the sequence obtained by transforming the above-mentioned added or average value sequence, as the monaural signal that is the signal obtained by mixing the two-channel digital audio signals (step S2121-12). The sample sequence obtained by giving the time difference represented by the characteristic parameter to the sample sequence of the digital audio signal of the second channel is, for example, the sample sequence obtained by advancing the sample sequence of the digital audio signal of the second channel by the number of time difference samples represented by the characteristic parameter.
[0056] The signal analysis unit 2121-m then obtains an extended code of the code serving as the performance characteristic parameter (step S2121-13). The extended code of the code serving as the performance characteristic parameter can be obtained by a well-known method. For example, the signal analysis unit 2121-m performs scalar quantization on the number of time difference samples of the input digital audio signals of two channels to obtain a code, and outputs the obtained code as the extended code. Or, for example, the signal analysis unit 2121-m outputs the binary number representing the number of time difference samples of the input digital audio signals of two channels itself as the extended code.
[0057] 〔Second example of the signal analysis unit 2121-m〕
[0058] As a second example, the operation of the signal analysis unit 2121-m for each frame in the case where information representing the intensity difference of each frequency band of the input digital audio signals of two channels is used as the characteristic parameter will be described. In addition, in the following, a specific example using a complex DFT (Discrete Fourier Transformation) will be described, but a well-known frequency domain transformation method other than the complex DFT can also be used.
[0059] The signal analysis unit 2121-m first performs an inverse DFT on the input digital audio signals of two channels respectively to obtain an inverse DFT coefficient sequence (step S2121-21). The inverse DFT coefficient sequence can also be obtained by well-known methods such as using a process of applying a window with overlap between frames and a process considering the symmetry of the complex numbers obtained by the inverse DFT. For example, if a frame consists of 128 samples, then an inverse DFT is performed on a sequence of 256 samples of the digital audio signal including the last 64 samples of the immediately preceding frame and the first 64 samples of the immediately following frame, and the first 128 complex number sequence among the obtained 256 complex number sequences is used as the inverse DFT coefficient sequence. Hereafter, let f be each integer from 1 to 128 and above, let each inverse DFT coefficient of the inverse DFT coefficient sequence of the first channel be V1(f), and let each inverse DFT coefficient of the inverse DFT coefficient sequence of the second channel be V2(f). The signal analysis unit 2121-m then obtains a sequence of values of the radius of each inverse DFT coefficient in the complex plane based on the inverse DFT coefficient sequences of the two channels (step S2121-22). The value of the radius of each inverse DFT coefficient of each channel in the complex plane corresponds to the intensity of each frequency bin of the digital audio signal of each channel. Hereafter, let the value of the radius of the inverse DFT coefficient V1(f) of the first channel in the complex plane be V1r(f), and let the value of the radius of the inverse DFT coefficient V2(f) of the second channel in the complex plane be V2r(f). The signal analysis unit 2121-m then obtains the average value of the ratio of the radius value of one channel to the radius value of the other channel for each frequency band, and obtains the sequence of average values as a characteristic parameter (step S2121-23). This sequence of average values is a characteristic parameter corresponding to the information representing the intensity difference of each frequency band of the input digital audio signals of the two channels. For example, if it is set to four bands, for the four bands where f ranges from 1 to 32, from 33 to 64, from 65 to 96, and from 97 to 128, the average values Mr(1), Mr(2), Mr(3), Mr(4) of 32 values obtained by dividing the radius value V1r(f) of the first channel by the radius value V2r(f) of the second channel are respectively obtained, and the sequence of average values {Mr(1), Mr(2), Mr(3), Mr(4)} is obtained as a characteristic parameter.
[0060] In addition, the number of bands only needs to be a value less than or equal to the number of frequency bins. That is, the same value as the number of frequency bins can be used as the number of bands, or 1 can be used as the number of bands. When using the same value as the number of frequency bins as the number of bands, the signal analysis unit 2121-m obtains the ratio of the radius value of one channel to the radius value of the other channel for each frequency bin, and the sequence of the obtained ratio values can be used as the feature parameter. When using 1 as the number of bands, the signal analysis unit 2121-m obtains the ratio of the radius value of one channel to the radius value of the other channel for each frequency bin, and the average value of the entire band of the obtained ratio values can be used as the feature parameter. In addition, when the number of bands is set to multiple, the number of frequency bins included in each frequency band is arbitrary. For example, the number of frequency bins included in the lower-frequency band can be less than the number of frequency bins included in the higher-frequency band.
[0061] In addition, the signal analysis unit 2121-m can also use the difference between the radius value of one channel and the radius value of the other channel instead of the ratio of the radius value of one channel to the radius value of the other channel. That is, in the case of the above example, the value obtained by subtracting the radius value V2r(f) of the second channel from the radius value V1r(f) of the first channel can be used instead of the value obtained by dividing the radius value V1r(f) of the first channel by the radius value V2r(f) of the second channel.
[0062] In addition, the signal analysis unit 2121-m obtains, for the sample sequence of the digital audio signal of the first channel and the sample sequence of the digital audio signal of the second channel, one of the sequence of the sum of the corresponding samples, the sequence of the average value of the corresponding samples, and the sequence obtained by transforming the above sum or average value sequence, as the mono signal (step S2121-24) which is the signal obtained by mixing the digital audio signals of the two channels. In addition, the signal analysis unit 2121-m can also obtain the average value VMr(f) of the radius and the average value VMθ(f) of the angle of each complex DFT coefficient V1(f) of the complex DFT coefficient sequence of the first channel and each complex DFT coefficient V2(f) of the complex DFT coefficient sequence of the second channel obtained in step S2121-21, and perform an inverse complex DFT on the sequence of the complex number VM(f) with a radius of VMr(f) and an angle of VMθ(f) in the complex plane to obtain the mono signal (step S2121-24’) which is the signal obtained by mixing the digital audio signals of the two channels.
[0063] The signal analysis unit 2121-m then obtains an extended code of the code that is the performance feature parameter (step S2121-25). The extended code of the code that is the performance feature parameter can be obtained by using a well-known method. For example, the signal analysis unit 2121-m performs vector quantization on the sequence of values obtained in step S2121-23 to obtain a code, and outputs the obtained code as the extended code. Or, for example, the signal analysis unit 2121-m performs scalar quantization on each value included in the sequence of values obtained in step S2121-23 to obtain codes, and combines the obtained codes to output them as the extended code. In addition, when the value obtained by the signal analysis unit 2121-m in step S2121-23 is a single value, it is sufficient to output the code obtained by performing scalar quantization on the single value as the extended code.
[0064] The time difference between the two-channel digital audio signals input as described in the first example of the signal analysis unit 2121-m, or the intensity difference of each frequency band of the two-channel digital audio signals input as described in the second example of the signal analysis unit 2121-m, depends on the position of the sound source. For general sound sources such as people or musical instruments, the position of the sound source changes little over time. Even when the position of the sound source changes over time, as long as the sound source does not move rapidly, the time difference or the intensity difference of each frequency band of the two-channel digital audio signals input does not change much.
[0065] Accordingly, the signal analysis unit 2121-m can also obtain the average or weighted average of the feature parameters obtained from the two-channel digital audio signals input for each of a plurality of consecutive frames including the frame to be processed as the feature parameter, and output the extended code representing the obtained feature parameter. The weight for weighted average may be set to the maximum value for the frame to be processed, and set to a smaller value for frames farther away from the frame to be processed. In addition, if the feature parameters of frames more in the future than the frame to be processed are used, pre-reading is required and the delay increases. Therefore, it is better for the signal analysis unit 2121-m to use a plurality of consecutive frames on the past side including the frame to be processed. In addition, of course, when the feature parameter includes a plurality of elements such as information representing the intensity difference of each of a plurality of frequency bands, the average or weighted average of the feature parameters is a numerical sequence with the average value or weighted average value of each element of the feature parameter as an element.
[0066] In addition, for example, for the difference in the waveforms of the two-channel digital audio signals being input, that is, the sample sequence of the differences between the corresponding samples of the two-channel digital audio signals being input, even if the time of each sample is shifted by one sample, it will become a sample sequence completely different from the difference in the waveforms of the two-channel digital audio signals being input. Therefore, it is information with a high dependence on time, information with a high time resolution, and information with large temporal variations. Similarly, the phase difference between the two-channel digital audio signals being input, for example, the difference in the angles in the complex plane of each complex DFT coefficient V1(f) of the complex DFT coefficient sequence of the first channel obtained in step S2121-21 and the angles in the complex plane of each complex DFT coefficient V2(f) of the complex DFT coefficient sequence of the second channel, is information with a high dependence on time, information with a high time resolution, and information with large temporal variations.
[0067] That is, the characteristic parameters represented by the spreading code obtained by the signal analysis unit 2121-m are not parameters that represent the information in the differences of the two-channel digital audio signals being input that depends on the waveform of the audio signal emitted by the sound source, such as the difference in the waveforms of the two-channel digital audio signals being input or the phase difference between the two-channel digital audio signals being input as exemplified immediately above. Instead, they are parameters that represent the information in the differences of the two-channel digital audio signals being input that depends on the relative position in space between the sound source and the microphone, such as the time difference between the two-channel digital audio signals being input as exemplified in the first example of the signal analysis unit 2121-m or the intensity difference of each frequency band of the two-channel digital audio signals being input as exemplified in the second example of the signal analysis unit 2121-m. In short, the characteristic parameters represented by the spreading code obtained by the signal analysis unit 2121-m can also be said to be parameters that represent the characteristics of the differences in the two-channel digital audio signals being input and have a low time resolution, can also be said to be parameters that represent the characteristics of the differences in the two-channel digital audio signals being input and have small temporal variations, can also be said to be parameters that represent the characteristics of the differences in the two-channel digital audio signals being input and have a low dependence on time, and can also be said to be parameters that represent the characteristics of the inter-channel differences in the two-channel digital audio signals being input and depend on the relative position in space between the sound source and the microphone.
[0068] [[[Mono Encoding Unit 2122-m]]]
[0069] The mono - channel encoding unit 2122 - m encodes the input mono - channel signal frame - by - frame in a specific encoding method to obtain a mono - channel code and outputs it to the transmitting unit 213 - m. As the encoding method, an encoding method in which the bit rate of the mono - channel code is equal to or less than the communication capacity of the first communication line 410 - m is required. For example, an encoding method for telephone - band sound for mobile phones such as the 13.2 kbps mode of the 3GPP EVS standard (3GPP TS26.442) can be used.
[0070] That is, the encoding device 212 - m obtains, frame - by - frame, a mono - channel code representing a signal obtained by mixing the input digital audio signals of two channels, and an extended code representing characteristic parameters. The characteristic parameters are parameters that represent the difference between the channels of the input digital audio signals of two channels and have a low time resolution. In addition, as described later, the mono - channel code obtained by the encoding device 212 - m is a code included in the first code string and output to the first communication line, and the extended code obtained by the encoding device 212 - m is a code included in the second code string and output to the second communication line.
[0071] In addition, the encoding device 212 - m may use, as the extended code, a code representing the average or weighted average of the characteristic parameters obtained from the digital audio signals of two channels of the current frame, which is the frame to be processed, and the characteristic parameters obtained from the digital audio signals of two channels of a frame earlier than the current frame to be processed.
[0072] [Transmitting unit 213 - m]
[0073] The transmitting unit 213 - m outputs, frame - by - frame, the first code string including the mono - channel code input from the encoding device 221 - m to the first communication line 410 - m, and outputs the second code string including the extended code input from the encoding device 221 - m to the second communication line 510 - m (step S213).
[0074] The transmitting unit 213 - m outputs in such a way that it can be determined which frame's mono - channel code the first code string contains. For example, the transmitting unit 213 - m includes information that can determine the frame, such as the frame number or the time corresponding to the frame, as auxiliary information in the first code string and outputs it. Similarly, the transmitting unit 213 - m outputs in such a way that it can be determined which frame's extended code the second code string contains. For example, the transmitting unit 213 - m includes information that can determine the frame, such as the frame number or the time corresponding to the frame, as auxiliary information in the second code string and outputs it. In addition, in the audio signal receiving - side device 220 - m of this first embodiment, as well as in each of the following embodiments and variations, an example in which the frame number is included as auxiliary information in both the first code string and the second code string is described.
[0075] [Audio signal receiving - side device 220 - m]
[0076] The sound signal receiving-side device 220-m outputs, for example, for each specific time interval of 20 ms, that is, for each frame, sound based on the monaural code included in the first code string input from the first communication line 410-m and the extended code included in the second code string input from the second communication line 510-m.
[0077] [[Receiving unit 221-m]]
[0078] The receiving unit 221-m outputs, for each frame, to the decoding device 222-m the extended code among the monaural code included in the first code string input from the first communication line 410-m and the extended code included in the second code string input from the second communication line 510-m, whose frame number is closest to that of the monaural code (step S221).
[0079] Since the first communication line 410-m is a high-priority communication network for two-way calls, the first code string including the monaural code is input from the first communication line 410-m to the receiving unit 221-m so that the monaural code output in frame number order by the encoding device 212-m' of the sound signal transmitting-side device 210-m' of the multi-line support terminal device 200-m' (m' is an integer of 1 or more and M or less different from m) of the call partner can be output at time intervals of the frame length (i.e., at a specific time interval of, for example, 20 ms) in accordance with the frame number order. In addition, since the purpose of the telephone system 100 is to smoothly implement two-way calls, it is preferable that the receiving unit 221-m outputs the code output by the encoding device 212-m' of the sound signal transmitting-side device 210-m' of the call partner to the decoding device 222-m with as low a delay as possible. Therefore, the receiving unit 221-m outputs the monaural code included in the first code string output by the sound signal transmitting-side device 210-m' of the call partner to the decoding device 222-m at time intervals of the frame length in accordance with the frame number order output by the sound signal transmitting-side device 210-m' of the call partner, regardless of whether the second code string including the extended code having the same frame number as each monaural code is input to the receiving unit 221-m.
[0080] The second communication line 510-m is a communication network with low priority. Therefore, generally, the second code string of a certain frame output by the voice signal transmitting-side device 210-m' of the call partner is input to the receiving unit 221-m from the second communication line 510-m after the first code string of the frame is input to the receiving unit 221-m from the first communication line 410-m. That is, when the receiving unit 221-m outputs a monaural code to the decoding device 222-m, generally, the second code string including the extended code with the same frame number as the monaural code is not input to the receiving unit 221-m, and the extended code with the same frame number as the monaural code cannot be output to the decoding device 222-m. In addition, the second communication line 510-m is a communication network with low priority. Therefore, the second code strings of each frame output by the voice signal transmitting-side device 210-m' of the call partner are not necessarily input from the second communication line 510-m in the order of frame numbers. Of course, depending on the status of the second communication network 500, for example, when the second communication network 500 is idle, etc., the second code string of a certain frame output by the voice signal transmitting-side device 210-m' of the call partner may also be input to the receiving unit 221-m from the second communication line 510-m simultaneously with or before the first code string of the frame is input to the receiving unit 221-m from the first communication line 410-m. That is, there is also a situation where, when the receiving unit 221-m outputs a monaural code to the decoding device 222-m, the second code string including the extended code with the same frame number as the monaural code has already been input to the receiving unit 221-m, and the extended code with the same frame number as the monaural code can be output to the decoding device 222-m. Therefore, for each frame, the receiving unit 221-m outputs, instead of the extended code among the extended codes included in the second code string input from the second communication line 510-m and having the same frame number as the monaural code output to the decoding device 222-m, the extended code among the extended codes included in the second code string input from the second communication line 510-m and having the frame number closest to the monaural code output to the decoding device 222-m, to the decoding device 222-m. In other words, for each frame, the receiving unit 221-m outputs the extended code included in the second code string among the second code strings input from the second communication line 510-m and having the frame number closest to the first code string included in the monaural code output to the decoding device 222-m, to the decoding device 222-m.
[0081] Here, among the extended codes included in the second code string input from the second communication line 510-m, the extended code whose frame number is closest to the mono code output to the decoding device 222-m is, when the extended codes included in the second code string input from the second communication line 510-m include an extended code with the same frame number as the mono code output to the decoding device 222-m, the extended code with the same frame number as the mono code output to the decoding device 222-m among the extended codes included in the second code string input from the second communication line 510-m, and when the extended codes included in the second code string input from the second communication line 510-m do not include an extended code with the same frame number as the mono code output to the decoding device 222-m, the extended code whose frame number is closest to the mono code output to the decoding device 222-m (that is, among the extended codes included in the second code string input from the second communication line 510-m, the extended code whose frame number is not the same as the mono code output to the decoding device 222-m but is closest to the mono code output to the decoding device 222-m). The same applies to the embodiments or modification examples described later.
[0082] That is, the receiving unit 221-m outputs, for each frame, the mono code included in the first code string input from the first communication line 410-m and the extended code whose frame number is closest to the mono code among the extended codes included in the second code string input from the second communication line 510-m. Of course, the receiving unit 221-m outputs the mono codes in the order of frame numbers. More specifically, the receiving unit 221-m accepts the input of the first code string from the first communication line 410-m and the input of the second code string from the second communication line 510-m, and outputs, for each frame, the mono code included in the first code string input from the first communication line 410-m (that is, the mono codes in the order of frame numbers). When the extended codes included in the second code string input from the second communication line 510-m include an extended code with the same frame number as the mono code, it outputs the extended code with the same frame number as the mono code. When the extended codes included in the second code string input from the second communication line 510-m do not include an extended code with the same frame number as the mono code, it outputs the extended code whose frame number is closest to the mono code among the extended codes included in the second code string input from the second communication line (that is, among the extended codes included in the second code string input from the second communication line, the extended code whose frame number is different from the mono code but is closest to the mono code).
[0083] In addition, although not described in detail because it is well-known technology, the receiving unit 221-m has a storage unit (not shown). This storage unit accumulates a plurality of frames of code strings asynchronously received from each communication line due to communication involving fluctuations (shakiness) or retransmission control, etc. Although the code strings are not necessarily input to the receiving unit 221-m from each communication line at specific time intervals or in the order of frame numbers, the receiving unit 221-m can output any code as long as it is included in the code strings accumulated in the storage unit. That is, the receiving unit 221-m accepts the input of the first code string from the first communication line 410-m and stores it. After storing the input-completed first code string, it can output the stored first code string. In addition, the receiving unit 221-m accepts the input of the second code string from the second communication line 510-m and stores it. After storing the input-completed second code string, it can output the stored second code string. Thus, the receiving unit 221-m can extract the monaural code in the order of frame numbers, or extract the extended code whose frame number is closest to that monaural code, for each specific time interval, that is, for each frame.
[0084] [[Decoding device 222-m]]
[0085] The monaural code and the extended code output by the receiving unit 221-m are input to the decoding device 222-m for each frame. The decoding device 222-m obtains, for each frame, two-channel decoded digital audio signals corresponding to the input monaural code and extended code and outputs them to the playback unit 223-m (step S222).
[0086] Input to the decoding device 222-m are the monaural codes in the order of frame numbers respectively included in the first code string input from the first communication line 410-m in the order of frame numbers, and the extended codes whose frame numbers are closest to the respective monaural codes included in the second code string input from the second communication line 510-m. That is, the decoding device 222-m obtains, for each frame, two-channel decoded digital audio signals based on the monaural codes included in the first code string input from the first communication line 410-m and the extended codes whose frame numbers are closest to those monaural codes included in the second code string input from the second communication line 510-m, and outputs them. In addition, the monaural codes used by the decoding device 222-m are of course in the order of frame numbers.
[0087] In other words, input to the decoding device 222-m are the monaural codes in the order of frame numbers output by the encoding device 212-m' of the voice signal transmitting side device 210-m' of the call partner, and the extended codes whose frame numbers are closest to those monaural codes. That is, the decoding device 222-m obtains, for each frame, two-channel decoded digital audio signals based on the monaural codes in the order of frame numbers output by the encoding device 212-m' of the voice signal transmitting side device 210-m' of the call partner and the extended codes whose frame numbers are closest to those monaural codes, and outputs them to the playback unit 223-m.
[0088] Here, as the extended code input to the decoding device 222-m, when the extended code included in the second code string input from the second communication line 510-m includes a frame whose frame number is the same as the mono code included in the first code string input from the first communication line 410-m, it is the extended code whose frame number included in the second code string input from the second communication line 510-m is the same as the mono code of that frame. When the extended code included in the second code string input from the second communication line 510-m does not include a frame whose frame number is the same as the mono code included in the first code string input from the first communication line 410-m, it is the extended code whose frame number included in the second code string input from the second communication line 510-m is closest to the mono code of that frame (that is, the extended code whose frame number is different from the mono code of that frame but is closest to the mono code of that frame). This is the same in the embodiments or variations described later.
[0089] Accordingly, for each frame, when the extended code included in the second code string input from the second communication line 510-m includes a frame whose frame number is the same as the mono code (i.e., the mono code in frame number order) included in the first code string input from the first communication line 410-m, based on the mono code (i.e., the mono code in frame number order) included in the first code string input from the first communication line 410-m and the extended code whose frame number is the same as that mono code, the decoding device 222-m obtains and outputs the decoded digital audio signals for two channels. When the extended code included in the second code string input from the second communication line 510-m does not include a frame whose frame number is the same as the mono code (i.e., the mono code in frame number order) included in the first code string input from the first communication line 410-m, based on the mono code (i.e., the mono code in frame number order) included in the first code string input from the first communication line 410-m and the extended code whose frame number included in the second code string input from the second communication line 510-m is closest to that mono code (that is, the extended code whose frame number is different from that mono code but is closest to that mono code), the decoding device 222-m obtains and outputs the decoded digital audio signals for two channels.
[0090] [[[Mono decoding unit 2221-m]]]
[0091] A monaural code input per frame is input to the monaural decoding unit 2221-m of the decoding device 222-m. The monaural decoding unit 2221-m decodes the input monaural code in a specific decoding manner per frame, obtains a decoded digital audio signal of the monaural channel, and outputs it to the extended decoding unit 2222-m. As the specific decoding manner, a decoding manner corresponding to the encoding manner used in the monaural encoding unit 2122-m' of the encoding device 212-m' of the voice signal transmitting side device 210-m' of the call partner is used.
[0092] What is input to the monaural decoding unit 2221-m is the monaural code in the frame number order output by the encoding device 212-m' of the voice signal transmitting side device 210-m' of the call partner. That is, the monaural decoding unit 2221-m obtains, per frame, a decoded digital audio signal of the monaural channel in the frame number order encoded by the encoding device 212-m' of the voice signal transmitting side device 210-m' of the call partner and outputs it to the extended decoding unit 2222-m.
[0093] [[Extended decoding unit 2222-m]]
[0094] The decoded digital audio signal of the monaural channel output by the monaural decoding unit 2221-m and the extended code input to the decoding device 222-m are input to the extended decoding unit 2222-m per frame. The extended decoding unit 2222-m obtains, per frame, a decoded digital audio signal of two channels based on the input decoded digital audio signal of the monaural channel and the extended code, and outputs it to the playback unit 223-m.
[0095] The decoded digital sound signal of the mono input to the extended decoding unit 2222-m is input to the decoding unit 222-m in the frame number order encoded by the encoding device 212-m' of the voice signal transmitting side device 210-m' of the call partner. The extended code input is the extended code closest to the frame number of the decoded digital sound signal of the mono. That is, the extended decoding unit 2222-m obtains the decoded digital sound signals of two channels and outputs them to the playback unit 223-m for each frame, based on the decoded digital sound signal of the mono in the frame number order output by the encoding device 212-m' of the voice signal transmitting side device 210-m' of the call partner, and the extended code closest to the decoded digital sound signal of the mono. In addition, the extended code represents the characteristic parameters obtained by the encoding device 212-m' of the voice signal transmitting side device 210-m' of the multi-line support terminal device 200-m' of the call partner, and thus represents the following parameters, which represent the characteristics of the difference between the digital sound signals of two channels. That is, the extended decoding unit 2222-m regards the decoded digital sound signal of the mono input as a signal formed by mixing the decoded digital sound signals of two channels for each frame, regards the characteristic parameters obtained based on the extended code as information representing the characteristics of the difference between the digital sound signals of two channels, obtains the decoded digital sound signals of two channels, and outputs them to the playback unit 223-m.
[0096] 〔First example of extended decoding unit 2222-m〕
[0097] As a first example, the operation of the extended decoding unit 2222-m for each frame is described in the case where the characteristic parameter is information representing the time difference between the digital sound signals of two channels. The extended decoding unit 2222-m first obtains information representing the time difference as the characteristic parameter represented by the extended code based on the input extended code (step S2222-11). The extended decoding unit 2222-m obtains the characteristic parameter based on the extended code in a manner corresponding to the manner in which the signal analysis unit 2121-m' of the encoding device 212-m' of the voice signal transmitting side device 210-m' of the call partner obtains the extended code based on the characteristic parameter. The information representing the time difference as the characteristic parameter is, for example, the number of time difference samples. For example, the extended decoding unit 2222-m performs scalar decoding on the input extended code to obtain a scalar value corresponding to the input extended code as the number of time difference samples. Or, for example, the extended decoding unit 2222-m regards the input extended code as a binary number value and obtains a decimal number corresponding to the binary number as the number of time difference samples.
[0098] The extended decoding unit 2222-m then regards the input decoded digital audio signal of a single channel as a signal formed by mixing two decoded digital audio signals, and regards the characteristic parameter as information representing the time difference between the two decoded digital audio signals, and obtains and outputs two decoded digital audio signals (step S2222-12). More specifically, the extended decoding unit 2222-m obtains the sample sequence of the input decoded digital audio signal of a single channel itself, the sequence of values obtained by dividing the value of each sample of the sample sequence of the input decoded digital audio signal of a single channel by 2, and one of the sequences obtained by transforming one of the above sample sequences as the digital audio signal of the first channel and outputs it (step S2222-121). The extended decoding unit 2222-m further obtains the sample sequence obtained by delaying the digital audio signal of the first channel by the number of sample of the time difference represented by the characteristic parameter as the sample sequence of the digital audio signal of the second channel and outputs it (step S2222-122).
[0099] 〔Second Example of Extended Decoding Unit 2222-m〕
[0100] As a second example, the operation of each frame of the extended decoding unit 2222-m is described in the case where the characteristic parameter is information representing the intensity difference of each frequency band of the digital audio signals of two channels. The extended decoding unit 2222-m first decodes the input extended code to obtain information representing the intensity difference of each frequency band (step S2222-21). The extended decoding unit 2222-m obtains the characteristic parameter from the extended code in a manner corresponding to the manner in which the signal analysis unit 2121-m' of the encoding device 212-m' of the voice signal transmitting side device 210-m' of the call partner obtains the extended code based on the information representing the intensity difference of each frequency band. For example, the extended decoding unit 2222-m performs vector decoding on the input extended code to obtain the element values of the vector corresponding to the input extended code as information representing the intensity difference of each of the multiple frequency bands. Or, for example, the extended decoding unit 2222-m performs scalar decoding on each of the codes included in the input extended code to obtain information representing the intensity difference of each frequency band. In addition, when the number of frequency bands is 1, the extended decoding unit 2222-m performs scalar decoding on the input extended code to obtain information representing the intensity difference of one frequency band, that is, the entire frequency band.
[0101] The extended decoding unit 2222-m then regards the input decoded digital audio signal of a mono-channel as a signal formed by mixing two decoded digital audio signals, and regards the characteristic parameter as information representing the intensity difference of each frequency band of the two decoded digital audio signals, and obtains and outputs two decoded digital audio signals based on the input decoded digital audio signal of the mono-channel and the characteristic parameter obtained in step S2222-21 (step S2222-22). If the signal analysis unit 2121-m' of the encoding device 212-m' of the voice signal transmitting-side device 210-m' of the call partner performs the operation of the above specific example using the complex DFT, the extended decoding unit 2222-m performs the following operations.
[0102] The extended decoding unit 2222-m first performs an inverse complex DFT on the input decoded digital audio signal of a single channel to obtain a sequence of complex DFT coefficients (step S2222-221). Thereafter, each complex DFT coefficient of the sequence of complex DFT coefficients of the single channel obtained by the extended decoding unit 2222-m is set as MQ(f). The extended decoding unit 2222-m then obtains, based on the sequence of complex DFT coefficients of the single channel, the value of the radius MQr(f) of each complex DFT coefficient in the complex plane and the value of the angle MQθ(f) of each complex DFT coefficient in the complex plane (step S2222-222). The extended decoding unit 2222-m then obtains the value obtained by multiplying each value of the radius MQr(f) by the square root of the corresponding value among the characteristic parameters as the value of the radius VLQr(f) of each channel of the first channel, and obtains the value obtained by dividing each value of the radius MQr(f) by the square root of the corresponding value among the characteristic parameters as the value of the radius VRQr(f) of each channel of the second channel (step S2222-223). If it is the above example of four bands, the corresponding value among the characteristic parameters of each frequency bin is Mr(1) for f from 1 to 32, Mr(2) for f from 33 to 64, Mr(3) for f from 65 to 96, and Mr(4) for f from 97 to 128. In addition, when the signal analysis unit 2121-m' of the encoding device 212-m' of the voice signal transmitting side device 210-m' of the call partner uses the difference between the value of the radius of the first channel and the value of the radius of the second channel instead of the ratio of the value of the radius of the first channel to the value of the radius of the second channel, the extended decoding unit 2222-m obtains the value obtained by adding each value of the radius MQr(f) to the value obtained by dividing the corresponding value among the characteristic parameters by 2 as the value of the radius VLQr(f) of each channel of the first channel, and obtains the value obtained by subtracting the value obtained by dividing the corresponding value among the characteristic parameters by 2 from each value of the radius MQr(f) as the value of the radius VRQr(f) of each channel of the second channel. The extended decoding unit 2222-m then performs an inverse complex DFT on the sequence of complex numbers with a radius of VLQr(f) and an angle of MQθ(f) in the complex plane to obtain and output the decoded digital audio signal of the first channel, and performs an inverse complex DFT on the sequence of complex numbers with a radius of VRQr(f) and an angle of MQθ(f) in the complex plane to obtain and output the decoded digital audio signal of the second channel (step S2222-224).
[0103] [[Playback unit 223-m]]
[0104] The playback unit 223-m outputs the sound corresponding to the decoded digital audio signals of the two input channels (step S223).
[0105] The playback unit 223-m includes, for example, two DA conversion units and two speakers. The DA conversion unit converts the input decoded digital audio signal into an analog electrical signal and outputs it. The speakers generate sound corresponding to the analog electrical signal input from the DA conversion unit. The speakers may also be configured in stereo headphones or stereo earphones. In this case, for example, the playback unit 223-m associates the DA conversion units with the speakers one-to-one, and generates sounds (decoded audio signals) corresponding to the two decoded digital audio signals respectively from the two speakers.
[0106] In addition, all or part of the playback unit 223-m may not be configured inside the audio signal receiving device 220-m, but connected to the audio signal receiving device 220-m. For example, the playback unit 223-m of the audio signal receiving device 220-m may not have speakers, but output the two analog electrical signals obtained by the DA conversion unit of the playback unit 223-m of the audio signal receiving device 220-m to the speakers connected to the audio signal receiving device 220-m. Alternatively, the audio signal receiving device 220-m may not have the playback unit 223-m, but the decoding device 222-m of the audio signal receiving device 220-m outputs the two-channel decoded digital audio signal to a playback device such as a DA converter connected to the audio signal receiving device 220-m.
[0107] 〔Example of operation of audio signal receiving device 220-m〕
[0108] Figure 5 It is a diagram schematically showing the temporal relationship between the monaural code included in the first code string input from the first communication line 410-m to the audio signal receiving device 220-m, the extended code included in the second code string input from the second communication line 510-m to the audio signal receiving device 220-m, and the decoded audio signal output by the audio signal receiving device 220-m, excluding the processing delay depending on the processing capacity of the device. Figure 5 The horizontal axis of is the time axis. The serial number i in the parentheses is the frame number in the encoding device 212-m' of the audio signal transmitting device 210-m' of the multi-line support terminal device 200-m' of the other party in the call. CM(i) is the monaural code included in the first code string input from the first communication line 410-m to the audio signal receiving device 220-m. CE(i) is the extended code included in the second code string input from the second communication line 510-m to the audio signal receiving device 220-m. YS'(i) is the decoded audio signal output by the audio signal receiving device 220-m. Figure 5This is an example in which, although the second code string is input to the voice signal receiving device 220-m from the second communication line 510-m, which is a communication network with a lower priority, in the frame number order, the second code string is input 5 frames later than the first code string in the frame number order from the first communication line 410-m, which is a communication network with a higher priority.
[0109] When the receiving unit 221-m finishes receiving the first code string including the monaural code CM(6) with frame number 6 from the first communication line 410-m, it outputs the monaural code CM(6) included in the first code string input from the first communication line 410-m, and the extension code CE(1) included in the second code string with the frame number closest to the monaural code CM(6) among the second code strings input from the second communication line 510-m, to the decoding device 222-m. When the decoding device 222-m is input with the monaural code CM(6) and the extension code CE(1), it obtains two-channel decoded digital voice signals corresponding to the input monaural code CM(6) and extension code CE(1) and outputs them to the playback unit 223-m. The playback unit 223-m starts outputting two-channel decoded voice signals YS'(6) corresponding to the two-channel decoded digital voice signals input thereto from the moment when the two-channel decoded digital voice signals corresponding to the monaural code CM(6) and the extension code CE(1) are input. Thus, the voice signal receiving device 220-m can, at the moment when the receiving unit 221-m finishes receiving the first code string including the monaural code CM(6) with frame number 6 from the first communication line 410-m, obtain two-channel decoded voice signals YS'(6) based on the monaural code CM(6) with frame number 6 and the extension code CE(1) included in the second code string with the frame number closest thereto, and start outputting them.
[0110] Similarly, the voice signal receiving device 220-m also operates in such a way that, at the moment when the receiving unit 221-m finishes receiving the first code string including the monaural code CM(7) with frame number 7 from the first communication line 410-m, it obtains two-channel decoded voice signals YS'(7) based on the monaural code CM(7) with frame number 7 and the extension code CE(2) included in the second code string with the frame number closest thereto, and starts outputting them; at the moment when the receiving unit 221-m finishes receiving the first code string including the monaural code CM(8) with frame number 8 from the first communication line 410-m, it obtains two-channel decoded voice signals YS'(8) based on the monaural code CM(8) with frame number 8 and the extension code CE(3) included in the second code string with the frame number closest thereto, and starts outputting them...
[0111] Figure 6This is a diagram schematically showing the temporal relationship between the monaural code included in the first code string input from the first communication line 410-m to the sound signal receiving device, the extended code included in the second code string input from the second communication line 510-m to the sound signal receiving device 220-m, and the decoded sound signal output by the sound signal receiving device, excluding the processing delay that depends on the processing power of the device, in the case of using the technology of Patent Document 1. Figure 6 The horizontal axis, the serial number i in parentheses, CM(i), CE(i) are the same as Figure 5 YS(i) is the decoded sound signal output by the sound signal receiving device using the technology of Patent Document 1. Figure 6 is also the same as Figure 5 Similarly, this is an example where, although the second code string is input to the sound signal receiving device from the second communication line 510-m, which is a communication network with a lower priority, in the order of frame numbers, the second code string is input 5 frames later than the first code string in the order of frame numbers from the first communication line 410-m, which is a communication network with a higher priority. Figure 6 This is an example where the above-mentioned restricted time in the sound signal receiving device using the technology of Patent Document 1 is a time of 5 frames.
[0112] The sound signal receiving device using the technology of Patent Document 1 obtains the decoded sound signal YS(6) of two channels corresponding to the monaural code CM(6) input from the first communication line 410-m and the extended code CE(6) input from the second communication line 510-m exactly 5 frames after the monaural code CM(6) is input, and starts to output. Similarly, the sound signal receiving device using the technology of Patent Document 1 also operates in such a way that, based on the monaural code CM(7) of frame number 7 and the extended code CE(7) of frame number 7 input from the second communication line 510-m at the moment 5 frames after receiving the monaural code CM(7) from the first communication line 410-m, it obtains the decoded sound signal YS(7) of two channels and starts to output, and based on the monaural code CM(8) of frame number 8 and the extended code CE(8) of frame number 8 input from the second communication line 510-m at the moment 5 frames after receiving the monaural code CM(8) from the first communication line 410-m, it obtains the decoded sound signal YS(8) of two channels and starts to output...
[0113] 〔Effect〕
[0114] According to Figure 6 and Figure 5It can also be determined that in the technology of Patent Document 1, in order to obtain a decoded sound signal with high sound quality, compared with obtaining a decoded sound signal with the minimum sound quality, there is a delay of 5 frames. However, in the technology of the first embodiment, compared with the case of obtaining a decoded sound signal with the minimum sound quality, the delay time is not increased significantly, that is, a decoded sound signal with high sound quality can be obtained with a delay time that does not cause a sense of disharmony during two-way calls.
[0115] <Second Embodiment>
[0116] In the first embodiment, the spreading code for each frame is obtained and output, but it is also possible to obtain the spreading code only once for multiple frames and output it. This method will be described as the second embodiment.
[0117] The difference between the second embodiment and the first embodiment lies in the operations of the signal analysis unit 2121-m and the transmission unit 213-m of the encoding device 212-m of the sound signal transmission side device 210-m. Hereinafter, the differences between the second embodiment and the first embodiment will be described.
[0118] [[Signal Analysis Unit 2121-m]]
[0119] Similar to the signal analysis unit 2121-m of the first embodiment, for each frame, the signal analysis unit 2121-m obtains a mono signal, which is a signal obtained by mixing the two-channel digital sound signals to be input, and outputs it. However, different from the signal analysis unit 2121-m of the first embodiment, for only the pre-determined frames among multiple frames, the spreading code representing the characteristic parameters is obtained and output. The characteristic parameters are parameters that represent the difference between the two-channel digital sound signals to be input and have little temporal variation.
[0120] For example, the signal analysis unit 2121-m obtains the characteristic parameters according to the two-channel digital sound signals to be input for the frames with odd frame numbers, obtains the spreading code representing the characteristic parameters and outputs it. However, for the frames with even frame numbers, the signal analysis unit 2121-m does not obtain the characteristic parameters, does not obtain the spreading code representing the characteristic parameters, and does not output it. In addition, in the case where the signal analysis unit 2121-m adopts a structure that uses the characteristic parameters when obtaining the mono signal, the signal analysis unit 2121-m uses the two-channel digital sound signals to be input for the frames for which the characteristic parameters have not been obtained, and the characteristic parameters corresponding to the latest spreading code among the spreading codes that have already been output, to obtain the mono signal.
[0121] Alternatively, for example, the signal analysis unit 2121-m obtains characteristic parameters from the digital audio signals of two channels input thereto for frames with odd frame numbers, but does not obtain or output the extension code representing the characteristic parameters. For frames with even frame numbers, the signal analysis unit 2121-m obtains characteristic parameters from the digital audio signals of two channels input thereto, obtains the extension code representing the average or weighted average of the characteristic parameters of the immediately preceding frame and the characteristic parameters of this frame, and outputs the extension code, where the immediately preceding frame does not obtain or output the extension code representing the characteristic parameters. The weight for weighted average may be set to a value larger than the weight of the immediately preceding frame for this frame.
[0122] In the above two examples, the extension code is obtained and output once for two frames, but it may also be configured to obtain and output the extension code once for three or more frames, or it may be configured to obtain the extension code for predetermined frames among multiple frames and output the extension code.
[0123] That is, the encoding device 212-m of the second embodiment obtains, for each frame, the monophonic code representing the signal obtained by mixing the digital audio signals of two channels input thereto, and for predetermined frames among multiple frames, obtains the extension code representing the characteristic parameters, where the characteristic parameters are the parameters representing the difference between channels of the digital audio signals of two channels input thereto and have low time resolution.
[0124] Alternatively, the encoding device 212-m of the second embodiment obtains, for each frame, the monophonic code representing the signal obtained by mixing the digital audio signals of two channels input thereto, obtains, for each frame, the characteristic parameters as the parameters representing the difference between channels of the digital audio signals of two channels input thereto and having low time resolution, and for predetermined frames among multiple frames, obtains the extension code representing the average or weighted average of the characteristic parameters obtained in each frame after the immediately preceding predetermined frame. The weight for weighted average may be set to the maximum value for this frame and to smaller values for frames farther away from this frame.
[0125] In addition, as will be described later, the monophonic code obtained by the encoding device 212-m is the code included in the first code string and output to the first communication line, and the extension code obtained by the encoding device 212-m is the code included in the second code string and output to the second communication line.
[0126] [[Transmission unit 213-m]]
[0127] The transmitting unit 213-m is the same as the transmitting unit 213-m in the first embodiment. For each frame, it outputs the first code string, which is a code string including the input monaural code, to the first communication line 410-m. However, different from the transmitting unit 213 in the first embodiment, for the frames with the input spread code, that is, only for the predetermined frames among the multiple frames, it outputs the second code string, which is a code string including the input spread code, to the second communication line 510-m.
[0128] 〔Effect〕
[0129] As described in the first embodiment, the spread code used in the sound signal receiving side device 220-m is the spread code whose frame number is closest to the monaural code. Therefore, the spread code with the same frame number as the monaural code is not necessarily input to the sound signal receiving side device 220-m. In addition, the characteristic parameter is originally a parameter with little temporal variation. Thus, according to this embodiment, by adopting a structure that obtains and outputs the spread code only once for multiple frames, compared with the first embodiment, the quality of the decoded sound signal will not deteriorate significantly, the computational processing amount of the signal analysis unit 2121-m can be reduced, and in addition, compared with the first embodiment, the code amount for transmitting the characteristic parameter can be reduced.
[0130] <Third Embodiment>
[0131] In the first embodiment, the sound signal receiving side device 220-m obtains the spread code for decoding for each frame. However, the sound signal receiving side device 220-m can also obtain the spread code for decoding only once for multiple frames. This method will be described as the third embodiment.
[0132] The difference between the sound signal receiving side device 220-m in the third embodiment and the sound signal receiving side device 220-m in the first embodiment lies in the operations of the receiving unit 221-m and the extended decoding unit 2222-m of the decoding device 222-m. In the following, the differences between the third embodiment and the first embodiment will be described.
[0133] [[Receiving Unit 221-m]]
[0134] The receiving unit 221-m, similar to the receiving unit 221-m in the first embodiment, outputs the monaural code included in the first code string input from the first communication line 410-m to the decoding device 222-m for each frame. However, different from the receiving unit 221-m in the first embodiment, for only the predetermined frames among the multiple frames, it obtains and outputs the extended code with the frame number closest to the monaural code among the extended codes included in the input second code string. That is, more specifically, the receiving unit 221-m obtains and outputs the extended code with the frame number closest to the monaural code among the extended codes included in the input second code string from a storage unit (not shown) within the receiving unit 221-m for only the predetermined frames among the multiple frames.
[0135] [[[Extended decoding unit 2222-m]]]
[0136] Similar to the extended decoding unit 2222-m in the first embodiment, for each frame, the monaural decoded digital audio signal output by the monaural decoding unit 2221-m is input to the extended decoding unit 2222-m. However, different from the extended decoding unit 2222-m in the first embodiment, for only the predetermined frames among the multiple frames, the extended code is input to the extended decoding unit 2222-m. For the predetermined frames among the multiple frames, that is, the frames to which the extended code is also input, similar to the extended decoding unit 2222-m in the first embodiment, based on the input monaural decoded digital audio signal and the extended code, the decoded digital audio signal of two channels is obtained and output. For the frames other than the predetermined frames among the multiple frames, that is, the frames to which the extended code is not input, different from the extended decoding unit 2222-m in the first embodiment, based on the input monaural decoded digital audio signal and the latest extended code among the extended codes that have already been input, the decoded digital audio signal of two channels is obtained and output.
[0137] That is, the decoding device 222-m obtains and outputs decoded digital audio signals for two channels based on the monaural code included in the first code string input from the first communication line 410-m and the extended code whose frame number is closest to the monaural code included in the second code string input from the second communication line 510-m for a predetermined frame among a plurality of frames. For frames other than the predetermined frame, the decoding device 222-m obtains and outputs decoded digital audio signals for two channels based on the monaural code included in the first code string input from the first communication line 410-m and the latest extended code used in the predetermined frame. Specifically, for a predetermined frame among a plurality of frames, when the extended code included in the second code string input from the second communication line 510-m includes an extended code whose frame number is the same as the monaural code (i.e., the monaural code in the order of frame numbers) included in the first code string input from the first communication line 410-m, the decoding device 222-m obtains and outputs decoded digital audio signals for two channels based on the monaural code (i.e., the monaural code in the order of frame numbers) included in the first code string input from the first communication line 410-m and the extended code whose frame number is the same as the monaural code. When the extended code included in the second code string input from the second communication line 510-m does not include an extended code whose frame number is the same as the monaural code (i.e., the monaural code in the order of frame numbers) included in the first code string input from the first communication line 410-m, the decoding device 222-m obtains and outputs decoded digital audio signals for two channels based on the monaural code (i.e., the monaural code in the order of frame numbers) included in the first code string input from the first communication line 410-m and the extended code whose frame number is closest to the monaural code (i.e., the extended code whose frame number is not the same as the monaural code but is closest to the monaural code) included in the second code string input from the second communication line 510-m. For frames other than the predetermined frame, the decoding device 222-m obtains and outputs decoded digital audio signals for two channels based on the monaural code (i.e., the monaural code in the order of frame numbers) included in the first code string input from the first communication line 410-m and the latest extended code used in the predetermined frame.
[0138] More specifically, the mono decoding unit 2221-m of the decoding device 222-m decodes the mono codes included in the first code string input from the first communication line 410-m for each frame to obtain a decoded digital audio signal of mono. The expansion decoding unit 2222-m of the decoding device 222-m regards the decoded digital audio signal of mono as a signal formed by mixing the decoded digital audio signals of two channels for a predetermined frame among a plurality of frames, and regards the characteristic parameters obtained based on the expansion code whose frame number included in the second code string input from the second communication line 510-m is closest to the mono code included in the first code string input from the first communication line 410-m as the information representing the characteristics of the difference between channels in the decoded digital audio signals of two channels, and obtains and outputs the decoded digital audio signals of two channels. In addition, since the expansion decoding unit 2222-m uses the characteristic parameters obtained based on the expansion code in the predetermined frame, it can store the characteristic parameters in advance and use them in frames other than the predetermined frame. That is, in frames other than the predetermined frame, the expansion decoding unit 2222-m regards the decoded digital audio signal of mono as a signal formed by mixing the decoded digital audio signals of two channels, regards the latest characteristic parameters obtained in the predetermined frame as the information representing the characteristics of the difference between channels in the decoded digital audio signals of two channels, and obtains and outputs the decoded digital audio signals of two channels.
[0139] That is, the mono decoding unit 2221-m of the decoding device 222-m decodes the mono codes (i.e., the mono codes in the order of frame numbers) included in the first code string input from the first communication line 410-m for each frame to obtain a decoded digital audio signal of mono. The extended decoding unit 2222-m of the decoding device 222-m, for a predetermined frame among a plurality of frames, when the extended code included in the second code string input from the second communication line 510-m includes an extended code with the same frame number as the mono code (i.e., the mono code in the order of frame numbers) included in the first code string input from the first communication line 410-m, regards the decoded digital audio signal of mono as a signal formed by mixing the decoded digital audio signals of two channels, regards the characteristic parameters obtained based on the extended code with the same frame number as the mono code as information representing the characteristics of the difference between the two channels in the decoded digital audio signals of two channels, obtains the decoded digital audio signals of two channels and outputs them. When the extended code included in the second code string input from the second communication line 510-m does not include an extended code with the same frame number as the mono code (i.e., the mono code in the order of frame numbers) included in the first code string input from the first communication line 410-m, regards the decoded digital audio signal of mono as a signal formed by mixing the decoded digital audio signals of two channels, regards the characteristic parameters obtained based on the extended code with the frame number closest to the mono code (i.e., the extended code whose frame number is not the same as the mono code but is the closest to the mono code) included in the second code string input from the second communication line 510-m as information representing the characteristics of the difference between the two channels in the decoded digital audio signals of two channels, obtains the decoded digital audio signals of two channels and outputs them. For frames other than the predetermined frames, regards the decoded digital audio signal of mono as a signal formed by mixing the decoded digital audio signals of two channels, regards the latest characteristic parameters obtained in the predetermined frames as information representing the characteristics of the difference between the two channels in the decoded digital audio signals of two channels, obtains the decoded digital audio signals of two channels and outputs them.
[0140] <Modification Example of the Third Embodiment>
[0141] In addition, instead of the third embodiment, the extended decoding unit 2222-m may perform the same operation as the first embodiment, and the receiving unit 221-m outputs, for a predetermined frame among a plurality of frames, the mono code included in the first code string input from the first communication line 410-m and the extended code with the frame number closest to the mono code among the extended codes included in the second code string input from the second communication line 510-m, and for frames other than the predetermined frames among a plurality of frames, outputs the mono code included in the first code string input from the first communication line 410-m and the latest extended code among the extended codes that have been output.
[0142] More specifically, the receiving unit 221-m may also output the monaural code and the extended code having the same frame number as the monaural code when the extended code included in the second code string input from the second communication line 510-m includes an extended code having the same frame number as the monaural code (i.e., the monaural code in the order of frame numbers) among a plurality of frames. When the extended code included in the second code string input from the second communication line 510-m does not include an extended code having the same frame number as the monaural code (i.e., the monaural code in the order of frame numbers) included in the first code string input from the first communication line 410-m, the monaural code (i.e., the monaural code in the order of frame numbers) included in the first code string input from the first communication line 410-m and the extended code having the frame number closest to the monaural code among the extended codes included in the second code string input from the second communication line 510-m (i.e., the extended code included in the second code string input from the second communication line 510-m that has a frame number different from the monaural code but is closest to the monaural code) are output. For frames other than the predetermined frame among the plurality of frames, the monaural code (monaural code in the order of frame numbers) included in the first code string input from the first communication line 410-m and the latest extended code among the already output extended codes are output.
[0143] 〔Effect〕
[0144] As described in the first embodiment, the extended code used in the sound signal receiving side device 220-m is the extended code with the frame number closest to the monaural code. Therefore, an extended code with the same frame number as the monaural code is not necessarily input to the extended decoding unit 2222-m. In addition, the characteristic parameter is originally a parameter with little temporal variation. Thus, according to the present embodiment and this modification example, by adopting a structure that obtains the extended code only once for a plurality of frames, the quality of the decoded sound signal is not significantly deteriorated compared to the first embodiment, and the arithmetic processing amount of the receiving unit 221-m or the amount of output information can be reduced.
[0145] <Fourth Embodiment>
[0146] As the characteristic parameter used by the sound signal receiving side device 220-m of the first embodiment when obtaining two decoded digital sound signals, the average or weighted average of the characteristic parameter represented by the extended code input in the frame to be processed and the characteristic parameter of the past frame may also be used. This method will be described as the fourth embodiment.
[0147] The difference between the fourth embodiment and the first embodiment lies in the operation of the extended decoding unit 2222-m of the decoding device 222-m of the sound signal receiving-side device 220-m. Hereinafter, the differences between the fourth embodiment and the first embodiment will be described. Hereinafter, the extended decoding unit 2222-m that processes each frame is referred to as the current frame at this moment, and the past frames are referred to as past frames.
[0148] [[[Extended decoding unit 2222-m]]]
[0149] Similar to the extended decoding unit 2222-m of the first embodiment, for each frame, a monaural decoded digital sound signal output by the monaural decoding unit 2221-m and an extended code input to the decoding device 222-m are input to the extended decoding unit 2222-m. The extended decoding unit 2222-m includes a storage unit (not shown). In the storage unit, the feature parameters obtained by the extended decoding unit 2222-m in the past frames are stored. The extended decoding unit 2222-m obtains a decoded digital sound signal of two channels based on the input monaural decoded digital sound signal, the input extended code, and the feature parameters of the past frames stored in the storage unit for each frame, and outputs it to the playback unit 223-m. Specifically, the extended decoding unit 2222-m performs the following steps S2222-31 to step S2222-35 for each frame.
[0150] The extended decoding unit 2222-m first obtains the characteristic parameters represented by the input extended code (step S2222-31), and stores the obtained characteristic parameters in the storage unit (step S2222-32). The extended decoding unit 2222-m then reads out K (K is an integer greater than or equal to 1) of the characteristic parameters of the past frames stored in the storage unit (step S2222-33). For example, the characteristic parameters of the past K past frames consecutive to the current frame are read out. The extended decoding unit 2222-m then obtains the average or weighted average of the K past frame characteristic parameters read from the storage unit and the current frame characteristic parameters (step S2222-34). The weight for the weighted average only needs to be set to the maximum value for the current frame characteristic parameter and to a smaller value for frames farther from the current frame. The extended decoding unit 2222-m then regards the input monaural decoded digital sound signal and the average or weighted average of the characteristic parameters obtained in step S2222-34 as a signal formed by mixing two decoded digital sound signals, regards the average or weighted average of the characteristic parameters obtained in step S2222-34 as information representing the difference characteristics of the two decoded digital sound signals, obtains two decoded digital sound signals, and outputs them to the playback unit 223-m (step S2222-35). Additionally, instead of step S2222-32 where the characteristic parameters represented by the extended code are stored in the storage unit, the extended decoding unit 2222-m of the decoding device 222-m in the sound signal receiving side device 220-m of the third embodiment may store the average or weighted average obtained in step S2222-34 as the characteristic parameter of the current frame in the storage unit. Furthermore, since only K characteristic parameters of the past frames need to be stored in the storage unit of the extended decoding unit 2222-m, the characteristic parameters of more than K + 1 past frames can be deleted from the storage unit during the processing of the next frame of the current frame.
[0151] <Fourth Embodiment Variation>
[0152] Similar to the sound signal receiving side device 220-m of the first embodiment, in the sound signal receiving side device 220-m of the third embodiment, as the characteristic parameters used when obtaining two decoded digital sound signals, the average or weighted average of the characteristic parameters represented by the extended code input in the frame to be processed and the characteristic parameters of the past frames can also be used. That is, in the extended decoding unit 2222-m of the decoding device 222-m in the sound signal receiving side device 220-m of the third embodiment, for a predetermined frame among multiple frames, as the characteristic parameters used when obtaining two decoded digital sound signals, the average or weighted average of the characteristic parameters represented by the extended code input in the frame to be processed and the characteristic parameters of the past frames can also be used. This method will be described as a variation of the fourth embodiment.
[0153] The difference between the modified example of the fourth embodiment and the third embodiment lies in the operation of the extended decoding unit 2222-m of the decoding device 222-m of the sound signal receiving side device 220-m. Hereinafter, the differences between the modified example of the fourth embodiment and the third embodiment will be described. Hereinafter, the extended decoding unit 2222-m that processes each frame is referred to as the current frame for the frame being processed at this moment, and the past frames are referred to as past frames.
[0154] [[[Extended decoding unit 2222-m]]]
[0155] Similar to the extended decoding unit 2222-m of the third embodiment, for each frame, a monaural decoded digital sound signal output by the monaural decoding unit 2221-m is input to the extended decoding unit 2222-m, and an extended code is input to the extended decoding unit 2222-m only for a predetermined frame among a plurality of frames. The extended decoding unit 2222-m includes a storage unit (not shown). In the storage unit, at least the average or weighted average of the feature parameters obtained by the extended decoding unit 2222-m in the past frames is stored, and sometimes the feature parameters represented by the extended codes of the past frames are also stored.
[0156] The extended decoding unit 2222-m performs the following steps S2222-41 to step S2222-46 for a predetermined frame among a plurality of frames, that is, a frame for which an extended code is also input.
[0157] The extended decoding unit 2222-m first obtains the characteristic parameters represented by the input extended code according to the input extended code (step S2222-41), and stores the obtained characteristic parameters in the storage unit (step S2222-42). The extended decoding unit 2222-m then reads out K (K is an integer greater than or equal to 1) of the characteristic parameters of the past frames stored in the storage unit (step S2222-43). For example, the characteristic parameters of the K past frames closest to the current frame are read out. Since the characteristic parameters are stored in the storage unit only for the frames for which the extended code is also input, the read-out characteristic parameters are the characteristic parameters of the K frames consecutive to the current frame among the frames for which the extended code is also input. The extended decoding unit 2222-m then obtains the average or weighted average of the characteristic parameters of the K past frames read out from the storage unit and the characteristic parameters of the current frame (step S2222-44), and stores the obtained average or weighted average of the characteristic parameters in the storage unit (step S2222-45). The weight for the weighted average may be set to the maximum value for the characteristic parameters of the current frame and to smaller values for the frames farther from the current frame. The extended decoding unit 2222-m then regards the input mono decoded digital audio signal and the average or weighted average of the characteristic parameters obtained in step S2222-44 as a signal formed by mixing two decoded digital audio signals, regards the average or weighted average of the characteristic parameters obtained in step S2222-44 as information representing the difference between the two decoded digital audio signals, obtains the two decoded digital audio signals, and outputs them to the playback unit 223-m (step S2222-46). In addition, the extended decoding unit 2222-m may not perform step S2222-42 of storing the characteristic parameters represented by the extended code in the storage unit, but may read out, in step S2222-43, the average or weighted average stored in the storage unit in step S2222-45 as the characteristic parameters of the past frames. Further, since only K characteristic parameters of the past frames need to be stored in the storage unit of the extended decoding unit 2222-m, the characteristic parameters of more than K + 1 past frames that are traced back can be deleted from the storage unit during the processing of the next frame of the current frame. Further, only the latest object among the average or weighted average of the characteristic parameters obtained in step S2222-44 needs to be stored in advance in the storage unit of the extended decoding unit 2222-m, so the average or weighted average of the characteristic parameters stored in the storage unit at the time of performing step S2222-45 can be deleted from the storage unit.
[0158] The extended decoding unit 2222-m according to a modification of the fourth embodiment performs the following steps S2222-47 to S2222-48 for frames other than predetermined frames among a plurality of frames, that is, frames for which the extended code is not input.
[0159] The extended decoding unit 2222-m first reads out the average or weighted average of the latest characteristic parameters stored in the storage unit from the storage unit (step S2222-47). The extended decoding unit 2222-m then regards the input decoded digital sound signal of a single channel as a signal formed by mixing two decoded digital sound signals, and regards the average or weighted average of the characteristic parameters obtained in step S2222-47 as information representing the difference between the two decoded digital sound signals, and obtains two decoded digital sound signals and outputs them to the playback unit 223-m (step S2222-48).
[0160] [Effect]
[0161] Although the characteristic parameters are statistically parameters with small temporal variations, they reflect the characteristics of the sound signals of each frame. Therefore, it is rare that they are exactly the same values across multiple frames. In addition, there are sometimes significant differences in values between frames. Thus, in the sound signal receiving side device 220-m, compared with using the characteristic parameters represented by a certain spreading code different from the original spreading code of the frame, by using the average or weighted average of the characteristic parameters represented by multiple spreading codes close in time as in the fourth embodiment and the modification example, etc., it is possible to suppress sudden changes or abnormal sounds between channels in the decoded sound signal.
[0162] <Fifth Embodiment>
[0163] In the first embodiment, the sound signal receiving side device 220-m obtains the decoded digital sound signals of two channels using the mono code and the spreading code with the closest frame number for each frame. However, for a frame without a spreading code within a specific restricted time range from the mono code, it is also possible to decode the mono code and use the obtained decoded digital sound signal as the decoded digital sound signals of two channels. This method will be described as the fifth embodiment.
[0164] The difference between the fifth embodiment and the first embodiment lies in the operations of the receiving unit 221-m and the decoding device 222-m of the sound signal receiving side device 220-m. In addition, in the decoding device 222-m, the part that performs different operations in the fifth embodiment from the first embodiment is the extended decoding unit 2222-m. Hereinafter, the differences between the fifth embodiment and the first embodiment will be described.
[0165] [Receiving Unit 221-m]
[0166] The receiving unit 221-m outputs, for a frame in which the difference in frame numbers between the mono code included in the first code string input from the first communication line 410-m and the extended code closest in frame number to the mono code among the extended codes included in the second code string input from the second communication line 510-m is less than a predetermined value, the mono code included in the first code string input from the first communication line 410-m and the extended code closest in frame number to the mono code among the extended codes included in the second code string input from the second communication line 510-m. For a frame in which the difference in frame numbers is not less than the predetermined value, the receiving unit 221-m outputs the mono code included in the first code string input from the first communication line 410-m. Specifically, the receiving unit 221-m performs the following steps S221-11 to S221-15 for each frame.
[0167] The receiving unit 221-m outputs the mono code included in the first code string input from the first communication line 410-m to the decoding device 222-m (step S221-11). The receiving unit 221-m then obtains the frame number of the mono code output in step S221-11 (step S221-12). The receiving unit 221-m then obtains the extended code included in the second code string closest in frame number to the frame number of the mono code obtained in step S221-12 among the second code strings input from the second communication line 510-m, and the frame number of this extended code (step S221-13). The receiving unit 221-m then determines whether the difference in frame numbers between the mono code obtained in step S221-12 and the frame number of the extended code obtained in step S221-13 is less than a predetermined value (step S221-14). The receiving unit 221-m then, when the difference in frame numbers between the mono code and the extended code in step S221-14 is less than the predetermined value, outputs the extended code to the decoding device 222-m (step S221-15). The receiving unit 221-m does not output the extended code when the difference in frame numbers between the mono code and the extended code in step S221-14 is not less than the predetermined value. That is, the receiving unit 221-m only needs to output the mono code when the difference in frame numbers between the mono code and the extended code in step S221-14 is not less than the predetermined value.
[0168] Here, the predetermined value is a value of 2 or more. That is, the receiving unit 221-m outputs, for a frame in which the difference in frame numbers between the mono code (i.e., the mono code in the order of frame numbers) included in the first code string input from the first communication line 410-m and the extended code included in the second code string input from the second communication line 510-m is 0 (i.e., a frame in which the second code string input from the second communication line 510-m includes an extended code with the same frame number as the mono code included in the first code string input from the first communication line 410-m), the mono code (i.e., the mono code in the order of frame numbers) included in the first code string input from the first communication line 410-m and the extended code with the same frame number as the mono code among the extended codes included in the second code string input from the second communication line 510-m. For a frame in which the difference in frame numbers is greater than 0 and less than the predetermined value, the receiving unit 221-m outputs the mono code (i.e., the mono code in the order of frame numbers) included in the first code string input from the first communication line 410-m and the extended code with the frame number closest to the mono code (i.e., among the extended codes included in the second code string input from the second communication line 510-m, the extended code with a frame number that is not the same as the mono code but is closest to the mono code). For a frame in which the difference in frame numbers is not less than the predetermined value, the receiving unit 221-m outputs only the mono code (i.e., the mono code in the order of frame numbers) included in the first code string input from the first communication line 410-m.
[0169] [[Decoding device 222-m]]
[0170] The decoding device 222-m is input with the mono code output by the receiving unit 221-m at a fixed rate per frame, and sometimes the extended code output by the receiving unit 221-m. The decoding device 222-m obtains, for each frame, two-channel decoded digital audio signals corresponding to the input mono code and extended code or the input mono code, and outputs them to the playback unit 223-m. Specifically, for a frame in which the difference in frame numbers is less than the predetermined value, the decoding device 222-m obtains and outputs two-channel decoded digital audio signals based on the mono code output by the receiving unit 221-m and the extended code output by the receiving unit 221-m. For a frame in which the difference in frame numbers is not less than the predetermined value, the decoding device 222-m outputs the mono digital signal based on the mono code output by the receiving unit 221-m as two-channel decoded digital audio signals as they are.
[0171] [[[Extended decoding unit 2222-m]]]
[0172] The extended decoding unit 2222-m inputs the decoded digital audio signal of the mono channel output from the mono channel decoding unit 2221-m at a certain input per frame, and sometimes inputs the extended code input to the decoding device 222-m. For the frame to which the decoded digital audio signal of the mono channel and the extended code are input, the extended decoding unit 2222-m obtains the decoded digital audio signal of two channels according to the input decoded digital audio signal of the mono channel and the extended code through the same operation as the extended decoding unit 2222-m in the first embodiment, and outputs it to the playback unit 223-m. For the frame to which only the decoded digital audio signal of the mono channel is input, the extended decoding unit 2222-m obtains the input decoded digital audio signal of the mono channel as the decoded digital audio signal of two channels as it is, and outputs it to the playback unit 223-m.
[0173] That is, for the frame in which the difference in frame numbers between the mono code included in the first code string input from the first communication line 410-m and the extended code with the frame number closest to the mono code included in the second code string input from the second communication line 510-m is less than a predetermined value, based on the mono code and the extended code with the frame number closest to the mono code, the decoded digital audio signal of two channels is obtained and output. For the frame in which the difference in frame numbers is not less than the predetermined value, the decoded digital audio signal based on the mono code included in the first code string input from the first communication line 410-m is output as the decoded digital audio signal of two channels as it is.
[0174] More specifically, for a frame in which the difference in frame numbers between the mono code included in the first code string input from the first communication line 410-m (i.e., the mono code in frame number order) and the extended code with the frame number closest to this mono code included in the second code string input from the second communication line 510-m is 0 (i.e., a frame in the second code string input from the second communication line 510-m that includes an extended code with the same frame number as the mono code included in the first code string input from the first communication line 410-m), based on this mono code and the extended code with the same frame number as this mono code, a decoded digital audio signal for two channels is obtained and output. For a frame in which the above-mentioned difference in frame numbers is greater than 0 and less than a predetermined value, based on the mono code included in the first code string input from the first communication line 410-m (i.e., the mono code in frame number order) and the extended code with the frame number closest to this mono code (i.e., among the extended codes included in the second code string input from the second communication line 510-m, the extended code that, although having a different frame number from this mono code, has the closest frame number to this mono code), a decoded digital audio signal for two channels is obtained and output. For a frame in which the above-mentioned difference in frame numbers is not less than the predetermined value, the decoded digital audio signal based on the mono code included in the first code string input from the first communication line 410-m (i.e., the mono code in frame number order) is output as the decoded digital audio signal for two channels.
[0175] <Variation of the Fifth Embodiment>
[0176] The structure and operation of the audio signal receiving side device 220-m of the fifth embodiment based on the structure of the audio signal receiving side device 220-m of the first embodiment have been described above. However, it is also possible to configure and operate the audio signal receiving side device 220-m of the fifth embodiment based on one of the third embodiment, the fourth embodiment, and their variations.
[0177] [Effect]
[0178] The encoding device 212-m' of the voice signal transmitting side device 210-m' of the multi-line support terminal device 200-m' of the call partner encodes in frames for each specific time interval. Therefore, the difference between the frame number of the mono code and the frame number of the extended code corresponds to the time difference of the digital voice signal encoded by the encoding device 212-m' of the voice signal transmitting side device 210-m' of the multi-line support terminal device 200-m' of the call partner. For example, if the frame length is 20 ms and the frame number difference is 150, there is a 3-second time difference between the digital voice signal obtained with the mono code and the digital voice signal obtained with the extended code. Even for a parameter with little temporal variation, if the times are very different, the value may vary greatly. Thus, in the case where there is a time difference to such an extent that the characteristic parameters represented by the extended code are very different, there may be a large error in the segmentation of the signals between channels in the decoded voice signals of the two channels that reflect the characteristics of the difference between the two channels. According to this fifth embodiment, for frames with a large difference in frame numbers between the mono code included in the first code string received from the first communication line and the extended code with the frame number closest to that mono code among the extended codes included in the second code string received from the second communication line, no difference is given to the decoded voice signals of the two channels, and a large error in the segmentation of the signals between channels of the decoded voice signals can be suppressed. For example, if it is assumed that the characteristic parameters are very different when the time difference is 400 ms or more, then when the frame length is 20 ms, if the frame number difference becomes 20 or more, the characteristic parameters are very different, so the above-mentioned predetermined value can be set to 20, for example.
[0179] <Sixth Embodiment>
[0180] The voice signal receiving side device 220-m may also, based on the average value of the time difference between the first code string input from the first communication line 410-m measured within a specific time range and the second code string input from the second communication line 510-m with the same frame number as the first code string, use the decoded digital voice signal obtained by decoding the mono code as the decoded digital voice signals of the two channels when the average value of the time difference is not within a predetermined limit time. This method will be described as the sixth embodiment.
[0181] The difference between the sixth embodiment and the first embodiment lies in the operations of the receiving unit 221-m and the decoding device 222-m of the voice signal receiving side device 220-m. In addition, in the decoding device 222-m, the part that performs different operations in the sixth embodiment from the first embodiment is the extended decoding unit 2222-m. Hereinafter, the differences between the sixth embodiment and the first embodiment will be described.
[0182] [[Receiving Unit 221-m]]
[0183] A first code string output by the voice signal transmitting-side device 210-m' of the calling party is input from the first communication line 410-m to the receiving unit 221-m, and a second code string output by the voice signal transmitting-side device 210-m' of the calling party is input from the second communication line 510-m to the receiving unit 221-m. Since the second communication line is a communication network with a lower priority, generally, the second code string of a certain frame output by the voice signal transmitting-side device 210-m' of the calling party is input from the second communication line 510-m to the receiving unit 221-m after the first code string of the frame is input from the first communication line 410-m to the receiving unit 221-m.
[0184] The receiving unit 221-m first determines whether the difference in the reception times of the first code string and the second code string corresponding to the first code string received from the second communication line 510-m in the group formed by the first code string received from the first communication line 410-m is less than a preset limit time Tmax among multiple groups. Additionally, the limit time Tmax is, for example, 400 ms.
[0185] For example, the receiving unit 221-m performs the following steps S221-21 to S221-24. The receiving unit 221-m reads the frame number for a preset number of first code strings starting from the reception of the first code string, measures the reception time, and stores the frame number and the reception time of the first code string in association with each other in a storage unit (not shown) within the receiving unit 221-m (step S221-21). In addition, for the received second code string, the receiving unit 221-m reads the frame number, and when one of the read frame number and the frame number stored in the storage unit is the same, measures the reception time, and also stores the reception time of the second code string in association with the frame number stored in the storage unit and the reception time of the first code string in the storage unit (step S221-22). The receiving unit 221-m then uses the frame numbers, the reception time of the first code string, and the reception time of the second code string stored in association in the storage unit to obtain the average value of the values obtained by subtracting the reception time of the first code string from the reception time of the second code string for each frame number among the above-mentioned preset number (step S221-23). The receiving unit 221-m then determines whether the average value obtained in step S221-23 is less than the preset limit time Tmax (step S221-24).
[0186] Receiving unit 221-m Then, in the case where the average value is less than the limit time Tmax in the above determination, for the subsequent frames, the monophonic code included in the first code string input from the first communication line 410-m and the extended code included in the second code string input from the second communication line 510-m, among which the extended code with the frame number closest to the monophonic code is output to the decoding device 222-m. In the case where the average value is not less than the limit time Tmax in the above determination, for the subsequent frames, the monophonic code included in the first code string input from the first communication line 410-m is output to the decoding device 222-m. The receiving unit 221-m does not output the extended code for the subsequent frames in the case where the average value is not less than the limit time Tmax in the above determination. That is, the receiving unit 221-m only needs to output the monophonic code in the case where the average value is not less than the limit time Tmax in the above determination.
[0187] That is, for the group formed by the first code string received by the receiving unit 221-m from the first communication line 410-m and the second code string received from the second communication line 510-m corresponding to the first code string, in the case where the difference in the reception times of the first code string and the second code string is less than the predetermined limit time Tmax on average among multiple groups, for the subsequent frames, in the case where the extended code included in the second code string input from the second communication line 510-m includes an extended code with the same frame number as the monophonic code included in the first code string input from the first communication line 410-m (i.e., the monophonic code in the frame number order), the monophonic code and the extended code with the same frame number as the monophonic code are output to the decoding device 222-m. In the case where the extended code included in the second code string input from the second communication line 510-m does not include an extended code with the same frame number as the monophonic code included in the first code string input from the first communication line 410-m (i.e., the monophonic code in the frame number order), among the monophonic code included in the first code string input from the first communication line 410-m (i.e., the monophonic code in the frame number order) and the extended code included in the second code string input from the second communication line 510-m, the extended code with the frame number closest to the monophonic code (i.e., among the extended codes included in the second code string input from the second communication line 510-m, although the frame number is not the same as the monophonic code but is the closest to the monophonic code) is output to the decoding device 222-m. In the case where the above average value is not less than the limit time Tmax, for the subsequent frames, only the monophonic code included in the first code string input from the first communication line 410-m (i.e., the monophonic code in the frame number order) is output to the decoding device 222-m.
[0188] In addition, until the above-mentioned determination is completed, the receiving unit 221-m may output nothing, may output the monaural code and the extended code to the decoding device 222-m in the same manner as in the first embodiment, may output only the monaural code to the decoding device 222-m without outputting the extended code, or, as in the fifth embodiment, always output the monaural code to the decoding device 222-m and output the extended code to the decoding device 222-m only when the difference in frame numbers between the monaural code and the extended code is small.
[0189] [[decoding device 222-m]]
[0190] In the case where the average value in the above-mentioned determination made by the receiving unit 221-m is less than a predetermined limit time Tmax, as in the decoding device 222-m of the first embodiment, the monaural code and the extended code are input to the decoding device 222-m frame by frame. On the other hand, in the case where the average value in the above-mentioned determination made by the receiving unit 221-m is not less than the predetermined limit time Tmax, the monaural code output by the receiving unit 221-m is input to the decoding device 222-m frame by frame, and the extended code is not input.
[0191] In addition, until the above-mentioned determination made by the receiving unit 221-m is completed, nothing is input to the decoding device 222-m, or only the monaural code is input without inputting the extended code, or the monaural code and the extended code are input. The decoding device 222-m obtains two-channel decoded digital audio signals corresponding to the input monaural code and extended code or the input monaural code frame by frame, and outputs them to the playback unit 223-m.
[0192] [[[extended decoding unit 2222-m]]]
[0193] When the extended decoding unit 2222-m is input with a monaural decoded digital audio signal and an extended code, that is, when the average value in the above-mentioned determination is less than a predetermined limit time Tmax, frame by frame, based on the input monaural decoded digital audio signal and extended code, through the same operation as the extended decoding unit 2222-m of the first embodiment, two-channel decoded digital audio signals are obtained and output to the playback unit 223-m. When the extended decoding unit 2222-m is input with a monaural decoded digital audio signal, that is, when the average value in the above-mentioned determination is not less than the predetermined limit time Tmax, the input monaural decoded digital audio signal is obtained as two-channel decoded digital audio signals as it is and output to the playback unit 223-m.
[0194] That is, when the average value of the differences in reception times between the first code string received from the first communication line 410-m and the second code string received from the second communication line 510-m corresponding to the first code string among a plurality of groups is less than a predetermined limit time Tmax, the decoding device 222-m obtains and outputs decoded digital audio signals for two channels based on the monaural code included in the first code string input from the first communication line 410-m and the extended code whose frame number is closest to the monaural code among the extended codes included in the second code string input from the second communication line 510-m. When the above average value is not less than the limit time Tmax, the monaural decoded digital audio signal based on the monaural code included in the first code string input from the first communication line 410-m is output as the decoded digital audio signals for two channels as it is.
[0195] More specifically, for a group consisting of the first code string received from the first communication line 410-m and the second code string received from the second communication line 510-m corresponding to the first code string, when the average value of the differences in reception times between the first code string and the second code string among a plurality of groups is less than a predetermined limit time Tmax, for a frame of the extended code included in the second code string input from the second communication line 510-m whose frame number is the same as the monaural code (i.e., the monaural code in the order of frame numbers) included in the first code string input from the first communication line 410-m, the decoding device 222-m obtains and outputs decoded digital audio signals for two channels based on the monaural code and the extended code whose frame number is the same as the monaural code. For a frame of the extended code included in the second code string input from the second communication line 510-m that does not include an extended code whose frame number is the same as the monaural code (i.e., the monaural code in the order of frame numbers) included in the first code string input from the first communication line 410-m, the decoding device 222-m obtains and outputs decoded digital audio signals for two channels based on the monaural code (i.e., the monaural code in the order of frame numbers) included in the first code string input from the first communication line 410-m and the extended code whose frame number is closest to the monaural code (i.e., among the extended codes included in the second code string input from the second communication line 510-m, the extended code whose frame number is not the same as the monaural code but is closest to the monaural code). When the above average value is not less than the limit time Tmax, the monaural decoded digital audio signal based on the monaural code (i.e., the monaural code in the order of frame numbers) included in the first code string input from the first communication line 410-m is output as the decoded digital audio signals for two channels as it is.
[0196] In addition, until the above-described determination performed by the reception unit 221-m ends, the extended decoding unit 2222-m, with respect to the frame of the decoded digital audio signal and the extended code input in mono, based on the input mono decoded digital audio signal and the extended code, through the same operation as the extended decoding unit 2222-m of the first embodiment, obtains a decoded digital audio signal of two channels and outputs it to the playback unit 223-m, or obtains the input mono decoded digital audio signal as it is as a decoded digital audio signal of two channels and outputs it to the playback unit 223-m, or does not output anything.
[0197] <Modification Example of the Sixth Embodiment>
[0198] The above has described the audio signal reception-side device 220-m of the sixth embodiment and its operation based on the structure of the audio signal reception-side device 220-m of the first embodiment, but it is also possible to configure the audio signal reception-side device 220-m of the sixth embodiment based on one of the third to fifth embodiments and their modification examples and make it operate. Further, in the above example, the period from the start of receiving the first code string until a predetermined number of first code strings are received is used as a specific time range, but the specific time range may have any moment as the starting point. For example, an interval starting from a certain moment after the start of receiving the first code string may be used as the specific time range, or each of the intervals starting from multiple moments after the start of receiving the first code string may be set as the specific time range.
[0199] 〔Effect〕
[0200] As also described in the fifth embodiment, even for characteristic parameters with little temporal variation, if the moments are very different, the values may vary greatly. Thus, when it is determined that there is a time difference between the first communication line and the second communication line such that the characteristic parameters represented by the extended code are greatly different, there may be a large error in the segmentation of the signals between the channels in the decoded audio signals of the two channels that reflect the differences in the characteristics of the two channels. According to this sixth embodiment, by not applying a difference to the decoded audio signals of the two channels when the difference between the moment when the first code string of the same frame is received from the first communication line and the moment when the second code string is received from the second communication line is large, a large error in the segmentation of the signals between the channels of the decoded audio signals can be suppressed.
[0201] <Seventh Embodiment>
[0202] The voice signal receiving side device 220-m may also, based on the average value of the time difference between the first code string input from the first communication line 410-m and measured within a specific time range and the second code string input from the second communication line 510-m having the same frame number as the first code string, and when the average value of the time difference is within a predetermined limit time, use a monaural code and an extended code having the same frame number as the monaural code as the decoded digital voice signals for two channels. This method will be described as the seventh embodiment.
[0203] The difference between the seventh embodiment and the first embodiment lies in the operation of the receiving unit 221-m of the voice signal receiving side device 220-m. Hereinafter, the differences between the seventh embodiment and the first embodiment will be described.
[0204] [[Receiving unit 221-m]]
[0205] A first code string output by the voice signal transmitting side device 210-m' of the calling party is input from the first communication line 410-m to the receiving unit 221-m, and a second code string output by the voice signal transmitting side device 210-m' of the calling party is input from the second communication line 510-m to the receiving unit 221-m. Since the second communication line is a communication network with a lower priority, generally, the second code string of a certain frame output by the voice signal transmitting side device 210-m' of the calling party is input from the second communication line 510-m to the receiving unit 221-m after the first code string of that frame is input from the first communication line 410-m to the receiving unit 221-m.
[0206] The receiving unit 221-m first determines whether the average value of the difference in the reception times of the first code string and the second code string corresponding to the first code string in the group formed by the first code string received from the first communication line 410-m and the second code string received from the second communication line 510-m is less than a predetermined limit time Tmin among multiple groups. Additionally, the limit time Tmin is, for example, a value twice the frame length. That is, if the frame length is 20 ms, the limit time Tmin is, for example, 40 ms.
[0207] For example, the receiving unit 221-m performs the following steps S221-31 to S221-34. The receiving unit 221-m reads out the frame number for a predetermined number of first code strings starting from the reception of the first code string, measures the reception time, and stores the frame number and the reception time of the first code string in an associated manner in a storage unit (not shown) within the receiving unit 221-m (step S221-31). In addition, for the received second code string, the receiving unit 221-m reads out the frame number, measures the reception time when one of the read frame number and the frame number stored in the storage unit is the same, and stores the reception time of the second code string in an associated manner with the frame number stored in the storage unit and the reception time of the first code string in the storage unit (step S221-32). Next, the receiving unit 221-m uses the frame number, the reception time of the first code string, and the reception time of the second code string stored in an associated manner in the storage unit to obtain the average value of the values obtained by subtracting the reception time of the first code string from the reception time of the second code string for each frame number among the above-mentioned predetermined number (step S221-33). Then, the receiving unit 221-m determines whether the average value obtained in step S221-33 is less than a predetermined limit time Tmin (step S221-34).
[0208] Next, when the average value is less than the limit time Tmin in the above determination, for the subsequent frames, the receiving unit 221-m outputs to the decoding device 222-m the extended code whose frame number is the same as that of the monophonic code included in the first code string input from the first communication line 410-m and the extended code included in the second code string input from the second communication line 510-m. When the average value is not less than the limit time Tmin in the above determination, for the subsequent frames, the receiving unit 221-m outputs to the decoding device 222-m the extended code whose frame number is closest to that of the monophonic code among the monophonic code included in the first code string input from the first communication line 410-m and the extended code included in the second code string input from the second communication line 510-m. It is assumed that, on average, it takes the time of the average value obtained in step S221-33 from the reception of the first code string from the first communication line 410-m until the reception of the second code string from the second communication line 510-m for this frame. Therefore, the receiving unit 221-m needs to perform operations so that the time from the reception of the first code string from the first communication line 410-m until the output to the decoding device 222-m becomes the average value obtained in step S221-33 or a value larger than it.
[0209] That is, the receiving unit 221-m, for a group formed by the first code string received from the first communication line 410-m and the second code string received from the second communication line 510-m corresponding to the first code string, when the difference in the reception times of the first code string and the second code string is less than the average value among multiple groups of a preset limit time Tmin, for subsequent frames, the mono code included in the first code string input from the first communication line 410-m (i.e., the mono code in the frame number order), and among the extended codes included in the second code string input from the second communication line 510-m, the extended code with the same frame number as the mono code, are output to the decoding device 222-m. When the above average value is not less than the limit time Tmin, for subsequent frames, when the extended code included in the second code string input from the second communication line 510-m includes an extended code with the same frame number as the mono code (i.e., the mono code in the frame number order) included in the first code string input from the first communication line 410-m, the mono code and the extended code with the same frame number as the mono code are output to the decoding device 222-m. When the extended code included in the second code string input from the second communication line 510-m does not include an extended code with the same frame number as the mono code (i.e., the mono code in the frame number order) included in the first code string input from the first communication line 410-m, the mono code included in the first code string (i.e., the mono code in the frame number order), and among the extended codes included in the second code string input from the second communication line 510-m, the extended code with the frame number closest to the mono code (i.e., among the extended codes included in the second code string input from the second communication line 510-m, the extended code whose frame number is not the same as the mono code but is closest to the mono code), are output to the decoding device 222-m.
[0210] The operation of the decoding device 222-m of the sound signal receiving side device 220-m in the seventh embodiment is the same as the operation of the decoding device 222-m of the sound signal receiving side device 220-m in the first embodiment. The decoding device 222-m obtains and outputs the decoded digital sound signals of two channels based on the mono code output by the receiving unit 221-m and the extended code output by the receiving unit 221-m. Among them, the extended code output by the receiving unit 221-m in the seventh embodiment may be different from the extended code output by the receiving unit 221-m in the first embodiment according to the situation. Therefore, the decoding device 222-m specifically performs the following operations.
[0211] That is, the decoding device 222-m, for a group formed by a first code string received from the first communication line 410-m and a second code string received from the second communication line 510-m corresponding to the first code string, when the difference in reception times between the first code string and the second code string is less than the average value among multiple groups of a predetermined limit time Tmin, based on the monaural code included in the first code string input from the first communication line 410-m and the extended code with the same frame number as the monaural code included in the second code string input from the second communication line 510-m, obtains and outputs the decoded digital audio signals for two channels. When the above average value is not less than the limit time Tmin, based on the monaural code included in the first code string input from the first communication line 410-m and the extended code with the frame number closest to the monaural code among the extended codes included in the second code string input from the second communication line 510-m, obtains and outputs the decoded digital audio signals for two channels.
[0212] More specifically, the decoding device 222-m, for a group formed by a first code string received from the first communication line 410-m and a second code string received from the second communication line 510-m corresponding to the first code string, when the difference in reception times between the first code string and the second code string is less than the average value among multiple groups of a predetermined limit time Tmin, based on the monaural code (i.e., the monaural code in frame number order) included in the first code string input from the first communication line 410-m and the extended code with the same frame number as the monaural code included in the second code string input from the second communication line 510-m, obtains and outputs the decoded digital audio signals for two channels. When the above average value is not less than the limit time Tmin, for frames in which the extended code included in the second code string input from the second communication line 510-m includes an extended code with the same frame number as the monaural code (i.e., the monaural code in frame number order) included in the first code string input from the first communication line 410-m, based on the monaural code and the extended code with the same frame number as the monaural code, obtains and outputs the decoded digital audio signals for two channels. For frames in which the extended code included in the second code string input from the second communication line 510-m does not include an extended code with the same frame number as the monaural code (i.e., the monaural code in frame number order) included in the first code string input from the first communication line 410-m, based on the monaural code (i.e., the monaural code in frame number order) included in the first code string input from the first communication line 410-m and the extended code with the frame number closest to the monaural code (i.e., among the extended codes included in the second code string input from the second communication line 510-m, the extended code whose frame number is not the same as the monaural code but is the closest to the monaural code), obtains and outputs the decoded digital audio signals for two channels.
[0213] In addition, until the above-described determination performed by the receiving unit 221-m is completed, for example, the receiving unit 221-m may output the monaural code and the extension code to the decoding device 222-m in the same manner as in the first embodiment, and the decoding device 222-m may obtain the decoded digital audio signal of two channels using the monaural code and the extension code and output it to the playback unit 223-m in the same manner as in the first embodiment.
[0214] <Modification Example of the Seventh Embodiment>
[0215] The above has described the audio signal receiving side device 220-m of the seventh embodiment and its operation based on the structure of the audio signal receiving side device 220-m of the first embodiment. However, it is also possible to configure and operate the audio signal receiving side device 220-m of the seventh embodiment based on one of the third to fifth embodiments and their modification examples. In addition, in the above example, the period from when the first code string starts to be received until a predetermined number of first code strings are received is used as a specific time range. However, the specific time range may have any moment as a starting point. For example, an interval starting from a certain moment after the first code string starts to be received may be used as the specific time range, or each interval starting from multiple moments after the first code string starts to be received may be set as the specific time range.
[0216] 〔Effect〕
[0217] Even for characteristic parameters with little temporal variation, if the moments are different, the values may be slightly different. Thus, if it is possible to perform decoding using the characteristic parameters of the same frame by slightly increasing the delay, it may be possible to obtain a decoded audio signal with high sound quality. Therefore, in the seventh embodiment, a limit time, which is a predetermined value, is set for the average value of a specific time range of the difference between the moment when the first code string of the same frame is received from the first communication line and the moment when the second code string is received from the second communication line. When it is less than the limit time, on the basis of deliberately increasing the delay slightly, the monaural code and the extension code of the same frame as the monaural code are used as the decoded digital audio signal of two channels, whereby a decoded digital audio signal with high sound quality can be obtained.
[0218] <Eighth Embodiment>
[0219] The sound signal receiving side device 220-m may also be based on the average value of the time difference between the first code string input from the first communication line 410-m measured within a specific time range and the second code string input from the second communication line 510-m having the same frame number as the first code string. When the average value of the time difference is less than the first limit time, a monaural code and an extended code having the same frame number as the monaural code are used to obtain decoded digital sound signals for two channels. When the average value of the time difference is equal to or greater than a predetermined second limit time which is greater than the first limit time, the decoded digital sound signal obtained by decoding the monaural code is used as the decoded digital sound signals for two channels. When the average value of the time difference is equal to or greater than the first limit time and less than the second limit time, a monaural code and an extended code closest to the monaural code in terms of frame number are used to obtain decoded digital sound signals for two channels. In short, the sixth embodiment and the seventh embodiment may also be combined and implemented. This method will be described as the eighth embodiment.
[0220] The difference between the eighth embodiment and the first embodiment lies in the operations of the receiving unit 221-m and the decoding device 222-m of the sound signal receiving side device 220-m. Among them, the operation of the decoding device 222-m of the sound signal receiving side device 220-m is the same as that of the decoding device 222-m in the sixth embodiment. Hereinafter, the operation of the receiving unit 221-m, which is different from both the first embodiment and the sixth embodiment in the eighth embodiment, will be described.
[0221] [[Receiving unit 221-m]]
[0222] A first code string output by the voice signal transmitting side device 210-m' of the calling party is input from the first communication line 410-m to the receiving unit 221-m, and a second code string output by the voice signal transmitting side device 210-m' of the calling party is input from the second communication line 510-m to the receiving unit 221-m. Since the second communication line is a communication network with a lower priority, generally, the second code string of a certain frame output by the voice signal transmitting side device 210-m' of the calling party is input to the receiving unit 221-m from the second communication line 510-m after the first code string of the frame is input to the receiving unit 221-m from the first communication line 410-m.
[0223] The receiving unit 221-m first determines, for a group formed by a first code string received from the first communication line 410-m and a second code string received from the second communication line 510-m corresponding to the first code string, whether the difference in the reception times of the first code string and the second code string is less than a first predetermined limit time Tmin, which is a first predetermined value, greater than or equal to a second predetermined limit time Tmax, which is a second predetermined value greater than the first limit time Tmin, or greater than or equal to the first limit time Tmin and less than the second limit time Tmax. Further, the first limit time Tmin is, for example, a value that is twice the frame length. That is, if the frame length is 20 ms, the first limit time Tmin is, for example, 40 ms. In addition, the second limit time Tmax is, for example, 400 ms.
[0224] For example, the receiving unit 221-m performs the following steps S221-41 to S221-44. The receiving unit 221-m reads out the frame number for a predetermined number of first code strings from the start of reception of the first code string, measures the reception time, and stores the frame number and the reception time of the first code string in association with each other in a storage unit (not shown) within the receiving unit 221-m (step S221-41). Further, for the received second code string, the receiving unit 221-m reads out the frame number, and when one of the read frame number and the frame number stored in the storage unit matches, measures the reception time, and also stores the reception time of the second code string in association with the frame number stored in the storage unit and the reception time of the first code string in the storage unit (step S221-42). Next, the receiving unit 221-m uses the frame number, the reception time of the first code string, and the reception time of the second code string stored in association in the storage unit to obtain an average value of the values obtained by subtracting the reception time of the first code string from the reception time of the second code string for each frame number among the above-mentioned predetermined number (step S221-43). The receiving unit 221-m then determines whether the average value obtained in step S221-43 is less than the first predetermined limit time Tmin, greater than or equal to the second predetermined limit time Tmax, which is greater than the first limit time Tmin, or greater than or equal to the first limit time Tmin and less than the second limit time Tmax (step S221-44).
[0225] When the average value is less than the first limit time Tmin in the above determination, the receiving unit 221-m outputs, for subsequent frames, the extended code whose frame number is the same as that of the monophonic code included in the first code string input from the first communication line 410-m and the extended code included in the second code string input from the second communication line 510-m, among the extended codes included in the second code string input from the second communication line 510-m. When the average value is equal to or greater than the first limit time Tmin and less than the second limit time Tmax in the above determination, the receiving unit 221-m outputs, for subsequent frames, the extended code whose frame number is closest to that of the monophonic code, among the monophonic code included in the first code string input from the first communication line 410-m and the extended code included in the second code string input from the second communication line 510-m, to the decoding device 222-m. When the average value is not less than the second limit time Tmax in the above determination, the receiving unit 221-m outputs, for subsequent frames, the monophonic code included in the first code string input from the first communication line 410-m to the decoding device 222-m. When the average value is not less than the second limit time Tmax in the above determination, the receiving unit 221-m does not output the extended code for subsequent frames. That is, when the average value is not less than the second limit time Tmax in the above determination, the receiving unit 221-m only needs to output the monophonic code. Here, it is assumed that the time from receiving the first code string from the first communication line until receiving the second code string from the second communication line for this frame is the average value obtained in step S221-43. Therefore, the receiving unit 221-m needs to operate so that the time from receiving the first code string from the first communication line until outputting to the decoding device 222-m becomes the average value obtained in step S221-43 or a value greater than it.
[0226] That is, the receiving unit 221-m, for a group of a first code string received from the first communication line 410-m and a second code string received from the second communication line 510-m corresponding to the first code string, when the difference in the reception times of the first code string and the second code string is less than the average value among multiple groups of a predetermined limit time Tmin, for subsequent frames, the mono code (i.e., the mono code in the order of frame numbers) included in the first code string input from the first communication line 410-m and the extension code included in the second code string input from the second communication line 510-m, and among the extension codes, the extension code with the same frame number as the mono code, are output to the decoding device 222-m. When the above average value is equal to or greater than the first limit time Tmin and less than the second limit time Tmax, for subsequent frames, when the extension code included in the second code string input from the second communication line 510-m includes an extension code with the same frame number as the mono code (i.e., the mono code in the order of frame numbers) included in the first code string input from the first communication line 410-m, the mono code and the extension code with the same frame number as the mono code are output to the decoding device 222-m. When the extension code included in the second code string input from the second communication line 510-m does not include an extension code with the same frame number as the mono code (i.e., the mono code in the order of frame numbers) included in the first code string input from the first communication line 410-m, the mono code (i.e., the mono code in the order of frame numbers) included in the first code string input from the first communication line 410-m and the extension code included in the second code string input from the second communication line 510-m, and among the extension codes, the extension code with the frame number closest to the mono code (i.e., among the extension codes included in the second code string input from the second communication line 510-m, the extension code whose frame number is not the same as the mono code but is the closest to the mono code) are output to the decoding device 222-m. When the above average value is not less than the second limit time Tmax, for subsequent frames, only the mono code (i.e., the mono code in the order of frame numbers) included in the first code string input from the first communication line 410-m is output to the decoding device 222-m.
[0227] In addition, until the above determination is completed, the receiving unit 221-m may output nothing, may output the mono code and the extension code to the decoding device 222-m in the same manner as in the first embodiment, may output only the mono code to the decoding device 222-m without outputting the extension code, or may, in the same manner as in the fifth embodiment, always output the mono code to the decoding device 222-m and output the extension code to the decoding device 222-m only when the difference in the frame numbers of the mono code and the extension code is small.
[0228] The operation of the decoding device 222-m of the sound signal receiving side device 220-m in the eighth embodiment is the same as that of the decoding device 222-m of the sound signal receiving side device 220-m in the sixth embodiment. Among them, the extended code output by the receiving unit 221-m in the eighth embodiment may be different from the extended code output by the receiving unit 221-m in the sixth embodiment depending on the situation. Therefore, the decoding device 222-m specifically performs the following operations.
[0229] That is, when the average value in the above determination is less than the first limit time Tmin, and when the average value in the above determination is equal to or greater than the first limit time Tmin and less than the second limit time Tmax, for the subsequent frames, based on the monaural code output by the receiving unit 221-m and the extended code output by the receiving unit 221-m, two-channel decoded digital sound signals are obtained and output. When the average value in the above determination is equal to or greater than the second limit time Tmax, for the subsequent frames, the monaural decoded digital sound signal based on the monaural code output by the receiving unit 221-m is output as two-channel decoded digital sound signals as it is.
[0230] More specifically, for a group consisting of the first code string received from the first communication line 410-m and the second code string received from the second communication line 510-m corresponding to the first code string, when the difference in the reception times of the first code string and the second code string is less than a predetermined first limit time Tmin among multiple groups, based on the monaural code included in the first code string input from the first communication line 410-m and the extended code with the same frame number as the monaural code included in the second code string input from the second communication line 510-m, two-channel decoded digital sound signals are obtained and output. When the above average value is equal to or greater than a predetermined second limit time Tmax greater than the first limit time Tmin, the monaural decoded digital sound signal based on the monaural code included in the first code string input from the first communication line 410-m is output as two-channel decoded digital sound signals as it is. When the above average value is equal to or greater than the first limit time Tmin and less than the second limit time Tmax, based on the monaural code included in the first code string input from the first communication line 410-m and the extended code with the frame number closest to the monaural code included in the second code string input from the second communication line 510-m, two-channel decoded digital sound signals are obtained and output.
[0231] More specifically, the decoding device 222-m, for a group consisting of the first code string received from the first communication line 410-m and the second code string received from the second communication line 510-m corresponding to the first code string, when the difference in the reception times of the first code string and the second code string is less than the average value among multiple groups of a predetermined first limit time Tmin, based on the monaural code included in the first code string input from the first communication line 410-m (i.e., the monaural code in the frame number order), and the extended code with the same frame number as the monaural code included in the second code string input from the second communication line 510-m, obtains and outputs the decoded digital audio signals for two channels. When the above average value is greater than or equal to a predetermined second limit time Tmax that is larger than the first limit time Tmin, the monaural decoded digital audio signal based on the monaural code included in the first code string input from the first communication line 410-m (i.e., the monaural code in the frame number order) is output as the decoded digital audio signals for two channels as it is. When the above average value is greater than or equal to the first limit time Tmin and less than the second limit time Tmax, for the frames in the extended code included in the second code string input from the second communication line 510-m that include the extended code with the same frame number as the monaural code included in the first code string input from the first communication line 410-m (i.e., the monaural code in the frame number order), based on the monaural code and the extended code with the same frame number as the monaural code, obtains and outputs the decoded digital audio signals for two channels. For the frames in the extended code included in the second code string input from the second communication line 510-m that do not include the extended code with the same frame number as the monaural code included in the first code string input from the first communication line 410-m (i.e., the monaural code in the frame number order), based on the monaural code included in the first code string input from the first communication line 410-m (i.e., the monaural code in the frame number order), and the extended code with the frame number closest to the monaural code included in the second code string input from the second communication line 510-m (i.e., among the extended codes included in the second code string input from the second communication line 510-m, the extended code whose frame number is not the same as the monaural code but is the closest to the monaural code), obtains and outputs the decoded digital audio signals for two channels.
[0232] In addition, until the above determination performed by the receiving unit 221-m ends, nothing is input to the decoding device 222-m, or a monaural code is input without inputting an extended code, or a monaural code and an extended code are input. The decoding device 222-m obtains the decoded digital audio signals for two channels corresponding to the input monaural code and extended code or the input monaural code for each frame, and outputs them to the playback unit 223-m.
[0233] <Modification Example of the Eighth Embodiment>
[0234] The above has described the voice signal receiving side device 220-m of the eighth embodiment based on the structure of the voice signal receiving side device 220-m of the first embodiment and its operation. However, it is also possible to configure and operate the voice signal receiving side device 220-m of the eighth embodiment based on one of the third to fifth embodiments and their modified examples. In addition, in the above example, the period from the start of receiving the first code string until a predetermined number of first code strings are received is used as a specific time range. However, the specific time range can also be set using any moment. For example, an interval starting from a certain moment after the start of receiving the first code string can be used as the specific time range, or each interval starting from multiple moments after the start of receiving the first code string can be set as the specific time range.
[0235] 〔Effect〕
[0236] According to this eighth embodiment, when the difference between the moment when the first code string of the same frame is received from the first communication line and the moment when the second code string is received from the second communication line is large, a large error in the segmentation of the signals between the channels of the decoded voice signal is suppressed, and when the above difference is small, a high-quality decoded voice signal can be obtained.
[0237] <Ninth Embodiment>
[0238] In a multi-site control device (Multipoint Control Unit (MCU)) for a telephone conference at multiple sites, digital voice signals respectively corresponding to the voice signals of two different sites can also be used as digital voice signals of two channels, and the same operations as those of the voice signal transmitting side device 210-m of the above embodiments can be performed. This method will be described as the ninth embodiment.
[0239] <Multi-site control device 600>
[0240] The multi-site control device 600 is as Figure 7 shown, and includes a receiving unit 610, a monaural decoding unit 620, a location selection unit 630, a signal analysis unit 640, a monaural encoding unit 650, and a transmitting unit 660. Hereinafter, an example will be described in which the multi-site control device 600 is connected to terminal devices at P sites (P is an integer of 3 or more), and transmits the voice signals of up to two of the P-1 sites from location m2 to location m P to the multi-line support terminal device 200-m1. The multi-site control device 600 performs, for example, for each frame as a specific time interval of 20 ms, Figure 8 and the processes of step S610 to step S660 exemplified below.
[0241] [Receiving unit 610]
[0242] Input P - 1 first code strings output via the first communication line to the receiving unit 610 by the multi - line support terminal device 200 - m else (where else is each integer from 2 to P). The receiving unit 610 outputs the mono - channel codes respectively included in the P - 1 first code strings to the mono - channel decoding unit 620 (step S610).
[0243] [Mono - channel decoding unit 620]
[0244] The mono - channel decoding unit 620 decodes the P - 1 mono - channel codes input from the receiving unit 610 respectively in a specific decoding manner to obtain decoded mono - channel signals which are decoded digital sound signals as mono - channels and outputs them to the location selection unit 630 (step S620). Regarding the specific decoding manner, as described in the first embodiment.
[0245] [Location selection unit 630]
[0246] The location selection unit 630 selects two decoded mono - channel signals from among the P - 1 decoded mono - channel signals input from the mono - channel decoding unit 620 based on a pre - determined selection criterion and outputs them to the signal analysis unit 640 (step S630). As the pre - determined selection criterion, as long as a criterion for selecting decoded mono - channel signals of locations with high importance can be pre - determined and the location selection unit 630 can perform the selection. For example, if the power of the sound signal is used as the selection criterion, the location selection unit 630 outputs, for each frame, the decoded mono - channel signal with the largest power and the decoded mono - channel signal with the second - largest power among the P - 1 decoded mono - channel signals input to the signal analysis unit 640.
[0247] [Signal analysis unit 640]
[0248] The signal analysis unit 640 obtains a mono signal, which is a signal obtained by mixing the two input decoded mono signals, and outputs it to the mono encoding unit 650. The signal analysis unit 640 also obtains an extended code representing characteristic parameters and outputs it to the transmission unit 660. The characteristic parameters are parameters that represent the difference between the two input decoded mono signals and have little temporal variation (step S640). The signal analysis unit 640 may perform the same operations as the signal analysis unit 2121-m of the encoding device 212-m on the sound signal transmission side device 210-m of the multi-line support terminal device 200-m in the first embodiment. In the case of this ninth embodiment, since the two input decoded mono signals correspond to sound signals emitted at different locations, as characteristic parameters, it is better to use the information representing the intensity difference for each frequency band shown in the second example than the information representing the time difference shown in the first example of the signal analysis unit 2121-m. Additionally, information representing the power ratio or difference between the two input decoded mono signals may also be used as characteristic parameters.
[0249] [Mono Encoding Unit 650]
[0250] The mono encoding unit 650 encodes the input mono signal in a specific encoding manner to obtain a mono code and outputs it to the transmission unit 660 (step S650). The specific encoding manner is as described in the first embodiment.
[0251] [Transmission Unit 660]
[0252] The transmission unit 660 outputs, for each frame, a code string containing the mono code input from the mono encoding unit 650, i.e., the first code string, to the multi-line support terminal device 200-m1 via the first communication line, and outputs a code string containing the extended code input from the signal analysis unit 640, i.e., the second code string, to the multi-line support terminal device 200-m1 via the second communication line (step S660).
[0253] [Effect]
[0254] By causing the multi-location control device 600 to perform the operations of this ninth embodiment, in the multi-line support terminal device 200-m1, it is possible to virtually assign the sound signals from two locations to the left and right and play them, and it is possible to clarify which location the speech is from or whether it is a speech from different locations.
[0255] <Variation Example of the Ninth Embodiment>
[0256] In the site selection unit 630 of the multi-site control device 600 of the ninth embodiment, since the power is used to select two decoded mono signals, the spreading code may be obtained by the site selection unit 630 instead of the signal analysis unit 640. This method is described as a modified example of the ninth embodiment, and the differences from the ninth embodiment are described.
[0257] <Multi-location control device 600>
[0258] A multi-location control device 600 according to a modified example of the ninth embodiment is as follows: Figure 9 As shown in FIG. 1 , the multi-site control device 600 includes a signal mixing unit 670 instead of the signal analyzing unit 640 included in the ninth embodiment. The multi-site control device 600 performs the signal mixing operation for each frame. Figure 10 The illustrated processing is steps S610 to S630, step S670, and step S650 to S660. Among them, the essential difference from the ninth embodiment is step S630 performed by the site selection unit 630 and step S670 performed by the signal mixing unit 670. Step S660 performed by the transmission unit 660 is the same as the ninth embodiment except that the spread code is input from the site selection unit 630 instead of the signal analysis unit 640.
[0259] [Location selection unit 630]
[0260] The location selection unit 630 selects the decoded mono signal with the largest power and the decoded mono signal with the second largest power from the P-1 decoded mono signals input from the mono decoding unit 620, and outputs them to the signal analysis unit 640. Furthermore, the ratio or difference between the powers of the two selected decoded mono signals is obtained as a characteristic parameter, and an extended code as a code representing the obtained characteristic parameter is obtained and output to the sending unit 660 (step S630).
[0261] [Signal mixing unit 670]
[0262] Signal mixing section 670 obtains a monaural signal that is a signal obtained by mixing the two input decoded monaural signals, based on the two input decoded monaural signals, and outputs the signal to monaural encoding section 650 (step S670).
[0263] In addition, in order to emphasize the virtual left and right distribution of the voice signals at two locations in the multi-line support terminal device 200-m1, the location selection unit 630 may also obtain, as a characteristic parameter, information for determining the location of the side with the greater power among the two selected decoded mono signals, obtain an extended code that is a code representing the obtained characteristic parameter, and output it to the transmission unit 660. In this case, in the extended decoding unit 2222-m1 of the decoding device 222-m1 of the voice signal receiving side device 220-m1 of the multi-line support terminal device 200-m1, it is only necessary to obtain the decoded digital voice signals of two channels so that the voice signals are positioned at the left and right positions predetermined for each location. In addition, in this case, the signal mixing unit 670 may either select the side with the greater power among the two input decoded mono signals and output it to the mono encoding unit 650, or may not have a signal mixing unit 670 at all, and the location selection unit 630 may only select and output the decoded mono signal with the greatest power.
[0264] <Tenth Embodiment>
[0265] In the above-described embodiments and modification examples, for the sake of simplicity of explanation, an example of processing the voice signals of two channels of the multi-line support terminal device 200-m has been described. However, the number of channels is not limited to this, and any number of 2 or more is acceptable. If the number of channels is set to C (C is an integer of 2 or more), then the above-described embodiments and modification examples can be implemented by replacing two channels with C (C is an integer of 2 or more) channels.
[0266] For example, it is only necessary for the sound collection unit 211-m of the voice signal transmission side device 210-m of the multi-line support terminal device 200-m to include C microphones and C AD conversion units, and it is only necessary for the encoding device 212-m of the voice signal transmission side device 210-m of the multi-line support terminal device 200-m to obtain a mono code and an extended code based on the input digital voice signals of C channels. Specifically, the encoding device 212-m encodes the signal obtained by mixing the input digital voice signals of C channels in a specific first encoding method to obtain a mono code, and it is only necessary to obtain an extended code that includes a code representing the following information, which is information equivalent to the difference between channels of the input digital voice signals of C channels. The information equivalent to the difference between channels of the digital voice signals of C channels is, for example, information equivalent to the difference between the digital voice signal of each of the C-1 channels other than the reference channel and the digital voice signal of the reference channel.
[0267] In addition, the decoding device 222-m of the sound signal receiving side device 220-m of the multi-line support device 200-m may obtain and output the decoded digital sound signals of C channels based on the input mono code and the extension code. Specifically, the mono decoding unit 2221-m of the decoding device 222-m decodes the input mono code to obtain the decoded digital sound signal of mono, and the extension decoding unit 2222-m of the decoding device 222-m regards the decoded digital sound signal of mono as a signal formed by mixing the decoded digital sound signals of C channels, regards the characteristic parameters obtained based on the input extension code as the information representing the characteristics of the difference between channels in the decoded digital sound signals of C channels, and may obtain and output the decoded digital sound signals of C channels. In addition, in this case, the playback unit 223-m of the sound signal receiving side device 220-m of the multi-line terminal device 200-m may also include a maximum of C DA conversion units and a maximum of C speakers.
[0268] <Other Embodiments>
[0269] <<The Mode of Also Including a Telephone Line Dedicated Terminal Device in the Telephone System>>
[0270] When the telephone line dedicated terminal device 300-n is also included in the telephone system 100, the telephone line dedicated terminal device 300-n performs the following well-known operations.
[0271] <Telephone Line Dedicated Terminal Device 300-n>
[0272] The telephone line dedicated terminal device 300-n is, for example, an existing type of mobile phone or an existing type of smart phone, and as Figure 11 shown, includes a sound signal transmitting side device 310-n and a sound signal receiving side device 320-n. The sound signal transmitting side device 310-n includes a sound collecting unit 311-n, an encoding device 312-n, and a transmitting unit 313-n. The sound signal receiving side device 320-n includes a receiving unit 321-n, a decoding device 322-n, and a playback unit 323-n. The sound signal transmitting side device 310-n of the telephone line dedicated terminal device 300-n performs Figure 12 and the processing of steps S311 to S313 exemplified below, and the sound signal receiving side device 320-n of the telephone line dedicated terminal device 300-n performs Figure 13 and the processing of steps S321 to S323 exemplified below.
[0273] [Sound Signal Transmitting Side Device 310-n]
[0274] The sound signal transmitting side device 310-n obtains, for example, a code string, i.e., a first code string, including a monophonic code corresponding to a digital sound signal of one channel for each specific time interval of 20 ms, i.e., for each frame, and outputs it to the first communication line 420-n.
[0275] [[Receiving unit 311-n]]
[0276] The receiving unit 311-n includes one microphone and one AD conversion unit. The microphone picks up the sound generated in the spatial domain around the microphone, converts it into an analog electrical signal, and outputs it to the AD conversion unit. The AD conversion unit converts the input analog electrical signal into a digital sound signal, such as a PCM signal with a sampling frequency of 8 kHz, and outputs it. That is, the receiving unit 311-n outputs a digital sound signal of one channel corresponding to the sound picked up by one microphone to the encoding device 312-n (step S311).
[0277] [[Encoding device 312-n]]
[0278] The encoding device 312-n encodes the digital sound signal of one channel input from the receiving unit 311-n for each frame in the above-mentioned specific encoding method to obtain a monophonic code and outputs it to the transmitting unit 313-n (step S312).
[0279] [[Transmitting unit 313-n]]
[0280] The transmitting unit 313-n outputs, for each frame, a code string including the monophonic code input from the encoding device 312-n, i.e., the first code string, to the first communication line 420-n (step S313).
[0281] [Sound signal receiving side device 320-n]
[0282] The sound signal receiving side device 320-n outputs, for example, a sound based on the monophonic code included in the first code string input from the first communication line 420-n for each specific time interval of 20 ms, i.e., for each frame.
[0283] [[Receiving unit 321-n]]
[0284] The receiving unit 321-n outputs, for each frame, the monophonic code included in the first code string input from the first communication line 420-n to the decoding device 322-n (step S321).
[0285] [[Decoding device 322-n]]
[0286] Input the mono code output by the receiving unit 321-n to the decoding device 322-n on a per-frame basis. The decoding device 322-n decodes the input mono code in the above-mentioned specific decoding manner on a per-frame basis, obtains one decoded digital audio signal, and outputs it to the playback unit 323-n (step S322).
[0287] [[Playback unit 323-n]]
[0288] The playback unit 323-n outputs the sound corresponding to the one decoded digital audio signal input thereto (step S323).
[0289] The playback unit 323-n includes, for example, one DA conversion unit and one speaker. The DA conversion unit converts the input decoded digital audio signal into an analog electrical signal and outputs it. The speaker generates the sound corresponding to the analog electrical signal input from the DA conversion unit. The speaker may also be configured in a stereo headset or stereo headphones. In the case of using the speakers provided in the stereo headset or stereo headphones, that is, two speakers, for example, the playback unit 323-n inputs the electrical signal output by the DA conversion unit to the two speakers, and the sound corresponding to the one decoded digital audio signal (decoded audio signal) is generated from the two speakers.
[0290] 〔Effect〕
[0291] In the dedicated telephone line terminal device 300-n, the same encoding method and decoding method as those of the multi-line support terminal device 200-m are also used. Therefore, in the dedicated telephone line terminal device 300-n, compatibility is ensured so that a decoded audio signal with a minimum sound quality can be obtained. On the basis of this, in the multi-line support terminal device 200-m, a decoded audio signal with high sound quality can be obtained at almost the same delay time as when obtaining a decoded audio signal with a minimum sound quality, that is, a delay time that does not cause a sense of disharmony during two-way calls.
[0292] <<There is also a code format that is neither a mono code nor an extended code>>
[0293] The sound signal transmitting side device 210-m of the multi-line support terminal device 200-m can also obtain a code (additional code) that is neither the above-mentioned mono code nor the above-mentioned extended code and output it. Specifically, it may also be that the encoding device 212-m further obtains the additional code and outputs it to the transmitting unit 213-m, and the transmitting unit 213-m outputs the additional code input from the encoding device 212-m to one of the first communication line 410-m and the second communication line 510-m. The additional code is, for example, a code representing the characteristics of the high-band component of the signal formed by mixing the digital audio signals of C (C is an integer of 2 or more) input channels.
[0294] Similarly, a code (additional code) that is neither the above-mentioned monaural code nor the above-mentioned extended code can also be input to the sound signal receiving side device 220-m of the multi-line support terminal device 200-m. The sound signal receiving side device 220-m of the multi-line support terminal device 200-m also uses the additional code to obtain a decoded sound signal and output it. Specifically, it can also be that the receiving unit 221-m outputs the additional code input from one of the first communication line 410-m and the second communication line 510-m to the decoding device 222-m, and the decoding device 222-m also uses the additional code input from the receiving unit 221-m to obtain a decoded sound signal.
[0295] <Program and Recording Medium>
[0296] The processing of each unit of the multi-line support terminal device 200-m can also be implemented by a computer. In other words, the processing of each step of the encoding method in the multi-line support terminal device 200-m and the decoding method in the multi-line support terminal device 200-m can also be executed by a computer. In this case, the processing of each step is described by a program. And by the computer executing this program, the processing of each step is implemented on the computer. Figure 14 It is a diagram showing an example of the functional structure of a computer for implementing the above-mentioned processing. This processing can be implemented by causing the recording unit 2020 to read in a program for causing the computer to function as the above-mentioned device, and causing the control unit 2010, the input unit 2030, the output unit 2040, etc. to operate.
[0297] The programs describing these processing contents can each be previously recorded on a recording medium readable by a computer. As a recording medium readable by a computer, for example, it can also be any medium such as a magnetic recording device, an optical disc, a magneto-optical recording medium, a semiconductor memory, etc.
[0298] In addition, the processing of each unit can be constituted by executing a specific program on a computer, or at least a part of these processes can be implemented in hardware.
[0299] Obviously, other changes can also be appropriately made without departing from the spirit of the present invention.
Claims
1. A method for receiving and decoding a voice signal, which is performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, includes: a receiving step of determining, for a group consisting of a first code string received from the first communication line and a second code string received from the second communication line corresponding to the first code string, whether an average value of differences in reception times of the first code string and the second code string among multiple groups is less than a predetermined limit time Tmin; in a case where the average value is less than the limit time Tmin in the determination, for frames after the determination, outputting an extended code having the same frame number as a monaural code included in the first code string input from the first communication line and an extended code included in the second code string input from the second communication line; in a case where the average value is not less than the limit time Tmin in the determination, for frames after the determination, outputting an extended code having a frame number closest to the monaural code among the monaural code included in the first code string input from the first communication line and the extended code included in the second code string input from the second communication line; and a decoding step of, for frames after the determination, obtaining and outputting a decoded digital voice signal of C channels based on the monaural code output in the receiving step and the extended code output in the receiving step, where C is an integer of 2 or more.
2. A method for receiving and decoding a voice signal, which is performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, includes: a receiving step of determining, for a group consisting of a first code string received from the first communication line and a second code string received from the second communication line corresponding to the first code string, whether an average value of differences in reception times of the first code string and the second code string among multiple groups is less than a predetermined limit time Tmax; in a case where the average value is less than the limit time Tmax in the determination, for frames after the determination, outputting an extended code having a frame number closest to the monaural code among the monaural code included in the first code string input from the first communication line and the extended code included in the second code string input from the second communication line; in a case where the average value is not less than the limit time Tmax in the determination, for frames after the determination, outputting the monaural code included in the first code string input from the first communication line; and a decoding step of, in a case where the average value is less than the limit time Tmax in the determination, for frames after the determination, obtaining and outputting a decoded digital voice signal of C channels based on the monaural code output in the receiving step and the extended code output in the receiving step, where C is an integer of 2 or more, in a case where the average value is not less than the limit time Tmax in the determination, for frames after the determination, outputting the decoded digital voice signal based on the monaural code output in the receiving step as the decoded digital voice signal of C channels.
3. A method for receiving and decoding a voice signal, which is performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, includes: A receiving step of determining, for a group formed by a first code string received from the first communication line and a second code string received from the second communication line corresponding to the first code string, whether the difference in the reception times of the first code string and the second code string is less than a first preset limit time Tmin, greater than or equal to a second preset limit time Tmax which is greater than the first limit time Tmin, or greater than or equal to the first limit time Tmin and less than the second limit time Tmax among multiple groups. In the case where the average value is less than the first limit time Tmin in the determination, for frames after the determination, output an extended code with the same frame number as the mono code included in the first code string input from the first communication line and the extended code included in the second code string input from the second communication line. In the case where the average value is greater than or equal to the first limit time Tmin and less than the second limit time Tmax in the determination, for frames after the determination, output an extended code with the frame number closest to the mono code among the mono code included in the first code string input from the first communication line and the extended code included in the second code string input from the second communication line. In the case where the average value is greater than or equal to the second limit time Tmax in the determination, for frames after the determination, output the mono code included in the first code string input from the first communication line; and A decoding step of obtaining and outputting a decoded digital voice signal of C channels based on the mono code output in the receiving step and the extended code output in the receiving step for frames after the determination in the case where the average value is less than the first limit time Tmin in the determination and in the case where the average value is greater than or equal to the first limit time Tmin and less than the second limit time Tmax in the determination, where C is an integer greater than or equal to 2. In the case where the average value is greater than or equal to the second limit time Tmax in the determination, for frames after the determination, output the decoded digital voice signal based on the mono code output in the receiving step as the decoded digital voice signal of C channels.
4. A method for receiving and decoding a voice signal, which is performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, includes: Receiving step: For a frame in which the difference in frame numbers between the mono code included in the first code string input from the first communication line and the extended code closest to the mono code among the extended codes included in the second code string input from the second communication line is less than a predetermined value, output the mono code included in the first code string input from the first communication line and the extended code closest to the mono code among the extended codes included in the second code string input from the second communication line. For a frame in which the difference is not less than the predetermined value, output the mono code included in the first code string input from the first communication line; and Decoding step: For a frame in which the difference is less than the predetermined value, based on the mono code output in the receiving step and the extended code output in the receiving step, obtain and output a decoded digital audio signal for C channels, where C is an integer of 2 or more. For a frame in which the difference is not less than the predetermined value, output the decoded digital audio signal based on the mono code output in the receiving step as the decoded digital audio signal for C channels.
5. A method for decoding an audio signal, performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, comprising: Decoding step: For a group consisting of a first code string received from the first communication line and a second code string received from the second communication line corresponding to the first code string, when the average value of the differences in the reception times of the first code string and the second code string among multiple groups is less than a predetermined limit time Tmin, based on the mono code included in the first code string input from the first communication line and the extended code with the same frame number as the mono code included in the second code string input from the second communication line, obtain and output a decoded digital audio signal for C channels, where C is an integer of 2 or more. When the average value is not less than the limit time Tmin, based on the mono code included in the first code string input from the first communication line and the extended code closest to the mono code among the extended codes included in the second code string input from the second communication line, obtain and output a decoded digital audio signal for C channels.
6. A method for decoding an audio signal, performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, comprising: Decoding step: For a group consisting of a first code string received from the first communication line and a second code string received from the second communication line corresponding to the first code string, when the average value of the differences in the reception times of the first code string and the second code string among multiple groups is less than a predetermined limit time Tmax, based on the mono code included in the first code string input from the first communication line and the extended code closest to the mono code among the extended codes included in the second code string input from the second communication line, obtain and output a decoded digital audio signal for C channels, where C is an integer of 2 or more. When the average value is not less than the limit time Tmax, the decoded digital audio signal based on the monaural code included in the first code string input from the first communication line is output as the decoded digital audio signal of C channels.
7. An audio signal decoding method performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, comprising: A decoding step, for a group consisting of a first code string received from the first communication line and a second code string received from the second communication line corresponding to the first code string, when the difference in the reception times of the first code string and the second code string among multiple groups is less than a predetermined first limit time Tmin, based on the monaural code included in the first code string input from the first communication line and the extended code with the same frame number as the monaural code included in the second code string input from the second communication line, obtain and output the decoded digital audio signal of C channels, where C is an integer greater than or equal to 2. When the average value is greater than or equal to a predetermined second limit time Tmax, which is greater than the first limit time Tmin, the decoded digital audio signal based on the monaural code included in the first code string input from the first communication line is output as the decoded digital audio signal of C channels. When the average value is greater than or equal to the first limit time Tmin and less than the second limit time Tmax, based on the monaural code included in the first code string input from the first communication line and the extended code with the frame number closest to the monaural code included in the second code string input from the second communication line, obtain and output the decoded digital audio signal of C channels.
8. An audio signal decoding method performed by a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, comprising: A decoding step, for frames in which the difference in frame numbers between the monaural code included in the first code string input from the first communication line and the extended code with the frame number closest to the monaural code included in the second code string input from the second communication line is less than a predetermined value, based on the monaural code and the extended code, obtain and output the decoded digital audio signal of C channels, where C is an integer greater than or equal to 2. For frames where the difference is not less than the predetermined value, the decoded digital audio signal based on the monaural code is output as the decoded digital audio signal of C channels.
9. An audio signal receiving-side device included in a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, comprising: A receiving unit, for a group consisting of a first code string received from the first communication line and a second code string received from the second communication line corresponding to the first code string, determine whether the average value of the difference in the reception times of the first code string and the second code string among multiple groups is less than a predetermined limit time Tmin. In the case where the average value in the determination is less than the limit time Tmin, for frames after the determination, output the extension code whose frame number is the same as the mono code included in the first code string input from the first communication line, and the extension code included in the second code string input from the second communication line. In the case where the average value in the determination is not less than the limit time Tmin, for frames after the determination, output the extension code whose frame number is closest to the mono code among the mono code included in the first code string input from the first communication line and the extension code included in the second code string input from the second communication line; and A decoding unit, for frames after the determination, obtains and outputs decoded digital audio signals for C channels based on the mono code output by the receiving unit and the extension code output by the receiving unit, where C is an integer of 2 or more.
10. An audio signal receiving-side device, included in a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, includes: A receiving unit that determines whether the difference in the reception times of the first code string received from the first communication line and the second code string received from the second communication line corresponding to the first code string is less than a predetermined limit time Tmax for a group formed by the first code string and the second code string. In the case where the average value in the determination is less than the limit time Tmax, for frames after the determination, output the extension code whose frame number is closest to the mono code among the mono code included in the first code string input from the first communication line and the extension code included in the second code string input from the second communication line. In the case where the average value in the determination is not less than the limit time Tmax, for frames after the determination, output the mono code included in the first code string input from the first communication line; and A decoding unit, in the case where the average value in the determination is less than the limit time Tmax, for frames after the determination, obtains and outputs decoded digital audio signals for C channels based on the mono code output by the receiving unit and the extension code output by the receiving unit, where C is an integer of 2 or more. In the case where the average value in the determination is not less than the limit time Tmax, for frames after the determination, output the decoded digital audio signal based on the mono code output by the receiving unit as the decoded digital audio signals for C channels.
11. An audio signal receiving-side device, included in a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, includes: A receiving unit determines, for a group consisting of a first code string received from the first communication line and a second code string received from the second communication line corresponding to the first code string, whether the difference in the reception times of the first code string and the second code string is less than a predetermined first limit time Tmin, greater than or equal to a predetermined second limit time Tmax that is greater than the first limit time Tmin, or greater than or equal to the first limit time Tmin and less than the second limit time Tmax among multiple groups. In the case where the average value is less than the first limit time Tmin in the determination, for frames after the determination, an extended code whose frame number is the same as that of the mono code included in the first code string input from the first communication line and the extended code included in the second code string input from the second communication line is output. In the case where the average value is greater than or equal to the first limit time Tmin and less than the second limit time Tmax in the determination, for frames after the determination, an extended code whose frame number is closest to that of the mono code included in the first code string input from the first communication line and the extended code included in the second code string input from the second communication line is output. In the case where the average value is greater than or equal to the second limit time Tmax in the determination, for frames after the determination, the mono code included in the first code string input from the first communication line is output; and A decoding unit, in the case where the average value is less than the first limit time Tmin in the determination and in the case where the average value is greater than or equal to the first limit time Tmin and less than the second limit time Tmax in the determination, for frames after the determination, obtains and outputs decoded digital audio signals for C channels based on the mono code output by the receiving unit and the extended code output by the receiving unit, where C is an integer greater than or equal to 2. In the case where the average value is greater than or equal to the second limit time Tmax in the determination, for frames after the determination, the decoded digital audio signal based on the mono code output by the receiving unit is output as the decoded digital audio signal for C channels.
12. An audio signal receiving-side device included in a terminal device connected to a first communication line and a second communication line having a lower priority than the first communication line, comprising: A receiving unit outputs, for frames in which the difference in frame numbers between the mono code included in the first code string input from the first communication line and the extended code whose frame number is closest to that of the mono code included in the second code string input from the second communication line is less than a predetermined value, the mono code included in the first code string input from the first communication line and the extended code whose frame number is closest to that of the mono code included in the second code string input from the second communication line. For frames in which the difference is not less than the predetermined value, the mono code included in the first code string input from the first communication line is output; and A decoding unit, for frames where the difference is less than the predetermined value, obtains and outputs decoded digital audio signals for C channels based on the mono code output by the receiving unit and the extended code output by the receiving unit, where C is an integer greater than or equal to 2. For frames where the difference is not less than the predetermined value, outputs the decoded digital audio signal based on the mono code output by the receiving unit as the decoded digital audio signals for C channels.
13. A decoding device, included in a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, includes: A decoding unit, for a group formed by a first code string received from the first communication line and a second code string received from the second communication line corresponding to the first code string, when the difference in the reception times of the first code string and the second code string among multiple groups is less than a predetermined limit time Tmin, obtains and outputs decoded digital audio signals for C channels based on the mono code included in the first code string input from the first communication line and the extended code with the same frame number as the mono code included in the second code string input from the second communication line, where C is an integer greater than or equal to 2. When the average value is not less than the limit time Tmin, obtains and outputs decoded digital audio signals for C channels based on the mono code included in the first code string input from the first communication line and the extended code with the frame number closest to the mono code included in the second code string input from the second communication line.
14. A decoding device, included in a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, includes: A decoding unit, for a group formed by a first code string received from the first communication line and a second code string received from the second communication line corresponding to the first code string, when the difference in the reception times of the first code string and the second code string among multiple groups is less than a predetermined limit time Tmax, obtains and outputs decoded digital audio signals for C channels based on the mono code included in the first code string input from the first communication line and the extended code with the frame number closest to the mono code included in the second code string input from the second communication line, where C is an integer greater than or equal to 2. When the average value is not less than the limit time Tmax, outputs the decoded digital audio signal based on the mono code included in the first code string input from the first communication line as the decoded digital audio signals for C channels.
15. A decoding device, included in a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, includes: A decoding unit, for a group consisting of a first code string received from the first communication line and a second code string received from the second communication line corresponding to the first code string, when the difference in the reception times of the first code string and the second code string among multiple groups is less than a preset first limit time Tmin, based on the monaural code included in the first code string input from the first communication line and the extended code with the same frame number as the monaural code included in the second code string input from the second communication line, obtains a decoded digital audio signal for C channels and outputs it, where C is an integer greater than or equal to 2. When the average value is greater than or equal to a preset second limit time Tmax that is greater than the first limit time Tmin, outputs the decoded digital audio signal based on the monaural code included in the first code string input from the first communication line as the decoded digital audio signal for C channels. When the average value is greater than or equal to the first limit time Tmin and less than the second limit time Tmax, based on the monaural code included in the first code string input from the first communication line and the extended code with the frame number closest to the monaural code included in the second code string input from the second communication line, obtains a decoded digital audio signal for C channels and outputs it.
16. A decoding device, included in a terminal device connected to a first communication line and a second communication line with a lower priority than the first communication line, comprising: A decoding unit, for frames where the difference in frame numbers between the monaural code included in the first code string input from the first communication line and the extended code with the frame number closest to the monaural code included in the second code string input from the second communication line is less than a preset value, based on the monaural code and the extended code, obtains a decoded digital audio signal for C channels and outputs it, where C is an integer greater than or equal to 2. For frames where the difference is not less than the preset value, outputs the decoded digital audio signal based on the monaural code as the decoded digital audio signal for C channels.
17. A computer program product, which stores a program for causing a computer to execute the audio signal reception and decoding method according to any one of claims 1 to 4.
18. A computer program product, which stores a program for causing a computer to execute the audio signal decoding method according to any one of claims 5 to 8.
19. A computer-readable recording medium that records a program for causing a computer to execute the audio signal reception and decoding method according to any one of claims 1 to 4.
20. A computer-readable recording medium that records a program for causing a computer to execute the audio signal decoding method according to any one of claims 5 to 8.
Citation Information
Patent Citations
Audio signal packet communication method, audio signal packet transmission method and reception method, apparatus for the same, program thereof and recording medium
JP2005117132A
Scalable encoding device and scalable encoding method
CN101111887A
Sound packet reproducing method, sound packet reproducing apparatus, sound packet reproducing program, and recording medium
CN1926824A