Signal processing apparatus, method, and program

By selecting different processing loads based on metadata in audio signal processing, the problem of excessive processing load in portable devices is solved, and high-quality audio signals can be generated under limited processing capabilities.

CN115315747BActive Publication Date: 2026-05-05SONY GROUP CORP
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SONY GROUP CORP
Filing Date
2021-03-19
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies involve excessive processing loads when processing multiple audio signals, making it impossible for portable devices to fully perform sound quality enhancement processing, especially when the number of targets is small, resulting in an overload of processing power.

Method used

Different sound quality enhancement processing methods are selected based on the metadata of the audio signal, including high-load, medium-load, and low-load processing. The appropriate processing method is selected according to the type and priority of the audio signal to reduce the overall processing load.

Benefits of technology

Even on devices with limited processing power, it can effectively improve the quality of audio signals and achieve high-quality sound signal processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115315747B_ABST
    Figure CN115315747B_ABST
Patent Text Reader

Abstract

The present technology relates to a signal processing device and method and program that can obtain a high sound quality signal even with a small processing amount. The signal processing device includes a selection section provided with a plurality of audio signals and selecting an audio signal to be subjected to sound quality enhancement processing, and a sound quality enhancement processing section that performs sound quality enhancement processing on the audio signal selected by the selection section. The present technology can be applied to a portable terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This technology relates to signal processing apparatus, methods, and procedures, and more specifically, to signal processing apparatus, methods, and procedures that can obtain high-quality audio signals even with a small processing volume. Background Technology

[0002] In the past, bandwidth expansion processing and dynamic range expansion processing have been known as processes for enhancing the sound quality of audio signals, that is, as processes for improving sound quality.

[0003] For example, as such bandwidth expansion processing, a technique has been proposed in which the filter coefficients of a bandpass filter with a high passband are calculated based on the low-frequency subband signal, and the flattened signal obtained from the low-frequency subband signal is filtered by using the filter coefficients to generate a high-frequency signal (see, for example, PTL1).

[0004] [List of Citations]

[0005] [Patent Literature]

[0006] [PTL1]: US Patent No. 9922660 Summary of the Invention

[0007] [Technical Issues]

[0008] Incidentally, if one attempts to perform sound quality enhancement processing on target audio signals, each corresponding to one of multiple targets, such that the processing is performed uniformly on the audio signals of all targets, then the number of times the processing needs to be performed must equal the number of targets.

[0009] Therefore, for example, in some cases, currently available platforms (such as smartphones, portable players, sound amplifiers, etc.) cannot fully perform this processing.

[0010] For example, if the number of targets is relatively small, such as 12, and if one attempts to perform sound quality enhancement processing on all 12 targets, the processing load becomes undesirably high, ranging from 1 GCPS (cycles per second) to 3 GCPS.

[0011] This technology was developed in light of this situation, and its purpose is to obtain high-quality audio signals even with a small processing load.

[0012] [Solution to the problem]

[0013] A signal processing apparatus according to one aspect of the present technology includes: a selection unit that is provided with a plurality of audio signals and selects an audio signal to be subjected to sound quality enhancement processing; and a sound quality enhancement processing unit that performs sound quality enhancement processing on the audio signal selected by the selection unit.

[0014] A signal processing method or procedure according to one aspect of the present technology includes the following steps: providing a plurality of audio signals; selecting an audio signal to be subjected to sound quality enhancement processing; and performing sound quality enhancement processing on the selected audio signal.

[0015] In one aspect of this technology, multiple audio signals are provided, an audio signal to be subjected to sound quality enhancement processing is selected, and sound quality enhancement processing is performed on the selected audio signal. Attached Figure Description

[0016] Figure 1 This is a diagram depicting an example configuration of a signal processing device.

[0017] Figure 2 This is a diagram illustrating a configuration example of the sound quality enhancement processing unit.

[0018] Figure 3 This is a diagram depicting a configuration instance of the dynamic range extension.

[0019] Figure 4 This is a diagram depicting a configuration example of the bandwidth extension section.

[0020] Figure 5 This is a diagram depicting a configuration instance of the dynamic range extension.

[0021] Figure 6 This is a diagram depicting a configuration example of the bandwidth extension section.

[0022] Figure 7 This is a diagram depicting a configuration example of the bandwidth extension section.

[0023] Figure 8 It is a flowchart used to illustrate the process of generating the reproduced signal.

[0024] Figure 9 This is a flowchart used to illustrate high-load sound quality enhancement processing.

[0025] Figure 10 This is a flowchart used to illustrate the medium-load sound quality enhancement process.

[0026] Figure 11 This is a flowchart used to illustrate low-load sound quality enhancement processing.

[0027] Figure 12 This is a diagram illustrating an example configuration of a signal processing device.

[0028] Figure 13 It is a flowchart used to illustrate the process of generating the reproduced signal.

[0029] Figure 14 This is a diagram depicting an example configuration of a signal processing device.

[0030] Figure 15 This is a diagram depicting an example configuration of a signal processing device.

[0031] Figure 16 It is a flowchart used to illustrate the process of generating the reproduced signal.

[0032] Figure 17 It is a diagram illustrating a configuration example of a computer. Detailed Implementation

[0033] The following describes the implementation of this technology with reference to the accompanying drawings.

[0034] <First Implementation Method>

[0035] <About this technology>

[0036] This technology aims to enable the acquisition of high-quality audio signals even with a small amount of processing by selecting different processes to be performed on the audio signal when performing sound quality enhancement on multi-channel audio represented by a target audio sound, using metadata and the like.

[0037] For example, in this technique, for each audio signal, the audio quality enhancement processing to be performed on the audio signal is selected based on metadata, etc. In other words, the audio signal to be subjected to audio quality enhancement processing is selected.

[0038] By doing so, the overall amount of processing required for sound quality enhancement can be reduced, and high-quality sound signals can be obtained even on platforms such as portable terminals with low processing power.

[0039] In recent years, there have been plans to allocate multi-channel audio represented by target audio sounds. In such audio allocation, for example, the MPEG (Moving Picture Experts Group)-H format could be used.

[0040] For example, as a sound quality enhancement process for compressed signals (audio signals) in MPEG-H format, dynamic range extension processing and bandwidth extension processing can be performed.

[0041] Here, dynamic range extension processing is the process of extending the dynamic range of an audio signal, that is, the bit count (quantization bit count) of a sample value of an audio signal. Furthermore, bandwidth extension processing is the process of adding high-frequency components to an audio signal that does not include high-frequency components.

[0042] Incidentally, it is impractical to perform high-processing-load audio quality enhancement processing and further improve the audio quality of all multiple audio signals.

[0043] Therefore, for example, this technology enables more appropriate sound quality improvement by performing sound quality enhancement processing on important audio signals (which requires high processing load but provides higher sound quality improvement) based on metadata of the audio signal, and performing sound quality enhancement processing on less important audio signals (which requires low processing load). That is, a signal with sufficiently high sound quality can be obtained even with a small amount of processing.

[0044] Note that the audio signal that is the object of sound quality enhancement can be any audio signal, but the following explanation assumes that multiple audio signals included in the predetermined content are the objects of sound quality enhancement.

[0045] Additionally, it is assumed that the multiple audio signals included in the content (as objects of sound quality enhancement) include audio signals of channels such as R or L, and audio signals of audio targets such as sounds (hereinafter referred to as targets).

[0046] Furthermore, it is assumed that each audio signal has metadata added to it, and that the metadata includes type information and priority information. Additionally, it is assumed that the metadata of the target's audio signal also includes location information indicating the target's position.

[0047] Type information is information that indicates the type of audio signal, such as, for example, the channel name of the audio signal (such as L or R), or the type of the target (such as voice or guitar), or more specifically, the type of sound source of the target.

[0048] Assume the priority information represents the priority of the audio signal, and here the priority is represented by a value from 1 to 10. Specifically, assume that the smaller the value representing the priority, the higher the priority. Therefore, in this example, priority "1" is the highest priority, and priority "10" is the lowest priority.

[0049] Furthermore, in the example described below, three distinct audio quality enhancement processes are prepared in advance: a high-load audio quality enhancement process, a medium-load audio quality enhancement process, and a low-load audio quality enhancement process. Then, based on metadata, the audio quality enhancement process to be performed on the audio signal is selected from these processes.

[0050] High-load audio enhancement processing is the audio enhancement processing that requires the highest processing load among the three audio enhancement processing methods but provides the highest audio quality improvement effect, and is particularly useful as an audio enhancement processing method for high-priority audio signals or audio signals of high importance.

[0051] As a specific example of high-load sound quality enhancement processing, dynamic range extension processing and bandwidth extension processing based on DNN (deep neural network) and other technologies, which are obtained in advance through machine learning, can be combined and executed.

[0052] Low-load audio enhancement processing is the audio enhancement processing that requires the lowest processing load and provides the lowest audio quality improvement effect among the three audio enhancement processing methods, and is particularly useful as an audio enhancement processing method for low-priority or low-importance types of audio signals.

[0053] As a specific example of low-load audio quality enhancement processing, for example, processing that requires extremely low load can be performed in combination, such as bandwidth expansion processing using predetermined coefficients or coefficients specified on the encoding side, simple bandwidth expansion processing that adds a signal such as white noise as a high-frequency component to the audio signal, or dynamic range expansion processing that uses filtering by predetermined coefficients.

[0054] Medium load audio enhancement processing is the audio enhancement processing that requires the second highest processing load among the three audio enhancement processing and provides the second highest audio quality improvement effect. It is particularly useful as an audio enhancement processing for audio signals of medium priority or medium importance.

[0055] As a specific example of medium-load sound quality enhancement processing, bandwidth expansion processing that generates high-frequency components through linear prediction, dynamic range expansion processing that uses filtering with predetermined coefficients, etc., can be performed in combination.

[0056] Note that although the number of distinct audio quality enhancement processes is three in the examples described below, the number of distinct audio quality enhancement processes can be any number, including two or more. Furthermore, the audio quality enhancement processes are not limited to dynamic range extension or bandwidth extension. Other processes may be performed, or only dynamic range extension or only bandwidth extension may be performed.

[0057] Here, a specific example is given. For instance, suppose there are eight audio signals, OB1 to OB7, which are to be used as the audio signals for sound quality enhancement.

[0058] In addition, the type and priority of each target are written as (type, priority).

[0059] Now assume that the metadata representations of target OB1 to target OB7 are of the following types and priorities: (voice, 1), (drum, 1), (guitar, 2), (bass, 3), (reverb, 9), (audience, 10), and (ambient sound, 10).

[0060] At this point, for example, on a platform with typical processing capabilities, high-load sound quality enhancement processing is performed on the audio signals of targets OB1 and OB2, which have the highest priority of "1". Furthermore, medium-load sound quality enhancement processing is performed on the audio signals of targets OB3 and OB4, which have priorities of "2" and "3", while low-load sound quality enhancement processing is performed on the audio signals of other targets (targets OB5 to OB7) with lower priorities.

[0061] In contrast, at a reproduction device (platform) with high processing power and capable of performing a greater number of processes to improve sound quality, high-load sound quality enhancement processing is performed on a greater number of target audio signals, unlike the previously mentioned examples.

[0062] For example, suppose the metadata representations of target OB1 to target OB7 are of the following types and priorities: (voice, 1), (drum, 2), (guitar, 2), (bass, 3), (reverb, 9), (audience, 10), and (ambient sound, 10).

[0063] At this point, high-load sound quality enhancement processing is performed on the audio signals of targets OB1 to OB3, which have high priorities "1" and "2", and medium-load sound quality enhancement processing is performed on the audio signals of targets OB4 and OB5, which have priorities "3" and "9". Then, low-load sound quality enhancement processing is performed only on the audio signals of targets OB6 and OB7, which have the lowest priority "10".

[0064] Furthermore, on platforms with processing power below typical levels, high-load sound quality enhancement processing is performed on fewer audio signals and is performed more efficiently compared to the two previously mentioned examples.

[0065] For example, suppose the metadata representations of target OB1 to target OB7 are of the following types and priorities: (voice, 1), (drum, 2), (guitar, 2), (bass, 3), (reverb, 9), (audience, 10), and (ambient sound, 10).

[0066] At this point, high-load sound quality enhancement processing is performed only on the audio signal of target OB1, which has the highest priority "1", and medium-load sound quality enhancement processing is performed on the audio signals of targets OB2 and OB3, which have priority "2". Then, low-load sound quality enhancement processing is performed on the audio signals of targets OB4 to OB7, which have a priority equal to or greater than "3".

[0067] As described above, in this technology, the sound quality enhancement processing to be performed on each audio signal is selected based on at least priority information or type information included in the metadata. By doing so, for example, the overall processing load can be set when performing sound quality enhancement according to the processing capacity of the reproduction device (platform), and sound quality enhancement, i.e., sound quality improvement, can be performed in any type of reproduction device.

[0068] <Configuration Example of Signal Processing Device>

[0069] Next, a more specific implementation of the above-described technology will be described.

[0070] Figure 1 This is a diagram illustrating a configuration example of an embodiment of a signal processing apparatus that applies this technology.

[0071] For example, in Figure 1 The signal processing device 11 described herein includes smartphones, portable players, sound amplifiers, personal computers, tablet computers, etc.

[0072] The signal processing device 11 includes a decoding unit 21, an audio selection unit 22, a sound quality enhancement processing unit 23, a renderer 24, and a reproducible signal generation unit 25.

[0073] For example, multiple audio signals are provided to the decoding unit 21, as well as encoded data obtained by the decoding unit 21 through encoding metadata of the audio signals. For example, the encoded data is in a predetermined encoding format, such as a bitstream of MPEG-H.

[0074] The decoding unit 21 performs decoding processing on the provided encoded data and provides the audio signal and the metadata of the audio signal obtained therefrom to the audio selection unit 22.

[0075] For each of the multiple audio signals provided from the decoding unit 21, and based on the metadata provided from the decoding unit 21, the audio selection unit 22 selects the audio signal to be subjected to sound quality enhancement processing, and provides the audio signal to the sound quality enhancement processing unit 23 according to the selection result.

[0076] In other words, the audio selection unit 22 is provided with multiple audio signals from the decoding unit 21, and also selects audio signals to undergo audio quality enhancement processing, such as high-load audio quality enhancement processing, based on metadata.

[0077] The audio selection unit 22 has selection units 31-1 to 31-m, and provides an audio signal and metadata of the audio signal to each of the selection units 31-1 to 31-m.

[0078] Specifically, in this example, the encoded data includes audio signals of n targets and audio signals of (mn) channels as audio signals of objects to be enhanced in sound quality. Then, the audio signals of the targets and their metadata are provided to selection units 31-1 to 31-n, and the audio signals of the channels and their metadata are provided to selection units 31-(n+1) to selection units 31-m.

[0079] Based on the metadata provided from the decoding unit 21, the selection units 31-1 to 31-m select the audio quality enhancement processing to be performed on the audio signal provided from the decoding unit 21 (i.e., the block to which the audio signal is output), and provide the audio signal to the block in the audio quality enhancement processing unit 23 according to the selection result.

[0080] Furthermore, the selection units 31-1 to 31-n provide the metadata of the target audio signal provided by the decoding unit 21 to the renderer 24 via the sound quality enhancement processing unit 23.

[0081] Note that, in the following cases where there is no need to specifically distinguish between selection sections 31-1 to 31-m, they are also referred to simply as selection section 31.

[0082] The sound quality enhancement processing unit 23 performs any one of three predetermined sound quality enhancement processes on each audio signal provided from the audio selection unit 22, and outputs the resulting audio signal as a high sound quality signal. The three sound quality enhancement processes mentioned here are the high-load sound quality enhancement process, the medium-load sound quality enhancement process, and the low-load sound quality enhancement process described above.

[0083] The sound quality enhancement processing unit 23 includes high-load sound quality enhancement processing units 32-1 to 32-m, medium-load sound quality enhancement processing units 33-1 to 33-m, and low-load sound quality enhancement processing units 34-1 to 34-m.

[0084] When an audio signal is provided from the selection units 31-1 to 31-m, the high-load sound quality enhancement processing units 32-1 to 32-m perform high-load sound quality enhancement processing on the provided audio signal and generate a high-quality sound signal.

[0085] The high-load sound quality enhancement processing units 32-1 to 32-n provide the high-quality sound signal of the target obtained through high-load sound quality enhancement processing to the renderer 24.

[0086] Furthermore, the high-load sound quality enhancement processing units 32-(n+1) to 32-m provide the high-quality sound signals of the channels obtained through the high-load sound quality enhancement processing to the reproduction signal generation unit 25.

[0087] Note that, in the absence of special distinction between the high-load sound quality enhancement processing units 32-1 to 32-m, they are also referred to simply as the high-load sound quality enhancement processing unit 32.

[0088] When an audio signal is provided from the selection units 31-1 to 31-m, the medium-load sound quality enhancement processing units 33-1 to 33-m perform medium-load sound quality enhancement processing on the provided audio signal and generate a high-quality sound signal.

[0089] The medium-load sound quality enhancement processing units 33-1 to 33-n provide the high-quality sound signal of the target obtained through the medium-load sound quality enhancement processing to the renderer 24.

[0090] Furthermore, the medium-load sound quality enhancement processing units 33-(n+1) to 33-m provide the high sound quality signal of the channel obtained by the medium-load sound quality enhancement processing to the reproduction signal generation unit 25.

[0091] In addition, where there is no need to specifically distinguish between the medium-load sound quality enhancement processing units 33-1 to 33-m, they are also referred to simply as the medium-load sound quality enhancement processing units 33.

[0092] When an audio signal is provided from the selection units 31-1 to 31-m, the low-load sound quality enhancement processing units 34-1 to 34-m perform low-load sound quality enhancement processing on the provided audio signal and generate a high-quality sound signal.

[0093] The low-load sound quality enhancement processing units 34-1 to 34-n provide the high-quality sound signal of the target obtained through low-load sound quality enhancement processing to the renderer 24.

[0094] Furthermore, the low-load sound quality enhancement processing units 34-(n+1) to 34-m provide the high sound quality signal of the channel obtained through the low-load sound quality enhancement processing to the reproduction signal generation unit 25.

[0095] In addition, where there is no need to specifically distinguish between the low-load sound quality enhancement processing units 34-1 to 34-m, they are also referred to simply as the low-load sound quality enhancement processing units 34.

[0096] Based on the metadata provided by the sound quality enhancement processing unit 23, the renderer 24 performs rendering processing on the target high sound quality signal provided by the high-load sound quality enhancement processing unit 32, the medium-load sound quality enhancement processing unit 33, and the low-load sound quality enhancement processing unit 34, according to the reproduction device such as the downstream speaker.

[0097] For example, at renderer 24, VBAP (Vector Based Amplitude Panning) is performed as a rendering process, and a target reproduction signal is obtained, which positions the sound of each target at a location represented by the location information contained in the target's metadata. The target reproduction signal is a multi-channel audio signal comprising (mn) audio channels.

[0098] The renderer 24 provides the target reproduction signal obtained through the rendering process to the reproduction signal generation unit 25.

[0099] The reproduction signal generation unit 25 performs synthesis processing of the target reproduction signal provided by the renderer 24 and the high sound quality signals of the channels provided by the high-load sound quality enhancement processing unit 32, the medium-load sound quality enhancement processing unit 33, and the low-load sound quality enhancement processing unit 34.

[0100] For example, in the synthesis process, the target reproduction signal and the high-quality sound signal of the same channel are added (synthesized) to generate (mn) channel reproduction signals. If these reproduction signals are reproduced at (mn) loudspeakers, the sound of each channel or the sound of each target is reproduced, that is, the sound of the content.

[0101] The regenerated signal generation unit 25 outputs the regenerated signal obtained through synthesis processing to the downstream side.

[0102] <Configuration Example of Sound Quality Enhancement Processing Unit>

[0103] Next, configuration examples of the high-load sound quality enhancement processing unit 32, the medium-load sound quality enhancement processing unit 33, and the low-load sound quality enhancement processing unit 34 will be described.

[0104] For example, such as Figure 2 The configuration shown includes a high-load sound quality enhancement processing unit 32, a medium-load sound quality enhancement processing unit 33, and a low-load sound quality enhancement processing unit 34. It should be noted that... Figure 2 An example is described where the renderer 24 is set downstream of the high-load sound quality enhancement processing unit 32 to the low-load sound quality enhancement processing unit 34.

[0105] exist Figure 2 In the example shown, the high-load sound quality enhancement processing unit 32 has a dynamic range extension unit 61 and a bandwidth extension unit 62.

[0106] The dynamic range extension unit 61 performs dynamic range extension processing on the audio signal provided from the selection unit 31 based on a DNN generated in advance through machine learning, and provides the audio signal obtained therefrom to the bandwidth extension unit 62.

[0107] The bandwidth extension unit 62 performs bandwidth extension processing on the audio signal provided by the dynamic range extension unit 61 based on a DNN pre-generated by machine learning, and provides the resulting high sound quality signal to the renderer 24.

[0108] The medium-load sound quality enhancement processing unit 33 has a dynamic range extension unit 71 and a bandwidth extension unit 72.

[0109] The dynamic range extension unit 71 performs dynamic range extension processing on the audio signal provided from the selection unit 31 through a multi-stage all-pass filter, and provides the audio signal obtained therefrom to the bandwidth extension unit 72.

[0110] The bandwidth expansion unit 72 performs bandwidth expansion processing on the audio signal provided from the dynamic range expansion unit 71 using linear prediction, and provides the resulting high sound quality signal to the renderer 24.

[0111] In addition, the low-load sound quality enhancement processing unit 34 has a dynamic range extension unit 81 and a bandwidth extension unit 82.

[0112] The dynamic range extension unit 81 performs dynamic range extension processing on the audio signal provided from the selection unit 31, similar to the dynamic range extension processing performed in the case of the dynamic range extension unit 71, and provides the audio signal obtained therefrom to the bandwidth extension unit 82.

[0113] On the audio signal provided from the dynamic range extension unit 81, the bandwidth extension unit 82 performs bandwidth extension processing using coefficients specified on the encoding side, and provides the resulting high sound quality signal to the renderer 24.

[0114] <Configuration Example of Dynamic Range Extension>

[0115] Moreover, the explanation below is as follows Figure 2 Configuration examples of the dynamic range extension unit 61, bandwidth extension unit 62, etc., described in the text.

[0116] Figure 3 This is a diagram illustrating a more detailed configuration example of the dynamic range extension unit 61.

[0117] Figure 3 The dynamic range extension unit 61 shown includes an FFT (Fast Fourier Transform) processing unit 111, a gain calculation unit 112, a differential signal generation unit 113, an IFFT (Inverse Fast Fourier Transform) processing unit 114, and a synthesis unit 115.

[0118] At the dynamic range extension section 61, the differential signal is predicted using prediction calculations of a DNN, and the differential signal and the audio signal are synthesized. The differential signal is the difference between the audio signal decoded at the decoding section 21 and the original sound signal before encoding the audio signal. By doing so, a high-quality audio signal that is closer to the original sound signal can be obtained.

[0119] The FFT processing unit 111 performs an FFT on the audio signal provided from the selection unit 31 and provides the resulting signal to the gain calculation unit 112 and the differential signal generation unit 113.

[0120] The gain calculation unit 112 includes a DNN obtained in advance through machine learning. That is, the gain calculation unit 112 holds the prediction coefficients obtained in advance through machine learning and used for calculation in the DNN, and uses them as a predictor to predict the envelope of the frequency characteristics of the differential signal.

[0121] Based on the retained prediction coefficients and the signal provided from the FFT processing unit 111, the gain calculation unit 112 calculates a gain value as a parameter for generating a differential signal corresponding to the audio signal, and provides the gain value to the differential signal generation unit 113. That is, the gain of the frequency envelope of the differential signal is calculated as a parameter for generating the differential signal.

[0122] Based on the signal provided from the FFT processing unit 111 and the gain value provided from the gain calculation unit 112, the differential signal generation unit 113 generates a differential signal and provides the differential signal to the IFFT processing unit 114. The IFFT processing unit 114 performs IFFT on the differential signal provided from the differential signal generation unit 113 and provides the resulting time-domain differential signal to the synthesis unit 115.

[0123] The synthesis unit 115 synthesizes the audio signal provided from the selection unit 31 and the differential signal provided from the IFFT processing unit 114, and provides the resulting audio signal to the bandwidth expansion unit 62.

[0124] <Configuration Example of Bandwidth Extension Section>

[0125] In addition, for example, Figure 2 The bandwidth extension unit 62 shown is configured as follows Figure 4 As shown in the image.

[0126] Figure 4 The bandwidth extension unit 62 shown includes a multi-phase low-pass filter 141, a delay circuit 142, a low-frequency extraction bandpass filter 143, a feature calculation circuit 144, a high-frequency subband power estimation circuit 145, a bandpass filter calculation circuit 146, an adder 147, a high-pass filter 148, a flattening circuit 149, a downsampling unit 150, a multi-phase level adjustment filter 151, and an adder 152.

[0127] The multiphase low-pass filter 141 performs filtering on the audio signal provided from the synthesis unit 115 of the dynamic range extension unit 61 using a multiphase low-pass filter, and provides the resulting low-frequency signal to the delay circuit 142.

[0128] In the multiphase low-pass filter 141, the low-frequency components of the signal are upsampled and extracted by filtering with the multiphase low-pass filter, and a low-frequency signal is obtained.

[0129] The delay circuit 142 delays the low-frequency signal provided by the multi-phase low-pass filter 141 by a certain delay time and provides the low-frequency signal to the adder 152.

[0130] The low-frequency extraction bandpass filter 143 includes bandpass filters 161-1 to 161-K, which have different passbands from each other.

[0131] The bandpass filter 161-k (nb1≤k≤K) allows signals from sub-bands of a predetermined passband that are low-frequency components of the audio signal provided by the synthesis unit 115 to pass through, and provides the signals in the predetermined frequency band obtained therefrom as low-frequency sub-band signals to the feature calculation circuit 144 and the flattening circuit 149. Therefore, in the low-frequency extraction bandpass filter 143, low-frequency sub-band signals from the K sub-bands included in the low-frequency range are obtained.

[0132] Note that, in the following cases, where there is no need to make a special distinction between bandpass filters 161-1 to bandpass filters 161-K, they are also referred to simply as bandpass filter 161.

[0133] The feature calculation circuit 144 calculates features based on multiple low-frequency subband signals provided from the bandpass filter 161 or audio signals provided from the synthesis unit 115, and provides these features to the high-frequency subband power estimation circuit 145.

[0134] The high-frequency subband power estimation circuit 145 includes a DNN pre-obtained through machine learning. That is, the high-frequency subband power estimation circuit 145 retains the prediction coefficients pre-obtained through machine learning and used for calculations in the DNN.

[0135] The high-frequency subband power estimation circuit 145 calculates an estimate of the high-frequency subband power for each high-frequency subband based on the maintained prediction coefficients and features provided from the feature calculation circuit 144, and provides this estimate to the bandpass filter calculation circuit 146. This high-frequency subband power is the power of the high-frequency subband signal. Hereinafter, the estimated high-frequency subband power is also referred to as pseudo-high-frequency subband power.

[0136] The bandpass filter calculation circuit 146 calculates the bandpass filter coefficients of the bandpass filter with the passband being the high-frequency subband based on the pseudo-high-frequency subband power from the multiple high-frequency subbands provided by the high-frequency subband power estimation circuit 145, and provides the bandpass filter coefficients to the adder 147.

[0137] The adder 147 adds the bandpass filter coefficients provided by the bandpass filter calculation circuit 146 to form a filter coefficient, and provides the filter coefficient to the high-pass filter 148.

[0138] By using a high-pass filter to filter the filter coefficients provided from the adder 147, the high-pass filter 148 removes low-frequency components from the filter coefficients and provides the resulting filter coefficients to the polyphase configuration level adjustment filter 151. That is, the high-pass filter 148 only allows the high-frequency components of the filter coefficients to pass through.

[0139] By flattening and summing the low-frequency subband signals from multiple low-frequency subbands provided by the bandpass filter 161, the flattening circuit 149 generates a flattened signal and provides the flattened signal to the downsampling unit 150.

[0140] The downsampling unit 150 performs downsampling on the flattening signal provided from the flattening circuit 149 and provides the downsampled flattening signal to the multiphase configuration level adjustment filter 151.

[0141] By using the filtering coefficients provided by the high-pass filter 148 to filter the flattened signal provided by the downsampling unit 150, the multiphase configuration level adjustment filter 151 generates a high-frequency signal and provides the high-frequency signal to the adder 152.

[0142] The adder 152 adds the low-frequency signal provided by the delay circuit 142 and the high-frequency signal provided by the multi-phase configuration level adjustment filter 151 to form a high-quality sound signal and provides the high-quality sound signal to the renderer 24 or the reproduction signal generation unit 25.

[0143] The high-frequency signal obtained in the multiphase configuration level adjustment filter 151 is a high-frequency component signal that is not included in the original audio signal, i.e., a high-frequency component signal that was undesirably lost during the encoding of the audio signal. Therefore, by combining such a high-frequency signal with a low-frequency signal that is a low-frequency component of the original audio signal, a signal including components in a wider frequency band can be obtained, i.e., a high-quality sound signal with higher sound quality.

[0144] <Configuration Example of Dynamic Range Extension>

[0145] in addition, Figure 2 The dynamic range extension unit 71 of the medium-load sound quality enhancement processing unit 33 shown is, for example, as... Figure 5 It is constructed as shown.

[0146] Figure 5 The dynamic range extension unit 71 shown includes all-pass filters 191-1 to 191-3, a gain adjustment unit 192, and an adder unit 193. In this example, the three all-pass filters 191-1 to 191-3 are connected in a cascaded manner.

[0147] The all-pass filter 191-1 filters the audio signal provided from the selection unit 31 and provides the resulting audio signal to the downstream all-pass filter 191-2.

[0148] The full-pass filter 191-2 performs filtering on the audio signal provided from the full-pass filter 191-1 and provides the resulting audio signal to the downstream full-pass filter 191-3.

[0149] The all-pass filter 191-3 filters the audio signal provided from the all-pass filter 191-2 and provides the resulting audio signal to the gain adjustment unit 192.

[0150] Note that, in the following cases, where there is no need to make special distinctions between all-pass filters 191-1 to 191-3, they are also referred to simply as all-pass filter 191.

[0151] The gain adjustment unit 192 adjusts the gain of the audio signal provided from the all-pass filter 191-3 and provides the gain-adjusted audio signal to the adder 193.

[0152] The addition unit 193 generates an audio signal with improved sound quality (i.e., expanded dynamic range) by adding the audio signal provided from the gain adjustment unit 192 and the audio signal provided from the selection unit 31, and provides the audio signal to the bandwidth expansion unit 72.

[0153] Because the processing performed at the dynamic range extension section 71 is filtering and gain adjustment, these processes can be performed using computational processes smaller than (or lower than) those in a DNN (such as in...). Figure 3 The processing load is achieved by those (those executed at the dynamic range extension section 61 described in the text).

[0154] <Configuration Example of Bandwidth Extension Section>

[0155] In addition, for example, Figure 2 The bandwidth extension unit 72 shown is configured as follows Figure 6 As shown in the image.

[0156] Figure 6 The bandwidth extension unit 72 shown includes a multi-phase low-pass filter 221, a delay circuit 222, a low-frequency extraction bandpass filter 223, a feature calculation circuit 224, a high-frequency subband power estimation circuit 225, a bandpass filter calculation circuit 226, an adder 227, a high-pass filter 228, a flattening circuit 229, a downsampling unit 230, a multi-phase level adjustment filter 231, and an adder 232.

[0157] In addition, the low-frequency extraction bandpass filter 223 has bandpass filters 241-1 to 241-K.

[0158] Note that because the multiphase configuration low-pass filter 221 to feature calculation circuit 224 and band-pass filter calculation circuit 226 to adder 232 have the same configuration, and perform the same... Figure 4 The multiphase configuration low-pass filter 141 to feature calculation circuit 144 and band-pass filter calculation circuit 146 to adder 152 of the bandwidth extension section 62 shown operate the same, so their description is omitted.

[0159] Furthermore, because bandpass filters 241-1 to 241-K also have the same characteristics as... Figure 4 The bandpass filters 161-1 to 161-K of the bandwidth extension unit 62 shown have the same configuration and perform the same operation, so their description is omitted.

[0160] Note that, in the following cases, where there is no need to make a special distinction between bandpass filters 241-1 to bandpass filters 241-K, they are also referred to simply as bandpass filters 241.

[0161] exist Figure 6 The bandwidth extension unit 72 described in the text is related to the bandwidth extension unit 72 in the text. Figure 4 The bandwidth extension unit 62 described herein differs only in its operation in the high-frequency subband power estimation circuit 225, and is otherwise identical to the bandwidth extension unit 62 in terms of configuration and operation.

[0162] The high-frequency subband power estimation circuit 225 retains the coefficients obtained in advance through statistical learning, and calculates the pseudo-high-frequency subband power based on the retained coefficients and the features provided from the feature calculation circuit 224, and provides the pseudo-high-frequency subband power to the bandpass filter calculation circuit 226. For example, at the high-frequency subband power estimation circuit 225, the high-frequency component, more specifically, the pseudo-high-frequency subband power, is calculated by using linear prediction of the retained coefficients.

[0163] The linear prediction at the high-frequency subband power estimation circuit 225 can be achieved with a smaller processing load compared to the prediction made by computation in the DNN at the high-frequency subband power estimation circuit 145.

[0164] <Configuration Example of Bandwidth Extension Section>

[0165] in addition, Figure 2 The dynamic range extension unit 81 of the low-load sound quality enhancement processing unit 34 shown has, for example, a function similar to... Figure 5 The dynamic range extension unit 71 shown has the same configuration. Alternatively, in the low-load sound quality enhancement processing unit 34, the dynamic range extension unit 81 may not be specifically provided.

[0166] in addition, Figure 2 The bandwidth extension unit 82 of the low-load sound quality enhancement processing unit 34 shown is, for example, as... Figure 7 It is constructed as shown.

[0167] Figure 7 The bandwidth extension unit 82 shown includes a subband segmentation circuit 271, a feature calculation circuit 272, a high-frequency decoding circuit 273, a high-frequency subband power calculation circuit 274, a high-frequency signal generation circuit 275, and a synthesis circuit 276.

[0168] It should be noted that the bandwidth extension section 82 has... Figure 7 In the configuration described herein, the encoded data provided to the decoding unit 21 includes high-frequency encoded data, and the high-frequency encoded data is provided to the high-frequency decoding circuit 273. The high-frequency encoded data is data obtained by encoding indices used to obtain the high-frequency subband power estimation coefficients described later.

[0169] The sub-band segmentation circuit 271 uniformly divides the audio signal provided by the dynamic range extension unit 81 into multiple low-frequency sub-band signals with a predetermined bandwidth, and provides the multiple low-frequency sub-band signals to the feature calculation circuit 272 and the decoding high-frequency signal generation circuit 275.

[0170] Based on the low-frequency subband signal provided by the subband segmentation circuit 271, the feature calculation circuit 272 calculates the features and provides the features to the decoding high-frequency subband power calculation circuit 274.

[0171] The high-frequency decoding circuit 273 decodes the provided high-frequency encoded data and provides the high-frequency subband power estimation coefficients corresponding to the obtained index to the high-frequency subband power calculation circuit 274.

[0172] For each of the multiple metrics, at the high-frequency decoding circuit 273, the high-frequency subband power estimation coefficients are recorded in association with that metric.

[0173] In this case, on the encoding side of the audio signal, an index representing the high-frequency subband power estimation coefficients most suitable for bandwidth expansion processing at the bandwidth expansion section 82 is selected, and the selected index is encoded. Then, the high-frequency encoded data obtained through encoding is stored in a bitstream and provided to the signal processing device 11.

[0174] Therefore, the high-frequency decoding circuit 273 selects a high-frequency subband power estimation coefficient from a plurality of pre-recorded high-frequency subband power estimation coefficients, which is represented by an index obtained by decoding high-frequency encoded data, and provides the coefficient to the high-frequency subband power calculation circuit 274.

[0175] The high-frequency subband power calculation circuit 274 calculates the high-frequency subband power based on the features provided by the feature calculation circuit 272 and the high-frequency subband power estimation coefficients provided by the high-frequency decoding circuit 273, and provides the high-frequency subband power to the high-frequency signal generation circuit 275.

[0176] The high-frequency signal generation circuit 275 generates a high-frequency signal based on the low-frequency subband signal provided by the subband segmentation circuit 271 and the high-frequency subband power provided by the high-frequency subband power calculation circuit 274, and provides the high-frequency signal to the synthesis circuit 276.

[0177] The synthesis circuit 276 synthesizes the audio signal provided by the dynamic range extension unit 81 and the high-frequency signal provided by the decoding high-frequency signal generation circuit 275, and provides the high sound quality signal obtained therefrom to the renderer 24 or the reproduction signal generation unit 25.

[0178] The high-frequency signal obtained in the decoding high-frequency signal generation circuit 275 is a high-frequency component signal that is not included in the original audio signal. Therefore, by synthesizing such a high-frequency signal with the original audio signal, a high-quality sound signal with higher sound quality that includes components in a wider frequency band can be obtained.

[0179] As described above, because the bandwidth extension unit 82 predicts the high-frequency signal using high-frequency sub-band power estimation coefficients represented by the provided index in the bandwidth extension process, it is similar to... Figure 6 Compared to the case of the bandwidth extension unit 72 described in the previous section, prediction can be achieved with a smaller processing load.

[0180] <Explanation of Reproduced Signal Generation and Processing>

[0181] Next, the operation of the signal processing device 11 will be explained.

[0182] That is, refer to the following Figure 8 The flowchart illustrates the reproduction signal generation process performed by the signal processing device 11. When the decoding unit 21 decodes the provided encoded data, the reproduction signal generation process begins, and the audio signal and metadata obtained through decoding are provided to the selection unit 31.

[0183] In step S11, based on the metadata provided from the decoding unit 21, the selection unit 31 selects the sound quality enhancement processing to be performed on the audio signal provided from the decoding unit 21.

[0184] That is, for example, the selection unit 31 selects a process as a sound quality enhancement process, which is one of high-load sound quality enhancement processing, medium-load sound quality enhancement processing and low-load sound quality enhancement processing, based on the priority information and type information included in the provided metadata.

[0185] Specifically, for example, in step S11, high-load sound quality enhancement processing is selected when the priority indicated by the priority information is equal to or lower than a predetermined value, or when the type indicated by the type information is a specific type such as the center channel or voice.

[0186] Note that although at least one of the priority information or type information is used for the selection of sound quality enhancement processing, in addition to these, sound quality enhancement processing can be selected by using information representing the processing capabilities of the signal processing device 11, etc.

[0187] Specifically, for example, when the processing capability indicated by the information representing processing capability is equal to or higher than a predetermined value, the selection priority value of high-load sound quality enhancement processing, etc., is changed, thereby increasing the number of audio signals selected for high-load sound quality enhancement processing.

[0188] In step S12, the selection unit 31 determines whether to perform high-load sound quality enhancement processing.

[0189] For example, if high-load sound quality enhancement processing is selected as the selection result in step S11, it is determined in step S12 that high-load sound quality enhancement processing will be performed.

[0190] If it is determined in step S12 that high-load sound quality enhancement processing will be performed, the selection unit 31 will provide the audio signal provided by the decoding unit 21 to the high-load sound quality enhancement processing unit 32, and thereafter, the processing will proceed to step S13.

[0191] In step S13, the high-load sound quality enhancement processing unit 32 performs high-load sound quality enhancement processing on the audio signal provided from the selection unit 31, and outputs the resulting high-quality sound signal. Note that details of the high-load sound quality enhancement processing will be mentioned later.

[0192] For example, if the target signal is an audio signal with enhanced sound quality, the high-load sound quality enhancement processing unit 32 provides the obtained high-quality sound signal to the renderer 24. In this case, the selection unit 31 provides the renderer 24 with location information included in the metadata provided from the decoding unit 21 via the sound quality enhancement processing unit 23.

[0193] Conversely, when the audio signal with enhanced sound quality is a channel signal, the high-load sound quality enhancement processing unit 32 provides the obtained high sound quality signal to the reproduction signal generation unit 25.

[0194] After performing high-load sound quality enhancement processing and generating a high-quality sound signal, the processing proceeds to step S17.

[0195] Furthermore, if it is determined in step S12 that high-load sound quality enhancement processing will not be performed, in step S14, the selection unit 31 determines whether to perform medium-load sound quality enhancement processing.

[0196] For example, if medium-load sound quality enhancement processing is selected as the selection result in step S11, then in step S14, it is determined to perform medium-load sound quality enhancement processing.

[0197] If it is determined in step S14 that medium-load sound quality enhancement processing will be performed, the selection unit 31 will provide the audio signal provided by the decoding unit 21 to the medium-load sound quality enhancement processing unit 33, and thereafter the processing will proceed to step S15.

[0198] In step S15, the medium-load sound quality enhancement processing unit 33 performs medium-load sound quality enhancement processing on the audio signal provided from the selection unit 31, and outputs the resulting high-quality sound signal. It should be noted that details of the medium-load sound quality enhancement processing will be discussed later.

[0199] For example, if the target signal is an audio signal with enhanced sound quality, the medium-load sound quality enhancement processing unit 33 provides the obtained high-quality sound signal to the renderer 24. In this case, the selection unit 31 provides the renderer 24 with location information included in the metadata provided from the decoding unit 21 via the sound quality enhancement processing unit 23.

[0200] Conversely, when the audio signal with enhanced sound quality is a channel signal, the medium-load sound quality enhancement processing unit 33 provides the obtained high sound quality signal to the reproduction signal generation unit 25.

[0201] After performing the load-enhanced sound quality processing and generating a high-quality sound signal, the processing proceeds to step S17.

[0202] Furthermore, if it is determined in step S14 that medium-load audio quality enhancement processing will not be performed, i.e., low-load audio quality enhancement processing will be performed, the processing proceeds to step S16. In this case, the selection unit 31 provides the audio signal from the decoding unit 21 to the low-load audio quality enhancement processing unit 34.

[0203] In step S16, the low-load sound quality enhancement processing unit 34 performs low-load sound quality enhancement processing on the audio signal provided from the selection unit 31 and outputs the resulting high-quality sound signal. Note that details of the low-load sound quality enhancement processing will be discussed later.

[0204] For example, if the target signal is an audio signal with enhanced sound quality, the low-load sound quality enhancement processing unit 34 provides the obtained high-quality sound signal to the renderer 24. In this case, the selection unit 31 provides the renderer 24 with location information included in the metadata provided from the decoding unit 21 via the sound quality enhancement processing unit 23.

[0205] Conversely, when the audio signal with enhanced sound quality is a channel signal, the low-load sound quality enhancement processing unit 34 provides the obtained high sound quality signal to the reproduction signal generation unit 25.

[0206] After performing low-load sound quality enhancement processing and generating a high-quality sound signal, the processing proceeds to step S17.

[0207] After performing the processing of step S13, step S15 or step S16, the processing of step S17 is performed.

[0208] In step S17, the audio selection unit 22 determines whether all audio signals provided from the decoding unit 21 have been processed.

[0209] For example, in step S17, if the selection of sound quality enhancement processing for the provided audio signals is performed in the selection units 31-1 to 31-m, and sound quality enhancement processing is performed in the sound quality enhancement processing unit 23 based on the selection result, it is determined that all audio signals have been processed. In this case, a high sound quality signal corresponding to all audio signals has been generated.

[0210] If it is determined in step S17 that not all audio signals have been processed, the process returns to step S11 and the above process is repeated.

[0211] For example, if the selection unit 31-n has not yet performed the processing of step S11, the processing of steps S11 to S16 described above is performed on the audio signal provided to the selection unit 31-n. In addition, specifically, in the sound selection unit 22, the selection unit 31 performs the processing of steps S11 to S16 in parallel.

[0212] Conversely, if it is determined in step S17 that all audio signals have been processed, then the processing proceeds to step S18.

[0213] In step S18, the renderer 24 performs rendering processing on a total of n high-quality signals provided by the high-load sound quality enhancement processing unit 32, the medium-load sound quality enhancement processing unit 33, and the low-load sound quality enhancement processing unit 34 in the sound quality enhancement processing unit 23.

[0214] For example, by performing VBAP based on the target's location information and high-quality sound signal provided by the sound quality enhancement processing unit 23, the renderer 24 generates a target reproduction signal and provides the target reproduction signal to the reproduction signal generation unit 25.

[0215] In step S19, the reproduction signal generation unit 25 synthesizes the target reproduction signal provided by the renderer 24 and the high sound quality signals of the channels provided by the high-load sound quality enhancement processing unit 32, the medium-load sound quality enhancement processing unit 33, and the low-load sound quality enhancement processing unit 34, and generates a reproduction signal.

[0216] The reproduced signal generation unit 25 outputs the obtained reproduced signal to the downstream side, and then the reproduced signal generation process ends.

[0217] In the manner described above, based on priority and type information included in the metadata, the signal processing device 11 selects from a plurality of sound quality enhancement processes that require different processing loads to perform a sound quality enhancement process on each audio signal, and performs the sound quality enhancement process according to the selection result. By doing so, the overall processing load can be reduced, and even with a small processing load, i.e., a small amount of processing, a reproduced signal with sufficiently high sound quality can be obtained.

[0218] <Explanation of High-Load Sound Quality Enhancement Processing>

[0219] Here, a more detailed explanation is provided for reference. Figure 8 The high-load sound quality enhancement process in step S13, the medium-load sound quality enhancement process in step S15, and the low-load sound quality enhancement process in step S16 are explained.

[0220] First, refer to Figure 9 The flowchart illustrates the process performed by the high-load sound quality enhancement processing unit 32. Figure 8 Step S13 corresponds to the high-load sound quality enhancement process.

[0221] In step S41, the FFT processing unit 111 performs an FFT on the audio signal provided from the selection unit 31 and provides the signal obtained therefrom to the gain calculation unit 112 and the differential signal generation unit 113.

[0222] In step S42, based on the retained prediction coefficients and the signal provided from the FFT processing unit 111, the gain calculation unit 112 calculates the gain value for generating the differential signal and provides the gain value to the differential signal generation unit 113. In step S42, based on the prediction coefficients and the signal provided from the FFT processing unit 111, calculations in the DNN are performed, and the gain value of the frequency envelope of the differential signal is calculated.

[0223] In step S43, the differential signal generation unit 113 generates a differential signal based on the signal provided from the FFT processing unit 111 and the gain value provided from the gain calculation unit 112, and provides the differential signal to the IFFT processing unit 114. For example, in step S43, the differential signal is generated by adjusting the gain of the signal provided from the FFT processing unit 111 based on the gain value.

[0224] In step S44, the IFFT processing unit 114 performs IFFT on the differential signal provided from the differential signal generation unit 113, and provides the resulting differential signal to the synthesis unit 115.

[0225] In step S45, the synthesis unit 115 synthesizes the audio signal provided from the selection unit 31 and the differential signal provided from the IFFT processing unit 114, and provides the audio signal obtained therefrom to the multiphase configuration low-pass filter 141, feature calculation circuit 144 and band-pass filter 161 of the bandwidth expansion unit 62.

[0226] In step S46, the multiphase low-pass filter 141 performs filtering on the audio signal provided from the synthesis unit 115 using the multiphase low-pass filter, and provides the resulting low-frequency signal to the delay circuit 142.

[0227] In addition, the delay circuit 142 delays the low-frequency signal provided by the multi-phase low-pass filter 141 by a certain length of delay time, and then provides the low-frequency signal to the adder 152.

[0228] In step S47, by allowing signals from the low-frequency subbands of the audio signal provided from the synthesis unit 115 to pass through, the bandpass filter 161 divides the audio signal into multiple low-frequency subband signals and provides the multiple low-frequency subband signals to the feature calculation circuit 144 and the flattening circuit 149.

[0229] In step S48, the feature calculation circuit 144 calculates a feature based on at least one of a plurality of low-frequency subband signals provided from the bandpass filter 161 or an audio signal provided from the synthesis unit 115, and provides the feature to the high-frequency subband power estimation circuit 145.

[0230] In step S49, the high-frequency subband power estimation circuit 145 calculates the pseudo high-frequency subband power for each high-frequency subband based on the pre-held prediction coefficients and the features provided from the feature calculation circuit 144, and provides the pseudo high-frequency subband power to the bandpass filter calculation circuit 146.

[0231] In step S50, the bandpass filter calculation circuit 146 calculates the bandpass filter coefficients based on the pseudo-high-frequency subband power in the multiple high-frequency subbands provided by the high-frequency subband power estimation circuit 145, and provides the bandpass filter coefficients to the adder 147.

[0232] In addition, the adder 147 adds the bandpass filter coefficients supplied from the bandpass filter calculation circuit 146 to form a filter coefficient, and provides the filter coefficient to the high-pass filter 148.

[0233] In step S51, the high-pass filter 148 performs filtering on the filtering coefficients provided from the adder 147 using the high-pass filter, and provides the resulting filtering coefficients to the multiphase configuration level adjustment filter 151.

[0234] In step S52, the flattening circuit 149 generates a flattened signal by flattening and summing the low-frequency subband signals from the multiple low-frequency subbands provided by the bandpass filter 161, and provides the flattened signal to the downsampling unit 150.

[0235] In step S53, the downsampling unit 150 downsamples the flattening signal provided from the flattening circuit 149 and provides the downsampled flattening signal to the multiphase configuration level adjustment filter 151.

[0236] In step S54, the flattened signal provided from the downsampling unit 150 is filtered by using the filtering coefficients provided from the high-pass filter 148, the multiphase configuration level adjustment filter 151 generates a high-frequency signal and provides the high-frequency signal to the adder 152.

[0237] In step S55, by adding the low-frequency signal provided from the delay circuit 142 and the high-frequency signal provided from the multi-phase configuration level adjustment filter 151 together, the adder 152 generates and outputs a high-quality sound signal. After generating the high-quality sound signal in this way, the high-load sound quality enhancement process ends, and thereafter, the process continues to... Figure 8 Step S17 in the process.

[0238] As described above, the high-load audio quality enhancement processing unit 32 combines high-load dynamic range extension processing and bandwidth extension processing, but can obtain a high-quality audio signal and generate a high-quality audio signal with high sound quality. By doing so, high-quality audio signals can be obtained for important audio signals, such as high-priority audio signals.

[0239] <Explanation of Medium-Load Sound Quality Enhancement Processing>

[0240] Next, refer to Figure 10 The flowchart in the diagram illustrates the corresponding process executed by the medium-load sound quality enhancement processing unit 33. Figure 8 The medium-load sound quality enhancement process in step S15.

[0241] In step S81, the full-pass filter 191 filters the audio signal provided from the selection unit 31 using a multi-stage full-pass filter, and provides the resulting audio signal to the gain adjustment unit 192.

[0242] That is, in step S81, filtering is performed at all-pass filters 191-1 to 191-3.

[0243] In step S82, the gain adjustment unit 192 performs gain adjustment on the audio signal provided from the all-pass filter 191-3, and provides the gain-adjusted audio signal to the adder unit 193.

[0244] In step S83, the adder 193 adds the audio signal provided from the gain adjustment unit 192 and the audio signal provided from the selection unit 31 together, and provides the resulting audio signal to the multiphase configuration low-pass filter 221, feature calculation circuit 224 and band-pass filter 241 of the bandwidth extension unit 72.

[0245] After the processing in step S83 is performed, the processing in steps S84 to S86 is performed by configuring the multiphase low-pass filter 221, the band-pass filter 241, and the feature calculation circuit 224. It should be noted that because these processes are related to... Figure 9 The processes of steps S46 to S48 are similar, so their descriptions are omitted.

[0246] In step S87, based on the retained coefficients and the features provided by the feature calculation circuit 224, the high-frequency subband power estimation circuit 225 calculates the pseudo-high-frequency subband power through linear prediction and provides the pseudo-high-frequency subband power to the bandpass filter calculation circuit 226.

[0247] After the processing in step S87, the bandpass filter calculation circuit 226 to the adder 232 execute the processing in steps S88 to S93, and the medium-load sound quality enhancement processing ends. It should be noted that because these processes are related to... Figure 9 The processes in steps S50 to S55 are similar, so their descriptions are omitted. After the medium-load sound quality enhancement process is completed, the processing proceeds to... Figure 8 Step S17 in the process.

[0248] In the manner described above, the medium-load sound quality enhancement processing unit 33 combines dynamic range expansion processing and bandwidth expansion processing to enhance the sound quality of the target and channel audio signals. This dynamic range expansion processing and bandwidth expansion processing enable the acquisition of a signal with a relatively high level of sound quality using a medium load. By doing so, for audio signals with a certain degree of high priority, a signal with a relatively high level of sound quality can be obtained using a medium load, and so on.

[0249] <Instructions for Low-Load Sound Quality Enhancement Processing>

[0250] In addition, refer to Figure 11 The flowchart illustrates the corresponding process executed by the low-load sound quality enhancement processing unit 34. Figure 8 The low-load sound quality enhancement process in step S16.

[0251] It should be noted that the processing in steps S121 to S123 is different from... Figure 10 The processes in steps S81 to S83 are similar, so their descriptions are omitted.

[0252] After the processing in step S123 is performed, the audio signal obtained by the processing in step S123 is provided from the dynamic range extension unit 81 to the sub-band splitting circuit 271 and the synthesis circuit 276 of the bandwidth extension unit 82, and the processing in step S124 is performed.

[0253] In step S124, the subband segmentation circuit 271 segments the audio signal provided by the dynamic range extension unit 81 into multiple low-frequency subband signals, and provides the multiple low-frequency subband signals to the feature calculation circuit 272 and the decoding high-frequency signal generation circuit 275.

[0254] In step S125, based on the low-frequency subband signal provided by the subband segmentation circuit 271, the feature calculation circuit 272 calculates features and provides the features to the decoding high-frequency subband power calculation circuit 274.

[0255] In step S126, the high-frequency decoding circuit 273 decodes the provided high-frequency encoded data and outputs (provides) the high-frequency subband power estimation coefficients corresponding to the obtained index to the high-frequency subband power calculation circuit 274.

[0256] In step S127, the decoding high-frequency subband power calculation circuit 274 calculates the high-frequency subband power based on the features provided by the feature calculation circuit 272 and the high-frequency subband power estimation coefficients provided by the high-frequency decoding circuit 273, and provides the high-frequency subband power to the decoding high-frequency signal generation circuit 275. For example, in step S127, the high-frequency subband power is calculated by determining the sum of the features multiplied by the high-frequency subband power estimation coefficients.

[0257] In step S128, the high-frequency signal generation circuit 275 generates a high-frequency signal based on the low-frequency subband signal provided by the subband segmentation circuit 271 and the high-frequency subband power provided by the high-frequency subband power calculation circuit 274, and provides the high-frequency signal to the synthesis circuit 276. For example, in step S128, the low-frequency subband signal is frequency-modulated and gain-adjusted based on the low-frequency subband signal and the high-frequency subband power, and a high-frequency signal is generated.

[0258] In step S129, the synthesis circuit 276 synthesizes the audio signal provided from the dynamic range extension unit 81 and the high-frequency signal provided from the decoding high-frequency signal generation circuit 275, and outputs the resulting high-quality audio signal. After generating the high-quality audio signal in this manner, the low-load audio quality enhancement process ends, and thereafter, the process proceeds to... Figure 8 Step S17 in the process.

[0259] As described above, the low-load sound quality enhancement processing unit 34 is capable of performing dynamic range extension processing and bandwidth extension processing for sound quality enhancement at a low load, and enhancing the sound quality of the target and channel audio signals. By doing so, sound quality enhancement is performed on less important audio signals (such as audio signals with low priority) at a low load, and the overall processing load can be reduced.

[0260] <Second Implementation Method>

[0261] <Configuration Example of Signal Processing Device>

[0262] As described above, in the high-load sound quality enhancement processing unit 32, the prediction coefficients used in the calculation of the DNN obtained in advance by machine learning are used to estimate (predict) the gain of the frequency envelope and the pseudo-high frequency subband power.

[0263] At this point, if the type of audio signal can be identified, prediction coefficients for each type can also be learned. By doing so, using prediction coefficients based on the type of audio signal, it is possible to more accurately and additionally predict the gain and pseudo-high-frequency subband power of the frequency envelope with a smaller processing load.

[0264] Specifically, if the prediction coefficients (i.e., DNNs) for each type of audio signal are machine learning-based, gain values ​​and pseudo-high-frequency subband power can be predicted more accurately with smaller-scale DNNs, reducing the processing load.

[0265] On the other hand, if there are no issues with handling the load, the same DNN, i.e., the same prediction coefficients, can be used independently of the type of audio signal. In this case, for example, if typical stereo audio content from various sound sources, also known as complete packages, is used for machine learning of the prediction coefficients, then that is sufficient.

[0266] The predictive coefficients generated by machine learning using audio content (e.g., complete packages) that includes sounds from various sources, and used collectively for all types of prediction coefficients, are also referred to below as general predictive coefficients.

[0267] In the first embodiment described above, because the metadata of each audio signal includes type information indicating the type of the audio signal, the type of the audio signal can be identified. Therefore, for example, as... Figure 12 As shown, sound quality enhancement can be performed by selecting prediction coefficients based on type information. It should be noted that... Figure 12 It has with Figure 1 In the case of corresponding parts, the same reference numerals are given, and their explanations are appropriately omitted.

[0268] exist Figure 12The signal processing apparatus 11 described herein includes a decoding unit 21, an audio selection unit 22, a sound quality enhancement processing unit 23, a renderer 24, and a reproducible signal generation unit 25.

[0269] In addition, the sound selection unit 22 has selection units 31-1 to 31-m.

[0270] In addition, the sound quality enhancement processing unit 23 includes ordinary sound quality enhancement processing units 302-1 to 302-m, high-load sound quality enhancement processing units 32-1 to 32-m, and coefficient selection units 301-1 to 301-m.

[0271] therefore, Figure 12 The signal processing device 11 shown is Figure 1 The only difference in the signal processing device 11 shown is the configuration of the sound quality enhancement processing unit 23; otherwise, they are the same.

[0272] The coefficient selection units 301-1 to 301-m pre-store the predicted coefficients for machine learning and calculation in the DNN for each type of audio signal, and these coefficient selection units 301-1 to 301-m are provided with metadata from the decoding unit 21.

[0273] The prediction coefficients described here are prediction coefficients used for processing in the high-load sound quality enhancement processing unit 32, more specifically in the gain calculation unit 112 of the dynamic range extension unit 61 and the high-frequency subband power estimation circuit 145 of the bandwidth extension unit 62.

[0274] Coefficient selection units 301-1 to 301-m select prediction coefficients of a type represented by type information included in the metadata provided from the decoding unit 21 from prediction coefficients corresponding to one of a variety of types that are stored in advance, and provide the prediction coefficients to high-load audio quality enhancement processing units 32-1 to 32-m. That is, for each audio signal, prediction coefficients to be used for high-load audio quality enhancement processing performed on the audio signal are selected.

[0275] Note that unless there is a specific need to distinguish between the coefficient selection sections 301-1 and 301-m below, they are also referred to simply as coefficient selection section 301.

[0276] The ordinary sound quality enhancement processing units 302-1 to 302-m are basically configured in the same way as the high-load sound quality enhancement processing unit 32.

[0277] It should be noted that the configuration of the blocks corresponding to the gain calculation unit 112 and the high-frequency subband power estimation circuit 145 in the ordinary sound quality enhancement processing units 302-1 to 302-m, i.e., the DNN configuration, is different from that in the high-load sound quality enhancement processing unit 32. These blocks maintain the aforementioned general prediction coefficients.

[0278] In addition, for example, in the ordinary sound quality enhancement processing unit 302-1 to the ordinary sound quality enhancement processing unit 302-m, the DNN configuration and the like may vary depending on whether the input audio signal is the target signal or the channel signal.

[0279] After the audio signal is provided from the selection units 31-1 to 31-m, the ordinary sound quality enhancement processing units 302-1 to 302-m perform sound quality enhancement processing based on the pre-held audio signal and the general prediction coefficient, and provide the high sound quality signal obtained therefrom to the renderer 24 or the reproduction signal generation unit 25.

[0280] Furthermore, unless otherwise specified below, the general sound quality enhancement processing units 302-1 to 302-m will also be referred to simply as the general sound quality enhancement processing unit 302. Additionally, the sound quality enhancement processing performed in the general sound quality enhancement processing unit 302 will be specifically referred to below as general sound quality enhancement processing.

[0281] Thus, in Figure 12 In the example shown, each selection unit 31 selects either the normal sound quality enhancement processing unit 302 or the high-load sound quality enhancement processing unit 32 as the destination for providing the audio signal based on the priority information and type information contained in the metadata.

[0282] <Explanation of Reproduced Signal Generation and Processing>

[0283] Next, refer to the following: Figure 13 The flowchart in the document is explained by... Figure 12 The signal processing device 11 described herein performs the reproducible signal generation process.

[0284] In step S161, based on the metadata provided from the decoding unit 21, the selection unit 31 selects the sound quality enhancement processing to be performed on the audio signal provided from the decoding unit 21.

[0285] For example, if the type represented by the type information included in the metadata is a type in which the prediction coefficients are pre-held at the coefficient selection unit 301, the selection unit 31 selects the high-load sound quality enhancement process. Conversely, if the type represented by the type information is a type in which the prediction coefficients are not held at the coefficient selection unit 301, the normal sound quality enhancement process is selected.

[0286] In step S162, the selection unit 31 determines whether high-load sound quality enhancement processing was selected in step S161, that is, whether to perform high-load sound quality enhancement processing.

[0287] If it is determined in step S162 that high-load sound quality enhancement processing will be performed, the selection unit 31 will provide the audio signal provided by the decoding unit 21 to the high-load sound quality enhancement processing unit 32, and thereafter, the processing will proceed to step S163.

[0288] In step S163, the coefficient selection unit 301 selects a prediction coefficient of a type represented by type information included in the metadata provided from the metadata of the decoding unit 21 from each prediction coefficient corresponding to one of the multiple types held in advance, and provides the prediction coefficient to the high-load sound quality enhancement processing unit 32.

[0289] Here, prediction coefficients are selected by machine learning for the type and will be used in each of the gain calculation unit 112 and the high-frequency subband power estimation circuit 145, and the prediction coefficients are provided to the gain calculation unit 112 and the high-frequency subband power estimation circuit 145.

[0290] After selecting the prediction coefficients, the process proceeds to step S164. That is, in step S164, a reference... Figure 9 The description refers to the high-load sound quality enhancement processing.

[0291] It should be noted that in step S42, the gain calculation unit 112 calculates the gain value for generating the differential signal based on the prediction coefficients provided by the coefficient selection unit 301 and the signal provided by the FFT processing unit 111. Furthermore, in step S49, the high-frequency subband power estimation circuit 145 calculates the pseudo-high-frequency subband power based on the prediction coefficients provided by the coefficient selection unit 301 and the features provided by the feature calculation circuit 144.

[0292] Furthermore, if it is determined in step S162 that high-load sound quality enhancement processing will not be performed, that is, if it is determined that normal sound quality enhancement processing will be performed, the selection unit 31 will provide the audio signal provided by the decoding unit 21 to the normal sound quality enhancement processing unit 302, and thereafter, the processing will proceed to step S165.

[0293] In step S165, the normal sound quality enhancement processing unit 302 performs normal sound quality enhancement processing on the audio signal provided from the selection unit 31, and provides the high sound quality signal obtained therefrom to the renderer 24 or the reproduction signal generation unit 25.

[0294] In general sound quality enhancement processing, the basic process involves execution and reference. Figure 9 The high-load sound quality enhancement process described herein is similar to the process used to generate a high-quality sound signal.

[0295] It should be noted, for example, in the case of ordinary sound quality enhancement processing and corresponding to Figure 9 In step S42 of the process, the pre-held universal prediction coefficients are used to calculate the gain value for generating the differential signal. Furthermore, in conjunction with... Figure 9 In step S49 of the process, the pre-held universal prediction coefficients are used to calculate the pseudo-high frequency subband power.

[0296] After performing the processing in step S164 or step S165 as described above, the processing in steps S166 to S168 is performed, and the signal generation process ends. Because these processes are related to... Figure 8 The processes in steps S17 to S19 are similar, so their descriptions are omitted.

[0297] In the manner described above, based on priority and type information contained in the metadata, the signal processing device 11 selectively performs ordinary sound quality enhancement processing or high-load sound quality enhancement processing, and generates a reproduced signal. By doing so, a reproduced signal with sufficiently high sound quality can be obtained even under a small processing load, i.e., a small processing amount. Specifically, in this example, by preparing prediction coefficients for each type of audio signal, a high-quality reproduced signal can be obtained with a small processing load.

[0298] <First Variation of the Second Embodiment>

[0299] <Configuration Example of Signal Processing Device>

[0300] It is important to note that when referring to Figure 12 In the example provided, either high-load sound quality enhancement processing or normal sound quality enhancement processing is selected as the sound quality enhancement process. However, this is not the only example, and any two or more of the following can be selected: high-load sound quality enhancement processing, medium-load sound quality enhancement processing, low-load sound quality enhancement processing, and normal sound quality enhancement processing.

[0301] For example, when any one of high-load sound quality enhancement processing, medium-load sound quality enhancement processing, low-load sound quality enhancement processing, and normal sound quality enhancement processing is selected as the sound quality enhancement processing, the signal processing device 11 is configured as follows: Figure 14 As shown in the image. It should be noted that in... Figure 14 It has in Figure 1 or Figure 12 In the case of [the preceding text], the corresponding parts are given the same reference numerals, and their explanations are appropriately omitted.

[0302] exist Figure 14 The signal processing apparatus 11 described herein includes a decoding unit 21, an audio selection unit 22, a sound quality enhancement processing unit 23, a renderer 24, and a reproducible signal generation unit 25.

[0303] In addition, the sound selection unit 22 has selection units 31-1 to 31-m.

[0304] In addition, the sound quality enhancement processing unit 23 includes ordinary sound quality enhancement processing units 302-1 to 302-m, medium-load sound quality enhancement processing units 33-1 to 33-m, low-load sound quality enhancement processing units 34-1 to 34-m, high-load sound quality enhancement processing units 32-1 to 32-m, and coefficient selection units 301-1 to 301-m.

[0305] therefore, Figure 14 The signal processing device 11 shown is Figure 1 or Figure 12 The only difference in the signal processing device 11 shown is the configuration of the sound quality enhancement processing unit 23; all other configurations are the same.

[0306] In this example, based on the metadata provided by the decoding unit 21, the selection unit 31 selects the sound quality enhancement processing to be performed on the audio signal provided by the decoding unit 21.

[0307] That is, the selection unit 31 selects high-load sound quality enhancement processing, medium-load sound quality enhancement processing, low-load sound quality enhancement processing or normal sound quality enhancement processing, and provides the audio signal to the high-load sound quality enhancement processing unit 32, medium-load sound quality enhancement processing unit 33, low-load sound quality enhancement processing unit 34 or normal sound quality enhancement processing unit 302 according to the selection result.

[0308] <Third Implementation Method>

[0309] <Configuration Example of Signal Processing Device>

[0310] Furthermore, if a coefficient selection unit 301 is provided in the sound quality enhancement processing unit 23, and the type of audio signal cannot be identified due to reasons such as the lack of type information in the metadata, the prediction coefficient cannot be selected in the coefficient selection unit 301, and high-load sound quality enhancement processing cannot be performed.

[0311] Therefore, for example, a metadata generation unit can be provided that generates metadata based on audio signals. Specifically, in the example explained below, the type of audio signal is identified based on the audio signal, and type information representing the identification result is generated as metadata.

[0312] In this case, the signal processing device 11 is configured, for example, as follows: Figure 15 As shown. It should be noted that, in Figure 15 It has in Figure 12 In the case of [the preceding text], the corresponding parts are given the same reference numerals, and their explanations are appropriately omitted.

[0313] exist Figure 15 The signal processing apparatus 11 described herein includes a decoding unit 21, an audio selection unit 22, a sound quality enhancement processing unit 23, a renderer 24, and a reproducible signal generation unit 25.

[0314] In addition, the audio selection unit 22 includes selection units 31-1 to 31-m and metadata generation units 341-1 to 341-m.

[0315] In addition, the sound quality enhancement processing unit 23 includes ordinary sound quality enhancement processing units 302-1 to 302-m, high-load sound quality enhancement processing units 32-1 to 32-m, and coefficient selection units 301-1 to 301-m.

[0316] therefore, Figure 15 The signal processing device 11 shown is Figure 12 The only difference in the signal processing device 11 shown is the configuration of the audio selection unit 22, and the configurations in other aspects are the same.

[0317] For example, metadata generation units 341-1 to 341-m are type classifiers of DNNs pre-generated through machine learning, and type prediction coefficients for implementing the type classifier are pre-stored. That is, by learning the type prediction coefficients through machine learning, a type classifier such as a DNN can be obtained.

[0318] Based on the pre-stored type prediction coefficients and the audio signal provided from the decoding unit 21, the metadata generation units 341-1 to 341-m perform calculations through the type classifier to identify (estimate) the type of the audio signal. For example, at the type classifier, type identification is performed based on the frequency characteristics of the audio signal, etc.

[0319] Metadata generation units 341-1 to 341-m generate type information representing the identification result of the type, i.e., metadata, and provide the type information to selection units 31-1 to 31-m and coefficient selection units 301-1 to 301-m.

[0320] Note that, in the absence of any special distinction required for metadata generation units 341-1 to 341-m, they are also referred to simply as metadata generation unit 341.

[0321] Furthermore, the type classifier included in the metadata generation unit 341 can be a type classifier that outputs information about the type of the input audio signal, indicating which of several types the audio signal belongs to, or multiple type classifiers, each corresponding to a specific type, and can output information indicating whether the input audio signal is of a specific type. For example, in the case where a type classifier is prepared for each type, the audio signal is input to the type classifier, and type information is generated based on the output of each type classifier.

[0322] Furthermore, although the example described here includes a normal sound quality enhancement processing unit 302 and a high-load sound quality enhancement processing unit 32 in the sound quality enhancement processing unit 23, a medium-load sound quality enhancement processing unit 33 and a low-load sound quality enhancement processing unit 34 may also be provided.

[0323] <Explanation of Reproduced Signal Generation and Processing>

[0324] Next, refer to the following: Figure 16 The flowchart in the document explains the process. Figure 15 The signal processing device 11 described herein performs the reproducible signal generation process.

[0325] In step S201, based on the pre-stored type prediction coefficients and the audio signal provided from the decoding unit 21, the metadata generation unit 341 identifies the type of the audio signal and generates type information indicating the identification result. The metadata generation unit 341 provides the generated type information to the selection unit 31 and the coefficient selection unit 301.

[0326] It should be noted that, more specifically, at the metadata generation unit 341, the processing at step S201 is performed only if the metadata obtained at the decoding unit 21 does not include type information. Here, we will continue the explanation assuming that the metadata does not include type information.

[0327] In step S202, based on priority information included in the metadata provided from the decoding unit 21 and type information provided from the metadata generation unit 341, the selection unit 31 selects the sound quality enhancement process to be performed on the audio signal provided from the decoding unit 21. Here, either high-load sound quality enhancement processing or normal sound quality enhancement processing is selected as the sound quality enhancement process.

[0328] After selecting the sound quality enhancement process, the processes in steps S203 to S209 are executed, and the signal generation process ends. Because these processes are related to... Figure 13 The processes in steps S162 to S168 are similar, so their descriptions are omitted. It should be noted that in step S204, the coefficient selection unit 301 selects prediction coefficients based on the type information provided from the metadata generation unit 341.

[0329] In the manner described above, the signal processing device 11 generates type information based on the audio signal, and selects sound quality enhancement processing based on the type information and priority information. By doing so, type information can be generated even when the metadata does not include type information, and sound quality enhancement processing and prediction coefficients can be selected. Therefore, a high-quality sound reproduction signal can be obtained even with a small processing load.

[0330] <Computer Configuration Examples>

[0331] Incidentally, the above series of processes can be executed via hardware or software. In the case where the processes are executed by software, the program within the software is installed on the computer. Here, "computer" includes computers with dedicated hardware, general-purpose personal computers, such as personal computers capable of performing various types of functions by installing various types of programs on them.

[0332] Figure 17 It is a block diagram describing a hardware configuration instance of a computer that performs the above series of processes through a program.

[0333] In a computer, the CPU (Central Processing Unit) 501, ROM (Read-Only Memory) 502, and RAM (Random Access Memory) 503 are interconnected via a bus 504.

[0334] Bus 504 is further connected to input / output interface 505. Input / output interface 505 is connected to input unit 506, output unit 507, recording unit 508, communication unit 509 and driver 510.

[0335] Input unit 506 includes a keyboard, mouse, microphone, image capture element, etc. Output unit 507 includes a display, speaker, etc. Recording unit 508 includes a hard disk, non-volatile memory, etc. Communication unit 509 includes a network interface, etc. Driver 510 drives removable recording media 511 such as hard disk, optical disk, magneto-optical disk, or semiconductor memory.

[0336] In a computer configured in this way, for example, the CPU 501 loads the program recorded on the recording unit 508 into the RAM 503 via the input / output interface 505 and the bus 504, and executes the program, thereby performing the series of processes described above.

[0337] For example, a program executed by a computer (CPU 501) can be configured to be recorded on a removable recording medium 511, such as a packaging medium. Furthermore, the program can be provided via a wired transmission medium or a wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0338] At the computer, by attaching the removable recording medium 511 to the drive 510, the program can be installed on the recording unit 508 via the input / output interface 505. Alternatively, the program can be received at the communication unit 509 via a cable transmission medium or a wireless transmission medium and installed on the recording unit 508. Or, the program can be pre-installed on the ROM 502 or the recording unit 508.

[0339] It should be noted that the program executed by the computer may be a program that executes processing in chronological order as described in this specification, or it may be a program that executes processing in parallel or at necessary time intervals, such as when those processing is called.

[0340] Furthermore, the implementation of this technology is not limited to the above-described implementation, but can be changed in various ways without departing from the spirit of this technology.

[0341] For example, this technology can be configured as cloud computing, in which a function is shared among multiple devices via a network and processed by multiple devices that cooperate with each other.

[0342] Furthermore, in addition to being executed on a single device, each step explained in the flowchart above can be shared and executed by multiple devices.

[0343] Furthermore, in cases where a step includes multiple processes in addition to those executed on a single device, the multiple processes included in a single step can be shared among multiple devices and executed by multiple devices.

[0344] In addition, this technology may also have the following configurations.

[0345] (1) A signal processing apparatus, comprising:

[0346] The selection unit is provided with multiple audio signals and selects the audio signal to be processed for sound quality enhancement; and

[0347] The sound quality enhancement processing unit performs the sound quality enhancement processing on the audio signal selected by the selection unit.

[0348] (2) The signal processing apparatus according to (1), wherein the selection unit selects the audio signal to be subjected to the sound quality enhancement processing based on the metadata of the audio signal.

[0349] (3) The signal processing apparatus according to (2), wherein the metadata includes priority information representing the priority of the audio signal.

[0350] (4) The signal processing apparatus according to (2) or (3), wherein the metadata includes type information indicating the type of the audio signal.

[0351] (5) The signal processing apparatus according to any one of (2) to (4), further comprising:

[0352] The metadata generation unit generates the metadata based on the audio signal.

[0353] (6) The signal processing apparatus according to any one of (1) to (5), wherein, for each of the audio signals, the selection unit selects from a plurality of mutually different sound quality enhancement processes the sound quality enhancement process to be performed on the audio signal.

[0354] (7) The signal processing apparatus according to (6), wherein the sound quality enhancement processing includes dynamic range extension processing or bandwidth extension processing.

[0355] (8) The signal processing apparatus according to (6), wherein the sound quality enhancement processing includes prediction coefficients obtained by machine learning and dynamic range extension processing or bandwidth extension processing based on the audio signal.

[0356] (9) The signal processing apparatus according to (8) further comprises:

[0357] The coefficient selection unit retains the prediction coefficients for each type of audio signal and selects the prediction coefficients to be used for sound quality enhancement processing from the retained prediction coefficients based on type information indicating the type of audio signal.

[0358] (10) The signal processing apparatus according to (6), wherein the sound quality enhancement processing includes bandwidth expansion processing based on the audio signal to generate high-frequency components by linear prediction.

[0359] (11) The signal processing apparatus according to (6), wherein the sound quality enhancement processing includes bandwidth expansion processing of adding white noise to the audio signal.

[0360] (12) The signal processing apparatus according to any one of (1) to (11), wherein the audio signal includes an audio signal of a channel or an audio signal of an audio target.

[0361] (13) A signal processing method executed by a signal processing apparatus, the signal processing method comprising:

[0362] Provides multiple audio signals and allows selection of the audio signal to undergo sound quality enhancement processing; and

[0363] Perform sound quality enhancement processing on the selected audio signal.

[0364] (14) A program that causes a computer to perform a process, the process comprising:

[0365] The steps of providing multiple audio signals and selecting the audio signal to be subjected to sound quality enhancement processing; and

[0366] The steps involve performing sound quality enhancement processing on the selected audio signal.

[0367] [List of Reference Numbers]

[0368] 11: Signal processing device

[0369] 22: Audio Selection Department

[0370] 23: Sound Quality Enhancement Processing Department

[0371] 24: Renderer

[0372] 25: Reproduced Signal Generation Unit

[0373] 32-1~32-m, 32: High-load sound quality enhancement processing unit

[0374] 33-1~33-m, 33: Medium-load sound quality enhancement processing unit

[0375] 34-1~34-m, 34: Low-load sound quality enhancement processing unit

[0376] 301-1~301-m, 301: Coefficient Selection Section

[0377] 341-1~341-m, 341: Metadata Generation Department

Claims

1. A signal processing apparatus, comprising: The selection unit is provided with multiple audio signals and selects the audio signal to be subjected to sound quality enhancement processing. The sound quality enhancement processing includes high-load sound quality enhancement processing, low-load sound quality enhancement processing, and medium-load sound quality enhancement processing. For each audio signal, the selection unit selects the sound quality enhancement processing to be performed on the audio signal from the high-load sound quality enhancement processing, the low-load sound quality enhancement processing, and the medium-load sound quality enhancement processing. The sound quality enhancement processing unit performs the selected sound quality enhancement processing on the audio signal selected by the selection unit. The sound quality enhancement processing includes dynamic range extension processing or bandwidth extension processing of the audio signal based on prediction coefficients obtained through machine learning, and... The coefficient selection unit maintains the prediction coefficients for each type of audio signal and selects the prediction coefficients to be used for the high-load sound quality enhancement processing from the maintained plurality of prediction coefficients based on type information representing the type of the audio signal.

2. The signal processing apparatus according to claim 1, wherein, The selection unit selects the audio signal to be subjected to the sound quality enhancement processing based on the metadata of the audio signal.

3. The signal processing apparatus according to claim 2, wherein, The metadata includes priority information indicating the priority of the audio signal.

4. The signal processing apparatus according to claim 2, wherein, The metadata includes type information indicating the type of the audio signal.

5. The signal processing apparatus according to claim 2, further comprising: The metadata generation unit generates the metadata based on the audio signal.

6. The signal processing apparatus according to claim 1, wherein, The sound quality enhancement process includes bandwidth expansion processing based on the audio signal to generate high-frequency components through linear prediction.

7. The signal processing apparatus according to claim 1, wherein, The sound quality enhancement process includes bandwidth expansion processing that adds white noise to the audio signal.

8. The signal processing apparatus according to claim 1, wherein, The audio signal includes audio signals from the vocal tract or audio signals from the audio target.

9. A signal processing method executed by a signal processing device, the signal processing method comprising: Multiple audio signals are provided, and the audio signal to be subjected to sound quality enhancement processing is selected. The sound quality enhancement processing includes high-load sound quality enhancement processing, low-load sound quality enhancement processing, and medium-load sound quality enhancement processing. For each audio signal, select the audio quality enhancement process to be performed on the audio signal from the high-load audio quality enhancement process, the low-load audio quality enhancement process, and the medium-load audio quality enhancement process; Perform the selected sound quality enhancement processing on the selected audio signal. The sound quality enhancement processing includes dynamic range extension processing or bandwidth extension processing of the audio signal based on prediction coefficients obtained through machine learning, and... Prediction coefficients are maintained for each type of audio signal, and prediction coefficients to be used for the high-load sound quality enhancement processing are selected from the maintained plurality of prediction coefficients based on type information representing the type of the audio signal.

10. A computer-readable storage medium storing a program that causes a computer to perform the signal processing method according to claim 9.

Citation Information

Patent Citations

  • Device for expanding frequency band of input signal via up-sampling

    US9922660B2

  • Device, method, and program for expanding frequency band

    CN105745706A

  • Device, method, and program for audio reproduction

    JP2006350132A

  • Sound processing device

    JP2011203483A

  • Method and apparatus for controlling enhancement of low-bitrate coded audio

    WO2020047298A1