Audio signal processing method and device, computer equipment and readable storage medium
By constructing a filter bank to convert the sampling rate of the audio signal, the problem of high computational complexity in the prior art is solved, and the efficiency and accuracy of audio signal processing are improved.
Patent Information
- Application Number
- CN202511233801.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-01
AI Technical Summary
In the prior art, the computational complexity of audio signal sampling rate conversion is high, resulting in low processing efficiency and possible processing failure.
By determining the conversion factor sequence and filter type of the audio signal, constructing the filter identification sequence, and sequentially combining the filter groups, the sampling rate conversion of the audio signal is achieved, reducing the data volume and processing complexity.
The efficiency and accuracy of audio signal processing are improved, the computational complexity is reduced, and the overall processing efficiency and accuracy are improved.
Smart Images

Figure CN120748416A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing technology, and in particular to a method, apparatus, computer device, and readable storage medium for processing audio signals. Background Art
[0002] Audio sampling rate conversion is the process of converting an audio signal from one sampling rate to another. Related technologies often perform low-pass filtering and phase compensation on upsampled audio signals before downsampling them. However, these methods increase computational complexity, reducing audio signal processing efficiency and even leading to audio signal processing failure. Summary of the Invention
[0003] Based on this, it is necessary to provide an audio signal processing method, apparatus, computer equipment, computer-readable storage medium and computer program product that can improve the processing efficiency of audio signals in order to address the above technical problems.
[0004] In a first aspect, the present application provides a method for processing an audio signal, the method comprising:
[0005] Acquire an audio signal, and determine a current sampling rate and a target sampling rate of the audio signal;
[0006] determining, based on the current sampling rate and the target sampling rate, a conversion factor sequence for the audio signal, wherein the conversion factor sequence includes at least one conversion factor, and the conversion factor is used to represent a change in data volume after a filter performs conversion processing on input signal data;
[0007] determining a filter type for the audio signal according to a filter constraint, wherein the filter constraint is obtained based on at least one of an audio feature of the audio signal, a signal indicator feature corresponding to the target sampling rate, and a local audio processing resource feature;
[0008] Constructing a filter identification sequence for the audio signal according to the filter type; the filter identification sequence includes filter identifications arranged in sequence, the number of which is the same as the number of the conversion factors, the filter identifications being used to identify filters on the local end that belong to the filter type;
[0009] According to the conversion factors in the conversion factor sequence, output conversion configuration is performed on the filters represented by the filter identifiers in the filter identifier sequence to obtain filter configuration results;
[0010] Based on the filter configuration result, the local filters are sequentially combined to obtain a filter group, and the audio signal is subjected to sampling rate conversion processing by the filter group to obtain a target signal that meets the target sampling rate.
[0011] In one embodiment, determining a conversion factor sequence for the audio signal based on the current sampling rate and the target sampling rate includes:
[0012] Determining an interpolation factor and a decimation factor for the audio signal based on the current sampling rate and the target sampling rate, and obtaining a conversion factor according to the interpolation factor and the decimation factor;
[0013] The conversion factors are combined to obtain a conversion factor sequence for the audio signal.
[0014] In one embodiment, constructing a filter identification sequence for the audio signal according to the filter type includes:
[0015] determining the number of factors of the conversion factor included in the conversion factor sequence;
[0016] Determining, based on the filter types, a number of filter identifiers equal to the number of factors;
[0017] The filter identifiers are combined to obtain a filter identifier sequence for the audio signal.
[0018] In one embodiment, the conversion factor includes an interpolation factor and a decimation factor, the factor quantity includes a first factor quantity of the interpolation factor and a second factor quantity of the decimation factor; the filter type includes an interpolation type and a decimation type;
[0019] The determining, based on the filter type, a number of filter identifiers that is equal to the number of factors, includes:
[0020] Determining, based on the interpolation type, first filter identifiers whose number is equal to the number of the first factors;
[0021] determining, based on the decimation type, second filter identifiers, the same number as the number of the second factors;
[0022] According to the first filter identifier and the second filter identifier, filter identifiers whose number is the same as the number of the factors are obtained.
[0023] In one embodiment, the conversion factor includes an interpolation factor and a decimation factor, and the filter type includes an interpolation type and a decimation type;
[0024] The step of performing output conversion configuration on the filters represented by the filter identifiers in the filter identifier sequence according to the conversion factors in the conversion factor sequence to obtain filter configuration results includes:
[0025] When the interpolation factor indicates that the filter of the interpolation type meets the trigger configuration condition, performing output conversion configuration on a first filter according to the interpolation factor to obtain a first configuration result, where the first filter is a filter of the interpolation type represented by the filter identifier in the filter identifier sequence;
[0026] When the decimation factor indicates that the filter of the decimation type meets the trigger configuration condition, performing output conversion configuration on a second filter according to the decimation factor to obtain a second configuration result, where the second filter is a filter of the decimation type represented by the filter identifier in the filter identifier sequence;
[0027] A filter configuration result is obtained according to the first configuration result and the second configuration result.
[0028] In one embodiment, the filter type includes an interpolation type and a decimation type;
[0029] The step of sequentially combining the filters at the local end based on the filter configuration result to obtain a filter bank includes:
[0030] Determining a target filter at a local end and a filter cascade order of the target filter according to the filter configuration result, the filter cascade order including a first cascade order of a first filter of the interpolation type and a second cascade order of a second filter of the decimation type in the target filter;
[0031] cascade-connecting the first filters in the first cascade order to obtain an interpolation filter bank;
[0032] cascade-connecting the second filters according to the second cascade order to obtain a decimation filter bank;
[0033] The interpolation filter group and the decimation filter group are connected in sequence to obtain a filter group.
[0034] In one embodiment, performing sampling rate conversion on the audio signal by the filter bank includes:
[0035] Obtaining the number of taps of each target filter in the filter bank;
[0036] Determining, according to the number of taps of each target filter, an audio processing algorithm corresponding to each target filter;
[0037] The target filters are used to perform sampling rate conversion processing on the audio signal based on the cascade order of the target filters and the audio processing algorithms corresponding to the target filters.
[0038] In a second aspect, the present application further provides an audio signal processing device, the device comprising:
[0039] An audio acquisition module, configured to acquire an audio signal and determine a current sampling rate and a target sampling rate of the audio signal;
[0040] a conversion factor sequence determining module, configured to determine a conversion factor sequence for the audio signal based on the current sampling rate and the target sampling rate, wherein the conversion factor sequence includes at least one conversion factor, and the conversion factor is used to represent a change in data volume after the filter performs conversion processing on the input signal data;
[0041] a filter type determination module, configured to determine a filter type for the audio signal based on a filter constraint condition, wherein the filter constraint condition is obtained based on at least one of an audio feature of the audio signal, a signal indicator feature corresponding to the target sampling rate, and a local audio processing resource feature;
[0042] a filter identification sequence determination module, configured to construct a filter identification sequence for the audio signal according to the filter type; the filter identification sequence comprising filter identifications arranged in sequence, the number of which being equal to the number of the conversion factors, the filter identifications being used to identify filters belonging to the filter type at the local end;
[0043] a filter configuration module, configured to perform output conversion configuration on the filters represented by the filter identifiers in the filter identifier sequence according to the conversion factors in the conversion factor sequence, to obtain a filter configuration result;
[0044] The audio processing module is used to sequentially combine the local filters to obtain a filter group based on the filter configuration result, and perform sampling rate conversion processing on the audio signal through the filter group to obtain a target signal that meets the target sampling rate.
[0045] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the first aspect when executing the computer program.
[0046] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the above-mentioned first aspect when executed by a processor.
[0047] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which implements the steps in the above-mentioned first aspect when executed by a processor.
[0048] In the audio signal processing method, apparatus, computer device, computer-readable storage medium, and computer program product provided by the present application, the audio signal processing method obtains an audio signal and determines a current sampling rate and a target sampling rate of the audio signal; based on the current sampling rate and the target sampling rate, determines a conversion factor sequence for the audio signal, the conversion factor sequence including at least one conversion factor, the conversion factor being used to characterize a change in the amount of data after the filter converts and processes the input signal data; determines a filter type for the audio signal based on filter constraints, the filter constraints being based on at least one of an audio feature of the audio signal, a signal indicator feature corresponding to the target sampling rate, and a local audio processing resource feature; constructs a filter identifier sequence for the audio signal according to the filter type; the filter identifier sequence includes filter identifiers arranged in sequence, the number of which is the same as the number of conversion factors, the filter identifiers being used to characterize filters belonging to the filter type at the local end; performs output conversion configuration on the filters represented by the filter identifiers in the filter identifier sequence according to the conversion factors in the conversion factor sequence, thereby obtaining a filter configuration result; sequentially combines the local filters to obtain a filter bank based on the filter configuration result, and performs sampling rate conversion processing on the audio signal using the filter bank to obtain a target signal that meets the target sampling rate. It can be seen that the present application can convert the current sampling rate and target sampling rate corresponding to the audio signal to be processed to obtain the corresponding conversion factor sequence, and the local end resource of the audio signal can be combined to determine the type of filter that can be adapted and available, and then based on the determined conversion factor sequence and filter type, a filter group for realizing audio signal processing can be obtained, which involves the cascade connection relationship of each filter and the conversion configuration of each filter. Finally, the audio signal is gradually sampled and converted by each filter in the filter group to obtain the target signal output by the filter group that meets the target sampling rate. In the audio signal processing method provided by the present application, the number of conversion factors obtained based on the current sampling rate and target sampling rate of the audio signal to be processed includes multiple cases, which makes the corresponding filters also include the case of using multiple cases. The conversion of the audio signal sampling rate is gradually realized by multiple cascaded filters, which is beneficial to reducing the amount of data and processing complexity of the audio signal processed by filters at all levels, thereby improving the efficiency and accuracy of audio signal processing by filters at all levels, and thus improving the overall processing efficiency and processing accuracy of the audio signal. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.
[0050] Figure 1 FIG1 is an application environment diagram of a method for processing an audio signal in one embodiment;
[0051] Figure 2 1 is a flow chart of a method for processing an audio signal in one embodiment;
[0052] Figure 3 is a flowchart of sub-steps of a method for processing an audio signal in one embodiment;
[0053] Figure 4 is a structural block diagram of an audio signal processing device in one embodiment;
[0054] Figure 5 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment.
[0055] Description of reference numerals:
[0056] 102 - terminal, 104 - server, 400 - audio signal processing device, 41 - audio acquisition module, 42 - conversion factor sequence determination module, 43 - filter type determination module, 44 - filter identification sequence determination module, 45 - filter configuration module, 46 - audio processing module. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0058] It should be noted that the terms "first", "second", etc. used in this application may be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "including" and "having" used in this application and any variations thereof are intended to cover non-exclusive inclusions. The term "plurality" used in this application refers to two or more. The term "and / or" used in this application refers to one of the solutions or any combination of multiple solutions.
[0059] The audio signal processing method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store data that server 104 needs to process. The data storage system can be integrated on server 104, or placed on the cloud or other network servers. The data storage system can be used to store information such as the filter type of each filter and the characteristics of local audio processing resources. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart car devices, projectors, etc. Portable wearable devices can include smart watches, smart bracelets, head-mounted devices, etc. Head-mounted devices can include virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. Server 104 can be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services.
[0060] In an exemplary embodiment, Figure 2 As shown, a method for processing an audio signal is provided, and the method is applied to Figure 1 The terminal 102 in the example is used as an example to illustrate the process, including the following steps 201 to 206. Among them:
[0061] Step 201: Acquire an audio signal and determine the current sampling rate and target sampling rate of the audio signal.
[0062] Among them, the acquired audio signal is also the audio signal to be processed, specifically the audio signal to be subjected to audio sampling rate conversion. Among them, the current sampling rate of the audio signal is the sampling rate of the audio signal to be subjected to sampling rate conversion; the target sampling rate of the audio signal is the sampling rate of the audio signal obtained after the sampling rate conversion. Here, the current sampling rate and the target sampling rate of the audio signal are both known, or can be obtained by analyzing the audio signal itself, and at least one of the storage device and the final playback device of the audio signal. Any of the current sampling rate and the target sampling rate may be an integer or may not be an integer, and this application does not limit this.
[0063] Exemplarily, for an audio signal to be subjected to audio sampling rate conversion, after obtaining the audio signal, its corresponding current audio sampling rate and target audio sampling rate are first determined, so as to facilitate the subsequent determination of the filters to be used, the configuration information of each filter, and the connection relationship between the filters based on the current sampling rate and the target sampling rate.
[0064] Step 202: Determine a conversion factor sequence for the audio signal based on the current sampling rate and the target sampling rate. The conversion factor sequence includes at least one conversion factor, which is used to represent the change in data volume after the filter converts the input signal data.
[0065] The conversion factor sequence may include only one conversion factor or multiple conversion factors, which is not limited in this application. The number of conversion factors may be determined based on the current sampling rate and the target sampling rate, and may be further determined in association with the number of filters included in the terminal. If required, it may also be determined in combination with other factors (such as the audio processing resources of the terminal). The conversion factor is used to represent the change in the amount of data after the filter converts the input signal data. For example, when the conversion factor is 3, it means that when 1 signal data is input to the corresponding filter, 3 signal data will be output accordingly. For example, when the conversion factor is 4, it means that when 1 signal data is input to the corresponding filter, 4 signal data will be output accordingly. In addition, when the conversion factor is 3, it may also mean that when 3 signal data are input to the corresponding filter, 1 signal data will be output accordingly. When the conversion factor is 5, it may also mean that when 5 signal data are input to the corresponding filter, 1 signal data will be output accordingly.
[0066] Exemplarily, after determining the current sampling rate and target sampling rate of the audio signal to be converted, at least one conversion factor adapted to the audio signal can be obtained based on the numerical conversion of the current sampling rate and the target sampling rate, and then a conversion factor sequence corresponding to the audio signal can be composed based on these conversion factors.
[0067] Step 203: Determine the filter type for the audio signal based on the filter constraint condition, where the filter constraint condition is obtained based on at least one of the audio characteristics of the audio signal, the signal index characteristics corresponding to the target sampling rate, and the audio processing resource characteristics of the local end.
[0068] The filter type refers to the filter level, such as Class A filter, Class B filter, etc. The audio features of the audio signal may include, for example, the signal spectrum, signal bandwidth, channel bandwidth, signal frequency, and other features of the audio signal; the signal index features corresponding to the target sampling rate refer to the features required for the target audio signal obtained after the sampling rate conversion, such as signal-to-noise ratio and passband ripple; the local audio processing resource features may include, for example, the computing resources of the MCU (Microcontroller Unit) and the storage resources of the storage unit, etc.
[0069] For example, when the sampling rate of an audio signal needs to be processed, at least one of the audio characteristics of the audio signal to be processed, the signal index characteristics required to be achieved by the target audio signal, and the audio processing resource characteristics of the local source used for audio processing can be obtained first, and then the filter type of the filter required to be used for the audio signal processing this time can be determined based on at least one of these three items.
[0070] Step 204 : construct a filter identification sequence for the audio signal according to the filter type. The filter identification sequence includes filter identifications arranged in sequence, the number of which is the same as the number of conversion factors. The filter identification is used to identify the filter of the local end belonging to the filter type.
[0071] The filter identifier sequence includes the filter identifiers of the filters required for processing the audio data. The filter identifier can be, for example, a filter ID (identity document) that identifies the filter. The number of filter identifiers included in the filter identifier sequence is the same as the number of conversion factors corresponding to the audio signal to be processed.
[0072] Exemplarily, after determining the filter types that can be used to process audio signals, the filter types suitable for each conversion factor can be determined based on the conversion factors corresponding to the audio signal, and then the filter identifiers of each filter to be used can be determined. The filter identifier sequence is formed by combining the determined filter identifiers that match the number of conversion factors.
[0073] Step 205 : According to the conversion factors in the conversion factor sequence, output conversion configuration is performed on the filters represented by the filter identifiers in the filter identifier sequence to obtain a filter configuration result.
[0074] Exemplarily, after determining the conversion factor sequence and filter identifier sequence corresponding to the audio signal to be processed through the above steps, it is necessary to match the conversion factors in the adapted conversion factor sequence with the filter identifiers in the filter identifier sequence to obtain a corresponding conversion factor combined with a filter identifier combination; in each set of conversion factor combined with a filter identifier combination, the conversion factor is used to configure the working function of the filter pointed to by the relevant filter identifier. For example, when the conversion factor is 3, the working function of the corresponding configured filter is: when the filter inputs 1 signal data, it will output 3 signal data, or when the filter inputs 3 signal data, it will output 1 signal data; for example, when the conversion factor is 5, the working function of the corresponding configured filter is: when the filter inputs 1 signal data, it will output 5 signal data, or when the filter inputs 5 signal data, it will output 1 signal data. By matching each conversion factor in the conversion factor sequence with each filter identifier in the filter identifier sequence one by one, the filter configuration result for the audio signal to be processed can be obtained.
[0075] Step 206 : Based on the filter configuration result, the local filters are sequentially combined to obtain a filter bank, and the audio signal is subjected to sampling rate conversion processing by the filter bank to obtain a target signal that meets the target sampling rate.
[0076] The purpose of sequentially combining the filters at the local end is to determine a connection method of the multiple filters when there are multiple filters for processing the audio signal, specifically, to determine a cascade connection method of the multiple filters.
[0077] Exemplarily, after determining the filters required for sampling rate conversion of the audio signal to be processed and the configuration information of each filter, the electrical connection method between the multiple filters can be further determined, specifically the cascade connection method, and then based on the determined cascade connection method, the multiple filters are cascaded to form a filter group, and then the audio signal is processed in an orderly manner one level at a time through each filter in the filter group to obtain the target signal with sampling rate conversion that is finally output by the filter group, and the target signal has the target sampling rate required by the audio data.
[0078] In the audio signal processing method provided by the present application, an audio signal is obtained and a current sampling rate and a target sampling rate of the audio signal are determined; based on the current sampling rate and the target sampling rate, a conversion factor sequence for the audio signal is determined, the conversion factor sequence includes at least one conversion factor, and the conversion factor is used to characterize the change in data volume after the filter converts and processes the input signal data; the filter type for the audio signal is determined according to the filter constraint condition, and the filter constraint condition is obtained based on at least one of the audio characteristics of the audio signal, the signal index characteristics corresponding to the target sampling rate, and the audio processing resource characteristics of the local end; a filter identification sequence for the audio signal is constructed according to the filter type; the filter identification sequence includes filter identifications arranged in sequence and the same number as the number of conversion factors, and the filter identification is used to characterize the filters belonging to the filter type on the local end; according to the conversion factors in the conversion factor sequence, output conversion configuration is performed on the filters represented by the filter identifications in the filter identification sequence to obtain a filter configuration result; based on the filter configuration result, the filters on the local end are sequentially combined to obtain a filter group, and the audio signal is sampled and converted by the filter group to obtain a target signal that meets the target sampling rate. It can be seen that the present application can convert the current sampling rate and target sampling rate corresponding to the audio signal to be processed to obtain the corresponding conversion factor sequence, and the local end resource of the audio signal can be combined to determine the type of filter that can be adapted and available, and then based on the determined conversion factor sequence and filter type, a filter group for realizing audio signal processing can be obtained, which involves the cascade connection relationship of each filter and the conversion configuration of each filter. Finally, the audio signal is gradually sampled and converted by each filter in the filter group to obtain the target signal output by the filter group that meets the target sampling rate. In the audio signal processing method provided by the present application, the number of conversion factors obtained based on the current sampling rate and target sampling rate of the audio signal to be processed includes multiple cases, which makes the corresponding filters also include the case of using multiple cases. The conversion of the audio signal sampling rate is gradually realized by multiple cascaded filters, which is beneficial to reducing the amount of data and processing complexity of the audio signal processed by filters at all levels, thereby improving the efficiency and accuracy of audio signal processing by filters at all levels, and thus improving the overall processing efficiency and processing accuracy of the audio signal.
[0079] It should be added that the multiple filters configured in the terminal 102 of the above embodiment can be multiple physical filters; and then, based on demand, the multiple physical filters are cascaded to form a filter group, so that the audio signal is processed in an orderly manner one level at a time through each physical filter in the physical filter group, and the target signal of the audio signal is obtained by the physical filter group after sampling rate conversion.
[0080] However, configuring multiple physical filters in the terminal 102 is only an optional implementation provided by the present application, and the present application is not limited thereto. Alternatively, the terminal 102 may be configured with only one physical filter. When multiple (multi-stage) filters are required to sequentially process the audio signal, the parameter configuration information of each stage of the filter may be updated through the parameter control module of the filter. For example, when the same filter is required to implement three-stage processing of the audio signal, three configuration information (first configuration information, second configuration information, and third configuration information) corresponding to the three-stage filters may be determined based on the filter configuration result. Then, the parameter control module of the filter first configures the filter parameters based on the first configuration information to perform a first sampling rate conversion on the received audio signal to be processed to obtain a first sampling rate converted audio signal. Further, the parameter control module of the filter configures the filter parameters based on the second configuration information to perform a corresponding second sampling rate conversion on the first sampling rate converted audio signal to obtain a second sampling rate converted audio signal. Finally, the parameter control module of the filter configures the filter parameters based on the third configuration information to perform a corresponding third sampling rate conversion on the second sampling rate converted audio signal to obtain a third sampling rate converted audio signal. The third sampling rate converted audio signal is also the target signal meeting the target sampling rate. In this embodiment, it is allowed that only one filter is provided in the terminal 102. By multiplexing one filter and configuring appropriate parameters for it in each use, multi-stage orderly processing of an audio signal can be achieved through one filter.
[0081] In an exemplary embodiment, the above-mentioned determination of the conversion factor sequence for the audio signal based on the current sampling rate and the target sampling rate may specifically include: determining the interpolation factor and the decimation factor for the audio signal based on the current sampling rate and the target sampling rate, and obtaining the conversion factor based on the interpolation factor and the decimation factor; and combining the conversion factors to obtain the conversion factor sequence for the audio signal.
[0082] The conversion factors may include only interpolation factors, or only decimation factors, or both interpolation factors and decimation factors. The interpolation factors are used to characterize that the amount of data after the adaptive filter performs conversion processing on the input signal data is increased; the decimation factors are used to characterize that the amount of data after the adaptive filter performs conversion processing on the input signal data is reduced. When the interpolation factor takes a value of 1, it is not necessary to increase or decrease the amount of the relevant audio data. Therefore, when the interpolation factor takes a value of 1, the filter is not suitable. Correspondingly, when the decimation factor takes a value of 1, it is not necessary to increase or decrease the amount of the relevant audio data. Therefore, when the decimation factor takes a value of 1, the filter is not suitable. The conversion factor sequence may include only interpolation factors, only decimation factors, or both decimation factors and interpolation factors.
[0083] Exemplarily, when the current sampling rate and target sampling rate of the audio signal to be processed are obtained, the interpolation factor and decimation factor suitable for processing the audio data can be determined based on the value of the current sampling rate and the value of the target sampling rate, so as to obtain a conversion factor sequence suitable for processing the audio data based on the determined combination of the interpolation factor and the decimation factor.
[0084] An embodiment is provided in which multiple conversion factors are determined based on the current sampling rate and the target sampling rate of the audio signal, and the conversion factor sequences formed therefrom may include multiple conversion factors; the number and values of the conversion factors included in these multiple conversion factor sequences are the same, but the order of the multiple conversion factors in each conversion factor sequence may be different.
[0085] Another embodiment is provided in which a conversion factor sequence composed of multiple conversion factors determined based on the current sampling rate and the target sampling rate of the audio signal can be unique, and the order of the multiple conversion factors in the conversion factor sequence is the most preferred order that has been determined.
[0086] Among them, the most preferred conversion factor sequence can be determined based on at least one of a preset method or a preset screening method. This application does not specifically limit the preset method and preset screening method here. Users can select and set the method for determining the most preferred conversion factor sequence based on needs or actual experience.
[0087] In this embodiment, the adapted interpolation factor and decimation factor are determined by the current sampling rate and target sampling rate of the audio signal to be processed, and then a conversion factor sequence suitable for the audio signal to be processed is obtained based on the determined combination of the interpolation factor and decimation factor, which can ensure the accuracy of the conversion factors included in the conversion factor sequence and the subsequent matching accuracy of the corresponding filter.
[0088] In an exemplary embodiment, the present application also provides a method for determining the interpolation factor and decimation factor for an audio signal based on the current sampling rate and the target sampling rate, including: determining the interpolation coefficient and decimation coefficient corresponding to the current sampling rate and the target sampling rate based on the simplest integer ratio of the current sampling rate and the target sampling rate; obtaining at least one factor corresponding to each interpolation coefficient as the corresponding at least one interpolation factor, and obtaining at least one factor corresponding to each decimation coefficient as the corresponding at least one decimation factor.
[0089] The at least one factor corresponding to each interpolation coefficient may or may not be the minimum factor, and this application does not impose any specific restrictions on this. Correspondingly, the at least one factor corresponding to each decimation coefficient may or may not be the minimum factor, and this application does not impose any specific restrictions on this. The conversion factors included in the conversion factor sequence may, for example, be arranged in ascending order of numerical value.
[0090] In an exemplary embodiment, the above-mentioned construction of a filter identification sequence for an audio signal according to filter type specifically includes: determining the number of conversion factors included in the conversion factor sequence; determining, based on the filter type, filter identifications whose number is the same as the number of factors; and combining the filter identifications to obtain a filter identification sequence for the audio signal.
[0091] The number of factors refers to the number of conversion factors included in the conversion factor sequence.
[0092] Exemplarily, when the conversion factor sequence and filter type adapted for audio data are determined, the number of conversion factors in the conversion factor sequence can be further obtained, and then, the corresponding filter is adapted to each conversion factor based on demand to determine a filter identifier equal to the number of factors, and finally, these determined filter identifiers are combined to obtain a filter identifier sequence adapted for the audio signal.
[0093] An embodiment is provided in which, based on the number of conversion factors and filter types adapted to the audio signal, a filter identification sequence composed of determined filter identifications may include multiple filter identifications; the number of filter identifications and filter types included in these multiple filter identification sequences are the same, but the order of the multiple filter identifications in each filter identification sequence may be different.
[0094] Another embodiment is provided in which, based on the number of conversion factors and filter types adapted to the audio signal, the filter identification sequence composed of the determined filter identifications can be unique, and the order of multiple filter identifications in the filter identification sequence is the most preferred order that has been determined.
[0095] Among them, the most preferred filter identification sequence can be determined based on at least one of a preset method or a preset screening method. This application does not specifically limit the preset method and preset screening method here. The user can select and set the method for determining the most preferred filter identification sequence based on needs or actual experience.
[0096] In this embodiment, the filter identifier adapted for the audio data to be processed is determined by the number of factors and filter types adapted for the conversion factors of the audio data to be processed, and then the filter identifiers are combined to obtain a filter identifier sequence adapted for the audio data to be processed. This ensures that the number of filters corresponding to the determined filter identifiers is sufficient for processing the audio signal to be processed, and also ensures the subsequent matching effect of the conversion factors and filters adapted for the audio signal.
[0097] In an exemplary embodiment, the conversion factor includes an interpolation factor and a decimation factor, the factor number includes a first factor number of the interpolation factor and a second factor number of the decimation factor; and the filter type includes an interpolation type and a decimation type.
[0098] The first number of factors is the number of interpolation factors adapted to the audio data to be processed, and the second number of factors is the number of decimation factors adapted to the audio data to be processed; either the first number of factors or the second number of factors may be 0. An interpolation filter is an interpolation filter, which is used to add data to the input to output an output result with a larger amount of data than the input data. A decimation filter is a decimation filter, which is used to subtract data from the input to output an output result with a smaller amount of data than the input data.
[0099] In one embodiment, the conversion factor sequence can be composed based on a combination of an interpolation factor sequence and an extraction factor sequence, wherein the interpolation factors included in the interpolation factor sequence can be arranged in ascending order of numerical value, and the extraction factors included in the extraction factor sequence can be arranged in ascending order of numerical value, or in descending order of numerical value. Then, the two sequences are integrated in the order of the interpolation factor sequence first and the extraction factor sequence second to obtain the conversion factor sequence.
[0100] The above-mentioned method of determining the filter identifier whose number is the same as the number of factors based on the filter type specifically includes: determining the first filter identifier whose number is the same as the number of first factors based on the interpolation type; determining the second filter identifier whose number is the same as the number of second factors based on the extraction type; and obtaining the filter identifier whose number is the same as the number of factors based on the first filter identifier and the second filter identifier.
[0101] Exemplarily, after obtaining the interpolation factors and decimation factors adapted to the audio signal to be processed, as well as the interpolation type and decimation type of the filter adapted to the audio signal to be processed, the first filter identifiers whose number is the same as the number of first factors (the number of interpolation factors) can be determined based on the interpolation type. At the same time, the second filter identifiers whose number is the same as the number of second factors (the number of decimation factors) can be determined based on the decimation type. Then, based on the combination of the first filter identifiers and the second filter identifiers, the filter identifiers whose number is the same as the number of factors in the conversion factor sequence can be obtained. The determined filter identifiers can be used to combine to obtain a filter identifier sequence adapted to the audio signal to be processed.
[0102] Among them, when the number of interpolation factors is 1 and the value of the interpolation factor is also 1, it means that there is no need to perform interpolation processing on the audio signal to be processed. In this case, the interpolation filter can be omitted in the processing of the audio signal to be processed, so the corresponding first filter identifier is "empty"; correspondingly, when the number of decimation factors is 1 and the value of the decimation factor is also 1, it means that there is no need to perform decimation processing on the audio signal to be processed. In this case, the decimation filter can be omitted in the processing of the audio signal to be processed, so the corresponding second filter identifier is "empty".
[0103] In this embodiment, by first determining a first filter identifier (the filter identifier of the interpolation filter) that is adapted to the audio signal to be processed, and a second filter identifier (the filter identifier of the extraction filter) that is adapted to the audio signal to be processed, and then combining the first filter identifier and the second filter identifier to obtain filter identifiers whose number is the same as the number of factors in the conversion factor sequence that is adapted to the audio data to be processed, and by determining the filter identifier of the interpolation filter and the filter identifier of the extraction filter in steps, it is beneficial to improve the accuracy of all the filter identifiers that are adapted to the audio signal to be processed, and thus it is beneficial to improve the accuracy of the target signal that is finally obtained after the sampling rate conversion of the audio signal to be processed.
[0104] In an exemplary embodiment, the conversion factor includes an interpolation factor and a decimation factor, and the filter type includes an interpolation type and a decimation type; the above-mentioned output conversion configuration is performed on the filters represented by the filter identifiers in the filter identification sequence according to the conversion factors in the conversion factor sequence to obtain a filter configuration result, specifically including: when the interpolation factor indicates that the filter of the interpolation type meets the trigger configuration condition, the output conversion configuration is performed on the first filter according to the interpolation factor to obtain a first configuration result, and the first filter is the filter of the interpolation type represented by the filter identifier in the filter identification sequence; when the decimation factor indicates that the filter of the decimation type meets the trigger configuration condition, the output conversion configuration is performed on the second filter according to the decimation factor to obtain a second configuration result, and the second filter is the filter of the decimation type represented by the filter identifier in the filter identification sequence; the filter configuration result is obtained based on the first configuration result and the second configuration result.
[0105] Among them, "not meeting the trigger configuration conditions" associated with "meeting the trigger configuration conditions" includes two situations, one of which is: the number of interpolation factors is 1, and the value of the interpolation factor is also 1, and the other is: the number of decimation factors is 1, and the value of the decimation factor is also 1. The reason corresponds to the above: when the number of interpolation factors is 1, and the value of the interpolation factor is also 1, it means that there is no need to perform interpolation processing on the audio signal to be processed. In this case, the interpolation filter may not be used in the processing of the audio signal to be processed, so the corresponding first filter identifier may be "empty"; correspondingly, when the number of decimation factors is 1, and the value of the decimation factor is also 1, it means that there is no need to perform decimation processing on the audio signal to be processed. In this case, the decimation filter may not be used in the processing of the audio signal to be processed, so the corresponding second filter identifier may be "empty".
[0106] It should be added that, when the number of interpolation factors is 1 and the value of the interpolation factor is also 1, the interpolation filter is not used to process the audio signal, which can reduce the number of filters required to be used in the audio signal processing process; similarly, when the number of decimation factors is 1 and the value of the decimation factor is also 1, the decimation filter is not used to process the audio signal, which can reduce the number of filters required to be used in the audio signal processing process.
[0107] Exemplarily, when the conversion factors in the conversion factor sequence include both interpolation factors and decimation factors, and the filter identifiers in the filter identifier sequence involve both interpolation filters and decimation filters, the interpolation filters and decimation filters are configured separately. Specifically, you can choose to configure the functions of the interpolation filters first, or you can choose to configure the functions of the decimation filters first, or you can choose to configure the functions of both types of filters simultaneously, and this application does not limit this.
[0108] Specifically, when the interpolation factor indicates that an interpolation-type filter meets the trigger configuration condition, that is, when the interpolation factor is not "1 in number and 1 in value", output conversion configuration can be performed for each first filter (the filter of the interpolation type represented by the filter identifier in the filter identifier sequence) based on the interpolation factors involved, to obtain a first configuration result, which includes a combination of each interpolation factor and the corresponding first filter. And when the decimation factor indicates that a decimation-type filter meets the trigger configuration condition, that is, when the decimation factor is not "1 in number and 1 in value", output conversion configuration can be performed for each second filter (the filter of the decimation type represented by the filter identifier in the filter identifier sequence) based on the decimation factors involved, to obtain a second configuration result, which includes a combination of each decimation factor and the corresponding second filter. Finally, by integrating the first configuration result and the second configuration result, a filter configuration result suitable for the audio signal to be processed can be obtained.
[0109] In this embodiment, by configuring the interpolation factor of the interpolation filter and the decimation factor of the decimation filter respectively, and then integrating the two configuration results, the configuration accuracy of the final filter configuration result can be improved, which is beneficial to improving the accuracy of the target signal obtained after the sampling rate conversion of the processed audio signal.
[0110] Among them, when the conversion factor sequence determined in the previous step is one and the filter identification sequence determined is one, the conversion factors in the conversion factor sequence are combined and matched one by one with the filter identifications in the filter identification sequence in order to obtain the final filter configuration result.
[0111] In the case where a single conversion factor sequence is determined in the preceding step but multiple filter identifier sequences are determined, the most suitable filter identifier sequence may be determined first, and then each conversion factor in the conversion factor sequence may be sequentially combined and matched with each filter identifier in the filter identifier sequence to obtain a final filter configuration result. Alternatively, each conversion factor in the conversion factor sequence may be sequentially combined and matched with each filter identifier in multiple filter identifier sequences to obtain multiple filter configuration sub-results, and then a final filter configuration result may be selected from the multiple filter configuration sub-results.
[0112] In the case where multiple conversion factor sequences are determined in the above steps but only one filter identifier sequence is determined, the most suitable conversion factor sequence may be determined first, and then the conversion factors in the conversion factor sequence may be sequentially combined and matched with the filter identifiers in the filter identifier sequence. Alternatively, each filter identifier in the filter identifier sequence may be sequentially combined and matched with the conversion factors in multiple conversion factor sequences, resulting in multiple filter configuration sub-results, and then a final filter configuration result may be selected from the multiple filter configuration sub-results.
[0113] In the case where multiple conversion factor sequences and multiple filter identifier sequences are determined in the preceding steps, the most suitable conversion factor sequence and the most suitable filter identifier sequence may be determined separately, and then each conversion factor in the conversion factor sequence may be sequentially combined and matched with each filter identifier in the filter identifier sequence to obtain a final filter configuration result. Alternatively, each conversion factor in multiple conversion factor sequences may be sequentially combined and matched with each filter identifier in multiple filter identifier sequences to obtain multiple filter configuration sub-results, and then a final filter configuration result may be selected from the multiple filter configuration sub-results.
[0114] In an exemplary embodiment, the filter types include interpolation types and decimation types; the above-mentioned method of combining the local filters in sequence based on the filter configuration result to obtain a filter group specifically includes: determining the local target filter and the filter cascade order of the target filter according to the filter configuration result, the filter cascade order including the first cascade order of the first filter of the interpolation type in the target filter and the second cascade order of the second filter of the decimation type; cascading the first filter in the first cascade order to obtain an interpolation filter group; cascading the second filter in the second cascade order to obtain a decimation filter group; and connecting the interpolation filter group and the decimation filter group in a certain order to obtain a filter group.
[0115] Exemplarily, according to the filter configuration result, a target filter is determined among multiple filters on the local end that needs to participate in the sampling rate conversion of the audio data to be processed. When the target filter includes both a first filter of the interpolation type and a second filter of the decimation type, the first cascade order of the first filter of the interpolation type and the second cascade order of the second filter of the decimation type can be determined respectively; and further, the relevant multiple first filters are connected in the first cascade order to obtain an interpolation filter group, and the relevant multiple second filters are connected in the second cascade order to obtain a decimation filter group. Finally, the obtained interpolation filter group and the decimation filter group are connected in the order of the interpolation filter group in front and the decimation filter group in the back to obtain the final filter group.
[0116] By setting the filter group connection order with the interpolation filter group in front and the decimation filter group in the back, during the sampling rate conversion of the audio signal to be processed, the audio signal is first interpolated and increased, and then the audio signal is decimated and reduced. This method is conducive to improving the accuracy of the final target signal.
[0117] In this embodiment, by first determining the cascade connection order of each interpolation filter in the interpolation filter group and the cascade connection order of each decimation filter in the decimation filter group, and then obtaining the filter group in the order of connection with the interpolation filter group in front and the decimation filter group in the back, the cascade adaptability of the filter group finally obtained can be improved, which is beneficial to improving the accuracy of the target signal obtained after the sampling rate conversion of the audio signal to be processed.
[0118] In an exemplary embodiment, please refer to Figure 3 , performing sampling rate conversion processing on the audio signal through the filter bank, including steps 301 to 303, wherein: step 301, obtaining the number of taps of each target filter in the filter bank; step 302, determining the audio processing algorithm corresponding to each target filter according to the number of taps of each target filter; step 303, using each target filter, based on the cascade order of each target filter and the audio processing algorithm corresponding to each target filter, performing sampling rate conversion processing on the audio signal.
[0119] The number of taps of the filter refers to the number of filter coefficients, which can be used to determine a related audio processing algorithm used by the related filter to process input data.
[0120] Exemplarily, the number of taps of each target filter in the filter group can be obtained first, and then the audio processing algorithm corresponding to each target filter can be determined based on the number of taps of each target filter and the filter type (interpolation filter and decimation filter) of each target filter. In this way, when the audio signal is transmitted to each target filter, each target filter can use the determined audio processing algorithm to process the corresponding audio signal to output the required audio processing result; that is, when the audio signal to be processed starts to undergo sampling rate conversion, each target filter will be used to perform sampling rate conversion on the audio signal step by step based on the cascade order of each target filter and the audio processing algorithm corresponding to each target filter to obtain the target signal that finally completes the sampling rate conversion.
[0121] In this embodiment, by determining the audio processing algorithm suitable for each target filter based on the number of taps of each target filter, it is beneficial to ensure that each target filter can perform the most appropriate processing on the audio signal when receiving the audio signal to output the required processing result, which can improve the accuracy of the target signal that finally completes the sampling rate conversion.
[0122] Table 1 below provides several optional embodiments of the audio signal processing method provided herein. Table 1 shows the relationship between the interpolation coefficient (Interpolation-1) and the decimation coefficient (Decimation-1) for various input sampling rates (current sampling rate, Input Fs) and output sampling rates (target sampling rate, Output Fs) of 48 kHz (kilohertz). The unit of Input Fs can also be kHz.
[0123] Table 1
[0124]
[0125] Furthermore, Table 2 corresponding to Table 1 is provided, which shows a relationship table between the input sampling rate, output sampling rate, conversion factor sequence, and filter identification sequence corresponding to the above embodiment. Wherein, when the input sampling rate is 88.2 Hz, the corresponding "2245" in Interpolation is the 4 interpolation factors included in the interpolation factor sequence in the corresponding conversion factor sequence, the corresponding "147" in Decimation is the 4 decimation factors included in the decimation factor sequence in the corresponding conversion factor sequence, and the corresponding "ABAAC" in Interpolation and Decimation together is the corresponding filter identification sequence, where "ABAA" can be the filter identification sequence of the interpolation filter included in the filter identification sequence, and "C" can be the filter identification sequence of the decimation filter included in the filter identification sequence.
[0126] Among them, A represents A-level filter, B represents B-level filter, C represents C-level filter, and D represents D-level filter. Both A-level filter and B-level filter are interpolation filters, and both C-level filter and D-level filter are decimation filters.
[0127] Table 2
[0128]
[0129] In Table 2, "feature" indicates the result. "1 step no decim" means that one interpolation filter and no decimation filters are required, "1 step decim" means that one interpolation filter and one decimation filter are required, "nostep 1 decim" means that no interpolation filters and one decimation filter are required, "2 step no decim" means that two interpolation filters and no decimation filters are required, and "multi step decim" means that multiple interpolation filters and one (or more) decimation filters are required. For example, the number 147 can be split into "3 and 49," in which case two decimation filters can be used.
[0130] Table 3 is further provided, which shows a reference factor table for class A filters, class B filters, class C filters, and class D filters.
[0131] Table 3
[0132]
[0133] Here, "dB" stands for decibel.
[0134] The interpolation mentioned above, also known as upsampling, is achieved using a polyphase filter. Decimation can be easily accomplished by selecting only the useful data from the total data. To reduce the number of taps in the finite impulse response (FIR) filter, a multi-step filter is used.
[0135] The sampling rate conversion formula is: Output Fs = (Input Fs × L) / M, where L is the interpolation factor and M is the decimation factor. The interpolation factor, L, is used to upsample the signal using a polyphase filter, expanding the signal bandwidth and inserting zeros, followed by low-pass filtering to eliminate images. The decimation factor, M, is used to reduce the sample size by an interval after limiting the bandwidth using an anti-aliasing filter to avoid information loss.
[0136] For the embodiments shown in Table 1 above provided by this application, the input sampling rate classification and processing strategies include at least the following three categories, where:
[0137] The first type involves integer multiple interpolation (no decimation required). For example, for input Fs of 8kHz, 16kHz, 12kHz, or 24kHz, the processing strategy is single-step interpolation to directly match 48kHz. Specifically, the conversion from 8kHz to 48kHz involves using an interpolation factor of 6, labeled 6, and performing direct polyphase filtering without decimation.
[0138] The second type is simple fractional conversion (single interpolation + decimation). For example, if the input Fs is 32kHz, 64kHz, or 96kHz, the processing strategy is to achieve ratio conversion through single interpolation and decimation. The conversion from 32kHz to 48kHz is as follows:
[0139] Interpolation 3 (32kHz to 96kHz, labeled 3A): Polyphase filtering to suppress images;
[0140] Decimation 2 (96kHz to 48kHz, labeled 2C): downsampling after anti-aliasing filtering.
[0141] The third category is complex fractional multiplication conversion (multi-level decomposition). For example, if the input Fs is 11.025kHz, 22.05kHz, 44.1kHz, or 88.2kHz, the processing strategy is to decompose the large ratio L / M into multiple levels of small ratios to optimize the filter complexity. The conversion of 11.025kHz to 48kHz is as follows:
[0142] Total ratio: L / M=640 / 147 (i.e. 48 / 11.025≈4.3548 / 11.025≈4.35);
[0143] The decomposition step consists of an interpolation phase and a decimation phase, where:
[0144] Interpolation stage: 640=5×128640=5×128;
[0145] Interpolation 5 (11.025 to 55.125kHz, labeled 5A);
[0146] Interpolation 128 (55.125 to 7056kHz, multi-level combinations marked 4A, 4B, etc.).
[0147] Extraction stage: 147=3×49147=3×49;
[0148] Decimation 3 (7056 to 2352kHz, labeled 3C);
[0149] Decimation 49 (2352 to 48kHz, labeled 49C).
[0150] The corresponding advantage is that each filter stage only needs to process a small ratio, significantly reducing the number of FIR taps (e.g., 128-tap instead of 640-tap). Letter suffixes (e.g., A, B, C) represent different filter configurations or stages. For example: 3A: interpolation factor 3, using a first-stage polyphase filter (A-stage filter). 2C: decimation factor 2, using a third-stage antialiasing filter (C-stage filter). The combination of numbers and letters (e.g., 4A, 5A) indicates the decomposition steps of the multi-stage interpolation / decimation.
[0151] Through a multi-stage interpolation and decimation strategy, the system efficiently and flexibly converts any input sampling rate to a fixed 48kHz output, while balancing computational resources with signal quality. Key design principles include: ratio decomposition, which converts large ratios into combinations of smaller integer ratios to reduce single-stage complexity; polyphase filtering optimization, which reduces interpolation computation and improves real-time performance; and anti-aliasing, which ensures that the signal is bandwidth-limited before decimation to avoid distortion.
[0152] The analysis of using a polyphase filter for sample rate conversion in multi-stage interpolation includes the core process of multi-stage interpolation of a polyphase filter. Multi-stage interpolation is achieved by cascading multiple polyphase filters, each with a different interpolation factor. The overall interpolation factor is the product of the interpolation factors at each stage. The following is a breakdown of the three-stage process provided by the user:
[0153] First stage: 3x interpolation (phase = 3, taps = 8)
[0154] Input: original signal x(n), sampling rate Fs.
[0155] Output: interpolated signal y(n), sampling rate increased to 3Fs.
[0156] Implementation formula:
[0157] ;
[0158] ;
[0159] .
[0160] Key points include:
[0161] Each input sample generates 3 output samples (corresponding to an interpolation factor of 3).
[0162] An 8-tap filter is used, with different phases (0, 1, 2) corresponding to different coefficient groups hk_0, hk_1, hk_2, where the k value ranges from 0 to 7 and is used to suppress the image frequency.
[0163] Historical data storage: The first seven input samples (x(n-1) to x(n-7)) must be retained for calculation.
[0164] Second stage: 2x interpolation (phase = 2, taps = 6)
[0165] Input: First-stage output y(n), sampling rate 3Fs.
[0166] Output: quadratic interpolation signal z(n), sampling rate increased to 6Fs.
[0167] Implementation formula:
[0168] ;
[0169] .
[0170] Key points include:
[0171] An interpolation factor of 2 doubles the sampling rate.
[0172] A 6-tap filter is used, where phases 0 and 1 correspond to different coefficients pk_0, pk_1, where k ranges from 0 to 5.
[0173] Advantages of cascading: The sampling rate of the first stage has been increased, and the transition band of the second stage filter is wider, allowing the use of a lower order (6 taps is lower than the first stage 8 taps).
[0174] Level 3: 4x interpolation (Phase = 4, Taps = 4)
[0175] Input: Second stage output z(n), sampling rate 6Fs.
[0176] Output: Final signal d(n), with sampling rate increased to 24Fs (total interpolation factor 3×2×4=24).
[0177] Implementation formula:
[0178] ;
[0179] ;
[0180] ;
[0181] .
[0182] Key points include:
[0183] With an interpolation factor of 4, the sampling rate is increased to 24 times the original frequency.
[0184] A 4-tap filter is used, and phases 0-3 correspond to coefficient groups ck_0 to ck_3, where k ranges from 0 to 3.
[0185] Decimation integration: If subsequent decimation is required, anti-aliasing filtering and decimation can be combined in the final stage.
[0186] For example, an embodiment of converting an input of 11.025 kHz to an output of 48 kHz is provided. Specifically, if 11.025 kHz needs to be converted to 48 kHz, the total conversion ratio is 48 / 11.025≈4.3548 / 11.025≈4.35, which needs to be decomposed into multiple levels of interpolation and decimation, where:
[0187] Interpolation stage:
[0188] Interpolation 5 (11.025 to 55.125kHz) to Interpolation 128 (55.125 to 7056kHz).
[0189] Multi-level implementation: According to the provided three-level process, the interpolation factors can be decomposed into smaller integers (e.g., 5 = 5, 128 = 2 × 2 × 2 × 2 × 8).
[0190] Extraction phase:
[0191] Decimation 3 (7056 to 2352kHz) to Decimation 49 (2352 to 48kHz).
[0192] Merged decimation: decimation is performed directly after the last level of interpolation filtering to reduce the number of calculation steps.
[0193] Exemplarily, the advantages of multi-stage processing include reduced computational complexity, resource optimization, and flexible adaptation, where:
[0194] Reduced computational complexity. For example, single-stage 24x interpolation requires designing a 24-phase, very high-tap filter (e.g., 192 taps), which is computationally expensive. Three-stage decomposition (3×2×4): The total number of taps is only 8+6+4=18, significantly reducing the computational complexity (approximately 1 / 10 of a single-stage filter).
[0195] Resource optimization can involve multi-phase structures and historical data reuse, where: Multi-phase structure: Each stage only calculates non-zero input samples to avoid zero-value multiplication (such as zero values inserted during interpolation). Historical data reuse: The output of the intermediate stage is directly used as the input of the next stage, reducing storage requirements;
[0196] Flexible adaptation: For example, by adjusting the interpolation factors at each level (such as 3, 2, 4 or 5, 128, 3), different input sampling rates can be adapted to the target output (such as 48kHz). Decimation and merging: Anti-aliasing filtering and decimation are added at the last stage (such as decimating from 24Fs to 48kHz), avoiding additional processing steps.
[0197] In summary, the polyphase filter decomposes the large-ratio sampling rate conversion into multiple small-ratio steps through staged interpolation and decimation, significantly optimizing computational efficiency and resource utilization. Key design points include:
[0198] Interpolation factorization: select small integer combinations whose product equals the total ratio (e.g., 3 × 2 × 4 = 24);
[0199] Polyphase filter design: Each stage uses an independent coefficient set to adapt to the current interpolation factor and signal bandwidth;
[0200] Historical data management: store necessary historical samples to ensure filter continuity;
[0201] Decimation integration: Combines anti-aliasing filtering and decimation in the final stage to simplify the process.
[0202] Polyphase filters offer several key features: the output of the previous step serves as the input for the next step; one input, multiple outputs; and the data required for storage includes FIR coefficients and historical data, with the latter representing the results of intermediate filtering steps. In polyphase filter design, each stage uses an independent coefficient set to adapt to the current interpolation factor and signal bandwidth, reducing interpolation computation and improving real-time performance.
[0203] In addition, with respect to the filter constraints involved in this application, the filter constraints include the audio characteristics of the audio signal (signal spectrum characteristics), the signal index characteristics (performance indicators) corresponding to the target sampling rate, and the audio processing resource characteristics (resource constraints) of the local end, and Tables 4, 5, and 6 are further provided; wherein, Table 4 shows information related to the signal spectrum characteristics, Table 5 shows information related to the performance indicators, and Table 6 shows information related to the resource constraints.
[0204] Table 4
[0205]
[0206] Table 5
[0207]
[0208] Table 6
[0209]
[0210] Based on the above information, an embodiment of converting Bluetooth audio from 44.1kHz to 48kHz with limited resources is provided, wherein:
[0211] Signal characteristics: Bandwidth 20Hz-20kHz, full frequency band must be retained;
[0212] Resource constraints: RAM < 2KB (kilobytes), MAC units are limited; RAM stands for random access memory, also known as "random access memory"; MAC stands for Media Access Control.
[0213] The decision-making process includes:
[0214] First-level interpolation: Input 44.1kHz corresponds to Class A (wide transition band, Δf ≥ 8.8kHz); the physical meaning of Δf is the same-level selection function, which is the width of the transition band;
[0215] Secondary interpolation: Upsampling to 88.2kHz corresponds to Class B (Δf ≥ 8.8kHz, taps ≤ 6);
[0216] The extraction phase includes:
[0217] It is necessary to consider the main anti-aliasing and strict filtering, but the resources are insufficient, so the C level is abandoned;
[0218] Use Class D instead (Δf ≥ 7.2kHz, tap = 4).
[0219] In summary, the final solution is: 5A (44.1k converted to 220.5k) transmitted to 2B (220.5k converted to 441k) transmitted to 2D (441k converted to 220.5k) transmitted to 2D (220.5k converted to 110.25k) transmitted to 2D (110.25k converted to 55.125k) transmitted to 2D (55.125k converted to 27.56k), which does not meet the requirements. The revised solution is: 5A converted to 2B×2 converted to 3C (main anti-aliasing) transmitted to 2D (final extraction).
[0220] In summary, the selection of filter level requires a three-dimensional trade-off, among which:
[0221] Spectral dimension: Class A: low sampling rate and wide transition band; Class C: high sampling rate and strict anti-aliasing;
[0222] Resource dimension: Level D: extremely resource-constrained scenarios; Level B: balancing performance and complexity;
[0223] Performance dimension: SNR>90dB: mandatory Class C;
[0224] Low latency: Prioritize A / D level;
[0225] Golden rule: Interpolation chain: A to B to B... (sampling rate increases, transition band decreases); decimation chain: C to D to D... (sampling rate decreases, filtering simplifies).
[0226] It should be understood that, although the various steps in the flow charts involved in the above-mentioned embodiments are displayed in order of magnitude of the step numbers, these steps are not necessarily executed in order of magnitude of the step numbers. Unless clearly stated herein, the execution of these steps is not strictly restricted in order, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flow charts involved in the above-mentioned embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily carried out in order, but can be performed in turn or alternately with at least a portion of the steps or stages in other steps or other steps. It should be understood that, in different embodiments, the various steps can be freely combined as needed, and the various non-contradictory schemes formed by the combination all fall within the scope of protection of this application.
[0227] Based on the same inventive concept, embodiments of the present application further provide an audio signal processing device for implementing the aforementioned audio signal processing method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more audio signal processing device embodiments provided below can be found in the limitations of the audio signal processing method described above and are not further elaborated here.
[0228] In an exemplary embodiment, Figure 4As shown, an audio signal processing device 400 is provided, comprising: an audio acquisition module 41, a conversion factor sequence determination module 42, a filter type determination module 43, a filter identification sequence determination module 44, a filter configuration module 45 and an audio processing module 46, wherein: the audio acquisition module 41 is used to acquire an audio signal and determine the current sampling rate and the target sampling rate of the audio signal; the conversion factor sequence determination module 42 is used to determine a conversion factor sequence for the audio signal based on the current sampling rate and the target sampling rate, the conversion factor sequence including at least one conversion factor, the conversion factor being used to characterize the change in the amount of data after the filter performs conversion processing on the input signal data; the filter type determination module 43 is used to determine the filter type for the audio signal according to the filter constraint condition, the filter constraint condition being based on the audio signal. The invention relates to a method for determining a filter identification sequence for an audio signal according to a filter type; a filter identification sequence determining module 44 is used to construct a filter identification sequence for an audio signal according to a filter type; the filter identification sequence includes filter identifications arranged in sequence, the number of which is the same as the number of conversion factors, and the filter identifications are used to represent filters belonging to the filter type at the local end; a filter configuration module 45 is used to perform output conversion configuration on the filters represented by the filter identifications in the filter identification sequence according to the conversion factors in the conversion factor sequence, and obtain a filter configuration result; an audio processing module 46 is used to sequentially combine the filters at the local end based on the filter configuration result to obtain a filter group, and perform sampling rate conversion processing on the audio signal through the filter group to obtain a target signal that meets the target sampling rate.
[0229] In an exemplary embodiment, the conversion factor sequence determination module 42 is used to determine a conversion factor sequence for the audio signal based on the current sampling rate and the target sampling rate, specifically to: determine the interpolation factor and the decimation factor for the audio signal based on the current sampling rate and the target sampling rate, and obtain the conversion factor based on the interpolation factor and the decimation factor; and combine the conversion factors to obtain a conversion factor sequence for the audio signal.
[0230] In an exemplary embodiment, the filter identification sequence determination module 44 is used to construct a filter identification sequence for the audio signal according to the filter type, specifically for: determining the number of factors of the conversion factors included in the conversion factor sequence; determining the same number of filter identifications as the number of factors based on the filter type; and combining the filter identifications to obtain a filter identification sequence for the audio signal.
[0231] In an exemplary embodiment, the conversion factor includes an interpolation factor and a decimation factor, and the number of factors includes a first factor number of the interpolation factor and a second factor number of the decimation factor; the filter type includes an interpolation type and a decimation type; the above-mentioned filter identifier sequence determination module 44 is used to determine, based on the filter type, a number of filter identifiers that is the same as the number of factors, specifically for: determining, based on the interpolation type, a first filter identifier that is the same as the number of first factors; determining, based on the decimation type, a second filter identifier that is the same as the number of second factors; and obtaining, based on the first filter identifier and the second filter identifier, a number of filter identifiers that is the same as the number of factors.
[0232] In an exemplary embodiment, the conversion factor includes an interpolation factor and a decimation factor, and the filter type includes an interpolation type and a decimation type; the above-mentioned filter configuration module 45 is used to perform output conversion configuration on the filters represented by the filter identifiers in the filter identification sequence according to the conversion factors in the conversion factor sequence to obtain a filter configuration result, specifically for: when the interpolation factor indicates that the filter of the interpolation type meets the trigger configuration condition, the output conversion configuration is performed on the first filter according to the interpolation factor to obtain a first configuration result, and the first filter is the filter of the interpolation type represented by the filter identifier in the filter identification sequence; when the decimation factor indicates that the filter of the decimation type meets the trigger configuration condition, the output conversion configuration is performed on the second filter according to the decimation factor to obtain a second configuration result, and the second filter is the filter of the decimation type represented by the filter identifier in the filter identification sequence; the filter configuration result is obtained based on the first configuration result and the second configuration result.
[0233] In an exemplary embodiment, the filter types include interpolation types and decimation types; the audio processing module 46 is used to sequentially combine the local filters based on the filter configuration result to obtain a filter group, specifically for: determining the local target filter and the filter cascade order of the target filter according to the filter configuration result, the filter cascade order including the first cascade order of the first filter of the interpolation type in the target filter and the second cascade order of the second filter of the decimation type; cascading the first filter in the first cascade order to obtain an interpolation filter group; cascading the second filter in the second cascade order to obtain an decimation filter group; and connecting the interpolation filter group and the decimation filter group in a certain order to obtain a filter group.
[0234] In an exemplary embodiment, the audio processing module 46 is used to perform sampling rate conversion processing on the audio signal through a filter group, specifically for: obtaining the number of taps of each target filter in the filter group; determining the audio processing algorithm corresponding to each target filter based on the number of taps of each target filter; using each target filter, based on the cascade order of each target filter and the audio processing algorithm corresponding to each target filter, to perform sampling rate conversion processing on the audio signal.
[0235] Each module in the above-mentioned audio signal processing device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0236] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal or a server, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data such as filter types. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection.
[0237] Furthermore, the present application may use a terminal as the execution subject of the audio signal processing method, may use a server as the execution subject of the audio signal processing method, or may implement the execution of the audio signal processing method based on a combination of a terminal and a server.
[0238] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0239] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements steps related to audio signal processing when executing the computer program.
[0240] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, steps related to processing of an audio signal are implemented.
[0241] In one embodiment, a computer program product is provided, comprising a computer program, which implements steps related to processing of an audio signal when executed by a processor.
[0242] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0243] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.
[0244] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0245] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for processing an audio signal, characterized in that: The method comprises: Acquire an audio signal, and determine a current sampling rate and a target sampling rate of the audio signal; determining, based on the current sampling rate and the target sampling rate, a conversion factor sequence for the audio signal, wherein the conversion factor sequence includes at least one conversion factor, and the conversion factor is used to represent a change in data volume after a filter performs conversion processing on input signal data; determining a filter type for the audio signal according to a filter constraint, wherein the filter constraint is obtained based on at least one of an audio feature of the audio signal, a signal indicator feature corresponding to the target sampling rate, and a local audio processing resource feature; Constructing a filter identification sequence for the audio signal according to the filter type; the filter identification sequence includes filter identifications arranged in sequence, the number of which is the same as the number of the conversion factors, the filter identifications being used to identify filters on the local end that belong to the filter type; According to the conversion factors in the conversion factor sequence, output conversion configuration is performed on the filters represented by the filter identifiers in the filter identifier sequence to obtain filter configuration results; Based on the filter configuration result, the local filters are sequentially combined to obtain a filter group, and the audio signal is subjected to sampling rate conversion processing by the filter group to obtain a target signal that meets the target sampling rate.
2. The method according to claim 1, characterized in that The determining, based on the current sampling rate and the target sampling rate, a sequence of conversion factors for the audio signal, comprises: Determining an interpolation factor and a decimation factor for the audio signal based on the current sampling rate and the target sampling rate, and obtaining a conversion factor according to the interpolation factor and the decimation factor; The conversion factors are combined to obtain a conversion factor sequence for the audio signal.
3. The method according to claim 1, characterized in that The constructing a filter identification sequence for the audio signal according to the filter type includes: determining the number of factors of the conversion factor included in the conversion factor sequence; Determining, based on the filter types, a number of filter identifiers equal to the number of factors; The filter identifiers are combined to obtain a filter identifier sequence for the audio signal.
4. The method according to claim 3, characterized in that The conversion factor includes an interpolation factor and a decimation factor, the factor quantity includes a first factor quantity of the interpolation factor and a second factor quantity of the decimation factor; the filter type includes an interpolation type and a decimation type; The determining, based on the filter type, a number of filter identifiers that is equal to the number of factors, includes: Determining, based on the interpolation type, first filter identifiers whose number is equal to the number of the first factors; determining, based on the decimation type, second filter identifiers, the same number as the number of the second factors; According to the first filter identifier and the second filter identifier, filter identifiers whose number is the same as the number of the factors are obtained.
5. The method according to claim 1, wherein The conversion factor includes an interpolation factor and a decimation factor, and the filter type includes an interpolation type and a decimation type; The step of performing output conversion configuration on the filters represented by the filter identifiers in the filter identifier sequence according to the conversion factors in the conversion factor sequence to obtain filter configuration results includes: When the interpolation factor indicates that the filter of the interpolation type meets the trigger configuration condition, performing output conversion configuration on a first filter according to the interpolation factor to obtain a first configuration result, where the first filter is a filter of the interpolation type represented by the filter identifier in the filter identifier sequence; When the decimation factor indicates that the filter of the decimation type meets the trigger configuration condition, performing output conversion configuration on a second filter according to the decimation factor to obtain a second configuration result, where the second filter is a filter of the decimation type represented by the filter identifier in the filter identifier sequence; A filter configuration result is obtained according to the first configuration result and the second configuration result.
6. The method according to claim 1, characterized in that The filter types include interpolation type and decimation type; The step of sequentially combining the filters at the local end based on the filter configuration result to obtain a filter bank includes: Determining a target filter at a local end and a filter cascade order of the target filter according to the filter configuration result, the filter cascade order including a first cascade order of a first filter of the interpolation type and a second cascade order of a second filter of the decimation type in the target filter; cascade-connecting the first filters in the first cascade order to obtain an interpolation filter bank; cascade-connecting the second filters according to the second cascade order to obtain a decimation filter bank; The interpolation filter group and the decimation filter group are connected in sequence to obtain a filter group.
7. The method according to claim 1, characterized in that The performing sampling rate conversion processing on the audio signal by the filter bank includes: Obtaining the number of taps of each target filter in the filter bank; Determining, according to the number of taps of each target filter, an audio processing algorithm corresponding to each target filter; The target filters are used to perform sampling rate conversion processing on the audio signal based on the cascade order of the target filters and the audio processing algorithms corresponding to the target filters.
8. An audio signal processing device, characterized in that: The device comprises: An audio acquisition module, configured to acquire an audio signal and determine a current sampling rate and a target sampling rate of the audio signal; a conversion factor sequence determining module, configured to determine a conversion factor sequence for the audio signal based on the current sampling rate and the target sampling rate, wherein the conversion factor sequence includes at least one conversion factor, and the conversion factor is used to represent a change in data volume after the filter performs conversion processing on the input signal data; a filter type determination module, configured to determine a filter type for the audio signal according to a filter constraint condition, wherein the filter constraint condition is obtained based on at least one of an audio feature of the audio signal, a signal indicator feature corresponding to the target sampling rate, and a local audio processing resource feature; a filter identification sequence determination module, configured to construct a filter identification sequence for the audio signal according to the filter type; the filter identification sequence comprising filter identifications arranged in sequence, the number of which being equal to the number of the conversion factors, the filter identifications being used to identify filters belonging to the filter type at the local end; a filter configuration module, configured to perform output conversion configuration on the filters represented by the filter identifiers in the filter identifier sequence according to the conversion factors in the conversion factor sequence, to obtain a filter configuration result; The audio processing module is used to sequentially combine the local filters to obtain a filter group based on the filter configuration result, and perform sampling rate conversion processing on the audio signal through the filter group to obtain a target signal that meets the target sampling rate.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Signal sampling rate conversion method and device, electronic equipment and storage medium
CN118449522A
Sampling rate converter, sampling rate conversion method and communication system
CN118476160A
Temporal interpolation of adjacent spectra
EP2562751A1
Method and system for utilizing rate conversion filters to reduce mixing complexity during multipath multi-rate audio processing
KR1020080049688A
Sample rate conversion with pitch-based interpolation filters
US10224062B1