Psychoacoustic analysis method, apparatus, device, and storage medium
By selecting some masking sources to participate in the masking threshold calculation in psychoacoustic analysis, the problems of large computational load and high complexity in the existing technology are solved, and the computational load and complexity are reduced, while maintaining the accuracy and flexibility of the analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-17
- Publication Date
- 2026-03-24
AI Technical Summary
Current psychoacoustic analysis involves a large amount of computation and high computational complexity, making it difficult to effectively reduce the computational load.
By identifying multiple masking sources for the audio signal and analyzing the masking threshold based on some of these sources, the number of masking sources involved in the calculation is reduced, thus lowering the computational complexity.
It effectively reduces the computational load and complexity of psychoacoustic analysis while maintaining the accuracy and flexibility of the analysis.
Smart Images

Figure CN116391226B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of communication, and in particular, to a psychoacoustic analysis method and device, equipment and storage medium. BACKGROUND
[0002] The psychoacoustic model can be applied to the field of audio coding. Using the psychoacoustic model can help to remove the redundancy in the audio signal that is less than the auditory threshold, thereby reducing the signal irrelevant to the auditory perception, improving the subjective quality of audio coding, reducing the quantization code rate and quantization noise. The psychoacoustic model can also be applied to the audio digital watermarking technology, to hide the watermark in the audio signal that cannot be perceived by the human ear, and to recover the hidden information in the decoding section. The psychoacoustic model can also be applied to the field of sound quality evaluation, to objectively evaluate the quality of sound through the psychoacoustic model, and to provide a method for improving sound quality.
[0003] In the related art, in the psychoacoustic analysis process, the input audio signal is processed, all tonal maskers and non-tonal maskers contained in the audio signal are extracted, and then the masking threshold of the audio signal is analyzed based on all the extracted tonal maskers and non-tonal maskers. SUMMARY
[0004] The embodiments of the present disclosure provide a psychoacoustic analysis method, device, equipment, chip system, storage medium, computer program and computer program product, which can be applied to the field of communication technology, and can effectively reduce the calculation amount of psychoacoustic analysis, thereby reducing the calculation complexity.
[0005] In a first aspect, the embodiments of the present disclosure provide a psychoacoustic analysis method, which comprises: determining a plurality of maskers of an audio signal; and analyzing a masking threshold of the audio signal according to part of the plurality of maskers.
[0006] In a second aspect, the embodiments of the present disclosure provide a communication device having part or all of the functions of the method of the first aspect, such as the function of the communication device, which can have part or all of the functions of the embodiments of the present disclosure, or can have the function of implementing any one of the embodiments of the present disclosure independently. The functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more units or modules corresponding to the above functions.
[0007] Optionally, in one embodiment of this disclosure, the communication device may include a transceiver module and a processing module, wherein the processing module is configured to support the communication device in performing the corresponding functions in the above-described method. The transceiver module is used to support communication between the communication device and other devices. The communication device may also include a storage module, which is coupled to the transceiver module and the processing module, and stores the necessary computer programs and data of the communication device.
[0008] As an example, the processing module can be a processor, the transceiver module can be a transceiver or a communication interface, and the storage module can be a memory.
[0009] Thirdly, embodiments of this disclosure provide a communication device including a processor that, when the processor invokes a computer program in memory, executes the psychoacoustic analysis method described in the first aspect.
[0010] Fourthly, embodiments of this disclosure provide a communication device including a processor and a memory, the memory storing a computer program; the processor executes the computer program stored in the memory to cause the communication device to perform the psychoacoustic analysis method described in the first aspect above.
[0011] Fifthly, embodiments of this disclosure provide a communication device including a processor and an interface circuit. The interface circuit is used to receive code instructions and transmit them to the processor, which is used to execute the code instructions to cause the device to perform the psychoacoustic analysis method described in the first aspect.
[0012] Sixthly, embodiments of this disclosure provide a communication system, which includes the communication device described in the second aspect, or the communication device described in the third aspect, or the communication device described in the fourth aspect, or the communication device described in the fifth aspect.
[0013] In a seventh aspect, embodiments of this disclosure provide a computer-readable storage medium for storing instructions for use by a terminal device, which, when executed, cause the terminal device to perform the psychoacoustic analysis method described in the first aspect.
[0014] Eighthly, embodiments of this disclosure provide a readable storage medium for storing instructions for use by a network device, which, when executed, cause the network device to perform the psychoacoustic analysis method described in the first aspect.
[0015] Ninthly, this disclosure also provides a computer program product including a computer program that, when run on a computer, causes the computer to perform the psychoacoustic analysis method described in the first aspect above.
[0016] In a tenth aspect, this disclosure provides a chip system including at least one processor and an interface for supporting terminal devices and / or network devices in implementing the functions involved in the first aspect, such as determining or processing at least one of the data and information involved in the above methods.
[0017] In one possible design, the chip system further includes a memory for storing computer programs and data necessary for the terminal device. The chip system may consist of chips or may include chips and other discrete components.
[0018] In the eleventh aspect, this disclosure provides a computer program that, when run on a computer, causes the computer to perform the psychoacoustic analysis method described in the first aspect above.
[0019] In summary, the psychoacoustic analysis method, apparatus, device, chip system, storage medium, computer program, and computer program product provided in the embodiments of this disclosure can achieve the following technical effects:
[0020] By identifying multiple masking sources for the audio signal, and then analyzing the masking threshold of the audio signal based on some of the masking sources, the computational load of psychoacoustic analysis can be effectively reduced, thereby lowering the computational complexity, since some masking sources are selected from all the masking sources of the audio signal to participate in the analysis and calculation of the masking threshold. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments or background art of this disclosure, the accompanying drawings used in the embodiments or background art of this disclosure will be described below.
[0022] Figure 1 This is a schematic diagram of the architecture of a communication system provided in an embodiment of the present disclosure;
[0023] Figure 2 This is a flowchart illustrating a psychoacoustic analysis method provided in an embodiment of this disclosure;
[0024] Figure 3 This is a flowchart illustrating another psychoacoustic analysis method provided in this embodiment of the present disclosure;
[0025] Figure 4 This is a flowchart illustrating another psychoacoustic analysis method provided in this embodiment of the present disclosure;
[0026] Figure 5 This is a flowchart illustrating another psychoacoustic analysis method provided in this embodiment of the present disclosure;
[0027] Figure 6 This is a flowchart illustrating another psychoacoustic analysis method provided in this embodiment of the present disclosure;
[0028] Figure 7a This is a flowchart illustrating another psychoacoustic analysis method provided in this embodiment of the present disclosure;
[0029] Figure 7b This is a flowchart illustrating another psychoacoustic analysis method provided in this embodiment of the present disclosure;
[0030] Figure 7c This is a flowchart illustrating another psychoacoustic analysis method provided in this embodiment of the present disclosure;
[0031] Figure 8 This is a schematic diagram of the masking expansion function in an embodiment of this disclosure;
[0032] Figure 9 This is a schematic diagram of the architecture of the psychoacoustic analysis method in this embodiment of the disclosure;
[0033] Figure 10a This is a schematic diagram of experimental statistics on the time taken to enable cross-masking in this embodiment of the present disclosure;
[0034] Figure 10b This is a schematic diagram of experimental statistics on the time taken to disable cross-masking in an embodiment of this disclosure;
[0035] Figure 11 This is a schematic diagram of the structure of a communication device provided in an embodiment of the present disclosure;
[0036] Figure 12 This is a schematic diagram of another communication device provided in an embodiment of this disclosure;
[0037] Figure 13 This is a schematic diagram of the structure of a chip provided in an embodiment of this disclosure. Detailed Implementation
[0038] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this disclosure as detailed in the appended claims.
[0039] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. The singular forms “a” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0040] It should be understood that although the terms first, second, third, etc., may be used to describe various information in embodiments of this disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of embodiments of this disclosure, and similarly, second information may also be referred to as first information. Depending on the context, the words “if” and “suppose” as used herein may be interpreted as “when”, “when”, or “in response to a determination”.
[0041] To facilitate understanding, the terminology used in this disclosure will be introduced first.
[0042] 1. Psychoacoustic model.
[0043] The psychoacoustic model is a mathematical representation of the statistical properties of human hearing, which explains the physiological principles behind various human auditory sensations.
[0044] 2. To conceal.
[0045] In audiology, masking refers to the increase in the human ear's threshold for perceiving one sound due to the presence of another sound.
[0046] 3. A masking source is a sound source that produces a masking effect. A tone masking source is a masking source that produces a tone component that produces a masking effect, while a non-tone masking source is a masking source that produces a non-tone component (such as noise) that produces a masking effect.
[0047] To better understand the psychoacoustic analysis method disclosed in this embodiment, the communication system to which this embodiment applies is first described below.
[0048] Please see Figure 1 , Figure 1 This is a schematic diagram of the architecture of a communication system provided in an embodiment of the present disclosure. The communication system may include, but is not limited to, a network device and a terminal device. Figure 1 The number and form of devices shown are for illustrative purposes only and do not constitute a limitation on the embodiments of this disclosure. In actual applications, two or more network devices and two or more terminal devices may be included. Figure 1 The communication system shown is exemplified by a network device 101 and a terminal device 102.
[0049] It should be noted that the technical solutions of this disclosure can be applied to various communication systems. For example, Long Term Evolution (LTE) systems, 5th Generation (5G) mobile communication systems, 5G New Radio (NR) systems, or other future new mobile communication systems.
[0050] The network device 101 in this disclosure is a network-side entity used for transmitting or receiving signals. For example, the network device 101 can be an evolved NodeB (eNB), a transmission reception point (TRP), a next-generation NodeB (gNB) in an NR system, a private network system, a base station in other future mobile communication systems, or an access node in a wireless fidelity (WiFi) system. This disclosure does not limit the specific technology or device form used in the network device.
[0051] The network device provided in this embodiment can be composed of a central unit (CU) and a distributed unit (DU). The CU can also be called a control unit. By adopting the CU-DU structure, the protocol layer of the network device, such as a base station, can be separated. Some of the protocol layer functions are centrally controlled by the CU, while the remaining part or all of the protocol layer functions are distributed in the DU, which is centrally controlled by the CU.
[0052] The terminal device 102 in this embodiment is a user-side entity used to receive or transmit signals, such as a mobile phone. The terminal device can also be referred to as a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. Terminal devices can be communication-enabled vehicles, smart cars, mobile phones, wearable devices, tablets, computers with wireless transceiver capabilities, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, wireless terminal devices in industrial control, wireless terminal devices in self-driving, wireless terminal devices in remote medical surgery, wireless terminal devices in smart grids, wireless terminal devices in transportation safety, wireless terminal devices in smart cities, wireless terminal devices in smart homes, and so on.
[0053] The embodiments disclosed herein do not limit the specific technology or device form used in the terminal device.
[0054] It is understood that the communication system described in the embodiments of this disclosure is for the purpose of more clearly illustrating the technical solutions of the embodiments of this disclosure, and does not constitute a limitation on the technical solutions provided in the embodiments of this disclosure. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this disclosure are also applicable to similar technical problems.
[0055] The psychoacoustic analysis method in this disclosure can be applied to the aforementioned network devices, or to terminal devices, or to both network devices and terminal devices, and can also be applied to any other system that may be capable of performing psychoacoustic analysis on audio signals, such as streaming media transmission systems and OTT (Over The Top) media transmission systems that provide various application services to users via the Internet. There are no limitations on this.
[0056] In related technologies, psychoacoustic analysis involves processing the input audio signal to extract all tone-masking sources and non-tone-masking sources. Then, a masking threshold is calculated based on these extracted sources. In this approach, the computational complexity depends on the number of masking sources involved; a higher number of sources results in greater computational complexity. Therefore, this embodiment identifies multiple masking sources for the audio signal. Then, based on a subset of these sources, the masking threshold is analyzed. By selecting a subset of masking sources from all sources for threshold calculation, the computational load of psychoacoustic analysis is effectively reduced, thereby lowering computational complexity.
[0057] It should be noted that the psychoacoustic analysis method provided in any embodiment of this application can be executed alone, or can be executed together with possible implementation methods in other embodiments, or can be executed together with any technical solution in related technologies.
[0058] The psychoacoustic analysis method and apparatus provided in this disclosure will now be described in detail with reference to the accompanying drawings. Figure 2 This is a flowchart illustrating a psychoacoustic analysis method provided in an embodiment of this disclosure, which is executed by a communication system. The psychoacoustic analysis method in this embodiment can be applied to a communication system, and there are no limitations on its application.
[0059] like Figure 2 As shown, the method may include, but is not limited to, the following steps:
[0060] S201: Identify multiple masking sources for the audio signal.
[0061] In some embodiments, an audio signal may be input to a communication device, which may perform psychoacoustic analysis on the audio signal to identify multiple masking sources of the audio signal, without limitation.
[0062] In some embodiments, all masking sources contained in the audio signal may be identified first, and then all of the multiple masking sources may be selected for processing; there is no limitation on this.
[0063] In some embodiments, the masking source may include at least one of tone masking sources and non-tone masking sources. For example, the number of tone masking sources may be one or more, and the number of non-tone masking sources may also be one or more. Non-tone masking sources may also be referred to as noise masking sources, and there is no limitation thereto.
[0064] In some embodiments, the power spectrum and sound pressure level of the input audio signal can be calculated first. Then, based on the calculation results, the tonal and non-tonal components contained in the audio signal can be found. Then, based on the calculated tonal and non-tonal components, tonal masking sources and / or non-tonal masking sources can be extracted to determine multiple masking sources contained in the audio signal. There is no limitation on this.
[0065] S202: Analyze the masking threshold of the audio signal based on a subset of masking sources from multiple masking sources.
[0066] Among them, partial masking sources refer to masking sources selected from multiple masking sources. The selected masking sources participate in the calculation of the masking threshold. Among the multiple masking sources, there may be one or more tone masking sources and one or more non-tone masking sources. The selected partial masking sources can be tone masking sources or non-tone masking sources, without any restrictions.
[0067] In some embodiments, masking sources can be extracted one by one from the audio signal. During the extraction of each masking source, it is determined whether to select the masking source as a partial masking source. Alternatively, all masking sources can be extracted from the audio signal. After extraction is completed, partial masking sources can be selected from multiple masking sources. Of course, other possible methods can also be used to determine partial masking sources, and there are no restrictions on this.
[0068] After identifying multiple masking sources for the audio signal, some masking sources can be selected from these sources, and the masking threshold of the audio signal can be analyzed based on the selected partial masking sources. There are no restrictions on this.
[0069] In some embodiments of this disclosure, partial masking sources can also be determined from multiple masking sources to support the analysis of the masking threshold of audio signals based on the selected partial masking sources. For example, partial masking sources can be selected from multiple masking sources based on a set selection strategy; or, partial masking sources can be selected from multiple masking sources based on artificial intelligence; of course, partial masking sources can also be determined from multiple masking sources in any other possible way, without limitation.
[0070] In this embodiment, multiple masking sources of the audio signal are determined, and then the masking threshold of the audio signal is analyzed based on some of the masking sources. Since some masking sources are selected from all the masking sources of the audio signal to participate in the analysis and calculation of the masking threshold, the amount of computation in psychoacoustic analysis can be effectively reduced, thereby reducing the computational complexity.
[0071] Figure 3This is a flowchart illustrating another psychoacoustic analysis method provided in this embodiment, which is executed by a communication system. The psychoacoustic analysis method in this embodiment can be applied to a communication system, and there are no limitations on its application.
[0072] like Figure 3 As shown, the method may include, but is not limited to, the following steps:
[0073] S301: Determine multiple masking sources for the audio signal, wherein the masking sources include: tone masking sources and non-tone masking sources.
[0074] That is to say, in some embodiments of this disclosure, the masking source may include one or more tone masking sources and one or more non-tone masking sources. Some masking sources may be selected from one or more tone masking sources and one or more non-tone masking sources, and there is no limitation thereto.
[0075] S302: If the critical frequency band to which the non-tone masking source belongs contains a tone masking source, then obtain the frequency distance between the tone masking source and the non-tone masking source.
[0076] It is understood that audio signals may cover multiple critical frequency bands, and different masking sources may correspond to the same or different critical frequency bands. Therefore, in some embodiments, the selection of some masking sources can be achieved based on whether the tone masking source and the non-tone masking source correspond to the critical frequency band.
[0077] In some embodiments, it can be determined whether each critical frequency band contains a tone masking source. If the critical frequency band contains a tone masking source, the tone masking source and non-tone masking sources in the critical frequency band can be selected by referring to the tone masking source. There are no restrictions on this.
[0078] In some embodiments, there may be multiple critical frequency bands. In this case, the selection and processing of tone masking sources and non-tone masking sources in the corresponding critical frequency band can be realized based on whether each critical frequency band contains tone masking sources. There is no limitation on this.
[0079] In some embodiments, if the critical frequency band to which the non-tone masking source belongs contains a tone masking source, the frequency distance between the tone masking source and the non-tone masking source is obtained. Then, based on the frequency distance, the tone masking source and the non-tone masking source in the critical frequency band are selected. There are no restrictions on this.
[0080] Frequency distance refers to the distance between tone-masked and non-tone-masked sources in the frequency dimension, and the unit can be the bark scale, which is the unit of measurement used in the critical band principle.
[0081] In other words, we can first determine whether a non-tone masking source's critical frequency band contains a tone masking source. If the non-tone masking source's critical frequency band contains a tone masking source, we can obtain the frequency distance between the tone masking source and the non-tone masking source to select the masking source.
[0082] In some embodiments, when performing the step of obtaining the frequency distance between tone-masking sources and non-tone-masking sources, it may be possible to determine the frequency domain position of the tone-masking source within the critical frequency band and determine the frequency domain position of the non-tone-masking source within the critical frequency band, and then perform a difference processing on the frequency domain position of the tone-masking source within the critical frequency band and the frequency domain position of the non-tone-masking source within the critical frequency band to obtain the frequency distance between the tone-masking source and the non-tone-masking source, without limitation.
[0083] In other embodiments, the frequency distance between tone-masked sources and non-tone-masked sources can be in bark scale, which is a unit of measurement used in the critical band principle, and there is no limitation thereto.
[0084] S303: Based on frequency distance, determine tone masking sources and / or non-tone masking sources as partial masking sources.
[0085] After obtaining the frequency distance between the tone masking source and the non-tone masking source, the tone masking source and / or the non-tone masking source can be determined as a partial masking source based on the frequency distance.
[0086] In some embodiments, the selection of tone masking sources and / or non-tone masking sources can be determined based on frequency distance and set rules; alternatively, the selection of tone masking sources and / or non-tone masking sources can be determined by combining a masking source selection model; of course, any other possible methods can be used to determine whether to select tone masking sources and / or non-tone masking sources, and there are no restrictions on this.
[0087] In other embodiments, determining whether to select a tone masking source and / or a non-tone masking source may, for example, determine whether to select a tone masking source as a partial masking source; or, determine whether to select a non-tone masking source as a partial masking source; or, it may also determine whether to select both a tone masking source and a non-tone masking source as partial masking sources, without limitation.
[0088] In other embodiments, if the audio signal contains multiple critical frequency bands, the above processing can be performed on each critical frequency band containing a tone masking source. That is, for each critical frequency band containing a tone masking source, it is determined whether to select a tone masking source and / or a non-tone masking source as a partial masking source based on the frequency distance between the tone masking source and the non-tone masking source contained therein, without any limitation.
[0089] In other embodiments, the critical frequency band may not contain a tone masking source, or it may contain one tone masking source, or it may contain multiple tone masking sources. The critical frequency band may also contain one non-tone masking source. In this embodiment, if the critical frequency band contains multiple tone masking sources, the frequency distance between each tone masking source and the non-tone masking source in the critical frequency band can be calculated, and it can be determined whether to select each tone masking source and the non-tone masking source. There are no restrictions on this.
[0090] S304: Analyze the masking threshold of the audio signal based on a subset of masking sources from multiple masking sources.
[0091] After identifying tone masking sources and / or non-tone masking sources as partial masking sources, the masking threshold of the audio signal can be analyzed based on the partial masking sources among multiple masking sources.
[0092] In this embodiment, by selecting a portion of the masking sources from all masking sources of the audio signal to participate in the analysis and calculation of the masking threshold, the computational load of psychoacoustic analysis can be effectively reduced, thereby lowering the computational complexity. If the masking sources include tone masking sources and non-tone masking sources, then when tone masking sources are included within the critical frequency band to which the non-tone masking sources belong, the frequency distance between tone masking sources and non-tone masking sources is obtained. Based on the frequency distance, tone masking sources and / or non-tone masking sources are determined as partial masking sources. This enables the rapid and accurate selection of partial masking sources from multiple masking sources, ensuring the accuracy of psychoacoustic analysis based on the selected partial masking sources.
[0093] Figure 4 This is a flowchart illustrating another psychoacoustic analysis method provided in this embodiment, which is executed by a communication system. The psychoacoustic analysis method in this embodiment can be applied to a communication system, and there are no limitations on its application.
[0094] like Figure 4 As shown, the method may include, but is not limited to, the following steps:
[0095] S401: Determine multiple masking sources for the audio signal, wherein the masking sources include: non-tone masking sources.
[0096] In some embodiments, the multiple masking sources contained in the audio signal may all be non-tonal masking sources, and different non-tonal masking sources may belong to the same or different critical frequency bands.
[0097] S402: If the critical frequency band to which the non-tone masking source belongs does not contain tone masking sources, then the non-tone masking source is determined to be a partial masking source.
[0098] In other embodiments, if the audio signal contains multiple non-tone masking sources, it can be analyzed whether each non-tone masking source's critical frequency band contains a tone masking source. If no tone masking source is found, the non-tone masking source is directly determined to be a partial masking source. If a tone masking source is found, then based on the above... Figure 3 The method steps in the illustrated embodiment determine whether to select the non-tone masking source as a partial masking source, without limitation.
[0099] S403: Analyze the masking threshold of the audio signal based on a subset of masking sources from multiple masking sources.
[0100] In some embodiments, if there is no tone masking source within the critical frequency band to which the non-tone masking source belongs, then after determining that the non-tone masking source is a partial masking source, the masking threshold of the audio signal can be analyzed based on the partial masking source.
[0101] In this embodiment, by selecting a subset of masking sources from all masking sources in the audio signal to participate in the masking threshold analysis and calculation, the computational load of psychoacoustic analysis can be effectively reduced, thereby lowering computational complexity. When no tone masking source is present within the critical frequency band of a non-tone masking source, determining the non-tone masking source as a partial masking source effectively improves the flexibility of partial masking source selection, is effectively applicable to the personalized distribution of masking sources, and flexibly applies to the psychoacoustic analysis of personalized audio signals.
[0102] Figure 5 This is a flowchart illustrating another psychoacoustic analysis method provided in this embodiment, which is executed by a communication system. The psychoacoustic analysis method in this embodiment can be applied to a communication system, and there are no limitations on its application.
[0103] like Figure 5 As shown, the method may include, but is not limited to, the following steps:
[0104] S501: Determine multiple masking sources for the audio signal, wherein the masking sources include: tone masking sources and non-tone masking sources.
[0105] S502: If the critical frequency band to which the non-tone masking source belongs contains a tone masking source, then obtain the frequency distance between the tone masking source and the non-tone masking source.
[0106] S503: Obtain the distance threshold.
[0107] The distance threshold can refer to the frequency distance threshold used to determine whether to select a tone-masking source or a non-tone-masking source as a partial masking source. The distance threshold can be measured in bars.
[0108] In some embodiments of this disclosure, a set distance threshold can be obtained, thereby effectively improving the efficiency of obtaining the distance threshold.
[0109] For example, in this embodiment of the disclosure, the distance threshold can be set to a fixed value of 0.5 bark, without limitation.
[0110] In some other embodiments of this disclosure, a distance threshold corresponding to the critical frequency band can also be obtained, so that the distance threshold can be effectively adapted to the critical frequency band, thereby improving the accuracy of selecting tone-masking sources and non-tone-masking sources.
[0111] Among them, the distance threshold corresponding to the critical frequency band refers to the critical frequency band for selecting partial masking sources, such as the critical frequency band to which the non-tone masking source belongs in step S502, and there is no restriction on it.
[0112] For example, a corresponding distance threshold can be set for each critical frequency band without any restrictions.
[0113] In some other embodiments of this disclosure, the distance threshold can be determined based on the frequency of the critical band, thereby improving the flexibility of distance threshold acquisition and effectively adapting to the psychoacoustic analysis of personalized audio signals.
[0114] For example, in this embodiment of the disclosure, a distance threshold of 0.3 bark can be set for the first 15 critical frequency bands and a distance threshold of 0.6 bark can be set for the last 10 critical frequency bands. The selection of the distance threshold is related to the frequency, and the higher the frequency of the critical frequency band, the larger the selected threshold.
[0115] For example, the human ear can perceive sound frequencies ranging from 20 Hz to 20 kHz, with the highest sensitivity to sounds in the 1 kHz to 3 kHz range. Therefore, using a smaller distance threshold in the low-frequency critical band preserves more masking source components, while selecting a larger distance threshold in the high-frequency critical band removes more masking source components, which is more beneficial to subjective auditory perception.
[0116] S504: If the frequency distance is greater than or equal to the distance threshold, then the tone masking source and the non-tone masking source are identified as partially masked sources.
[0117] After obtaining the frequency distance between the tone masking source and the non-tone masking source and obtaining the distance threshold, the tone masking source and the non-tone masking source can be determined as partial masking sources if the frequency distance is greater than or equal to the distance threshold. That is, in the critical frequency band where the tone masking source exists, it is determined whether the frequency distance between the position (bark scale) of the tone masking source in the critical frequency band and the noise masking source (non-tone masking source) of the critical frequency band is less than the set distance threshold. If the frequency distance is greater than or equal to the distance threshold, the tone masking source and the non-tone masking source can be directly selected as partial masking sources, that is, the tone masking source and the non-tone masking source are selected to participate in the calculation of the masking threshold without any restrictions.
[0118] S505: Analyze the masking threshold of an audio signal based on a subset of masking sources from multiple masking sources.
[0119] In this embodiment, by selecting a portion of the masking sources from all masking sources of the audio signal to participate in the masking threshold analysis and calculation, the computational load of psychoacoustic analysis can be effectively reduced, thereby lowering the computational complexity. When a tone masking source is included within the critical frequency band of a non-tone masking source, the frequency distance between the tone masking source and the non-tone masking source is obtained, and a distance threshold is acquired. If the frequency distance is greater than or equal to the distance threshold, the tone masking source and the non-tone masking source are determined to be partial masking sources. This achieves accurate selection of partial masking sources from multiple masking sources, improving the accuracy of psychoacoustic analysis.
[0120] Figure 6 This is a flowchart illustrating another psychoacoustic analysis method provided in this embodiment, which is executed by a communication system. The psychoacoustic analysis method in this embodiment can be applied to a communication system, and there are no limitations on its application.
[0121] like Figure 6 As shown, the method may include, but is not limited to, the following steps:
[0122] S601: Determine multiple masking sources for an audio signal, wherein the masking sources include: tone masking sources and non-tone masking sources.
[0123] S602: If the critical frequency band to which the non-tone masking source belongs contains a tone masking source, then obtain the frequency distance between the tone masking source and the non-tone masking source.
[0124] S603: Obtain the distance threshold.
[0125] S604: If the frequency distance is less than the distance threshold, the tone masking source and / or the non-tone masking source are determined to be partially masked sources based on the frequency distance, the first sound pressure level of the tone masking source, and the second sound pressure level of the non-tone masking source.
[0126] In some embodiments, in a critical frequency band where a tone masking source exists, it is determined whether the frequency distance between the location (bark scale) of the tone masking source in the critical frequency band and the noise masking source (non-tone masking source) in the critical frequency band is less than a set distance threshold. If the frequency distance is less than the distance threshold, the sound pressure level of the tone masking source and the sound pressure level of the non-tone masking source can be further calculated. Then, by combining the frequency distance, the sound pressure level of the tone masking source and the sound pressure level of the non-tone masking source, the tone masking source and / or the non-tone masking source are determined to be partial masking sources, without any limitation.
[0127] The sound pressure level of the tone-masked source can be referred to as the first sound pressure level, and the sound pressure level of the non-tone-masked source can be referred to as the second sound pressure level.
[0128] The methods for calculating the sound pressure level of tone-masked sources and the sound pressure level of non-tone-masked sources can be found in relevant technologies, and will not be elaborated here.
[0129] In some embodiments, the selection of a tone-masking source and / or a non-tone-masking source can be determined based on frequency distance, the first sound pressure level of the tone-masking source, and the second sound pressure level of the non-tone-masking source, combined with set rules; or, the selection of a tone-masking source and / or a non-tone-masking source can be determined by combining a masking source selection model; of course, any other possible method can be used to determine whether to select a tone-masking source and / or a non-tone-masking source, without limitation.
[0130] S605: Analyze the masking threshold of an audio signal based on a subset of masking sources from multiple masking sources.
[0131] In this embodiment, by selecting a portion of the masking sources from all masking sources of the audio signal to participate in the masking threshold analysis and calculation, the computational load of psychoacoustic analysis can be effectively reduced, thereby lowering the computational complexity. When a tone masking source is included within the critical frequency band of a non-tone masking source, the frequency distance between the tone masking source and the non-tone masking source is obtained, and a distance threshold is acquired. If the frequency distance is less than the distance threshold, the tone masking source and / or the non-tone masking source are determined as partial masking sources based on the frequency distance, the first sound pressure level of the tone masking source, and the second sound pressure level of the non-tone masking source. This achieves accurate selection of partial masking sources from multiple masking sources, improving the accuracy of psychoacoustic analysis.
[0132] Figure 7a This is a flowchart illustrating another psychoacoustic analysis method provided in this embodiment, which is executed by a communication system. The psychoacoustic analysis method in this embodiment can be applied to a communication system, and there are no limitations thereto.
[0133] like Figure 7aAs shown, the method may include, but is not limited to, the following steps:
[0134] S701a: Identify multiple masking sources for an audio signal, wherein the masking sources include: tone masking sources and non-tone masking sources.
[0135] S702a: If the critical frequency band to which the non-tone masking source belongs contains a tone masking source, then obtain the frequency distance between the tone masking source and the non-tone masking source.
[0136] S703a: Obtain the distance threshold.
[0137] S704a: If the frequency distance is less than the distance threshold, then obtain the first sound pressure level of the tone-masked source and the second sound pressure level of the non-tone-masked source.
[0138] In some embodiments, in a critical frequency band where a tone masking source exists, it is determined whether the frequency distance between the location (bark scale) of the tone masking source in the critical frequency band and the noise masking source (non-tone masking source) in the critical frequency band is less than a set distance threshold. If the frequency distance is less than the distance threshold, the sound pressure level of the tone masking source and the sound pressure level of the non-tone masking source can be further calculated. Then, by combining the frequency distance, the sound pressure level of the tone masking source and the sound pressure level of the non-tone masking source, the tone masking source and / or the non-tone masking source are determined to be partial masking sources, without any limitation.
[0139] The sound pressure level of the tone-masked source can be referred to as the first sound pressure level, and the sound pressure level of the non-tone-masked source can be referred to as the second sound pressure level.
[0140] The methods for calculating the sound pressure level of tone-masked sources and the sound pressure level of non-tone-masked sources can be found in relevant technologies, and will not be elaborated here.
[0141] S705a: Identify pitch-masked sources and non-pitch-masked sources as partially masked sources, wherein the first sound pressure level is equal to the second sound pressure level.
[0142] In other words, in some possible embodiments, if the frequency distance is less than the distance threshold and the first sound pressure level of the tone masking source is the same as the second sound pressure level of the non-tone masking source, then both the tone masking source and the non-tone masking source can be selected to participate in the calculation of the masking threshold.
[0143] S706a: Analyze the masking threshold of an audio signal based on a subset of masking sources from multiple masking sources.
[0144] Therefore, in this embodiment, when the frequency distance is less than the distance threshold and the first sound pressure level of the tone masking source is the same as the second sound pressure level of the non-tone masking source, it is possible to accurately select some masking sources from multiple masking sources, thereby improving the accuracy of psychoacoustic analysis.
[0145] Figure 7b This is a flowchart illustrating another psychoacoustic analysis method provided in this embodiment, which is executed by a communication system. The psychoacoustic analysis method in this embodiment can be applied to a communication system, and there are no limitations thereto.
[0146] like Figure 7b As shown, the method may include, but is not limited to, the following steps:
[0147] S701b: Determine multiple masking sources for an audio signal, wherein the masking sources include: tone masking sources and non-tone masking sources.
[0148] S702b: If the critical frequency band to which the non-tone masking source belongs contains a tone masking source, then obtain the frequency distance between the tone masking source and the non-tone masking source.
[0149] S703b: Obtain the distance threshold.
[0150] S704b: If the frequency distance is less than the distance threshold, then obtain the first sound pressure level of the tone-masked source and the second sound pressure level of the non-tone-masked source.
[0151] S705b: Determine a first attenuation value based on frequency distance, and determine a tone masking source and / or a non-tone masking source as a partial masking source based on the first attenuation value, a first sound pressure level, and a second sound pressure level, wherein the first sound pressure level is greater than the second sound pressure level, and the first attenuation value is the attenuation value of the masking ability of the tone masking source at the location of the non-tone masking source.
[0152] In other words, when the frequency distance between the tone-masking source and the non-tone-masking source in the critical frequency band is less than the distance threshold, and the first sound pressure level of the tone-masking source is greater than the second sound pressure level of the non-tone-masking source, the attenuation value of the masking ability of the tone-masking source at the location of the non-tone-masking source can be calculated. This attenuation value can be called the first attenuation value. Then, the tone-masking source and / or the non-tone-masking source are determined as partial masking sources based on the first attenuation value, the first sound pressure level, and the second sound pressure level.
[0153] In some embodiments, a first attenuation value of the masking capability of a tone masking source at a non-tone masking source location can be calculated based on the following formula:
[0154] The approximate expression for the masking spread function of a strong masking component at its adjacent positions is:
[0155]
[0156] The unit of s(Δb) is dB (decibels), and the unit of Δb is bark. It represents the distance to the masking source (an optional example of distance, such as frequency distance), and the value range is defined as (-0.7 bark, 0.7 bark).
[0157] Within the specified range of values, its function graph is as follows: Figure 8 As shown, Figure 8 This is a schematic diagram of the masking spread function in an embodiment of this disclosure. s(Δb) can be approximated as a piecewise line segment s ′ (Δb), s ′ The expression for (Δb) is:
[0158]
[0159] In some embodiments, a possible value of Δb represents the frequency distance between tone-masking sources and non-tone-masking sources in the critical band. The first attenuation value s can then be calculated by substituting the frequency distance into the above formula. ′ (Δb), where Δb is the frequency distance. The above formula can determine how much the masking ability of a stronger masking component changes at the location of another masking component. s(Δb) can be calculated using only the distance on the bark scale.
[0160] In some embodiments of this disclosure, when performing the step of determining a tone-masking source and / or a non-tone-masking source as a partial masking source based on a first attenuation value, a first sound pressure level, and a second sound pressure level, the tone-masking source may be determined as a partial masking source if the sum of the first sound pressure level and the first attenuation value is greater than the second sound pressure level. This is not a limitation.
[0161] In some other embodiments of this disclosure, when performing the step of determining the tone masking source and / or the non-tone masking source as a partial masking source based on the first attenuation value, the first sound pressure level, and the second sound pressure level, the tone masking source and the non-tone masking source may be determined as partial masking sources if the sum of the first sound pressure level and the first attenuation value is less than or equal to the second sound pressure level. This is not a limitation.
[0162] For example, the first sound pressure level and the first attenuation value can be summed to obtain the sum of the first sound pressure level and the first attenuation value. If the sum of the first sound pressure level and the first attenuation value is greater than the second sound pressure level, then the tone masking source is determined to be a partially masked source. If the sum of the first sound pressure level and the first attenuation value is less than or equal to the second sound pressure level, then the tone masking source and the non-tone masking source are determined to be partially masked sources.
[0163] S706b: Analyze the masking threshold of an audio signal based on a subset of masking sources from multiple masking sources.
[0164] Therefore, in this embodiment of the present disclosure, it is possible to select a tone masking source or a non-tone masking source as a partial masking source based on the attenuation of the masking ability of the tone masking source at the location of the non-tone masking source, thereby effectively improving the accuracy of the selection of partial masking sources and thus effectively improving the accuracy of psychoacoustic analysis.
[0165] Figure 7c This is a flowchart illustrating another psychoacoustic analysis method provided in this embodiment, which is executed by a communication system. The psychoacoustic analysis method in this embodiment can be applied to a communication system, and there are no limitations thereto.
[0166] like Figure 7c As shown, the method may include, but is not limited to, the following steps:
[0167] S701c: Determine multiple masking sources for an audio signal, wherein the masking sources include: tone masking sources and non-tone masking sources.
[0168] S702c: If the critical frequency band to which the non-tone masking source belongs contains a tone masking source, then obtain the frequency distance between the tone masking source and the non-tone masking source.
[0169] S703c: Obtain distance threshold.
[0170] S704c: If the frequency distance is less than the distance threshold, then obtain the first sound pressure level of the tone-masked source and the second sound pressure level of the non-tone-masked source.
[0171] S705c: Determines a second attenuation value based on frequency distance, and determines a tone masking source and / or a non-tone masking source as a partial masking source based on the second attenuation value, the first sound pressure level, and the second sound pressure level. The second attenuation value is the attenuation value of the masking ability of the non-tone masking source at the position of the tone masking source, and the first sound pressure level is less than the second sound pressure level.
[0172] In other words, when the frequency distance between the tone-masking source and the non-tone-masking source in the critical frequency band is less than the distance threshold, and the first sound pressure level of the tone-masking source is less than the second sound pressure level of the non-tone-masking source, the attenuation value of the masking ability of the non-tone-masking source at the position of the tone-masking source can be calculated. This attenuation value can be called the second attenuation value. Then, based on the second attenuation value, the first sound pressure level, and the second sound pressure level, the tone-masking source and / or the non-tone-masking source are determined to be partially masked sources.
[0173] In some embodiments, a second attenuation value of the masking capability of the non-tone masking source at the location of the tone masking source can be calculated based on the above formula:
[0174] In some embodiments, a possible value of Δb represents the frequency distance between tone-masking sources and non-tone-masking sources in the critical band. The second attenuation value s can then be calculated by substituting the frequency distance into the above formula. ′ (Δb), where Δb is the frequency distance. The above formula can determine how much the masking ability of a stronger masking component changes at the location of another masking component. s(Δb) can be calculated using only the distance on the bark scale.
[0175] In some embodiments, when performing the step of determining a tone-masking source and / or a non-tone-masking source as a partial masking source based on a second attenuation value, a first sound pressure level, and a second sound pressure level, the non-tone-masking source may be determined as a partial masking source if the sum of the second sound pressure level and the second attenuation value is greater than the first sound pressure level; there is no limitation on this.
[0176] In other embodiments, when performing the step of determining the tone masking source and / or non-tone masking source as a partial masking source based on the second attenuation value, the first sound pressure level, and the second sound pressure level, the tone masking source and non-tone masking source can be determined as partial masking sources if the sum of the second sound pressure level and the second attenuation value is less than or equal to the first sound pressure level, without limitation.
[0177] For example, the second sound pressure level and the second attenuation value can be summed to obtain the sum of the second sound pressure level and the second attenuation value. If the sum of the second sound pressure level and the second attenuation value is greater than the first sound pressure level, then the non-tone masking source is determined to be a partially masked source. If the sum of the second sound pressure level and the second attenuation value is less than or equal to the first sound pressure level, then the tone masking source and the non-tone masking source are determined to be partially masked sources.
[0178] S706c: Analyzes the masking threshold of an audio signal based on a subset of masking sources from multiple masking sources.
[0179] Therefore, in this embodiment of the present disclosure, it is possible to select either a tone masking source or a non-tone masking source as a partial masking source based on the attenuation of the masking ability of the non-tone masking source at the position of the tone masking source, thereby effectively improving the accuracy of the selection of partial masking sources and thus effectively improving the accuracy of psychoacoustic analysis.
[0180] The method for selecting partial masking sources in the above embodiments of this disclosure can also be called a cross-masking algorithm, as illustrated below:
[0181] like Figure 9 As shown, Figure 9This is a schematic diagram of the architecture of the psychoacoustic analysis method in this embodiment. After extracting the tone masking source and / or non-tone masking source of the audio signal, a cross-masking algorithm is executed to select a portion of the masking sources from multiple masking sources, and a masking threshold is calculated based on the selected portion of the masking sources.
[0182] The calculation process for cross-masking is as follows:
[0183] ① In the critical frequency band where a tone masking source exists, determine whether the distance between the position of the tone masking source (bark scale) in the critical frequency band and the position of the noise masking source in the critical frequency band is less than a set threshold. If it is less, proceed to the next step.
[0184] ② Determine which sound pressure level is greater, the tone masking source or the noise masking source. Consider the attenuation s(Δb) of the larger masking ability at the location of the smaller one. Determine whether the larger sound pressure level +s(Δb) or the smaller sound pressure level is greater. If the former is greater, the latter is masked. In other words, in this case, the latter will be excluded from the calculation of the final masking threshold.
[0185] ③ If there are other tone masking sources in the critical frequency band, repeat process ① and ②; otherwise, proceed to the next critical frequency band containing tone masking sources.
[0186] In this embodiment, the threshold is set to a fixed value of 0.5 bark, and the corresponding processing flow is as follows.
[0187] ① In the critical frequency band where a tone masking source exists, determine whether the distance between the location of the tone masking source in the critical frequency band (on a bark scale) and the location of the noise masking source in the critical frequency band is less than ±0.5 bark. If it is less, proceed to the next step.
[0188] ② Determine which sound pressure level is higher, the tone masking source or the noise masking source. Consider the attenuation of the masking ability of the higher source at the location of the lower source. Determine the sound pressure level of the higher source + s ′ Which is larger, (Δb) or the smaller sound pressure level? If the former is larger, the latter will be masked. In other words, in this case, the latter will be excluded from the calculation of the final masking threshold.
[0189] ③ If there are other tone signals in the critical frequency band, repeat process ① and ②; otherwise, proceed to the next critical frequency band containing tone signals.
[0190] like Figure 10a and 10b As shown, Figure 10a This is a schematic diagram illustrating experimental statistics on the time taken to enable cross-masking in this embodiment of the present disclosure. Figure 10bThis is a schematic diagram of experimental statistics on the time taken when cross-masking is turned off in this embodiment of the present disclosure. "times" means the time required for the encoder to run the psychoacoustic model in each frame. Each frame outputs the time that the program has spent on the psychoacoustic model. The last output represents the total time spent on the psychoacoustic model when the encoder finishes encoding. Figure 10a and 10b These are the number of clock units consumed when running 1476 frames of audio with cross masking enabled and disabled, respectively. As you can see, although the number of clock units required to run the program fluctuates, enabling cross masking is on average about 150 clock units faster than disabling it.
[0191] In another embodiment of this disclosure, considering that the human ear can perceive sound in the frequency range of 20Hz-20kHz, and is most sensitive to sound in the frequency range of 1kHz-3kHz, a smaller distance threshold is used in the low-frequency critical band to retain more masking source components, while a larger distance threshold is selected in the high-frequency critical band to remove more masking source components, which is beneficial to the subjective auditory perception.
[0192] In another embodiment of this disclosure, a distance threshold of 0.3 bark can be set for the first 15 critical frequency bands and a distance threshold of 0.6 bark can be set for the last 10 critical frequency bands. The selection of the distance threshold is related to the frequency, and the higher the frequency of the critical frequency band, the larger the selected distance threshold.
[0193] The corresponding processing flow is as follows.
[0194] ① In the critical frequency band where a tone masking source exists, determine whether the distance between the position of the tone masking source in the critical frequency band (bark scale) and the position of the noise masking source in the critical frequency band is less than a set distance threshold (the distance threshold is related to the critical frequency band in which it is located). If it is less than the threshold, proceed to the next step.
[0195] ② Determine which sound pressure level is higher, the tone masking source or the noise masking source. Consider the attenuation of the masking ability of the higher source at the location of the lower source. Determine the sound pressure level of the higher source + s ′ Which is greater, (Δb) or the smaller sound pressure level? If the former is greater, the latter will be masked. In other words, in this case, the latter will be excluded from the calculation of the final masking threshold.
[0196] ③ If there are other tone masking sources in the critical frequency band, repeat process ① and ②; otherwise, proceed to the next critical frequency band containing tone masking sources.
[0197] Therefore, in this embodiment of the disclosure, the number of unnecessary masking source components involved in the masking threshold calculation is reduced by using the cross-masking algorithm, thereby reducing the overall computational complexity of the psychoacoustic model, while the masking threshold remains almost unchanged, and the power consumption of the device is also effectively reduced.
[0198] Figure 11 This is a schematic diagram of the structure of a communication device provided in an embodiment of this disclosure. Figure 11 The communication device 110 shown may include a transceiver module 1101 and a processing module 1102. The transceiver module 1101 may include a sending module and / or a receiving module. The sending module is used to implement the sending function, and the receiving module is used to implement the receiving function. The transceiver module 1101 can implement the sending function and / or the receiving function.
[0199] The communication device 110 may be a terminal device (such as the terminal device in the foregoing method embodiments), a device within a terminal device, or a device that can be used in conjunction with a terminal device. Alternatively, the communication device 110 may be a network device (such as the network device in the foregoing method embodiments), a device within a network device, or a device that can be used in conjunction with a network device.
[0200] Communication device 110, the device comprising:
[0201] Processing module 1102 is used to determine multiple masking sources of the audio signal; and to analyze the masking threshold of the audio signal based on some of the masking sources.
[0202] In this embodiment, multiple masking sources of the audio signal are determined, and then the masking threshold of the audio signal is analyzed based on some of the masking sources. Since some masking sources are selected from all the masking sources of the audio signal to participate in the analysis and calculation of the masking threshold, the amount of computation in psychoacoustic analysis can be effectively reduced, thereby reducing the computational complexity.
[0203] Figure 12 This is a schematic diagram of another communication device provided in an embodiment of this disclosure. The communication device 120 can be a network device, a terminal device, a chip, chip system, or processor that supports the network device in implementing the above methods, or a chip, chip system, or processor that supports the terminal device in implementing the above methods. This device can be used to implement the methods described in the above method embodiments; for details, please refer to the descriptions in the above method embodiments.
[0204] The communication device 120 may include one or more processors 1201. The processor 1201 may be a general-purpose processor or a dedicated processor, such as a baseband processor or a central processing unit (CPU). The baseband processor can be used to process communication protocols and communication data, while the CPU can be used to control the communication device (e.g., base station, baseband chip, terminal equipment, terminal equipment chip, DU or CU, etc.), execute computer programs, and process data from the computer programs.
[0205] Optionally, the communication device 120 may further include one or more memories 1202, on which computer programs 1204 may be stored, and the processor 1201 may store computer programs 1203. The processor 1201 executes computer programs 1204 and / or computer programs 1203 to cause the communication device 120 to perform the methods described in the above method embodiments. Optionally, the memories 1202 may also store data. The communication device 120 and the memories 1202 may be provided separately or integrated together.
[0206] Optionally, the communication device 120 may also include a transceiver 1205 and an antenna 1206. The transceiver 1205 may be referred to as a transceiver unit, transceiver, or transceiver circuit, etc., and is used to implement the transmission and reception functions. The transceiver 1205 may include a receiver and a transmitter. The receiver may be referred to as a receiver or receiving circuit, etc., and is used to implement the receiving function; the transmitter may be referred to as a transmitter or transmitting circuit, etc., and is used to implement the transmitting function.
[0207] Optionally, the communication device 120 may further include one or more interface circuits 1207. The interface circuits 1207 are used to receive code instructions and transmit them to the processor 1201. The processor 1201 executes the code instructions to cause the communication device 120 to perform the methods described in the above method embodiments.
[0208] In one implementation, the processor 1201 may include a transceiver for implementing receiving and transmitting functions. For example, the transceiver may be a transceiver circuit, an interface, or an interface circuit. The transceiver circuit, interface, or interface circuit for implementing receiving and transmitting functions may be separate or integrated. The aforementioned transceiver circuit, interface, or interface circuit can be used for reading and writing code / data, or it can be used for transmitting or relaying signals.
[0209] In one implementation, processor 1201 may store computer program 1203, which runs on processor 1201 and causes communication device 120 to perform the methods described in the above method embodiments. Computer program 1203 may be embedded in processor 1201, in which case processor 1201 may be implemented in hardware.
[0210] In one implementation, the communication device 120 may include circuitry capable of performing the functions of transmitting, receiving, or communicating as described in the aforementioned method embodiments. The processor and transceiver described in this disclosure can be implemented on integrated circuits (ICs), analog ICs, radio frequency integrated circuits (RFICs), mixed-signal ICs, application-specific integrated circuits (ASICs), printed circuit boards (PCBs), electronic devices, etc. The processor and transceiver can also be manufactured using various IC process technologies, such as complementary metal-oxide-semiconductor (CMOS), n-metal-oxide-semiconductor (NMOS), positive-channel metal-oxide-semiconductor (PMOS), bipolar junction transistors (BJTs), bipolar CMOS (BiCMOS), silicon-germanium (SiGe), gallium arsenide (GaAs), etc.
[0211] The communication device described in the above embodiments may be a network device or a terminal device, but the scope of the communication device described in this disclosure is not limited thereto, and the structure of the communication device may vary. Figure 12 The communication device can be a standalone device or part of a larger device. For example, the communication device can be:
[0212] (1) Independent integrated circuit IC, or chip, or chip system or subsystem;
[0213] (2) A collection of one or more ICs, optionally including storage components for storing data and computer programs;
[0214] (3) ASIC, such as modem;
[0215] (4) Modules that can be embedded in other devices;
[0216] (5) Receivers, terminal equipment, smart terminal equipment, cellular phones, wireless equipment, handheld devices, mobile units, vehicle-mounted equipment, network equipment, cloud equipment, artificial intelligence equipment, etc.
[0217] (6) Others, etc.
[0218] For cases where the communication device can be a chip or a chip system, please refer to [link / reference]. Figure 13 The diagram shows the structure of the chip. Figure 13 The chip shown includes a processor 1301 and an interface 1302. There can be one or more processors 1301, and multiple interfaces 1302.
[0219] For cases where the chip is used to implement the functions of the communication system (e.g., including terminal devices and / or network devices) in the embodiments of this application:
[0220] Processor 1301, used to implement Figure 2 - Figure 1 The steps in 0.
[0221] Optionally, the chip also includes a memory 1303, which is used to store necessary computer programs and data.
[0222] Those skilled in the art will also understand that the various illustrative logical blocks and steps listed in the embodiments of this disclosure can be implemented by electronic hardware, computer software, or a combination of both. Whether such functionality is implemented in hardware or software depends on the specific application and the overall system design requirements. Those skilled in the art can implement the functionality using various methods for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of this disclosure.
[0223] This disclosure also provides a communication system, which includes the aforementioned... Figure 11 The communication device in the embodiment, or the system including the aforementioned Figure 12 The communication device in the embodiment.
[0224] This disclosure also provides a readable storage medium having instructions stored thereon that, when executed by a computer, implement the functions of any of the above method embodiments.
[0225] This disclosure also provides a computer program product that, when executed by a computer, implements the functions of any of the above method embodiments.
[0226] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer programs. When a computer program is loaded and executed on a computer, it generates, in whole or in part, the flow or function according to the embodiments of this disclosure. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer program can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, a computer program can be transferred from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).
[0227] Those skilled in the art will understand that the various numerical designations such as "first," "second," etc., used in this disclosure are merely for the convenience of description and are not intended to limit the scope of the embodiments of this disclosure, nor do they indicate the order of events.
[0228] At least one of the features described in this disclosure can also be described as one or more, and multiple features can be two, three, four or more, and this disclosure does not impose any limitations. In the embodiments of this disclosure, for a technical feature, the technical features in that technical feature are distinguished by "first", "second", "third", "A", "B", "C" and "D", etc., and there is no sequential order or size order among the technical features described by "first", "second", "third", "A", "B", "C" and "D".
[0229] The correspondences shown in the tables of this disclosure can be configured or predefined. The values of the information in each table are merely examples and can be configured to other values; this disclosure is not limiting. When configuring the correspondences between information and parameters, it is not necessarily required to configure all the correspondences shown in each table. For example, the correspondences shown in some rows of the tables in this disclosure may not be configured. Furthermore, appropriate modifications and adjustments can be made based on the above tables, such as splitting, merging, etc. The names of the parameters shown in the headers of the above tables can also use other names that the communication device can understand, and the values or representations of the parameters can also be other values or representations that the communication device can understand. In the implementation of the above tables, other data structures can also be used, such as arrays, queues, containers, stacks, linear lists, pointers, linked lists, trees, graphs, structures, classes, heaps, hash tables, or hash tables, etc.
[0230] The predefined terms in this disclosure can be understood as defined, predefined, stored, pre-stored, pre-negotiated, pre-configured, solidified, or pre-burned.
[0231] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0232] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0233] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A psychoacoustic analysis method, characterized in that, The method includes: Identify multiple masking sources for the audio signal; Based on a subset of the multiple masking sources, the masking threshold of the audio signal is analyzed; The masking source includes: a non-tone masking source; wherein, determining the partial masking source from the plurality of masking sources includes: if the critical frequency band to which the non-tone masking source belongs does not contain a tone masking source, then the non-tone masking source is determined to be the partial masking source.
2. The method as described in claim 1, characterized in that, The masking source also includes: tone masking source.
3. The method according to any one of claims 1-2, characterized in that, The method further includes: From the plurality of masking sources, the partial masking sources are determined.
4. The method according to any one of claims 1-3, characterized in that, The masking source further includes: a tone masking source; wherein, determining the partial masking source from the plurality of masking sources includes: If the critical frequency band to which the non-tone masking source belongs contains the tone masking source, then the frequency distance between the tone masking source and the non-tone masking source is obtained; Based on the frequency distance, the tone masking source and / or the non-tone masking source are determined as the partial masking source.
5. The method as described in claim 4, characterized in that, The step of determining the tone masking source and / or the non-tone masking source as the partial masking source based on the frequency distance includes: Obtain the distance threshold; If the frequency distance is greater than or equal to the distance threshold, then the tone masking source and the non-tone masking source are determined to be the partial masking sources; If the frequency distance is less than the distance threshold, then the tone masking source and / or the non-tone masking source are determined to be the partial masking source based on the frequency distance, the first sound pressure level of the tone masking source, and the second sound pressure level of the non-tone masking source.
6. The method as described in claim 5, characterized in that, The distance threshold obtained includes at least one of the following: Get the set distance threshold; Obtain the distance threshold corresponding to the critical frequency band; The distance threshold is determined based on the frequency of the critical band.
7. The method according to any one of claims 5-6, characterized in that, The step of determining the tone-masking source and / or the non-tone-masking source as the partial masking source based on the frequency distance, the first sound pressure level of the tone-masking source, and the second sound pressure level of the non-tone-masking source includes at least one of the following: The tone masking source and the non-tone masking source are identified as the partial masking sources, wherein the first sound pressure level is equal to the second sound pressure level; A first attenuation value is determined based on the frequency distance, and the tone masking source and / or the non-tone masking source are determined as the partial masking source based on the first attenuation value, the first sound pressure level, and the second sound pressure level, wherein the first sound pressure level is greater than the second sound pressure level, and the first attenuation value is the attenuation value of the masking ability of the tone masking source at the location of the non-tone masking source. A second attenuation value is determined based on the frequency distance, and the tone masking source and / or the non-tone masking source are determined as the partial masking source based on the second attenuation value, the first sound pressure level, and the second sound pressure level, wherein the second attenuation value is the attenuation value of the masking ability of the non-tone masking source at the position of the tone masking source, and the first sound pressure level is less than the second sound pressure level.
8. The method as described in claim 7, characterized in that, The step of determining the tone masking source and / or the non-tone masking source as the partial masking source based on the first attenuation value, the first sound pressure level, and the second sound pressure level includes at least one of the following: If the sum of the first sound pressure level and the first attenuation value is greater than the second sound pressure level, then the tone masking source is determined to be the partial masking source; If the sum of the first sound pressure level and the first attenuation value is less than or equal to the second sound pressure level, then the tone masking source and the non-tone masking source are determined to be the partial masking sources.
9. The method as described in claim 7, characterized in that, The step of determining the tone masking source and / or the non-tone masking source as the partial masking source based on the second attenuation value, the first sound pressure level, and the second sound pressure level includes at least one of the following: If the sum of the second sound pressure level and the second attenuation value is greater than the first sound pressure level, then the non-tone masking source is determined to be the partial masking source; If the sum of the second sound pressure level and the second attenuation value is less than or equal to the first sound pressure level, then the tone masking source and the non-tone masking source are determined to be the partial masking sources.
10. A communication device, characterized in that, The device includes: The processing module is used to determine multiple masking sources of the audio signal and analyze the masking threshold of the audio signal based on a portion of the multiple masking sources. The masking source includes: a non-tone masking source; wherein, determining the partial masking source from the plurality of masking sources includes: if the critical frequency band to which the non-tone masking source belongs does not contain a tone masking source, then the non-tone masking source is determined to be the partial masking source.
11. A communication system, characterized in that, The communication system includes network devices and terminal devices, wherein the network devices perform the method as described in any one of claims 1-9, and / or the terminal devices perform the method as described in any one of claims 1-9.
12. A computer-readable storage medium for storing instructions that, when executed, cause the method of any one of claims 1-9 to be implemented.
Citation Information
Patent Citations
Device and process for use in encoding audio data
US20040243397A1