Audio processing method and device and storage medium

By performing equal-length slicing on the original audio data and collecting EEG signals, and adjusting the gamma frequency band energy of the audio slices, the problem of poor AD audio treatment effect is solved, the treatment effect is improved, and the auditory experience is taken into account.

CN120640199APending Publication Date: 2025-09-12LIAO TECH (TIANJIN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510671982.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The AD audio therapy in the existing technology is relatively poor and lacks organic integration with music signals, resulting in low patient acceptance and affecting the treatment effect.

Method used

By slicing the original audio data into equal lengths, determining the masking threshold curve, and combining it with EEG signal acquisition, the gamma frequency band energy of the audio slice is adjusted, taking into account the user's auditory experience and enhancing the gamma frequency band energy of the next audio slice.

Benefits of technology

It improves the effect of AD audio therapy, takes into account the user's auditory experience, and enhances the pertinence and effectiveness of treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120640199A_ABST
    Figure CN120640199A_ABST
Patent Text Reader

Abstract

The invention discloses an audio processing method and device and a storage medium. Wherein the original audio data is subjected to equal-length slicing to obtain a plurality of audio slices, a masking threshold curve corresponding to the current audio slice is determined, and the masking threshold curve is used for indicating a first masking threshold corresponding to the frequency index of each masker; playing the current audio slice to the user, collecting an electroencephalogram signal of the user, and determining an electroencephalogram signal slice corresponding to the current audio slice; and according to the spectrum characteristics of the gamma frequency band in the electroencephalogram signal slice and the masking threshold curve, enhancing the energy of the gamma frequency band in the next audio slice of the current audio slice to obtain the enhanced next audio slice, so that enhancement of the energy of the gamma frequency band in the audio slices and auditory experience of a user can be considered; therefore, the treatment effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio technology, and in particular to an audio processing method, device, and storage medium. Background Art

[0002] With the rapid development of digital multimedia technology, the importance of audio signal processing and optimization in the field of medical rehabilitation has become increasingly prominent.

[0003] In the field of neurodegenerative disease treatment, Alzheimer's disease (AD), a common cognitive impairment in the elderly, has attracted significant attention for its non-drug interventions. Studies have found that 40 Hz audio stimulation can effectively activate gamma oscillations in the brain, promoting neuronal activity and thus improving patients' cognitive function. This is particularly true when patients are in a calm state, when gamma frequency activity is low. By precisely controlling the intensity of specific audio frequencies, neuronal activity can be stimulated in a targeted manner, slowing disease progression.

[0004] In existing technologies, AD treatment through audio therapy uses a single 40 Hz pure tone signal. This method lacks an organic combination with music signals, which may lead to low patient acceptance and reduce the treatment effect.

[0005] With respect to the technical problem of poor AD audio treatment effect existing in the above-mentioned prior art, no effective solution has been proposed so far. Summary of the Invention

[0006] The embodiments of the present disclosure provide an audio processing method, device, and storage medium to at least solve the technical problem of poor AD audio treatment effect in the prior art.

[0007] According to one aspect of an embodiment of the present disclosure, an audio processing method is provided, including: slicing original audio data of equal length to obtain a plurality of audio slices; determining a masking threshold curve corresponding to a current audio slice, the masking threshold curve being used to indicate a first masking threshold corresponding to a frequency index of each masker; playing the current audio slice to a user while collecting the user's electroencephalogram (EEG) signal to determine an EEG signal slice corresponding to the current audio slice; and enhancing the energy of a gamma frequency band in an audio slice next to the current audio slice based on spectral characteristics of the gamma frequency band in the EEG signal slice and the masking threshold curve to obtain an enhanced next audio slice.

[0008] According to another aspect of an embodiment of the present disclosure, a storage medium is further provided, the storage medium including a stored program, wherein when the program is run, a processor executes any one of the above methods.

[0009] According to another aspect of an embodiment of the present disclosure, an audio processing device is also provided, including: an audio slicing module, used to slice the original audio data into equal lengths to obtain a plurality of audio slices; a masking threshold curve determination module, used to determine the masking threshold curve corresponding to the current audio slice, the masking threshold curve being used to indicate a first masking threshold corresponding to the frequency index of each masker; an EEG signal acquisition module, used to play the current audio slice to the user, and at the same time collect the user's EEG signal to determine the EEG signal slice corresponding to the current audio slice; and an audio adjustment module, which enhances the energy of the gamma frequency band in the next audio slice of the current audio slice according to the spectral characteristics of the gamma frequency band in the EEG signal slice and the masking threshold curve to obtain the enhanced next audio slice.

[0010] According to another aspect of an embodiment of the present disclosure, an audio processing device is also provided, including: a processor; and a memory, connected to the processor, for providing the processor with instructions for processing the following processing steps: slicing the original audio data of equal length to obtain a plurality of audio slices; determining a masking threshold curve corresponding to the current audio slice, the masking threshold curve being used to indicate a first masking threshold corresponding to the frequency index of each masker; playing the current audio slice to the user, and simultaneously collecting the user's EEG signal to determine the EEG signal slice corresponding to the current audio slice; and enhancing the energy of the gamma frequency band in the next audio slice of the current audio slice based on the spectral characteristics of the gamma frequency band in the EEG signal slice and the masking threshold curve to obtain the enhanced next audio slice.

[0011] In the embodiment of the present disclosure, when the current audio slice is played, the energy of the next audio slice in the gamma frequency band can be enhanced by the spectral characteristics of the gamma frequency band in the EEG signal slice corresponding to the current audio slice (specifically, the frequency with the largest amplitude in the gamma band). However, in order to take into account the user's auditory experience, the energy enhancement operation of the next audio slice can be auditorily shielded by using a masking threshold curve. Thus, when the next audio slice is played to the user, the user will not directly feel the energy enhancement operation. Therefore, the present method can take into account both the energy enhancement of the gamma frequency band of the audio slice and the user's auditory experience, and by applying the present method to AD audio treatment, the treatment effect can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of this application. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:

[0013] Figure 1 is a hardware structure block diagram of a computing device for implementing the method according to embodiment 1 of the present disclosure;

[0014] Figure 2 is a flowchart of the audio processing method according to the first aspect of embodiment 1 of the present disclosure;

[0015] Figure 3 This is a schematic diagram of an audio slice time domain signal according to the first aspect provided by Embodiment 1 of the present disclosure;

[0016] Figure 4 This is a schematic diagram of a frequency domain subband signal according to the first aspect provided by Embodiment 1 of the present disclosure;

[0017] Figure 5 is a schematic diagram of a masking threshold curve according to the first aspect provided in Example 1 of the present disclosure;

[0018] Figure 6 is a schematic diagram of an audio processing device according to the first aspect of embodiment 2 of the present disclosure; and

[0019] Figure 7 It is a schematic diagram of the audio processing device according to the first aspect of embodiment 3 of the present disclosure. DETAILED DESCRIPTION

[0020] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present disclosure.

[0021] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0022] Example 1

[0023] According to this embodiment, an embodiment of a method for audio processing is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0024] The method embodiment provided in this embodiment can be executed in a mobile terminal, a computer terminal, a server or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computing device for implementing an audio processing method. Figure 1 As shown, the computing device may include one or more processors (the processor may include but is not limited to a microprocessor MCU or a programmable logic device FPGA, etc.), a memory for storing data, a transmission device for communication functions, and an input / output interface. The memory, transmission device, and input / output interface are connected to the processor via a bus. In addition, it may also include: a display, a keyboard, and a cursor control device connected to the input / output interface. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0025] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computing device. As described in the embodiments of the present disclosure, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0026] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the audio processing method in the embodiment of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the audio processing method of the above-mentioned application. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computing device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0027] The transmission device is used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by a communications provider of the computing device. In one embodiment, the transmission device includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0028] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computing device.

[0029] It should be noted that, in some optional embodiments, the above Figure 1 The computing device shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the computing device described above.

[0030] Under the above operating environment, according to the first aspect of this embodiment, an audio processing method is provided, which can be Figure 1 The computing device implementation shown. Figure 2 A schematic diagram showing the process of the method is shown in FIG. Figure 2 As shown, the method includes:

[0031] S202: Slice the original audio data into equal length slices to obtain a plurality of audio slices;

[0032] S204: Determine a masking threshold curve corresponding to the current audio slice, where the masking threshold curve is used to indicate a first masking threshold corresponding to a frequency index of each masker;

[0033] S206: Play the current audio slice to the user, and simultaneously collect the user's EEG signal to determine the EEG signal slice corresponding to the current audio slice;

[0034] S208: According to the spectrum characteristics of the gamma frequency band in the EEG signal slice and the masking threshold curve, the energy of the gamma frequency band in the next audio slice of the current audio slice is enhanced to obtain the enhanced next audio slice.

[0035] Specifically, the computing device may slice the original audio data into equal lengths to obtain a plurality of audio slices (S202). The original audio data mentioned here may be audio data containing normal music.

[0036] The purpose of slicing the raw audio data into equal lengths is to pre-process the original audio file for subsequent spectrum analysis and energy regulation. Each audio slice will serve as the basic unit for subsequent processing. Subsequent processing will be performed in the order of the slices. As will be mentioned later, the computing device can play each audio slice sequentially to the user. When playing the current audio slice to the user, the computing device adjusts the next audio slice after the current audio slice by collecting the corresponding EEG signal and the masking threshold curve determined by the psychoacoustic model. The specific adjustment method will be described in detail later. Figure 3 A schematic diagram of the audio slice time domain signal provided in the first aspect of embodiment 1 of the present disclosure.

[0037] The computing device may determine a masking threshold curve corresponding to the current audio slice, the masking threshold curve being used to indicate first masking thresholds corresponding to frequency indexes of respective maskers (S204). The first masking threshold corresponding to a frequency index of a masker is used to indicate that when a frequency of an audio signal is at the frequency index of the masker, the amplitude of the audio signal is lower than the first masking threshold and will be masked by the listener.

[0038] Optionally, the operation of determining the masking threshold curve corresponding to the current audio slice specifically includes: performing sub-band filtering on the time domain signal of the current audio slice to determine the frequency domain sub-band signal corresponding to the current audio slice; converting the frequency domain sub-band signal to a critical band rate, and determining the first masking threshold corresponding to the frequency index of each masker based on the frequency domain sub-band signal converted to the critical band rate to obtain the above-mentioned masking threshold curve.

[0039] The above-mentioned sub-band filtering of the time domain signal of the current audio slice can be specifically performed by using sub-band filtering technology to decompose the time domain signal of the current audio slice into 32 frequency domain sub-band signals. i Traverse according to the window size of 512, and use j to represent the window index. Use the decomposition window coefficient C[k] with a length of 512 and the current audio slice s i Multiply the corresponding window values ​​in to get the windowed sample sequence B i,j [k], this process is to reduce spectrum leakage, the formula is as follows:

[0040] B i,j [k]=C[k]·s i,j [k], (k=0,1,...,511)

[0041] Secondly, the sample points are grouped and calculated: In order to preliminarily aggregate the frequency domain features and facilitate the division of subbands, the windowed sample sequence is grouped and summed to obtain a fused sequence Y with a length of 64 i,j , the formula is as follows:

[0042]

[0043] Finally, the frequency domain sub-band signal is calculated and decomposed according to different sub-bands. The analysis matrix acts as a filter and is used with the analysis matrix M and Y. i,j [k] weighted summation to obtain the frequency domain subband signal S i,j [k]:

[0044]

[0045] The analysis matrix M is:

[0046]

[0047] Figure 4 This is a schematic diagram of a frequency domain subband signal according to the first aspect provided by Embodiment 1 of the present disclosure;

[0048] exist Figure 4 The frequency domain subband signal of an audio slice is shown in Figure 4 It can be seen from the figure that there are 32 sub-bands in the embodiment 1 of the present disclosure, corresponding to k=0 to 31, and Figure 4 The amplitude corresponding to each sub-band is S i,j [k].

[0049] Then, the frequency domain subband signal S i,j [k] is converted to the critical band rate, from the frequency f (unit is Hz) to the critical band rate (unit is Bark), and the frequency domain subband signal z converted to the critical band rate is obtainedi,j [k], the formula is as follows:

[0050]

[0051] Therefore, the first masking threshold corresponding to the frequency index of each masker can be determined according to the frequency domain subband signal converted to the critical band rate to obtain the masking threshold curve.

[0052] Optionally, the operation of determining the first masking threshold corresponding to the frequency index of each masker based on the frequency domain subband signal converted to the critical band rate includes: judging whether the frequency domain subband signal corresponding to the local maximum is a tonal component or a non-tonal component based on the amplitude difference between the local maximum and the adjacent points in the frequency domain subband signal converted to the critical band rate; determining a first masking function corresponding to the tonal component and a second masking function corresponding to the non-tonal component; and determining the first masking threshold corresponding to the frequency index of each masker based on the first masking function and the second masking function.

[0053] First, distinguish z i,j The masker in [k] is a tonal component or a non-tonal component (noise). Calculate z i,j If the amplitude of the two adjacent points around a local maximum frequency is at least 7dB lower than that of the local maximum, the point can be recorded as a tonal component; otherwise, it is classified as a non-tonal component. i The set of frequency indices of all the tonal components in the jth window is U T , the frequency index set determined to be the tonal component is U NT .

[0054] Then, the masking degree of the tonal component and the non-tonal component themselves is calculated. The calculation methods of the two are different, as shown in the following formula.

[0055] av tm (p) = -6.025 - 0.275z p ,(z p ∈U T )

[0056] av nm (p) = -2.025 - 0.175z p ,(z p ∈U NT )

[0057] The masking functions for tonal and non-tonal components are the same, as shown in the following equation.

[0058]

[0059] Among them, z qis the index of the masker, z p is the index of the masked object. ΔZ=z q -z p The first masking function (tone masking function) corresponding to the tonal component and the second masking function (non-tone masking function) corresponding to the non-tonal component can be calculated by the following formula.

[0060]

[0061] Comprehensive masking threshold T i,j (z q ) can be determined by the first masking function corresponding to the tonal component, the second masking function corresponding to the non-tonal component, and the silence threshold curve, as shown in the following formula.

[0062]

[0063] in is the silence threshold curve, calculated by the following formula:

[0064]

[0065] Among them, due to T i,j (z q ) q The value range of is also 0 to 31 (the same as k), so we can directly use the above formula z q Replace it with k, thus transforming it into T i,j (k). Then, for the current audio slice s i The average value of the masking threshold corresponding to the same frequency index in the sub-band frequency domain signal of each window is obtained to obtain the masking threshold curve T i (k), such as Figure 5 shown.

[0066] Figure 5 Schematic diagram of the masking threshold curve according to the first aspect of embodiment 1 of the present disclosure. Figure 5 The first masking thresholds corresponding to frequency indices 0 to 31 of different sub-bands are shown.

[0067] The computing device may play the current audio slice to the user, and simultaneously collect the user's electroencephalogram (EEG) signal to determine an EEG signal slice corresponding to the current audio slice ( S206 ).

[0068] In this manual, the user needs to wear a professional EEG signal acquisition device. When collecting the patient's EEG signal x in real time, play each audio slice at the same time, and slice the EEG signal in time intervals equal to the length of the audio slice. Assume that the current audio slice s i The corresponding EEG signal slice is x i.

[0069] Then, the computing device may enhance the energy of the gamma frequency band in the next audio slice of the current audio slice according to the spectrum characteristics of the gamma frequency band in the EEG signal slice and the masking threshold curve to obtain the enhanced next audio slice (S208).

[0070] Optionally, based on the spectral characteristics of the gamma frequency band in the EEG signal slice and the masking threshold curve, the operation of enhancing the energy of the gamma frequency band in the next audio slice of the current audio slice specifically includes: extracting the gamma frequency band characteristics of the EEG signal slice to determine the EEG signal frequency characteristics corresponding to the maximum amplitude of the EEG signal slice in the gamma frequency band; determining the second masking threshold corresponding to the EEG signal frequency characteristics in the masking threshold curve; and enhancing the energy of the gamma frequency band in the next audio slice of the current audio slice based on the EEG signal frequency characteristics and the second masking threshold.

[0071] That is, when playing the current audio slice s i When the current audio slice s i Corresponding EEG signal slice x i The spectrum characteristics of the mid-gamma band (specifically, the frequency with the largest amplitude in the gamma band) are used to determine the next audio slice s i+1 The energy in the gamma frequency band is enhanced, but in order to take into account the user's auditory experience, the required masking threshold (the second masking threshold mentioned above) is determined by the masking threshold curve, and the s i+1 The operation of enhancing the energy in the gamma frequency band is masked auditorily. i+1 When the user is not directly aware of the i+1 Perform energy-boosting maneuvers.

[0072] Among them, when extracting gamma frequency band features, fast Fourier transform (FFT) can be used for spectrum analysis. The fast Fourier transform formula is as follows:

[0073]

[0074] In this manual, the analysis of EEG signals focuses on the gamma frequency band (usually ranging from approximately 30Hz to 100Hz), and extracts key features of EEG signals in this frequency band, such as the average power, power spectral density, frequency peak of the gamma frequency band, and the fluctuation of the gamma frequency band energy at different times.

[0075] For EEG signal slice x i , we can determine the EEG signal slice x i The frequency corresponding to the maximum amplitude in the gamma frequency band is calculated as follows:

[0076] fm i =argmax(X i [f])

[0077] Using the EEG signal slice x obtained by the above calculation i The frequency fm corresponding to the maximum value in the gamma frequency band i Since the masking threshold in the masking threshold curve determined above corresponds to the frequency index k of each frequency domain sub-band signal, according to the calculated frequency fm i , determine the frequency index k of its corresponding frequency domain subband signal i (where k i =0~31), and T i (k i ) as the second masking threshold.

[0078] It can be mapped to the next audio slice s i+1 Energy intensity E in the spectrum i+1 (fm i ) is enhanced and limited to the second masking threshold T i (k i )Down.

[0079] The following is provided to slice the next audio i+1 Energy intensity E in the spectrum i+1 (fm i ) to enhance the specific method.

[0080] Optionally, based on the second masking threshold and the EEG signal frequency characteristics, the operation of enhancing the energy of the gamma frequency band in the next audio slice of the current audio slice specifically includes:

[0081] The audio signal amplitude of the audio signal point corresponding to the EEG signal frequency feature in the next audio slice is enhanced, and the audio signal amplitude is limited to below the second masking threshold.

[0082] Optionally, based on the second masking threshold and the EEG signal frequency characteristics, the operation of enhancing the energy of the gamma frequency band in the next audio slice of the current audio slice specifically includes:

[0083] An additional audio signal is inserted at a frequency point corresponding to the frequency feature of the EEG signal in the next audio slice, and the amplitude of the additional audio signal is limited to be below and close to the second masking threshold.

[0084] That is, in order to achieve the next audio slice s i+1 Energy intensity E in the spectrum i+1 (fm i ) is enhanced and limited to the second masking threshold Ti (k i ) below, you can use s i+1 Chinese FM i The audio signal amplitude of the corresponding audio signal point is enhanced, and the audio signal amplitude is limited to below the second masking threshold. For example, the next audio slice s can be directly i+1 mid-frequency FM i The audio signal amplitude of the audio signal point is adjusted to the second masking threshold T i (k i ).

[0085] However, in actual applications there may be a next audio slice s i+1 Chinese FM i The audio signal amplitude of the corresponding audio signal point is weak (even the audio signal amplitude is close to 0). In this case, the next audio slice s i+1 Chinese FM i A signal (ie, the additional audio signal) is inserted at the corresponding frequency point, and the amplitude of the additional audio signal is limited to be below and close to the second masking threshold.

[0086] If this method is used to perform signal energy enhancement, the next audio slice s can be determined first. i+1 Chinese FM i The signal amplitude of the corresponding frequency point, if it is determined that the signal amplitude is smaller than the preset amplitude (the preset amplitude can be a smaller amplitude preset artificially), it can be enhanced by inserting a signal (that is, inserting an additional audio signal at the frequency point corresponding to the frequency feature of the EEG signal in the next audio slice, and limiting the amplitude of the additional audio signal to below the second masking threshold and close to the second masking threshold).

[0087] It should be noted that, since the next audio slice needs to be adjusted with reference to the spectral characteristics of the previous audio slice, the above method can only adjust the audio slice starting from the second audio slice. Therefore, if the first audio slice needs to be adjusted, the frequency for which energy enhancement is required in this audio slice can be manually set. For example, fm0 = 40 Hz can be set to directly enhance the signal energy corresponding to the 40 Hz frequency in the first audio slice.

[0088] When the current audio slice is played, the adjustment of the next audio slice is completed, and the current audio slice is played, then the enhanced next audio slice can be played to the user in sequence. Among them, this method can be applied to AD audio treatment, and the user in this method can be a patient with AD.

[0089] In addition, reference Figure 1As shown, according to a second aspect of this embodiment, a storage medium is provided, wherein the storage medium includes a stored program, wherein when the program is run, a processor executes any one of the above methods.

[0090] Therefore, according to this embodiment, it is possible to take into account both the energy enhancement of the gamma frequency band of the audio slice and the user's auditory experience. Therefore, by applying this method in AD audio treatment, the treatment effect can be improved.

[0091] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0092] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0093] Example 2

[0094] Figure 6 FIG. 6 shows an audio processing device 600 according to the first aspect of this embodiment, which corresponds to the method according to the first aspect of embodiment 1. Figure 6 As shown, the apparatus 600 includes: an audio slicing module 610, configured to slice the original audio data into equal lengths to obtain a plurality of audio slices;

[0095] a masking threshold curve determining module 620, configured to determine a masking threshold curve corresponding to a current audio slice, the masking threshold curve being configured to indicate a first masking threshold corresponding to a frequency index of each masker;

[0096] The EEG signal acquisition module 630 is used to play the current audio slice to the user, collect the user's EEG signal at the same time, and determine the EEG signal slice corresponding to the current audio slice; and

[0097] The audio adjustment module 640 enhances the energy of the gamma frequency band in the next audio slice of the current audio slice according to the spectrum characteristics of the gamma frequency band in the EEG signal slice and the masking threshold curve to obtain the adjusted next audio slice.

[0098] Optionally, the masking threshold curve determination module 620 is specifically used to perform sub-band filtering on the time domain signal of the current audio slice to determine the frequency domain sub-band signal corresponding to the current audio slice; convert to a critical band rate based on the frequency domain sub-band signal; and determine a first masking threshold corresponding to the frequency index of each masker based on the frequency domain sub-band signal converted to the critical band rate to obtain a masking threshold curve.

[0099] Optionally, the masking threshold curve determination module 620 is specifically used to determine whether the frequency domain subband signal corresponding to the local maximum is a tonal component or a non-tonal component based on the amplitude difference between the local maximum and the adjacent points in the frequency domain subband signal converted to the critical band rate; determine a first masking function corresponding to the tonal component and a second masking function corresponding to the non-tonal component; and determine a first masking threshold corresponding to the frequency index of each masker based on the first masking function and the second masking function.

[0100] Optionally, the audio adjustment module 640 is specifically used to extract gamma frequency band features of the EEG signal slice, thereby determining the EEG signal frequency features corresponding to the maximum amplitude of the EEG signal slice in the gamma frequency band; determining a second masking threshold in the masking threshold curve corresponding to the EEG signal frequency features; and enhancing the energy of the gamma frequency band in the next audio slice of the current audio slice based on the EEG signal frequency features and the second masking threshold.

[0101] Optionally, the audio adjustment module 640 is specifically configured to enhance the audio signal amplitude of an audio signal point corresponding to the EEG signal frequency feature in the next audio slice, while limiting the audio signal amplitude to below a second masking threshold.

[0102] Optionally, the audio adjustment module 640 is specifically configured to insert an additional audio signal at a frequency point corresponding to the EEG signal frequency feature in the next audio slice, and limit the amplitude of the additional audio signal to be below and close to the second masking threshold.

[0103] Therefore, according to this embodiment, it is possible to take into account both the energy enhancement of the gamma frequency band of the audio slice and the user's auditory experience. Therefore, by applying this method in AD audio treatment, the treatment effect can be improved.

[0104] Example 3

[0105] Figure 7FIG. 7 shows an audio processing device 700 according to the first aspect of this embodiment, which corresponds to the method according to the first aspect of embodiment 1. Figure 7 As shown, the device 700 includes: a processor 710; and a memory 720, which is connected to the processor 710 and is used to provide the processor 710 with instructions for processing the following processing steps: slicing the original audio data with equal length to obtain a plurality of audio slices; determining a masking threshold curve corresponding to the current audio slice, the masking threshold curve being used to indicate a first masking threshold corresponding to the frequency index of each masker; playing the current audio slice to the user, and simultaneously collecting the user's EEG signal to determine the EEG signal slice corresponding to the current audio slice; and enhancing the energy of the gamma frequency band in the next audio slice of the current audio slice based on the spectral characteristics of the gamma frequency band in the EEG signal slice and the masking threshold curve to obtain the enhanced next audio slice.

[0106] Optionally, the operation of determining the masking threshold curve corresponding to the current audio slice specifically includes: performing sub-band filtering on the time domain signal of the current audio slice to determine the frequency domain sub-band signal corresponding to the current audio slice; converting the frequency domain sub-band signal to a critical band rate; and determining a first masking threshold corresponding to the frequency index of each masker based on the frequency domain sub-band signal converted to the critical band rate to obtain a masking threshold curve.

[0107] Optionally, the operation of determining the first masking threshold corresponding to the frequency index of each masker based on the frequency domain subband signal converted to the critical band rate includes: judging whether the frequency domain subband signal corresponding to the local maximum is a tonal component or a non-tonal component based on the amplitude difference between the local maximum and the adjacent points in the frequency domain subband signal converted to the critical band rate; determining a first masking function corresponding to the tonal component and a second masking function corresponding to the non-tonal component; and determining the first masking threshold corresponding to the frequency index of each masker based on the first masking function and the second masking function.

[0108] Optionally, the operation of adjusting the energy of the gamma frequency band in the next audio slice of the current audio slice according to the spectral characteristics of the gamma frequency band in the EEG signal slice and the masking threshold curve specifically includes: extracting the gamma frequency band characteristics of the EEG signal slice to determine the EEG signal frequency characteristics corresponding to the maximum amplitude of the EEG signal slice in the gamma frequency band; determining the second masking threshold corresponding to the EEG signal frequency characteristics in the masking threshold curve; and enhancing the energy of the gamma frequency band in the next audio slice of the current audio slice according to the EEG signal frequency characteristics and the second masking threshold.

[0109] Optionally, based on the second masking threshold and the EEG signal frequency characteristics, the operation of enhancing the energy of the gamma frequency band in the next audio slice of the current audio slice specifically includes:

[0110] The audio signal amplitude of the audio signal point corresponding to the EEG signal frequency feature in the next audio slice is enhanced, and the audio signal amplitude is limited to below the second masking threshold.

[0111] Optionally, based on the second masking threshold and the EEG signal frequency characteristics, the operation of enhancing the energy of the gamma frequency band in the next audio slice of the current audio slice specifically includes:

[0112] An additional audio signal is inserted at a frequency point corresponding to the frequency feature of the EEG signal in the next audio slice, and the amplitude of the additional audio signal is limited to be below and close to the second masking threshold.

[0113] Therefore, according to this embodiment, it is possible to take into account both the energy enhancement of the gamma frequency band of the audio slice and the user's auditory experience. Therefore, by applying this method in AD audio treatment, the treatment effect can be improved.

[0114] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0115] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0116] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0117] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0118] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0119] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.

[0120] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. An audio processing method, characterized in that: include: Slice the original audio data into equal length slices to obtain a number of audio slices; Determine a masking threshold curve corresponding to the current audio slice, where the masking threshold curve is used to indicate a first masking threshold corresponding to a frequency index of each masker; Playing the current audio slice to the user, while collecting the user's EEG signal, and determining the EEG signal slice corresponding to the current audio slice; as well as According to the spectrum characteristics of the gamma frequency band in the EEG signal slice and the masking threshold curve, the energy of the gamma frequency band in the next audio slice of the current audio slice is enhanced to obtain the enhanced next audio slice.

2. The method according to claim 1, characterized in that The operation of determining the masking threshold curve corresponding to the current audio slice specifically includes: Performing sub-band filtering on the time domain signal of the current audio slice to determine a frequency domain sub-band signal corresponding to the current audio slice; Converting the frequency domain subband signal to a critical band rate; According to the frequency domain subband signal converted to the critical band rate, a first masking threshold corresponding to the frequency index of each masker is determined to obtain the masking threshold curve.

3. The method according to claim 2, characterized in that The operation of determining a first masking threshold corresponding to a frequency index of each masker according to the frequency domain subband signal converted to the critical band rate includes: Determining whether the frequency domain subband signal corresponding to the local maximum value is a tonal component or a non-tonal component based on an amplitude difference between a local maximum value and an adjacent point in the frequency domain subband signal converted to the critical band rate; determining a first masking function corresponding to the tonal component and a second masking function corresponding to the non-tonal component; and A first masking threshold corresponding to a frequency index of each masker is determined according to the first masking function and the second masking function.

4. The method according to claim 1, wherein The operation of enhancing the energy of the gamma frequency band in the next audio slice of the current audio slice according to the spectrum characteristics of the gamma frequency band in the EEG signal slice and the masking threshold curve specifically includes: Performing gamma frequency band feature extraction on the EEG signal slice to determine the EEG signal frequency feature corresponding to the maximum amplitude of the EEG signal slice in the gamma frequency band; Determining a second masking threshold in the masking threshold curve corresponding to the EEG signal frequency feature; and The energy of the gamma frequency band in the next audio slice of the current audio slice is enhanced according to the EEG signal frequency feature and the second masking threshold.

5. The method according to claim 4, characterized in that The operation of enhancing the energy of the gamma frequency band in the next audio slice of the current audio slice according to the second masking threshold and the EEG signal frequency feature specifically includes: The audio signal amplitude of the audio signal point corresponding to the EEG signal frequency feature in the next audio slice is enhanced, and the audio signal amplitude is limited to below the second masking threshold.

6. The method according to claim 4, characterized in that The operation of enhancing the energy of the gamma frequency band in the next audio slice of the current audio slice according to the second masking threshold and the EEG signal frequency feature specifically includes: An additional audio signal is inserted into the next audio slice at a frequency point corresponding to the frequency feature of the electroencephalogram signal, and an amplitude of the additional audio signal is limited to be below and close to the second masking threshold.

7. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is run, the processor executes the method according to any one of claims 1 to 6.

8. An audio processing device, characterized in that: include: The audio slicing module is used to slice the original audio data into equal length slices to obtain several audio slices; a masking threshold curve determining module, configured to determine a masking threshold curve corresponding to a current audio slice, wherein the masking threshold curve is configured to indicate a first masking threshold corresponding to a frequency index of each masker; An EEG signal acquisition module is used to play the current audio slice to the user, collect the user's EEG signal at the same time, and determine the EEG signal slice corresponding to the current audio slice; as well as The audio adjustment module enhances the energy of the gamma frequency band in the next audio slice of the current audio slice according to the spectrum characteristics of the gamma frequency band in the EEG signal slice and the masking threshold curve to obtain the enhanced next audio slice.

9. The device according to claim 8, characterized in that The masking threshold curve determination module is specifically used to perform sub-band filtering on the time domain signal of the current audio slice to determine the frequency domain sub-band signal corresponding to the current audio slice; convert the frequency domain sub-band signal to a critical band rate according to the frequency domain sub-band signal; and determine the first masking threshold corresponding to the frequency index of each masker according to the frequency domain sub-band signal converted to the critical band rate to obtain the masking threshold curve.

10. An audio processing device, characterized in that: include: processor; as well as A memory, connected to the processor, configured to provide the processor with instructions for processing the following processing steps: Slice the original audio data into equal length slices to obtain a number of audio slices; Determine a masking threshold curve corresponding to the current audio slice, where the masking threshold curve is used to indicate a first masking threshold corresponding to a frequency index of each masker; Playing the current audio slice to the user, while collecting the user's EEG signal, and determining the EEG signal slice corresponding to the current audio slice; as well as According to the spectrum characteristics of the gamma frequency band in the EEG signal slice and the masking threshold curve, the energy of the gamma frequency band in the next audio slice of the current audio slice is enhanced to obtain the enhanced next audio slice.