Audio feature extraction method and device and electronic equipment

Through the DMA controller, the audio data and window function data are directly processed during the audio feature extraction process, which solves the problem of excessive CPU computing burden in traditional methods, improves the speed and efficiency of audio processing, and reduces system power consumption and cost.

CN120071960APending Publication Date: 2025-05-30ZHUHAI HUGE IC CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510066808.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional audio feature extraction methods require a large amount of calculation and judgment by the CPU, which consumes too much computing power, limits the speed and efficiency of audio processing, and increases the power consumption and cost of the system.

Method used

By using a DMA controller to read the audio data and window function data directly from the data memory and perform necessary processing, such as multiplying the window function without CPU intervention. The solution includes a data memory, a DMA controller, a multiplier, an index maintenance unit, a ROM, a frame data assembly unit and a feature extraction unit. Through the coordinated work of these components, efficient extraction of audio features is achieved.

Benefits of technology

It significantly reduces the burden on the CPU in the audio feature extraction process, improves the speed and efficiency of audio processing, reduces the overall power consumption of the system, and may reduce the required hardware resources, thereby reducing the system cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071960A_ABST
    Figure CN120071960A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an audio feature extraction method and device and electronic equipment, and relates to the field of audio processing. According to the invention, the DMA controller directly reads the audio and window function data from the data memory and carries out necessary processing such as multiplication with the window function, CPU intervention is not needed, the burden of the CPU is reduced, and the audio feature extraction speed and efficiency are improved. The DMA controller quickly accesses the memory, data reading and primary processing are completed faster than a CPU, system power consumption and hardware resource requirements are reduced, and cost reduction is facilitated. According to the technical scheme, the original audio sequence is divided into the multiple overlapped frame time periods in advance, the index, the data sampling point address and the window function sampling point address are set for each frame, accurate control over audio data processing and feature extraction is achieved, and the accuracy and reliability of audio feature extraction are improved. In this way, the CPU can focus on other tasks, the DMA controller focuses on rapid processing of the audio data, and the performance and efficiency of the whole system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing, and particularly to a method, apparatus, and electronic device for extracting audio features. Background Art

[0002] In the existing audio processing technology, the extraction of audio features is a crucial link, which is of great significance for various applications such as audio recognition, audio classification, and audio compression. However, there is a significant problem with traditional audio feature extraction methods, that is, the computing power consumption of the CPU is too large.

[0003] Specifically, in the traditional audio feature extraction process, the CPU first needs to read the complete original audio sequence from the memory. This step itself requires a certain amount of computing resources and time, especially for long-duration audio files, and its reading process is even more time-consuming and laborious. Immediately afterwards, the CPU needs to perform frame segmentation on the read original audio sequence. Frame segmentation is to cut the continuous audio signal into multiple shorter and relatively independent processing units for subsequent audio analysis. Although this step seems simple, in fact, the CPU needs to perform complex calculations and judgments to determine the start position and length of each frame. After the frame segmentation is completed, the CPU also needs to perform windowing on each frame. The purpose of windowing is to reduce the boundary effect between frames, so that the audio signal can smoothly transition between frames. However, windowing also requires the CPU to perform a large number of mathematical operations to calculate the shape and parameters of each window function. Finally, after the frame segmentation and windowing are completed, the CPU can perform the extraction of audio features. This step is the most complex and time-consuming part of the entire audio processing process, and it requires the CPU to deeply analyze and calculate the processed audio data to extract useful audio features.

[0004] In summary, due to the need for the CPU to perform a large number of calculations and judgments, the traditional audio feature extraction method consumes too much computing power. This not only limits the speed and efficiency of audio processing, but also increases the power consumption and cost of the system. Therefore, how to reduce the burden on the CPU when calculating audio features and improve the speed and efficiency of audio processing has become an urgent problem in the current audio processing technology. Summary of the Invention

[0005] Embodiments of this application provide a method, apparatus, and electronic device for extracting audio features, which can solve the problem of the equalization effect being affected by quantization errors during the switching of EQ equalization processing in the prior art. The technical solutions are as follows:

[0006] In a first aspect, an embodiment of the present application provides an apparatus for extracting audio features, including: a data memory, a DMA controller, a multiplier, an index maintenance unit, a ROM, a frame data assembly unit, a feature extraction unit, and a register; the DMA controller is respectively connected to the data memory, the ROM, and the index maintenance unit, two output terminals of the DMA controller are respectively connected to the multiplier, the multiplier is connected to the frame data assembly unit, the frame data assembly unit is respectively connected to the index maintenance unit and the feature extraction unit, and the feature extraction unit is connected to the register;

[0007] Among them, the DMA controller is configured to read the current frame index, the current audio data sampling point address, and the current window function sampling point address from the index maintenance unit, determine the sampling point set corresponding to the current frame time period in the original audio sequence stored in the data memory according to the current frame index, then read the current audio data sampling point from the sampling point set according to the current audio data sampling point address, read the corresponding current window function sampling point from the ROM according to the current window function sampling point address, and send the current audio data sampling point and the current window function sampling point to the multiplier; wherein, the original audio sequence is pre-divided into multiple overlapping frame time periods, each frame time period is set with a frame index, each frame time period includes multiple audio data sampling points, and each audio data sampling point is set with an audio data sampling point address;

[0008] The index maintenance unit is configured to update the current frame index, the current audio data sampling point address, and the current window function sampling point address;

[0009] The multiplier is configured to multiply the current audio data sampling point and the current window function sampling point to obtain a multiplication result, and send the multiplication result to the frame data assembly unit;

[0010] The frame data assembly unit is configured to assemble the multiplication results corresponding to all audio data sampling points within each frame time period into an audio data frame, and send the assembled audio data frame to the audio feature extraction unit;

[0011] The audio feature extraction unit is configured to extract the audio feature information of the audio data frame and write the audio feature information into the register.

[0012] In a second aspect, an embodiment of the present application provides a method for extracting audio features, including:

[0013] The DMA controller reads the current frame index, the current audio data sampling point address, and the current window function sampling point address from the index maintenance unit, determines the set of sampling points corresponding to the current frame time period in the original audio sequence stored in the data memory according to the current frame index, then reads the current audio data sampling point from the set of sampling points according to the current audio data sampling point address, reads the corresponding current window function sampling point from the ROM according to the current window function sampling point address, and sends the current audio data sampling point and the current window function sampling point to the multiplier; wherein, the original audio sequence is pre-divided into multiple overlapping frame time periods, each frame time period is set with a frame index, each frame time period includes multiple audio data sampling points, and each audio data sampling point is set with an audio data sampling point address;

[0014] The index maintenance unit updates the current frame index, the current audio data sampling point address, and the current window function sampling point address;

[0015] The multiplier multiplies the current audio data sampling point and the current window function sampling point to obtain a multiplication result, and sends the multiplication result to the frame data assembly unit;

[0016] The frame data assembly unit assembles the multiplication results corresponding to all audio data sampling points within each frame time period into an audio data frame, and sends the assembled audio data frame to the audio feature extraction unit;

[0017] The audio feature extraction unit extracts the audio feature information of the audio data frame and writes the audio feature information into the register.

[0018] In a third aspect, an embodiment of the present application provides an electronic device, which may be a device with audio processing functions such as a tablet computer, a personal computer, a mobile phone, etc. It may include: a processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the above method steps.

[0019] The beneficial effects brought by the technical solutions provided by some embodiments of the present application at least include:

[0020] The DMA controller directly reads the audio data and window function data from the data memory and performs necessary processing (such as multiplying with the window function) without the intervention of the CPU. In this way, the CPU can focus on other tasks, thus significantly reducing the burden on the CPU during the audio feature extraction process. Since the DMA controller can access the memory quickly and efficiently, compared with the CPU, it can complete the data reading and preliminary processing more rapidly. This leads to a significant improvement in the speed and efficiency of the entire audio feature extraction process. By reducing the use of the CPU, the overall power consumption of the system is reduced. In addition, due to the improvement in processing speed, the required hardware resources (such as the CPU and memory) may also be reduced accordingly, which helps to reduce the cost of the system. By pre-dividing the original audio sequence into multiple overlapping frame time periods and setting the frame index, audio data sampling point address, and window function sampling point address for each frame time period, this technical solution can more precisely control the processing and feature extraction of audio data. This helps to improve the accuracy and reliability of audio feature extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 is a schematic structural diagram of an audio feature extraction device provided by an embodiment of the present application;

[0023] Figure 2 is a schematic flowchart of an audio feature extraction method provided by an embodiment of the present application;

[0024] Figure 3 is a schematic diagram of the principle of frame division provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the drawings.

[0026] It should be noted that the audio feature extraction method provided by the present application is generally executed by an audio feature extraction device. Correspondingly, the audio feature extraction device is generally set in the audio feature extraction device.

[0027] Figure 1 shows an exemplary system architecture to which the audio feature extraction method or audio feature extraction device of the present application can be applied.

[0028] Such asFigure 1 As shown, it may include: a data memory, a DMA controller, a multiplier, an index maintenance unit, a ROM, a frame data assembly unit, a feature extraction unit, and registers.

[0029] Among them, the data memory is connected to the DMA controller, the ROM (read only memory) is connected to the DMA controller, the DMA controller is connected to the index maintenance unit. The first output terminal and the second output terminal of the DMA controller are respectively used to output audio data sampling points and window function sampling points. The first output terminal and the second output terminal are connected to the multiplier. The output terminal of the multiplier is connected to the frame data assembly unit. The frame data assembly unit is connected to the index maintenance unit. The output terminal of the frame data assembly unit is connected to the feature extraction unit. The output terminal of the feature extraction unit is connected to the registers.

[0030] The data memory is used to store the original audio sequence. The original audio sequence is composed of multiple audio data sampling points. The original audio sequence is pre-divided into multiple overlapping frame time periods. Each frame time period is set with a frame index. Each frame time period includes multiple audio data sampling points. Each audio data sampling point is set with an audio data sampling point address.

[0031] It should be understood that Figure 1 the number of each unit or module in [[ ]] is only illustrative. According to the implementation requirements, it can be any number.

[0032] Please refer to Figure 2 , which provides a schematic flow diagram of a method for extracting audio features in an embodiment of the present application. As Figure 2 shown, the method in the embodiment of the present application may include the following steps:

[0033] S201. The DMA controller reads the current frame index, the current audio data sampling point address, and the current window function sampling point address from the index maintenance unit. Determines the sampling point set corresponding to the current frame time period in the original audio sequence stored in the data memory according to the current frame index. Then reads the current audio data sampling point from the sampling point set according to the current audio data sampling point address, and reads the corresponding current window function sampling point from the ROM according to the current window function sampling point address, and sends the current audio data sampling point and the current window function sampling point to the multiplier.

[0034] Among them, the original audio sequence is pre-divided into multiple overlapping frame time periods. Each frame time period is set with a frame index. Each frame time period includes multiple audio data sampling points. Each audio data sampling point is set with an audio data sampling point address. The overlapping part of two adjacent frame time periods, that is, the frame shift, is 1 / 3 - 1 / 2 of the frame time period.

[0035] For example, refer toFigure 3 The framed schematic diagram shown, the input original audio sequence input data is divided into 3 frame time periods, the length of the frame time period is winlen, the length of the frame shift frame_shift is winlen / 2, and winlen also represents the duration of the window function.

[0036] Furthermore, the format of the original audio sequence is PCM. PCM audio data is usually stored in units of samples, and each sample represents the amplitude of the audio signal at a certain moment. The resolution of the sample (i.e., the number of bits per sample) can be 8 bits, 16 bits, 24 bits, or 32 bits, etc., depending on the precision requirements of the audio data. Samples are usually collected at a fixed sampling rate, such as 44.1 kHz, 48 kHz, 96 kHz, etc. The higher the sampling rate, the better the audio quality, but the larger the data volume.

[0037] The DMA controller accesses the index maintenance unit to read the current frame index (frame_index), the current audio data sample point address (audio_sample_address), and the current window function sample point address (window_sample_address). According to the read frame_index, the DMA controller accesses the data memory (the data memory can be RAM), in which the original audio sequence is stored. The sequence has been pre-divided into multiple overlapping frame time periods, and each frame time period is identified by the frame index. Determine the start and end positions of the current frame time period in the original audio sequence, so as to extract all the audio data sample point sets within this time period. Using audio_sample_address, the DMA controller reads the current audio data sample (audio_sample) from the sample point set extracted in step 2. Using window_sample_address, the DMA controller reads the corresponding current window function sample (window_sample) from the ROM (read-only memory). The window function is usually used for weighting or smoothing operations in audio processing. The DMA controller sends the read audio_sample and window_sample to the multiplier to prepare for multiplication.

[0038] Furthermore, the window function is a predefined array or function, and its length (i.e., the number of window function sample points) is fixed. The length of the window function is designed to be greater than or equal to the number of audio data sample points within each frame time period.

[0039] Audio data frame:

[0040] The original audio sequence is divided into multiple overlapping frame time periods, and each frame time period contains a certain number of audio data sampling points. Each frame time period is convolved (or multiplied) with a window function to smooth the edges of the data frame and reduce spectral leakage. When the window function length is equal to the audio data frame length, each sampling point of the window function is multiplied by the corresponding sampling point of the audio data frame. When the window function length is greater than the audio data frame length, there are several processing methods:

[0041] a. Truncate the window function: Only use the first N sampling points of the window function (N is equal to the audio data frame length) to multiply with the audio data frame. This is the most common and simple method.

[0042] b. Zero-pad the audio data frame: Add zero-valued sampling points at the end of the audio data frame to make its length match that of the window function. This method may cause edge effects because the zero-valued sampling points do not contribute effective signal energy.

[0043] c. Use partially overlapping window functions: Although not common, theoretically, a window function application strategy can be designed where the window functions partially overlap between adjacent frames, and the total length of the window function is greater than the length of any single audio data frame. This method may require more complex implementation and additional computing resources.

[0044] S202. The index maintenance unit updates the current frame index, the current audio data sampling point address, and the current window function sampling point address.

[0045] Among them, after completing the operation of S301, the index maintenance unit updates frame_index, audio_sample_address, and window_sample_address according to preset rules.

[0046] The specific update process is as described below.

[0047] 1. Update of the current frame index: At the end of each audio processing cycle, the index maintenance unit calculates the next frame index according to the preset frame size and overlap amount.

[0048] The frame size is the number of sampling points contained in each audio frame, and it is usually determined according to the requirements of audio processing.

[0049] The overlap amount is the number of overlapping sampling points between two adjacent frames, and it is used to smooth the transition between frames and reduce boundary effects.

[0050] The new frame index (new_frame_index) can be calculated based on the current frame index (current_frame_index), frame size (frame_size), and overlap amount (overlap): new_frame_index = (current_frame_index + frame_size - overlap) (Note: The calculation method needs to be adjusted according to the actual situation to ensure the correctness of the index).

[0051] If the new_frame_index exceeds the total number of frames of the original audio sequence, it may be necessary to loop back to the starting frame or perform other appropriate processing.

[0052] 2. Update of the current audio data sample point address: According to the new frame index, the index maintenance unit calculates the new audio data sample point address (new_audio_sample_address).

[0053] This usually involves multiplying the new frame index by the number of bytes occupied by each sample point (or calculating according to the specific storage format).

[0054] The new_audio_sample_address will point to the first audio data sample point in the next frame time period.

[0055] 3. Update of the current window function sample point address: The update of the window function sample point address may be relatively simple because the window function is usually fixed or changes according to a predetermined pattern.

[0056] If the window function is fixed, it may simply repeat the previous address each time it is updated (especially when the window function size matches the frame size).

[0057] If the window function is variable (e.g., changes with the frame index), the index maintenance unit needs to calculate the new window function sample point address (new_window_sample_address) according to the definition and change rule of the window function.

[0058] Finally, the index maintenance unit writes the calculated new_frame_index, new_audio_sample_address, and new_window_sample_address back to the corresponding registers or memory for use in the next audio processing cycle.

[0059] S203. The multiplier multiplies the current audio data sample point and the current window function sample point to obtain a multiplication result, and sends the multiplication result to the frame data assembly unit.

[0060] Among them, the multiplier receives audio_sample and window_sample from the DMA controller. It performs a multiplication operation to calculate the product of audio_sample and window_sample, that is, the multiplied_result, and sends the multiplied_result to the frame data assembly unit.

[0061] S204. The frame data assembly unit assembles the multiplied results corresponding to all audio data sampling points within each frame time period into an audio data frame, and sends the assembled audio data frame to the audio feature extraction unit.

[0062] Among them, the frame data assembly unit receives each multiplied_result from the multiplier. Within a frame time period, for all audio data sampling points, the frame data assembly unit will receive the corresponding multiplied_result. The frame data assembly unit assembles these multiplied_results in sequence into a complete audio data frame (audio_frame), and sends the assembled audio_frame to the audio feature extraction unit.

[0063] S205. The feature extraction unit extracts the audio feature information of the audio data frame and writes the audio feature information into the register.

[0064] Among them, the audio feature extraction unit receives audio_frame from the frame data assembly unit. It analyzes the audio_frame using a predefined algorithm or model to extract the audio feature information. The extracted audio feature information is written into the register for subsequent processing or analysis.

[0065] Furthermore, the audio feature information includes one or more of: energy, zero-crossing rate, pitch period, and formant.

[0066] Among them, energy is a feature related to the amplitude of the audio signal, which reflects the intensity of the signal in the time domain. The energy feature has extensive applications in voice activity detection (VAD). For example, by comparing the energy of the signal to determine whether it is speech or noise. In addition, the energy feature also contains rich emotional information. For example, the energy of a person's speech is usually lower when they are sad.

[0067] The zero-crossing rate refers to the number of times the audio signal crosses zero within a unit time, that is, the frequency at which the signal changes from positive to negative or from negative to positive. The zero-crossing rate is widely used in the fields of speech recognition and music information retrieval. It has higher value for high-impact sounds such as metal and rock. Generally, the higher the zero-crossing rate, the higher the frequency.

[0068] The pitch period is the time for one complete wavelength of a mechanical wave, and it is one of the most important features in voiced signals. When pronouncing a voiced sound, the air flow passes through the glottis to generate a quasi-periodic pulsed air flow, which excites the vocal tract to produce a voiced sound, carrying most of the mechanical wave energy in speech. The wavelength of this mechanical wave is called the fundamental wave, and the corresponding period is called the pitch period. The extraction of the pitch period has wide applications in speech compression coding, speech analysis and synthesis, and speech recognition. Accurately estimating and extracting the pitch period is crucial for speech signal processing, which directly affects the authenticity of synthesized speech, the recognition rate of speech recognition, and the correct rate of speech compression coding.

[0069] Formants are the integer multiples of the fundamental frequency corresponding frequency components, used to reflect the physical characteristics of a person's vocal tract. When a quasi-periodic pulse excitation enters the vocal tract, it will drive the air to resonate, generating a set of resonance frequencies, called formant frequencies or simply formants for short. Formant features include formant frequencies and bandwidths, etc. Formant information is contained in the speech spectrum envelope and has wide applications in speech signal synthesis and speech recognition. The envelope, position, and adjacent distances of formants together form the timbre characteristics. For example, when the distance between formants is close, it sounds thick and rough, and when the distance is far, it sounds clear.

[0070] In some embodiments of the present application, a plurality of sets of window function sampling points of different types are stored in the ROM, and the DMA controller selects the corresponding set of window function sampling points according to the application scenario of the original audio sequence.

[0071] Window functions are mainly used in audio signal processing to control the boundary effects of signals, so as to reduce spectral leakage and improve frequency resolution. Common window functions include rectangular window, Hanning window, Hamming window, Blackman window, and Kaiser window, etc. Each window function has its unique spectral characteristics and time-domain characteristics. When the DMA controller selects a window function, the following factors are usually considered:

[0072] Application scenario: Different application scenarios have different requirements for window functions. For example, in applications that require high-resolution spectral analysis, a window function with a narrower main lobe and lower side lobes (such as the Kaiser window) may be selected. While in applications that require fast processing, a window function with simpler calculations (such as the rectangular window or Hanning window) may be selected.

[0073] Signal characteristics: The characteristics of the original audio sequence also affect the selection of the window function. For example, if the signal contains many high-frequency components, a window function with better high-frequency suppression ability may be required.

[0074] Computational overhead: Different window functions have different computational overheads. When selecting a window function, it is necessary to balance the computational overhead and performance requirements.

[0075] Multiple sets of window function sampling points of different types are pre-stored in the ROM. The DMA controller selects a suitable set of window function sampling points from the ROM according to the application scenario and signal characteristics. The DMA controller reads the sampling points of the selected window function and sends them to the multiplier together with the corresponding audio data sampling points for multiplication. The result of the multiplication is sent to the frame data assembly unit for assembly, and finally the audio feature extraction unit extracts the feature information.

[0076] In some embodiments of the present application, after the DMA controller finishes processing all the sampling points within the current frame time period corresponding to the current frame index, it sends an interrupt instruction to the CPU, and the CPU reads the audio feature information of the audio data frame indicated by the current frame index in the register.

[0077] Among them, after the DMA controller finishes the processing task of all the sampling points within the current frame time period corresponding to the current frame index, it will enter a specific interrupt preparation stage. In this stage, the DMA controller will sort out and confirm the processed audio data frames and their related information to ensure the integrity and accuracy of the data.

[0078] The DMA controller sends an interrupt instruction to the CPU through an internal interrupt mechanism. This interrupt instruction is a signal that tells the CPU that the current DMA controller has completed the specified task. The interrupt instruction usually contains information about the processed data frame, such as the frame index, the length of the data frame, or the storage location, etc., so that the CPU can accurately locate and process this data.

[0079] When the CPU receives the interrupt instruction from the DMA controller, it will temporarily interrupt the currently executing task and jump to the Interrupt Service Routine (ISR) for processing. The interrupt service program is a piece of code written in advance for responding to and processing the interrupt request from the DMA controller.

[0080] In the interrupt service program, the CPU will locate the audio feature information of the audio data frame indicated by the current frame index stored in the register according to the information provided in the interrupt instruction. The CPU reads this audio feature information and performs necessary processing or storage operations. These audio feature information may include spectral features, energy features, fundamental frequencies, etc., depending on the requirements of the application scenario.

[0081] After finishing reading and processing the audio feature information, the CPU will exit the interrupt service program and resume the previously interrupted task to continue execution. This can ensure the continuity and stability of the system, and at the same time avoid affecting the execution of other tasks due to long-term processing of interrupt requests.

[0082] There may be multiple interrupt sources in the system. Therefore, it is necessary to reasonably set the interrupt priorities to ensure that the interrupt requests of the DMA controller can be promptly responded to. When processing interrupt requests, it is necessary to ensure the synchronization and integrity of the data. For example, a lock mechanism or semaphore can be used to avoid data competition and access conflicts. The interrupt handling code should be as concise and efficient as possible to reduce the CPU resource occupancy and latency.

[0083] In summary, through the interrupt interaction between the DMA controller and the CPU, real-time processing and feature extraction of audio data can be achieved. This mechanism not only improves the processing efficiency of the system but also ensures the accuracy and real-time nature of the data.

[0084] The embodiments of the present application include the following beneficial effects:

[0085] The DMA controller directly reads the audio data and window function data from the data memory and performs necessary processing (such as multiplying with the window function) without the intervention of the CPU. In this way, the CPU can focus on other tasks, thus significantly reducing the burden on the CPU during the audio feature extraction process. Since the DMA controller can access the memory quickly and efficiently, compared with the CPU, it can complete the data reading and preliminary processing more rapidly. This results in a significant improvement in the speed and efficiency of the entire audio feature extraction process. By reducing the use of the CPU, the overall power consumption of the system is reduced. In addition, due to the improvement in processing speed, the required hardware resources (such as the CPU and memory) may also be reduced accordingly, which helps to reduce the cost of the system. By pre-dividing the original audio sequence into multiple overlapping frame time periods and setting the frame index, audio data sampling point address, and window function sampling point address for each frame time period, this technical solution can more precisely control the processing and feature extraction of audio data. This helps to improve the accuracy and reliability of audio feature extraction.

[0086] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.

[0087] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. An audio feature extraction device, characterized in that: include: Data memory, DMA controller, multiplier, index maintenance unit, ROM, frame data assembly unit, feature extraction unit, register; The DMA controller is connected to the data storage, the ROM, and the index maintenance unit respectively, the two output ends of the DMA controller are connected to the multiplier respectively, the multiplier is connected to the frame data assembly unit, the frame data assembly unit is connected to the index maintenance unit and the feature extraction unit respectively, and the feature extraction unit is connected to the register; Wherein, the DMA controller is used to read the current frame index, the current audio data sampling point address and the current window function sampling point address from the index maintenance unit, determine the sampling point set corresponding to the current frame time period in the original audio sequence stored in the data storage device according to the current frame index, then read the current audio data sampling point in the sampling point set according to the current audio data sampling point address, read the corresponding current window function sampling point in the ROM according to the current window function sampling point address, and send the current audio data sampling point and the current window function sampling point to the multiplier; wherein the original audio sequence is pre-divided into a plurality of overlapping frame time periods, each frame time period is provided with a frame index, each frame time period includes a plurality of audio data sampling points, and each audio data sampling point is provided with an audio data sampling point address; An index maintenance unit, used to update the current frame index, the current audio data sampling point address and the current window function sampling point address; A multiplier, used for multiplying the current audio data sampling point and the current window function sampling point to obtain a multiplication result, and sending the multiplication result to the frame data assembly unit; a frame data assembling unit, configured to assemble the multiplication results corresponding to all audio data sampling points in each frame time period into an audio data frame, and send the assembled audio data frame to the audio feature extraction unit; The audio feature extraction unit is used to extract audio feature information of the audio data frame and write the audio feature information into the register.

2. The device according to claim 1, characterized in that The number of window function sampling points included in the window function is greater than or equal to the number of audio data sampling points in the frame time period.

3. The device according to claim 2, characterized in that The overlap amount between two adjacent frame periods is 1 / 2 of the frame period.

4. The device according to claim 1, 2 or 3, characterized in that: The ROM stores a plurality of different types of window function sampling point sets, and the DMA controller selects a corresponding window function sampling point set according to the application scenario of the original audio sequence.

5. The device according to claim 4, characterized in that After the DMA controller completes processing of all sampling points in the current frame time period corresponding to the current frame index, it sends an interrupt instruction to the CPU, and the CPU reads the audio feature information of the audio data frame indicated by the current frame index in the register.

6. The device according to claim 5, characterized in that The audio feature information includes: one or more of energy, zero-crossing rate, pitch period and resonance peak.

7. The device according to claim 4 or 5, characterized in that The format of the original audio sequence is PCM.

8. A method for extracting audio features, characterized in that: include: The DMA controller reads the current frame index, the current audio data sampling point address and the current window function sampling point address from the index maintenance unit, determines the sampling point set corresponding to the current frame time period in the original audio sequence stored in the data storage device according to the current frame index, then reads the current audio data sampling point in the sampling point set according to the current audio data sampling point address, reads the corresponding current window function sampling point in the ROM according to the current window function sampling point address, and sends the current audio data sampling point and the current window function sampling point to the multiplier; wherein the original audio sequence is pre-divided into a plurality of overlapping frame time periods, each frame time period is provided with a frame index, each frame time period includes a plurality of audio data sampling points, and each audio data sampling point is provided with an audio data sampling point address; The index maintenance unit updates the current frame index, the current audio data sampling point address and the current window function sampling point address; The multiplier multiplies the current audio data sampling point and the current window function sampling point to obtain a multiplication result, and sends the multiplication result to the frame data assembly unit; The frame data assembly unit assembles the multiplication results corresponding to all audio data sampling points in each frame time period into an audio data frame, and sends the assembled audio data frame to the audio feature extraction unit; The audio feature extraction unit extracts audio feature information of the audio data frame and writes the audio feature information into a register.

9. A computer storage medium, characterized in that The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method steps according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps as claimed in any one of claims 1 to 7.