An ai audio processor, method and storage medium thereof

By employing adaptive repair techniques in frequency domain mapping, trajectory generation, and spectrum reconstruction modules, the problem of spectrum drift in AI audio processors under complex environments has been solved, achieving fine audio recovery and speech feature restoration under dynamic interference.

CN122493876APending Publication Date: 2026-07-31GUANGDONG CHANGE INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG CHANGE INTELLIGENT TECH CO LTD
Filing Date
2026-05-08
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing AI audio processors struggle to dynamically track resonant acoustic features in complex environments, leading to abnormal spectral drift, disrupting the resonant structure within the target speech segment, and failing to meet the requirements for fine audio recovery under dynamic interference.

Method used

The frequency domain mapping module filters the candidate set of formants, the trajectory generation module generates a stable trajectory set, the anomaly correction module corrects the frequency deviation, the spectrum reconstruction module merges the trajectory and envelope estimation, the audio generation module performs inverse short-time Fourier transform to generate the speaker driving waveform, and combines multi-frequency energy extremum capture and inter-frame difference sign product filtering to adaptively repair spectrum drift.

Benefits of technology

It deeply suppresses redundant interference noise in complex environments, improves the fidelity of speech feature restoration, preserves dynamic acoustic structure, and enhances output clarity and realism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493876A_ABST
    Figure CN122493876A_ABST
Patent Text Reader

Abstract

This invention relates to the field of speech enhancement technology, specifically to an AI audio processor, method, and storage medium. The AI ​​audio processor includes: a frequency domain mapping module, a trajectory generation module, an anomaly correction module, a spectrum reconstruction module, and an audio generation module. The beneficial effects of this invention are as follows: by capturing energy extrema at multiple frequency points to establish a resonance candidate set, selecting stable trajectories based on the sign product of the differences between adjacent frames, accurately locking the core frequency band and investigating spectral drift, performing deviation determination and extremum re-search by calculating the difference in the energy distribution corresponding to the preceding time frame, completing adaptive repair of abnormal frequencies, merging the corrected frequency points and stable trajectories for envelope estimation and difference segment reconstruction, applying fixed attenuation to non-reconstructed regions and inversely recovering the waveform, deeply suppressing redundant interference noise while preserving the dynamic acoustic structure, and effectively improving the speech feature restoration accuracy in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech enhancement technology, and in particular to an AI audio processor, method, and storage medium thereof. Background Technology

[0002] The field of speech enhancement technology involves technologies for processing speech signals affected by noise interference, distortion, and reverberation during acquisition, transmission, and playback. It mainly includes core aspects such as time-domain waveform analysis, frequency-domain spectral feature extraction, noise component estimation, speech and noise separation, and post-enhanced speech reconstruction. Typically, it involves acquiring the original speech signal using a microphone array, obtaining spectral information through short-time Fourier transform, determining speech and non-speech segments by combining speech activity detection, adjusting the target speech segment based on noise spectrum estimation results, and recovering the time-domain signal through inverse transform to meet the speech quality processing requirements of communication, recording, and interactive applications.

[0003] One traditional AI audio processor method and its storage medium refer to a hardware and software combined processing scheme built for speech signal processing needs. It performs frame-by-frame windowing processing on the input speech signal, extracts Mel frequency cepstral coefficients or linear prediction coefficients as feature parameters, uses a pre-trained acoustic model to distinguish speech from noise, adjusts or replaces the spectral amplitude, and combines Wiener filtering or spectral subtraction to suppress noise bands. Finally, it synthesizes the output speech through overlapping and addition. At the same time, the processing steps are stored in the storage medium in the form of program instructions for the processor to call.

[0004] Existing technologies often rely on global spectral features for unified adjustment and replacement when dealing with complex speech processing needs. They lack the ability to dynamically track the changes in resonant acoustic features in the frequency domain over time. The fixed mechanism is difficult to identify abnormal spectral drift caused by interference in a variable environment. Forcibly applying filtering rules can easily destroy the resonant structure in the target speech segment. It fails to combine the energy distribution state of consecutive time frames to constrain specific frequency bands, resulting in excessive loss of core acoustic details during the operation, which seriously weakens the clarity and realism of the output and cannot meet the requirements of fine audio recovery under dynamic interference. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing an AI audio processor, method, and storage medium thereof.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: an AI audio processor, the processor comprising: The frequency domain mapping module acquires discrete audio sequences and performs frequency domain feature mapping through the microphone of the AI ​​audio processor. It then filters out frequency points whose energy exceeds that of adjacent positions, constructs a candidate set of formants, and passes it to the trajectory generation module. The trajectory generation module calculates the Euclidean distance between adjacent frequency points based on the candidate set of formants and extracts the frequency point with the minimum distance. It then filters the frequency points whose sign product of the frequency difference corresponding to consecutive time frames is positive, generates a stable trajectory set, and passes it to the anomaly correction module. The anomaly correction module extracts the neighborhood energy distribution sequence that does not belong to the stable trajectory set, combines it with the cumulative calculation of the frequency energy difference of the previous time frame, filters the frequency points that exceed the preset deviation threshold and performs extreme value search, constructs the correction frequency set and passes it to the spectrum reconstruction module. The spectrum reconstruction module merges the corrected frequency set and the stable trajectory set and estimates the spectrum envelope, calculates the frequency change difference between adjacent time frames and filters out frequency points that exceed a preset judgment threshold, constructs a reconstructed spectrum set and transmits it to the audio generation module. The audio generation module attenuates the energy of the non-reconstructed frequency bands by a fixed ratio based on the reconstructed spectrum set and performs an inverse short-time Fourier transform to generate a speaker driving waveform.

[0007] As a further embodiment of the present invention, the candidate set of resonant peaks includes peak frequency coordinates, energy intensity, and peak shape parameters; the stable trajectory set includes trajectory continuity identifier, frequency evolution direction, and time correlation index; the corrected frequency set includes frequency offset, anomaly correction mark, and energy compensation result; the reconstructed spectrum set includes spectral envelope curve, frequency jump identifier, and spectral line reconstruction structure; and the loudspeaker driving waveform includes time-domain sampling point sequence, waveform amplitude envelope, and driving signal phase.

[0008] As a further aspect of the present invention, the frequency domain mapping module includes: The data analysis submodule acquires analog-to-digital conversion data from the microphone terminal of the AI ​​audio processor, collects continuous sampling voltage sequences and arranges them according to a fixed sampling period, maps the voltage of each sampling point according to the sampling time, divides the continuous sampling voltage sequence into discrete time sequences and normalizes them to obtain a discrete audio amplitude set. The spectrum transformation submodule, based on the audio discrete amplitude set, divides the frame according to the fixed frame length and frame shift parameters, performs complex form weighted superposition operation on the amplitude of the sampling points in each frame and calculates the square of the amplitude of the corresponding frequency component, and arranges the squares of the amplitudes of the frequency components of multiple frames according to frequency to obtain multi-frame frequency energy data. The peak filtering submodule, based on the multi-frame frequency energy data, performs adjacent frequency index difference comparison on the energy at multiple frequency index positions within each time frame, selects the frequency index where the current frequency energy simultaneously exceeds the energy values ​​on both sides and records it, thus obtaining a candidate set of resonant peaks.

[0009] As a further aspect of the present invention, the trajectory generation module includes: The frequency matching submodule, based on the formant candidate set, reads the coordinates of frequency points in the current time frame and the previous time frame point by point and constructs a frequency coordinate pair sequence, calls the frequency coordinate pair sequence to calculate the Euclidean distance, binds the matching relationship according to the minimum distance index and rearranges it to generate a frequency matching pair sequence. The sign determination submodule extracts the corresponding frequencies from three consecutive time frames of the sequence based on the frequency matching, performs subtraction on the frequencies of adjacent time frames to obtain a difference sequence, extracts the sign bits of adjacent terms in the difference sequence and performs multiplication, determines the sign consistency based on the product result, and obtains the sign product sequence. The trajectory filtering submodule synchronously traverses the symbol product sequence and the corresponding frequency point index sequence, filters and marks the index positions with product values ​​greater than zero, and reassembles the corresponding time frame frequency points in chronological order, connects the sequences and numbers them to generate a stable trajectory set.

[0010] As a further aspect of the present invention, the anomaly correction module includes: The neighborhood energy submodule extracts the neighborhood energy distribution sequence for frequency points in the candidate set of resonant peaks that do not belong to the stable trajectory set, obtains the spectral amplitude data corresponding to multiple frequency points, performs sliding segment accumulation according to a preset frequency window, and performs normalized amplitude conversion to generate a frequency neighborhood energy sequence. The deviation accumulation submodule calls the corresponding frequency energy data in the previous time frame based on the frequency neighborhood energy sequence, performs frequency point difference accumulation calculation and constructs a time series offset array, and compares and filters the offset array corresponding to each frequency point with the preset deviation threshold to obtain the abnormal frequency offset set. The extreme value positioning submodule extracts the corresponding frequency points based on the abnormal frequency offset set and reads the continuous energy value sequence within the surrounding preset frequency range. It performs the maximum amplitude search within the interval and determines the corresponding frequency coordinates. It then combines the extreme value positions of multiple frequency points for summary mapping and establishes a corrected frequency set.

[0011] As a further aspect of the present invention, the deviation threshold is determined by collecting a multi-frame frequency energy sequence of initial background noise, calculating the standard deviation of the frequency energy fluctuation amplitude within a preset silent time frame to extract the basic fluctuation value, and multiplying the basic fluctuation value with a preset tolerance coefficient.

[0012] As a further aspect of the present invention, the spectrum reconstruction module includes: The resonance sequence submodule performs point-by-point pairing data extraction and superposition based on the modified frequency set and the stable trajectory set, rearranges and merges them according to time to obtain a continuous frequency trajectory array and accumulates the amplitude, divides the interval according to the set frequency bandwidth, records multiple corresponding time markers, and generates a resonance feature sequence. The envelope estimation submodule sorts the global frequency index based on the resonance feature sequence, extracts the amplitude of adjacent frequency points to construct a continuous amplitude vector sequence, performs sliding window averaging along the frequency axis, and extracts the amplitude data corresponding to the center frequency position of multiple sliding windows to obtain a continuous spectrum envelope sequence. The differential filtering submodule extracts multiple corresponding frequency position data according to the continuous spectrum envelope sequence in time frame order, calculates the difference between the same frequency points in adjacent time frames and constructs a frequency difference array, filters the coordinates of frequency points that exceed the preset change threshold and reassembles them to obtain the reconstructed spectrum set.

[0013] As a further aspect of the present invention, the audio generation module includes: The frequency band attenuation submodule extracts the energy value array corresponding to the non-reconstructed frequency band based on the reconstructed spectrum set, performs point-by-point multiplication operation in combination with the preset fixed attenuation ratio coefficient, calculates the corresponding amplitude after proportional attenuation at multiple frequency index positions, and arranges and reassembles them according to the original frequency index order to obtain the frequency band attenuation amplitude sequence. The spectrum splicing submodule calls the frequency band attenuation amplitude sequence and the energy array of the corresponding frequency band of the reconstructed spectrum set to align the frequency index, writes the proportionally attenuated amplitude data into the index interval of the non-reconstructed frequency band, merges the retained original reconstructed frequency band amplitude data and splices the sequence to generate the corrected spectrum sequence. The time-domain transformation submodule extracts multiple frames of frequency domain complex data from the corrected spectrum sequence, performs discrete inverse short-time Fourier integration, obtains time-domain discrete sampling points corresponding to multiple time frames, overlaps and adds the multiple time-domain sampling points according to time identifiers, and outputs them to the AI ​​audio processor to establish the speaker driving waveform.

[0014] An AI audio processing method is used to execute an AI audio processor, the AI ​​audio processing method comprising: S1: Acquire discrete audio sequences and perform frequency domain feature mapping through the microphone end of the AI ​​audio processor, filter out frequency points whose energy exceeds that of adjacent positions, and construct a candidate set of formants; S2: Calculate the Euclidean distance between adjacent frequency points based on the candidate set of formants and extract the frequency point with the minimum distance. Select the frequency points whose sign product of the frequency difference corresponding to consecutive time frames is positive and generate a stable trajectory set. S3: Extract the neighborhood energy distribution sequence that does not belong to the stable trajectory set, combine it with the cumulative calculation of the frequency energy difference of the previous time frame, filter the frequency points that exceed the preset deviation threshold and perform extreme value search to construct the corrected frequency set; S4: Merge the corrected frequency set and the stable trajectory set and estimate the spectral envelope, calculate the frequency change difference between adjacent time frames and filter out frequency points that exceed the preset judgment threshold, and construct the reconstructed spectrum set; S5: Based on the reconstructed spectrum set, the energy of the non-reconstructed frequency band is attenuated by a fixed ratio and an inverse short-time Fourier transform is performed to generate the loudspeaker driving waveform.

[0015] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of an AI audio processor as described above.

[0016] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, a resonance candidate set is established by capturing the energy extrema of multiple frequency points, and a stable trajectory is selected based on the sign product of the difference between adjacent frames. The core frequency band is accurately locked and spectral drift is investigated. In conjunction with the energy distribution of the preceding time frame, the deviation is determined and the extrema are re-searched by difference calculation. The abnormal frequency is adaptively repaired. The corrected frequency point and the stable trajectory are merged to perform envelope estimation and difference segment reconstruction. A fixed attenuation is applied to the non-reconstruction area and the waveform is restored in reverse. While preserving the dynamic acoustic structure, redundant interference noise is deeply suppressed, which effectively improves the speech feature restoration in complex environments. Attached Figure Description

[0017] Figure 1 This is a system schematic diagram of the present invention; Figure 2 This is a schematic diagram of the system framework of the present invention; Figure 3 This is a flowchart of the frequency domain mapping module in this invention; Figure 4 This is a flowchart of the trajectory generation module in this invention; Figure 5 This is a flowchart of the anomaly correction module in this invention; Figure 6 This is a flowchart of the spectrum reconstruction module in this invention; Figure 7 This is a flowchart of the audio generation module in this invention; Figure 8 This is a flowchart of the method of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0019] Please see Figure 1 This invention provides a technical solution: an AI audio processor, the processor comprising: The frequency domain mapping module acquires discrete audio sequences and performs frequency domain feature mapping through the microphone of the AI ​​audio processor. It then filters out frequency points whose energy exceeds that of adjacent positions, constructs a candidate set of formants, and passes it to the trajectory generation module. The trajectory generation module calculates the Euclidean distance between adjacent frequency points based on the formant candidate set and extracts the frequency point with the minimum distance. It then filters the frequency points whose sign product of the frequency difference corresponding to consecutive time frames is positive, generates a stable trajectory set, and passes it to the anomaly correction module. The anomaly correction module extracts the neighborhood energy distribution sequence of the unassigned stable trajectory set, combines it with the cumulative calculation of the frequency energy difference of the previous time frame, filters the frequency points that exceed the preset deviation threshold and performs extreme value search, constructs the correction frequency set and passes it to the spectrum reconstruction module. The spectrum reconstruction module merges the corrected frequency set and the stable trajectory set and estimates the spectrum envelope, calculates the frequency change difference between adjacent time frames and filters out frequency points that exceed the preset judgment threshold, constructs the reconstructed spectrum set and passes it to the audio generation module. The audio generation module attenuates the energy of the non-reconstructed frequency bands by a fixed ratio based on the reconstructed spectrum set and performs an inverse short-time Fourier transform to generate the speaker driving waveform.

[0020] The candidate set of formants includes peak frequency coordinates, energy intensity, and peak shape parameters; the stable trajectory set includes trajectory continuity identifiers, frequency evolution direction, and time correlation index; the corrected frequency set includes frequency offset, anomaly correction markers, and energy compensation results; the reconstructed spectrum set includes spectral envelope curves, frequency jump identifiers, and spectral line reconstruction structure; and the loudspeaker driving waveform includes time-domain sampling point sequence, waveform amplitude envelope, and driving signal phase.

[0021] Specifically, such as Figure 2 , 3 As shown, the frequency domain mapping module includes: The data analysis submodule acquires analog-to-digital conversion data from the microphone terminal of the AI ​​audio processor, collects continuous sampling voltage sequences and arranges them according to a fixed sampling period, maps the voltage of each sampling point according to the sampling time, divides the continuous sampling voltage sequence into discrete time sequences and normalizes them to obtain a discrete audio amplitude set. The AI ​​audio processor directly extracts analog-to-digital conversion (ADC) data from the microphone terminal of the intelligent audio processor. The ADC unit continuously acquires the analog voltage signal from the microphone terminal at a fixed sampling frequency of 48,000 Hz, converting it into a raw, continuously sampled voltage sequence in digital form. The data arrangement unit linearly arranges this continuous sampled voltage sequence according to a fixed sampling period of 20.8 microseconds, constructing a one-dimensional initial voltage matrix. The time mapping component then performs absolute time mapping on each sampled voltage based on its entry timestamp, with the mapping reference value set to time 0 at startup, thus assigning a unique time attribute to each voltage data point. The discretization processing unit, based on the mapped absolute time axis, strictly truncates the continuous sampled voltage sequence at equal intervals according to the aforementioned fixed sampling period, dividing it into a discrete time sequence. The normalization component then performs max-min normalization operations on the voltage values ​​of each sampled point within the discrete time sequence. The normalization component extracts the current sampling point voltage value from the discrete-time series and obtains the highest and lowest sampling voltage values ​​within a 200-millisecond statistical window. It then calculates the difference between the current and lowest sampling voltage values ​​to obtain the first difference, and the difference between the highest and lowest sampling voltage values ​​to obtain the second difference. Finally, it calculates the ratio of the first and second differences and outputs the standard normalized amplitude of the current sampling point. For example, if the sensor detects a current sampling point voltage of 1.5 volts, and the lowest sampling voltage within the window is -2.0 volts and the highest is 3.0 volts, the first difference after subtraction is 3.5 volts, and the second difference is 5.0 volts. Substituting these values ​​into a division operation yields a standard normalized amplitude of 0.7 for this sampling point. For data streams containing multiple sets of continuous voltage values, the same iterative processing logic is used to obtain all corresponding standard amplitude results. The discretization processing unit outputs all processed standard normalized amplitude values, which are then combined to generate the final audio discrete amplitude set.

[0022] The spectrum transformation submodule, based on the discrete amplitude set of audio, divides the frame according to the fixed frame length and frame shift parameters, performs complex form weighted superposition operation on the amplitude of the sampling points in each frame and calculates the square of the amplitude of the corresponding frequency component, and arranges the square of the amplitude of the frequency components of multiple frames according to frequency to obtain multi-frame frequency energy data. The system receives the generated discrete audio amplitude set and performs truncating and segmentation processing on it according to pre-set fixed frame length and fixed frame shift parameters. Specifically, the fixed frame length parameter is set to 1024 sampling points, and the fixed frame shift parameter is set to 256 sampling points. The value of this fixed frame shift parameter was determined based on multiple overlap rate experiments, which showed that a 75% overlap rate optimally captures rapid changes in speech phonemes. The windowing processing component then applies a Hamming window function to the amplitude values ​​of the 1024 sampling points within each frame for smoothing and weighting, attenuating abrupt amplitude changes at frame boundaries. The frequency calculation unit then performs a discrete Fourier transform on the smoothed amplitude values ​​of each sampling point. Specifically, it multiplies each sampling point amplitude value with a complex rotation factor of the corresponding frequency, and then sums the 1024 product results for the entire frame to generate complex-form spectral data. The energy extraction unit obtains the complex number result obtained by summation, extracts its real and imaginary parts separately, multiplies the real part by itself to obtain the square of the real part, multiplies the imaginary part by itself to obtain the square of the imaginary part, and finally sums the squares of the real and imaginary parts to obtain the squared amplitude value of the corresponding frequency component. In the aforementioned test example where the real part equals 4.0 and the imaginary part equals 3.0, substituting the multiplication and addition logic, we obtain a squared real part value of 16.0 and a squared imaginary part value of 9.0. The sum of the two yields a squared amplitude value of 25.0. The full-frequency calculation component repeatedly performs the above calculations until all frequency components in the range from 0 Hz to 24000 Hz are covered. Finally, the data arrangement component arranges the squared amplitude values ​​of all frequency components within multiple frames in a two-dimensional matrix according to the frequency values ​​from low to high, outputting multi-frame frequency energy data.

[0023] The peak filtering submodule, based on multi-frame frequency energy data, performs differential comparison of adjacent frequency indices on the energy at multiple frequency index positions within each time frame, selects the frequency indices whose current frequency energy simultaneously exceeds the energy values ​​on both sides and records them, thus obtaining a candidate set of resonant peaks; The system reads multi-frame frequency energy data from the output and the differential comparison unit extracts the energy data of the multi-frequency index positions within each time frame. For the energy value of the current frequency index, this unit obtains the energy values ​​of the low-frequency index adjacent to its left and the high-frequency index adjacent to its right. The differential operation component subtracts the current frequency energy value from the low-frequency index energy value to generate the left difference, and subtracts the current frequency energy value from the high-frequency index energy value to generate the right difference. The determination component simultaneously checks the sign of the left and right differences. Only when both the left and right differences are positive is the current frequency energy determined to exceed the adjacent energy values ​​on both sides. For example, if the calculated current frequency energy value is 25.0, the low-frequency index energy value is 18.0, and the high-frequency index energy value is 20.0, after performing the difference operation, the left difference is 7.0, and the right difference is 5.0. Since both 7.0 and 5.0 are greater than 0, the current frequency energy is determined to be a local maximum. The threshold verification component further compares the current frequency energy value with the environmental dynamic energy benchmark value. The environmental dynamic energy baseline is obtained by calculating the average of all frequency energy values ​​within the current time frame and multiplying it by a weighting factor of 1.5. This weighting factor of 1.5 is determined based on extensive statistical experiments of voiced and unvoiced sounds and has the ability to filter out most background noise spurious peaks. The candidate recording component selects the frequency index of the local maximum point where the energy value is greater than the environmental dynamic energy baseline and records it in the final formant candidate set.

[0024] Table 1. Comparison Results of Frequency-Energy Difference As shown in Table 1, the local extrema in the sequence can be accurately located by the difference operation of adjacent data.

[0025] Specifically, such as Figure 2 , 4 As shown, the trajectory generation module includes: The frequency matching submodule, based on the formant candidate set, reads the coordinates of frequency points in the current time frame and the previous time frame point by point and constructs a frequency coordinate pair sequence. It then calls the frequency coordinate pair sequence to calculate the Euclidean distance, binds the matching relationship according to the minimum distance index, and rearranges it to generate a frequency matching pair sequence. The generated formant candidate set data is read. The candidate set extraction unit extracts frequency point data from the previous time frame and the current time frame. The coordinate construction component obtains the absolute frequency value of each frequency point as the horizontal axis data, and simultaneously obtains the previously generated 25.0 equal amplitude squared value as the vertical axis data, combining them to generate a frequency coordinate pair sequence of the previous time frame and the current time frame. The distance calculation component performs distance calculation for a single reference coordinate in the previous time frame, traversing all target coordinates in the current time frame. The calculation logic is as follows: subtract the horizontal axis data of the target coordinate from the horizontal axis data of the reference coordinate to obtain the first difference value; multiply the first difference value by itself to obtain the first squared value; subtract the vertical axis data of the target coordinate from the vertical axis data of the reference coordinate and multiply it by itself to obtain the second squared value; add the first squared value and the second squared value; and perform a square root operation on the sum to obtain the distance value between the two coordinate points. For example, in the previous frame, the horizontal axis of a reference coordinate is 1000 Hz and the vertical axis is 25.0. In the current frame, the horizontal axis of a target coordinate is 1050 Hz and the vertical axis is 24.0. The first difference is 50, and its square value is 2500. The second difference is -1.0, and its square value is 1.0. The sum of these two values ​​is 2501. Taking the square root yields a distance value of 50.01. The matching and binding unit compares and filters all distance values, extracts the target coordinate corresponding to the smallest value, establishes a connection between the two values, and outputs them sequentially as a frequency matching pair sequence. The advantage of this operation logic is that it integrates spectral features through two-dimensional spatial measurement. Comparing this distance result of 50.01 with the set conventional distance upper limit of 100.0, the value is within a reasonable matching range, representing a high confidence level in the point binding.

[0026] The sign determination submodule extracts the corresponding frequencies from three consecutive time frames of the sequence based on frequency matching, performs subtraction on the frequencies of adjacent time frames to obtain a difference sequence, extracts the sign bits of adjacent terms in the difference sequence and performs multiplication, and determines the sign consistency based on the product result to obtain the sign product sequence. The receiver receives a frequency matching pair sequence. Based on the established matching relationship, the frequency extraction unit sequentially extracts the actual frequency values ​​bound together within consecutive time frames (1st, 2nd, and 3rd). The subtraction operation component performs differential logic calculations on the extracted consecutive frequency values. This component first subtracts the frequency value of the 1st time frame from the frequency value of the 2nd time frame to generate the first difference value, then subtracts the frequency value of the 2nd time frame from the frequency value of the 3rd time frame to generate the second difference value. These difference values ​​are then combined to generate a preliminary difference sequence. The symbol quantization extraction unit then performs discrete quantization on each value in the difference sequence, assigning a non-numerical state. The quantization standard is set as follows: a difference value greater than 0 is assigned a value of 1; a difference value less than 0 is assigned a value of -1; and a difference value equal to 0 is directly assigned a value of 0. The consistency operation component extracts the quantized symbol values ​​of two adjacent time intervals and performs multiplication. Taking the extracted values ​​from the aforementioned example, let the frequency of the first time frame be 1000 Hz, the second time frame be 1050 Hz, and the third time frame be 1120 Hz. Subtraction yields a first difference value of 50 and a second difference value of 70. According to the quantization standard, 50 corresponds to an assignment value of 1, and 70 also corresponds to an assignment value of 1. Multiplying these two assignment values ​​results in a final product of 1, which is recorded in the sign product sequence. The advantage of this operational logic is that direct multiplication using discrete sign bits significantly reduces the hardware resource overhead of floating-point operations. Comparing the resulting product value of 1 with the baseline threshold of 0 confirms that the frequency maintains a unidirectional shift pattern within three consecutive frames.

[0027] The trajectory filtering submodule synchronously traverses the symbol product sequence and the corresponding frequency point index sequence, filters and marks the index positions with product values ​​greater than zero, and reassembles the corresponding time frame frequency points in chronological order, connects and numbers the sequences, and generates a stable trajectory set. Simultaneously, the output symbol product sequence and the corresponding frequency point index position sequence are retrieved. The synchronous traversal component reads the product values ​​and their associated time frame index numbers one by one according to the order of data arrangement. The conditional comparison and filtering unit obtains the product value at the current index position and introduces a pre-set zero value comparison benchmark. The device compares the read product value with the zero value comparison benchmark. When the product value is determined to be strictly greater than 0, the retention execution unit adds an allowance mark to the corresponding frequency point index position. When the product value is equal to or less than 0, a rejection operation is performed. For example, substituting the product value 1 obtained in the aforementioned calculation step, since 1 is greater than the set benchmark, the component assigns a retention mark to the frequency evolution path index from 1000 Hz to 1120 Hz. The recombination and connection unit retrieves all index positions carrying retention marks and extracts the specific frequency coordinate points within the corresponding continuous time frames. Subsequently, this unit splices and recombines the aforementioned filtered frequency points according to the chronological order of the time axis, assigning them continuously increasing Arabic numerals starting from 1 as independent numbers. The combined dataset is output as a stable trajectory set. The advantage of this operational logic is that it prevents the spread of errors caused by high-frequency random jitter through a strict consistency check mechanism.

[0028] Specifically, such as Figure 2 , 5 As shown, the anomaly correction module includes: The neighborhood energy submodule extracts the neighborhood energy distribution sequence for frequency points in the candidate set of resonant peaks that do not belong to the stable trajectory set, obtains the spectral amplitude data corresponding to multiple frequency points, performs sliding segment accumulation according to the preset frequency window, and performs normalized amplitude conversion to generate the frequency neighborhood energy sequence. The system retrieves isolated frequency point data from the candidate set of formants that are not assigned to the stable trajectory set. For each isolated frequency point, the amplitude acquisition unit accurately extracts the continuous spectral amplitude data of the corresponding surrounding area from the original spectral data. The sliding accumulation component sets the preset frequency window size parameter to 50 Hz and the sliding step parameter to 10 Hz. This component extracts amplitude data from a central isolated frequency point, for example, 1200 Hz, and extracts amplitude data from five equidistant sampling points within the interval of 1175 Hz to 1225 Hz. In a specific example, the amplitude data of these five sampling points are 12.5, 14.0, 15.5, 13.5, and 11.5, respectively. The sliding accumulation component sequentially adds these five amplitude data, outputting an initial local accumulated energy value of 67.0. The normalization conversion unit then processes this initial local accumulated energy value. The parameter configuration component obtains the global maximum window energy value of 100.0 and the global minimum window energy value of 10.0 within the current time frame through a full-frame scan. The normalization conversion unit subtracts the global maximum window energy value of 100.0 from the global minimum window energy value of 10.0, outputting a denominator range value of 90.0. This unit then subtracts the currently acquired local accumulated energy value of 67.0 from the global minimum window energy value of 10.0, outputting a numerator difference value of 57.0. The division operation component divides the numerator difference value of 57.0 by the denominator range value of 90.0, outputting a standardized normalized amplitude conversion result of 0.63. This unit iteratively processes all isolated frequency points, generating a continuous frequency neighborhood energy sequence. The advantage of this operation logic is that it eliminates transient abrupt interference at single frequency points through the normalization operation mechanism of continuous accumulation of local window energy combined with global extreme value difference.

[0029] The deviation accumulation submodule calls the corresponding frequency energy data in the previous time frame based on the frequency neighborhood energy sequence, performs frequency-point difference accumulation calculation and constructs a time series offset array, and compares and filters the offset array corresponding to each frequency point with the preset deviation threshold to obtain the abnormal frequency offset set. Based on the input frequency neighborhood energy sequence, the historical energy data at the same frequency index position in the previous time frame is read. The difference calculation unit extracts the normalized amplitude data of the current time frame and the historical energy data of the previous time frame for each frequency point. This unit subtracts the historical energy data of the previous time frame from the normalized amplitude data of the current time frame, outputting a single-step energy difference value. The multi-frame accumulation component extracts the single-step energy difference values ​​within five consecutive time frames and performs continuous addition. In the actual calculation scenario, the five consecutive frame difference values ​​extracted for the 1200 Hz frequency point are 0.15, 0.12, 0.18, 0.10, and 0.14, respectively. The multi-frame accumulation component sums these five values, outputting a total offset value of 0.69. This component performs the same logic for each frequency point, arranging the generated total offset values ​​in frequency index order to construct a complete time series offset array. The comparison and judgment component reads the benchmark deviation threshold parameter from the preset configuration library. The baseline deviation threshold parameter is calculated by multiplying the statistical mean of normal speech baseband fluctuations (0.25) by a fixed tolerance multiplier of 2.0, with a fixed value of 0.50. The comparison and judgment component extracts the total offset value of each item in the offset array and performs a quantified comparison with the baseline deviation threshold of 0.50. For the aforementioned calculated total offset value of 0.69, since 0.69 is strictly greater than 0.50, the filtering and retention unit assigns an anomaly mark to this 1200 Hz frequency point and centrally extracts all frequency points with anomaly marks to generate an abnormal frequency offset set.

[0030] Table 2 Abnormal Frequency Offset Judgment Data Table As shown in Table 2, the energy drift state at different frequency points was accurately distinguished through threshold comparison. The advantage of this operational logic is that it amplifies the hidden frequency drift trend by using a multi-frame interpolation operation combined with a fixed threshold control mechanism. Comparing the calculated 0.69 with 0.50, the value is found to be out of limit, confirming that abnormal energy ionization has occurred at this frequency point.

[0031] The extreme value location submodule extracts the corresponding frequency points based on the abnormal frequency offset set and reads the continuous energy value sequence within the surrounding preset frequency range. It performs the maximum amplitude search within the interval and determines the corresponding frequency coordinates. It then combines the extreme value positions of multiple frequency points for summary mapping and establishes a corrected frequency set. The coordinate data of anomalous frequency points are extracted one by one from the input set of anomalous frequency offsets. The interval definition unit uses the extracted anomalous frequency points as the center anchor point, subtracting 100 Hz to the left to generate a lower search bound, and adding 100 Hz to the right to generate an upper search bound, thus defining a preset frequency range with a total width of 200 Hz. The data reading component extracts all continuous energy value sequences within this preset frequency range from the underlying spectrum matrix. The maximum value comparison component performs a successive comparison operation on all energy values ​​within this sequence. For the aforementioned 1200 Hz anomalous point example, there are 5 significant test points within the corresponding 1100 Hz to 1300 Hz interval, with energy values ​​of 21.5, 23.2, 28.4, 24.6, and 20.1, respectively. The maximum value comparison component compares 21.5 with 23.2 and retains 23.2, compares 23.2 with 28.4 and retains 28.4, and continues to traverse the entire interval, ultimately locking 28.4 as the highest energy value within the interval. The coordinate mapping unit retrieves the horizontal axis parameter corresponding to the highest energy value of 28.4, confirming its precise location at 1235 Hz. The replacement and rewriting component discards the original abnormal frequency point of 1200 Hz and replaces it in place with the newly acquired true extreme value coordinates at 1235 Hz. The output component collects all the new frequency coordinates that have undergone replacement and rewriting, reassembles them into a one-dimensional matrix in ascending order of value, establishes the final corrected frequency set, and outputs it to the downstream storage area. The advantage of this operation logic is that it corrects the frequency domain misalignment problem caused by the initial positioning through adaptive wideband exhaustive search of extreme values.

[0032] Specifically, such as Figure 2 , 6 As shown, the spectrum reconstruction module includes: The resonance sequence submodule performs point-by-point pairing data extraction and superposition based on the corrected frequency set and the stable trajectory set, rearranges and merges them according to time to obtain a continuous frequency trajectory array and accumulates the amplitude, divides the interval according to the set frequency bandwidth, records multiple corresponding time markers, and generates a resonance feature sequence. The system reads the constructed corrected frequency set and stable trajectory set. The pairing and overlay component extracts the target isolated frequency point from the corrected frequency set for a specific time frame, such as the previously obtained 1235 Hz corrected frequency point and its corresponding amplitude of 28.4. Simultaneously, it extracts adjacent consecutive trajectory frequency points within the same time frame from the stable trajectory set, such as 1230 Hz and its corresponding amplitude of 25.1. This component adds the amplitude data of the two adjacent frequency points, i.e., 28.4 and 25.1, outputting an overlaid amplitude data of 53.5. The rearrangement and merging unit arranges all frequency points that have undergone overlay processing in one-dimensional order according to time, combining them to establish a continuous frequency trajectory array. The accumulation operation component extracts the overlaid amplitude data of the 1235 Hz frequency position within three consecutive time frames from this continuous frequency trajectory array, for example, 53.5, 50.2, and 48.6 respectively, and continuously adds these three data points to obtain the total amplitude data of the local trajectory, which is 152.3. The interval division unit reads the preset bandwidth parameter and sets it to 100 Hz, dividing the effective speech spectrum range from 0 Hz to 4000 Hz into 40 independent fixed intervals. The mapping and recording component locates the 1200 Hz to 1300 Hz frequency band where the aforementioned 1235 Hz is located, and accurately maps and stores the total amplitude data of this local trajectory (152.3) into the designated storage slot within this interval. Simultaneously, it calls the clock component to read the time stamp parameter corresponding to the current data, such as the exact timestamps of frames 5 to 7. The aggregation and generation unit packages the accumulated amplitude and time stamp of each interval into a unified package, generates a structured resonance feature sequence according to the temporal sequence, and outputs it to the data bus. The advantage of this operation logic is that it strengthens the energy characterization of the core resonance band through the direct superposition and fusion mechanism of the amplitudes of isolated correction points and stable trajectory points.

[0033] The envelope estimation submodule sorts the global frequency index based on the resonance feature sequence, extracts the amplitude of adjacent frequency points to construct a continuous amplitude vector sequence, and performs sliding window averaging along the frequency axis to extract the amplitude data corresponding to the center frequency position of multiple sliding windows, thus obtaining a continuous spectrum envelope sequence. The system receives the generated resonance feature sequence data as input. The global sorting component reads all frequency indices within the resonance feature sequence and performs a monotonically increasing sorting operation strictly according to the frequency values ​​in ascending order. The vector construction component extracts the cumulative amplitude data of adjacent frequency points sequentially based on the sorted frequency indices. Taking five consecutive uniformly distributed test frequency points in the 1200 Hz to 1300 Hz interval as an example, the extracted corresponding cumulative amplitude data are 152.3 obtained earlier, and subsequently associated 145.6, 138.2, 130.5, and 125.4. This component concatenates the aforementioned five amplitude data to construct a local continuous amplitude vector sequence. The moving average calculation unit reads the preset moving window length parameter, which is strictly defined as containing three frequency points. The unit aligns the starting point of the window with the first three terms of the aforementioned vector sequence, namely 152.3, 145.6, and 138.2, and performs a summation operation to obtain a local sum of 436.1. Then, the unit divides this local sum of 436.1 by the window length parameter 3, outputting a smoothed moving average amplitude of 145.36. The extraction component precisely locates the center frequency position within this moving window, corresponding to the second frequency point at 1225 Hz, and assigns the previously calculated average amplitude of 145.36 to this center frequency position. The moving average calculation unit continues to slide the window towards higher frequencies according to a single frequency point step size, covering the new data combination 145.6, 138.2, and 130.5. Similarly, a summation operation is performed to obtain a total of 414.3, and a division operation is performed to output a new average amplitude of 138.10, which is then assigned a new center frequency position of 1250 Hz. The sequence recombination component performs a two-dimensional combination of time and frequency domain data on all newly generated center frequency positions and their corresponding moving average amplitude data to generate a continuous spectral envelope sequence for output.

[0034] The differential filtering submodule extracts multiple corresponding frequency position data according to the continuous spectrum envelope sequence in time frame order, calculates the difference between the same frequency points in adjacent time frames and constructs a frequency difference array, filters the coordinates of the frequency points that exceed the preset change threshold and reassembles them to obtain the reconstructed spectrum set; The envelope data for corresponding frequency positions is extracted from the continuous spectral envelope sequence according to a strict chronological order of time frames. The difference calculation unit locks onto the same fixed frequency position, such as 1225 Hz, and extracts the envelope amplitude data of the current time frame (the 6th frame), which is the previously calculated 145.36. Simultaneously, it reads the envelope amplitude data of the adjacent previous time frame (the 5th frame) at the same frequency point, for example, 135.16. This unit subtracts the historical envelope amplitude data of the previous time frame (135.16) from the current time frame's envelope amplitude data (145.36), outputting a frequency difference value of 10.20 in the time dimension. The array construction component performs the same subtraction logic on all frequency points within the monitoring range, arranging and combining all the acquired frequency difference values ​​according to their original frequency index positions to construct a complete frequency difference array covering the global frequency band. The threshold comparison unit reads the built-in preset change threshold parameter. The specific logic for setting this parameter is as follows: First, 100 segments of pure silent environment test data with a total duration of 50 seconds are collected. Inter-frame difference calculations are performed to obtain the average value, resulting in a baseline fluctuation value of 2.10. Then, this baseline fluctuation value of 2.10 is multiplied by a fixed safety redundancy factor of 2.0, with the specific benchmark threshold set to 4.20. This unit directly compares the calculated frequency difference value of 10.20 with the preset change threshold of 4.20. Since 10.20 is significantly greater than the set benchmark parameter of 4.20, the filtering and retention component determines that the acoustic energy corresponding to the 1225 Hz frequency point has undergone a drastic change and retains the coordinate data of this frequency point. The reconstruction unit centrally summarizes all the marked abruptly changed frequency point coordinates that meet the retention conditions within the current time frame, reorganizes them in order from low frequency to high frequency, thereby obtaining the reconstructed spectrum set after removing stable redundant frequency point data. The advantage of this operation logic is that it quickly locks the transient motion edge of the vocal organs by performing a difference comparison operation through direct subtraction of the same-frequency envelopes of adjacent time frames. The obtained 10.20 difference result was included in the dynamic feature dataset, confirming that the data recorded extremely clear voice opening and closing actions.

[0035] Specifically, such as Figure 2 , 7 As shown, the audio generation module includes: The frequency band attenuation submodule extracts the energy value array corresponding to the non-reconstructed frequency band based on the reconstructed spectrum set, performs point-by-point multiplication operation with the preset fixed attenuation ratio coefficient, calculates the corresponding amplitude after proportional attenuation at multiple frequency index positions, and arranges and reassembles them according to the original frequency index order to obtain the frequency band attenuation amplitude sequence. The generated reconstructed spectrum set is used for reverse mapping to directly extract the energy value array corresponding to the non-reconstructed frequency bands not included in the reconstructed spectrum set. The identification unit scans the global frequency index and locks the non-reconstructed interval, for example, from 2500 Hz to 3000 Hz. The data extraction component extracts a specific isolated noise frequency point, such as 2500 Hz, from this interval and reads its initial energy amplitude as 45.8. The coefficient configuration component pre-stores a fixed attenuation ratio coefficient. The specific value setting process for this attenuation ratio coefficient is as follows: the system collects 100 sets of background noise segments in a silent environment, calculates the root mean square amplitude, takes the reciprocal of the root mean square amplitude, and multiplies it by a fixed scaling constant of 0.5. The actual recorded value is 0.35. The multiplication unit receives the aforementioned initial energy amplitude of 45.8 and the attenuation ratio coefficient of 0.35, performs a direct multiplication operation between the initial energy amplitude of 45.8 and the attenuation ratio coefficient of 0.35, and outputs the target amplitude of 16.03 after proportional attenuation at the corresponding frequency index position. This unit iterates through all frequency nodes within the non-reconstructed frequency bands, performing the same multiplication operation to calculate the corresponding amplitude data after proportional attenuation at multiple frequency index positions in batches. The sequence reconstruction component strictly arranges and merges all calculated target amplitudes in ascending order of the original frequency coordinates, ultimately assembling them to generate a complete frequency band attenuation amplitude sequence. The advantage of this operational logic is that it forcibly suppresses non-core energy bands through static proportional multiplication operations, thereby reducing the overall sound pressure level of broadband background noise.

[0036] The spectrum splicing submodule calls the frequency band attenuation amplitude sequence and the energy array of the corresponding frequency band in the reconstructed spectrum set to align the frequency index, writes the proportionally attenuated amplitude data into the index interval of the non-reconstructed frequency band, merges the retained original reconstructed frequency band amplitude data and splices the sequence to generate the corrected spectrum sequence; The generated reconstructed spectrum set and the newly constructed band attenuation amplitude sequence are read. The alignment mapping component extracts all frequency index coordinates for the complete frequency domain space of the current time frame. This component extracts the high-energy active frequency points retained in the reconstructed spectrum set, such as the 1225 Hz node locked in the reconstruction stage and its corresponding true envelope amplitude of 145.36, and precisely anchors it in index slot 1225 of the output matrix. The write control unit then reads the data in the band attenuation amplitude sequence, extracts the previously calculated 2500 Hz node and its corresponding attenuated amplitude of 16.03, and directly overwrites this value into the blank index interval of index 2500 of the output matrix. This unit continuously scans all free nodes and writes the full amplitude data after proportional attenuation processing into the correct index intervals corresponding to all non-reconstructed frequency bands one by one. After verifying that all frequency slots in the output storage matrix of the merged splicing component have been fully assigned values, the high-energy amplitude data retained in the original reconstructed frequency band and the newly written attenuation amplitude data are deeply concatenated and fused, and arranged continuously and seamlessly according to a monotonically increasing frequency coordinate pattern from low to high. The sequence generation unit performs information loop verification on the arranged dataset, and after confirming that there are no missing frequency coordinates, it generates a structurally complete corrected spectrum sequence. The advantage of this operation logic is that, through strict absolute frequency index alignment and partitioned block directional overwriting mechanism, it ensures the accurate physical splicing of the reconstructed high-frequency features and attenuation noise floor features in the frequency domain dimension. The amplitude difference between adjacent abrupt frequency points after splicing, for example, the difference of 129.33 between 145.36 and 16.03, is compared with the preset spectrum continuity discontinuity threshold of 150.0. This value does not exceed the preset discontinuity warning range, confirming that the global spectrum shape after splicing and recombination maintains a reasonable energy transition gradient.

[0037] The time-domain transformation submodule extracts multiple frames of frequency domain complex data from the corrected spectrum sequence, performs discrete inverse short-time Fourier integration, obtains the time-domain discrete sampling points corresponding to multiple time frames, overlaps and adds the multiple time-domain sampling points according to the time identifier, and outputs them to the AI ​​audio processor to establish the speaker driving waveform. The system continuously extracts frequency domain complex data corresponding to multiple frames from the input corrected spectrum sequence. The data separation unit extracts the complex form of a specific frequency for a single independent time frame, such as frame 6, separating the real part value (e.g., 4.5) and the imaginary part value (e.g., 2.0). The discrete inverse transform unit reads the preset time-domain sampling rate setting parameter of 16000 Hz. This unit multiplies the previously extracted real part value 4.5 with the cosine basis term (e.g., 0.8) of the corresponding frequency node in the current phase to obtain the real part projection value 3.6. Simultaneously, it multiplies the imaginary part value 2.0 with the sine basis term (e.g., 0.6) of the current phase to obtain the imaginary part projection value 1.2. Then, it directly adds the real and imaginary part projection values, outputting the single-point time-domain discrete sampling value of 4.8. The multi-frame overlap component extracts the time-domain waveform data stream of adjacent time frames, locking 128 preset overlap sampling points in the end region of frame 5 and 128 overlap sampling points in the beginning region of frame 6. This component directly adds the residual value of 3.2 at the 10th sample point at the end of frame 5 to the sample point value of 4.8 calculated at the corresponding position at the beginning of frame 6, generating a smooth transition sample point value of 8.0 for the overlapping and blending region. The timing splicing unit strictly concatenates all time-domain sample points that have undergone overlapping and addition processing according to their timestamps, constructing a continuous and complete analog level value string. The output drive component converts this continuous value string into a physical drive voltage command stream and sends it to the AI ​​audio processing chip, ultimately establishing a speaker drive waveform that can be directly mapped to the audio hardware. The advantage of this operation logic is that by combining separate orthogonal projection addition with a multi-frame overlapping addition mechanism, it eliminates phase abrupt changes and audible clicking sounds between adjacent data frames. The calculated peak value of 8.0 in the overlapping region is compared with the speaker hardware drive upper limit protection value of 15.0. This value is within the safe operating range, confirming that the output drive waveform will not cause overload clipping of the physical device.

[0038] Please see Figure 8 An AI audio processing method, based on the aforementioned AI audio processor, includes the following steps: S1: Acquire discrete audio sequences and perform frequency domain feature mapping through the microphone end of the AI ​​audio processor, filter out frequency points whose energy exceeds that of adjacent positions, and construct a candidate set of formants; S2: Calculate the Euclidean distance between adjacent frequency points based on the candidate set of formants and extract the frequency point with the minimum distance. Select the frequency points whose sign product of the frequency difference corresponding to consecutive time frames is positive and generate a stable trajectory set. S3: Extract the neighborhood energy distribution sequence of the unassigned stable trajectory set, combine it with the cumulative calculation of the frequency energy difference of the previous time frame, filter the frequency points that exceed the preset deviation threshold and perform extreme value search to construct the corrected frequency set; S4: Merge the corrected frequency set and the stable trajectory set and estimate the spectral envelope, calculate the frequency change difference between adjacent time frames and filter out frequency points that exceed the preset judgment threshold, and construct the reconstructed spectral set; S5: Based on the reconstructed spectrum set, the energy of the non-reconstructed frequency band is attenuated by a fixed ratio and inverse short-time Fourier transform is performed to generate the loudspeaker driving waveform.

[0039] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. An AI audio processor, characterized by, The processor includes: The frequency domain mapping module acquires discrete audio sequences and performs frequency domain feature mapping through the microphone of the AI ​​audio processor. It then filters out frequency points whose energy exceeds that of adjacent positions, constructs a candidate set of formants, and passes it to the trajectory generation module. The trajectory generation module calculates the Euclidean distance between adjacent frequency points based on the candidate set of formants and extracts the frequency point with the minimum distance. It then filters the frequency points whose sign product of the frequency difference corresponding to consecutive time frames is positive, generates a stable trajectory set, and passes it to the anomaly correction module. The anomaly correction module extracts the neighborhood energy distribution sequence that does not belong to the stable trajectory set, combines it with the cumulative calculation of the frequency energy difference of the previous time frame, filters the frequency points that exceed the preset deviation threshold and performs extreme value search, constructs the correction frequency set and passes it to the spectrum reconstruction module. The spectrum reconstruction module merges the corrected frequency set and the stable trajectory set and estimates the spectrum envelope, calculates the frequency change difference between adjacent time frames and filters out frequency points that exceed a preset judgment threshold, constructs a reconstructed spectrum set and transmits it to the audio generation module. The audio generation module attenuates the energy of the non-reconstructed frequency bands by a fixed ratio based on the reconstructed spectrum set and performs an inverse short-time Fourier transform to generate a speaker driving waveform.

2. The AI audio processor of claim 1, wherein, The candidate set of resonant peaks includes peak frequency coordinates, energy intensity, and peak shape parameters; the stable trajectory set includes trajectory continuity identifier, frequency evolution direction, and time correlation index; the corrected frequency set includes frequency offset, anomaly correction marker, and energy compensation result; the reconstructed spectrum set includes spectral envelope curve, frequency jump identifier, and spectral line reconstruction structure; and the loudspeaker driving waveform includes time-domain sampling point sequence, waveform amplitude envelope, and driving signal phase.

3. The AI ​​audio processor according to claim 1, characterized in that, The frequency domain mapping module includes: The data analysis submodule acquires analog-to-digital conversion data from the microphone terminal of the AI ​​audio processor, collects continuous sampling voltage sequences and arranges them according to a fixed sampling period, maps the voltage of each sampling point according to the sampling time, divides the continuous sampling voltage sequence into discrete time sequences and normalizes them to obtain a discrete audio amplitude set. The spectrum transformation submodule, based on the audio discrete amplitude set, divides the frame according to the fixed frame length and frame shift parameters, performs complex form weighted superposition operation on the amplitude of the sampling points in each frame and calculates the square of the amplitude of the corresponding frequency component, and arranges the squares of the amplitudes of the frequency components of multiple frames according to frequency to obtain multi-frame frequency energy data. The peak filtering submodule, based on the multi-frame frequency energy data, performs adjacent frequency index difference comparison on the energy at multiple frequency index positions within each time frame, selects the frequency index where the current frequency energy simultaneously exceeds the energy values ​​on both sides and records it, thus obtaining a candidate set of resonant peaks.

4. The AI ​​audio processor according to claim 1, characterized in that, The trajectory generation module includes: The frequency matching submodule, based on the formant candidate set, reads the coordinates of frequency points in the current time frame and the previous time frame point by point and constructs a frequency coordinate pair sequence, calls the frequency coordinate pair sequence to calculate the Euclidean distance, binds the matching relationship according to the minimum distance index and rearranges it to generate a frequency matching pair sequence. The sign determination submodule extracts the corresponding frequencies from three consecutive time frames of the sequence based on the frequency matching, performs subtraction on the frequencies of adjacent time frames to obtain a difference sequence, extracts the sign bits of adjacent terms in the difference sequence and performs multiplication, determines the sign consistency based on the product result, and obtains the sign product sequence. The trajectory filtering submodule synchronously traverses the symbol product sequence and the corresponding frequency point index sequence, filters and marks the index positions with product values ​​greater than zero, and reassembles the corresponding time frame frequency points in chronological order, connects the sequences and numbers them to generate a stable trajectory set.

5. An AI audio processor according to claim 1, characterized in that, The anomaly correction module includes: The neighborhood energy submodule extracts the neighborhood energy distribution sequence for frequency points in the candidate set of resonant peaks that do not belong to the stable trajectory set, obtains the spectral amplitude data corresponding to multiple frequency points, performs sliding segment accumulation according to a preset frequency window, and performs normalized amplitude conversion to generate a frequency neighborhood energy sequence. The deviation accumulation submodule calls the corresponding frequency energy data in the previous time frame based on the frequency neighborhood energy sequence, performs frequency point difference accumulation calculation and constructs a time series offset array, and compares and filters the offset array corresponding to each frequency point with the preset deviation threshold to obtain the abnormal frequency offset set. The extreme value positioning submodule extracts the corresponding frequency points based on the abnormal frequency offset set and reads the continuous energy value sequence within the surrounding preset frequency range. It performs the maximum amplitude search within the interval and determines the corresponding frequency coordinates. It then combines the extreme value positions of multiple frequency points for summary mapping and establishes a corrected frequency set.

6. An AI audio processor according to claim 5, characterized in that, The deviation threshold is determined by collecting a multi-frame frequency energy sequence of initial background noise, calculating the standard deviation of the frequency energy fluctuation amplitude within a preset silent time frame to extract the basic fluctuation value, and multiplying the basic fluctuation value with a preset tolerance coefficient.

7. An AI audio processor according to claim 1, characterized in that, The spectrum reconstruction module includes: The resonance sequence submodule performs point-by-point pairing data extraction and superposition based on the modified frequency set and the stable trajectory set, rearranges and merges them according to time to obtain a continuous frequency trajectory array and accumulates the amplitude, divides the interval according to the set frequency bandwidth, records multiple corresponding time markers, and generates a resonance feature sequence. The envelope estimation submodule sorts the global frequency index based on the resonance feature sequence, extracts the amplitude of adjacent frequency points to construct a continuous amplitude vector sequence, performs sliding window averaging along the frequency axis, and extracts the amplitude data corresponding to the center frequency position of multiple sliding windows to obtain a continuous spectrum envelope sequence. The differential filtering submodule extracts multiple corresponding frequency position data according to the continuous spectrum envelope sequence in time frame order, calculates the difference between the same frequency points in adjacent time frames and constructs a frequency difference array, filters the coordinates of frequency points that exceed the preset change threshold and reassembles them to obtain the reconstructed spectrum set.

8. An AI audio processor according to claim 1, characterized in that, The audio generation module includes: The frequency band attenuation submodule extracts the energy value array corresponding to the non-reconstructed frequency band based on the reconstructed spectrum set, performs point-by-point multiplication operation in combination with the preset fixed attenuation ratio coefficient, calculates the corresponding amplitude after proportional attenuation at multiple frequency index positions, and arranges and reassembles them according to the original frequency index order to obtain the frequency band attenuation amplitude sequence. The spectrum splicing submodule calls the frequency band attenuation amplitude sequence and the energy array of the corresponding frequency band of the reconstructed spectrum set to align the frequency index, writes the proportionally attenuated amplitude data into the index interval of the non-reconstructed frequency band, merges the retained original reconstructed frequency band amplitude data and splices the sequence to generate the corrected spectrum sequence. The time-domain transformation submodule extracts multiple frames of frequency domain complex data from the corrected spectrum sequence, performs discrete inverse short-time Fourier integration, obtains time-domain discrete sampling points corresponding to multiple time frames, overlaps and adds the multiple time-domain sampling points according to time identifiers, and outputs them to the AI ​​audio processor to establish the speaker driving waveform.

9. An AI audio processing method, characterized in that, The AI ​​audio processing method is used to execute the AI ​​audio processor according to any one of claims 1 to 8, and the AI ​​audio processing method includes: S1: Acquire discrete audio sequences and perform frequency domain feature mapping through the microphone end of the AI ​​audio processor, filter out frequency points whose energy exceeds that of adjacent positions, and construct a candidate set of formants; S2: Calculate the Euclidean distance between adjacent frequency points based on the candidate set of formants and extract the frequency point with the minimum distance. Select the frequency points whose sign product of the frequency difference corresponding to consecutive time frames is positive and generate a stable trajectory set. S3: Extract the neighborhood energy distribution sequence that does not belong to the stable trajectory set, combine it with the cumulative calculation of the frequency energy difference of the previous time frame, filter the frequency points that exceed the preset deviation threshold and perform extreme value search to construct the corrected frequency set; S4: Merge the corrected frequency set and the stable trajectory set and estimate the spectral envelope, calculate the frequency change difference between adjacent time frames and filter out frequency points that exceed the preset judgment threshold, and construct the reconstructed spectrum set; S5: Based on the reconstructed spectrum set, the energy of the non-reconstructed frequency band is attenuated by a fixed ratio and an inverse short-time Fourier transform is performed to generate the loudspeaker driving waveform.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of an AI audio processor according to any one of claims 1 to 8.