A voice box control method, system, device and storage medium
By performing short-time power feature calculation and long-time power change trajectory analysis on the speech signal, and combining adaptive time-domain smoothing technology and error feedback algorithm, a high-precision voice control signal is generated, which solves the problem of inaccurate response in existing voice interaction systems and improves the system's adaptability and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN FRIDA LCD CO LTD
- Filing Date
- 2025-09-30
- Publication Date
- 2026-06-05
AI Technical Summary
Existing voice interaction systems struggle to respond accurately to rapid changes in voice signals, resulting in output control strategies failing to fully reflect the dynamic characteristics of the voice signal, thus affecting the system's adaptability and user experience.
By acquiring real-time intensity change monitoring streams, short-term power feature calculations and long-term power change trajectory analysis are performed. Combined with adaptive time-domain smoothing technology, fine correction is carried out to generate an optimized time-domain feature dataset. Then, through real-time error feedback and environmental adaptive correction algorithms, secondary adjustments are made to finally generate a high-precision voice control signal.
It achieves accurate response to voice signals, improves the robustness and accuracy of the voice control system under noise interference, and provides a smoother and more natural interactive experience.
Smart Images

Figure CN121260154B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech signal processing technology, and in particular to a speech box control method, system, device and storage medium. Background Technology
[0002] Currently, voice interaction technology, as an important branch of artificial intelligence, plays a crucial role in smart homes, in-vehicle systems, and wearable devices. Its core lies in achieving natural and smooth human-computer interaction through precise processing of voice signals. With the widespread adoption of smart devices, users are placing higher demands on the response speed and interactive experience of voice control systems, making the accuracy and dynamic adaptability of voice interaction a key focus of technological development.
[0003] In existing technologies, many voice control methods rely on a single signal processing approach when handling dynamic voice signals, such as based on a fixed sampling frequency or simple power analysis. However, this approach struggles to adapt to rapid changes in voice input intensity, resulting in inaccurate output responses. Because voice signal intensity changes are instantaneous and diverse, the system must be able to perceive and respond accurately in real time. Fixed processing methods fail to capture the dynamic trends of voice signals on both short- and long-term scales, making the system prone to response lag or over-adjustment when faced with rapidly changing voice input. More critically, existing methods lack refined processing of temporal characteristics when analyzing voice power at multiple levels, causing the output control strategy to fail to fully reflect the dynamic characteristics of the voice signal, thus affecting the system's adaptability and user experience.
[0004] In summary, existing technologies suffer from insufficient accuracy in voice interaction systems. Summary of the Invention
[0005] This invention provides a voice box control method, system, device, and storage medium to solve the technical problem of insufficient accuracy in voice interaction systems.
[0006] Firstly, in order to solve the above-mentioned technical problems, the present invention provides a voice box control method, comprising:
[0007] Acquire real-time intensity change monitoring stream and initial intensity change data stream, perform short-time power characteristic calculation on the initial intensity change data stream, and obtain a preliminary short-time power trend curve;
[0008] Based on the preliminary short-term power trend curve, the long-term power change trajectory is calculated. If the long-term power change trajectory exceeds the preset dynamic sampling threshold, fine-tuning is performed to obtain an optimized time-domain feature dataset.
[0009] The optimized time-domain feature dataset is used to calibrate the output parameters to obtain the initial response control pulse signal;
[0010] Calculate the signal matching error between the real-time intensity change monitoring stream and the initial response control pulse signal. If the error exceeds the preset error range, adjust the initial response control pulse signal a second time to obtain the final response control pulse signal.
[0011] The parameters of the final response control pulse signal are reconfigured to obtain the updated short-time power trend curve;
[0012] The updated short-term power trend curve is subjected to time-domain feature correction and frequency optimization step adjustment to obtain the voice box interactive control pulse signal.
[0013] Secondly, the present invention provides a voice box control system, comprising:
[0014] The data acquisition module acquires real-time intensity change monitoring stream and initial intensity change data stream, performs short-time power characteristic calculation on the initial intensity change data stream, and obtains a preliminary short-time power trend curve.
[0015] The time-domain feature optimization module calculates the long-term power change trajectory based on the preliminary short-term power trend curve. If the long-term power change trajectory exceeds the preset dynamic sampling threshold, it performs fine-tuning correction to obtain an optimized time-domain feature dataset.
[0016] The pulse generation module calibrates the output parameters of the optimized time-domain feature dataset to obtain the initial response control pulse signal;
[0017] The secondary adjustment module calculates the signal matching error between the real-time intensity change monitoring stream and the initial response control pulse signal. If the error exceeds the preset error range, the initial response control pulse signal is adjusted secondary to obtain the final response control pulse signal.
[0018] The parameter reconfiguration module reconfigures the parameters of the final response control pulse signal to obtain an updated short-time power trend curve.
[0019] The iterative control module performs time-domain feature correction and frequency optimization step adjustment on the updated short-time power trend curve to obtain the voice box interactive control pulse signal.
[0020] Thirdly, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement a voice box control method as described in any one of the above.
[0021] Fourthly, the present invention also provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform any of the above-described voice box control methods.
[0022] Compared with the prior art, the present invention has the following beneficial effects:
[0023] (1) This invention performs short-time power feature calculation and long-time power change trajectory analysis on the initial intensity change data stream, and uses adaptive temporal smoothing technology to finely correct trajectories that exceed the dynamic sampling threshold, which can effectively filter out instantaneous changes and distortions in the speech signal. This multi-level analysis and correction method combining short-time and long-time can generate an optimized temporal feature dataset that accurately reflects the dynamic characteristics of speech, laying the foundation for the subsequent generation of high-precision control signals and overcoming the problem that existing technologies cannot adapt to rapid changes in speech intensity.
[0024] (2) This invention compares the initial response control pulse signal with the intensity change monitoring stream for error. When the error exceeds the preset range, a secondary adjustment is performed using an environmental adaptive correction algorithm. This closed-loop adjustment mechanism based on real-time error feedback can dynamically compensate the control signal according to actual working conditions such as background noise and dynamic response delay, so that the final generated response signal can actively adapt to the complex and ever-changing acoustic environment, significantly improving the robustness and accuracy of voice control under noise interference.
[0025] (3) This invention feeds back the final response signal to the front end, constructing an intelligent decision-making closed loop driven by artificial intelligence and possessing true learning capabilities. In the parameter reconfiguration stage, by introducing a reinforcement learning agent, the system can learn an optimal dynamic acquisition strategy oriented towards long-term benefits. It no longer relies on fixed iterative formulas, but instead learns online to intelligently explore the optimal combination of acquisition parameters (such as sampling step size, frequency range, etc.) under different system states, in order to continuously maximize the stability of future data streams, thus achieving forward-looking and optimal decision-making in the acquisition stage. In the frequency step size fine-tuning stage, by adopting the gradient boosting decision tree (GBDT) regression model, the system can accurately predict the frequency step size that best matches the current voice command type and signal characteristics. This achieves continuous iterative improvement and long-term stability of control performance, bringing a smoother and more natural interactive experience. Attached Figure Description
[0026] Figure 1 This is a schematic flowchart of a voice box control method provided in the first embodiment of the present invention;
[0027] Figure 2 This is a schematic diagram of a voice box control system provided in the second embodiment of the present invention. Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] Reference Figure 1 The first embodiment of the present invention provides a voice box control method, including the following steps:
[0030] S11, acquire the real-time intensity change monitoring stream and the initial intensity change data stream, perform short-time power characteristic calculation on the initial intensity change data stream, and obtain a preliminary short-time power trend curve;
[0031] S12, Based on the preliminary short-term power trend curve, calculate the long-term power change trajectory. If the long-term power change trajectory exceeds the preset dynamic sampling threshold, perform fine correction to obtain an optimized time-domain feature dataset.
[0032] S13, perform output parameter calibration on the optimized time-domain feature dataset to obtain the initial response control pulse signal;
[0033] S14, calculate the signal matching error between the real-time intensity change monitoring stream and the initial response control pulse signal. If it exceeds the preset error range, adjust the initial response control pulse signal a second time to obtain the final response control pulse signal.
[0034] S15, perform parameter reconfiguration on the final response control pulse signal to obtain the updated short-time power trend curve;
[0035] S16, perform time-domain feature correction and frequency optimization step adjustment on the updated short-time power trend curve to obtain the voice box interactive control pulse signal.
[0036] In step S11, real-time intensity change monitoring stream and initial intensity change data stream are acquired, and short-time power characteristics are calculated on the initial intensity change data stream to obtain a preliminary short-time power trend curve, including:
[0037] The raw sound data stream is obtained by continuously acquiring voice input signals through the acquisition module.
[0038] Background noise separation is performed on the original sound data stream to obtain decibel intensity distribution data, and the decibel intensity distribution data is segmented to obtain segmented decibel intensity distribution data;
[0039] If the fluctuation of the segmented decibel intensity distribution data exceeds the preset fluctuation threshold, a smoothing correction is performed to obtain the initial intensity change data stream.
[0040] The initial intensity change data stream is monitored in real time. If an interruption in the data stream is detected, interpolation is performed to complete the data stream and obtain a complete monitoring stream. The stability of the complete monitoring stream is then monitored to obtain a real-time intensity change monitoring stream.
[0041] The original audio data stream is subjected to multi-level sampling and segmentation to obtain a sub-segment sequence;
[0042] Short-time power calculation is performed on each sub-segment in the sub-segment sequence to obtain a power feature sequence. If there is a sub-segment in the power feature sequence that exceeds a preset power feature value threshold, the fluctuation frequency of the sub-segment is processed by Fourier transform to obtain frequency distribution characteristics.
[0043] Based on the frequency distribution characteristics and the power characteristic sequence, a preliminary short-time power trend curve is obtained by curve fitting.
[0044] It should be noted that the purpose of background noise separation and segmentation processing of the raw audio data stream is to eliminate and reduce interference from non-target speech signals, and to divide the continuous data stream into logical processing units for subsequent fluctuation analysis. The specific implementation process is as follows: First, spectral subtraction is used to process the raw audio data stream, filtering out steady-state or slowly varying background noise, thereby extracting the decibel intensity distribution data that mainly reflects the energy changes of the speech signal. Next, based on the characteristics of environmental interference and data analysis requirements, data stream processing is performed, for example, by setting a fixed time window (e.g., every 100 milliseconds), cutting the continuous decibel intensity data stream into a series of independent segments, obtaining segmented decibel intensity distribution data, which facilitates subsequent segment-by-segment stability assessment.
[0045] For example, suppose a user issues a voice command in a smart home control scenario. The device's built-in acquisition module starts working continuously at a sampling rate of 16kHz and a sampling precision of 16 bits. Within 3 seconds of the user saying "turn on the living room lights," the acquisition module will capture a data sequence containing 3*16000=48000 sampling points. This sequence is the raw audio data stream containing all environmental sounds, including the user's voice, background fan noise, and distant television sounds. Taking the above 48000 sampling points of raw audio data stream as an example, the system first applies an adaptive filter to separate approximately 20 dB of steady-state background fan noise. After processing, a decibel intensity distribution data that mainly reflects the change in voice intensity is obtained. Subsequently, the system segments the data stream in 100-millisecond increments. Since the total duration is 3000 milliseconds, the data stream is divided into 30 consecutive data segments. Among them, the decibel intensity of segments corresponding to the user's voice may fluctuate wildly between 50 and 70 decibels, while the decibel intensity of segments corresponding to speech gaps may drop back to around 30 decibels.
[0046] It should be noted that smoothing corrections for data segments with fluctuations exceeding a preset threshold aim to eliminate abnormal data points caused by sudden noise, momentary jitter in signal acquisition equipment, or sudden environmental changes. This ensures the overall stability of the data stream and the accuracy of its trend, resulting in a more reliable initial intensity change data stream. The specific implementation process is as follows: The system first sets a preset fluctuation threshold. Then, it detects the fluctuation amplitude (variance and peak-to-valley difference) of each data segment. If the fluctuation of a segment exceeds the threshold, a Gaussian filter is used to correct that segment, weakening the impact of extreme values and making its trend smoother.
[0047] For example, suppose the preset fluctuation threshold is 0.2 units. During real-time monitoring, the system detects that the intensity value of a data segment abruptly changes from 4.8 units to 5.2 units, with a fluctuation of 0.4 units exceeding the preset threshold. At this point, the system triggers a smoothing correction mechanism and calls an environmental adaptive correction algorithm. This algorithm combines the current environmental noise level (e.g., 3.1 dB) and temperature fluctuation (e.g., 2.5 degrees Celsius), and through weighted calculation (based on preset weights, such as: error correction value = 0.6 × noise influence + 0.4 × temperature influence) to derive a comprehensive correction coefficient. This coefficient is used to adjust the signal baseline, correcting the abnormal fluctuation point to a more reasonable value, thereby forming a stable and reliable initial intensity change data stream.
[0048] It should be noted that real-time monitoring and processing of the initial intensity change data stream aims to ensure its continuity and stability, preventing data loss or transmission delays from affecting the accuracy of the final analysis. The implementation process is as follows: First, the intensity change data stream is continuously monitored. Once data point loss or transmission interruption is detected, a time series prediction model is immediately activated to estimate and fill in the missing data based on the data trends before and after the interruption point, forming a complete monitoring stream. Subsequently, the rate of change is checked to ensure it remains within a preset stable range, ultimately outputting a high-quality, continuous, and stable real-time intensity change monitoring stream.
[0049] For example, when monitoring the initial intensity change data stream, the system detects a signal transmission delay of 0.03 seconds, exceeding the standard delay threshold of 0.01 seconds, which is considered a brief data stream interruption in the system. To compensate, the system immediately employs a time series prediction model (ARIMA model, with parameters p=2, d=1, q=1) to compensate for the delay and predicts that the delay of the next cycle will be shortened to 0.02 seconds, thereby generating an adjusted timestamp sequence to complete the data. Next, the system performs a stability check on the completed data stream to ensure that its overall fluctuation error is controlled within, for example, 0.1 units, thus confirming it as a stable and usable real-time intensity change monitoring stream.
[0050] It should be noted that multi-level sampling and segmentation of the raw audio data stream are performed to analyze the characteristics of the speech signal at different time scales and decompose the long data stream into basic units that are easier to perform refined power calculations and feature extraction. The implementation process is as follows: First, the raw signal is examined at different sampling rates and time resolutions using a layered acquisition module to capture macroscopic and microscopic signal changes. Then, a time-domain segmentation method is used to cut the entire raw audio data stream into multiple continuous, potentially overlapping small segments ("frames") according to a preset, fixed time interval (i.e., "frame length"). The ordered set of these frames constitutes the sub-segment sequence.
[0051] For example, for a raw audio data stream with a sampling rate of 16kHz, the system sets the frame length for analysis to 25 milliseconds and the frame shift to 10 milliseconds (meaning there is a 15-millisecond overlap between two adjacent sub-segments). Accordingly, a 1-second audio data stream (containing 16,000 samples) will be divided into approximately 100 sub-segments. Each sub-segment is 25 milliseconds long, corresponding to 400 samples (16,000 * 0.025 = 400). These sub-segments are arranged chronologically, forming a sub-segment sequence containing 100 elements, preparing for subsequent segment-by-segment short-time power calculations.
[0052] It should be noted that performing short-time power calculation and Fourier transform on the sub-segment sequence is to quantify the energy intensity of the speech signal within each extremely short time interval and to conduct in-depth frequency domain analysis on segments with particularly high energy to distinguish whether they are caused by effective speech or burst noise. The implementation process is as follows: First, a segmented power extraction technique (calculating the sum of squares of the signal within each sub-segment) is used to calculate the short-time power of each sub-segment in the sequence, thus obtaining a power feature sequence corresponding one-to-one with the sub-segment sequence. Then, a preset power feature value threshold is set, and the power feature sequence is iterated. Once the power value of a sub-segment exceeds the threshold, it is considered that the segment has significant energy fluctuations, and a Fourier transform (FFT) is immediately applied to convert it from a time-domain signal to a frequency-domain signal, analyzing its frequency components to obtain frequency distribution characteristics.
[0053] For example, the system calculates a power feature sequence consisting of 5 sub-segments, with power values (in milliwatts) of P=[2.5,3.1,2.8,3.5,2.9]. Assume the preset power feature threshold is 3.0 milliwatts. The system detects that the power values (3.1 and 3.5) of the 2nd and 4th sub-segments exceed the threshold. For the 4th sub-segment with a power of 3.5 milliwatts, the system performs a Fourier transform on its corresponding original audio signal. Analysis reveals that the frequency fluctuation range of its input signal is between 10 and 15 Hz, exceeding the stable range of 12 Hz, and the dominant frequency component is 13.2 Hz. This frequency distribution characteristic will be recorded for subsequent comprehensive judgment and curve fitting.
[0054] It should be noted that curve fitting based on frequency distribution characteristics and power feature sequences aims to integrate time-domain energy information (power sequence) and frequency-domain component information (frequency features) to generate a preliminary short-time power trend curve that accurately and smoothly reflects the changes in real speech energy over time. The implementation process is as follows: First, the power feature sequence is used as the basic data points. Then, the previously obtained frequency distribution characteristics are used to correct or weight these data points. For example, if a high-power segment is identified as noise by the Fourier transform, its weight in the fitting can be appropriately reduced. Finally, a linear regression algorithm is used to generate a continuous and smooth curve between the corrected power data points; this curve is the preliminary short-time power trend curve.
[0055] For example, the system uses a power feature sequence P=[2.5,3.1,2.8,3.5,2.9] as a basis, combined with analyzed frequency distribution characteristics (e.g., identifying the 3.5 mW power point as primarily generated by non-speech frequency howling). Before fitting, the system first corrects the data, using a sliding window to calculate the local mean power, resulting in a smoothed sequence [2.96,3.26,3.06,3.36,3.16]. Subsequently, a linear regression algorithm is used to fit this corrected sequence, calculating a trend line, for example, obtaining a trend slope k=0.05, indicating that despite fluctuations, the overall power shows a slow upward trend over time. This line generated by the algorithm, reflecting the core trend, is the final preliminary short-term power trend curve.
[0056] In step S12, based on the preliminary short-term power trend curve, the long-term power change trajectory is calculated. If the long-term power change trajectory exceeds a preset dynamic sampling threshold, fine-tuning is performed to obtain an optimized time-domain feature dataset, including:
[0057] The preliminary short-time power trend curve is segmented in the time domain to obtain the power characteristic distribution.
[0058] Based on the preset power trend offset data, the power characteristic distribution is corrected for power value. If the corrected power value exceeds the preset power threshold, it is marked as an abnormal fluctuation point, and the initial fluctuation range is obtained.
[0059] Based on the initial fluctuation range, the continuity features of the long-term power change trajectory are extracted, and the continuity features are denoised to obtain the long-term power change trajectory.
[0060] The long-term power change trajectory is compared with a preset dynamic sampling threshold. If the fluctuation amplitude of the long-term power change trajectory exceeds the dynamic sampling threshold, it is marked as an abnormal power point, and an abnormal power detection result is obtained.
[0061] Based on the abnormal power detection results, the smoothing parameters of the marked abnormal power points are adjusted to obtain a smoothed signal sequence, and time-domain features are extracted from the smoothed signal sequence to obtain a time-domain feature dataset.
[0062] The time-domain feature dataset is corrected to obtain an optimized time-domain feature dataset.
[0063] It should be noted that the purpose of performing time-domain segmentation on the preliminary short-term power trend curve is to decompose the continuous, macroscopic power trend curve into a series of discrete, statistically significant time windows, in order to analyze the specific distribution of power characteristics within different time periods. The implementation process is as follows: First, a time-domain sliding window algorithm is used, setting a fixed-length time window and allowing this window to slide across the preliminary short-term power trend curve with a certain step size. For each data segment covered by the window, the system calculates its key statistical indicators, such as the average power, variance, and peak power, thereby transforming a continuous curve into a dataset containing the power characteristic distribution within each time period.
[0064] For example, suppose the initial short-term power trend curve is a continuous data record of power changes over 24 hours, sampled every minute, resulting in 1440 data points. The system uses a time-domain sliding window algorithm for processing, setting the window width to 10 minutes and the step size to 5 minutes. The first window covers the data from minute 1 to minute 10, and the system calculates the average power for these 10 minutes to be 52 watts, with a power variance of 3.5. Subsequently, the window slides forward 5 minutes to cover the data from minute 6 to minute 15, and the average power and variance are calculated again. This process is repeated until the window has covered all the data, ultimately resulting in a power characteristic distribution consisting of 287 ((1440-10) / 5+1) window statistics.
[0065] It should be noted that the correction based on preset power trend offset data is to remove known, normal systematic deviations (such as trend drift caused by equipment aging, periodic load changes, etc.) from the measurement data before analyzing fluctuations, thereby more accurately identifying true, unexpected abnormal fluctuations. The implementation process is as follows: The system pre-stores a set of power trend offset data, which describes the normal power drift under different conditions. When analyzing the power characteristic values of each time window, the system first queries the corresponding offset data based on the current conditions (such as time, equipment status) and performs a weighted correction on the original power value. Then, the corrected power value is compared with a preset power threshold. If it exceeds the threshold, the time window is marked as an "abnormal fluctuation point." The set of all marked points constitutes the initial fluctuation range.
[0066] For example, in the above example, the calculated average power for a 10-minute window is 85 watts. The system queries the preset power trend offset database and learns that during this period, due to equipment preheating, the normal power offset value is +5 watts. The system then corrects the measured value of 85 watts to 80 watts (85-5). Assuming the preset power threshold is 75 watts, since the corrected value of 80 watts still exceeds 75 watts, the system marks this time point as an abnormal fluctuation point. By performing this operation on all windows, the system identifies which time periods have potential anomalies, forming a preliminary fluctuation range.
[0067] It should be noted that the purpose of extracting and denoising the continuous features of long-term power variation trajectories is to filter out high-frequency noise and random disturbances in the data, restoring a smooth curve that truly reflects the long-term trend of power variation. The implementation process is as follows: First, the system connects all power value points within the initial fluctuation range in chronological order, forming an original continuous feature curve that may contain spikes and noise. Next, a weighted moving average filtering algorithm is used to process this original curve. The filtering algorithm replaces the original value of a point with the weighted average of the points and their neighbors, effectively reducing the impact of sudden noise and obtaining a smooth, clean long-term power variation trajectory.
[0068] For example, the system extracts a continuous power curve containing several noise peaks from the initial fluctuation range. To denoise, the system employs a weighted moving average filtering algorithm, setting the window width to 5 data points and the preset weighting coefficients to [0.1, 0.2, 0.4, 0.2, 0.1]. For any point on the curve with a power value of P(t), its smoothed new value P'(t) will be calculated as: P'(t) = 0.1 × P(t-2) + 0.2 × P(t-1) + 0.4 × P(t) + 0.2 × P(t+1) + 0.1 × P(t+2). By performing this rolling calculation on the entire sequence, noise in the original curve is effectively suppressed, ultimately generating a long-term power change trajectory that clearly shows the main trend.
[0069] It should be noted that comparing the long-term power variation trajectory with a preset dynamic sampling threshold aims to quantitatively assess the long-term power fluctuation amplitude and formally and decisively identify key points where the fluctuation intensity exceeds the system's acceptable range. The implementation process is as follows: The system first calculates the fluctuation amplitude of the long-term power variation trajectory by calculating the rate of change of the trajectory's slope. Then, the system compares the calculated fluctuation amplitude with a preset "dynamic sampling threshold," which represents the maximum fluctuation tolerance the system can tolerate for stable operation. Once the trajectory's fluctuation amplitude exceeds this threshold, the corresponding power point or time period is marked as an "abnormal power point," and the set of all these marks constitutes the final abnormal power detection result.
[0070] For example, when analyzing a 24-hour long-term power variation trajectory, the system sets a preset dynamic sampling threshold of ±10% for fluctuation. By calculating the rate of change of the trajectory's slope, the system finds that between 3:00 AM and 3:30 AM, the slope changes drastically, reaching a fluctuation of -15%. Since the absolute value of 15% exceeds the preset 10% threshold, the system marks the power data points within this period as abnormal power points. This marking result (i.e., the time period in which the fluctuation exceeded the threshold) is the abnormal power detection result output by the system, providing a clear target for subsequent fine-tuning.
[0071] It should be noted that the targeted smoothing and feature extraction of outliers based on abnormal power detection results aims to correct only the confirmed drastic fluctuations without affecting normal data, and to extract standardized quantitative features from the corrected high-quality signal for further analysis. The implementation process is as follows: The system first locates all data segments marked as "abnormal power points." Then, an adaptive temporal smoothing technique is used, dynamically adjusting the smoothing algorithm parameters (window size and weights, where the window size W and fluctuation amplitude A satisfy: W = W0 × (1 + α|A|), where W0 is the baseline window and α is the adjustment coefficient (preset to 0.5)) based on the local fluctuation characteristics of the signal. This process processes these outliers to obtain a smoothed signal sequence. Finally, the system extracts a series of key temporal features from this smoothed sequence, such as mean, peak value, and standard deviation, and combines these feature values to form a temporal feature dataset.
[0072] For example, based on the abnormal power detection results, the system locates the abnormal power point between 3:00 AM and 3:30 AM. The system then activates adaptive temporal smoothing technology, dynamically adjusting the standard 5-minute smoothing window to 10 minutes to more effectively smooth out this period of sharp fluctuations. After processing, a locally smoother signal sequence is obtained. Next, the system performs temporal feature extraction on this optimized complete sequence containing 1440 data points, calculating the daily average power to be 50.3 watts, the peak power to be 75.6 watts, and the power standard deviation to be 4.8. These calculated values {average power: 50.3, peak power: 75.6, power standard deviation: 4.8} constitute a structured temporal feature dataset.
[0073] It should be noted that the purpose of the final data correction on the time-domain feature dataset is to compensate for systematic distortions that may be introduced during the entire signal processing flow, ensuring that the final output feature data has the highest accuracy and fidelity, and can accurately reflect the true state of the physical world. The implementation process is as follows: the system calculates the distortion ratio or mean square error of the signal by comparing the smoothed signal with the original signal or an ideal reference signal. Then, based on this quantified distortion, a preset correction model (based on a linear function fitted by least squares) is applied to fine-tune each feature value in the time-domain feature dataset, thereby offsetting the effects of distortion and obtaining the final optimized time-domain feature dataset.
[0074] For example, the system compares the smoothed signal sequence with the original signal sequence and calculates the mean square error between them, resulting in a signal distortion ratio of 3.2%. To correct this distortion, the system applies a correction model based on the least squares method. For instance, the correction formula for the average power is: optimized average power = 1.015 * original average power - 0.2. Substituting the average power of 50.3 watts calculated from the time-domain feature dataset into the formula, the optimized average power is obtained as 50.85 watts. After similar corrections are applied to all other features in the dataset, the distortion ratio is recalculated and reduced to below 1.5%, meeting the system's accuracy requirements. This finally corrected feature set is the optimized time-domain feature dataset that can be directly used for training the subsequent power prediction model.
[0075] In step S13, the optimized time-domain feature dataset is subjected to output parameter calibration to obtain an initial response control pulse signal, including:
[0076] The optimized time-domain feature dataset is subjected to time-axis alignment processing to obtain initial time-axis calibration parameters;
[0077] Based on the initial time axis calibration parameters, the optimized time-domain feature dataset is amplitude adjusted, and the adjusted value is compared with the preset target response value. If the deviation exceeds the preset amplitude threshold range, iterative correction is performed to obtain the final amplitude adjustment value.
[0078] Based on the final amplitude adjustment value, a pulse signal is generated to obtain an unverified initial response control pulse signal. The waveform of the unverified initial response control pulse signal is then smoothed to obtain the initial response control pulse signal.
[0079] It should be noted that the purpose of time-axis alignment processing on the optimized time-domain feature dataset is to eliminate or compensate for any potential time delay or phase difference between the acquired signal and a standard reference, ensuring that subsequent amplitude adjustments and signal generation are performed on a time-synchronized basis. This is a prerequisite for achieving precise control. The implementation process is as follows: First, the system loads a preset standard reference signal as the alignment target. This standard reference signal is a standard sine wave signal with ideal time characteristics, acting as a "standard time ruler" to calibrate the acquired signal. Then, using a cross-correlation algorithm from signal processing, the optimized time-domain feature dataset is compared with the reference signal. By calculating the correlation coefficient at different time offsets, the time point that maximizes the similarity between the two signals is found, and this time offset is determined as the initial time-axis calibration parameter.
[0080] For example, suppose we have an optimized dataset containing 1000 sampling points that records the vibration amplitude of a device. For alignment, the system introduces a standard sine wave with a frequency of 10Hz and an amplitude of 4.0 as a reference signal. By running a cross-correlation algorithm, the system finds that the cross-correlation coefficient between the acquired dataset and the reference signal reaches its maximum when the entire dataset is shifted backward by 0.02 seconds. Therefore, this 0.02-second value is determined as the initial time axis calibration parameter. Applying this parameter, the system shifts the timestamps of the entire dataset, significantly reducing the mean square error between the aligned dataset and the reference signal from 2.5 to 0.8.
[0081] It should be noted that the purpose of amplitude adjustment and iterative correction is to scale the dynamic range of the signal to a range that meets the requirements of the control system, and through closed-loop feedback, ensure that the adjusted signal can accurately approximate the preset target response value, thereby eliminating deviations. The implementation process is as follows: First, the system scales the time-aligned dataset based on an initial adjustment coefficient. Then, the adjusted data value is compared point-by-point with a preset target response value, and the deviation between the two is calculated. If the deviation exceeds a preset amplitude threshold, the system performs iterative correction. This operation recalculates and updates the adjustment coefficient using a PID control algorithm based on the magnitude and direction of the current deviation, and then performs amplitude adjustment and comparison again using the new coefficient. This "adjustment-comparison-readjustment" process is repeated until the deviation is corrected to within the threshold range; the adjustment coefficient used at this point is determined as the final amplitude adjustment value.
[0082] For example, after time alignment, the system analyzes the data and finds that the average amplitude is 3.8, while the preset target response value is 4.0. The system first sets an initial adjustment coefficient of 1.05, multiplies all data points by this coefficient, and the adjusted average amplitude becomes 3.99. At this time, the calculated deviation is 0.01 (4.0-3.99). Assuming the preset amplitude threshold is ±0.05, since 0.01 is within the threshold range, iterative correction will not be initiated, and the final amplitude adjustment value is 1.05. However, if the target response value is 4.2, the initial adjusted deviation is 0.21, exceeding the threshold, and the system will initiate iterative correction, possibly increasing the adjustment coefficient to 1.1. After recalculation, the average amplitude is found to be 4.18, and the deviation is reduced to 0.02, which meets the threshold requirement. Therefore, 1.1 becomes the final amplitude adjustment value.
[0083] It should be noted that the purpose of pulse signal generation and waveform smoothing is to convert continuous, calibrated amplitude values into discrete pulse signals that can be directly used by digital circuits or actuators, and to optimize the waveform of this signal to reduce the impact or electromagnetic interference that may be caused by high-frequency noise and sharp edges. The implementation process is as follows: First, a pulse width modulation (PWM) algorithm is used to generate a pulse signal based on the final amplitude adjustment value. The magnitude of the amplitude value is linearly mapped to the pulse duty cycle. The generated signal is a series of unverified initial response control pulse signals with ideal square wave edges. Subsequently, to avoid these sharp edges, the system performs waveform smoothing on the pulse signal, for example, by using a low-pass filter or a specific waveform shaping algorithm to make the right-angled edges of the square wave smoother, thus obtaining a final stable and smooth initial response control pulse signal.
[0084] For example, the system generates a pulse signal using a pulse width modulation (PWM) algorithm based on the final calibrated amplitude data sequence. The reference frequency of the pulse signal is set to 50Hz. The algorithm linearly maps the calibrated amplitude values (e.g., ranging from -6.0 to +6.0) to a duty cycle (0% to 100%). An amplitude value of +3.0 is converted into a pulse with a 75% duty cycle. Thus, the entire amplitude data sequence is converted into a series of binary pulses with continuously varying duty cycles. This original pulse sequence is then fed into a digital filter for waveform smoothing, changing the rising and falling edges of each pulse from vertical to a gentle slope, effectively reducing high-order harmonics in the signal and ultimately outputting a high-quality initial response control pulse signal that meets the control requirements.
[0085] In step S14, the signal matching error between the real-time intensity change monitoring stream and the initial response control pulse signal is calculated. If it exceeds a preset error range, the initial response control pulse signal is adjusted a second time to obtain the final response control pulse signal, including:
[0086] The real-time intensity change monitoring stream is extracted, and the extracted data stream is decomposed in the frequency domain to obtain a first set of frequency features;
[0087] If the matching error between the first frequency feature set and the initial response control pulse signal exceeds a preset error range, then environmental adaptive correction is performed on the first frequency feature set to obtain a second frequency feature set.
[0088] The dynamic response delay and input fluctuation frequency in the second frequency feature set are obtained, and the dynamic response delay and the input fluctuation frequency are adjusted a second time to obtain the first control pulse signal;
[0089] Based on the first control pulse signal, the final response control pulse signal is generated through weighted average filtering calculation.
[0090] It should be noted that the purpose of data stream extraction and frequency domain decomposition of the real-time intensity change monitoring stream is to convert the real-time, continuous time-domain signal into a frequency-domain representation. This facilitates the analysis of the signal's intrinsic frequency components and provides a precise benchmark for subsequent signal matching and correction. The implementation process is as follows: First, the data segments to be analyzed are extracted from the complete real-time intensity change monitoring stream, forming a discrete time-series data stream. Next, the Fast Fourier Transform (FFT) algorithm is applied to this data stream for frequency domain analysis. This algorithm can calculate the amplitude and phase of each frequency component contained in the signal. The set of these frequencies and their corresponding amplitudes and phases constitutes the first set of frequency features.
[0091] For example, the system extracts the most recent second of data from a real-time intensity change monitoring stream, which consists of 48,000 sampling points. The system applies a Fast Fourier Transform (FFT) to these 48,000 points, converting them from a time-domain signal to a frequency-domain signal. Analysis shows that within the frequency range of 20Hz to 8kHz, the main frequency components of the signal are concentrated around 10kHz, with a power spectral density of 0.85W. There is also a power frequency noise component at 60Hz with a power spectral density of 0.1W. This dataset, containing all frequency components from 20Hz to 8kHz and their corresponding power spectral densities, constitutes the first frequency feature set.
[0092] It should be noted that the purpose of matching error judgment and environmental adaptive correction is to verify whether the initially generated control signal matches the real environmental signal, and when there is a significant deviation between the two, to actively filter out noise and interference in the real environmental signal, obtaining a purer frequency characteristic that better represents the essence of the target signal, laying the foundation for subsequent precise adjustments. The implementation process is as follows: First, the frequency characteristics of the initial response control pulse signal are compared with the first frequency feature set to calculate a matching error value. Then, this error value is compared with a preset error range. If it exceeds the range, an environmental adaptive correction algorithm, such as a Kalman filter, is activated. This algorithm combines an environmental model (such as known noise levels and types) to filter the first frequency feature set, suppressing or eliminating interference components, and outputting a corrected, cleaner second frequency feature set.
[0093] For example, the system compares the expected intensity (equivalent to energy in the frequency domain) of the initial response control pulse signal (4.8 units) with the total energy of 5.2 units calculated from the first frequency feature set of the real-time monitoring stream. It finds a matching error of 0.4 units, exceeding the preset error range of 0.2 units. The system then triggers an environmental adaptive correction mechanism. A Kalman filter is invoked, using the current ambient noise level (3.1 dB) and temperature fluctuation (2.5 degrees Celsius) as input parameters to process the first frequency feature set. The filter effectively suppresses frequency components related to ambient noise, outputting a second frequency feature set whose total energy is corrected to 4.9 units, closer to the target.
[0094] It should be noted that extracting key parameters from the corrected frequency set and performing secondary adjustments aims to fine-tune the two most crucial dynamic characteristics of the control signal—time and frequency—to generate a highly accurate intermediate control signal in both timing and frequency. The implementation process is as follows: First, the dynamic response delay of the system is calculated from the phase information of the second frequency feature set, and the main input fluctuation frequencies are determined from its amplitude spectrum. Then, for these two extracted parameters, a quadratic polynomial interpolation algorithm is used to calculate more ideal delay time and frequency values based on a preset optimization objective. The optimization objective is a quantifiable performance indicator, such as "maximizing signal stability," "minimizing the matching error with the reference model," or "maximizing output power." These parameters, after secondary adjustment, are used to define a new control signal, namely the first control pulse signal.
[0095] For example, the system analyzes the second frequency feature set and finds that the current dynamic response delay is 0.03 seconds, exceeding the standard value of 0.01 seconds; simultaneously, the dominant frequency of the input fluctuation is 13.2 Hz, which also deviates from the stable range of 12 Hz. The system initiates a secondary adjustment procedure, performing calculations using a quadratic polynomial interpolation algorithm. This quadratic polynomial interpolation process can be understood as follows: the algorithm not only considers the current value (e.g., delay of 0.03 seconds) and the final target value (delay of 0.01 seconds), but also introduces a third reference point (e.g., a standard calibration point or the previous measurement value) to construct a smooth parabolic correction path. Unlike direct linear adjustment, this parabolic path can calculate a more stable and reasonable intermediate adjustment target based on the magnitude of the current error. After calculation by this algorithm, the optimal dynamic response delay is found to be 0.015 seconds, and the most stable input fluctuation frequency is 13 Hz. Based on these two new parameters after secondary adjustment (delay of 0.015 seconds, frequency of 13 Hz), the system generates the first control pulse signal.
[0096] It should be noted that the purpose of weighted average filtering calculation based on the first control pulse signal is to transform the parameter-optimized intermediate signal into a smooth, stable final signal that can be directly output to the actuator, ensuring a smooth transition of control actions and avoiding system oscillations caused by signal abrupt changes. The implementation process is as follows: the system constructs an output waveform based on the precise parameters defined by the first control pulse signal (such as adjusted delay, frequency, and amplitude). This waveform is then fed into a weighted average filter. This filter comprehensively considers the signal values at the current moment and several past moments, calculating the final output value at the current moment through a weighted summation. This process effectively smooths out any minute jitter or discontinuities in the signal, thereby generating the final response control pulse signal.
[0097] For example, the system generates an initial waveform based on the parameters of the first control pulse signal (intensity 5.0 units, delay 0.015 seconds, frequency 13 Hz). To ensure the smoothness of the output, this waveform is fed into a 3-point weighted average filter with weight coefficients of [0.25, 0.5, 0.25]. For any point P(t) on the waveform, its final output value P_final(t) will be calculated as 0.25*P(t-1) + 0.5*P(t) + 0.25*P(t+1). Here, P(t-1) is the waveform value at the point before time t, and P(t+1) is the waveform value at the point after time t. After the entire filtering process, the final response control pulse signal output by the system not only fully meets the requirements in terms of core parameters (intensity 5.0 units, delay 0.015 seconds, frequency 13 Hz), but also has smooth waveform edges, and the overall signal matching error is successfully reduced to within 0.1 units, meeting the final requirements of system stability and accuracy.
[0098] In step S15, the parameters of the final response control pulse signal are reconfigured to obtain an updated short-time power trend curve, including:
[0099] The sampling frequency step size and frequency range of the acquisition module are adjusted according to the final response control pulse signal to obtain the adjusted frequency configuration data;
[0100] Based on the frequency configuration data, the acquisition module is reconfigured to obtain the updated sampling frequency parameters;
[0101] Using the updated sampling frequency parameters, the short-time power data stream is continuously acquired to obtain the acquisition results;
[0102] The power fluctuation value is extracted from the collected results, and it is determined whether the power fluctuation value meets the preset stability condition to obtain the preliminary results of the power analysis.
[0103] A short-time power trend curve is constructed from the preliminary results of the power analysis to obtain an updated short-time power trend curve.
[0104] It should be noted that adjusting the parameters of the acquisition module in reverse based on the final response control pulse signal aims to establish an intelligent closed-loop feedback system with self-optimization and learning capabilities. This invention designs this parameter reconfiguration stage as the core of the system's intelligent decision-making, using a pre-trained reinforcement learning (RL) agent, such as an actor-critic model, to achieve dynamic optimization of the acquisition parameters, rather than relying on fixed iterative formulas. The implementation process is as follows: First, the RL agent inputs various performance indicators of the final response control pulse signal, such as power deviation ratio, signal-to-noise ratio, and signal stability, as a composite "state." Based on the current state, the RL agent aims to output an "action" that maximizes future rewards. This action is not a single parameter but a set of optimized acquisition parameters, including but not limited to: a new sampling frequency step size, a frequency range requiring focus, and wavelet transform decomposition levels for multi-level sampling. The training objective of the RL agent is to maximize a long-term accumulated "reward" signal. The reward signal is defined as a stability indicator of the data stream acquired in the next acquisition cycle, such as the reciprocal of the power fluctuation value; that is, the smaller the signal fluctuation, the higher the reward value. In this way, the system, through continuous online learning and trial and error, intelligently grasps the complex mapping relationship between the current system state and the optimal front-end acquisition strategy, thereby acquiring high-quality data more efficiently and achieving truly adaptive control with learning capabilities.
[0105] It is worth noting that the parameter reconfiguration process is implemented through a pre-trained Actor-Critic reinforcement learning model, the specific training and usage steps of which are described below. The model is trained in a speech box environment based on historical data, aiming to teach the model an optimal acquisition parameter control strategy. The core elements of the training process include the state space, action space, reward function, and learning algorithm. The state space, as the model's input, consists of a series of quantized features that comprehensively describe the performance of the "final response control impulse signal." These features are constructed into a feature vector, including multiple dimensions such as the power deviation ratio, the signal-to-noise ratio (SNR), the quantized value of signal stability, and parameters characterizing the dynamic response delay of the signal. The action space, as the model's output, defines a series of parameter adjustment actions that the RL agent can execute. This is a multi-dimensional, continuous and discrete mixed action space, specifically including: adjusting the "sampling frequency step size" (continuous value), selecting the "key frequency range" (discrete interval selection), and setting the "wavelet transform decomposition level" (discrete integer value). To guide model learning, the system defines a reward function. After the agent performs an action (i.e., applies a new set of acquisition parameters), the system runs an acquisition cycle and evaluates the quality of the newly acquired data stream. The reward value is set as the reciprocal of the power fluctuation value of the data stream, combined with an additional term related to the improvement in signal-to-noise ratio (SNR). That is, the more stable the newly acquired signal and the higher the SNR, the higher the single-step reward for the agent. The learning algorithm is trained using the Actor-Critic algorithm. This algorithm consists of two neural networks: an Actor Network and a Critic Network. The Actor Network is responsible for learning the policy, i.e., directly outputting an optimal action (a set of acquisition parameters) based on the input state feature vector. The Critic Network is responsible for evaluating the quality of the action chosen by the Actor Network. It predicts the long-term cumulative reward that may be obtained after performing an action in the current state by learning a value function.
[0106] During training iterations, the "actor" selects an action based on the current state, and the "critic" rates this action. If the rating (i.e., the value assessment) is higher than expected, it indicates a good action, and the "actor" updates its network parameters to increase the probability of selecting that action in similar states in the future. Conversely, it decreases the probability. Through the collaborative work and continuous interaction of these two networks, the policy (i.e., the "actor" network) continuously optimizes in the direction that maximizes long-term cumulative rewards until convergence, thereby generating the final reinforcement learning model.
[0107] For example, when parameter reconfiguration is required, the system extracts and transforms the performance indicators of the real-time generated "final response control pulse signal" into a real-time state feature vector with the same format as during training. This real-time state vector is input into the Actor Network in the trained reinforcement learning model. Based on the current state, the Actor Network performs a forward propagation calculation, and its output is the optimal set of acquisition parameters for the current operating condition (e.g., {step size: 1.26kHz, range: 8kHz-12kHz, decomposition level: 4}). This set of parameters is then adopted by the system to generate a completely new acquisition module configuration scheme, thereby completing an intelligent closed-loop feedback and parameter reconfiguration.
[0108] It should be noted that the purpose of reconfiguring the parameters of the acquisition module is to apply the theoretical configuration data calculated in the previous step to the actual data acquisition hardware or software, enabling the system to immediately execute the optimized scheme in the next acquisition. The implementation process is as follows: the system sends the adjusted frequency configuration data (including the new frequency step size and range) to the control interface of the acquisition module. After receiving these instructions, the acquisition module updates its internal operating parameters. Furthermore, based on the signal accuracy requirements, the system can also combine the new frequency configuration to determine a more suitable multi-level sampling scheme (such as the decomposition level of wavelet transform) and update these parameters accordingly. All these parameters successfully loaded and activated by the module constitute the updated sampling frequency parameters.
[0109] For example, the system sends the frequency configuration data {step size: 1.26kHz, range: 8kHz-12kHz} to the acquisition module. The acquisition module then updates its internal settings. Simultaneously, to more precisely analyze the newly locked frequency range, the system decides to use the Daubechies4 wavelet basis function, which has good time-frequency localization characteristics in signal processing, and increases the decomposition level of the wavelet transform from 3 to 4 levels to better separate low-frequency and high-frequency components in the signal. Once the acquisition module confirms that all parameters have been updated, its currently effective operating parameters {step size: 1.26kHz, range: 8-12kHz, decomposition level: 4} constitute the updated sampling frequency parameters.
[0110] It should be noted that the purpose of using the updated sampling frequency parameters for a new round of data acquisition is to verify the effect of the parameter adjustment and obtain the latest data on the system response under optimized conditions. This is the core execution step in the entire closed-loop iterative control process. The implementation process is as follows: the system activates the acquisition module, enabling it to operate entirely according to the updated sampling frequency parameters (including step size, range, and sampling level). The module continuously acquires data from the input signal within a specified time, converting the acquired raw analog signal into a digital data stream. This unprocessed raw data stream acquired under the new, optimized configuration is the result of this acquisition.
[0111] For example, after the parameters of the acquisition module are updated to {step size: 1.26kHz, range: 8-12kHz, decomposition level: 4}, the system immediately starts a new round of data acquisition. Assuming the system feedback period is 0.2 seconds, the acquisition module will focus on the frequency range of 8kHz to 12kHz within these 0.2 seconds, performing signal scanning and power data acquisition with a step size of 1.26kHz. All the sampled data generated during this period are combined into a data sequence, which is the acquisition result of this round, and it will be used for subsequent stability analysis.
[0112] It should be noted that the purpose of extracting and judging power fluctuation values from the acquisition results is to quickly and quantitatively evaluate the quality of newly acquired data, especially to verify whether the signal stability has been improved as expected after parameter adjustments. The implementation process is as follows: First, the system extracts the power value sequence within the key time period from the acquisition results. Then, an index representing the degree of fluctuation, i.e., the power fluctuation value (such as standard deviation), is calculated for this sequence. Finally, this calculated fluctuation value is compared with a percentage based on the system's rated power (such as 0.02W). The result of the comparison—a judgment of "meets" or "does not meet" stability—constitutes the preliminary result of the power analysis.
[0113] For example, the system analyzed the acquisition results from the most recent 0.2 seconds and extracted the power value sequence. The calculated power fluctuation value (measured by standard deviation) of this sequence was 0.018W. Assuming the system's preset stability condition is that the fluctuation value must be less than 0.02W, since the calculated 0.018W is less than the threshold of 0.02W, the system determines that the currently acquired data meets the stability condition. Therefore, the preliminary result of this power analysis is marked as "stable," indicating that the parameter adjustments in the previous round were effective.
[0114] It should be noted that the purpose of constructing the updated short-term power trend curve based on the preliminary analysis results is to formally integrate the verified and stable new data into a smooth and representative trend curve. This new curve will serve as the benchmark for the next round of control and decision-making. The implementation process is as follows: After confirming that the preliminary power analysis results are "stable," the system uses the data points collected in this study. Then, using a curve construction algorithm (combining low-frequency components obtained from multi-level sampling, weighting and summing the low-frequency components according to preset weights, for example, the original data weight is 0.35, the low-frequency component weight is 0.65, and the summation result is corrected), these discrete data points are fitted into a continuous and smooth curve. This final generated graphical or data curve, reflecting the latest system state, is the updated short-term power trend curve.
[0115] For example, after the initial power analysis results are confirmed as "stable," the system uses the power data points collected this time to construct a new trend curve. The system utilizes the previously set four-level wavelet transform to extract the low-frequency component of the signal and uses this low-frequency component to perform benchmark correction and smoothing on the original data points. In this way, minute noise in the data is filtered out, ultimately generating a smooth curve with a deviation controlled within 0.01W. This newly generated curve is stored by the system as the "updated short-time power trend curve," thus completing a full closed-loop iteration and providing a more accurate basis for signal generation and control in the next cycle.
[0116] In step S16, the updated short-time power trend curve undergoes time-domain feature correction and frequency optimization step adjustment to obtain the voice box interactive control pulse signal, including:
[0117] The input signal for speech processing and the feedback period and trend interval data of the updated short-time power trend curve are acquired.
[0118] The feedback cycle and the trend interval are compared and analyzed. If they are not synchronized, the time base is adjusted to obtain synchronized cycle interval data.
[0119] The synchronized periodic interval data is subjected to feature offset extraction to obtain feature offset. If the feature offset exceeds a preset offset threshold, parameter correction is performed to obtain the adjusted time domain feature value.
[0120] The adjusted time-domain feature values are dynamically divided and fine-tuned to obtain frequency step size data that meets the requirements;
[0121] Based on the required frequency step data, the input signal for voice processing is encoded to generate the voice box interactive control pulse signal.
[0122] It should be noted that acquiring the input signal for speech processing and the period and interval data of the new curve is intended to aggregate the two major information sources—the latest external instructions (speech signal) and the description of the system's most stable internal state (metadata of the trend curve)—at the beginning of the final control loop, preparing for subsequent synchronization and correction. The implementation process is as follows: The system first captures the real-time speech input signal that needs final processing. Simultaneously, the system queries and extracts metadata associated with the "updated short-term power trend curve" generated in the previous S15 step from pre-established data storage. This metadata specifically includes the curve's "feedback period" (i.e., the frequency of curve updates) and "trend interval" (i.e., the time span represented by each curve).
[0123] For example, suppose the user issues a new voice command, "Next song." This audio stream is captured by the system and used as the "input signal for voice processing" in this instance. Simultaneously, the system accesses its internal data storage and retrieves the attribute information of the most recently generated short-term power trend curve in S15. The query results show that the feedback period of this curve is set to 200 milliseconds, meaning the system generates such a curve every 200 milliseconds based on new data; and its trend interval is 100 milliseconds, indicating that each curve is derived from data analysis over the past 100 milliseconds.
[0124] It's important to note that comparing and synchronizing the feedback cycle and trend interval aims to ensure that the system state data used for analysis (trend interval) and the rhythm of system actions (feedback cycle) are perfectly aligned in time. This prevents control commands from being generated based on outdated or irrelevant data due to time base drift or misalignment. The process involves comparing the start and end timestamps of the extracted feedback cycle and trend interval. If a discrepancy is found in their starting points or beats, the system performs a timeline alignment operation. This operation uses one as a reference (feedback cycle) and shifts or fine-tunes the timestamp of the other, ensuring their analysis windows completely overlap, thus obtaining perfectly synchronized cycle interval data.
[0125] For example, the feedback period obtained by the system is from time point t=1000ms to t=1200ms, while the latest trend interval data corresponds to time points t=990ms to t=1090ms. The system detects a 10ms delay between the two, indicating a lack of synchronization. At this point, a time axis alignment operation is triggered. Using the feedback period of t=1000ms as a baseline, it adjusts the trend interval analysis window to t=1000ms to t=1100ms, thus ensuring that the data used for subsequent analysis strictly corresponds to the current control period. This adjusted interval [1000ms, 1100ms] is the synchronized period interval data.
[0126] It should be noted that extracting and correcting the feature offset from the synchronized data serves as a final, ultra-precise fine-tuning process. This captures and corrects any subtle but persistent systematic deviations that may exist in the system, ensuring the final output time-domain characteristics achieve the highest possible accuracy. The process involves: the system performing waveform analysis. Within the synchronized periodic intervals, a key time-domain characteristic of the actual signal (such as average amplitude) is compared to a theoretically ideal value. The difference between the two is calculated, which is the feature offset. This offset is then compared to a very small, preset offset threshold. If the threshold is exceeded, parameter correction is performed. Based on the magnitude and direction of the offset, a precise correction value is calculated and applied to the original time-domain characteristics to obtain the final adjusted time-domain feature value.
[0127] For example, within a 100ms periodic interval after synchronization, the system analyzes and finds the average signal amplitude to be 1.008 units, while the theoretical ideal value for this stage should be 1.000 units. The calculated characteristic offset is +0.008. Assuming a preset offset threshold of ±0.005, since 0.008 exceeds this threshold, the system triggers parameter correction. The correction process calculates that a correction value of -0.008 needs to be applied. After applying this correction value, the final adjusted time-domain characteristic value is 1.000, completely eliminating the minor system deviation.
[0128] It should be noted that the purpose of dynamically dividing and fine-tuning the final adjusted time-domain feature values is to transform the final, refined correction result in the time domain into a corresponding adjustment of the frequency domain control parameters (i.e., frequency step size), so that the generation of the control signal can also achieve optimal performance at the spectral level. The implementation process is as follows: the system uses a pre-trained offline Gradient Boosting Decision Trees (GBDT) regression model to replace the traditional iterative calculation formula, thereby achieving intelligent and refined fine-tuning of the frequency step size. The input to this model is a feature vector containing multi-dimensional information. Its core features include: the final corrected time-domain feature values obtained in the previous step (such as average amplitude, peak value, etc.); and the current voice command type identified by a parallel speech recognition module (e.g., encoded as category labels such as "play music," "next song," "turn up volume," etc.).
[0129] The model's output is a single, continuously optimized frequency step size (in kHz). By training on a dataset containing a large number of speech command samples and corresponding optimal frequency step size labels, the GBDT model learns the specific requirements of different command types for spectral detail. For example, simple control commands may allow for a wider step size, while music playback-related commands require a finer step size to ensure high fidelity. This approach transforms frequency step size determination from an indirect adjustment based on deviation ratios to a direct prediction based on specific interaction requirements, thus obtaining frequency step size data that ultimately meets the needs.
[0130] It is worth noting that the fine-tuning process of the frequency step size is achieved through a pre-trained Gradient Boosting Decision Trees (GBDT) regression machine learning model. The specific training and usage steps are as follows: The model's training data comes from a pre-built database of features and optimal parameters containing a large number of speech samples. Each data sample contains two parts: a quantized feature vector (as model input) consisting of "temporal feature values" and "speech command type"; and an optimal frequency step size value (as model label / output) pre-calibrated by audio engineers based on listening tests and signal fidelity analysis for that specific condition. The feature vector may include: adjusted temporal feature average amplitude, peak value, signal standard deviation, zero-crossing rate, and multiple dimensions such as the one-hot encoded speech command category (e.g., "play," "pause," "next track," etc.).
[0131] The construction of the GBDT model is an iterative, additive process. First, the model initializes a simple base learner (typically the mean of all labels). In each iteration, the algorithm calculates the residuals (i.e., errors) between the current ensemble model's predictions and the true labels. Then, the algorithm trains a new, simple decision tree to fit these residuals. This new tree is added to the ensemble model, its contribution scaled by a pre-defined learning rate. This process continues, with the model progressively and greedily optimizing in the direction of the negative gradient of the loss function by continuously fitting the residuals from the previous iteration. The iteration process continues until a pre-defined stopping condition is met (e.g., reaching the maximum number of trees, or performance on the validation set no longer improving), thus generating the final GBDT model.
[0132] For example, suppose the system recognizes the user's voice command "next song" in parallel and simultaneously acquires the finely calibrated set of temporal feature values obtained in the previous step. The system integrates this information into a feature vector and inputs it into a pre-trained GBDT regression model. The model performs forward inference based on its internal set of decision trees and directly outputs an optimal prediction value. For instance, based on the mid-to-high frequency clarity requirements of the "next song" command and the current signal characteristics, the model predicts that the frequency step size that best balances resolution and high-frequency response is 2.45 kHz. This predicted value of 2.45 kHz is then determined as the final frequency step size data that meets the requirements and is used for subsequent encoding processing. When fine-tuning of the frequency step size is required, the system extracts and converts the currently calibrated "temporal feature values" and the real-time recognized "voice command type" data into a real-time feature vector with the same format as during training. This real-time feature vector is input into the trained GBDT model. The model sequentially passes it through all the internally integrated decision trees and performs a weighted sum of the prediction results from each tree.
[0133] It's important to note that encoding the voice input based on the final determined frequency step size data aims to transform the raw analog voice signal into a final, hardware-executable digital pulse signal that has undergone comprehensive, multi-dimensional optimization and calibration. This signal represents the ultimate embodiment of the system's intelligence and adaptive capabilities. The process is as follows: the system loads the required frequency step size data into the encoder. Then, the user's original voice input signal is fed into the encoder. The encoder uses this finely tuned frequency step size as its core quantization parameter to encode the input signal, generating a binary pulse sequence. This sequence is the final output voice-to-text interaction control pulse signal.
[0134] For example, the system configures the finalized 2.45kHz frequency step size data into a pulse code modulator (PCM Encoder). The raw audio stream of the user's "next track" command is then input into this encoder. During analog-to-digital conversion and quantization, the encoder strictly adheres to the 2.45kHz step size. Ultimately, the encoder outputs a series of binary pulse signals representing the "next track" command, which have undergone time-domain correction, frequency optimization, and multiple closed-loop feedback adjustments. This signal is sent to the voice box's digital signal processor (DSP) for execution, thereby playing the next song with the highest fidelity and fastest response time.
[0135] In summary, this invention constructs a full-process control method integrating "multi-level dynamic analysis, real-time error feedback adjustment, and closed-loop self-optimization." First, it performs in-depth analysis and correction of the short-time power and long-time trajectory of the voice signal. Then, it introduces a secondary adjustment mechanism based on real-time environmental monitoring. Finally, it uses the output signal to back-optimize the front-end acquisition parameters. This solves the technical problems of low response accuracy and poor adaptability of existing voice control methods in complex dynamic environments. It realizes continuous self-optimization and accurate generation of voice control pulse signals, significantly improving the accuracy, robustness, and fluency of human-computer interaction.
[0136] Reference Figure 2 The second embodiment of the present invention provides a voice box control system, including:
[0137] The data acquisition module acquires real-time intensity change monitoring stream and initial intensity change data stream, performs short-time power characteristic calculation on the initial intensity change data stream, and obtains a preliminary short-time power trend curve.
[0138] The time-domain feature optimization module calculates the long-term power change trajectory based on the preliminary short-term power trend curve. If the long-term power change trajectory exceeds the preset dynamic sampling threshold, it performs fine-tuning correction to obtain an optimized time-domain feature dataset.
[0139] The pulse generation module calibrates the output parameters of the optimized time-domain feature dataset to obtain the initial response control pulse signal;
[0140] The secondary adjustment module calculates the signal matching error between the real-time intensity change monitoring stream and the initial response control pulse signal. If the error exceeds the preset error range, the initial response control pulse signal is adjusted secondary to obtain the final response control pulse signal.
[0141] The parameter reconfiguration module reconfigures the parameters of the final response control pulse signal to obtain an updated short-time power trend curve.
[0142] The iterative control module performs time-domain feature correction and frequency optimization step adjustment on the updated short-time power trend curve to obtain the voice box interactive control pulse signal.
[0143] It should be noted that the voice box control system provided in this embodiment of the invention is used to execute all the process steps of the voice box control method in the above embodiment. The working principle and beneficial effects of the two are one-to-one, so they will not be described again.
[0144] This invention also provides an electronic device. The electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a voice box control program. When the processor executes the computer program, it implements the steps described in the various embodiments of the voice box control method above, for example... Figure 1 The step S11 shown. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above system embodiments, such as the data acquisition module.
[0145] Wherein, if the modules / units integrated in the electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0146] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the system embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0147] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A voice box control method, characterized in that, include: Acquire real-time intensity change monitoring stream and initial intensity change data stream, perform short-time power characteristic calculation on the initial intensity change data stream, and obtain a preliminary short-time power trend curve; Based on the preliminary short-term power trend curve, the long-term power change trajectory is calculated. If the long-term power change trajectory exceeds the preset dynamic sampling threshold, fine-tuning is performed to obtain an optimized time-domain feature dataset. The optimized time-domain feature dataset is used to calibrate the output parameters to obtain the initial response control pulse signal; Calculate the signal matching error between the real-time intensity change monitoring stream and the initial response control pulse signal. If the error exceeds the preset error range, adjust the initial response control pulse signal a second time to obtain the final response control pulse signal. The parameters of the final response control pulse signal are reconfigured to obtain the updated short-time power trend curve; The updated short-term power trend curve is subjected to time-domain feature correction and frequency optimization step adjustment to obtain the voice box interactive control pulse signal.
2. The voice box control method according to claim 1, characterized in that, The process of acquiring real-time intensity change monitoring stream and initial intensity change data stream, and performing short-time power characteristic calculation on the initial intensity change data stream to obtain a preliminary short-time power trend curve includes: The raw sound data stream is obtained by continuously acquiring voice input signals through the acquisition module. Background noise separation is performed on the original sound data stream to obtain decibel intensity distribution data, and the decibel intensity distribution data is segmented to obtain segmented decibel intensity distribution data; If the fluctuation of the segmented decibel intensity distribution data exceeds the preset fluctuation threshold, a smoothing correction is performed to obtain the initial intensity change data stream. The initial intensity change data stream is monitored in real time. If an interruption in the data stream is detected, interpolation is performed to complete the data stream and obtain a complete monitoring stream. The stability of the complete monitoring stream is then monitored to obtain a real-time intensity change monitoring stream. The original audio data stream is subjected to multi-level sampling and segmentation to obtain a sub-segment sequence; Short-time power calculation is performed on each sub-segment in the sub-segment sequence to obtain a power feature sequence. If there is a sub-segment in the power feature sequence that exceeds a preset power feature value threshold, the fluctuation frequency of the sub-segment is processed by Fourier transform to obtain frequency distribution characteristics. Based on the frequency distribution characteristics and the power characteristic sequence, a preliminary short-time power trend curve is obtained by curve fitting.
3. The voice box control method according to claim 1, characterized in that, The step involves calculating the long-term power change trajectory based on the preliminary short-term power trend curve. If the long-term power change trajectory exceeds a preset dynamic sampling threshold, fine-tuning is performed to obtain an optimized time-domain feature dataset, including: The preliminary short-time power trend curve is segmented in the time domain to obtain the power characteristic distribution. Based on the preset power trend offset data, the power characteristic distribution is corrected for power value. If the corrected power value exceeds the preset power threshold, it is marked as an abnormal fluctuation point, and the initial fluctuation range is obtained. Based on the initial fluctuation range, the continuity features of the long-term power change trajectory are extracted, and the continuity features are denoised to obtain the long-term power change trajectory. The long-term power change trajectory is compared with a preset dynamic sampling threshold. If the fluctuation amplitude of the long-term power change trajectory exceeds the dynamic sampling threshold, it is marked as an abnormal power point, and an abnormal power detection result is obtained. Based on the abnormal power detection results, the smoothing parameters of the marked abnormal power points are adjusted to obtain a smoothed signal sequence, and time-domain features are extracted from the smoothed signal sequence to obtain a time-domain feature dataset. The time-domain feature dataset is corrected to obtain an optimized time-domain feature dataset.
4. The voice box control method according to claim 1, characterized in that, The step of calibrating the output parameters of the optimized time-domain feature dataset to obtain the initial response control pulse signal includes: The optimized time-domain feature dataset is subjected to time-axis alignment processing to obtain initial time-axis calibration parameters; Based on the initial time axis calibration parameters, the optimized time-domain feature dataset is amplitude adjusted, and the adjusted value is compared with the preset target response value. If the deviation exceeds the preset amplitude threshold range, iterative correction is performed to obtain the final amplitude adjustment value. Based on the final amplitude adjustment value, a pulse signal is generated to obtain an unverified initial response control pulse signal. The waveform of the unverified initial response control pulse signal is then smoothed to obtain the initial response control pulse signal.
5. The voice box control method according to claim 1, characterized in that, The step of calculating the signal matching error between the real-time intensity change monitoring stream and the initial response control pulse signal, and if it exceeds a preset error range, then performing a secondary adjustment on the initial response control pulse signal to obtain the final response control pulse signal, includes: The real-time intensity change monitoring stream is extracted, and the extracted data stream is decomposed in the frequency domain to obtain a first set of frequency features; If the matching error between the first frequency feature set and the initial response control pulse signal exceeds a preset error range, then environmental adaptive correction is performed on the first frequency feature set to obtain a second frequency feature set. The dynamic response delay and input fluctuation frequency in the second frequency feature set are obtained, and the dynamic response delay and the input fluctuation frequency are adjusted a second time to obtain the first control pulse signal; Based on the first control pulse signal, the final response control pulse signal is generated through weighted average filtering calculation.
6. The voice box control method according to claim 1, characterized in that, The step of reconfiguring the parameters of the final response control pulse signal to obtain the updated short-time power trend curve includes: The sampling frequency step size and frequency range of the acquisition module are adjusted according to the final response control pulse signal to obtain the adjusted frequency configuration data; Based on the frequency configuration data, the acquisition module is reconfigured to obtain the updated sampling frequency parameters; Using the updated sampling frequency parameters, the short-time power data stream is continuously acquired to obtain the acquisition results; The power fluctuation value is extracted from the collected results, and it is determined whether the power fluctuation value meets the preset stability condition to obtain the preliminary results of the power analysis. A short-time power trend curve is constructed from the preliminary results of the power analysis to obtain an updated short-time power trend curve.
7. The voice box control method according to claim 1, characterized in that, The updated short-time power trend curve is subjected to time-domain feature correction and frequency optimization step adjustment to obtain the voice box interaction control pulse signal, including: The input signal for speech processing and the feedback period and trend interval data of the updated short-time power trend curve are acquired. The feedback cycle and the trend interval are compared and analyzed. If they are not synchronized, the time base is adjusted to obtain synchronized cycle interval data. The synchronized periodic interval data is subjected to feature offset extraction to obtain feature offset. If the feature offset exceeds a preset offset threshold, parameter correction is performed to obtain the adjusted time domain feature value. The adjusted time-domain feature values are dynamically divided and fine-tuned to obtain frequency step size data that meets the requirements; Based on the required frequency step data, the input signal for voice processing is encoded to generate the voice box interactive control pulse signal.
8. A voice box control system, characterized in that, include: The data acquisition module acquires real-time intensity change monitoring stream and initial intensity change data stream, performs short-time power characteristic calculation on the initial intensity change data stream, and obtains a preliminary short-time power trend curve. The time-domain feature optimization module calculates the long-term power change trajectory based on the preliminary short-term power trend curve. If the long-term power change trajectory exceeds the preset dynamic sampling threshold, it performs fine-tuning correction to obtain an optimized time-domain feature dataset. The pulse generation module calibrates the output parameters of the optimized time-domain feature dataset to obtain the initial response control pulse signal; The secondary adjustment module calculates the signal matching error between the real-time intensity change monitoring stream and the initial response control pulse signal. If the error exceeds the preset error range, the initial response control pulse signal is adjusted secondary to obtain the final response control pulse signal. The parameter reconfiguration module reconfigures the parameters of the final response control pulse signal to obtain an updated short-time power trend curve. The iterative control module performs time-domain feature correction and frequency optimization step adjustment on the updated short-time power trend curve to obtain the voice box interactive control pulse signal.
9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement a voice box control method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform a voice box control method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-pulse analysis speech processing system and method
CN1153566A
Improving perceived quality of dereverberation
CN116964665A