A data processing method and related apparatus
Patent Information
- Application Number
- CN202611200502.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-07
- Publication Date
- 2026-10-09
AI Technical Summary
[0016]借由上述技术方案,本申请提供的数据处理方法及装置中,因为仿真音频信号与终端在第一使用条件下输出的音频信号相似,使得音频处理模型能够学习到终端在第一使用条件下输出的音频信号的特征,提高音频处理模型对终端输出的音频信号的处理效果。并且,在音频处理模型的第一使用条件发生变化后,从所有预设频响特征曲线中可以获取与变化后的第一使用条件匹配的第一频响特征曲线,省去采集音频信号、分析频响特征曲线等流程,节省人力成本和时间成本。并且第一特征向量可以表示一类终端共性的特征向量,由此该类终端可以使用同一个预设频响特征曲线,针对该类终端同样省去采集音频信号、分析频响特征曲线等流程,节省人力成本和时间成本。
Smart Images

Figure CN122889005A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a data processing method and related apparatus. Background Technology
[0002] With technological advancements, AI (Artificial Intelligence) models / ML (Machine Learning) models can be used as audio processing models for terminals, performing tasks such as noise reduction on the audio signals output by the terminal. However, obtaining training data is a pressing issue that needs to be addressed when training these audio processing models. Summary of the Invention
[0003] In view of the above problems, this application provides a data processing method and related apparatus to achieve the goal of saving labor and time costs. The specific solution is as follows: The first aspect of this application provides a data processing method, including: The first usage condition of the audio processing model is obtained, and the second usage condition of each preset frequency response characteristic curve is obtained. The preset frequency response characteristic curve is obtained in advance using the first feature vector under the second usage condition. The first feature vector is a feature vector representing the common characteristics of a class of terminals obtained from the second feature vector under multiple third usage conditions. The second feature vector is obtained from the third feature vector of the first audio signal output by the terminal under the third usage conditions. The second usage condition is obtained from the multiple third usage conditions. Based on the first usage condition of the audio processing model and the second usage condition of each preset frequency response characteristic curve, a first frequency response characteristic curve matching the first usage condition is obtained from all preset frequency response characteristic curves; The standard audio signal is processed using at least the first frequency response characteristic curve to obtain a simulated audio signal, which is the training data of the audio processing model.
[0004] In one implementation, the preset frequency response characteristic curve is obtained in advance using a first feature vector under the second usage condition, wherein the first feature vector is obtained based on second feature vectors under multiple third usage conditions, including: All second feature vectors are clustered to divide them into different clusters, and the usage conditions of the clusters are obtained according to the third usage conditions corresponding to the second feature vectors located in the same cluster. The usage conditions of the clusters are used to indicate the usage scenarios and / or terminal statistics corresponding to the clusters. Based on all the second feature vectors in the same cluster, obtain the feature vector located at the centroid of the cluster, and the feature vector located at the centroid of the cluster is the first feature vector; The first feature vector is processed using a preset frequency response simulation model to obtain a preset frequency response feature curve corresponding to the first feature vector. The usage conditions corresponding to the cluster to which the first feature vector belongs are used as the second usage conditions of the preset frequency response feature curve.
[0005] In one implementation, the second feature vector is obtained based on the third feature vector of the first audio signal output by the terminal under the third usage condition, including: Multiple audio data pairs are acquired for each third usage condition. The audio data pairs include a first standard audio signal and a first audio signal. The first standard audio signal and the first audio signal in the same audio data pair are used to calculate the frequency response characteristic curve under the third usage condition. Feature extraction is performed on each first audio signal to obtain a third feature vector for each first audio signal; The third feature vector of each first audio signal is processed using a preset coding model to obtain the second feature vector of each first audio signal. The preset coding model is trained based on the frequency response feature curve under the third usage condition and the second feature vector of each first audio signal.
[0006] In one implementation, obtaining a first frequency response feature curve matching the first usage condition from all preset frequency response feature curves based on the first usage condition of the audio processing model and the second usage condition of each preset frequency response feature curve includes: Obtain the matching degree between the first usage condition and each second usage condition; The preset frequency response feature curve corresponding to the second usage condition that meets the preset matching condition is used as the first frequency response feature curve.
[0007] In one implementation, the first usage condition includes a usage scenario and terminal information, and the second usage condition includes a usage scenario, the number of covered terminals, and model distribution information, wherein the model distribution information is used to indicate the terminal information corresponding to the cluster to which the second usage condition belongs; obtaining the matching degree between the first usage condition and each second usage condition includes: If the usage scenario and / or terminal information in the first usage condition are recorded in the second usage condition, the matching degree between the first usage condition and the second usage condition is determined based on at least one of the model distribution information and the number of covered terminals.
[0008] In one implementation, if the first frequency response characteristic curve is not obtained according to the first usage condition and the second usage condition, the method further includes: obtaining a second audio signal output by the terminal after the second standard audio signal is processed under the first usage condition; The frequency response characteristic curve of the terminal is obtained based on the second standard audio signal and the second audio signal; Calculate the similarity between the frequency response characteristic curve of the terminal and each preset frequency response characteristic curve; Based on the similarity between the frequency response characteristic curve of the terminal and each preset frequency response characteristic curve, and the second usage conditions, at least one first frequency response characteristic curve is determined from the preset frequency response characteristic curves.
[0009] In one implementation, after obtaining the second audio signal output by the terminal after processing the second standard audio signal under the first usage condition, the method further includes: Feature extraction is performed on the second audio signal to obtain a third feature vector of the second audio signal; The third feature vector of the second audio signal is processed using a preset encoding model to obtain the second feature vector of the second audio signal. The first usage condition corresponding to the second audio signal is the third usage condition of the second feature vector of the second audio signal. The preset frequency response feature curve and the second usage condition corresponding to the preset frequency response feature curve are updated using the second feature vector of the second audio signal and the third usage condition of the second feature vector of the second audio signal.
[0010] In one implementation, after obtaining the first frequency response characteristic curve, the method further includes: determining a first number of standard audio signals processed by the first frequency response characteristic curve and a second number of standard audio signals processed by the second frequency response characteristic curve, wherein the second frequency response characteristic curve is a preset frequency response characteristic curve that does not match the first usage conditions, and the first number is greater than the second number. The step of processing the standard audio signal using at least the first frequency response characteristic curve to obtain the simulated audio signal includes: processing the first number of standard audio signals using the first frequency response characteristic curve to obtain the first number of simulated audio signals. The second number of standard audio signals are processed using the second frequency response characteristic curve to obtain the second number of simulated audio signals.
[0011] In one implementation, determining the first number of standard audio signals processed by the first frequency response characteristic curve and the second number of standard audio signals processed by the second frequency response characteristic curve includes: A first weight of the first frequency response characteristic curve and a second weight of the second frequency response characteristic curve are determined. The first weight is greater than the second weight. Either the first weight or the second weight is used to indicate the number of standard audio signals processed by the frequency response characteristic curve corresponding to that weight. The relationship between the first weight and the second weight is used to represent the relationship between the first quantity and the second quantity.
[0012] A second aspect of this application provides a data processing apparatus, comprising: The acquisition unit is used to acquire the first usage conditions of the audio processing model and the second usage conditions of each preset frequency response characteristic curve. The preset frequency response characteristic curve is obtained in advance using the first feature vector under the second usage conditions. The first feature vector is a feature vector representing the common characteristics of a class of terminals obtained from the second feature vectors under multiple third usage conditions. The second feature vector is obtained from the third feature vector of the first audio signal output by the terminal under the third usage conditions. The second usage conditions are obtained from the multiple third usage conditions. The matching unit is configured to obtain a first frequency response feature curve that matches the first usage condition from all preset frequency response feature curves, based on the first usage condition of the audio processing model and the second usage condition of each preset frequency response feature curve; The processing unit is configured to process the standard audio signal using at least the first frequency response characteristic curve to obtain a simulated audio signal, wherein the simulated audio signal is the training data of the audio processing model.
[0013] A third aspect of this application provides a computer program product including computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the data processing method described in the first aspect or any implementation thereof.
[0014] A fourth aspect of this application provides an electronic device, including at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the data processing method of the first aspect or any implementation thereof.
[0015] The fifth aspect of this application provides a computer storage medium carrying one or more computer programs, which, when executed by an electronic device, enable the electronic device to perform the data processing method described in the first aspect or any implementation thereof.
[0016] By employing the above technical solution, the data processing method and apparatus provided in this application, because the simulated audio signal is similar to the audio signal output by the terminal under the first usage condition, enables the audio processing model to learn the characteristics of the audio signal output by the terminal under the first usage condition, thereby improving the processing effect of the audio processing model on the audio signal output by the terminal. Furthermore, after the first usage condition of the audio processing model changes, a first frequency response characteristic curve matching the changed first usage condition can be obtained from all preset frequency response characteristic curves, eliminating the need for audio signal acquisition and frequency response characteristic curve analysis, thus saving manpower and time costs. Moreover, the first feature vector can represent a common feature vector of a class of terminals, thereby allowing this class of terminals to use the same preset frequency response characteristic curve, similarly eliminating the need for audio signal acquisition and frequency response characteristic curve analysis for this class of terminals, saving manpower and time costs. Attached Figure Description
[0017] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0018] Figure 1 A flowchart of a data processing method provided in this application; Figure 2 A flowchart of another data processing method provided in this application; Figure 3 A flowchart of another data processing method provided in this application; Figure 4 A schematic diagram illustrating a scenario for the data processing method provided in this application; Figure 5 and Figure 6 The effect comparison chart provided for this application; Figure 7 A schematic diagram of the structure of a data processing device provided in this application; Figure 8 A schematic diagram of another data processing device provided in this application. Detailed Implementation
[0019] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.
[0020] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0021] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements but may include other elements not explicitly listed or inherent to such processes, methods, systems, products, or apparatus.
[0022] The audio processing model used by the terminal can be an AI model or a ML model. The terminal utilizes the AI or ML model to perform noise reduction, pitch correction, and acoustic echo cancellation (AEC) on the output audio signal. The audio processing model can obtain training data in the following ways: One approach is to use standard audio signals from a standard audio source library as training data. These standard audio signals are artificially generated or recorded by standard recording equipment, and differ from the audio signals output by the terminal. As a result, when an audio processing model trained using standard audio signals processes the audio signals output by the terminal, the processing effect of the audio processing model is poor or the processing effect of the audio processing model does not match the actual effect of the terminal.
[0023] Another approach involves using professional audio analysis software (such as Sound Check) to control the terminal to play a fixed test tone (e.g., a sweep signal), simultaneously acquiring the audio signal output from the terminal's microphone. The professional audio analysis software then analyzes the audio signal to derive the terminal's frequency response characteristic curve. This frequency response characteristic curve is simulated using a software algorithm module. Standard audio signals from a standard audio library are processed by this software algorithm module, and the processed standard audio signals are used as training data. However, the training data is obtained using the frequency response characteristic curve under a specific usage condition. A usage condition can include the usage scenario and / or terminal information, and the audio processing model trained using the training data is only applicable to that usage condition. When the usage conditions (such as the usage scenario and / or terminal information) change, it is necessary to re-acquire audio signals, analyze the frequency response characteristic curve, simulate the frequency response characteristic curve using the software algorithm module, and process the standard audio signals using the software algorithm module again. This process consumes significant manpower and time.
[0024] To address the aforementioned problems, this application provides a data processing method and apparatus. The method involves pre-obtaining preset frequency response characteristic curves under different second usage conditions. After acquiring the first usage conditions of the audio processing model, based on the first usage conditions of the audio processing model and the second usage conditions of each preset frequency response characteristic curve, a first frequency response characteristic curve matching the first usage conditions is obtained from all preset frequency response characteristic curves. At least the first frequency response characteristic curve is used to process the standard audio signal to obtain a simulated audio signal, making the simulated audio signal similar to the audio signal output by the terminal under the first usage conditions. The simulated audio signal can be used as training data for the audio processing model. Because the simulated audio signal is similar to the audio signal output by the terminal under the first usage conditions, the audio processing model can learn the characteristics of the audio signal output by the terminal under the first usage conditions, improving the processing effect of the audio processing model on the audio signal output by the terminal. Furthermore, when the first usage conditions of the audio processing model change, a first frequency response characteristic curve matching the changed first usage conditions can be obtained from all preset frequency response characteristic curves, eliminating the need for audio signal acquisition and frequency response characteristic curve analysis, thus saving manpower and time costs. Furthermore, the first feature vector can represent a common feature vector of a class of terminals. As a result, this class of terminals can use the same preset frequency response feature curve. For this class of terminals, the process of collecting audio signals and analyzing frequency response feature curves is also eliminated, saving manpower and time costs.
[0025] First, let me explain the terminology used in this application: The first usage condition can be the matching target of the audio processing model, such as the specific scenario requirements that the audio processing model hopes to simulate during training. For example, the first usage condition can be used to indicate the usage scenario and / or terminal information.
[0026] The third usage condition can be the scenario in which the terminal collects audio signals, such as a strong wind outdoors.
[0027] The second usage condition is obtained based on multiple third usage conditions, such as by clustering the third usage conditions. For example, the second usage condition can be a typical usage scenario after clustering. A typical usage scenario can be considered a frequently occurring usage scenario, such as "indoor - comfortable and gentle sound model cluster" or "outdoor - handheld noise cluster".
[0028] The second feature vector can be the feature vector of the audio signal collected by the terminal under the third usage condition. For example, it can be a low-dimensional feature vector, or it can be called a low-dimensional coded feature.
[0029] The first feature vector is obtained based on multiple second feature vectors, such as those obtained by clustering multiple second feature vectors. The first feature vector can be a feature of the centroid obtained by clustering, or simply the cluster centroid feature.
[0030] The data processing method of this application embodiment will now be described in detail with reference to the accompanying drawings. (Refer to...) Figure 1 , Figure 1 An optional flow of a data processing method provided in this application embodiment may include the following steps: S101. Obtain the first usage conditions of the audio processing model and the second usage conditions of each preset frequency response characteristic curve. The preset frequency response characteristic curve is obtained in advance using the first feature vector under the second usage conditions. The first feature vector is obtained based on the second feature vector under multiple third usage conditions. The second feature vector is obtained based on the third feature vector of the first audio signal output by the terminal under the third usage conditions. The second usage conditions are obtained based on multiple third usage conditions.
[0031] In this embodiment, the first usage condition is used to indicate the usage scenario and / or terminal information. The usage scenario can indicate the current environment of the terminal. The usage environment can be an outdoor environment or an indoor environment. An outdoor environment can be a user holding the terminal outdoors, a strong wind and noise environment outdoors, a high humidity environment outdoors, etc. An indoor environment can be a specific location where the terminal is placed indoors, an indoor music environment, etc. The terminal information is used to point to a terminal. For example, the terminal information can include the terminal's model, etc.
[0032] When the first usage condition is used to indicate the usage scenario and / or terminal information, the audio processing model is used to process the audio signal output by the terminal under the usage scenario. The audio processing model needs to learn the characteristics of the audio signal under the first usage condition. Therefore, the audio processing model needs to be trained using the audio signal under the first usage condition.
[0033] If the first usage condition indicates a usage scenario, the audio processing model is used to process the audio signal output by at least one terminal in that usage scenario. The audio processing model needs to learn the characteristics of the audio signals output by different terminals in that usage scenario, and therefore needs to be trained using the audio signals output by at least one terminal in that usage scenario. If the first usage condition indicates terminal information, the audio processing model is used to process the audio signal output by that terminal in at least one usage scenario. The audio processing model needs to learn the audio signals output by that terminal in different usage scenarios, and therefore needs to be trained using the audio signals output by that terminal in at least one usage scenario.
[0034] The third usage condition has the same function as the first usage condition; for an explanation of the third usage condition, please refer to the first usage condition. This embodiment pre-obtains the third feature vector of the first audio signal output by the terminal under the third usage condition, and uses the third feature vector to obtain the second feature vector under the third usage condition. In one implementation, the process of obtaining the second feature vector includes: obtaining multiple audio data pairs under each third usage condition; each audio data pair includes a first standard audio signal and a first audio signal; the first audio signal is the audio signal output by the terminal after processing the first standard audio signal under the third usage condition; the first standard audio signal and the first audio signal in the same audio data pair are used to calculate the frequency response characteristic curve under the third usage condition; performing feature extraction on each first audio signal to obtain the third feature vector of each first audio signal; and processing the third feature vector of each first audio signal using a preset encoding model to obtain the second feature vector of each first audio signal. The preset encoding model is trained based on the frequency response characteristic curve under the third usage condition and the second feature vector of each first audio signal.
[0035] Optionally, the first standard audio signal and the first audio signal in the same audio data pair can be processed using cross-correlation or frequency domain division to obtain the frequency response characteristic curve under the third usage condition. The third feature vector is a high-dimensional feature vector obtained by feature extraction from the first audio signal, such as extracting the Fourier transform amplitude spectrum, Mel spectrum, log-Mel energy, and linear predictive coding coefficients from the first audio signal, to represent the features of the first audio signal in the multi-scale frequency domain. The high-dimensional feature vector is encoded using a preset coding model to obtain a low-dimensional feature vector, which can be used as the second feature vector to obtain a low-dimensional representation of the first audio signal under the third usage condition.
[0036] After obtaining the second feature vector of the first audio signal under each third usage condition, the first feature vector is obtained by clustering all the second feature vectors. One implementation is as follows: all the second feature vectors are clustered to divide them into different clusters, and the usage conditions of the cluster are obtained according to the third usage conditions corresponding to the second feature vectors located in the same cluster. The usage conditions of the cluster are used to indicate the usage scenario and / or terminal statistics information corresponding to the cluster. The first feature vector located at the centroid of the cluster is obtained according to all the second feature vectors in the same cluster. Thus, similar second feature vectors are clustered into a cluster by clustering, and the feature vector that serves as the centroid of the cluster is determined as the first feature vector. Therefore, the first feature vector can be a centroid vector representing the common characteristics of a certain type of terminal, that is, the first feature vector can represent the feature vector common to a certain type of terminal.
[0037] The cluster usage conditions are obtained by statistically analyzing the third usage conditions corresponding to the second feature vectors located in the same cluster. For example, if second feature vectors under the same usage scenario are clustered into a cluster, then the usage condition of that cluster is the usage scenario corresponding to the second feature vector. Second feature vectors of terminals with similar structures may be clustered into a cluster, and the usage condition of that cluster is used to indicate terminal statistical information. The terminal statistical information includes model distribution information and the number of covered terminals, etc., so that the model distribution information points to the terminal, and the number of covered terminals indicates the number of terminals adapted to the cluster.
[0038] After obtaining the first feature vector, the first feature vector is processed using a preset frequency response simulation model to obtain the preset frequency response feature curve corresponding to the first feature vector. The usage conditions corresponding to the cluster to which the first feature vector belongs are used as the second usage conditions of the preset frequency response feature curve, thereby obtaining the preset frequency response feature curve under each second usage condition in advance.
[0039] The preset frequency response simulation model is trained based on audio data pairs collected under different usage conditions. Each audio data pair includes a standard audio signal and an actual audio signal. The actual audio signal is the standard audio signal processed by the terminal and output by the terminal. The standard audio signal and the actual audio signal in an audio data pair can be used to calculate the actual frequency response characteristic curve. Feature extraction is performed on the actual audio signal to obtain a high-dimensional feature vector. This high-dimensional feature vector is then processed using a preset encoding model to obtain a low-dimensional feature vector. The preset frequency response simulation model is trained using the actual frequency response characteristic curve and the low-dimensional feature vector. After training the preset frequency response simulation model, the first feature vector is processed using the preset frequency response simulation model to obtain the preset frequency response characteristic curve under the second usage condition. This preset frequency response characteristic curve represents the physical acoustic characteristics of the terminal under the second usage condition.
[0040] S102. Based on the first usage conditions of the audio processing model and the second usage conditions of each preset frequency response characteristic curve, obtain the first frequency response characteristic curve that matches the first usage conditions from all preset frequency response characteristic curves.
[0041] Because the characteristics of audio signals differ under different first usage conditions, it is very difficult to collect a large number of audio signals under the first usage conditions. To solve this problem, this embodiment obtains preset frequency response characteristic curves for different second usage conditions in advance. The preset frequency response characteristic curves can represent the physical acoustic characteristics of the terminal under the second usage conditions. Processing the standard audio signal using the preset frequency response characteristic curves is equivalent to processing the standard audio signal using the terminal under the second usage conditions. Thus, the actual audio signal output by the terminal under the second usage conditions can be obtained using the preset frequency response characteristic curves.
[0042] When it is necessary to collect audio signals under the first usage condition, this embodiment can obtain a second usage condition that matches the first usage condition, and use the preset frequency response characteristic curve under the second usage condition as the first frequency response characteristic curve that matches the first usage condition. That is, the first frequency response characteristic curve can represent the physical acoustic characteristics of the terminal under the first usage condition. After the standard audio signal is processed by the first frequency response characteristic curve, the resulting simulated audio signal can be used as the actual audio signal output by the terminal under the first usage condition, thus eliminating the need to manually collect audio signals under the first usage condition.
[0043] In one implementation, obtaining a first frequency response characteristic curve matching the first usage condition includes: obtaining a matching degree between the first usage condition and each second usage condition; and using a preset frequency response characteristic curve corresponding to a second usage condition whose matching degree satisfies a preset matching condition as the first frequency response characteristic curve. The matching degree indicates the degree of matching or similarity between the first and second usage conditions. The matching degree and the preset matching condition are used to determine whether the preset frequency response characteristic curve under the second usage condition can be used as the frequency response characteristic curve under the first usage condition. When the preset frequency response characteristic curve can be used as the frequency response characteristic curve under the first usage condition, the preset frequency response characteristic curve is the first frequency response characteristic curve matching the first usage condition.
[0044] The preset matching condition is used to indicate at least one of the matching degree range and the minimum matching degree when the preset frequency response characteristic curve can serve as the first frequency response characteristic curve. For example, if the matching degree is greater than the minimum matching degree, it is determined that the preset frequency response characteristic curve can serve as the first frequency response characteristic curve; or, if the matching degree is within the matching degree range, it is determined that the preset frequency response characteristic curve can serve as the first frequency response characteristic curve.
[0045] In some examples, the first usage condition includes the usage scenario and terminal information, and the second usage condition includes the usage scenario, the number of covered terminals, and model distribution information, whereby the model distribution information indicates the terminal information corresponding to the cluster to which the second usage condition belongs; obtaining the matching degree between the first usage condition and each second usage condition includes: If the usage scenario and / or terminal information in the first usage condition are recorded in the second usage condition, the matching degree between the first and second usage conditions is determined based on at least one of the model distribution information and the number of covered terminals. It is understood that: terminal information is used to indicate the terminal model, and model distribution information can record the terminal model. When the model indicated by the terminal information is recorded in the model distribution information, it is determined that the terminal information in the first usage condition is recorded in the second usage condition. When the usage scenario in the first usage condition is the same as the usage scenario in the second usage condition, it is determined that the usage scenario in the first usage condition is recorded in the second usage condition.
[0046] When determining the usage scenario and / or terminal information recorded in the first usage condition in the second usage condition, the number and model distribution information of covered terminals in the second usage condition are obtained. For example, the greater the number of covered terminals and / or the wider or more concentrated the model distribution indicated by the model distribution information, the greater the matching degree between the first and second usage conditions, making it more likely that the preset frequency response characteristic curve corresponding to the second usage condition will be selected as the first frequency response characteristic curve. For example, when the model distribution information records the model of the terminal indicated by the terminal information, the greater the matching degree between the second and first usage conditions to which this model distribution information belongs.
[0047] S103. The standard audio signal is processed using at least the first frequency response characteristic curve to obtain a simulated audio signal, which serves as training data for the audio processing model. In one implementation, the standard audio signal and the first frequency response characteristic curve are convolved in the time domain (or multiplied in the frequency domain) to obtain the simulated audio signal.
[0048] From the above Figure 1As shown in the flowchart, this embodiment can pre-obtain preset frequency response characteristic curves under different second usage conditions. After obtaining the first usage conditions of the audio processing model, based on the first usage conditions of the audio processing model and the second usage conditions of each preset frequency response characteristic curve, a first frequency response characteristic curve matching the first usage conditions is obtained from all preset frequency response characteristic curves. The standard audio signal is then processed using at least the first frequency response characteristic curve to obtain a simulated audio signal, making the simulated audio signal similar to the audio signal output by the terminal under the first usage conditions. The simulated audio signal can be used as training data for the audio processing model. Because the simulated audio signal is similar to the audio signal output by the terminal under the first usage conditions, the audio processing model can learn the characteristics of the audio signal output by the terminal under the first usage conditions, improving the processing effect of the audio processing model on the audio signal output by the terminal. Furthermore, after the first usage conditions of the audio processing model change, a first frequency response characteristic curve matching the changed first usage conditions can be obtained from all preset frequency response characteristic curves, eliminating the need for audio signal acquisition and frequency response characteristic curve analysis, thus saving manpower and time costs.
[0049] Please see Figure 2 This illustrates an optional flow of another data processing method provided in an embodiment of this application, which may include the following steps: S201. Obtain the first usage conditions of the audio processing model and the second usage conditions of each preset frequency response characteristic curve. The preset frequency response characteristic curve is obtained in advance using the first feature vector under the second usage conditions. The first feature vector is obtained based on the second feature vector under multiple third usage conditions. The second feature vector is obtained based on the third feature vector of the first audio signal output by the terminal under the third usage conditions. The second usage conditions are obtained based on multiple third usage conditions.
[0050] S202. Based on the first usage conditions of the audio processing model and the second usage conditions of each preset frequency response characteristic curve, obtain the first frequency response characteristic curve that matches the first usage conditions from all preset frequency response characteristic curves.
[0051] In this embodiment, steps S201 and S202 are the same as steps S101 and S102 described above, and will not be repeated here.
[0052] S203. Determine the first number of standard audio signals processed by the first frequency response characteristic curve and the second number of standard audio signals processed by the second frequency response characteristic curve. The second frequency response characteristic curve is a preset frequency response characteristic curve that does not match the first usage conditions. The first number is greater than the second number.
[0053] It is understood that the training data for the audio processing model includes positive and negative samples, with the number of positive samples being greater than the number of negative samples. Positive samples can be obtained using a first frequency response feature curve, and negative samples can be obtained using a second frequency response feature curve. Furthermore, to ensure that the number of positive samples is greater than the number of negative samples, this embodiment needs to determine a first number of standard audio signals processed by the first frequency response feature curve and a second number of standard audio signals processed by the second frequency response feature curve, with the first number being greater than the second number.
[0054] In one implementation, one way to determine the first quantity and the second quantity is to determine a first weight of the first frequency response characteristic curve and a second weight of the second frequency response characteristic curve, wherein the first weight is greater than the second weight, and either the first weight or the second weight is used to indicate the number of standard audio signals processed by the frequency response characteristic curve corresponding to that weight, and the relationship between the first weight and the second weight is used to represent the relationship between the first quantity and the second quantity.
[0055] For example, each preset frequency response characteristic curve has a base weight. If a preset frequency response characteristic curve is used as the first frequency response characteristic curve, then the base weight of the first frequency response characteristic curve is increased, and the increased base weight becomes the first weight. If a preset frequency response characteristic curve is used as the second frequency response characteristic curve, then the base weight of the second frequency response characteristic curve is decreased, and the decreased base weight becomes the second weight. As another example, either the first weight or the second weight can be determined based on the terminal statistics information in the second usage conditions. For instance, the larger the number of covered terminals in the terminal statistics information, and / or the wider or more concentrated the model distribution indicated by the model distribution information, the larger the weight value.
[0056] For example, when the first usage condition is used to refer to a terminal of a specific brand (such as a walkie-talkie manufacturer or a mobile phone), the first weight of the first frequency response characteristic curve with a high proportion of that brand is increased based on the model distribution information. As another example, when the first usage condition is used to refer to a specific usage scenario, the first weight of the first frequency response characteristic curve belonging to that specific usage scenario is increased based on the usage scenario in the second usage condition. For instance, if the specific usage scenario is an outdoor handheld scenario, the first weight of the first frequency response characteristic curve belonging to the outdoor handheld scenario is increased.
[0057] It should be noted here that if there are multiple first frequency response characteristic curves matching the first usage condition, the first weights of each of the multiple first frequency response characteristic curves can be the same or different. In scenarios where the first weights of the multiple first frequency response characteristic curves are different, this embodiment can set the first weights based on the terminal statistics information in the second usage condition. For example, the larger the number of covered terminals in the terminal statistics information, and / or the wider or more concentrated the model distribution indicated by the model distribution information, the larger the value of the first weight. Taking the first usage condition as indicating multiple terminals under multiple usage scenarios as an example, under this first usage condition, the audio signals of different terminals under different usage scenarios are uniformly sampled by default. Under uniform sampling, multiple first frequency response characteristic curves can be obtained, and then the first weight of each first frequency response characteristic curve is configured according to the terminal statistics information. For example, the larger the number of covered terminals, the more common the physical acoustic characteristics represented by the first frequency response characteristic curve are, and thus the larger the first weight of the first frequency response characteristic curve is. Therefore, different first weights are set for different first frequency response characteristic curves according to the number of covered terminals.
[0058] S204. Process a first number of standard audio signals using a first frequency response characteristic curve to obtain a first number of simulated audio signals, and process a second number of standard audio signals using a second frequency response characteristic curve to obtain a second number of simulated audio signals.
[0059] For example, a first audio subset and a second audio subset are obtained from a standard audio library. The first audio subset includes a first number of standard audio signals, and the second audio subset includes a second number of standard audio signals. Each standard audio signal in the first audio subset is convolved in the time domain (or multiplied in the frequency domain) with a first frequency response characteristic curve to obtain a first number of simulated audio signals. Similarly, each standard audio signal in the second audio subset is convolved in the time domain (or multiplied in the frequency domain) with a second frequency response characteristic curve to obtain a second number of simulated audio signals.
[0060] Because the first frequency response characteristic curve can represent the physical acoustic characteristics of the terminal under the first usage condition, while the second frequency response characteristic curve cannot, the simulated audio signal obtained using the first frequency response characteristic curve is similar to the audio signal output by the terminal under the first usage condition. Therefore, the simulated audio signal obtained using the first frequency response characteristic curve can be labeled as the audio signal output by the terminal under the first usage condition. Conversely, the simulated audio signal obtained using the second frequency response characteristic curve differs significantly from the audio signal output by the terminal under the first usage condition. Therefore, the simulated audio signal obtained using both the first and second frequency response characteristic curves can be used as training data for the audio processing model. Training the audio processing model improves its ability to learn the accurate and realistic audio feature data output by the terminal under the first usage condition, thereby enhancing the generalization ability and processing performance of the audio processing model.
[0061] Please see Figure 3 This illustrates an optional flow of another data processing method provided in an embodiment of this application, which may include the following steps: S301. Obtain the first usage conditions of the audio processing model and the second usage conditions of each preset frequency response characteristic curve. The preset frequency response characteristic curve is obtained in advance using the first feature vector under the second usage conditions. The first feature vector is obtained based on the second feature vector under multiple third usage conditions. The second feature vector is obtained based on the third feature vector of the first audio signal output by the terminal under the third usage conditions. The second usage conditions are obtained based on multiple third usage conditions.
[0062] S302. Based on the first usage conditions of the audio processing model and the second usage conditions of each preset frequency response characteristic curve, obtain the first frequency response characteristic curve that matches the first usage conditions from all preset frequency response characteristic curves.
[0063] S303. If a first frequency response characteristic curve is obtained according to the first and second usage conditions, the standard audio signal is processed using at least the first frequency response characteristic curve to obtain a simulated audio signal, which serves as training data for the audio processing model.
[0064] For a description of steps S301 to S303 above, please refer to the above. Figure 1 and Figure 2 The relevant descriptions in the illustrated embodiments will not be repeated here.
[0065] S304. If the first frequency response characteristic curve is not obtained according to the first and second usage conditions, then the second audio signal output by the terminal after the second standard audio signal is obtained under the first usage conditions.
[0066] S305. Based on the second standard audio signal and the second audio signal, obtain the frequency response characteristic curve of the terminal, and calculate the similarity between the frequency response characteristic curve of the terminal and each preset frequency response characteristic curve.
[0067] S306. Based on the similarity between the terminal's frequency response characteristic curve and each preset frequency response characteristic curve, and the second usage conditions, determine at least one first frequency response characteristic curve from the preset frequency response characteristic curves.
[0068] If the usage scenario and terminal information in the first usage condition are not recorded in the second usage condition, it is determined that there is no first preset frequency response characteristic curve that matches the first usage condition among all preset frequency response characteristic curves, that is, it is determined that no first frequency response characteristic curve has been obtained according to the first and second usage conditions.
[0069] If no preset frequency response characteristic curve matches the first usage condition among all preset frequency response characteristic curves, the control terminal plays a second standard audio signal under the first usage condition to acquire the second audio signal output by the terminal. The second standard audio signal and the second audio signal are processed using cross-correlation or frequency domain division to obtain the terminal's frequency response characteristic curve. Then, the similarity between the terminal's frequency response characteristic curve and each preset frequency response characteristic curve is calculated. Simultaneously, referring to the model distribution information in the second usage condition, at least one first frequency response characteristic curve is determined from all preset frequency response characteristic curves. For example, at least one first frequency response characteristic curve is selected from all preset frequency response characteristic curves based on preset brand conditions and / or preset similarity conditions. Alternatively, at least one first frequency response characteristic curve is selected from all preset frequency response characteristic curves based on the principle of prioritizing the same brand and / or a similarity greater than the preset similarity.
[0070] For example, in the initial test example of the new terminal, the terminal information is not recorded in the model distribution information of the second usage condition. If the usage scenario in the first usage condition is different from that in the second usage condition, then in the initial test example of the new terminal, there is no first preset frequency response characteristic curve that matches the first usage condition among all the preset frequency response characteristic curves. Therefore, in the initial test example of the new terminal, the frequency response characteristic curve of the new terminal is extracted using the second standard audio signal and the second audio signal. The similarity is calculated with all the pre-obtained preset frequency response characteristic curves. At the same time, based on the preset brand condition, the closest candidate frequency response characteristic curves are quickly matched from the preset frequency response characteristic curves, and the first frequency response characteristic curve is determined from the candidate frequency response characteristic curves. The determined first frequency response characteristic curve is the curve that matches the frequency response characteristic curve of the terminal. Therefore, the first frequency response characteristic curve can still represent the physical acoustic characteristics of the terminal, and the training data of the audio processing model can be obtained using the first frequency response characteristic curve.
[0071] The methods for determining the first frequency response characteristic curve from multiple candidate frequency response characteristic curves include, but are not limited to, at least one of the following: Multiple candidate frequency response feature curves are all designated as the first frequency response feature curve; candidate frequency response feature curves within the same brand that meet preset similarity conditions are designated as the first frequency response feature curve; candidate frequency response feature curves belonging to the same enterprise and meeting preset similarity conditions are designated as the first frequency response feature curve. The preset similarity conditions include, but are not limited to, at least one: maximum similarity; similarity greater than a preset threshold; similarity within a preset range.
[0072] S307. At least the first frequency response characteristic curve is used to process the standard audio signal to obtain a simulated audio signal, which serves as the training data for the audio processing model.
[0073] In the above Figure 3 In the illustrated process, although it is necessary to acquire the second audio signal output by the terminal under the first usage condition, the above... Figure 3 The process described only requires collecting a few second audio signals to roughly extract the terminal's frequency response characteristic curve. Based on the similarity between the terminal's frequency response characteristic curve and a preset frequency response characteristic curve, as well as the second usage conditions, a first frequency response characteristic curve is determined from all preset frequency response characteristic curves. Then, the first frequency response characteristic curve is used to process the standard audio signal to obtain a simulated audio signal as training data for the audio processing model. Although the user needs to manually collect a few second audio signals in the early stages, the above... Figure 3The process shown can still automatically match the first frequency response feature curve, and can use the first frequency response feature curve to obtain the training data of the audio processing model. While improving the processing effect of the audio processing model, it also achieves the goal of saving manpower and time costs.
[0074] exist Figure 3 Based on the process shown, the data processing method provided in this embodiment further includes: extracting features from the second audio signal to obtain a third feature vector of the second audio signal; processing the third feature vector of the second audio signal using a preset encoding model to obtain a second feature vector of the second audio signal, wherein the first usage condition corresponding to the second audio signal is the third usage condition of the second feature vector of the second audio signal; and updating the preset frequency response feature curve and the second usage condition corresponding to the preset frequency response feature curve using the second feature vector of the second audio signal and the third usage condition of the second feature vector of the second audio signal.
[0075] The updating of the preset frequency response characteristic curve and the second usage conditions corresponding to the preset frequency response characteristic curve includes: All second feature vectors are re-clustered to divide them into different clusters. The usage conditions of a cluster are obtained based on the third usage conditions corresponding to the second feature vectors within the same cluster. These usage conditions indicate the usage scenario and / or terminal statistics corresponding to the cluster. A first feature vector located at the centroid of the cluster is obtained from all second feature vectors within the same cluster. The first feature vector is processed using a preset frequency response simulation model to obtain a preset frequency response feature curve corresponding to the first feature vector. The usage conditions corresponding to the cluster to which the first feature vector belongs are used as the second usage conditions of the preset frequency response feature curve. This achieves the clustering of the newly acquired second feature vectors into a specific cluster and the inclusion of the first usage conditions in the second usage conditions. When the audio processing model is trained again under these first usage conditions, a first frequency response feature curve matching the first usage conditions can be obtained from all preset frequency response feature curves, eliminating the need for matching based on the frequency response feature curve corresponding to the acquired second audio signal, thus saving time and manpower costs.
[0076] One point needs to be clarified here: After obtaining the second audio signal and the frequency response characteristic curve of the terminal, the preset coding model can be updated, and the updated preset coding model can be used to process the second audio signal.
[0077] As can be seen from the above embodiments, the data processing method provided in this embodiment processes standard audio signals using the first frequency response characteristic curve without increasing the amount of actual audio signal acquisition or collection, dynamically simulating and generating a real audio set. This is equivalent to designing a front-end adaptive frequency response algorithm module, without being limited by the actual structure, hardware, and software influences of the terminal. This front-end adaptive frequency response algorithm module can automatically extract the frequency response characteristic curve of the terminal through a small amount of real recordings, in order to pre-generate and construct a diverse "target frequency response feature library," which includes preset frequency response characteristic curves corresponding to different second usage conditions. This front-end adaptive frequency response algorithm module can dynamically generate training data adapted to the audio processing model training using the data processing method described above, according to specific needs (such as the first usage condition). This fundamentally solves the limitations of existing solutions, adaptively fitting the audio signal that the terminal actually performs normally, thereby ensuring that the audio provided for the audio processing model's training is closest to the actual effect of the terminal, maximizing the realism and accuracy of the audio processing model's training data, enabling the audio processing model to learn data closer to reality, and improving the accuracy of the audio processing model.
[0078] like Figure 4 The diagram illustrates that hundreds of thousands of standard audio signals from a standard audio library are input into the front-end adaptive frequency response algorithm module. This module processes the standard audio signals using a first frequency response characteristic curve, outputting a simulated audio signal that closely approximates the actual audio signal output by the terminal. This simulated audio signal serves as input to the AI algorithm module. The AI algorithm model learns from the simulated audio signal to acquire the audio characteristics of the actual audio signal output by the terminal. Specifically, the AI algorithm module stores the audio processing model, and the simulated audio signal is used to train this model, enabling it to learn the audio characteristics of the actual audio signal output by the terminal.
[0079] In combination with the above Figures 1 to 4 As can be seen, the data processing method provided in this embodiment includes: original data pair construction → feature extraction and model training → cluster analysis and target curve library construction → dynamic frequency response simulation. The original data pair construction → feature extraction and model training → cluster analysis and target curve library construction are for pre-generating a preset frequency response characteristic curve and the corresponding second usage conditions. Dynamic frequency response simulation is for obtaining a first frequency response characteristic curve that matches the first usage conditions, and using the first frequency response characteristic curve to obtain a simulated audio signal. The following describes each step: Step 1: Acquire raw data pairs (Data Pair Collection), that is, acquire audio data pairs under the third usage condition. The third usage condition can be the environment in which the terminal is located, referred to as the real usage condition—the “system input-output” acquisition of the terminal under the real usage condition: The goal of this step is to acquire the paired data (Xstandard, Ycaptured) of the terminal’s implicit frequency response characteristics through the process of the terminal “playing a standard audio signal - passing through the terminal’s MIC and cavity hardware integrated system, and then recording the audio signal after processing by the integrated system”. The frequency response characteristics of the terminal can indicate the physical acoustic characteristics of the terminal, and the frequency response characteristics of the terminal can be represented by a frequency response characteristic curve.
[0080] In step one, the Xstandard (the input standard audio signal, also known as the first standard audio signal) in the audio data pair is designed as a test signal covering the entire frequency band and with uniform energy distribution to ensure effective excitation of the terminal's frequency response characteristics. Xstandard can include, but is not limited to, at least one of the following signals: White noise: spectrally flat (e.g., frequency band between 20Hz and 20kHz), used for global frequency response estimation; Log sweep signal: frequency increases linearly with time (e.g., 20Hz → 20kHz, duration 10s), can accurately extract the frequency response of linear time-invariant (LTI) systems through inverse Fourier transform; Multi-tonal sweep: densely distributed sine waves (100Hz intervals) in specific frequency bands (e.g., 300Hz-3kHz human voice bandwidth, high frequencies above 10kHz), enhancing the capture of local frequency response details; Speech ultimatum: for example, selecting everyday colloquial phrases, covering typical sound sources in actual use, supplementing the frequency response characteristics of non-steady-state audio signals.
[0081] Ycaptured (the audio signal output from Xstandard after terminal processing, also known as the first audio signal) in the audio data pair is recorded using the terminal's built-in microphone or the user's commonly used microphone, ensuring consistency with the audio signal output by the terminal under the third usage condition. Recording can meet the following conditions: control ambient noise and avoid strong reflections, but a laboratory-grade acoustic environment is not mandatory, to preserve acoustic interference under the third usage condition. For example, recording in a semi-anechoic chamber or quiet room with a sound pressure level (SPL) less than or equal to 40 dB can preserve slight reverberation in the recording environment.
[0082] Data pair validity assurance: Each terminal collects a preset set of data pairs, such as collecting ≥5 sets of data pairs at different times and under different power conditions, to avoid accidental errors in a single recording; Xstandard and Ycaptured in the same data pair are time-aligned (matching the starting point through a cross-correlation algorithm) to ensure the time synchronization of input and output; After acquiring the data pairs, anomaly detection is performed to remove abnormal data. One anomaly detection method is to calculate the signal-to-noise ratio (SNR) of Ycaptured in the audio data pair. If the SNR is greater than a preset SNR, the audio data pair to which Ycaptured belongs is determined to be an abnormal data pair. Through anomaly detection, valid audio data pairs can be filtered out, and the valid audio data pairs are taken as the result of step one.
[0083] If the terminal moves or the user speaks during recording, the signal-to-noise ratio (SNR) of Ycaptured may be less than 30dB. Whether an audio data pair is an anomalous pair is determined by checking if the SNR of Ycaptured is greater than or equal to 30dB. For example, if the SNR of Ycaptured is greater than or equal to 30dB, the audio data pair to which it belongs is determined to be an anomalous pair.
[0084] Step 2: Feature Extraction and Model Training (e.g., FREN Model) – By using the high-dimensional feature vectors of Ycaptured in the audio data pair, we can obtain the low-dimensional feature vectors of Ycaptured in the audio data pair. The high-dimensional feature vectors can be called the third feature vectors, and the low-dimensional feature vectors can be called the second feature vectors. This allows us to reverse-engineer the core features (such as the low-dimensional feature vectors) that characterize the frequency response from the complex audio signal.
[0085] In step two, the FREN model (Frequency Response Extraction Network) is trained. Its input is the frequency domain features of Ycaptured, which are high-dimensional feature vectors. The output is a low-dimensional feature vector representing the frequency response characteristics of the terminal. Essentially, it learns the compressed representation of the "terminal system function H" (Ycaptured ≈ H·Xstandard). H (system function) is calculated by the pair (Xstandard, Ycaptured), such as by cross-correlation: H = IFFT(cross-correlation(Xstandard, Ycaptured)), or by frequency domain division: FFT(Ycaptured) / FFT(Xstandard) (with regularization).
[0086] The input feature design for the FREN model involves extracting multi-scale frequency domain features from Ycaptured. These features are combined to form a high-dimensional feature vector, which serves as the input to the FREN model. Ycaptured's multi-scale frequency domain features can capture both temporal dynamics and spectral details. Examples of Ycaptured's multi-scale frequency domain features include, but are not limited to: STFT (Short-Time Fourier Transform) amplitude spectrum: for example, a window length of 25ms (1024 points, sampling rate of 48kHz), a window shift of 10ms, to extract the time-frequency graph (time frame × frequency bins, such as 1000 frames × 240 bins); Mel spectrogram: for example, mapping linear frequencies to a Mel scale (20-8000Hz corresponds to 0-128 bins) to simulate the human ear's perception of frequency; Log-Mel energy: taking the logarithm of the Mel spectrogram (log(1+|S|)) to compress the dynamic range and highlight perceptually relevant spectral changes; LPC (Linear Predictive Coding) coefficients: for example, fitting the linear prediction residual of Ycaptured using a 12th-order LPC model to capture the formant characteristics of the frequency response.
[0087] Feature fusion is performed on the aforementioned frequency domain features: features of different dimensions are synchronized in time and aligned with the frame rate to generate a two-dimensional matrix. The rows of the two-dimensional matrix can be time data, and the columns can be all the concatenated feature dimensions. The two-dimensional matrix is a representation of high-dimensional feature vectors.
[0088] The FREN model employs an encoder-decoder structure. The encoder compresses high-dimensional feature vectors into low-dimensional feature vectors, retaining information relevant to the frequency response characteristics while discarding irrelevant information. For example, it retains information related to the inherent frequency response shape and resonance mode of the terminal, while discarding irrelevant information such as environmental noise, recording level, and transient interference. The decoder assists in training, such as reconstructing Ycaptured to constrain the effectiveness of the encoding. The encoder's input can be the high-dimensional feature vector of Ycaptured, and its output can be the low-dimensional feature vector. The decoder's input can be the low-dimensional feature vector, and its output can be the reconstructed feature vector. captured.
[0089] The structure of the encoder and decoder is not limited in this embodiment. For example, the encoder may include a Convolutional Neural Network (CNN) layer, a Long Short-Term Memory (LSTM) layer, and a fully connected layer. For instance, the encoder may include three CNN layers, two bidirectional LSTM layers, and one fully connected layer. The three CNN layers are cascaded, the two bidirectional LSTM layers are cascaded, the third layer of the three CNN layers is connected to the first layer of the two bidirectional LSTM layers, the second layer of the two bidirectional LSTM layers is connected to the fully connected layer, the first layer of the three CNN layers is the input to the encoder, and the output of the fully connected layer is the output of the encoder. Three CNN layers are used to extract local texture features. For example, the three CNN layers include convolution kernels, which are 3×3. Different sizes of convolution kernels are used to extract local texture features of the spectrum. Two LSTM layers are used to capture the long-term correlation of audio signals on the time axis, such as capturing temporal dependence. Specifically, it can capture the dynamic changes of the signal. The fully connected layer can output a low-dimensional feature vector. The low-dimensional feature vector condenses the terminal's frequency response, phase distortion, and nonlinear distortion features.
[0090] The decoder can include transposed convolutional (CNN) layers and fully connected layers, such as two cascaded transposed convolutional layers and one fully connected layer. The decoder reconstructs the low-dimensional feature vector into an approximate Y-captured spectrum, which is simply referred to as the reconstructed spectrum. Captured. Reconstructed during the training phase. The captured value is used to calculate the loss value.
[0091] The FREN model employs a multi-task joint loss function, including: spectral reconstruction loss (Lspectral), system function consistency loss (Lsystem), and perceptual loss (Lperceptual). The spectral reconstruction loss is used to calculate the difference between Ycaptured and reconstructed data in the audio data pair. The difference between Ycaptured and reconstructed in the feature space is specifically calculated by comparing Ycaptured with the reconstructed Y. The L1 loss between captured values constrains the ability of the low-dimensional feature vector to represent Ycaptured, ensuring that the encoder-decoder can accurately reconstruct the spectral features of the input. The weight λ1 of the spectral reconstruction loss Lspectral is usually set to 1.0, and the spectral reconstruction loss Lspectral can be used as a basic loss term.
[0092] The system function consistency loss Lsystem is used to ensure consistency between Ycaptured and refactored functions. The captured features not only exhibit spectral similarity but, more importantly, maintain identical frequency response characteristics. This is the core difference between the encoder in the FREN model and a regular autoencoder, ensuring that the low-dimensional feature vector accurately represents the terminal's frequency response characteristics. The system function consistency loss Lsystem is used to calculate the error between the first system function H and the second system function H, constraining the fidelity of the low-dimensional feature vector to the system function H to guarantee that the low-dimensional feature vector accurately represents the terminal's frequency response characteristics. The first system function H is calculated using Xstandard and Ycaptured, and the second system function H is calculated using Xstandard and Ycaptured... The weight λ2 of the system function consistency loss Lsystem is calculated. It is usually set to 0.5-1.0 and adjusted according to the training stability.
[0093] Perceptual loss Lperceptua is used to ensure Captured is similar to Ycaptured in human auditory perception. The perceptual loss Lperceptua can be evaluated based on a pre-trained PMSQE (Perceptual Mel-Spectrogram Quality Evaluation) model. The captured perceptual similarity is used as the perceptual loss Lperceptua; the weight λ3 of the perceptual loss Lperceptua is usually set to 0.3-0.5 to balance perceptual quality and technical indicators.
[0094] The loss value of the FREN model is Ltotal = λ1Lspectral + λ2Lsystem + λ3Lperceptua.
[0095] The training strategies for the FREN model include data augmentation, transfer learning, and end-to-end training. Data augmentation can add distortions to Ycaptured to simulate real-world scenarios, providing diverse inputs for transfer learning and allowing pre-trained knowledge to better adapt to the target task, thus improving model robustness. For example, Ycaptured can collect environmental noise, reverberation, etc.
[0096] Transfer learning can be achieved by using a pre-trained audio classification network (such as YAMNet) to extract general audio features in the initial stage, fine-tuning the FREN model to adapt it to a specific task, providing a good starting point for end-to-end training, accelerating convergence and improving stability.
[0097] End-to-end training can jointly optimize the encoder and decoder, and iteratively train to convergence (the validation set loss no longer decreases) using the Adam optimizer (learning rate 1e-4), thereby making full use of augmented data and transfer knowledge to achieve global optimum.
[0098] The synergistic effect of the three training strategies described above enabled the FREN model to achieve optimal performance in terms of system function estimation error, convergence speed, and generalization performance, laying a solid foundation for subsequent steps three (cluster analysis) and four (dynamic frequency response simulation). After training the FREN model, the encoder in the FREN model is used as the preset encoding model. When audio signals are subsequently acquired, the encoder in the FREN model is used to process the high-dimensional feature vectors of the audio signals to obtain the low-dimensional feature vectors.
[0099] Step 3: Cluster Analysis and Construction of the Target Curve Library—Unsupervised clustering is used to mine the implicit frequency response commonalities in data pairs, mapping low-dimensional feature vectors to several typical "pre-defined frequency response feature curves," solving the problem that a "single line" cannot cover diversity. The aim is to obtain the first feature vector based on the second feature vector under multiple third usage conditions, obtain the second usage conditions based on multiple third usage conditions, and pre-determine the pre-defined frequency response feature curves using the first feature vector. Step 3 includes steps 1) to 4). Step 1) Feature Vector Set Construction: Collect M sets of valid data pairs from N different terminals (different models, batches, and configurations), with each terminal corresponding to at least one low-dimensional feature vector V. i (Dimension d=256, output by the encoder of the FREN model), forming a set {V1, V2, ..., V m}
[0100] Step 2) Use an unsupervised clustering algorithm to cluster the low-dimensional feature vectors in the feature vector set: A hybrid strategy combining DBSCAN (Density-Based Spatial Clustering of Applications with Noise) and Hierarchical Clustering algorithms is employed. Both algorithms are used to cluster low-dimensional feature vectors within a feature vector set. DBSCAN automatically identifies outliers (such as faulty terminals or abnormal recordings) to prevent noise from interfering with the cluster structure. The Hierarchical Clustering algorithm refines the intra-cluster structure by constructing a dendritic graph using the cosine similarity (distance metric) of feature vectors and combining this with the silhouette score to determine the optimal number of clusters K (typically K=5-8, covering over 95% of terminals).
[0101] Step 3) Generation of the target frequency response feature vector, which can be called the first feature vector: For each cluster C k(Including M) k (a low-dimensional eigenvector), calculate its centroid V. k (center) This serves as the "target frequency response eigenvector" for this cluster. V k (center) The physical meaning is that the average of all low-dimensional eigenvectors in the cluster best represents the common characteristics of the frequency response of the terminals within the cluster. k (center) It can be called the first eigenvector.
[0102] Step 4) Construction and verification of the target curve library. The target curve library stores preset frequency response characteristic curves and the corresponding second usage conditions for each preset frequency response characteristic curve—setting each V k (center) Decoded into a preset frequency response characteristic curve H k The rationality of the preset frequency response characteristic curve is verified by combining actual test data from typical terminals, and the curves that pass the rationality verification are saved in the target curve library. Each H... k The preset frequency response simulation model is obtained by means of a pre-defined frequency response model. Essentially, this model is the inverse mapping from low-dimensional feature vectors to the system function H. Specifically, the preset frequency response simulation model can employ a neural network inverse model. Its input is a low-dimensional feature vector, and its output is the system function H, which can serve as the preset frequency response feature curve. The training data consists of H (calculated by Xstandard and Ycaptured) from the (Xstandard, H) pair obtained in step one, and the low-dimensional feature vector of Ystandard obtained in step one. Network structure: FCN (Fully Connected Network) or CNN (Convolutional Neural Network). In one implementation, the system function H is the filter parameter, such as FIR (Finite Impulse Response) coefficients or IIR (Infinite Impulse Response) denominator / numerator coefficients. The output dimension of the preset frequency response simulation model is adjusted according to the filter type; for example, if the FIR filter length is 512, then the output will be 512-dimensional coefficients.
[0103] H is obtained through a preset frequency response simulation model. k Then, calculate H. k The mean squared error (MSE) between the system function H and the actual system function H is used to determine H. If the MSE meets the preset mean squared error condition (e.g., MSE ≤ 0.5 dB), then H is determined. k The rationality verification was passed. Finally, the target curve library is stored as {H1, H2, ..., H...} k}, each H k The accompanying statistical information can be considered as a second condition of use. The second condition of use includes at least one of the following: the number of covered terminals, model distribution information, and usage scenarios.
[0104] In this embodiment, each H k The accompanying statistical information serves the following purposes: Application 1: Dynamic Sampling Weight Calculation: The number of covered terminals serves as the base weight. A larger number of covered terminals indicates a more common preset frequency response characteristic curve, resulting in a higher weight in the default uniform sampling. Default uniform sampling allows the audio processing model to adapt to most preset frequency response characteristic curves; higher weights mean more standard audio signals are processed by the preset frequency response characteristic curve. Model distribution information is used for targeted enhancement. When the training target is a terminal of a specific brand (e.g., a walkie-talkie manufacturer or a mobile phone), the H-value of the brand with a high proportion in the corresponding model distribution information is increased. k The sampling probability, that is, increasing the H that has a high proportion of this brand. k The number of standard audio signals processed. Scenario-guided training is used; for example, when the audio processing model is designed for outdoor noise reduction, the weights of the preset frequency response characteristic curves for "outdoor handheld" scenarios are increased, such as by 30-50%.
[0105] Application 2: Target curve library completeness assessment: Calculate coverage index based on model distribution information. If the coverage rate of a mainstream model series in the library is <60%, trigger the supplementary data collection process. Identify missing scenarios based on scene distribution. By analyzing the accompanying usage scenario statistics, identify typical scenarios with low coverage (such as "strong wind and noise environment") to guide targeted data collection.
[0106] Application 3: Monitoring the balance of terminal quantity distribution: When a certain H k If the number of covered terminals exceeds 30% of the total, it indicates that there may be over-clustering, triggering clustering parameter optimization.
[0107] Application 4: Rapid adaptation mechanism for new terminals: During the initial testing of a new terminal, its rough frequency response characteristics are extracted and compared with those of each H in the target curve library. k Calculate similarity, while also referencing model distribution information (e.g., prioritizing brands), and quickly match the 3-5 closest candidate frequency responses. If the new terminal's model is already recorded in the model distribution information, directly use the H frequency response with the highest percentage for that model. k This reduces the adaptation time for new terminals. By combining usage scenario tags, the system recommends the optimal preset frequency response characteristic curve for new terminals in specific scenarios. For example, for law enforcement recorders in "high humidity environments," the preset frequency response characteristic curve for humidity-resistant scenarios is preferred.
[0108] Application 5: Enhanced model generalization ability: During the training of AI algorithm models, adaptive batching is designed based on the number of covered terminals: the proportion of samples of high-frequency response feature categories in the minibatch is proportional to the number of covered terminals.
[0109] Application 6: Design adversarial training based on model distribution information: Design adversarial losses between different brand models to improve the model's invariance to brand differences.
[0110] Application 7: Design scenario adversarial training based on usage scenarios: Build a scenario classifier and force the AI algorithm model to learn scenario-independent feature representations through gradient inversion layers.
[0111] From each H k As can be seen from the use of the accompanying statistical information, the statistical information is not only stored as metadata, but is also deeply integrated into the decision-making system of dynamic frequency response simulation, forming a "data-feature-application" closed loop. This is the main innovation that distinguishes it from the traditional static frequency response library.
[0112] Step 4: Dynamic Frequency Response Simulation – Generating high-fidelity simulation data for AI training, i.e., generating simulated audio signals that closely approximate the actual output of the terminal for the audio processing model. This step utilizes preset frequency response characteristic curves from the target curve library to dynamically generate simulated audio signals adapted to the training of the audio processing model (such as an AI algorithm model), solving the problem that "fixed curves" cannot cover the diversity of real-world data.
[0113] In step four, a first frequency response feature curve matching the first usage condition is selected from all preset frequency response feature curves in the target curve library. The first usage condition can be determined based on the training requirements of the AI algorithm model. For example, if the training requirement of the AI algorithm model is to improve the model's robustness to high-frequency noise, then the first usage condition is a high-frequency noise scenario. If the training requirement of the AI algorithm model is to improve low-frequency noise reduction capability, then the first usage condition is a low-frequency noise scenario. The first frequency response feature curve is selected based on the first usage condition determined by the training requirements of the AI algorithm model.
[0114] For example, if the training requirement of an AI algorithm model is to improve low-frequency noise reduction capabilities, it is necessary to increase the number of simulated audio signals in low-frequency noise scenarios. The preset frequency response characteristic curve for the corresponding low-frequency noise scenario is then the first frequency response characteristic curve, and the number of standard audio signals processed by the first frequency response characteristic curve is increased. Given the known training requirements of the AI algorithm model, adversarial training also needs to be introduced. Adversarial data is introduced during AI algorithm model training. This adversarial data can be simulated audio signals obtained using a second frequency response characteristic curve that does not match the first usage conditions, thereby introducing adversarial training during AI algorithm model training and enhancing the model's generalization ability. If the training requirement of the AI algorithm model is to improve low-frequency noise reduction capabilities, then simulated audio signals in ultra-high frequency usage scenarios need to be introduced.
[0115] After determining the first and second frequency response characteristic curves, a large number of standard audio signals are selected. The standard audio signals are then convolved in the time domain (or multiplied in the frequency domain) with the determined frequency response characteristic curves to obtain the simulated audio signals. sim Furthermore, after obtaining the simulated audio signal... sim The authenticity of the data can then be verified, such as through calculations. sim The frequency response difference between the simulated audio signal and the actual terminal recording (Ycaptured) is minimized (e.g., MSE ≤ 1dB) to ensure a high degree of consistency between the simulated audio signal and the real data. The simulated audio signal... sim The corresponding labeled data (such as audio signal labels and noise type labels) are input into the AI algorithm model (such as noise reduction network and AEC network) for learning and training, which improves the AI algorithm model's ability to learn more realistic and accurate feature data, enhances the AI algorithm model's generalization ability to real terminals, and improves the accuracy of the AI algorithm model.
[0116] As can be seen from steps one through four above, step four forms an organic whole with the first three steps, and their relationship is as follows: Dependency Architecture: I. Data Dependency: Step four is fully dependent on the target curve library {H1,...,H} built in step three. k The construction of this library is deeply dependent on the FREN model trained in step two and the raw data collected in step one; secondly, model dependence: the training of the preset frequency response simulation model used in step three depends on the system function H extracted in step one, and the accurate calculation of H depends on the (Xstandard, Ycaptured) pair that is precisely synchronized in step one; thirdly, knowledge dependence: the dynamic feature selection strategy in step four depends on the statistical information obtained from the cluster analysis in step three, and this distribution directly reflects the feature vector space structure extracted in step two.
[0117] Feedback Optimization Closed Loop: 1. Performance Feedback Closed Loop: After training the AI algorithm model with the simulation data generated in step four, the performance on the actual terminal will be fed back to step one, guiding the focus of subsequent data collection (e.g., if the performance of a certain frequency band is insufficient, then increase the test signal of that frequency band); 2. Feature Optimization Closed Loop: Simulation data that fails verification in step four (e.g., MSE>1dB) triggers targeted retraining of the FREN model in step two, forming a continuous model optimization mechanism; 3. Library Expansion Closed Loop: When step four is applied to a new terminal type, if there is no matching frequency response in the target curve library, step one is triggered to supplement the collection, step two is retrained, and step three dynamically expands the library, realizing the self-evolution of the knowledge base.
[0118] Functional synergy mechanism: The real data collected in step one serves as the verification benchmark for the preset frequency response simulation model used in step three, ensuring the simulation's authenticity; the data generated in step four can be expanded into supplementary test signals for step one, forming a closed loop of acquisition-generation-verification; the encoding capability of the FREN model in step two directly affects the inverse mapping accuracy of the preset frequency response simulation model used in step three; the performance feedback of the preset frequency response simulation model used in step three guides the structural optimization of the FREN model; the target curve library constructed in step three is the data foundation for step four; the application effect of step four (such as the generalization performance of the AI algorithm model) directly evaluates the clustering quality of step three, forming an evaluation-optimization loop.
[0119] Computational resource allocation strategy: Offline-online division of labor: Steps one to three are offline processes, concentrating computing resources to build a high-quality frequency response library; Step four is divided into two parts: offline data generation and online inference. The former generates training data in batches, while the latter is a lightweight real-time application; Elastic resource allocation: When Step four detects that the performance of a certain type of terminal is insufficient, resources are dynamically adjusted, and Step one is prioritized to supplement the collection of data of that type, reflecting the system's adaptability; Incremental update mechanism: Steps one to three support incremental updates. Newly collected data does not require full retraining, only updating the relevant parts, ensuring that Step four continuously obtains the latest frequency response features.
[0120] In summary, the innovation of the data processing method provided in this embodiment lies in upgrading the acquisition of frequency response characteristic curves from "relying on laboratory measurements" to "adaptive learning and induction," and its advantages are reflected in the following aspects: 1. Breakthrough in thinking mode: From "static measurement" to "dynamic learning": The existing method is a linear process based on "measurement-modeling", which assumes that the frequency response characteristic curve of the terminal is fixed and can be accurately measured. This embodiment, however, believes that the frequency response characteristic curve of the terminal is a statistical representation of a complex system. By learning a large number of real data pairs, it summarizes the frequency response characteristic curves that cover a variety of coverage, which fundamentally solves the problems of "incomplete coverage and poor adaptability" of traditional methods.
[0121] 2. Integration of technical means: Deep integration of signal processing and machine learning; for example, this embodiment deeply integrates signal processing (such as frequency response analysis, filter design) with machine learning (such as feature extraction, clustering, neural network inverse model); replaces manual feature engineering with deep learning (such as existing methods rely on engineers' experience to design test sounds, while this embodiment automatically learns key features through AI algorithm models); replaces manual classification with unsupervised clustering (existing methods require manual division of terminal types, while this embodiment automatically determines the frequency response characteristic curve through statistical information); and replaces static curves with dynamic simulation (existing methods use a single line, while this embodiment generates diverse simulation data).
[0122] 3. Improved application performance: scalability, realism, and robustness
[0123] Scalability: Only a small amount of terminal data needs to be collected (≥5 sets per terminal) to generalize a frequency response feature library covering thousands of terminals, significantly reducing testing costs; Authenticity: Feature extraction is based on recordings of real user scenarios (rather than ideal laboratory environments), and the target frequency response feature library is closer to the frequency response fluctuations in actual use (such as changes in acoustic path caused by holding); Robustness: The AI algorithm model learns the frequency response features of hundreds of "virtual terminals", significantly improving its generalization ability to new and unknown terminals (experiments show that the noise reduction model trained in this embodiment improves the SNR of unknown terminals by 3-5dB compared to traditional solutions).
[0124] Figure 5 and Figure 6 These are comparison images of the effects. Figure 5 This is a comparison chart of AEC processing performance in a one-way duplex conversation scenario. Figure 6 This is a comparison chart of AEC processing effects in a duplex / dual-talk scenario. From Figure 5 and Figure 6 It can be seen that after processing by the data processing method provided in this implementation, the noise echo is cleaner and the effect is better.
[0125] The above describes a data processing method provided by an embodiment of this application. The following describes an apparatus for performing the above data processing method.
[0126] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application. Figure 7 As shown, the data processing device includes: an acquisition unit 10, a matching unit 20, and a processing unit 30.
[0127] The acquisition unit 10 is used to acquire the first usage conditions of the audio processing model and the second usage conditions of each preset frequency response characteristic curve. The preset frequency response characteristic curve is obtained in advance using the first feature vector under the second usage conditions. The first feature vector is a feature vector representing the common characteristics of a class of terminals, obtained from the second feature vector under multiple third usage conditions. The second feature vector is obtained from the third feature vector of the first audio signal output by the terminal under the third usage conditions. The second usage conditions are obtained from multiple third usage conditions.
[0128] The matching unit 20 is used to obtain a first frequency response feature curve that matches the first usage condition from all preset frequency response feature curves, based on the first usage condition of the audio processing model and the second usage condition of each preset frequency response feature curve.
[0129] The processing unit 30 is used to process the standard audio signal using at least the first frequency response characteristic curve to obtain a simulated audio signal, which serves as training data for the audio processing model.
[0130] In some examples, the data processing apparatus may further include a feature acquisition unit for clustering all second feature vectors to divide all second feature vectors into different clusters, and obtaining the usage conditions of the clusters based on the third usage conditions corresponding to the second feature vectors located in the same cluster. The usage conditions of the clusters are used to indicate the usage scenarios and / or terminal statistics corresponding to the clusters. Based on all second feature vectors in the same cluster, the feature vector located at the centroid of the cluster is obtained, and the feature vector located at the centroid of the cluster is the first feature vector. The first feature vector is processed using a preset frequency response simulation model to obtain a preset frequency response feature curve corresponding to the first feature vector, and the usage conditions corresponding to the cluster to which the first feature vector belongs are used as the second usage conditions of the preset frequency response feature curve.
[0131] In one possible implementation, a feature acquisition unit is used to acquire multiple audio data pairs under each third usage condition. The audio data pairs include a first standard audio signal and a first audio signal. The first standard audio signal and the first audio signal in the same audio data pair are used to calculate the frequency response feature curve under the third usage condition. Feature extraction is performed on each first audio signal to obtain a third feature vector of each first audio signal. The third feature vector of each first audio signal is processed using a preset coding model to obtain a second feature vector of each first audio signal. The preset coding model is trained based on the frequency response feature curve under the third usage condition and the second feature vector of each first audio signal.
[0132] In one possible implementation, the matching unit 20 is specifically used to obtain the matching degree between the first usage condition and each second usage condition; and to use the preset frequency response characteristic curve corresponding to the second usage condition whose matching degree satisfies the preset matching condition as the first frequency response characteristic curve.
[0133] The first usage condition includes usage scenario and terminal information, while the second usage condition includes usage scenario, number of covered terminals, and model distribution information. The model distribution information indicates the terminal information corresponding to the cluster to which the second usage condition belongs. Correspondingly, the matching unit 20 is specifically used to determine the matching degree between the first usage condition and the second usage condition if the usage scenario and / or terminal information in the first usage condition is recorded in the second usage condition, based on at least one of the model distribution information and the number of covered terminals.
[0134] There is a possibility that the matching unit 20 may not acquire the first frequency response characteristic curve based on the first and second usage conditions. In this case, the data processing device may include a feature acquisition unit 40 and a similarity calculation unit 50, such as... Figure 8 As shown.
[0135] The system includes: an acquisition unit 10, configured to acquire a second audio signal output by the terminal after processing the second standard audio signal under the first usage condition; a feature acquisition unit 40, configured to obtain the frequency response characteristic curve of the terminal based on the second standard audio signal and the second audio signal; a similarity calculation unit 50, configured to calculate the similarity between the terminal's frequency response characteristic curve and each preset frequency response characteristic curve; and a matching unit 20, configured to determine at least one first frequency response characteristic curve from the preset frequency response characteristic curves based on the similarity between the terminal's frequency response characteristic curve and each preset frequency response characteristic curve, and the second usage condition.
[0136] In some examples, the feature acquisition unit 40 is also used to extract features from the second audio signal to obtain a third feature vector of the second audio signal; process the third feature vector of the second audio signal using a preset coding model to obtain a second feature vector of the second audio signal, wherein the first usage condition corresponding to the second audio signal is the third usage condition of the second feature vector of the second audio signal; and update the preset frequency response feature curve and the second usage condition corresponding to the preset frequency response feature curve using the second feature vector of the second audio signal and the third usage condition of the second feature vector of the second audio signal.
[0137] In one possible implementation, the processing unit 30 is configured to determine a first number of standard audio signals processed by a first frequency response characteristic curve and a second number of standard audio signals processed by a second frequency response characteristic curve, wherein the second frequency response characteristic curve is a preset frequency response characteristic curve that does not match the first usage conditions, and the first number is greater than the second number; process the first number of standard audio signals using the first frequency response characteristic curve to obtain a first number of simulated audio signals; and process the second number of standard audio signals using the second frequency response characteristic curve to obtain a second number of simulated audio signals.
[0138] The processing unit 30 determines the first quantity and the second quantity by determining a first weight of the first frequency response characteristic curve and a second weight of the second frequency response characteristic curve. The first weight is greater than the second weight. Either the first weight or the second weight is used to indicate the number of standard audio signals processed by the frequency response characteristic curve corresponding to that weight. The relationship between the first weight and the second weight is used to represent the relationship between the first quantity and the second quantity.
[0139] Because the simulated audio signal is similar to the audio signal output by the terminal under the first usage condition, the audio processing model can learn the characteristics of the audio signal output by the terminal under the first usage condition, thus improving the processing effect of the audio processing model on the audio signal output by the terminal. Furthermore, after the first usage condition of the audio processing model changes, a first frequency response characteristic curve matching the changed first usage condition can be obtained from all preset frequency response characteristic curves, eliminating the need for audio signal acquisition and frequency response characteristic curve analysis, saving manpower and time costs. Moreover, the first feature vector can represent a common feature vector of a class of terminals, thus allowing this class of terminals to use the same preset frequency response characteristic curve, again eliminating the need for audio signal acquisition and frequency response characteristic curve analysis for this class of terminals, saving manpower and time costs.
[0140] For a detailed description of each unit in the above data processing device, please refer to the method embodiment section, which will not be repeated here.
[0141] This application also provides an electronic device, including at least one processor and a memory connected to the processor, wherein: the memory is used to store a computer program; the processor is used to execute the computer program so that the electronic device can implement any of the data processing methods provided in this application.
[0142] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the data processing methods provided in this application.
[0143] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the data processing methods provided in this application.
[0144] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0145] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0146] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0147] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. A data processing method, characterized in that, include: The first usage condition of the audio processing model is obtained, and the second usage condition of each preset frequency response characteristic curve is obtained. The preset frequency response characteristic curve is obtained in advance using the first feature vector under the second usage condition. The first feature vector is a feature vector representing the common characteristics of a class of terminals obtained from the second feature vector under multiple third usage conditions. The second feature vector is obtained from the third feature vector of the first audio signal output by the terminal under the third usage conditions. The second usage condition is obtained from the multiple third usage conditions. Based on the first usage condition of the audio processing model and the second usage condition of each preset frequency response characteristic curve, a first frequency response characteristic curve matching the first usage condition is obtained from all preset frequency response characteristic curves; The standard audio signal is processed using at least the first frequency response characteristic curve to obtain a simulated audio signal, which is the training data of the audio processing model.
2. The method according to claim 1, characterized in that, The preset frequency response characteristic curve is obtained in advance using a first feature vector under the second usage condition. The first feature vector is obtained based on second feature vectors under multiple third usage conditions, including: All second feature vectors are clustered to divide them into different clusters, and the usage conditions of the clusters are obtained according to the third usage conditions corresponding to the second feature vectors located in the same cluster. The usage conditions of the clusters are used to indicate the usage scenarios and / or terminal statistics corresponding to the clusters. Based on all the second feature vectors in the same cluster, obtain the feature vector located at the centroid of the cluster, and the feature vector located at the centroid of the cluster is the first feature vector; The first feature vector is processed using a preset frequency response simulation model to obtain a preset frequency response feature curve corresponding to the first feature vector. The usage conditions corresponding to the cluster to which the first feature vector belongs are used as the second usage conditions of the preset frequency response feature curve.
3. The method according to claim 2, characterized in that, The second feature vector is obtained based on the third feature vector of the first audio signal output by the terminal under the third usage condition, including: Multiple audio data pairs are acquired for each third usage condition. The audio data pairs include a first standard audio signal and a first audio signal. The first standard audio signal and the first audio signal in the same audio data pair are used to calculate the frequency response characteristic curve under the third usage condition. Feature extraction is performed on each first audio signal to obtain a third feature vector for each first audio signal; The third feature vector of each first audio signal is processed using a preset coding model to obtain the second feature vector of each first audio signal. The preset coding model is trained based on the frequency response feature curve under the third usage condition and the second feature vector of each first audio signal.
4. The method according to claim 1, characterized in that, The step of obtaining a first frequency response feature curve matching the first usage condition from all preset frequency response feature curves based on the first usage condition of the audio processing model and the second usage condition of each preset frequency response feature curve includes: Obtain the matching degree between the first usage condition and each second usage condition; The preset frequency response feature curve corresponding to the second usage condition that meets the preset matching condition is used as the first frequency response feature curve.
5. The method according to claim 4, characterized in that, The first usage condition includes usage scenario and terminal information; the second usage condition includes usage scenario, number of covered terminals and model distribution information, wherein the model distribution information is used to indicate the terminal information corresponding to the cluster to which the second usage condition belongs. The process of obtaining the matching degree between the first usage condition and each second usage condition includes: If the usage scenario and / or terminal information in the first usage condition are recorded in the second usage condition, the matching degree between the first usage condition and the second usage condition is determined based on at least one of the model distribution information and the number of covered terminals.
6. The method according to claim 1 or 4, characterized in that, If the first frequency response characteristic curve is not obtained according to the first and second usage conditions, the method further includes: obtaining a second audio signal output by the terminal after the second standard audio signal is processed by the terminal under the first usage conditions; The frequency response characteristic curve of the terminal is obtained based on the second standard audio signal and the second audio signal; Calculate the similarity between the frequency response characteristic curve of the terminal and each preset frequency response characteristic curve; Based on the similarity between the frequency response characteristic curve of the terminal and each preset frequency response characteristic curve, and the second usage conditions, at least one first frequency response characteristic curve is determined from the preset frequency response characteristic curves.
7. The method according to claim 6, characterized in that, After obtaining the second audio signal output by the terminal after processing the second standard audio signal under the first usage conditions, the method further includes: Feature extraction is performed on the second audio signal to obtain a third feature vector of the second audio signal; The third feature vector of the second audio signal is processed using a preset encoding model to obtain the second feature vector of the second audio signal. The first usage condition corresponding to the second audio signal is the third usage condition of the second feature vector of the second audio signal. The preset frequency response feature curve and the second usage condition corresponding to the preset frequency response feature curve are updated using the second feature vector of the second audio signal and the third usage condition of the second feature vector of the second audio signal.
8. The method according to claim 1, characterized in that, After obtaining the first frequency response characteristic curve, the method further includes: determining a first number of standard audio signals processed by the first frequency response characteristic curve and a second number of standard audio signals processed by the second frequency response characteristic curve, wherein the second frequency response characteristic curve is a preset frequency response characteristic curve that does not match the first usage conditions, and the first number is greater than the second number. The step of processing the standard audio signal using at least the first frequency response characteristic curve to obtain the simulated audio signal includes: processing the first number of standard audio signals using the first frequency response characteristic curve to obtain the first number of simulated audio signals. The second number of standard audio signals are processed using the second frequency response characteristic curve to obtain the second number of simulated audio signals.
9. The method according to claim 8, characterized in that, The determination of the first number of standard audio signals processed by the first frequency response characteristic curve and the second number of standard audio signals processed by the second frequency response characteristic curve includes: A first weight of the first frequency response characteristic curve and a second weight of the second frequency response characteristic curve are determined. The first weight is greater than the second weight. Either the first weight or the second weight is used to indicate the number of standard audio signals processed by the frequency response characteristic curve corresponding to that weight. The relationship between the first weight and the second weight is used to represent the relationship between the first quantity and the second quantity.
10. A data processing apparatus, characterized in that, include: The acquisition unit is used to acquire the first usage conditions of the audio processing model and the second usage conditions of each preset frequency response characteristic curve. The preset frequency response characteristic curve is obtained in advance using the first feature vector under the second usage conditions. The first feature vector is a feature vector representing the common characteristics of a class of terminals obtained from the second feature vectors under multiple third usage conditions. The second feature vector is obtained from the third feature vector of the first audio signal output by the terminal under the third usage conditions. The second usage conditions are obtained from the multiple third usage conditions. The matching unit is configured to obtain a first frequency response feature curve that matches the first usage condition from all preset frequency response feature curves, based on the first usage condition of the audio processing model and the second usage condition of each preset frequency response feature curve; The processing unit is configured to process the standard audio signal using at least the first frequency response characteristic curve to obtain a simulated audio signal, wherein the simulated audio signal is the training data of the audio processing model.
11. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to enable the electronic device to implement the data processing method as described in any one of claims 1 to 9.