Sound pressure level calibration methods, devices, equipment, chips, and storage media
By regularizing the speech segments to be processed to conform to a specific distribution before sound pressure level calibration, the calibration deviation caused by human factors is solved, and more accurate sound pressure level calibration of human speech corpora is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-03
- Publication Date
- 2026-03-06
AI Technical Summary
In existing technologies, the sound pressure level calibration process of human speech corpora is affected by factors such as the volume of different people speaking, the distance and angle between the mouth and the microphone, etc., resulting in insufficient calibration accuracy and failure to correctly reflect the long-term statistical level of human speech corpora.
If the first parameter value of the speech segment to be processed is not found to conform to the first distribution, it is normalized to conform to the second distribution before sound pressure level calibration is performed to ensure the accuracy of the target calibration result.
It improves the accuracy of sound pressure level calibration, enabling the target calibration results to accurately reflect the long-term statistical level of human speech data and reducing calibration deviations caused by human factors.
Smart Images

Figure CN116364110B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of acoustic technology, and in particular to a sound pressure level calibration method, apparatus, device, chip, and storage medium. Background Technology
[0002] In related technologies, the development of acoustic experiments such as Keyword Spotting (KWS) wake-up testing, speech recognition, and beamforming audio data acquisition has increased the difficulty of calibrating the signal-to-noise ratio (SNR) of human speech corpora. To reduce the difficulty of SNR calibration for human speech corpora, the sound pressure level of the human speech corpora can be calibrated before calibrating the SNR.
[0003] Considering that various factors, such as differences in speaking volume between individuals and variations in the distance and angle between the mouth and microphone, can affect the accuracy of sound pressure level calibration during data acquisition, improving the accuracy of sound pressure level calibration is a crucial issue that needs to be addressed. Summary of the Invention
[0004] This application provides a sound pressure level calibration method, apparatus, device, chip, and storage medium, which can improve the accuracy of sound pressure level calibration.
[0005] In a first aspect, embodiments of this application provide a sound pressure level calibration method, the method comprising:
[0006] Determine M speech segments to be processed, and calculate the first parameter value for each of the M speech segments to be processed;
[0007] If the first parameter value of at least one of the M speech segments to be processed does not conform to the first distribution, then the M speech segments to be processed are normalized so that the first parameter values of the N target speech segments obtained after normalization conform to the second distribution.
[0008] The sound pressure levels of the M speech segments to be processed are calibrated according to the second distribution to determine the target calibration result;
[0009] Where M and N are both positive integers, and N is less than or equal to M.
[0010] Secondly, embodiments of this application provide a sound pressure level calibration device, which includes a calculation unit, a normalization unit, and a determination unit, wherein:
[0011] The calculation unit is configured to determine M speech segments to be processed and calculate the first parameter value of each of the M speech segments to be processed.
[0012] The normalization unit is configured to normalize the M speech segments if the first parameter value of at least one of the M speech segments to be processed does not conform to the first distribution, so that the first parameter values of the N target speech segments obtained after normalization conform to the second distribution.
[0013] The unit is configured to calibrate the sound pressure levels of M speech segments to be processed according to the second distribution, and determine the target calibration result;
[0014] Where M and N are both positive integers, and N is less than or equal to M.
[0015] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to perform the sound pressure level calibration method described in the first aspect.
[0016] Fourthly, embodiments of this application provide a chip for implementing the sound pressure level calibration method described in the first aspect.
[0017] Specifically, the chip includes a processor for calling and running a computer program from a memory, causing a device equipped with the chip to perform the sound pressure level calibration method described in the first aspect above.
[0018] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by at least one processor, implements the sound pressure level calibration method described in the first aspect.
[0019] Sixthly, embodiments of this application provide a computer program product, including computer program instructions that cause a computer to execute the sound pressure level calibration method described in the first aspect.
[0020] In a seventh aspect, embodiments of this application provide a computer program that, when run on a computer, causes the computer to execute the sound pressure level calibration method described in the first aspect.
[0021] This application provides a sound pressure level calibration method. First, the first parameter value of each of the M speech segments to be processed is calculated. Then, if the first parameter value of at least one of the M speech segments does not conform to a first distribution, the M speech segments are normalized so that the first parameter values of the resulting N target speech segments conform to a second distribution. Finally, the sound pressure level of the M speech segments is calibrated according to the second distribution to determine the target calibration result. Thus, when the first parameter value of at least one of the M speech segments does not conform to the first distribution, normalization can be used to ensure that the first parameter values of the resulting N target speech segments conform to the second distribution. This eliminates the restriction on whether the first parameter values of the M speech segments conform to the first distribution, maintaining the diversity of the first parameter values of the M speech segments. Calibrating the sound pressure level of the M speech segments according to the second distribution improves the accuracy of the sound pressure level calibration, allowing the target calibration result to accurately reflect the long-term statistical level of the human speech corpus. Attached Figure Description
[0022] Figure 1 This is a test scenario diagram of a sound pressure level calibration result provided in an embodiment of this application.
[0023] Figure 2 This is a test block diagram of a sound pressure level calibration result provided in an embodiment of this application.
[0024] Figure 3 It is the probability density distribution of the root mean square amplitude of the same batch of human voice corpora.
[0025] Figure 4 This is a schematic flowchart of a sound pressure level calibration method provided in an embodiment of this application.
[0026] Figure 5A This is a schematic diagram of a scene before human voice corpus is detected and processed using speech activity detection, as provided in an embodiment of this application.
[0027] Figure 5B This is a schematic diagram of a scene after human voice corpus is detected and processed using speech activity detection, as provided in an embodiment of this application.
[0028] Figure 6 This is a schematic diagram showing the effect of filtering on different types of corpora provided in an embodiment of this application.
[0029] Figure 7 This is a detailed flowchart illustrating a sound pressure level calibration method provided in an embodiment of this application. Figure 1 .
[0030] Figure 8This is a detailed flowchart illustrating a sound pressure level calibration method provided in an embodiment of this application. Figure 2 .
[0031] Figure 9 This is a normal distribution diagram of the root mean square amplitude of each of N target speech segments provided in an embodiment of this application.
[0032] Figure 10 This is another normal distribution diagram of the root mean square amplitude of each of the N target speech segments provided in the embodiments of this application.
[0033] Figure 11 This is a detailed flowchart illustrating a sound pressure level calibration method provided in an embodiment of this application. Figure 3 .
[0034] Figure 12 This is a detailed flowchart illustrating a sound pressure level calibration method provided in an embodiment of this application. Figure 4 .
[0035] Figure 13 This is a schematic diagram of the composition of a sound pressure level calibration device provided in an embodiment of this application.
[0036] Figure 14 This is a schematic structural diagram of an electronic device provided in an embodiment of this application.
[0037] Figure 15 This is a schematic structural diagram of a chip provided in an embodiment of this application. Detailed Implementation
[0038] In order to gain a more detailed understanding of the features and technical content of the embodiments of this application, the implementation of the embodiments of this application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for reference and illustration only and are not intended to limit the embodiments of this application.
[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0040] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0041] It should also be noted that the terms "first, second, and third" used in the embodiments of this application are only used to distinguish similar objects and do not represent a specific order of objects. It is understood that "first, second, and third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0042] Furthermore, in the embodiments of this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0043] To facilitate understanding of the technical solutions of the embodiments of this application, the relevant technologies of the embodiments of this application are described below. The following relevant technologies are optional solutions and can be combined with the technical solutions of the embodiments of this application in any way, and they all fall within the protection scope of the embodiments of this application.
[0044] With the development of acoustic experiments such as KWS wake-up testing, speech recognition, and Beamforming audio data acquisition, the difficulty of calibrating the signal-to-noise ratio (SNR) of human speech corpora has increased. To reduce the difficulty of SNR calibration, the sound pressure level of the human speech corpora can be calibrated before calibrating the SNR.
[0045] Figure 1 This diagram illustrates a test scenario for sound pressure level calibration results provided in an embodiment of this application. Figure 1 As shown, the test scenario for sound pressure level calibration results mainly includes the following modules: a speech playback module 110, a test model 120, and background noise playback modules 130-150. The speech playback module 110 plays human speech data, and the test model 120 detects whether the speech playback module 110 can wake it up, thereby determining whether the sound pressure level calibration result of the human speech data is within a preset range. For example, if the human speech data can wake up the test model 120, it indicates that the sound pressure level calibration result of the human speech data is within the preset range; conversely, if the human speech data cannot wake up the test model 120, it indicates that the sound pressure level calibration result of the human speech data is not within the preset range. The background noise playback modules 130-150 simulate noise in a real test scenario, thereby making the test scenario for sound pressure level calibration results closer to the real test scenario.
[0046] Figure 2 This diagram illustrates a test block diagram for sound pressure level calibration results provided in an embodiment of this application. For example... Figure 2The test block diagram for the sound pressure level calibration results mainly includes the following modules: a corpus playback module 201, a model under test 202, and a background noise playback module 203. The corpus playback module 201 may include an artificial head 204, a power amplifier 205, a sound card 206, and a computer (Personal Computer, PC) 207. The model under test 202 may include a mobile phone 208. The background noise playback module 203 may include speakers 209 to 216.
[0047] First, the artificial head 204 emits human voice data. This data is amplified by a power amplifier 205, and then further converted from analog to digital signals by a sound card 206, allowing the computer 207 to successfully play the voice data. The played voice data is then transmitted to a mobile phone 208 to determine if it can wake up the phone, thus determining whether the sound pressure level calibration result is within a preset range. Speakers 209 to 216 simulate noise in a real test scenario, making the sound pressure level calibration result more closely resemble the actual test scenario.
[0048] When calibrating sound pressure levels using relevant technologies, a batch of human speech data is typically sampled first. Then, the sound pressure levels of each sampled data item are calibrated, and the average of these calibration results is taken as the predicted sound pressure level for the entire batch of human speech. These technologies for sound pressure level calibration have the following main drawbacks:
[0049] (1) The sampling process used in the sound pressure level calibration process will deviate from the overall true situation. The sampling method is very likely to make the calibrated sound pressure level too high or too low.
[0050] (2) The sound pressure level calibration process does not take into account the influence of the silent segment and the noise segment, which will cause the obtained sound pressure level calibration result to deviate from the expected sound pressure level calibration result;
[0051] (3) The sound pressure level calibration process did not take into account the root mean square (RMS) probability density distribution characteristics of the human voice corpus, resulting in a deviation between the sound pressure level calibration result and the expected sound pressure level calibration result. The unit of amplitude RMS is decibel (dB).
[0052] Considering that during the data acquisition process, factors such as differences in the volume of different people's speech, and the inconsistency in the distance and angle between the mouth and the microphone, can lead to uneven distribution of the amplitude RMS of the acquired human speech data, a large range between the minimum and maximum amplitude RMS of the same batch of human speech data, and the existence of silent and noisy segments in the human speech data, the accuracy of sound pressure level calibration is affected.
[0053] For example, Figure 3 This is the probability density distribution of the root mean square amplitude of the same batch of human voice corpora. For example... Figure 3 As shown, the minimum amplitude RMS of this batch of human voice data is approximately -53dB, and the maximum amplitude RMS is approximately -17dB. The difference between the minimum and maximum amplitude RMS is several decibels. If the power amplifier uses the same gain drive, it will affect the accuracy of the sound pressure level calibration, causing the sound pressure level calibration results to fail to correctly reflect the long-term statistical level of the human voice data, thus making it difficult to utilize... Figure 1 The test results obtained when testing the sound pressure level calibration results in the test scenario shown are inaccurate.
[0054] Based on this, this application provides a sound pressure level calibration method. First, the first parameter value of each of the M speech segments to be processed is calculated. Then, if the first parameter value of at least one of the M speech segments does not conform to a first distribution, the M speech segments are normalized so that the first parameter values of the resulting N target speech segments conform to a second distribution. Finally, the sound pressure level of the M speech segments is calibrated according to the second distribution to determine the target calibration result. Thus, when the first parameter value of at least one of the M speech segments does not conform to the first distribution, normalization can be used to make the first parameter values of the resulting N target speech segments conform to the second distribution. This eliminates the restriction on whether the first parameter values of the M speech segments conform to the first distribution, maintaining the diversity of the first parameter values of the M speech segments. Calibrating the sound pressure level of the M speech segments according to the second distribution improves the accuracy of the sound pressure level calibration, allowing the target calibration result to correctly reflect the long-term statistical level of the human speech corpus.
[0055] Figure 4 This is a schematic flowchart of a sound pressure level calibration method provided in an embodiment of this application, as shown below. Figure 4 As shown, the method may include the following steps.
[0056] S410, determine M speech segments to be processed, and calculate the first parameter value of each of the M speech segments to be processed.
[0057] Where M is a positive integer.
[0058] It should be noted that, in the embodiments of this application, the sound pressure level calibration method can be applied to a sound pressure level calibration device or an electronic device integrated with such a device. The electronic device can be implemented in various forms, such as smartphones, tablets, laptops, handheld computers, personal digital assistants (PDAs), portable media players (PMPs), etc., and is not limited thereto.
[0059] It should also be noted that the first parameter value can be the amplitude RMS or other parameter values. This application does not limit this. The following embodiments mainly use the amplitude RMS as the first parameter value for illustration.
[0060] For example, taking a speech segment to be processed as an example, the formula for calculating the RMS amplitude of the speech segment to be processed is as follows:
[0061] amplitude
[0062] Where L is the total number of samples in the speech segment to be processed, and x l Let be the amplitude of the l-th sample.
[0063] In some embodiments, determining M speech segments to be processed may include: acquiring M corpora; performing detection processing on the M corpora according to Voice Activity Detection (VAD) to determine the speech position information of each of the M corpora; and determining the M speech segments to be processed according to the speech position information of each of the M corpora.
[0064] The speech location information for each of the M corpora may include the start and end positions of the speech for each of the M corpora.
[0065] It should be noted that the M corpora can be M human voice corpora.
[0066] Furthermore, when acquiring M corpora, multiple wake words spoken by the same person at different speeds will be combined into a complete corpus.
[0067] It should also be noted that by using VAD to detect and process M corpora, the silent segments and most of the noise segments in each of the M corpora can be removed. Based on the above processing, the start and end positions of the speech in each of the M corpora can be determined, thereby identifying the M speech segments to be processed.
[0068] For example, Figure 5A This is a schematic diagram of a scene before processing human voice corpus using speech activity detection, as provided in an embodiment of this application. Figure 5B This is a schematic diagram of a scene after human voice corpus is detected and processed using speech activity detection, as provided in an embodiment of this application. Figure 5A The human voice corpus in the text not only contains speech segments, but also silence segments and noise segments; and Figure 5A compared to, Figure 5B The human voice corpus has removed silent segments and most noise segments, which is more conducive to the calculation of the RMS amplitude of speech segments.
[0069] By using VAD to detect and process M corpora, not only can the sound pressure level calibration be made more automated, but the time for sound pressure level calibration can also be saved, and the consistency of the sound pressure level calibration test process can be guaranteed.
[0070] S420, if the first parameter value of at least one of the M speech segments to be processed does not conform to the first distribution, then the M speech segments to be processed are normalized so that the first parameter values of the N target speech segments obtained after normalization conform to the second distribution.
[0071] Where N are all positive integers, and N is less than or equal to M.
[0072] It should be noted that the first distribution and / or the second distribution can be a normal distribution or other distributions. This application does not limit this. The following embodiments mainly use the first distribution and / or the second distribution as normal distributions as examples for illustration.
[0073] It should also be noted that the first distribution and the second distribution can be the same distribution or different distributions, and need to be distinguished according to the actual application scenario.
[0074] In some embodiments, the method further includes: determining the probability density distribution of the first parameter values of each of the M speech segments to be processed; and determining whether the first parameter values of each of the M speech segments to be processed conform to the first distribution based on the probability density distribution of the first parameter values of each of the M speech segments to be processed.
[0075] It should be noted that, in this case, the first distribution can be determined by the probability density distribution of the M speech segments to be processed.
[0076] For example, taking the first distribution as the first normal distribution, the expected value, standard deviation, and confidence interval of the first parameter value of most of the M speech segments to be processed can be determined by the probability density distribution of the first parameter value of each of the M speech segments to be processed; then the first distribution can be generated by the expected value, standard deviation, and confidence interval of the first parameter value of most of the speech segments.
[0077] S430, calibrate the sound pressure levels of the M speech segments to be processed according to the second distribution, and determine the target calibration result.
[0078] It should be understood that calibrating the sound pressure level of M speech segments to be processed can be understood as calibrating the sound pressure level of N target speech segments, and the sound pressure level calibration results of the N target speech segments can be used to represent the target calibration results of the M speech segments to be processed.
[0079] In some embodiments, calibrating the sound pressure levels of M speech segments to be processed according to the second distribution and determining the target calibration result may include: dividing the second distribution into intervals to obtain at least one second sub-interval, and determining the median of the at least one second sub-interval; averaging the medians of the at least one second sub-interval to determine a first mean; and calibrating the sound pressure levels of the M speech segments to be processed according to the first mean to determine the target calibration result.
[0080] Considering that directly calibrating the first mean using a sound pressure meter would lead to unstable target calibration results, in some embodiments, calibrating the sound pressure levels of the M speech segments to be processed based on the first mean to determine the target calibration result may include: determining the initial parameter values of the M speech segments to be processed; calibrating the sound pressure levels of the M speech segments to be processed based on the first mean and the initial parameter values to determine the target calibration result.
[0081] The initial parameter values for the M speech segments to be processed can be the initial parameter values for white noise in the M speech segments to be processed.
[0082] For example, the initial parameter value can be the initial amplitude RMS.
[0083] In some embodiments, calibrating the sound pressure levels of the M speech segments to be processed based on the first mean and the initial parameter values to determine the target calibration result may include: calibrating the sound pressure levels of the M speech segments to be processed based on the difference between the first mean and the initial parameter values to determine the target calibration result.
[0084] For example, assuming the first mean is X0 and the initial parameter value is Z0, the sound pressure level of the M speech segments to be processed can be calibrated based on X0-Z0, thereby determining the target calibration result.
[0085] Furthermore, considering that the effects of filtering (weighting) on speech segments and white noise differ when using a sound pressure meter to calibrate the sound pressure level, the effects of filtering must also be considered when using white noise to calibrate the sound pressure level of a speech segment.
[0086] For example, the filtering process can be performed by a bandpass filter.
[0087] For example, taking the first parameter value as the amplitude RMS, Figure 6 This is a schematic diagram illustrating the impact of a filtering process on different types of corpora, provided as an embodiment of this application. Figure 6 As shown, corpora A, C, and K are all processed by filtering. The filtering process has the greatest impact on the RMS amplitude of corpora A, followed by the RMS amplitude of corpora C, and the least impact on the RMS amplitude of corpora K.
[0088] In some embodiments, calibrating the sound pressure levels of M speech segments to be processed based on a first mean and initial parameter values to determine a target calibration result may include: filtering the initial parameter values to determine a third parameter value; subtracting the third parameter value from the initial parameter value to determine a second difference; filtering the first mean to determine a fourth parameter value; subtracting the fourth parameter value from the second difference to determine a third difference; and calibrating the sound pressure levels of the M speech segments to be processed based on the third difference to determine the target calibration result.
[0089] For example, assuming the initial parameter value is Z0, the third parameter value determined by filtering Z0 is Z1, then the second difference is Z1-Z0; the first mean is X0, the fourth parameter value determined by filtering X0 is X1, then the formula for calculating the third difference is as follows, so that the sound pressure level of the M speech segments to be processed can be calibrated according to formula (2), thereby determining the target calibration result.
[0090] The third difference = X1 - (Z1 - Z0) (2)
[0091] In this embodiment of the application, the method may further include: if the first parameter values of each of the M speech segments to be processed conform to a first distribution, then calibrating the sound pressure level of the M speech segments to be processed according to the first distribution, and determining the target calibration result.
[0092] It should be noted that since the first parameter values of the M speech segments to be processed already conform to the first distribution, there is no need to perform normalization processing on the M speech segments to be processed. In this case, the sound pressure level calibration can be performed directly on the sound pressure level of these M speech segments to be processed according to the first distribution, thereby determining the target calibration result.
[0093] This application provides a sound pressure level calibration method. When the first parameter value of at least one of the M speech segments to be processed does not conform to a first distribution, the M speech segments to be processed can be normalized so that the first parameter values of the N target speech segments obtained after normalization conform to a second distribution. The sound pressure level of the M speech segments to be processed is then calibrated according to the second distribution to determine the target calibration result. Thus, when the first parameter value of at least one of the M speech segments to be processed does not conform to the first distribution, normalization can be performed to make the first parameter values of the N target speech segments after normalization conform to the second distribution. This eliminates the restriction on whether the first parameter values of the M speech segments to be processed conform to the first distribution, maintaining the diversity of the first parameter values of the M speech segments to be processed. Calibrating the sound pressure level of the M speech segments to be processed according to the second distribution improves the accuracy of the sound pressure level calibration, allowing the target calibration result to correctly reflect the long-term statistical level of the human speech corpus.
[0094] Based on the aforementioned embodiments, there are three ways to normalize the M speech segments to be processed so that the first parameter values of the N target speech segments each conform to the second distribution. These methods are described below in conjunction with the specific implementations. Figures 7-11 Please provide an explanation.
[0095] Method #A: Delete at least one speech segment that does not conform to the first distribution, so that the first parameter values of each of the N target speech segments conform to the second distribution.
[0096] Figure 7 This is a detailed flowchart illustrating a sound pressure level calibration method provided in an embodiment of this application. Figure 1 ,like Figure 7 As shown, the method may include the following steps:
[0097] S710, delete at least one pending voice segment;
[0098] S720, take the speech segments other than at least one of the M speech segments to be processed as N target speech segments, so that the first parameter values of each of the N target speech segments conform to the second distribution.
[0099] It should be noted that the first distribution and the second distribution are the same distribution at this time.
[0100] It should also be noted that, since the first parameter value of at least one speech segment to be processed does not conform to the first distribution, after deleting at least one speech segment to be processed, the first parameter values of the speech segments (N target speech segments) other than at least one speech segment to be processed among the M speech segments to be processed conform to the first distribution.
[0101] In some embodiments, the number of at least one speech segment to be processed is less than or equal to a first threshold.
[0102] It should be noted that the first threshold can be set manually or in other ways, and this application embodiment does not limit this.
[0103] It should also be noted that the first threshold is usually a small value. In this case, deleting at least one speech segment to be processed can ensure that the first parameter values of the remaining speech segments out of the M speech segments to be processed conform to the first distribution, with minimal impact on the target calibration result. For example, if the first threshold is 2 and the total number of speech segments to be processed is 1000, deleting 1 or 2 speech segments from the 1000 speech segments to be processed will have almost no impact on the final target calibration result.
[0104] Based on method #A, by deleting at least one speech segment to be processed, the first parameter values of each of the N target speech segments can conform to the second distribution. This improves the accuracy of sound pressure level calibration when subsequently calibrating the sound pressure level of the M speech segments to be processed according to the second distribution, and further enables the target calibration results to correctly reflect the long-term statistical level of the human voice corpus.
[0105] Method #B: Normalize at least one speech segment that does not conform to the first distribution so that the first parameter values of each of the N target speech segments conform to the second distribution.
[0106] Figure 8 This is a detailed flowchart illustrating a sound pressure level calibration method provided in an embodiment of this application. Figure 2 ,like Figure 8 As shown, the method may include the following steps:
[0107] S810, Divide the first distribution into intervals to obtain at least one first sub-interval, and determine the median of each of the at least one first sub-interval;
[0108] S820, perform normalization processing on at least one speech segment to be processed based on the median of at least one first sub-interval, and determine at least one target speech segment;
[0109] S830, take at least one target speech segment and the speech segments other than at least one speech segment from the M speech segments to be processed as N target speech segments, so that the first parameter values of each of the N target speech segments conform to the second distribution.
[0110] It should be noted that the first distribution and the second distribution are the same distribution at this time.
[0111] It should also be noted that by normalizing at least one speech segment to be processed using the median of at least one first sub-interval, the first parameter value of at least one target speech segment obtained conforms to the first distribution. Furthermore, the first parameter value of at least one target speech segment, as well as the first parameter values of each of the speech segments in the M speech segments to be processed (excluding at least one speech segment to be processed), all conform to the first distribution.
[0112] For example, suppose the first distribution is a first normal distribution and the first parameter value is the amplitude RMS, such as Figure 9 As shown, the x-axis of the first normal distribution is the amplitude RMS (unit: dB), the y-axis is the probability density distribution, the expected value is μ, the standard deviation is σ, and the confidence interval is (a, b).
[0113] For example, such as Figure 10 As shown, in Figure 9 Based on the normal distribution plot shown, the interval of the first normal distribution is divided into P equal parts to obtain the sub-intervals of the first normal distribution. After P equal division, the RMS range of the amplitude of each sub-interval is shown below ( Figure 10 (The shaded area in the image):
[0114]
[0115] The formula for calculating the median of each subinterval is shown below:
[0116] amplitude
[0117] In some embodiments, normalizing at least one speech segment to be processed based on the median of at least one first sub-interval to determine at least one target speech segment may include: subtracting the first parameter value of the first speech segment to be processed from the median of at least one first sub-interval to determine at least one first difference value, and determining the minimum value among the absolute values of the at least one first difference value as a second parameter value; calculating a first gain value based on the median of the sub-interval to which the second parameter value belongs and the first parameter value of the first speech segment to be processed; and determining the first parameter value of the first target speech segment based on the first gain value and the first parameter value of the first speech segment to be processed.
[0118] The first speech segment to be processed is one of at least one speech segment to be processed, and the first target speech segment is one of at least one target speech segment.
[0119] For example, such as Figure 10 As shown, assuming the first parameter value of the first speech segment to be processed is y and the first difference value is D, the formula for calculating the first difference value D is as follows:
[0120]
[0121] Assume the value of k corresponding to the second parameter is k. best Then the median amplitude of the subinterval to which the second parameter value belongs. The calculation formula is as follows:
[0122] amplitude
[0123] In some embodiments, calculating a first gain value based on the median of the sub-interval to which the second parameter value belongs and the first parameter value of the first speech segment to be processed may include: performing a ratio operation on the median of the sub-interval to which the second parameter value belongs and the first parameter value of the first speech segment to be processed to determine the first gain value.
[0124] For example, such as Figure 10 As shown, assume that the first parameter value of the first speech segment to be processed is y, and the median value of the sub-interval to which the second parameter value belongs is the amplitude. The formula for calculating the first gain value m is shown below. At this point, the RMS amplitude of the first target speech segment is the amplitude.
[0125]
[0126] Furthermore, in some embodiments, determining the first parameter value of the first target speech segment based on the first gain value and the first parameter value of the first speech segment to be processed may include: multiplying the first gain value and the first parameter value of the first speech segment to be processed to determine the first parameter value of the first target speech segment.
[0127] It should be noted that, through the above processing method, the first parameter value of the first speech segment to be processed can be normalized to the first parameter value of the first target speech segment. By using a method similar to that used for the first speech segment to be processed, the other speech segments to be processed in at least one speech segment to be processed, excluding the first speech segment to be processed, can also be normalized to the median of the corresponding sub-interval of the first distribution, thereby ensuring that the first parameter value of the at least one target speech segment conforms to the first distribution.
[0128] Based on method #B, by normalizing at least one speech segment to be processed that does not conform to the first distribution, the first parameter values of each of the N target speech segments conform to the second distribution. Thus, when calibrating the sound pressure level of the M speech segments to be processed according to the second distribution, the accuracy of the sound pressure level calibration can be improved, and the target calibration results can correctly reflect the long-term statistical level of the human voice corpus.
[0129] Method #C: Normalize the M speech segments to be processed so that the first parameter values of the N target speech segments conform to the second distribution.
[0130] Figure 11 This is a detailed flowchart illustrating a sound pressure level calibration method provided in an embodiment of this application. Figure 3 ,like Figure 11 As shown, the method may include the following steps:
[0131] S1110, Divide the second distribution into intervals to obtain at least one second sub-interval, and determine the median of at least one second sub-interval;
[0132] S1120, Normalize the M speech segments to be processed based on the median of at least one second sub-interval to determine the M target speech segments;
[0133] S1130, M target speech segments are taken as N target speech segments, so that the first parameter values of each of the N target speech segments conform to the second distribution.
[0134] It should be noted that by normalizing the M speech segments to be processed using the median of at least one second sub-interval, the first parameter values of the resulting M target speech segments each conform to the second distribution. Thus, when the M target speech segments are used as N target speech segments, the first parameter values of the N target speech segments each also conform to the second distribution.
[0135] It should also be noted that the first distribution and the second distribution are different distributions at this time.
[0136] For example, assuming the second distribution is a second normal distribution and the first parameter value is the amplitude RMS, the amplitude RMS range of each sub-interval of the second normal distribution can be referred to the aforementioned formula (3), and the calculation formula of the median of each sub-interval of the second normal distribution can be referred to the aforementioned formula (4).
[0137] It should also be noted that the method of normalizing the M speech segments to be processed based on the median of at least one second sub-interval to determine the M target speech segments is similar to the method of normalizing at least one speech segment to be processed based on the median of at least one first sub-interval to determine at least one target speech segment in the aforementioned embodiment, and will not be described again here.
[0138] Using the above processing method, the M speech segments to be processed can be normalized, thereby normalizing the first parameter values of the M speech segments to be processed to the median of the corresponding sub-interval of the second distribution, so that the first parameter values of the obtained M target speech segments conform to the second distribution.
[0139] In some embodiments, the method further includes: determining Q test speech segments and calculating a first parameter value for each of the Q test speech segments; and determining a second distribution based on the expected value, standard deviation, and confidence interval of the first parameter value for each of the Q test speech segments.
[0140] Where Q is a positive integer.
[0141] It should be noted that this second distribution is not related to the probability density distribution of the M speech segments to be processed. The Q test speech segments are speech segments different from the M speech segments to be processed. Determining the second distribution through the Q test speech segments ensures that the first parameter values of the M target speech segments obtained after normalizing the M speech segments to be processed conform to the ideal second distribution, thus better meeting practical needs.
[0142] In other embodiments, the second distribution satisfies a preset distribution model.
[0143] It should be noted that when the second distribution satisfies the preset distribution model, the first parameter values of the obtained M target speech segments can better conform to the ideal second distribution and better meet the actual needs.
[0144] Based on method #C, by normalizing the M speech segments to be processed, the first parameter values of the N target speech segments each conform to the second distribution. Therefore, when calibrating the sound pressure level of the M speech segments to be processed according to the second distribution, the accuracy of the sound pressure level calibration can be improved, and the target calibration results can correctly reflect the long-term statistical level of the human voice corpus.
[0145] The sound pressure level calibration method provided in the above embodiments will be described in detail below in conjunction with specific application scenarios.
[0146] The embodiments of this application can segment M speech segments to be processed using VAD and automatically calculate the probability density distribution of the amplitude RMS of the M speech segments to be processed, so as to normalize the M speech segments to be processed to a reasonable distribution (the following example is based on the normal distribution), and then use equivalent white noise to calibrate the sound pressure level for speech testing.
[0147] The core of this application's embodiment lies in utilizing VAD to batch process M corpora, identifying the start and end positions of the speech in each of the M corpora as M speech segments to be processed, and calculating the RMS amplitude of each of the M speech segments. During the recording of the M corpora, multiple wake words spoken by the same person at different speeds are combined into a complete corpus, and the application of VAD significantly reduces the impact of silent segments. Based on the expectation, standard deviation, and confidence interval of the set of speech segments to be processed (such as M speech segments to be processed, or Q test speech segments), a normal distribution (such as a first normal distribution, or a second normal distribution) is established. The M speech segments to be processed are then normalized into this normal distribution according to certain rules, and equivalent white noise is used for sound pressure level calibration. Figure 12 This is a detailed flowchart illustrating a sound pressure level calibration method provided in an embodiment of this application. Figure 4 ,like Figure 12As shown, the method may include the following steps:
[0148] S1210 uses VAD to batch process the corpus and find the start and end positions of each speech segment to be processed.
[0149] S1220, calculate the RMS amplitude of each speech segment to be processed, and statistically analyze the probability density distribution of the RMS amplitude of the speech segment to be processed.
[0150] The formula for calculating the RMS amplitude of each speech segment to be processed can be found in the aforementioned formula (1), and will not be repeated here.
[0151] S1230, based on the expected value, standard deviation, and confidence interval of the amplitude RMS of each speech segment to be processed, formulate the normal distribution that the amplitude RMS of each target speech segment conforms to after normalization.
[0152] S1240, divide the confidence interval P into equal parts. The number of target speech segments in a subinterval is the product of the area of the subinterval and the total number of target speech segments. The RMS amplitude of the target speech segments in a subinterval is the median of the subinterval.
[0153] It should be noted that the area of a subinterval is the ratio of the area of that subinterval to the total area.
[0154] The RMS range of each sub-interval can be referred to the aforementioned formula (3), and the formula for calculating the median of each sub-interval can be referred to the aforementioned formula (4), which will not be repeated here.
[0155] For example, such as Figure 10 As shown, assuming the total number of target speech segments is X, the formula for calculating the number of target speech segments within a sub-interval is as follows:
[0156]
[0157] Where f(t) is the normal distribution function, Let be the area of the sub-interval.
[0158] To simplify the calculation, in practical applications, the area of irregular subintervals can be approximated by rectangles, so formula (8) can be simplified as follows:
[0159]
[0160] S1250, based on the principle of minimizing the difference (absolute value) between the amplitude RMS of the speech segment to be processed before normalization and the median of each sub-interval after normalization, the amplitude RMS of each speech segment to be processed is normalized to the median of the corresponding sub-interval according to different proportions to obtain each target speech segment.
[0161] The formula for calculating the RMS amplitude of each target speech segment can be found in the formulas (5) to (7) mentioned above, and will not be repeated here.
[0162] S1260 uses equivalent white noise to calibrate the normalized target speech segment and determines the target calibration result.
[0163] It should be noted that when calibrating sound pressure level using a sound pressure meter, the weighting method (filtering) has different effects on the target speech segment and white noise. Therefore, when calibrating the sound pressure level of the target speech segment using equivalent white noise, the influence of the weighting method also needs to be considered. The following example uses the A-weighted measurement method. The amplitude RMS (first mean) of the target speech segment after normalization is X0, and the amplitude RMS of the target speech segment after A-weighting becomes X1. The amplitude RMS (initial parameter value) of the white noise with the original amplitude RMS (initial parameter value) is Z0, and the amplitude RMS of the target white noise after A-weighting becomes Z1. The formula for calculating the amplitude RMS (third difference) of white noise for calibrating the sound pressure level of the target speech segment can be referred to the aforementioned formula (2), and will not be repeated here.
[0164] When calibrating the sound pressure level of the target speech segment in the anechoic chamber, equivalent calibration is performed using Gaussian white noise with an amplitude RMS of X1-(Z1-Z0) (the third difference).
[0165] The technical solutions provided in this application mainly include the following:
[0166] (1) Use VAD to automatically distinguish between human voice and non-human voice parts of the corpus, and further calculate the RMS of the amplitude of the human voice part and statistically analyze its probability density distribution.
[0167] (2) Determine the expected value, standard deviation and confidence interval of the amplitude RMS of the speech segment to be processed before normalization according to the test requirements, and determine the distribution that the amplitude RMS of the target speech segment after normalization should conform to.
[0168] (3) Based on the principle of minimizing the difference (absolute value) between the amplitude RMS of the speech segment to be processed before normalization and the median of each sub-interval after normalization, calculate the digital gain (first gain value) of the speech segment to be processed, and multiply the amplitude RMS of the speech segment to be processed by the digital gain to complete the normalization of the speech segment to be processed.
[0169] (4) Considering that the weighting method of the sound pressure meter has different effects on the target speech segment and white noise, the sound pressure level of the target speech segment is calibrated using equivalent white noise.
[0170] It should be noted that the embodiments of this application calibrate the sound pressure level of the target speech segment, that is, calibrate the sound pressure level of the speech segment to be processed. The sound pressure level calibration result of the target speech segment can be used to represent the target calibration result of the speech segment to be processed.
[0171] The technical solution provided in this application can automatically calculate the probability density distribution of the amplitude RMS of the speech segment to be processed, normalize the speech segment to be processed according to the test requirements, and use equivalent white noise for sound pressure level calibration. This technical solution effectively solves the problem that the target calibration results cannot accurately reflect the long-term statistical level of the human voice corpus, while maintaining the diversity of the sound pressure level of the human voice corpus (the first parameter value of each of the M speech segments to be processed); at the same time, the use of automated calibration saves calibration time and ensures the consistency of the testing process.
[0172] The preferred embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this application, various simple modifications can be made to the technical solutions of this application, and these simple modifications all fall within the protection scope of this application. For example, the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this application will not describe the various possible combinations separately. Furthermore, various different embodiments of this application can also be arbitrarily combined, as long as they do not violate the spirit of this application, they should also be considered as the content disclosed in this application. Moreover, without conflict, the various embodiments and / or the technical features in the various embodiments described in this application can be arbitrarily combined with the prior art, and the resulting technical solutions should also fall within the protection scope of this application.
[0173] It should also be understood that, in the various method embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0174] Based on the same inventive concept as the foregoing embodiments. Figure 13 This paper illustrates a schematic diagram of the composition of a sound pressure level calibration device provided in an embodiment of this application. Figure 13 As shown, the sound pressure level calibration device 1300 may include a calculation unit 1310, a normalization unit 1320, and a determination unit 1330, wherein:
[0175] The calculation unit 1310 is configured to determine M speech segments to be processed and calculate the first parameter value of each of the M speech segments to be processed.
[0176] The normalization unit 1320 is configured to normalize the M speech segments if the first parameter value of at least one of the M speech segments to be processed does not conform to the first distribution, so that the first parameter values of the N target speech segments obtained after normalization conform to the second distribution.
[0177] Unit 1330 is configured to calibrate the sound pressure levels of M speech segments to be processed according to the second distribution, and determine the target calibration result;
[0178] Where M and N are both positive integers, and N is less than or equal to M.
[0179] In some embodiments, such as Figure 13 As shown, the sound pressure level calibration device 1300 may further include an acquisition unit 1340, wherein:
[0180] Unit 1340 is configured to acquire M corpora.
[0181] The determining unit 1330 is further configured to perform detection processing on M corpora based on speech activity detection, determine the speech position information of each of the M corpora, and determine M speech segments to be processed based on the speech position information of each of the M corpora.
[0182] In some embodiments, such as Figure 13 As shown, the sound pressure level calibration device 1300 may further include a deletion unit 1350, wherein:
[0183] Deletion unit 1350 is configured to delete at least one voice segment to be processed;
[0184] The determining unit 1330 is further configured to take the speech segments other than at least one of the M speech segments to be processed as N target speech segments, so that the first parameter values of each of the N target speech segments conform to the second distribution.
[0185] In some embodiments, such as Figure 13 As shown, the sound pressure level calibration device 1300 may further include a dividing unit 1360, wherein:
[0186] The partitioning unit 1360 is configured to divide the first distribution into intervals to obtain at least one first sub-interval;
[0187] The determining unit 1330 is further configured to determine the median of at least one first sub-interval, and to perform normalization processing on at least one speech segment to be processed based on the median of at least one first sub-interval to determine at least one target speech segment; and to take the at least one target speech segment and the speech segments other than the at least one speech segment from the M speech segments to be processed as N target speech segments, so that the first parameter values of the N target speech segments conform to the second distribution.
[0188] In some embodiments, the determining unit 1330 is further configured to perform subtraction processing on the first parameter value of the first speech segment to be processed and the median of each of at least one first sub-interval, respectively, to determine at least one first difference value, and determine the minimum value among the absolute values of the at least one first difference value as the second parameter value; the calculating unit 1310 is further configured to calculate a first gain value based on the median of the sub-interval to which the second parameter value belongs and the first parameter value of the first speech segment to be processed; the determining unit 1330 is further configured to determine the first parameter value of the first target speech segment based on the first gain value and the first parameter value of the first speech segment to be processed; wherein, the first speech segment to be processed is one of at least one speech segment to be processed, and the first target speech segment is one of at least one target speech segment.
[0189] In some embodiments, the partitioning unit 1360 is further configured to partition the second distribution into intervals to obtain at least one second sub-interval; the determining unit 1330 is further configured to determine the median of at least one second sub-interval, and to perform normalization processing on the M speech segments to be processed based on the median of at least one second sub-interval to determine M target speech segments; and to use the M target speech segments as N target speech segments so that the first parameter values of each of the N target speech segments conform to the second distribution.
[0190] In some embodiments, the calculation unit 1310 is further configured to determine Q test speech segments and calculate the first parameter value of each of the Q test speech segments; the determination unit 1330 is further configured to determine a second distribution based on the expectation, standard deviation and confidence interval of the first parameter value of each of the Q test speech segments; wherein Q is a positive integer.
[0191] In some embodiments, the second distribution satisfies a preset distribution model.
[0192] In some embodiments, the division unit 1360 is further configured to divide the second distribution into intervals to obtain at least one second sub-interval, and determine the median of at least one second sub-interval; the determination unit 1330 is further configured to perform average processing on the median of at least one second sub-interval to determine a first mean; and to calibrate the sound pressure level of the M speech segments to be processed according to the first mean to determine the target calibration result.
[0193] In some embodiments, the determining unit 1330 is further configured to determine the initial parameter values of the M speech segments to be processed; calibrate the sound pressure levels of the M speech segments to be processed based on the first mean and the initial parameter values; and determine the target calibration result.
[0194] In some embodiments, the determining unit 1330 is further configured to: filter the initial parameter value to determine a third parameter value; subtract the third parameter value from the initial parameter value to determine a second difference; filter the first mean value to determine a fourth parameter value; subtract the fourth parameter value from the second difference to determine a third difference; and calibrate the sound pressure level of the M speech segments to be processed based on the third difference to determine the target calibration result.
[0195] In some embodiments, the determining unit 1330 is further configured to, if the first parameter values of each of the M speech segments to be processed conform to the first distribution, calibrate the sound pressure level of the M speech segments to be processed according to the first distribution, and determine the target calibration result.
[0196] In some embodiments, such as Figure 13 As shown, the sound pressure level calibration device 1300 may further include a judgment unit 1370, wherein:
[0197] The determining unit 1330 is also configured to determine the probability density distribution of the first parameter value of each of the M speech segments to be processed;
[0198] The judgment unit 1370 is configured to determine whether the first parameter values of the M speech segments to be processed conform to the first distribution based on the probability density distribution of the first parameter values of the M speech segments to be processed.
[0199] In some embodiments, the first parameter value is the root mean square of the amplitude.
[0200] This application provides a sound pressure level calibration device. When the first parameter value of at least one of the M speech segments to be processed does not conform to a first distribution, it can perform normalization processing to make the first parameter values of each of the normalized N target speech segments conform to a second distribution. This eliminates the restriction on whether the first parameter values of each of the M speech segments to be processed conform to the first distribution, thus maintaining the diversity of the first parameter values of each of the M speech segments to be processed. By calibrating the sound pressure level of the M speech segments to be processed according to the second distribution, the accuracy of the sound pressure level calibration can be improved, so that the target calibration results can correctly reflect the long-term statistical level of the human voice corpus.
[0201] Those skilled in the art should understand that the description of the sound pressure level calibration device in the embodiments of this application can be understood with reference to the description of the sound pressure level calibration method in the embodiments of this application.
[0202] Figure 14 This is a schematic structural diagram of an electronic device 1400 provided in an embodiment of this application. Figure 14As shown, the electronic device 1400 includes a processor 1410 and a memory 1420. The memory 1420 can store computer programs, and the processor 1410 can call and run the computer programs from the memory 1420 to implement the methods in the embodiments of this application.
[0203] The memory 1420 can be a separate device independent of the processor 1410, or it can be integrated into the processor 1410.
[0204] In some embodiments, such as Figure 14 As shown, the electronic device 1400 may also include a transceiver 1430, and the processor 1410 may control the transceiver 1430 to communicate with other devices. Specifically, it may send information or data to other devices or receive information or data sent by other devices.
[0205] The transceiver 1430 may include a transmitter and a receiver. The transceiver 1430 may further include an antenna, and the number of antennas may be one or more.
[0206] In some embodiments, this application also provides another electronic device, wherein the electronic device may include the sound pressure level calibration device 1300 as described in any of the foregoing embodiments.
[0207] Figure 15 This is a schematic structural diagram of a chip provided in an embodiment of this application. For example... Figure 15 As shown, chip 1500 includes processor 1510, which can call and run computer programs from memory to implement the methods in the embodiments of this application.
[0208] In some embodiments, such as Figure 15 As shown, chip 1500 may further include memory 1520. Processor 1510 can retrieve and run computer programs from memory 1520 to implement the methods described in this embodiment.
[0209] The memory 1520 can be a separate device independent of the processor 1510, or it can be integrated into the processor 1510.
[0210] In some embodiments, the chip 1500 may further include an input interface 1530. The processor 1510 can control the input interface 1530 to communicate with other devices or chips; specifically, it can acquire information or data sent by other devices or chips.
[0211] In some embodiments, the chip 1500 may further include an output interface 1540. The processor 1510 can control the output interface 1540 to communicate with other devices or chips, specifically, to output information or data to other devices or chips.
[0212] In some embodiments, the chip can be applied to the electronic device described in this application, which will not be elaborated further here for the sake of brevity.
[0213] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0214] It is understood that the processor in the embodiments of this application may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0215] It is also understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate Synchronous DRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM). It should be noted that the memories described herein are intended to include, but are not limited to, these and any other suitable types of memory.
[0216] It is also understood that the above-described memory is exemplary and not a limiting description. For example, the memory in the embodiments of this application may also be Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate Synchronous DRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Dynch Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM), etc. In other words, the memory in the embodiments of this application is intended to include, but is not limited to, these and any other suitable types of memory.
[0217] This application also provides a computer-readable storage medium for storing computer programs.
[0218] In some embodiments, the computer-readable storage medium may be applied to the electronic device in the embodiments of this application, and when the computer program is executed by at least one processor, it implements the corresponding processes implemented by the electronic device in the various methods of the embodiments of this application. For the sake of brevity, these will not be described in detail here.
[0219] This application also provides a computer program product, including computer program instructions.
[0220] In some embodiments, the computer program product can be applied to the electronic device in the embodiments of this application, and the computer program instructions cause the computer to execute the corresponding processes implemented by the electronic device in the various methods of the embodiments of this application. For the sake of brevity, they will not be described in detail here.
[0221] This application also provides a computer program.
[0222] In some embodiments, the computer program can be applied to the electronic device in the embodiments of this application. When the computer program is run on a computer, it causes the computer to execute the corresponding processes implemented by the electronic device in the various methods of the embodiments of this application. For the sake of brevity, it will not be described in detail here.
[0223] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0224] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0225] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0226] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0227] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0228] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0229] It should be noted that, in this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0230] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0231] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0232] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0233] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0234] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A sound pressure level calibration method, characterized by, The method comprises: determining M pieces of to-be-processed voice segments, and calculating first parameter values of the M pieces of to-be-processed voice segments; the first parameter value is a root mean square of amplitude; if the first parameter value of at least one piece of to-be-processed voice segment in the M pieces of to-be-processed voice segments does not conform to a first distribution, performing normalization processing on the M pieces of to-be-processed voice segments, so that first parameter values of N pieces of target voice segments obtained after normalization conform to a second distribution; the first distribution and the second distribution are normal distributions; the normalization processing comprises: deleting the at least one piece of to-be-processed voice segment that does not conform to the first distribution, or adjusting the first parameter value of the at least one piece of to-be-processed voice segment that does not conform to the first distribution to a median value of at least one first subinterval of the first distribution, or adjusting the first parameter values of the M pieces of to-be-processed voice segments to median values of at least one second subinterval of the second distribution; calibrating sound pressure levels of the M pieces of to-be-processed voice segments according to the second distribution, and determining a target calibration result; wherein M and N are positive integers, and N is less than or equal to M.
2. The method of claim 1, wherein, The determination of the M pieces of to-be-processed voice segments comprises: obtaining M pieces of corpus; detecting and processing the M pieces of corpus according to voice activity detection, and determining voice position information of the M pieces of corpus; determining the M pieces of to-be-processed voice segments according to the voice position information of the M pieces of corpus.
3. The method of claim 1, wherein, The normalization processing on the M pieces of to-be-processed voice segments so that the first parameter values of the N pieces of target voice segments conform to the second distribution comprises: deleting the at least one piece of to-be-processed voice segment, and taking voice segments other than the at least one piece of to-be-processed voice segment in the M pieces of to-be-processed voice segments as the N pieces of target voice segments, so that the first parameter values of the N pieces of target voice segments conform to the second distribution.
4. The method of claim 1, wherein, The normalization processing on the M pieces of to-be-processed voice segments so that the first parameter values of the N pieces of target voice segments conform to the second distribution comprises: dividing the first distribution into intervals to obtain the at least one first subinterval, and determining median values of the at least one first subinterval; performing normalization processing on the at least one piece of to-be-processed voice segment according to the median values of the at least one first subinterval, and determining at least one piece of target voice segment; taking the at least one piece of target voice segment and voice segments other than the at least one piece of to-be-processed voice segment in the M pieces of to-be-processed voice segments as the N pieces of target voice segments, so that the first parameter values of the N pieces of target voice segments conform to the second distribution.
5. The method of claim 4, wherein, The normalization processing on the at least one piece of to-be-processed voice segment according to the median values of the at least one first subinterval to determine at least one piece of target voice segment comprises: performing difference processing on the first parameter value of a first to-be-processed voice segment and the median values of the at least one first subinterval respectively to determine at least one first difference value, and determining a minimum value in absolute values of the at least one first difference value as a second parameter value; calculating a first gain value according to the median value of the subinterval to which the second parameter value belongs and the first parameter value of the first to-be-processed voice segment; determining a first parameter value of a first target speech segment according to the first gain value and a first parameter value of the first to-be-processed speech segment; wherein the first to-be-processed speech segment is one of the at least one to-be-processed speech segment, and the first target speech segment is one of the at least one target speech segment.
6. The method of claim 1, wherein, The method further comprises: dividing the second distribution into intervals to obtain at least one second sub-interval, and determining a median value of the at least one second sub-interval; performing normalization processing on the M to-be-processed speech segments according to the median value of the at least one second sub-interval to determine the N target speech segments, so that the first parameter values of the N target speech segments conform to the second distribution.
7. The method of claim 6, wherein, The method further comprises: determining Q test speech segments, and calculating a first parameter value of each of the Q test speech segments; determining the second distribution according to an expectation, a standard deviation, and a confidence interval of the first parameter values of the Q test speech segments, wherein Q is a positive integer.
8. The method of claim 6, wherein, The second distribution satisfies a preset distribution model.
9. The method according to any one of claims 1 to 8, characterized in that, The method further comprises: dividing the second distribution into intervals to obtain at least one second sub-interval, and determining a median value of the at least one second sub-interval; performing average processing on the median values of the at least one second sub-interval to determine a first average value; performing calibration on the sound pressure levels of the M to-be-processed speech segments according to the first average value to determine the target calibration result.
10. The method of claim 9, wherein, The method further comprises: determining an initial parameter value of the M to-be-processed speech segments; performing calibration on the sound pressure levels of the M to-be-processed speech segments according to the first average value and the initial parameter value to determine the target calibration result.
11. The method of claim 10, wherein, The method further comprises: performing filtering processing on the initial parameter value to determine a third parameter value, and performing difference processing on the third parameter value and the initial parameter value to determine a second difference value; performing filtering processing on the first average value to determine a fourth parameter value, and performing difference processing on the fourth parameter value and the second difference value to determine a third difference value; performing calibration on the sound pressure levels of the M to-be-processed speech segments according to the third difference value to determine the target calibration result.
12. The method according to any one of claims 1 to 8, characterized in that, The method further comprises: if the first parameter values of the M to-be-processed speech segments conform to a first distribution, performing calibration on the sound pressure levels of the M to-be-processed speech segments according to the first distribution to determine the target calibration result.
13. The method according to any one of claims 1 to 8, characterized in that, The method further comprises: determining a probability density distribution of the first parameter values of the M to-be-processed speech segments; According to the probability density distribution of the first parameter value of each of the M to-be-processed voice segments, it is determined whether the first parameter value of each of the M to-be-processed voice segments conforms to the first distribution.
14. A sound pressure level calibration device, characterized by The device comprises a calculation unit, a normalization unit and a determination unit, wherein: The calculation unit is configured to determine M to-be-processed voice segments, and calculate the first parameter value of each of the M to-be-processed voice segments; the first parameter value is the root mean square of amplitude; The normalization unit is configured to, if the first parameter value of at least one to-be-processed voice segment in the M to-be-processed voice segments does not conform to the first distribution, perform normalization processing on the M to-be-processed voice segments, so that the first parameter value of each of the N target voice segments obtained after normalization conforms to a second distribution; the first distribution and the second distribution are normal distributions; the normalization processing comprises deleting the at least one to-be-processed voice segment that does not conform to the first distribution, or adjusting the first parameter value of the at least one to-be-processed voice segment that does not conform to the first distribution to the median value of each of at least one first subinterval of the first distribution, or adjusting the first parameter value of the M to-be-processed voice segments to the median value of each of at least one second subinterval of the second distribution; The determination unit is configured to calibrate the sound pressure level of the M to-be-processed voice segments according to the second distribution, and determine a target calibration result. Wherein, M and N are positive integers, and N is less than or equal to M.
15. An electronic device, comprising: It comprises: A processor and a memory, the memory is used to store a computer program, the processor is used to call and run the computer program stored in the memory, and execute the method as claimed in any one of claims 1 to 13.
16. A chip, characterized by It comprises: A processor is used to call and run a computer program from a memory, so that a device installed with the chip executes the method as claimed in any one of claims 1 to 13.
17. A computer-readable storage medium, characterized in that, The computer storage medium stores a computer program, and the computer program is executed by at least one processor to implement the method as claimed in any one of claims 1 to 13.
Citation Information
Patent Citations
Processing method for local mutation generated by EQ calibration
CN110784792A
Sound pressure level calibration method and device
CN113450846A