Sound evaluation method, device, electronic device, storage medium and program product
Through the sound evaluation method and device, the voice quality evaluation of intelligent voice interaction products is automatically carried out, solving the problems of unified standards and low efficiency, and achieving efficient and unified voice quality detection and adjustment.
Patent Information
- Application Number
- CN202110248660.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-05
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-03-05
AI Technical Summary
Intelligent voice interaction products have unified evaluation standards and low efficiency in terms of voice acquisition, processing and output quality, resulting in unstable service quality.
Provide a sound evaluation method and device, obtain the evaluation target through the sound evaluation command, automatically collect sound and perform clipping detection, sound pickup evaluation and balance evaluation, output evaluation results and perform adjustment operations based on the results.
The sound evaluation operation process has been simplified, the sound evaluation standards have been unified, the self-test capability and test acceptance rate have been improved, the sound evaluation efficiency has been improved, and the investment cost of manpower and material resources has been reduced.
Smart Images

Figure CN115019830B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data evaluation, and particularly to a voice evaluation method, device, electronic device, storage medium, and program product. Background Art
[0002] With the development of Internet and artificial intelligence technologies, intelligent voice interaction products are increasingly widely used. For intelligent voice interaction products, the quality of voice collection, voice processing, and voice output will directly affect the service quality of intelligent voice interaction products. As the number of docking products of intelligent voice interaction products increases, there is an urgent need for a convenient, fast, and unified voice evaluation request scheme to evaluate factors such as the microphone pickup quality, the distortion of the speaker cavity design, and the speaker playback quality, and provide adjustment schemes to simplify the voice evaluation operation process, unify the voice evaluation standards, enhance the self-inspection ability of docking products, improve the test acceptance rate of docking products, improve the voice evaluation efficiency, and reduce the input cost of human and material resources for voice evaluation. Summary of the Invention
[0003] Embodiments of the present disclosure provide a voice evaluation method, device, electronic device, storage medium, and program product.
[0004] In a first aspect, embodiments of the present disclosure provide a voice evaluation method.
[0005] Specifically, the voice evaluation method includes:
[0006] Receiving a voice evaluation command and obtaining a voice evaluation target carried by the voice evaluation command;
[0007] In response to the voice evaluation command, collecting a voice to be measured;
[0008] Evaluating the voice to be measured according to the voice evaluation target to obtain a voice evaluation result;
[0009] Outputting the voice evaluation result.
[0010] Combined with the first aspect, in a first implementation manner of the first aspect of the present disclosure, the obtaining the voice evaluation target carried by the voice evaluation command includes:
[0011] Parsing the voice evaluation command according to a preset information transmission protocol to obtain the voice evaluation target.
[0012] Combined with the first aspect and the first implementation manner of the first aspect, in a second implementation manner of the first aspect of the embodiments of the present disclosure, the evaluating the voice to be measured according to the voice evaluation target to obtain a voice evaluation result includes:
[0013] When the sound evaluation target is clipping detection, perform clipping detection on the sound to be measured;
[0014] When the sound evaluation target is pickup evaluation, perform pickup evaluation on the sound to be measured;
[0015] When the sound evaluation target is equalization evaluation, perform equalization evaluation on the sound to be measured.
[0016] Combining the first aspect, the first implementation manner of the first aspect, and the second implementation manner of the first aspect, in the third implementation manner of the first aspect of the embodiments of the present disclosure, the performing clipping detection on the sound to be measured includes:
[0017] Set the volume detection range;
[0018] Within the volume detection range, traverse each volume to perform clipping detection on the sound to be measured.
[0019] Combining the first aspect, the first implementation manner of the first aspect, the second implementation manner of the first aspect, the third implementation manner of the first aspect, and the fourth implementation manner of the first aspect, in the fourth implementation manner of the first aspect of the embodiments of the present disclosure, the performing clipping detection on the sound to be measured includes:
[0020] Statistical amplitude histogram of the sound to be measured to obtain the maximum amplitude;
[0021] Divide the amplitude interval formed from 0 to the maximum amplitude into N amplitude sub - intervals, calculate the number of amplitude values falling into each amplitude sub - interval, and generate an amplitude value sequence;
[0022] Perform smoothing filtering on the amplitude value sequence to obtain a smoothed amplitude value sequence;
[0023] For a certain amplitude sub - interval in the smoothed amplitude value sequence, determine the local mean within its preset length window as the corresponding amplitude threshold;
[0024] Calculate the amplitude difference between the smoothed amplitude value of this amplitude sub - interval in the smoothed amplitude value sequence and the corresponding amplitude threshold;
[0025] Compare the amplitude difference with zero, and determine the clipping detection result according to the comparison result.
[0026] Combining the first aspect, the first implementation manner of the first aspect, the second implementation manner of the first aspect, the third implementation manner of the first aspect, the fourth implementation manner of the first aspect, and the fifth implementation manner of the first aspect, in the fifth implementation manner of the first aspect of the embodiments of the present disclosure, the performing pickup evaluation on the sound to be measured includes:
[0027] Obtain the sounds to be measured collected by different sound collection devices;
[0028] Calculate the correlation between the to-be-tested sounds, where the correlation is a time-domain correlation or a frequency-domain correlation;
[0029] Determine the sound pickup evaluation result according to the correlation.
[0030] Combined with the first aspect, the first implementation manner of the first aspect, the second implementation manner of the first aspect, the third implementation manner of the first aspect, the fourth implementation manner of the first aspect, and the fifth implementation manner of the first aspect, in the sixth implementation manner of the first aspect in the embodiments of the present disclosure, the equalization evaluation of the to-be-tested sounds includes:
[0031] Perform a fast Fourier transform on the to-be-tested sounds to obtain the frequency-domain data of the to-be-tested sounds;
[0032] Determine the intensity values of different frequency components according to the frequency-domain data of the to-be-tested sounds;
[0033] Calculate a harmonic distortion metric value according to the intensity values of different frequency components;
[0034] Obtain a preset harmonic distortion threshold, and determine the equalization evaluation result according to the comparison result between the harmonic distortion metric value and the preset harmonic distortion threshold.
[0035] Combined with the first aspect, the first implementation manner of the first aspect, the second implementation manner of the first aspect, the third implementation manner of the first aspect, the fourth implementation manner of the first aspect, the fifth implementation manner of the first aspect, and the sixth implementation manner of the first aspect, in the seventh implementation manner of the first aspect in the embodiments of the present disclosure, the output of the sound evaluation result includes:
[0036] Package the sound evaluation result according to a preset information transmission protocol and then output it.
[0037] Combined with the first aspect, the first implementation manner of the first aspect, the second implementation manner of the first aspect, the third implementation manner of the first aspect, the fourth implementation manner of the first aspect, the fifth implementation manner of the first aspect, the sixth implementation manner of the first aspect, and the seventh implementation manner of the first aspect, in the eighth implementation manner of the first aspect in the embodiments of the present disclosure, it further includes:
[0038] Execute a preset adjustment operation according to the sound evaluation result.
[0039] In a second aspect, an embodiment of the present disclosure provides a sound evaluation device.
[0040] Specifically, the sound evaluation device includes:
[0041] An acquisition module, configured to receive a voice evaluation command and acquire a voice evaluation target carried by the voice evaluation command;
[0042] An acquisition module, configured to collect a voice to be measured in response to the voice evaluation command;
[0043] An execution module, configured to evaluate the voice to be measured according to the voice evaluation target to obtain a voice evaluation result;
[0044] An output module, configured to output the voice evaluation result.
[0045] Combined with the second aspect, in the first implementation manner of the second aspect of the present disclosure, the part in the acquisition module that acquires the voice evaluation target carried by the voice evaluation command is configured to:
[0046] Parse the voice evaluation command according to a preset information transmission protocol to obtain the voice evaluation target.
[0047] Combined with the second aspect and the first implementation manner of the second aspect, in the second implementation manner of the second aspect of the embodiments of the present disclosure, the execution module is configured to:
[0048] When the voice evaluation target is clipping detection, perform clipping detection on the voice to be measured;
[0049] When the voice evaluation target is pick-up evaluation, perform pick-up evaluation on the voice to be measured;
[0050] When the voice evaluation target is equalization evaluation, perform equalization evaluation on the voice to be measured.
[0051] Combined with the second aspect, the first implementation manner of the second aspect, and the second implementation manner of the second aspect, in the third implementation manner of the second aspect of the embodiments of the present disclosure, the part that performs clipping detection on the voice to be measured is configured to:
[0052] Set a volume detection range;
[0053] Within the volume detection range, traverse each volume to perform clipping detection on the voice to be measured.
[0054] Combined with the second aspect, the first implementation manner of the second aspect, the second implementation manner of the second aspect, and the third implementation manner of the second aspect, in the fourth implementation manner of the second aspect of the embodiments of the present disclosure, the part that performs clipping detection on the voice to be measured is configured to:
[0055] Statistical amplitude histogram of the voice to be measured to obtain the maximum amplitude value;
[0056] Divide the amplitude range formed from 0 to the maximum amplitude value into N amplitude sub - ranges, calculate the number of amplitude values falling into each amplitude sub - range, and generate an amplitude value sequence;
[0057] Perform a smoothing filtering process on the amplitude value sequence to obtain a smoothed amplitude value sequence;
[0058] For a certain amplitude sub - range in the smoothed amplitude value sequence, determine the local mean within its preset - length window as the corresponding amplitude threshold;
[0059] Calculate the amplitude difference between the smoothed amplitude value of this amplitude sub - range in the smoothed amplitude value sequence and the corresponding amplitude threshold;
[0060] Compare the amplitude difference with zero, and determine the clipping detection result according to the comparison result.
[0061] Combining the second aspect, the first implementation manner of the second aspect, the second implementation manner of the second aspect, the third implementation manner of the second aspect, and the fourth implementation manner of the second aspect, in the fifth implementation manner of the second aspect of the embodiments of the present disclosure, the part for performing a pickup evaluation on the to - be - measured sound is configured to:
[0062] Obtain the to - be - measured sound collected by different sound collection devices;
[0063] Calculate the correlation degree between the to - be - measured sounds, where the correlation degree is a time - domain correlation degree or a frequency - domain correlation degree;
[0064] Determine the pickup evaluation result according to the correlation degree.
[0065] Combining the second aspect, the first implementation manner of the second aspect, the second implementation manner of the second aspect, the third implementation manner of the second aspect, the fourth implementation manner of the second aspect, and the fifth implementation manner of the second aspect, in the sixth implementation manner of the second aspect of the embodiments of the present disclosure, the part for performing an equalization evaluation on the to - be - measured sound is configured to:
[0066] Perform a fast Fourier transform on the to - be - measured sound to obtain the frequency - domain data of the to - be - measured sound;
[0067] Determine the intensity values of different frequency components according to the frequency - domain data of the to - be - measured sound;
[0068] Calculate a harmonic distortion metric value according to the intensity values of different frequency components;
[0069] Obtain a preset harmonic distortion threshold, and determine the equalization evaluation result according to the comparison result between the harmonic distortion metric value and the preset harmonic distortion threshold.
[0070] Combined with the second aspect, the first implementation manner of the second aspect, the second implementation manner of the second aspect, the third implementation manner of the second aspect, the fourth implementation manner of the second aspect, the fifth implementation manner of the second aspect, and the sixth implementation manner of the second aspect, in the seventh implementation manner of the second aspect of the embodiments of the present disclosure, the output module is configured to:
[0071] Encapsulate the voice evaluation result according to a preset information transmission protocol and then output it.
[0072] Combined with the second aspect, the first implementation manner of the second aspect, the second implementation manner of the second aspect, the third implementation manner of the second aspect, the fourth implementation manner of the second aspect, the fifth implementation manner of the second aspect, the sixth implementation manner of the second aspect, and the seventh implementation manner of the second aspect, in the eighth implementation manner of the second aspect of the embodiments of the present disclosure, it further includes:
[0073] An adjustment module, configured to perform a preset adjustment operation according to the voice evaluation result.
[0074] In a ninth aspect, an embodiment of the present disclosure provides an electronic device, including a memory and at least one processor, wherein the memory is used to store one or more computer instructions, and wherein the one or more computer instructions are executed by the at least one processor to implement the method steps of the above-mentioned voice evaluation method.
[0075] In a tenth aspect, an embodiment of the present disclosure provides a computer-readable storage medium for storing computer instructions used by a voice evaluation device, which includes computer instructions for executing the above-mentioned voice evaluation method related to the voice evaluation device.
[0076] In an eleventh aspect, an embodiment of the present disclosure provides a computer program product, including a computer program / instructions, wherein when the computer program / instructions are executed by a processor, the method steps of the above-mentioned voice evaluation method are implemented.
[0077] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0078] The above technical solution obtains a voice evaluation target through the interaction of voice evaluation commands, and automatically realizes multi-factor evaluation of voices based on the voice evaluation target. This technical solution can simplify the voice evaluation operation process, unify the voice evaluation standard, enhance the self-inspection ability of the docking product, improve the test acceptance rate of the docking product, improve the voice evaluation efficiency, and reduce the input cost of manpower and material resources for voice evaluation.
[0079] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings
[0080] In conjunction with the accompanying drawings, other features, objects, and advantages of the present disclosure will become more apparent through the following detailed description of non-limiting embodiments. In the drawings:
[0081] Figure 1 A flowchart showing a voice evaluation method according to an embodiment of the present disclosure;
[0082] Figure 2A A magnitude histogram showing no clipping according to an embodiment of the present disclosure;
[0083] Figure 2B A magnitude histogram showing hard clipping according to an embodiment of the present disclosure;
[0084] Figure 2C A magnitude histogram showing soft clipping according to an embodiment of the present disclosure;
[0085] Figure 3 A schematic diagram showing the application of gain according to an embodiment of the present disclosure;
[0086] Figure 4 An overall flowchart showing a voice evaluation method according to an embodiment of the present disclosure;
[0087] Figure 5 A structural block diagram showing a voice evaluation device according to an embodiment of the present disclosure;
[0088] Figure 6 A structural block diagram showing an electronic device according to an embodiment of the present disclosure;
[0089] Figure 7 It is a schematic structural diagram of a computer system suitable for implementing a voice evaluation method according to an embodiment of the present disclosure. Detailed Embodiments
[0090] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. In addition, for clarity, parts unrelated to the description of the exemplary embodiments are omitted in the drawings.
[0091] In the present disclosure, it should be understood that terms such as "including" or "having" are intended to indicate the presence of features, numbers, steps, actions, components, parts, or combinations thereof disclosed in this specification, and are not intended to exclude the possibility of the presence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0092] It should also be noted that, without conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other. The present disclosure will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0093] The technical solution provided by the embodiments of the present disclosure obtains a voice evaluation target through the interaction of voice evaluation commands, and automatically realizes multi-factor evaluation of voices based on the voice evaluation target. This technical solution can simplify the operation process of voice evaluation, unify the voice evaluation standard, enhance the self-inspection ability of the docking products, improve the test acceptance rate of the docking products, improve the voice evaluation efficiency, and reduce the input cost of human resources and material resources for voice evaluation.
[0094] Figure 1 The flowchart showing a voice evaluation method according to an embodiment of the present disclosure is as Figure 1 shown, and the voice evaluation method includes the following steps S101 - S104:
[0095] In step S101, a voice evaluation command is received, and the voice evaluation target carried by the voice evaluation command is obtained;
[0096] In step S102, in response to the voice evaluation command, the voice to be measured is collected;
[0097] In step S103, the voice to be measured is evaluated according to the voice evaluation target to obtain a voice evaluation result;
[0098] In step S104, the voice evaluation result is output.
[0099] As mentioned above, with the development of Internet and artificial intelligence technologies, intelligent voice interaction products are more and more widely used. For intelligent voice interaction products, the quality of voice collection, voice processing, and voice output will directly affect the service quality of intelligent voice interaction products. With the increase in the number of docking products of intelligent voice interaction products, there is an urgent need for a convenient, fast, and unified voice evaluation request scheme to evaluate factors such as microphone pickup quality, speaker cavity design distortion, and speaker playback quality, and provide adjustment schemes to simplify the voice evaluation operation process, unify the voice evaluation standard, enhance the self-inspection ability of the docking products, improve the test acceptance rate of the docking products, improve the voice evaluation efficiency, and reduce the input cost of human resources and material resources for voice evaluation.
[0100] Considering the above defects, in this embodiment, a voice evaluation method is proposed. This method obtains a voice evaluation target through the interaction of voice evaluation commands, and automatically realizes multi-factor evaluation of voices based on the voice evaluation target. This technical solution can simplify the voice evaluation operation process, unify the voice evaluation standard, enhance the self-inspection ability of the docking products, improve the test acceptance rate of the docking products, improve the voice evaluation efficiency, and reduce the input cost of human resources and material resources for voice evaluation.
[0101] In an embodiment of the present disclosure, the voice evaluation method can be applicable to voice evaluation devices such as computers, computing devices, and electronic devices for evaluating voices.
[0102] In an embodiment of the present disclosure, the voice evaluation command refers to a command sent by a voice evaluation requester, transmitted through a data network, carrying a voice evaluation target, and enabling the voice evaluation device to evaluate a certain factor of the voice according to the voice evaluation target. Among them, the voice evaluation target can be, for example: voice clipping detection, pickup quality evaluation, voice equalization evaluation, etc., to evaluate factors such as the pickup quality of pickups such as microphones, the distortion of the speaker cavity design, and the speaker playback quality. Among them, the voice evaluation command can be sent by the voice evaluation requester through a user interaction page. For example, a user interaction page can be set up, which displays voice evaluation buttons for different evaluation targets. When the user clicks on the voice evaluation button for a certain evaluation target, a voice evaluation command carrying this evaluation target can be sent.
[0103] In an embodiment of the present disclosure, the voice to be measured refers to the voice played for realizing the voice evaluation corresponding to the voice evaluation command. Among them, the voice to be measured can be either the voice played by the voice evaluation requester, the voice played by the voice evaluation requester at the request of the voice evaluation device, or the voice automatically played by the voice evaluation device. When the voice to be measured is the voice played at the request of the voice evaluation device, the voice evaluation device can display voice-to-be-played prompt information through the user interaction page, so that the voice evaluation requester plays the corresponding voice to be measured according to the voice-to-be-played prompt information.
[0104] In an embodiment of the present disclosure, the evaluation refers to the voice evaluation corresponding to the voice evaluation target. For example, if the voice evaluation target is voice clipping detection, the evaluation is clipping detection; if the voice evaluation target is pickup quality evaluation, the evaluation is pickup evaluation; if the voice evaluation target is voice equalization evaluation, the evaluation is equalization evaluation.
[0105] In an embodiment of the present disclosure, the voice evaluation result can include information such as test results, test scores, and adjustment suggestions. The voice evaluation requester can understand the current voice quality situation, the current voice problems, and how to make adjustments according to the voice evaluation result, so as to complete the automatic and rapid evaluation of products such as docking products.
[0106] In the above embodiments, after receiving the voice evaluation command sent by the voice evaluation requester, the voice evaluation device obtains the voice evaluation target carried by the voice evaluation command, and in response to the voice evaluation command, collects the voice to be measured, then performs a voice evaluation corresponding to the voice evaluation target on the voice to be measured according to the voice evaluation target, obtains a voice evaluation result, and finally outputs the voice evaluation result, such as returning it to the voice evaluation requester.
[0107] In an embodiment of the present disclosure, the step of obtaining the voice evaluation target carried by the voice evaluation command in step S101 may include the following steps:
[0108] Parse the voice evaluation command according to a preset information transmission protocol to obtain the voice evaluation target.
[0109] As mentioned above, the voice evaluation command is transmitted through a data network. That is to say, the sending of the voice evaluation command needs to follow a preset information transmission protocol, and operations such as format conversion and encapsulation are performed according to the preset information transmission protocol. Therefore, in this embodiment, after receiving the voice evaluation command, it is necessary to first parse it according to the preset information transmission protocol, such as data extraction or decapsulation operations, to obtain the voice evaluation target carried by the voice evaluation command.
[0110] In an embodiment of the present disclosure, step S102, that is, the step of performing an evaluation on the voice to be measured according to the voice evaluation target to obtain a voice evaluation result, may include the following steps:
[0111] When the voice evaluation target is clipping detection, perform clipping detection on the voice to be measured;
[0112] When the voice evaluation target is pickup evaluation, perform pickup evaluation on the voice to be measured;
[0113] When the voice evaluation target is equalization evaluation, perform equalization evaluation on the voice to be measured.
[0114] As mentioned above, the voice evaluation target may include: voice clipping detection, pickup quality evaluation, voice equalization evaluation, etc. Therefore, in this embodiment, when it is determined that the voice evaluation target is clipping detection, clipping detection is performed on the voice to be measured; when it is determined that the voice evaluation target is pickup evaluation, pickup evaluation is performed on the voice to be measured; when it is determined that the voice evaluation target is equalization evaluation, equalization evaluation is performed on the voice to be measured.
[0115] In an embodiment of the present disclosure, the step of performing clipping detection on the voice to be measured may include the following steps:
[0116] Set the volume detection range;
[0117] Within the described volume detection range, traverse each volume to perform clipping detection on the sound to be measured.
[0118] In order to determine at which volume the clipping of the sound to be measured occurs, this embodiment adopts a progressive volume control and a method of performing clipping detection at each volume. That is, first set the volume detection range, and then within the volume detection range, traverse each volume to perform clipping detection on the sound to be measured.
[0119] In an embodiment of the present disclosure, the step of performing clipping detection on the sound to be measured may include the following steps:
[0120] Statistically analyze the amplitude histogram of the sound to be measured to obtain the maximum amplitude value.
[0121] Divide the amplitude range formed from 0 to the maximum amplitude value into N amplitude sub - ranges, calculate the number of amplitude values falling into each amplitude sub - range, and generate an amplitude value sequence.
[0122] Perform a smoothing filtering process on the amplitude value sequence to obtain a smoothed amplitude value sequence.
[0123] For a certain amplitude sub - range in the smoothed amplitude value sequence, determine the local mean value within its preset length window as the corresponding amplitude threshold.
[0124] Calculate the amplitude difference between the smoothed amplitude value of this amplitude sub - range in the smoothed amplitude value sequence and the corresponding amplitude threshold.
[0125] Compare the amplitude difference with zero, and determine the clipping detection result according to the comparison result.
[0126] In this embodiment, the above - mentioned clipping detection steps are executed for each determined volume. Specifically:
[0127] First, statistically analyze the amplitude histogram of the sound to be measured to obtain the maximum amplitude value Max_Amplitude.
[0128] Then divide the amplitude range formed from 0 to the maximum amplitude value Max_Amplitude into N amplitude sub - ranges, calculate the number of amplitude values falling into each amplitude sub - range, and generate an amplitude value sequence Histogram[n_bins], where n_bins ∈ [0, N] represents the nth amplitude sub - range. If the sound to be measured does not undergo clipping, with n_bins as the abscissa and the amplitude value sequence Histogram[n_bins] as the ordinate, the corresponding amplitude histogram can be as Figure 2A shown. If the sound to be measured undergoes hard clipping, the corresponding amplitude histogram can be as Figure 2BAs shown, if soft clipping occurs in the sound to be measured, the corresponding amplitude histogram can be as Figure 2C shown, where Figure 2A , Figure 2B and Figure 2C the abscissas in are normalized. From the amplitude histogram distribution characteristics shown by Figure 2A , Figure 2B and Figure 2C , it can be known that whether hard clipping or soft clipping occurs, the amplitude histogram of the clipped sound data will have an obvious bulge at the end. Therefore, subsequent clipping detection can be implemented based on this feature without knowing the clipping threshold used during the clipping process.
[0129] Perform a smoothing filter process on the amplitude value sequence Histogram[n_bins] to eliminate the interference caused by small data jumps and obtain a smoothed amplitude value sequence Histogram_Sm[n_bins]. In an embodiment of the present disclosure, the following formula can be used to perform a smoothing filter process on the amplitude value sequence Histogram[n_bins]:
[0130] Histogram_Sm[n_bins] = Sm_Para * Histogram[n_bins] + (1.0 - Sm_Para) * Histogram[n_bins - 1],
[0131] where Histogram_Sm[n_bins] represents the smoothed amplitude value sequence obtained after the smoothing filter process, Sm_Para is the smoothing parameter, Sm_Para ∈ [0, 1]. The smaller the smoothing parameter, the smoother the amplitude value sequence, but the more details are lost. Therefore, in practical applications, the value of the smoothing parameter can be determined according to the actual application needs and the characteristics of the sound to be measured.
[0132] Then, for a certain amplitude sub-interval in the smoothed amplitude value sequence Histogram_Sm[n_bins], determine the local mean within its preset length window as the corresponding amplitude threshold. For example, assuming the preset length is M, for the k-th amplitude sub-interval in the smoothed amplitude value sequence Histogram_Sm[n_bins], its corresponding amplitude threshold can be expressed as:
[0133]
[0134] Calculate the amplitude difference between the smoothed amplitude value of this amplitude sub-interval in the smoothed amplitude value sequence and the corresponding amplitude threshold. For the k-th amplitude sub-interval in the smoothed amplitude value sequence Histogram_Sm[n_bins], the amplitude difference can be expressed as:
[0135] G_Sm_Avg[k] = Histogram_Avg[k] - Histogrem_Sm[k],
[0136] where G_Sm_Avg[k] represents the amplitude difference corresponding to the k-th amplitude sub-interval.
[0137] Combined with the above formula and Figure 2A 、 Figure 2B and Figure 2C it can be known that when there is no clipping, the overall amplitude histogram shows a downward trend, so the amplitude difference G_Sm_Avg[k] should be greater than zero. When clipping occurs, the tail section of the amplitude histogram will turn up, and the amplitude difference G_Sm_Avg[k] will be less than zero. Therefore, if X consecutive amplitude differences are less than zero, it can be determined that clipping occurs in the measured sound at this volume. Among them, the value of X can be determined according to the actual application needs and the characteristics of the measured sound. The present disclosure does not make a special limitation on its specific value. Subsequently, the volume value where clipping occurs and the clipping detection score obtained based on the volume value where clipping occurs are returned to the sound evaluation requester, so that the sound evaluation requester can adjust the relevant products according to the above clipping detection. Among them, the clipping detection score is mainly determined according to different clipping detection evaluation criteria. For example, if the clipping detection evaluation criterion stipulates that there is no clipping when the volume is at the maximum volume of 100, and if it is found through clipping detection that the product has clipping when the volume reaches 90, the clipping detection score for this time can be set to 90 points.
[0138] The above clipping detection method can realize the detection of hard clipping and soft clipping only relying on the amplitude distribution histogram, and without knowing the clipping threshold used during the clipping process, that is, it can detect the scenario with an unknown clipping threshold in the variable gain scenario. Therefore, it has high flexibility and a wide range of applications.
[0139] In an embodiment of the present disclosure, the step of performing a pickup evaluation on the measured sound may include the following steps:
[0140] Obtain the measured sound collected by different sound collection devices;
[0141] Calculate the correlation between the measured sounds, where the correlation is the time-domain correlation or the frequency-domain correlation;
[0142] Determine the pickup evaluation result according to the correlation.
[0143] Among them, the main purpose of the pickup evaluation is to detect whether the consistency between the sound data collected from each path meets the requirements when a product has two or more pickup devices such as microphones. In this embodiment, the correlation between the multi-path sound data is used to judge the consistency between the sound data of each path, and then the pickup evaluation result is obtained.
[0144] Among them, the calculation of the correlation can be performed either in the time domain or in the frequency domain, that is, the correlation can be either the time-domain correlation or the frequency-domain correlation. In practical applications, different correlation calculation methods can be selected according to the computing power of the sound evaluation device. For example, if the computing power of the sound evaluation device is low, the method of calculating the correlation in the time domain can be adopted. Specifically, assuming that there are two paths of sound data x[n] and y[n], the time-domain correlation of these two paths of sound data is calculated using the following formula:
[0145]
[0146] Among them, N represents the length of the sound data, and cross_corr represents the time-domain correlation of the two paths of sound data. The larger the value of the time-domain correlation cross_corr, such as higher than the preset time-domain correlation threshold, the better the consistency of the sound data obtained by the pickup device. On the contrary, the smaller the value of the time-domain correlation cross_corr, such as lower than the preset time-domain correlation threshold, the worse the consistency of the sound data obtained by the pickup device, and there is a problem with the consistency of the pickup of the pickup device. The sound evaluation requester needs to check the pickup device. Among them, the time-domain correlation threshold for judging the size of the time-domain correlation cross_corr value can be set according to the needs of actual applications.
[0147] If the computing power of the sound evaluation device is high, the method of calculating the correlation in the frequency domain can be adopted. Specifically, still assuming that there are two paths of sound data x[n] and y[n], first perform a fast Fourier transform (FFT) on the two paths of sound data to obtain the corresponding frequency-domain data x_freq[n] and y_freq[n], n = [0, N]. The frequency-domain data represents the energy distribution of the sound data in each frequency band. If the pickups of two pickup devices are consistent, the energy distributions of the corresponding frequency bands of the two paths of sound data should also be similar. Based on this, the frequency-domain data can be quantified into a vector composed of N binary numbers. Taking the frequency-domain data x_freq[n] as an example, the frequency-domain data x_freq[n] can be quantified into a vector binary_vector_x[n] composed of N binary numbers using the following formula, n = [0, N]:
[0148]
[0149]
[0150] Similarly, a vector binary_vector_y[n] composed of N binary numbers corresponding to the frequency-domain data y_freq[n] can be obtained. If the sound pickups of two sound pickup devices are consistent, the 1s and 0s distributions of the vectors binary_vector_x[n] and binary_vector_y[n] should be close enough. At this time, the accumulated sum of the differences between the two, that is, the energy distribution difference, can be used as the frequency-domain correlation degree of the two-channel sound data:
[0151]
[0152] Similar to the time-domain correlation degree, the greater the frequency-domain correlation degree, for example, higher than a preset frequency-domain correlation degree threshold, it indicates that the consistency of the sound data obtained by the sound pickup device is better. On the contrary, the smaller the frequency-domain correlation degree value, for example, lower than the preset frequency-domain correlation degree threshold, it indicates that the consistency of the sound data obtained by the sound pickup device is poor, and there is a problem with the consistency of the sound pickup of the sound pickup device. The requester of the sound evaluation needs to check the sound pickup device. Among them, the frequency-domain correlation degree threshold can be set according to the actual application needs.
[0153] In an embodiment of the present disclosure, the step of performing an equalization evaluation on the to-be-tested sound may include the following steps:
[0154] Perform a fast Fourier transform on the to-be-tested sound to obtain the frequency-domain data of the to-be-tested sound;
[0155] Determine the intensity values of different frequency components according to the frequency-domain data of the to-be-tested sound;
[0156] Calculate a harmonic distortion metric value according to the intensity values of different frequency components;
[0157] Obtain a preset harmonic distortion threshold, and determine the equalization evaluation result according to the comparison result between the harmonic distortion metric value and the preset harmonic distortion threshold.
[0158] In this embodiment, the harmonic distortion situation is obtained based on the analysis of frequency components for equalization evaluation. Specifically, the acquired sound to be measured is subjected to a fast Fourier transform to obtain the frequency-domain data of the sound to be measured. Among them, the sound data_in[n] to be measured is the sound obtained by the sound evaluation device collecting the swept-frequency sound or the limited single-frequency sound fteq_x played by the sound evaluation device itself. The sound data_in[n] collected by the sound evaluation device may contain not only the frequency component of freq_x but also frequency components such as harmonic distortion freq_x*2, freq_x*3, freq_x*4, etc. Subsequently, the harmonic distortion situation can be obtained by calculating and analyzing the intensities of the above frequency components; determine the intensity values of different frequency components according to the frequency-domain data of the sound to be measured. For example, assuming that the intensity of the frequency component corresponding to freq_x is H(f), and the intensity of the frequency component corresponding to each harmonic is H(nf), n = [2, m], where m is the number of harmonics for which the harmonic distortion metric value is calculated. For example, when m is set to 5, it is to calculate the harmonic distortion within 5 times; calculate the harmonic distortion metric value according to the intensity values of different frequency components. For example, the harmonic distortion metric value THD_Val can be calculated by the following formula:
[0159]
[0160] Obtain the preset harmonic distortion threshold THD_THRE, and determine the equalization evaluation result according to the comparison result between the harmonic distortion metric value and the preset harmonic distortion threshold. For example, if the harmonic distortion metric value THD_Val of freq_x < the preset harmonic distortion threshold THD_THRE, it indicates that the harmonic distortion at the freq_x frequency point is small and does not need to be suppressed; if the harmonic distortion metric value THD_Val of freq_x > the preset harmonic distortion threshold THD_THRE, it indicates that the harmonic distortion at the freq_x frequency point is large, and a gain gain_x needs to be applied to the corresponding frequency band during output for suppression. Among them, the gain gain_x can be calculated by the following formula:
[0161]
[0162] In practical applications, the product chip's underlying equalization setting interface can be called to apply the gain to the corresponding frequency band, as Figure 3 shown. Assuming that the frequency band to which the gain needs to be applied is band3 and the applied gain is gain_3, the gain graph obtained after applying the gain gain_3 to the frequency band band3 is as shown by the dashed line in Figure 3 In this way, the distorted frequency band can be suppressed to a certain extent during the output process, thereby reducing the interference of harmonics on echo cancellation, and further improving the reception accuracy of sound data.
[0163] In an embodiment of the present disclosure, step S103, that is, the step of outputting the voice evaluation result, may include the following steps:
[0164] Package the voice evaluation result according to a preset information transmission protocol and then output it.
[0165] As mentioned above, the voice evaluation command is transmitted through a data network. Therefore, the voice evaluation result also needs to be transmitted through the data network, that is, the transmission of the voice evaluation result also needs to follow the preset information transmission protocol, and operations such as format conversion and packaging are performed according to the preset information transmission protocol. Therefore, in this embodiment, before outputting the voice evaluation result, it needs to be packaged according to the preset information transmission protocol and then transmitted.
[0166] In an embodiment of the present disclosure, the method may further include the following steps:
[0167] Execute a preset adjustment operation according to the voice evaluation result.
[0168] Wherein, the preset adjustment operation refers to an operation performed based on the voice evaluation result for the purpose of optimizing the voice evaluation result. For example, the preset adjustment operation can be adjusting the speaker gain, adjusting hardware parameters, checking the product material quality, checking the hardware access situation, and so on.
[0169] The above embodiment provides a real-time and platformable voice evaluation solution. The voice evaluation platform built based on the voice evaluation solution enables the voice evaluation requester to not need to build a voice evaluation environment and voice peripherals by themselves, and can realize voice evaluation through the automatic execution of the voice evaluation platform, and can conveniently obtain the voice evaluation result through the voice evaluation platform, thereby reducing the voice evaluation threshold, and can also obtain adjustment suggestions corresponding to the voice evaluation result.
[0170] Figure 4 Show the overall flowchart of the voice evaluation method according to an embodiment of the present disclosure. As Figure 4 shown, the voice evaluation requester sends a voice evaluation command to the voice evaluation device through the user interaction page. The voice evaluation device responds to the received voice evaluation command sent by the voice evaluation requester, obtains the voice evaluation target carried by the voice evaluation command, and collects the voice to be measured; the voice evaluation device performs a voice evaluation corresponding to the voice evaluation target on the voice to be measured according to the voice evaluation target, such as clipping detection, pick-up evaluation, equalization evaluation, etc., to obtain a voice evaluation result; finally, the voice evaluation result is output and sent to the voice evaluation requester.
[0171] The following is an embodiment of the present disclosure device, which can be used to execute the embodiment of the present disclosure method.
[0172] Figure 5 The structural block diagram of a voice evaluation device according to an embodiment of the present disclosure is shown. This device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. As Figure 5 shown, the voice evaluation device includes:
[0173] An acquisition module 501, configured to receive a voice evaluation command and acquire the voice evaluation target carried by the voice evaluation command;
[0174] A collection module 502, configured to collect the voice to be measured in response to the voice evaluation command;
[0175] An execution module 503, configured to evaluate the voice to be measured according to the voice evaluation target to obtain a voice evaluation result;
[0176] An output module 504, configured to output the voice evaluation result.
[0177] As mentioned above, with the development of Internet and artificial intelligence technologies, the applications of intelligent voice interaction products are becoming more and more extensive. For intelligent voice interaction products, the quality of voice collection, voice processing, and voice output will directly affect the service quality of intelligent voice interaction products. With the increase in the number of docking products of intelligent voice interaction products, there is an urgent need for a convenient, fast, and unified voice evaluation request scheme to evaluate factors such as the microphone pickup quality, the distortion of the speaker cavity design, and the speaker playback quality, and provide adjustment schemes to simplify the voice evaluation operation process, unify the voice evaluation standards, enhance the self-inspection ability of docking products, improve the test acceptance rate of docking products, improve the voice evaluation efficiency, and reduce the input cost of human and material resources for voice evaluation.
[0178] Considering the above defects, in this embodiment, a voice evaluation device is proposed. This device obtains the voice evaluation target through the interaction of voice evaluation commands and automatically realizes the multi-factor evaluation of voices based on the voice evaluation target. This technical solution can simplify the voice evaluation operation process, unify the voice evaluation standards, enhance the self-inspection ability of docking products, improve the test acceptance rate of docking products, improve the voice evaluation efficiency, and reduce the input cost of human and material resources for voice evaluation.
[0179] In an embodiment of the present disclosure, the voice evaluation device can be implemented as a voice evaluation device such as a computer, a computing device, an electronic device, etc. for evaluating voices.
[0180] In an embodiment of the present disclosure, the voice evaluation command refers to a command issued by a voice evaluation requester, transmitted through a data network, carrying a voice evaluation target, and enabling the voice evaluation device to evaluate a certain factor of the voice according to the voice evaluation target. Among them, the voice evaluation target can be, for example: voice clipping detection, pickup quality evaluation, voice equalization evaluation, etc., to evaluate factors such as the pickup quality of pickups such as microphones, the distortion of the speaker cavity design, and the playback quality of the speaker. Among them, the voice evaluation command can be issued by the voice evaluation requester through a user interaction page. For example, a user interaction page can be set up, which displays voice evaluation buttons for different evaluation targets. When the user clicks the voice evaluation button for a certain evaluation target, a voice evaluation command carrying the evaluation target can be issued.
[0181] In an embodiment of the present disclosure, the voice to be measured refers to the voice played to achieve the voice evaluation corresponding to the voice evaluation command. Among them, the voice to be measured can be either the voice played by the voice evaluation requester, the voice played by the voice evaluation requester at the request of the voice evaluation device, or the voice automatically played by the voice evaluation device. When the voice to be measured is the voice played at the request of the voice evaluation device, the voice evaluation device can display a voice playback prompt message through the user interaction page, so that the voice evaluation requester plays the corresponding voice to be measured according to the voice playback prompt message.
[0182] In an embodiment of the present disclosure, the evaluation refers to the voice evaluation corresponding to the voice evaluation target. For example, if the voice evaluation target is voice clipping detection, the evaluation is clipping detection; if the voice evaluation target is pickup quality evaluation, the evaluation is pickup evaluation; if the voice evaluation target is voice equalization evaluation, the evaluation is equalization evaluation.
[0183] In an embodiment of the present disclosure, the voice evaluation result may include information such as test results, test scores, and adjustment suggestions. The voice evaluation requester can understand the current voice quality situation, the current voice problems, and how to make adjustments according to the voice evaluation result, so as to complete the automatic and rapid evaluation of products such as docking products.
[0184] In the above embodiment, after receiving the voice evaluation command issued by the voice evaluation requester, the voice evaluation device obtains the voice evaluation target carried by the voice evaluation command, and in response to the voice evaluation command, collects the voice to be measured, then performs a voice evaluation corresponding to the voice evaluation target on the voice to be measured according to the voice evaluation target, obtains a voice evaluation result, and finally outputs the voice evaluation result, for example, returns it to the voice evaluation requester.
[0185] In an embodiment of the present disclosure, the part of the obtaining module 501 that obtains the voice evaluation target carried by the voice evaluation command may be configured to:
[0186] Parse the voice evaluation command according to a preset information transmission protocol to obtain the voice evaluation target.
[0187] As mentioned above, the voice evaluation command is transmitted through a data network. That is to say, the sending of the voice evaluation command needs to follow a preset information transmission protocol, and operations such as format conversion and encapsulation are performed according to the preset information transmission protocol. Therefore, in this embodiment, after receiving the voice evaluation command, it is necessary to first parse it according to the preset information transmission protocol, such as operations like data extraction or decapsulation, to obtain the voice evaluation target carried by the voice evaluation command.
[0188] In an embodiment of the present disclosure, the execution module 503 may be configured to:
[0189] When the voice evaluation target is clipping detection, perform clipping detection on the voice to be measured;
[0190] When the voice evaluation target is pick-up evaluation, perform pick-up evaluation on the voice to be measured;
[0191] When the voice evaluation target is equalization evaluation, perform equalization evaluation on the voice to be measured.
[0192] As mentioned above, the voice evaluation target may include: voice clipping detection, pick-up quality evaluation, voice equalization evaluation, etc. Therefore, in this embodiment, when it is determined that the voice evaluation target is clipping detection, then perform clipping detection on the voice to be measured; when it is determined that the voice evaluation target is pick-up evaluation, then perform pick-up evaluation on the voice to be measured; when it is determined that the voice evaluation target is equalization evaluation, then perform equalization evaluation on the voice to be measured.
[0193] In an embodiment of the present disclosure, the part that performs clipping detection on the voice to be measured may be configured to:
[0194] Set a volume detection range;
[0195] Within the volume detection range, traverse each volume to perform clipping detection on the voice to be measured.
[0196] In order to determine at which volume the voice to be measured undergoes clipping, this embodiment adopts a method of progressive volume control and performs clipping detection at each volume. That is, first set the volume detection range, and then within the volume detection range, traverse each volume to perform clipping detection on the voice to be measured.
[0197] In an embodiment of the present disclosure, the part for performing clipping detection on the to-be-detected sound may be configured to:
[0198] Statistically analyze the amplitude histogram of the to-be-detected sound to obtain the maximum amplitude;
[0199] Divide the amplitude range formed from 0 to the maximum amplitude into N amplitude sub-ranges, calculate the number of amplitude values falling into each amplitude sub-range, and generate an amplitude value sequence;
[0200] Perform a smoothing filtering process on the amplitude value sequence to obtain a smoothed amplitude value sequence;
[0201] For a certain amplitude sub-range in the smoothed amplitude value sequence, determine the local mean within its preset length window as the corresponding amplitude threshold;
[0202] Calculate the amplitude difference between the smoothed amplitude value of this amplitude sub-range in the smoothed amplitude value sequence and the corresponding amplitude threshold;
[0203] Compare the amplitude difference with zero, and determine the clipping detection result according to the comparison result.
[0204] In this embodiment, the above clipping detection process is executed for each determined volume. Specifically:
[0205] First, statistically analyze the amplitude histogram of the to-be-detected sound to obtain the maximum amplitude Max_Amplitude.
[0206] Then divide the amplitude range formed from 0 to the maximum amplitude Max_Amplitude into N amplitude sub-ranges, calculate the number of amplitude values falling into each amplitude sub-range, and generate an amplitude value sequence Histogram[n_bins], where n_bins ∈ [0, N] represents the nth amplitude sub-range. If the to-be-detected sound does not undergo clipping, with n_bins as the abscissa and the amplitude value sequence Histogram[n_bins] as the ordinate, the corresponding amplitude histogram may be as Figure 2A shown. If the to-be-detected sound undergoes hard clipping, the corresponding amplitude histogram may be as Figure 2B shown. If the to-be-detected sound undergoes soft clipping, the corresponding amplitude histogram may be as Figure 2C shown, where Figure 2A 、 Figure 2B and Figure 2C The abscissas in are normalized, and from Figure 2A 、 Figure 2B and Figure 2CFrom the amplitude histogram distribution characteristics shown, whether hard clipping or soft clipping occurs, the amplitude histogram of the clipped audio data will show an obvious bulge at the end. Therefore, subsequent clipping detection can be implemented based on this feature without knowing the clipping threshold used during the clipping process.
[0207] Perform a smoothing filter process on the amplitude value sequence Histogram[n_bins] to eliminate the interference caused by small data jumps, obtaining a smoothed amplitude value sequence Histogram_Sm[n_bins]. In an embodiment of the present disclosure, the following formula can be used to perform the smoothing filter process on the amplitude value sequence Histogram[n_bins]:
[0208] Histogram_Sm[n_bins] = Sm_Para * Histogram[n_bins] + (1.0 - Sm_Para) * Histogram[n_bins - 1],
[0209] where Histogram_Sm[n_bins] represents the smoothed amplitude value sequence obtained after the smoothing filter process, Sm_Para is the smoothing parameter, Sm_Para ∈ [0, 1]. The smaller the smoothing parameter, the smoother the amplitude value sequence, but the more details are lost. Therefore, in practical applications, the value of the smoothing parameter can be determined according to the actual application requirements and the characteristics of the audio to be measured.
[0210] Then, for a certain amplitude sub-interval in the smoothed amplitude value sequence Histogram_Sm[n_bins], determine the local mean within the preset length window as the corresponding amplitude threshold. For example, assuming the preset length is M, for the k-th amplitude sub-interval in the smoothed amplitude value sequence Histogram_Sm[n_bins], the corresponding amplitude threshold can be expressed as:
[0211]
[0212] Calculate the amplitude difference between the smoothed amplitude value of this amplitude sub-interval in the smoothed amplitude value sequence and the corresponding amplitude threshold. For the k-th amplitude sub-interval in the smoothed amplitude value sequence Histogram_Sm[n_bins], the amplitude difference can be expressed as:
[0213] G_Sm_Avg[k] = Histogram_Avg[k] - Histogrem_Sm[k],
[0214] where G_Sm_Avg[k] represents the amplitude difference corresponding to the k-th amplitude sub-interval.
[0215] Combining the above formula and Figure 2A , Figure 2B and Figure 2C It can be seen that when clipping does not occur, the amplitude histogram shows an overall downward trend, and the amplitude difference G_Sm_Avg[k] should be greater than zero. When clipping occurs, the amplitude histogram will have an upward tail section, and the amplitude difference G_Sm_Avg[k] will be less than zero. Therefore, if X consecutive amplitude differences are less than zero, it can be determined that clipping has occurred in the sound to be tested at this volume, wherein the value of X can be determined based on the needs of the actual application and the characteristics of the sound to be tested, and the present disclosure does not specifically limit its specific value. Subsequently, the volume value where clipping occurs and the clipping detection score obtained based on the volume value where clipping occurs are returned to the sound evaluation requester, so that the sound evaluation requester can adjust the relevant products based on the above clipping detection. Among them, the clipping detection score is mainly determined according to different clipping detection evaluation standards. For example, if the clipping detection evaluation standard stipulates that no clipping occurs at the maximum volume of 100, if the clipping detection finds that the product has clipping when the volume reaches 90, then the clipping detection score can be set to 90 points.
[0216] The above clipping detection method can realize the detection of hard clipping and soft clipping by relying only on the amplitude distribution histogram, and does not need to know the clipping threshold used in the clipping processing, that is, it can detect the scene of unknown clipping threshold in the variable gain scene, so it has high flexibility and a wide range of application.
[0217] In one embodiment of the present disclosure, the part of performing sound pickup and evaluation on the sound to be tested may be configured as follows:
[0218] Acquire the sound to be tested collected by different sound collection devices;
[0219] Calculating the correlation between the sounds to be tested, wherein the correlation is a time domain correlation or a frequency domain correlation;
[0220] The sound pickup evaluation result is determined according to the correlation.
[0221] The main purpose of the sound pickup test is to detect whether the consistency between the various sound data collected when a product has two or more sound pickup devices such as microphones meets the requirements. In this embodiment, the correlation between the multiple sound data is used to determine the consistency between the various sound data, and then the sound pickup test result is obtained.
[0222] Among them, the calculation of the correlation can be carried out in the time domain or in the frequency domain, that is, the correlation can be either the time-domain correlation or the frequency-domain correlation. In practical applications, different correlation calculation methods can be selected according to the computing power of the sound evaluation device. For example, if the computing power of the sound evaluation device is low, the method of calculating the correlation in the time domain can be adopted. Specifically, assuming there are two paths of sound data x[n] and y[n], the time-domain correlation of these two paths of sound data is calculated using the following formula:
[0223]
[0224] Among them, N represents the length of the sound data, and cross_corr represents the time-domain correlation of the two paths of sound data. The larger the value of the time-domain correlation cross_corr, such as higher than the preset time-domain correlation threshold, the better the consistency of the sound data obtained by the sound pickup device. On the contrary, the smaller the value of the time-domain correlation cross_corr, such as lower than the preset time-domain correlation threshold, the worse the consistency of the sound data obtained by the sound pickup device, and there is a problem with the consistency of the sound pickup of the sound pickup device. The sound evaluation requester needs to check the sound pickup device. Among them, the time-domain correlation threshold for judging the size of the time-domain correlation cross_corr value can be set according to the needs of actual applications.
[0225] If the computing power of the sound evaluation device is high, the method of calculating the correlation in the frequency domain can be adopted. Specifically, still assuming there are two paths of sound data x[n] and y[n], first perform a fast Fourier transform (FFT) on the two paths of sound data to obtain the corresponding frequency-domain data x_freq[n] and y_freq[n], n = [0, N]. The frequency-domain data represents the energy distribution of the sound data in each frequency band. If the sound pickups of two sound pickup devices are consistent, the energy distributions of the corresponding frequency bands of the two paths of sound data should also be similar. Based on this, the frequency-domain data can be quantized into a vector composed of N binary numbers. Taking the frequency-domain data x_freq[n] as an example, the frequency-domain data x_freq[n] can be quantized into a vector binary_vector_x[n] composed of N binary numbers using the following formula, n = [0, N]:
[0226]
[0227]
[0228] Similarly, a vector binary_vector_y[n] composed of N binary numbers corresponding to the frequency-domain data y_freq[n] can be obtained. If the sound pickups of two sound pickup devices are consistent, the 1s and 0s distributions of the vectors binary_vector_x[n] and binary_vector_y[n] should be close enough. At this time, the cumulative sum of the differences between the two, that is, the energy distribution difference, can be used as the frequency-domain correlation degree of the two-channel sound data:
[0229]
[0230] Similar to the time-domain correlation degree, the larger the frequency-domain correlation degree, such as higher than a preset frequency-domain correlation degree threshold, indicates that the consistency of the sound data obtained by the sound pickup device is better. On the contrary, the smaller the frequency-domain correlation degree value, such as lower than the preset frequency-domain correlation degree threshold, indicates that the consistency of the sound data obtained by the sound pickup device is poor, and there is a problem with the consistency of the sound pickup of the sound pickup device. The requester of the sound evaluation needs to check the sound pickup device. Among them, the frequency-domain correlation degree threshold can be set according to the needs of actual applications.
[0231] In an embodiment of the present disclosure, the part for performing the equalization evaluation on the to-be-tested sound can be configured to:
[0232] Perform a fast Fourier transform on the to-be-tested sound to obtain the frequency-domain data of the to-be-tested sound;
[0233] Determine the intensity values of different frequency components according to the frequency-domain data of the to-be-tested sound;
[0234] Calculate a harmonic distortion metric value according to the intensity values of different frequency components;
[0235] Obtain a preset harmonic distortion threshold, and determine the equalization evaluation result according to the comparison result between the harmonic distortion metric value and the preset harmonic distortion threshold.
[0236] In this embodiment, the harmonic distortion situation is obtained based on the analysis of frequency components for equalization evaluation. Specifically, the collected sound to be measured is subjected to a fast Fourier transform to obtain the frequency domain data of the sound to be measured. Among them, the sound to be measured data_in[n] is the sound obtained by the sound evaluation device collecting the swept-frequency sound or limited single-frequency sound freq_x played by the sound evaluation device itself. The sound to be measured data_in[n] collected by the sound evaluation device may contain not only the frequency component of freq_x, but also frequency components such as harmonic distortion freq_x*2, freq_x*3, freq_x*4, etc. Subsequently, the harmonic distortion situation can be obtained by calculating and analyzing the intensities of the above frequency components; determine the intensity values of different frequency components according to the frequency domain data of the sound to be measured. For example, assuming that the intensity of the frequency component corresponding to freq_x is H(f), and the intensity of the frequency component corresponding to each harmonic is H(nf), n = [2, m], where m is the number of harmonics for which the harmonic distortion metric value is calculated. For example, when m is set to 5, it is to calculate the harmonic distortion within 5 times; calculate the harmonic distortion metric value according to the intensity values of different frequency components. For example, the harmonic distortion metric value THD_Val can be calculated by the following formula:
[0237]
[0238] Obtain the preset harmonic distortion threshold THD_THRE. According to the comparison result between the harmonic distortion metric value and the preset harmonic distortion threshold, determine the equalization evaluation result. For example, if the harmonic distortion metric value THD_Val of freq_x < the preset harmonic distortion threshold THD_THRE, it indicates that the harmonic distortion at the freq_x frequency point is small and does not need to be suppressed; if the harmonic distortion metric value THD_Val of freq_x > the preset harmonic distortion threshold THD_THRE, it indicates that the harmonic distortion at the freq_x frequency point is large, and a gain gain_x needs to be applied to the corresponding frequency band at the output for suppression. Among them, the gain gain_x can be calculated by the following formula:
[0239]
[0240] In practical applications, the product chip's underlying equalization setting interface can be called to apply the gain to the corresponding frequency band, as Figure 3 shown. Assuming that the frequency band for which the gain needs to be applied is band3 and the applied gain is gain_3, the gain graph obtained after applying the gain gain_3 to the frequency band band3 is the dotted line in Figure 3 In this way, the distorted frequency band can be suppressed to a certain extent during the output process, thereby reducing the interference of harmonics on echo cancellation and further improving the reception accuracy of sound data.
[0241] In an embodiment of the present disclosure, the output module 504 may be configured to:
[0242] Encapsulate the voice evaluation result according to a preset information transmission protocol and then output it.
[0243] As mentioned above, the voice evaluation command is transmitted through a data network. Therefore, the voice evaluation result also needs to be transmitted through the data network, that is, the sending of the voice evaluation result also needs to follow the preset information transmission protocol, and operations such as format conversion and encapsulation are performed according to the preset information transmission protocol. Therefore, in this embodiment, before outputting the voice evaluation result, it needs to be encapsulated according to the preset information transmission protocol and then transmitted.
[0244] In an embodiment of the present disclosure, the device may further include:
[0245] An adjustment module, configured to perform a preset adjustment operation according to the voice evaluation result.
[0246] Wherein, the preset adjustment operation refers to an operation performed based on the voice evaluation result for the purpose of optimizing the voice evaluation result. For example, the preset adjustment operation may be to adjust the speaker gain, adjust the hardware parameters, check the product material quality, check the hardware access situation, and so on.
[0247] The above embodiment provides a real-time and platformizable voice evaluation solution. The voice evaluation platform built based on the voice evaluation solution enables the voice evaluation requester to not need to build a voice evaluation environment and voice peripherals by itself, and can realize voice evaluation through the automatic execution of the voice evaluation platform, and can conveniently obtain the voice evaluation result through the voice evaluation platform, thereby reducing the voice evaluation threshold and obtaining adjustment suggestions corresponding to the voice evaluation result.
[0248] The present disclosure also discloses an electronic device, Figure 6 showing a structural block diagram of an electronic device according to an embodiment of the present disclosure, as Figure 6 shown, the electronic device 600 includes a memory 601 and a processor 602; wherein,
[0249] The memory 601 is used to store one or more computer instructions, and wherein, the one or more computer instructions are executed by the processor 602 to implement the above method steps.
[0250] Figure 7 is a schematic structural diagram of a computer system suitable for implementing the voice evaluation method according to an embodiment of the present disclosure.
[0251] As Figure 7As shown, computer system 700 includes a processing unit 701, which can perform various processes in the above embodiments according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage section 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the computer system 700 are also stored. The processing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0252] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read from it can be installed into the storage section 708 as needed. Among them, the processing unit 701 can be implemented as a processing unit such as a CPU, a GPU, a TPU, an FPGA, an NPU, etc.
[0253] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0254] The units or modules involved in the embodiments described in the present disclosure can be implemented in software or in hardware. The described units or modules can also be provided in a processor, and the names of these units or modules do not constitute a limitation to the units or modules themselves in some cases.
[0255] As another aspect, the present disclosure also provides a computer-readable storage medium, which may be the computer-readable storage medium included in the device in the above-described embodiments; or it may exist separately and be a computer-readable storage medium not assembled into the device. The computer-readable storage medium stores one or more programs, and the one or more programs are used by one or more processors to execute the methods described in the present disclosure.
[0256] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
Claims
1. A sound evaluation method, comprising: Receiving a sound evaluation command and obtaining a sound evaluation target carried by the sound evaluation command; Collecting a sound to be measured in response to the sound evaluation command; Evaluating the sound to be measured according to the sound evaluation target to obtain a sound evaluation result; Outputting the sound evaluation result; The evaluating the sound to be measured according to the sound evaluation target to obtain a sound evaluation result includes: Setting a volume detection range; Within the volume detection range, traversing each volume, counting the amplitude histogram of the sound to be measured, and obtaining a maximum amplitude value; Dividing the amplitude interval formed from 0 to the maximum amplitude value into N amplitude sub - intervals, calculating the number of amplitude values falling into each amplitude sub - interval, and generating an amplitude value sequence; Performing a smoothing filter process on the amplitude value sequence to obtain a smoothed amplitude value sequence; For a certain amplitude sub - interval in the smoothed amplitude value sequence, determining the local mean within a preset - length window thereof as the corresponding amplitude threshold; Calculating the amplitude difference between the smoothed amplitude value of this amplitude sub - interval in the smoothed amplitude value sequence and the corresponding amplitude threshold; If a continuous plurality of amplitude differences are less than zero, it is determined that clipping occurs in the sound to be measured at this volume.
2. The method according to claim 1, wherein the obtaining the sound evaluation target carried by the sound evaluation command includes: Parsing the sound evaluation command according to a preset information transmission protocol to obtain the sound evaluation target.
3. The method according to claim 1 or 2, wherein the evaluating the sound to be measured according to the sound evaluation target to obtain a sound evaluation result includes: When the sound evaluation target is a pick - up evaluation, performing a pick - up evaluation on the sound to be measured; When the sound evaluation target is an equalization evaluation, performing an equalization evaluation on the sound to be measured.
4. The method according to claim 3, wherein the performing a pick - up evaluation on the sound to be measured includes: Obtaining sounds to be measured collected by different sound collection devices; Calculating the correlation between the sounds to be measured, wherein the correlation is a time - domain correlation or a frequency - domain correlation; Determining a pick - up evaluation result according to the correlation.
5. The method according to claim 4, wherein the performing an equalization evaluation on the sound to be measured includes: Performing a fast Fourier transform on the sound to be measured to obtain frequency - domain data of the sound to be measured; Determining intensity values of different frequency components according to the frequency - domain data of the sound to be measured; Calculating a harmonic distortion metric value according to the intensity values of different frequency components; Obtaining a preset harmonic distortion threshold, and determining an equalization evaluation result according to the comparison result between the harmonic distortion metric value and the preset harmonic distortion threshold.
6. The method according to any one of claims 1 - 2, 4 - 5, wherein the outputting the sound evaluation result includes: Encapsulating the sound evaluation result according to a preset information transmission protocol and then outputting it.
7. The method according to any one of claims 1 - 2, 4 - 5, further comprising: Performing a preset adjustment operation according to the sound evaluation result.
8. A sound evaluation device, comprising: An acquisition module, configured to receive a voice evaluation command and acquire a voice evaluation target carried by the voice evaluation command; A collection module, configured to collect a voice to be measured in response to the voice evaluation command; An execution module, configured to evaluate the voice to be measured according to the voice evaluation target to obtain a voice evaluation result; The evaluating the voice to be measured according to the voice evaluation target to obtain a voice evaluation result includes: Setting a volume detection range; Within the volume detection range, traversing each volume, counting the amplitude histogram of the voice to be measured, and obtaining a maximum amplitude value; Dividing the amplitude interval formed from 0 to the maximum amplitude value into N amplitude sub-intervals, calculating the number of amplitude values falling into each amplitude sub-interval, and generating an amplitude value sequence; Performing a smoothing filtering process on the amplitude value sequence to obtain a smoothed amplitude value sequence; For a certain amplitude sub-interval in the smoothed amplitude value sequence, determining the local mean within a preset length window thereof as the corresponding amplitude threshold; Calculating the amplitude difference between the smoothed amplitude value of the amplitude sub-interval in the smoothed amplitude value sequence and the corresponding amplitude threshold; If a continuous plurality of amplitude differences are less than zero, it is determined that clipping occurs in the voice to be measured at the volume; An output module, configured to output the voice evaluation result.
9. The device according to claim 8, the part in the acquisition module that acquires the voice evaluation target carried by the voice evaluation command is configured to: Parsing the voice evaluation command according to a preset information transmission protocol to obtain the voice evaluation target.
10. The device according to claim 8 or 9, the execution module is configured to: When the voice evaluation target is a pick-up evaluation, performing a pick-up evaluation on the voice to be measured; When the voice evaluation target is an equalization evaluation, performing an equalization evaluation on the voice to be measured.
11. The device according to claim 10, the part that performs a pick-up evaluation on the voice to be measured is configured to: Acquiring voices to be measured collected by different voice acquisition devices; Calculate the correlation between the sounds to be measured, where, The correlation is a time-domain correlation or a frequency-domain correlation; Determining a pick-up evaluation result according to the correlation.
12. The device according to claim 11, the part that performs an equalization evaluation on the voice to be measured is configured to: Performing a fast Fourier transform on the voice to be measured to obtain voice frequency domain data; Determining the intensity values of different frequency components according to the voice frequency domain data; Calculating a harmonic distortion metric value according to the intensity values of different frequency components; Obtaining a preset harmonic distortion threshold, and determining an equalization evaluation result according to the comparison result between the harmonic distortion metric value and the preset harmonic distortion threshold.
13. The device according to any one of claims 8-9, 11-12, the output module is configured to: Encapsulating the voice evaluation result according to a preset information transmission protocol and then outputting it.
14. The device according to any one of claims 8-9, 11-12, further includes: An adjustment module, configured to perform a preset adjustment operation according to the voice evaluation result.
15. An electronic device, characterized in that, Comprising a memory and at least one processor; wherein the memory is used to store one or more computer instructions, and the one or more computer instructions are executed by the at least one processor to implement the method steps recited in any one of claims 1-7.
16. A computer-readable storage medium having computer instructions stored thereon, characterized in that, When the computer instructions are executed by the processor, the method steps recited in any one of claims 1-7 are implemented.
17. A computer program product, comprising a computer program / instructions, wherein, When the computer program / instructions are executed by the processor, the method steps recited in any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Signal clipping detection method and device, terminal and computer readable storage medium
CN112671376A
Audio clipping detection
US20140226829A1
Method and Apparatus for Testing Speaker, Electronic Device and Storage Medium
US20210058724A1