Sound evaluation system, sound evaluation method, and program

The sound evaluation system integrates black-box and interpretable models to enhance accuracy and interpretability in estimating sound-related indices, addressing the limitations of existing methods by combining first and second estimation units for improved index calculation.

JP7793104B1Active Publication Date: 2025-12-26CYBER AGENT
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025142478
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-12-26
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing sound evaluation methods using black-box machine learning models lack interpretability and accuracy in estimating sound-related indices, with unclear contributions from low-level speech descriptors.

Method used

A sound evaluation system that calculates sound-related indices using a combination of a first estimation unit, a second index acquisition unit, and an integration unit, where the first estimation unit uses a black-box model, the second unit acquires interpretable indices, and the integration unit combines these to enhance accuracy and interpretability.

Benefits of technology

The system achieves high estimation accuracy and explainability of sound-related indices by integrating black-box and interpretable models, allowing for analysis of index contributions and improving the validity of estimates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007793104000001_ABST
    Figure 0007793104000001_ABST
Patent Text Reader

Abstract

To achieve both high accuracy in estimating sound-related index values ​​and explainability of the obtained estimated values. [Solution] The sound evaluation system comprises a first estimation unit that calculates, based on sound data, a first estimated value of a first index, which is one or more indexes related to sound, for a sound indicated by the sound data; a second index acquisition unit that calculates, based on the sound data, a value of a second index, which is one or more indexes related to sound and different from the first index, for a sound indicated by the sound data; a second estimation unit that calculates, based on the value of the second index for the sound indicated by the sound data, a second estimated value of the first index, which is an estimated value for the sound indicated by the sound data; and an integration unit that calculates a third estimated value, which is an estimated value of the first index, based on the first estimated value and the second estimated value.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a sound evaluation system, a sound evaluation method, and a program. [Background technology]

[0002] Sound index values ​​may be estimated using a machine learning model such as a neural network (NN). For example, Non-Patent Document 1 describes a method for estimating a sound aesthetic evaluation score (AES) using a machine learning model. In the method described in Non-Patent Document 1, sound data is input to an encoder using a convolutional neural network (CNN), and the obtained features are further input to a transformer encoder to calculate features. In the method described in Non-Patent Document 1, the features output from the transformer encoder are input to a multi-layer perceptron (MLP), and an estimate is calculated for each indicator of the sound aesthetic evaluation score.

[0003] Furthermore, Non-Patent Document 2 describes a method for calculating a perceptual quality score of speech using a machine learning model. The method described in Non-Patent Document 2 generates features that combine speech embeddings based on a Speech Foundation Model (SFM) with low-level speech descriptors (LLDs) based on speech jitter, shimmer, and harmonic to noise ratio (HNR). In the method described in Non-Patent Document 2, the generated features are input into a three-layer fully connected neural network, and an attention module is applied to the obtained latent representation to calculate a perceptual quality score. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Andros Tjandra and 12 others, "Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound," [online], February 7, 2025, [Retrieved August 4, 2025], Internet<URL:https: / / arxiv.org / aab / 2502.05139> [Non-patent document 2] Whenty Ariyanti and 4 others, "Towards Robust Assessment of Pathological Voices via Combined Low-Level Descriptors and Foundation Model Representations," [online], May 30, 2025, [Retrieved August 4, 2025], Internet<URL:https: / / arxiv.org / abs / 2505.21356> Summary of the Invention [Problem to be solved by the invention]

[0005] In a method for calculating estimated values ​​of sound-related indices based on features calculated using a black-box machine learning model, the calculation of the features is black-boxed, and the interpretability of the resulting estimated values ​​is low. The interpretability of the estimated values ​​here means that information can be provided to evaluate the validity of the estimated values.

[0006] Even in a method in which features calculated using a black-box machine learning model and features combining low-level speech descriptors are input into a black-box machine learning model to calculate estimated values ​​for sound-related indicators, it is unclear to what extent the low-level speech descriptors contributed to the calculation of the estimated values, and the interpretability of the resulting estimates is low.

[0007] On the other hand, one approach to improve the interpretability of the estimated values ​​is to calculate the estimated values ​​of sound-related indices using human-designed interpretable indices. In this method, the correlation between the indices used in the calculation and the calculated indices (estimated values) can be examined, and the indices used in the calculation can be used as information to evaluate the validity of the calculated indices (estimated values).

[0008] However, with this method, if the correlation between the index used in the calculation and the index to be calculated is weak, the accuracy of estimating the sound-related index value may be low. It is preferable to achieve both high accuracy in estimating sound-related index values ​​and good interpretability for the estimated values ​​obtained.

[0009] An example of an object of the present disclosure is to provide a sound evaluation system, a sound evaluation method, and a program that can achieve both high estimation accuracy of sound-related index values ​​and explainability of the obtained estimated values. [Means for solving the problem]

[0010] According to a first aspect of the present disclosure, a sound evaluation system includes a first estimation unit that calculates, based on sound data, a first estimate of a first index, which is one or more indexes related to sound, for a sound indicated by the sound data; a second index acquisition unit that acquires, based on the sound data, a value of a second index, which is one or more indexes related to sound and different from the first index, for a sound indicated by the sound data; a second estimation unit that calculates, based on the value of the second index for the sound indicated by the sound data, a second estimate of the first index, which is an estimate of the sound indicated by the sound data; and an integration unit that calculates a third estimate of the first index, which is an estimate of the first index, based on the first estimate and the second estimate.

[0011] According to a second aspect of the present disclosure, a sound evaluation method includes a computer calculating, based on sound data, a first estimate of a first index, which is one or more indices related to sound, for a sound indicated by the sound data; calculating, based on the sound data, a value of a second index, which is one or more indices related to sound and different from the first index, for a sound indicated by the sound data; calculating, based on the value of the second index for the sound indicated by the sound data, a second estimate of the first index, which is an estimate of the sound indicated by the sound data; and calculating, based on the first estimate and the second estimate, a third estimate of the first index.

[0012] According to a third aspect of the present disclosure, a program causes a computer to execute the following operations: calculate, based on sound data, a first estimate of a first index, which is one or more indices related to sound, for a sound indicated by the sound data; calculate, based on the sound data, a value of a second index, which is one or more indices related to sound and different from the first index, for a sound indicated by the sound data; calculate, based on the value of the second index for the sound indicated by the sound data, a second estimate of the first index, which is an estimate of the sound indicated by the sound data; and calculate, based on the first estimate and the second estimate, a third estimate of the first index. [Effects of the Invention]

[0013] According to an aspect of the present disclosure, it is possible to achieve both high estimation accuracy of sound-related index values ​​and explainability of the obtained estimated values. [Brief explanation of the drawings]

[0014] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a sound evaluation system according to an embodiment. [Figure 2] FIG. 4 is a diagram illustrating an example of input and output of data in a processing unit according to the embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of the configuration of a first estimating unit according to the embodiment. [Figure 4]FIG. 4 is a diagram illustrating an example of the configuration of a second estimating unit according to the embodiment. [Figure 5] FIG. 10 is a diagram illustrating an example of a second index according to the embodiment. [Figure 6] FIG. 10 is a diagram illustrating an example of an evaluation of a model for calculating a first index value according to the embodiment. [Figure 7] FIG. 1 is a diagram illustrating an example of a configuration of a computer according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0015] Hereinafter, embodiments will be described with reference to the drawings.

[0016] Fig. 1 is a diagram showing an example of the configuration of a sound evaluation system according to an embodiment. In the configuration shown in Fig. 1, the sound evaluation system 100 includes a communication unit 110, a display unit 120, an operation input unit 130, a storage unit 140, and a processing unit 150. The processing unit 150 includes a first estimation unit 151, a second index acquisition unit 152, a second estimation unit 153, and an integration unit 154.

[0017] The sound evaluation system 100 estimates an index value related to sound. The index to be estimated is also referred to as a first index. The sound evaluation system 100 may be configured using a computer such as a personal computer (PC) or a workstation (WS). Alternatively, the sound evaluation system 100 may be configured using an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA).

[0018] The first index may be one index or a combination of multiple indexes. Furthermore, the first index is not limited to a specific index, and may be various indexes related to sound. For example, the first index may be an aesthetic evaluation score (AES) of sound based on a combination of the following indexes, but is not limited to this. ·Production Quality (PQ): The quality of the production sound. ·Production Complexity (PC): The complexity of the production sound. ·Content Enjoyment (CE): The enjoyment of the content. ·Content Usefulness (CU): The usefulness of the content.

[0019] Furthermore, the first index is not limited to an index directly related to the sound, but may be an index indirectly related to the sound. For example, the first index may be an index related to a person's operation on the sound to be evaluated or an item using the sound to be evaluated, such as the click-through rate (CTR) or conversion rate (CVR) of an advertisement using the sound to be evaluated.

[0020] The communication unit 110 communicates with other devices. For example, the communication unit 110 may receive sound data of a target for estimating a first index value (sound data indicating a sound of a target for estimating a first index value) from a device that stores the sound data. The sound that is the target of estimation of the first index value is also referred to as the target sound. The sound data that is the target of estimation of the first index value (sound data that indicates the target sound) is also referred to as the target sound data. The target sound is not limited to a specific type of sound. For example, the target sound may be speech, music, or environmental sound. The target sound may also be a generated sound (a sound generated by a generation AI).

[0021] The display unit 120 has a display screen such as a liquid crystal panel or an LED (Light Emitting Diode) panel, and displays various images. For example, the display unit 120 may display various data related to the estimation of the first index by the sound evaluation system 100, such as an estimated value of the first index and a second index value used to estimate the first index value.

[0022] The operation input unit 130 includes input devices such as a keyboard and a mouse, and accepts user operations. For example, the operation input unit 130 may be configured to accept various user operations related to the estimation of the first index by the sound evaluation system 100, such as a user operation to instruct the acquisition of target sound data.

[0023] The storage unit 140 stores various types of data. For example, the storage unit 140 may store various types of data related to the estimation of the first index by the sound evaluation system 100, such as target sound data and various machine learning models used to estimate the first index value. The storage unit 140 is configured using a storage device provided in the sound evaluation system 100 .

[0024] The processing unit 150 performs various processes by controlling each unit of the sound evaluation system 100. The functions of the processing unit 150 are performed, for example, by a CPU (Central Processing Unit) provided in the sound evaluation system 100 reading out a program from the storage unit 140 and executing it.

[0025] The first estimation unit 151 calculates an estimate of a first index for the target sound based on the target sound data. The estimate of the first index for the target sound calculated by the first estimation unit 151 is also referred to as a first estimate. The first estimation unit 151 may be configured as a black box. For example, the first estimation unit 151 may be configured with a black box machine learning model or a combination thereof, and may receive target sound data as input and output a first estimation value.

[0026] Since the first estimation unit 151 may be configured as a black box, various configurations can be adopted as the configuration of the first estimation unit 151. In particular, as the configuration of the first estimation unit 151, it is possible to adopt a configuration that is expected to calculate an estimated value of the first index with high accuracy, such as using a deep learning model.

[0027] The second index acquisition unit 152 acquires a second index value for the target sound. The second index here is one or more indexes related to the sound and is different from the first index. As the second index, an index whose calculation method or estimation method is predefined and whose meaning is interpretable by humans is used. An index whose calculation method or estimation method is predefined and whose meaning is interpretable by humans can be considered as an interpretable index designed by humans.

[0028] In the following, a case will be described as an example in which the second index acquisition unit 152 calculates a second index value for a target sound. However, the method by which the second index acquisition unit 152 acquires the second index value for the target sound is not limited to the method of calculating the value. For example, the second index acquisition unit 152 may acquire the second index value attached to the target sound data as metadata.

[0029] The second estimation unit 153 calculates an estimate of the first index for the target sound based on the second index value for the target sound. The estimate of the first index for the target sound calculated by the second estimation unit 153 is also referred to as a second estimate.

[0030] The second estimate can be referred to as an estimate of the first index for the target sound calculated based on the second index value, while the first estimate can be referred to as an estimate of the first index for the target sound calculated directly based on the target sound data without passing through the second index value.

[0031] The second estimation unit 153 may be configured to analyze the calculation process of the second estimated value. For example, the second estimation unit 153 may be configured using a regression model by XGboost. XGboost is a model that performs boosting using a decision tree, and is analyzable in that the calculation process is represented by a decision tree.

[0032] By being able to analyze the calculation process of the second estimated value, it is possible to analyze the extent to which the second index value (especially the value of each of multiple second indexes) contributes to the second estimated value, and this can be used as material for evaluating the validity of the second estimated value. The second index can be regarded as an explanatory variable in the regression analysis performed by the second estimating unit 153. The first index can be regarded as a response variable in the regression analysis performed by the second estimating unit 153.

[0033] The integration unit 154 calculates an estimate of the first index based on the first estimate (the estimate of the first index calculated by the first estimation unit 151 from the target sound data) and the second estimate (the estimate of the first index calculated by the second estimation unit 153 from the second index value). The estimated value of the first index calculated by the integration unit 154 is also referred to as a third estimated value. The third estimated value is used by the sound evaluation system 100 as the estimation result of the first index for the target sound.

[0034] Since both the first estimated value and the second estimated value are estimates of the first index for the target sound, it is expected that the structure of the second estimation unit can be made relatively simple, and the calculation process of the third estimated value can be analyzed. The integrating unit 154 may be configured using a linear regression model and may calculate the third estimated value as a weighted average of the first estimated value and the second estimated value using weighting coefficients obtained by machine learning (regression analysis), but is not limited to this. For example, the integrating unit 154 may obtain the third estimated value by a weighted sum of the first estimated value and the second estimated value, nonlinear integration, calculation of the median, MAX calculation, or MIN calculation.

[0035] The weighting coefficients used when calculating the weighted average of the first and second estimated values ​​can be considered to indicate the respective contributions of the first and second estimated values ​​to the third estimated value. Therefore, except for special cases such as when the weighting coefficient value for the second estimated value is zero, the second index value can be considered to contribute to the third index value.

[0036] In this way, it is thought that the second estimated value contributes to the third estimated value to a considerable extent, and by analyzing the degree of contribution of the values ​​of each index included in the second index to the second estimated value, it is possible to improve the explainability of which of the sound evaluation elements indicated by each index are used to estimate the first index. Furthermore, the analysis results of the contribution of the values ​​of each index included in the second index to the second estimated value can be used as information suggesting the correlation between the sound evaluation elements indicated by each index and the third index value.

[0037] The sound evaluation elements referred to here can be various elements depending on the individual indices included in the second index. For example, if the second index includes indices indicating sound quality, naturalness, and speech rate, these sound quality, naturalness, and speech rate correspond to the sound evaluation elements, but are not limited to this example.

[0038] FIG. 2 is a diagram showing an example of input and output of data in the processing unit 150. As shown in FIG.

[0039] The first estimation unit 151 receives the target sound data as input and outputs a first estimate (an estimate of a first index for the target sound calculated from the target sound data) to the integration unit 154. In learning the machine learning model that constitutes the first estimation unit 151, a combination of sound data and a first index value (correct answer data for the first estimation value) for the sound indicated by the sound data may be used as training data.

[0040] Learning here refers to adjusting the values ​​of the learning parameters of a machine learning model using training data. Learning parameters are parameters set in a machine learning model that are subject to adjustment during machine learning. Learning of a machine learning model can also be called training of the machine learning model.

[0041] The second index acquisition unit 152 receives the target sound data as input and outputs a second index value for the target sound to the second estimation unit 153. When the calculation method of the second index value is defined by a mathematical formula, the second index acquisition unit 152 may calculate the second index value using a calculation model (that does not involve machine learning). When the second index acquisition unit 152 calculates the second index value using a machine learning model, the learning of the machine learning model may use a combination of sound data and the second index value for the sound represented by the sound data as training data.

[0042] The second estimation unit 153 receives the second index value as input and outputs the second estimated value (the estimated value of the first index for the target sound calculated from the second index value) to the integration unit 154. In learning the machine learning model that constitutes the second estimation unit 153, a combination of sound data and a first index value (correct data for the second estimation value) for the sound indicated by the sound data may be used as training data.

[0043] The integration unit 154 receives the first estimated value and the second estimated value as input and outputs a third estimated value. In training the machine learning model that constitutes the integration unit 154, a combination of the first estimated value and second estimated value for a certain sound and the first index value for that sound (correct data for the third estimated value) may be used as training data.

[0044] Fig. 3 is a diagram illustrating an example of the configuration of the first estimating unit 151. In the configuration illustrated in Fig. 3, the first estimating unit 151 includes a CNN encoder 211, a transformer encoder 212, and a feature-first index converting unit 213.

[0045] The CNN encoder 211 is configured using a convolutional neural network and converts sound data into features. The CNN encoder 211 outputs the calculated features to the transformer encoder 212.

[0046] The transformer encoder 212 is configured using a transformer encoder, which is a type of deep learning model, and converts the input feature into a feature more suitable for calculating the second estimated value. The transformer encoder 212 outputs the calculated feature to the feature-first index conversion unit 213.

[0047] The combination of the CNN encoder 211 and the transformer encoder 212 constitutes the encoder of WavLM, which is a self-supervised speech processing model.

[0048] The feature quantity-first index conversion unit 213 converts the input feature quantity into a first estimated value. The feature quantity-first index conversion unit 213 outputs the calculated first estimated value to the integration unit 154. In the example of FIG. 3, the feature-first index conversion unit 213 is configured using a plurality of KANs (Kolmogorov-Arnold Networks).

[0049] KAN can be considered a type of neural network. While typical neural networks have fixed activation functions, KAN uses a learnable advantage function as the activation function. This allows KAN to perform expressive nonlinear modeling.

[0050] A GR-KAN (Group-Rational KAN) may be used as the KAN constituting the feature-first index conversion unit 213. In this case, the combination of the transformer encoder 212 and the GR-KAN in the example of Fig. 3 can be considered as a configuration in which the multilayer perceptron used in a general transformer is replaced with the GR-KAN.

[0051] A transformer that replaces a multilayer perceptron with a GR-KAN is called a Kolmogorov-Arnold Transformer (KAT), and it has been reported that it improves image recognition performance compared to the Vision Transformer. Furthermore, when replacing the multilayer perceptron in a transformer with GR-KAN to construct a KAT, the weights of GR-KAN are compatible with those of the multilayer perceptron, allowing GR-KAN to be seamlessly integrated into existing transformer architectures without retraining, improving flexibility and modeling capabilities.

[0052] When the first index is made up of a plurality of indexes, a KAN may be provided for each index, and the estimated value of each index may be calculated using a separate KAN. Furthermore, multiple KANs (with the same structure) that have been trained under different conditions may be used to calculate an estimated value of one index. The estimated values ​​calculated by each KAN may then be averaged or voted for by majority vote. For example, the integration unit 154 may calculate a third estimated value by taking a weighted average of the estimated values ​​(first estimated values) calculated by each of the multiple KANs and the estimated value (second estimated value) calculated by the second estimation unit 153.

[0053] The method for realizing different conditions during learning is not limited to a specific method. For example, the value of some hyperparameter, such as a pseudorandom number seed value used to determine the order of data during learning, may be set to a different value for each KAN. Also, the initial value of a learning parameter may be set to a different value for each KAN.

[0054] In this way, by calculating estimates by combining multiple models with different weights even if they have the same structure, it is expected that the prediction errors and biases of each model can be canceled out more than when using only a single model, making it possible to create a configuration that can calculate estimates with higher accuracy.

[0055] Furthermore, by calculating an estimated value by combining multiple models of the same structure with models of different structures, such as in a configuration in which a third estimated value is calculated using a first estimated value calculated by each of multiple KANs and a second estimated value calculated by the second estimation unit 153, it is expected that structural flexibility will be achieved compared to when only multiple models of the same structure are used or when only models of different structures are used, and that a configuration will be possible in which estimated values ​​can be calculated with higher accuracy.

[0056] The learning of the first estimator 151 may be performed by supervised learning or semi-supervised learning with iterative pseudo-labeling added. In this case, a trained model may be prepared as a teacher model for calculating the pseudo-labels. Here, the trained model prepared as the teacher model is not limited to a specific one. For example, a trained model using a multilayer perceptron (rather than a KAN) may be prepared as the teacher model, but is not limited to this.

[0057] The learning of the first estimating unit 151 can be performed, for example, by the following procedure. 1. Use a training model to assign pseudo-labels to unlabeled data. 2. Combine the pseudo-labeled data with the labeled data: add the newly pseudo-labeled data to the training data list that contains the original labeled data. 3. Train the student model (a model using KAN that is the subject of machine learning). 4. Using validation data, the predictive performance of the teacher model is compared with that of the student model. If the student model has better predictive performance than the teacher model, the student model is updated to the teacher model. For example, if a model with a different structure from the student model is prepared as the initial setting for the teacher model, the first update replaces the teacher model with a student model with a different structure. On the other hand, in the second and subsequent updates, the teacher model is replaced with a student model with the same structure but with more optimized parameters. 5. Return to 1.

[0058] 4 is a diagram illustrating an example of the configuration of the second estimating unit 153. In the configuration illustrated in FIG. The second index-first index converter 231 converts the input second index value into a second estimated value. The second index-first index converter 231 outputs the calculated second estimated value to the integrator 154. In the example of FIG. 4, the second index-first index conversion unit 231 is configured using a plurality of XGboost regressors.

[0059] The XGboost regressor is a regression model based on XGboost. By using a regression model to calculate the second estimate based on the second index value, the calculation can be analyzed, and in this respect, the interpretability of the obtained second estimate can be improved. When the first index is composed of a plurality of indexes, an XGboost regressor may be provided for each index, and the estimated value of each index may be calculated by a separate XGboost regressor. Furthermore, multiple XGboost regressors (with the same structure) trained under different conditions may be used to calculate an estimated value of one index. The estimated values ​​calculated by each XGboost regressor may be averaged or voted for by majority vote. For example, the integration unit 154 may calculate a third estimated value by taking a weighted average of the estimated value (first estimated value) calculated by the first estimation unit 151 and the estimated values ​​(second estimated values) calculated by the multiple XGboost regressors in the second estimation unit 153.

[0060] As described above, the first estimation unit 151 may be configured to have each of the multiple KANs calculate a first estimated value. The integration unit 154 may then calculate a third estimated value by taking a weighted average of the estimated values ​​(first estimated values) calculated by each of the multiple KANs in the first estimation unit 151 and the estimated values ​​(second estimated values) calculated by each of the multiple XGboost regressors in the second estimation unit 153.

[0061] As explained for the KAN of the first estimation unit 151, by calculating an estimated value by combining multiple models of the same structure with models of different structures, it is expected that structural flexibility will be achieved compared to when only multiple models of the same structure are used or when only models of different structures are used, and that a configuration will be possible that can calculate estimated values ​​with higher accuracy.

[0062] The secondary indicators may include or include some of the indicators listed in VERSA, which is a collection of indicators for the general evaluation of speech and audio signal quality.

[0063] Fig. 5 is a diagram showing an example of the second index. Fig. 5 shows some of the independent type indexes among the indexes shown in VERSA. The independent type indexes are indexes that do not require a reference signal. The "Name" column shows the name of the indicator. The "Domain" column indicates whether or not Speech, Audio, and Music are each subject to the indicator. The "Variants" column indicates the number of variations of the indicator. The "Range" column shows the range of the index. The "Model Based" column indicates whether a pre-trained model is used. A check mark indicates that a pre-trained model is used. The "Target Direction" column indicates what values ​​are considered good (desirable) for evaluation. An up arrow (↑) indicates that a higher value indicates a better evaluation.

[0064] In the example of Figure 5, 21 of the 24 VERSA independent type indices are used as second indices. However, the three indices excluded in Figure 5 may also be used as second indices. Also, indices of a type other than independent may be used as second indices, or indices other than VERSA may be used as second indices.

[0065] The second indicator may include the following indicators or some of them: Voice quality (Deep Noise Suppression; DNS) indicators DNSMOS P.835 / P.808: MOS (Mean Opinion Score, Mean Weighted Score) score estimation model compliant with ITU-T P.835 / 808 ●Indicators related to overall voice quality NISQA: Outputs a comprehensive score of the naturalness and quality of the voice. ●Indicators related to the naturalness of speech synthesis UTMOS / UTMOSv2: MOS score prediction (specialized for TTS evaluation) ●Indicators related to the impact of packet loss on audio PLCMOS: Quality evaluation after Packet Loss Concealment (PLC) processing Voice quality indicators SCORE (with Ref.): MOS prediction by contrastive learning Sound quality / clarity / signal-to-noise ratio indicators TS-PESQ / TS-STOI / TS-SNR: Deep Neural Network (DNN) models that mimic the output of PESQ (Perceptual Evaluation of Speech Quality), STOI (Short-Time Objective Intelligibility), and SNR (Signal to Noise Ratio). ●Speech enhancement indicators SE-SI-SNR / SE-CI-SDR: SNR comparison with the original clean audio / Comparison of distortion (Convolutive Transfer Function Invariant Signal-to-Distortion; CI-SDR) SE-SAR / SE-SDR: Evaluates artifacts (distortion) / Overall audio distortion level Clarity indicators SRMR: Evaluates clarity according to the degree of reverberation based on the degree of perceptual modulation ●Speech speed indicators SWR: Measures speaking rate (words per minute) Voice spoofing indicators SpoofS: Speaker spoofing detection performance ●Indicators related to the naturalness of the voice SSQA: Automatic score for subjective speaker quality assessment ●Indicators related to the naturalness of singing voice SingMOS: Subjective MOS evaluation of the naturalness of singing voices Indicators of the degree of agreement between sound and text expression PAM metric: Speech is input as a prompt into the language model and evaluated for naturalness and appropriateness. Indicators for evaluating the aesthetic factors of sound Audiobox-Aesthetics: Predicts subjective evaluation scores of sound aesthetic factors (PQ / PC / CE / CU)

[0066] FIG. 6 is a diagram illustrating an example of evaluation of a model for calculating a first index value. To evaluate the configuration of the sound evaluation system 100, -Configuration using multilayer perceptron ("MLP used" in Figure 6), Configuration of a sound evaluation system 100 using four KANs ("Sound Evaluation System" in Figure 6), A single system configuration of the first estimation unit 151 using one KAN ("KAN#1" in FIG. 6), A single system configuration of the first estimation unit 151 using four KANs ("KAN#1-#4" in FIG. 6), A single system consisting of the second index acquisition unit 152 and the second estimation unit 153 ("Use of second index" in Figure 6) For each of the above, the aesthetic evaluation score (AES) of the sound is Production Quality (PQ): Sound quality of the production sound Production Complexity (PC): Complexity of the production sound ·Content Enjoyment (CE): Enjoyment of content ·Content Usefulness (CU): Usefulness of content We estimated the accuracy of the estimated values ​​for each configuration. Figure 6 shows the results of the estimation accuracy evaluation.

[0067] MSE indicates the mean squared error, and the smaller the value, the better the accuracy. SRCC indicates Spearman's Rank Correlation Coefficient, and the larger the value, the better the accuracy.

[0068] "Utterance-level" and "System-level" respectively indicate the method of calculating the error in evaluating the synthesized speech. Utterance-level is a method in which the error is calculated one by one when a certain audio signal is given, and the final average value is taken. System-level is a method in which the average of the subjective evaluation results (true values) and the predicted results is calculated for each synthesis system, the error between the average values ​​is calculated, and the final average value between the systems is calculated.

[0069] For example, consider a case where there are three types of synthesis systems, each with 10 samples. Let the subjective evaluation corresponding to each sample be x and its predicted value be y. When the error is calculated using the mean square error, the utterance level is expressed as in equation (1).

[0070]

number

[0071] The system level is expressed as equation (2).

number

[0072] From the evaluation results shown in FIG. 6, it can be evaluated that the configuration of the sound evaluation system 100 using four KANs has the best estimation accuracy. Furthermore, the sound evaluation system 100 configured using four KANs achieved the best estimation accuracy in the AudioMOS Challenge 2025 (Track 2) (https: / / sites.google.com / view / voicemos-challenge / audiomos-challenge-2025), a competition to estimate aesthetic evaluation scores for sound.

[0073] As described above, the first estimation unit 151 calculates, based on the target sound data, a first estimated value that is an estimated value for the target sound indicated by the target sound data of the first index, which is one or more indexes related to sound. The second index acquisition unit 152 calculates, based on the target sound data, a value of a second index, which is one or more indexes related to sound and is different from the first index, for the target sound indicated by the target sound data. The second estimation unit 153 calculates a second estimated value, which is an estimated value of the first index for the target sound indicated by the target sound data, based on the value of the second index for the target sound indicated by the target sound data. The integration unit 154 calculates a third estimated value, which is an estimated value of the first index, based on the first estimated value and the second estimated value.

[0074] According to the sound evaluation system 100, it is possible to achieve both high estimation accuracy of sound-related index values ​​and explainability of the obtained estimated values. In particular, according to sound evaluation system 100, first estimator 151 may be configured as a black box, and various configurations can be adopted as the configuration of first estimator 151. In particular, as the configuration of first estimator 151, it is possible to adopt a configuration that is expected to calculate an estimated value of the first index with high accuracy, such as using a deep learning model. Furthermore, according to the sound evaluation system 100, it is possible to analyze the degree to which the second index value (especially the value of each of the multiple second indexes) contributes to the second estimated value, and use the analysis results as material for evaluating the validity of the second estimated value. Furthermore, according to sound evaluation system 100, because both the first estimated value and the second estimated value are estimates of the first index for the target sound, it is expected that the calculation process of the third estimated value can be analyzed using a relatively simple structure for the second estimator. By analyzing the calculation process of the third estimated value, it is possible to analyze the contribution of the second estimated value to the third estimated value. The contribution of the second estimated value to the third estimated value can be used as a value indicating the usefulness of the evaluation material for the validity of the second estimated value obtained by analyzing the calculation process of the second estimated value as evaluation material for the validity of the third estimated value, or as an evaluation material for the validity of the third estimated value.

[0075] The first estimation unit 151 is configured to include a neural network that extracts features contained in the sound data and calculates a first estimation value based on the features. The second estimation unit 153 calculates a second estimated value by regression analysis based on the value of the second index.

[0076] According to the sound evaluation system 100, it is possible to achieve both high estimation accuracy of sound-related index values ​​and explainability of the obtained estimated values. In particular, it is expected that the first estimating unit 151 can estimate sound-related index values ​​with high accuracy by using a neural network. Furthermore, the second estimation unit 153 calculates the second estimated value by regression analysis based on the value of the second index, and thus it is possible to analyze the relationship between the value of each index included in the second index and the second estimated value, and in this respect, it is possible to obtain the explainability of the third estimated value that has a correlation with the second estimated value.

[0077] In addition, at least one of the first estimation unit and the second estimation unit calculates an estimated value of the first index based on estimated values ​​of the first index output by multiple machine learning models that have been trained under different conditions. According to sound evaluation system 100, by calculating estimated values ​​by combining multiple models of the same structure with models of different structures, it is expected that structural flexibility will be achieved compared to when only multiple models of the same structure are used, or when only models of different structures are used, and that a configuration will be possible that will allow for more accurate calculation of estimated values.

[0078] 7 is a diagram showing an example of the configuration of a computer according to an embodiment. In the configuration shown in Fig. 7, a computer 700 includes a CPU (Central Processing Unit) 710, a main memory device 720, an auxiliary memory device 730, and an interface 740.

[0079] The above-described sound evaluation system 100, or a part thereof, may be implemented in a computer 700. In this case, the operations of the above-described processing units are stored in the form of a program in an auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program. Furthermore, the CPU 710 allocates a storage area in the main storage device 720 for the sound evaluation system 100 to perform processing in accordance with the program.

[0080] Communication between the sound evaluation system 100 and other devices is performed by the interface 740, which has a communication function, and performs communication under the control of the CPU 710. Interaction between the sound evaluation system 100 and a user is performed by the interface 740, which has an input device and an output device, presenting information to the user via the output device under the control of the CPU 710 and accepting user operations via the input device.

[0081] It is also possible to record a program for realizing all or part of the functions of sound evaluation system 100 on a computer-readable recording medium, and have a computer system load and execute the program to perform processing of each part. Note that the term "computer system" here includes the OS (Operating System) and hardware such as peripheral devices. Furthermore, "computer-readable recording media" refers to portable media such as flexible disks, optical magnetic disks, ROMs (Read Only Memory), and CD-ROMs (Compact Disc Read Only Memory), as well as storage devices such as hard disks built into computer systems. The program may be one that realizes part of the aforementioned functions, or may be one that can realize the aforementioned functions in combination with a program already stored in the computer system.

[0082] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, the above-described embodiments may be combined with other embodiments as appropriate. [Explanation of symbols]

[0083] 100-sound evaluation system 110 Communications Department 120 Display section 130 Operation input section 140 Storage section 150 Processing section 151 First Estimation Department 152 Second index acquisition part 153 Second Estimation Department 154 Integration Department 211 CNN Encoder 212 Transformer Encoder 213 Feature-to-first index conversion unit 231 Second index-first index conversion part

Claims

1. a first estimation unit that calculates, based on sound data, a first estimated value of a first index that is one or more indexes related to sound, the first estimated value being an estimated value of the sound indicated by the sound data; a second index acquisition unit that acquires, based on the sound data, a value of a second index that is one or more indexes related to sound and is different from the first index, for the sound indicated by the sound data; a second estimation unit that calculates a second estimated value of the first index for the sound indicated by the sound data based on the value of the second index for the sound indicated by the sound data; an integration unit that calculates a third estimate that is an estimate of the first index based on the first estimate and the second estimate; A sound evaluation system comprising:

2. the first estimation unit is configured to include a neural network that extracts features included in the sound data and calculates a first estimation value based on the features; the second estimation unit calculates the second estimated value by regression analysis based on the value of the second index. The sound evaluation system according to claim 1 .

3. at least one of the first estimator and the second estimator calculates an estimated value of the first index to be input to the integrator, based on estimated values ​​of the first index output by a plurality of machine learning models that have been trained under different conditions; The sound evaluation system according to claim 1 or 2.

4. The computer calculating, based on the sound data, a first estimate of a first index, which is one or more indexes related to the sound, for the sound indicated by the sound data; calculating, based on the sound data, a value of a second index related to the sound, the second index being different from the first index; calculating a second estimated value of the first index for the sound indicated by the sound data, based on the value of the second index for the sound indicated by the sound data; calculating a third estimate, which is an estimate of the first index, based on the first estimate and the second estimate; A sound evaluation method comprising:

5. On the computer, calculating, based on the sound data, a first estimate of a first indicator, which is one or more indicators related to the sound, the first estimate being an estimate of the sound indicated by the sound data; calculating, based on the sound data, a value of a second index related to the sound, the second index being different from the first index; calculating a second estimated value of the first index for the sound indicated by the sound data, based on the value of the second index for the sound indicated by the sound data; calculating a third estimate that is an estimate of the first index based on the first estimate and the second estimate; A program that executes the following.

Citation Information

Patent Citations

  • Target value estimation system, target value estimation method, and target value estimation program

    WO2017122798A1