Sound quality analysis method, sound quality analysis program, and computer-readable storage medium storing the sound quality analysis program
The sound quality analysis method employs machine learning techniques, including data augmentation and Lasso regression, to objectively quantify the impact of acoustic characteristics on sound quality, addressing the subjectivity of conventional methods and enhancing audio performance.
Patent Information
- Application Number
- JP2021046131
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-03-19
- Publication Date
- 2025-06-09
- Estimated Expiration
- 2041-03-19
AI Technical Summary
Conventional methods for evaluating sound quality are subjective and rely on human expertise, making it difficult to objectively and quantitatively identify the acoustic characteristics that contribute to sound quality judgments.
A sound quality analysis method using machine learning, specifically data augmentation and Lasso regression, to determine a linear classifier that objectively quantifies the influence of various acoustic characteristics on sound quality evaluations.
The method effectively clarifies which acoustic characteristics significantly contribute to sound quality evaluations, enabling objective and quantitative analysis and facilitating improvements in audio performance.
Smart Images

Figure 0007689669000005 
Figure 0007689669000006 
Figure 0007689669000007
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to a sound quality analysis method, a sound quality analysis program, and a computer-readable storage medium storing the sound quality analysis program.
Background Art
[0002] Patent Document 1 discloses an example of supervised learning. Specifically, Patent Document 1 discloses a configuration that automatically generates more teacher data than the given number of teacher data by performing data augmentation on the given number of teacher data.
[0003] Further, Patent Document 2 discloses another example of supervised learning. Specifically, the document classification device disclosed in Patent Document 2 is configured to generate correct data for machine learning by creating new cases from selected correct cases and adding the new correct cases to the correct cases for learning.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] Conventionally, when determining the quality of sound, for example, the power of an expert called a golden ear has been borrowed. However, the evaluation by humans has to be subjective based on hearing, sensibility, etc.
[0006] On the other hand, it is considered that various acoustic characteristics affect the judgment of sound quality. However, in subjective evaluation by humans, it is not easy to grasp which of these acoustic characteristics have what degree of influence.
[0007] The inventors of the present application have come up with the idea that in order to efficiently improve the acoustic performance of audio and the like, it is necessary to clarify the acoustic characteristics with a high degree of contribution and focus on improving those characteristics. As a result of intensive studies, the inventors of the present application have focused on the method of machine learning and newly come up with the idea of using the data augmentation described in Patent Documents 1 and 2 for a different purpose from the conventional one, leading to the conception of the present invention.
[0008] The present disclosure has been made in view of such points, and its object is to objectively and quantitatively clarify the acoustic characteristics that contribute to the evaluation of sound quality by humans and the like among various acoustic characteristics.
Means for Solving the Problems
[0009] The first aspect of the present disclosure relates to a sound quality analysis method for determining a linear classifier for classifying the sound quality of each of a plurality of acoustic data and performing an analysis of acoustic characteristics based on the linear classifier by using a computer including an arithmetic unit that executes a program and a storage unit that reads data.
[0010] According to the first aspect, the sound quality analysis method includes a sample amplification step in which the storage unit reads the plurality of acoustic data in a state where each acoustic characteristic is quantified, and the arithmetic unit performs data augmentation on the plurality of acoustic data to generate a plurality of augmented data to which data augmentation has been applied. The calculation unitA learning preparation step of labeling, for each acoustic data before data augmentation or each augmented data after data augmentation, a determination result of the quality of sound quality in units of acoustic data; and as weighting coefficients constituting the linear classifier, a plurality of influence coefficients corresponding to the respective acoustic characteristics are set, and the arithmetic unit performs lasso regression in a state where a pair of the plurality of augmented data and the sound quality determination result labeled in the learning preparation step is set as teacher data, and a machine learning step of learning the plurality of influence coefficients. , the plurality of acoustic data corresponds to audio sounds recorded in a plurality of types of vehicle cabins, and the acoustic characteristics are determined based on the recording data of the sound collection microphones set in the vehicle cabin and the vehicle interior environment of the vehicle cabin, and at least include one or more of the time centroid, the interaural level difference, the interaural time difference, the initial decay time, the initial lateral energy ratio, the speech transmission index, the C value, the D value, and the interaural correlation function 。
[0011] According to the first aspect, data augmentation is performed on a plurality of acoustic data to generate a plurality of augmented data. According to conventionally known findings, although data augmentation is known as a method suitable when the number of acoustic characteristics to be analyzed, that is, the number of weighting coefficients, is large, excessive data augmentation has been considered undesirable because it causes overfitting.
[0012] However, in the first aspect, data augmentation is performed to intentionally create overfitting or a situation close to it. As a result, while the classification performance for unknown acoustic data is suppressed, for the known acoustic data used as teacher data, the difference between the influence coefficients that contributed to the sound quality determination and other influence coefficients will be significantly enlarged.
[0013] In addition, by intentionally causing overfitting as described above and combining it with machine learning by lasso regression (a regression method in which weighting coefficients with small contribution degrees become zero), it is possible to clearly distinguish between influence coefficients that greatly contribute to the sound quality determination and those with smaller contributions. As a result, it is possible to objectively and quantitatively clarify which acoustic characteristics contribute to the evaluation of the sound quality labeled based on human subjective evaluation or the like.
[0014] Also, according to a second aspect of the present disclosure, in the sample amplification step, the arithmetic unit may standardize the plurality of extended data for each of the influence coefficients based on the average value and the standard deviation of the plurality of extended data.
[0015] Here, the "standardization of a plurality of extended data" refers to a process of converting a plurality of extended data into Z-values (z-scores) as standard scores.
[0016] According to the second aspect, by performing standardization based on the average value and the standard deviation of the extended data generated by data augmentation, even when acoustic data is newly added afterwards, it becomes possible to smoothly perform standardization reflecting the added acoustic data. As a result, it can be used as it is without changing the content of each step in the sound quality analysis method.
[0017] Also, according to a third aspect of the present disclosure, the sound quality analysis method includes a classification test step in which the arithmetic unit collates, for each of the plurality of extended data, the quality of the sound quality labeled in the learning preparation step and the quality of the sound quality classified by the linear classifier by inputting the plurality of extended data into the linear classifier. In the classification test step, the arithmetic unit determines classification error data indicating extended data in which the quality of the sound quality labeled and the quality of the sound quality classified by the linear classifier are different among the plurality of extended data. The arithmetic unit may update the linear classifier by re-executing the machine learning step with the classification error data excluded from the plurality of extended data.
[0018] According to the third aspect, the arithmetic unit re-executes the machine learning step with the classification error data excluded from the extended data. The classification error data corresponds to, for example, extended data that is determined to have good sound quality by a human or the like, but is determined to have poor sound quality by classification using a linear classifier. More specifically, when using a linear classifier that performs linear separation, such classification error data corresponds to extended data that crosses the boundary line characterizing the linear separation.
[0019] According to the third aspect, the arithmetic unit excludes such classification error data from the teacher data. By excluding the data that crosses the boundary line as described above, overfitting is further promoted. This is advantageous in objectively and quantitatively clarifying which acoustic characteristics contribute to the evaluation when evaluating the quality of sound.
[0020] Further, a fourth aspect of the present disclosure relates to a sound quality analysis method for determining a linear classifier for classifying the quality of each of a plurality of acoustic data by using a computer including a calculation unit that executes a program and a storage unit that reads data, and performing an analysis of acoustic characteristics based on the linear classifier
[0021] According to the fourth aspect, the sound quality analysis method includes a sample amplification step in which the storage unit reads the plurality of acoustic data in a state where each acoustic characteristic is quantified, and the calculation unit performs data expansion on the plurality of acoustic data to generate a plurality of expanded data subjected to data expansion; a learning preparation step in which the calculation unit labels the determination result of the sound quality of each acoustic data before data expansion or each expanded data after data expansion in units of acoustic data; a machine learning step in which the calculation unit performs lasso regression in a state where a plurality of influence coefficients corresponding to the respective acoustic characteristics are set as the weighting coefficients constituting the linear classifier, and a pair of the plurality of expanded data and the sound quality labeled in the learning preparation step is set as teacher data, thereby learning the plurality of influence coefficients; and a classification test step in which the calculation unit collates the sound quality labeled in the learning preparation step and the sound quality classified by the linear classifier for each of the plurality of expanded data by inputting the plurality of expanded data into the linear classifier. In the classification test step, the calculation unit determines classification error data indicating the expanded data in which the labeled sound quality and the sound quality classified by the linear classifier are different among the plurality of expanded data, and the calculation unit updates the linear classifier by re-executing the machine learning step with the classification error data excluded from the plurality of expanded data
[0022] Also, according to the 5 aspect of the present disclosure, the arithmetic unit may repeatedly execute the classification test step and the update of the linear classifier until the classification error data is no longer extracted.
[0023] The 5 aspect results in further promotion of overfitting. This is advantageous in objectively and quantitatively clarifying which acoustic characteristics contribute to the evaluation when evaluating the quality of sound.
[0024] Also, according to the 6According to an aspect, after executing the machine learning step, the calculation unit determines a small contribution coefficient indicating an influence coefficient determined to be less than a predetermined reference value through the Lasso regression among the plurality of influence coefficients, and the calculation unit may update the linear classifier by excluding the classification error data from the plurality of extended data and excluding the small contribution coefficient from the plurality of influence coefficients and then executing the machine learning step again.
[0025] By performing Lasso regression, influence coefficients with little contribution to the determination of good or bad become zero. As a result of intensive studies by the inventors of the present application, according to the obtained findings, by performing Lasso regression again with the small contribution coefficients such as the influence coefficients that have become zero excluded, it becomes possible to suppress the appearance of classification error data as described above. Thereby, a linear classifier with higher classification accuracy can be obtained.
[0026] Also 、 The plurality of acoustic data may be composed of audio acoustic data recorded for each vehicle type over a plurality of vehicle types.
[0027] In order to obtain audio acoustic data for each vehicle type, it is necessary to prepare a plurality of different vehicle bodies. Although it is required to record audio acoustics in a plurality of patterns in the vehicle interior to improve the accuracy of machine learning, it is not easy to prepare a plurality of different vehicle bodies.
[0028] On the other hand, as in the present disclosure, by performing data augmentation on such audio acoustic data, it becomes possible to perform high-precision machine learning even when there are few vehicle bodies to be analyzed. Also, in such audio acoustics, by clarifying the acoustic characteristics with a large contribution degree, it becomes possible to focus on creating specific acoustic characteristics when improving the acoustic performance of the audio. Thereby, it becomes possible to efficiently improve the performance of audio products. 。
[0029] Ma However, in the present disclosure, 7The aspect is related to a sound quality analysis program that determines a linear classifier for classifying the quality of each of a plurality of acoustic data by causing a computer including an arithmetic unit that executes a program and a storage unit that reads data to execute, and performs an analysis of acoustic characteristics based on the linear classifier.
[0030] And, according to the 7 aspect, the sound quality analysis program causes the computer to perform a sample amplification step of generating a plurality of augmented data subjected to data augmentation by causing the storage unit to read the plurality of acoustic data in a state where each acoustic characteristic is quantified and causing the arithmetic unit to execute data augmentation on the plurality of acoustic data, The calculation unit a learning preparation step of labeling a determination result of whether the sound quality is good or bad for each acoustic data before data augmentation or each augmented data after data augmentation in units of acoustic data, and a machine learning step of determining the plurality of influence coefficients by causing the arithmetic unit to perform Lasso regression in a state where a plurality of influence coefficients corresponding to the respective acoustic characteristics are set as the weighting coefficients constituting the linear classifier and a combination of the plurality of augmented data and the determination results labeled for each of the plurality of augmented data is set as teacher data. The plurality of acoustic data corresponds to audio sounds recorded in a plurality of types of vehicle cabins, and the acoustic characteristics are determined based on the recording data of the sound collection microphones set in the vehicle cabin and the in-vehicle environment in the vehicle cabin, and at least include one or more of the time centroid, the inter-aural level difference, the inter-aural time difference, the initial decay time, the initial lateral energy ratio, the speech transmission index, the C value, the D value, and the inter-aural correlation function.
[0031] According to the 7 aspect, in the sound quality goodness or badness labeled based on human subjective evaluation or the like, it is possible to objectively and quantitatively clarify which acoustic characteristics contribute to the evaluation.
[0032] Moreover, an eighth aspect of the present disclosure relates to a sound quality analysis program that determines a linear classifier for classifying the quality of each of a plurality of acoustic data by causing a computer including an arithmetic unit that executes a program and a storage unit that reads data to execute, and performs an analysis of acoustic characteristics based on the linear classifier.
[0033] According to the eighth aspect, the sound quality analysis program causes the computer to perform a sample amplification step of causing the storage unit to read the plurality of acoustic data in a state where each acoustic characteristic is quantified, and causing the arithmetic unit to perform data expansion on the plurality of acoustic data to generate a plurality of expanded data subjected to data expansion; a learning preparation step of causing the arithmetic unit to label, for each acoustic data before data expansion or each expanded data after data expansion, a determination result of the quality of sound in units of acoustic data; a machine learning step of setting a plurality of influence coefficients corresponding to the respective acoustic characteristics as weighting coefficients constituting the linear classifier, and causing the arithmetic unit to perform LASSO regression in a state where a combination of the plurality of expanded data and the determination results labeled for each of the plurality of expanded data is set as teacher data to determine the plurality of influence coefficients; and a classification test step of causing the arithmetic unit to collate, for each of the plurality of expanded data, the quality of sound labeled in the learning preparation step and the quality of sound classified by the linear classifier by inputting the plurality of expanded data into the linear classifier. In the classification test step, the arithmetic unit determines classification error data indicating expanded data in which the labeled quality of sound and the quality of sound classified by the linear classifier are different among the plurality of expanded data, and the arithmetic unit updates the linear classifier by re-executing the machine learning step in a state where the classification error data is excluded from the plurality of expanded data.
[0034] Also, a ninth aspect of the present disclosure relates to a computer-readable storage medium storing the sound quality analysis program.
[0035] According to this storage medium, it is possible to objectively and quantitatively clarify which acoustic characteristics contribute to the evaluation of the quality of sound quality labeled based on subjective evaluations by humans and the like.
Advantages of the Invention
[0036] As described above, according to the present disclosure, among various acoustic characteristics, it is possible to objectively and quantitatively clarify the acoustic characteristics that contribute to the evaluation of sound quality by humans and the like.
Brief Description of the Drawings
[0037]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Mode for Carrying Out the Invention
[0038] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that the following description is illustrative.
[0039] <Device Configuration> FIG. 1 is a diagram illustrating the hardware configuration of a sound quality analysis apparatus (specifically, a computer 1 constituting the sound quality analysis apparatus) according to the present disclosure. FIG. 2 is a diagram illustrating the software configuration thereof.
[0040] As illustrated in FIG. 1, the computer 1 includes a central processing unit (CPU) 3 that controls the entire computer 1, a read only memory (ROM) 5 that stores a boot program and the like, a random access memory (RAM) 7 that functions as a main memory, and a hard disk drive (HDD) 9 as a secondary storage device. Note that as the secondary storage device, a solid state drive (SSD) or the like can be used instead of the HDD 9.
[0041] Among these elements, the CPU 3 executes various programs. The CPU 3 functions as an arithmetic unit in the present embodiment. Further, the RAM 7 temporarily stores programs and various data executed by the CPU 3, and the HDD 9 continuously stores programs and data. The RAM 7 can also read data from the HDD 9 as necessary. The RAM 7 functions as a storage unit in the present embodiment.
[0042] The computer 1 also includes a display unit 11 composed of a CRT display, a liquid crystal display, an organic EL display, etc., a graphics memory (Video RAM: VRAM) 13 for storing image data displayed on the display unit 11, a keyboard 15 and a mouse 17 as a man-machine interface. The display unit 11 can display the calculation result by the CPU 3. Also, the computer 1 according to the present embodiment can transmit and receive data to and from an external device via a communication interface 21.
[0043] Note that, as will be described later, the sound quality analysis apparatus according to the present disclosure may be configured by a plurality of computers 1. In that case, the CPU 3 and the RAM 7 of each computer 1 can be regarded as the calculation unit and the storage unit according to the present disclosure, respectively.
[0044] As illustrated in FIG. 2, in the program memory of the HDD 9, an operating system (OS) 19, a sample amplification program 29A, a learning preparation program 29B, a machine learning program 29C, a classification test program 29D, an application program 39, etc. are stored.
[0045] Among these elements, the sample amplification program 29A, the learning preparation program 29B, the machine learning program 29C, and the classification test program 29D constitute the sound quality analysis program 29 in the present embodiment.
[0046] Here, the sound quality analysis program 29 is a program configured to cause the computer 1 to execute the sound quality analysis method according to the present embodiment, and each step constituting the method can be executed by the computer 1. The sound quality analysis program 29 is pre-stored in a computer-readable storage medium 18.
[0047] In the program memory of the HDD 9, the sample amplification program 29A, the learning preparation program 29B, the machine learning program 29C, and the classification test program 29D are each activated in response to commands input from the keyboard 15, the mouse 17, etc. At that time, the sample amplification program 29A etc. are loaded from the HDD 9 to the RAM 7 and are to be executed by the CPU 3. By the CPU 3 executing the sample amplification program 29A etc., the computer 1 functions as a sound quality analyzer.
[0048] On the other hand, a plurality of acoustic data 49 to be analyzed are stored in the data memory of the HDD 9. Each acoustic data 49 is defined as multi-variable data and is generated based on the sound recorded in different environments respectively. Also, each variable constituting each acoustic data 49 is given as a feature amount characterizing the acoustic characteristics of the recorded sound respectively. Each acoustic data 49 is loaded from the HDD 9 to the RAM 7 as necessary and is used for calculations by the CPU 3.
[0049] In addition, various data generated by executing the sample amplification program 29A, the learning preparation program 29B, the machine learning program 29C, and the classification test program 29D, as well as the execution results of the application program 39, are stored in the data memory of the HDD 9 or in the RAM 7 as the main memory as necessary.
[0050] Hereinafter, the specific methodology of the sound quality analysis method will be described in detail.
[0051] <Methodology> FIG. 3 is a flowchart illustrating the procedure of the sound quality analysis method. The method illustrated in FIG. 3 determines a linear classifier for classifying the sound quality of each of the plurality of acoustic data 49 by using the computer 1, and analyzes the acoustic characteristics based on the linear classifier.
[0052] The sound quality analysis method is basically implemented by sequentially executing a sample amplification phase S1 as a sample amplification step, a learning preparation phase S2 as a learning preparation step, a machine learning phase S3 as a machine learning step, and a classification test phase S4 as a classification test step. At this time, steps S3 and S4 will be repeatedly executed until the linear classifier reaches the desired performance, as will be described later.
[0053] Among these phases, the sample amplification phase S1 is implemented by the CPU 3 executing the aforementioned sample amplification program 29A. Similarly, the learning preparation phase S2 is implemented by the CPU 3 executing the learning preparation program 29B, the machine learning phase S3 is implemented by the CPU 3 executing the machine learning program 29C, and the classification test phase S4 is implemented by the CPU 3 executing the classification test program 29D.
[0054] Hereinafter, each phase constituting the sound quality analysis method will be described in order.
[0055] (Sample Amplification Phase S1) FIG. 4 is a flowchart illustrating the procedure of the sample amplification phase S1. Also, FIG. 8 is a diagram showing a specific example of the acoustic data 49 before data augmentation, and FIG. 9 is a diagram showing a specific example of the augmented data 59 after data augmentation. The flowchart illustrated in FIG. 4 shows the processing performed in step S1 of FIG. 3. That is, when the control process proceeds to step S1 in FIG. 3, the CPU 3 executes steps S11 to S16 in FIG. 4 according to the flow.
[0056] First, in step S11 of FIG. 4, the RAM 7 reads a plurality of acoustic data 49 in a state where each acoustic characteristic is quantified and sets this as the analysis target (sample).
[0057] More specifically, assuming that a plurality of acoustic data 49 is composed of n (n is an integer of 2 or more) acoustic data 49, as the plurality of acoustic data 49, acoustic data recorded in n environments can be used. As the acoustic characteristics in this case, P types of acoustic characteristics (P is an integer of 2 or more) can be used.
[0058] More specifically, as the plurality of acoustic data 49, audio acoustic data for each vehicle type (acoustic data corresponding to the audio recorded in the vehicle interior) recorded in a plurality (n types) of vehicle types can be used. As the plurality of types (P types) of acoustic characteristics, various properties characterizing the sound quality of the audio can be used. The acoustic characteristics in the present embodiment include at least one or more of a time centroid, an interaural level difference, an interaural time difference, an initial decay time, an initial lateral energy ratio, a speech transmission index, a C value, a D value, and an interaural correlation function. These acoustic characteristics can be determined using, for example, recording data of a sound collection microphone set in the vehicle interior and various transfer functions calculated based on the recording data and the vehicle interior environment.
[0059] The RAM 7 reads n×P pieces of data and sets this as an analysis target (sample). For example, as shown in FIG. 8, when the number of vehicle types is 2 (n = 2) consisting of vehicle type A and vehicle type B and the number of acoustic characteristics is 96 types (P = 96), the RAM 7 reads a total of 2×96 pieces of data by adding vehicle type A and vehicle type B. Hereinafter, the value of the data quantified for each acoustic characteristic is also referred to as a "feature amount".
[0060] Next, in step S12 of FIG. 4, the CPU 3 performs data augmentation on each acoustic data 49 to generate augmented data 59 subjected to data augmentation. Particularly in this embodiment, the CPU 3 performs data augmentation so that the number of each acoustic data 49 is multiplied equally (for example, so that the magnification is the same for each vehicle type). Assuming that the magnification at that time is M (M is an integer of 2 or more), the plurality of augmented data 59 will be composed of N (= M × n) augmented data 59 excluding acoustic characteristics. In this case, each augmented data 59 will be characterized by P types of feature amounts in the same manner as before the data augmentation. In the example shown in FIG. 9, among the two types of vehicle types, the number of data of vehicle type A and the number of data of vehicle type B will each be amplified by 10,000 times (M = 10,000). Note that the magnification M used during data augmentation is set to be larger than the total number P of feature amounts.
[0061] Note that in this embodiment, data augmentation can be performed by adding M sets of normal random numbers (random numbers generated to follow a normal distribution) to each feature amount in each acoustic data 49. Here, the average value characterizing the normal random number is set to use the value of the feature amount before augmentation. That is, the average value of the p-th feature amount in the augmented data 59 is set to match the value of the p-th feature amount in the acoustic data 49 for each vehicle type.
[0062] Next, in steps S13 to S15 of FIG. 4, the CPU 3 normalizes (Z-values) each augmented data 59 for each acoustic characteristic (influence coefficient) based on the augmented data 59 and the acoustic data 49.
[0063] For example, when normalizing based on the augmented data 59, the CPU 3 normalizes each of the plurality of augmented data 59 for each acoustic characteristic (influence coefficient) based on the average value and standard deviation of the plurality of augmented data 59. In this case, the CPU 3 converts each augmented data 59 so that the average value of the plurality of augmented data 59 becomes 0 for each acoustic characteristic (influence coefficient) and the standard deviation of the plurality of augmented data 59 becomes 1 for each acoustic characteristic (influence coefficient).
[0064] For example, let the p-th feature quantity before normalization in the i-th extended data 59 be X i,p and the feature quantity after normalization be x i,p Also, let the mean value and standard deviation related to the p-th feature quantity calculated based on the extended data 59 (i.e., the mean value and standard deviation based on N (= M × n) data) be μ p , σ p respectively. Then, [Equation 1] Normalization is performed so that the relationship of [Equation 1] holds (where i is an integer from 1 to N, and p is an integer from 1 to P).
[0065] On the other hand, when normalizing based on the acoustic data 49, the CPU 3 normalizes each of the plurality of extended data 59 for each acoustic characteristic (feature quantity) based on the mean value and standard deviation of the plurality of extended data 59.
[0066] For example, let the p-th feature quantity before normalization in the i-th extended data 59 be X i,p and the feature quantity after normalization be x i,p Also, let the mean value and standard deviation related to the p-th feature quantity calculated based on the acoustic data 49 (i.e., the mean value and standard deviation based on n feature quantities) be μ p ’ and σ p ’ respectively. Then, [Equation 2] It follows that the relationship of [Equation 2] holds (where i is an integer from 1 to N, and p is an integer from 1 to P).
[0067] Specifically, in step S13, the CPU 3 selects whether to perform normalization using the extended data 59. If the former is selected (step S13: YES), the process proceeds to step S14 to execute the normalization based on the above formula (1). On the other hand, if the latter is selected in step S13 (step S13: NO), the process proceeds to step S15 to execute the normalization based on the above formula (2). Note that step S13 is not essential. For example, instead of using steps S13 to S15, only one of step S14 or step S15 may be implemented.
[0068] Each feature amount x in the normalized extended data 59 i,p is used as the "p-th component of the input vector x i in machine learning". In other words, the sound quality analysis method according to the present embodiment is configured to use the feature amount x i,p as an input vector in machine learning.
[0069] The plurality of extended data 59 normalized by the CPU 3 are temporarily or continuously stored in the RAM 7, HDD 9, etc. The data thus stored is read as needed in the learning preparation phase S2, etc. When step S14 or step S15 is completed, the control process returns from the flow illustrated in FIG. 4 and proceeds to step S2 in FIG. 3.
[0070] Also, for convenience of explanation, the term "a plurality of normalized extended data" is simply referred to as "a plurality of extended data".
[0071] (Learning Preparation Phase S2) FIG. 5 is a flowchart illustrating the procedure of the learning preparation phase S2. The flowchart illustrated in FIG. 5 shows the processing performed in step S2 in FIG. 3. That is, when the control process proceeds to step S2 in FIG. 3, the CPU 3 executes steps S21 to S22 in FIG. 5 according to the flow.
[0072] First, in step S21 of FIG. 5, the RAM 7 reads a plurality of extended data 59 generated in the sample amplification phase S1 from the HDD 9 or the like.
[0073] Next, in step S22, in the learning preparation phase S2, for each extended data 59 after data expansion, more specifically, for each feature amount x given for each extended data 59 i,p the determination result of the sound quality, whether good or bad, is labeled in units of the acoustic data 49.
[0074] This labeling is performed independently of the setting of the feature amounts corresponding to the respective acoustic characteristics. The "determination of the sound quality, whether good or bad" here includes the determination by a human, such as a specialist called the so-called "golden ear".
[0075] As described above, the determination of the sound quality, whether good or bad, is labeled in units of the acoustic data 49. That is, in the present embodiment, as illustrated in FIGS. 8 and 9, the determination result (Good / Bad) labeled for the acoustic data 49 before data expansion is also retained in each extended data 59 after data expansion.
[0076] For example, as shown in FIG. 9, the extended data 59 (the vehicle type A in FIG. 9 1 ,...A m , related extended data 59) generated based on the acoustic data 49 related to the vehicle type A is labeled with the same determination result as that of the vehicle type A (in the example in the figure, "Good"). Similarly, the extended data 59 (the vehicle type B in FIG. 9 1 ,...B m , related extended data 59) generated based on the acoustic data 49 related to the vehicle type B is labeled with the same determination result as that of the vehicle type B (in the example in the figure, "Bad").
[0077] The labeling according to the present embodiment is configured to be performed by associating each extended data 59, more specifically, each feature amount x i,p with the discrete value Y i . The discrete value Y i is an index indicating the sound quality, whether good or bad. For example, "Yi When “=1”, it is defined as “Good (sound quality: good)”, and “Y” i When “=0”, it can be defined as “Bad (sound quality: bad)”, or it can be defined to have the opposite correspondence relationship.
[0078] For each feature quantity x i,p The goodness or badness of the sound quality labeled for it is used as a “label” in machine learning. That is, the sound quality analysis method according to this embodiment is configured to execute supervised learning as machine learning. In this case, as the training data, the above-described extended data 59, more specifically, the P-dimensional input vector x i and the discrete value Y as a label i and the pair (x i,p , Y i ) is configured to be used.
[0079] The training data (x i,p , Y i ) generated by the CPU 3 is temporarily or continuously stored in the RAM 7, HDD 9, etc. The data thus stored is read as needed in the machine learning phase S3, etc. When step S22 is completed, the control process returns from the flow illustrated in FIG. 5 and proceeds to step S3 in FIG. 3.
[0080] (Machine learning phase S3) FIG. 6 is a flowchart illustrating the procedure of the machine learning phase S3. The flowchart illustrated in FIG. 6 shows the processing performed in step S3 in FIG. 3. That is, when the control process proceeds to step S3 in FIG. 3, the CPU 3 executes steps S31 to S33 in FIG. 6 according to the flow.
[0081] First, in step S31 of FIG. 6, the RAM 7 reads various settings related to machine learning and completes the preliminary preparation for machine learning. The settings read in this step S31 include, for example, the initial values of the respective influence coefficients β that make up the function f indicating the linear classifier p , the training data (x i,p , Yi ) The magnitude of the number of divisions k in the division tolerance verification, the upper limit value, lower limit value, and step width of the regularization parameter in the Lasso regression, etc. are included.
[0082] Among the above settings, for example, the function f is modeled to represent a linear classifier. That is, the function f is preset to take as an argument a linear combination of the weighting coefficients and the feature quantity x i,p and. Here, as the weighting coefficients constituting the linear classifier, a plurality of influence coefficients β p corresponding to the acoustic characteristics are set respectively. Each influence coefficient β p represents the contribution degree of the acoustic characteristics corresponding to the influence coefficient β p to the determination of the quality of sound quality for the feature quantity x i,p corresponding thereto, and thus for the feature quantity x i,p corresponding thereto. That is, when the value of the influence coefficient β p is large, it can be considered that the corresponding acoustic characteristic has a relatively strong influence on the determination of the quality of sound quality, and when the influence coefficient β p is small, it can be considered that the corresponding acoustic characteristic has a relatively weak influence on the determination of the quality of sound quality.
[0083] As a specific function form of the function f, a linear function (so-called classification by linear separation) showing the boundary line that divides the quality of sound quality can be used. Instead of this, as the function form of the function f, a probability distribution function for probabilistically determining the quality of sound quality can also be used.
[0084] In the former case,
Equation
[0085] In Equation (3), the output result y of the function f i and the label Y that is preset to constitute the teacher data i correspond to "correct (successfully reproduced the golden year)" when their values match, and to "incorrect (failed to reproduce the golden year)" when they do not match. By using the teacher data (x i,p , Y i ) given for i = 1 to N to learn the value of the influence coefficient β p , a linear classifier that can reproduce the determination by the golden year can be determined.
[0086] Specifically, in step S32 following step S31, the CPU 3 performs Lasso regression with the aforementioned pair (x i,p , Y i ) as the teacher data to learn a plurality of influence coefficients β p . That is, the CPU 3 executes machine learning so as to minimize the error function S λ under L1 regularization using the regularization parameter λ.
[0087] In this case, the functional form of the error function S λ can be described as in the following Equation (4).
[0088]
Equation
[0089] Equation (4) is a specific example when using linear separation. When performing logistic regression, the term corresponding to the linear function is replaced with a sigmoid function or the like. In that case, machine learning through Bayesian estimation is performed.
[0090] More specifically, when using linear classification, cross-validation can be performed to optimize the value of the regularization parameter λ. At that time, it is required to divide the teacher data (x i,p , Y i ) with respect to the variable i for classifying the extended data 59. Here, in the present embodiment, the teacher data (xi,p , Y i ) is divided on the basis of acoustic data units, that is, for each vehicle type. In other words, the division of the teacher data (x i,p , Y i ) is performed so that the number of acoustic data 49 and thus the number for each vehicle type are the same.
[0091] For example, in the example shown in FIG. 9, when the number of divisions k is set to 10, all M pieces of augmented data 59 with vehicle type A as the augmentation source and all M pieces of augmented data 59 with vehicle type B as the augmentation source are each divided into 10 parts.
[0092] In the state divided as described above, the CPU 3 sets any one block of the divided teacher data (x i,p , Y i ) as test data for verification, and uses the other blocks of the divided teacher data (x i,p , Y i ) as teacher data as they are. For example, when the teacher data (x i,p , Y i ) is divided into 10 blocks, one of the blocks is set as test data, and the remaining 9 blocks are used as teacher data. In each vehicle type, the teacher data and the test data are set in the same way.
[0093] Then, the CPU 3 learns the function f using 9 blocks as teacher data. Specifically, the CPU 3 calculates the influence coefficient β p that minimizes the error function represented by the above formula (4). This calculation can be performed using, for example, the coordinate descent method.
[0094] When learning the function f, λ is fixed to a predetermined value within a previously set range. After learning the function f, the CPU 3 calculates the value of the error described in formula (4) by inputting one block as test data into the learned function f.
[0095] After that, the CPU 3 changes the combination of the nine data blocks used as teacher data and the one data block used as test data, and calculates the error represented by the above formula (4) again. In this way, each of the ten divided data blocks is set as test data once, and the arithmetic mean of the errors obtained at each setting (more generally, the error called "Classification error") is calculated.
[0096] After that, the CPU 3 increases or decreases the value of λ within a pre-set range and repeatedly calculates the arithmetic mean of the errors. Then, while changing the value of λ, the arithmetic mean of the errors is calculated while switching the combination of teacher data and test data, thereby determining the optimal value of λ that minimizes the arithmetic mean of the errors.
[0097] Then, in step S33 following step S32, the CPU 3 uses λ determined to be the optimal value and all the teacher data (x i,p , Y i ) before division, and recalculates a plurality of influence coefficients β p to tentatively determine the functional form of the linear classifier (function f).
[0098] The linear classifier tentatively determined by the CPU 3 and the extended data 59 used as teacher data in the determination (in the case of the first execution of the machine learning phase S3, all N extended data 59 are applicable) are temporarily or continuously stored in the RAM 7, HDD 9, etc. Then, the stored data is read as needed in the classification test phase S4. When step S33 is completed, the control process returns from the flow illustrated in FIG. 6 and proceeds to step S4 in FIG. 3.
[0099] (Classification Test Phase S4) FIG. 7 is a flowchart illustrating the procedure of the classification test phase S4. Further, FIG. 10 is a diagram for explaining thinning out such as weight coefficients in the classification test phase S4. The flowchart illustrated in FIG. 7 shows the processing performed in step S4 of FIG. 3. That is, when the control process proceeds to step S4 in FIG. 3, the CPU 3 executes steps S41 to S45 in FIG. 7 according to the flow.
[0100] First, in step S41 of FIG. 7, the RAM 7 reads the functional form of the classifier tentatively determined in the machine learning phase S3 and the extended data 59 at the current time used for the determination of the classifier (when the first execution of step S4, all N pieces of extended data 59 are applicable).
[0101] Next, by inputting the plurality of extended data 59 read in step S41 into the classifier (linear classifier) (step S42), for each of the plurality of extended data 59, the CPU 3 collates the quality of sound labeled in the learning preparation phase S2 and the quality of sound classified by the classifier (step S43).
[0102] That is, in this step S43, the CPU 3 i compares the output result y i of the tentatively determined function f with the value of the label Y associated with the i-th extended data 59, and determines whether or not they match.
[0103] As described above, when i the output result y i of the function f matches the value of the label Y preset as the teacher data, it corresponds to "success in reproducing the golden ear", and when they do not match, it corresponds to "failure in reproducing the golden ear".
[0104] Then, in step S44 following step S43, the CPU 3 performs the classification result (the output result y i ) for all the extended data 59 at the current time and the prior determination result (label Y i) Determine whether they match. If they match (step S44: YES), proceed to step S45. On the other hand, if at least a part does not match (step S44: NO), proceed to step S46.
[0105] When proceeding to step S45, the CPU 3 extracts acoustic characteristics contributing to the quality of sound based on the values of a plurality of influence coefficients β p determined in the machine learning phase S3.
[0106] Specifically, in this step S45, the CPU 3 sorts a plurality of influence coefficients β p in descending order and performs ranking. The ranking thus determined is visualized on the display unit 11 or stored in the HDD 9 or the like. For example, through the ranking displayed on the display unit 11, developers and the like can visually recognize the acoustic characteristics affecting the determination by the golden ear.
[0107] On the other hand, when proceeding from step S44 to step S46, the CPU 3 determines classification error data 59' indicating the extended data 59 in which the determination result of good or bad set as a label (label Y i ) and the classification result by the classifier (output result y of the function f i ) are different.
[0108] Then, in step S47 following step S46, the CPU 3 excludes the data regarded as the classification error data 59' from all N pieces of extended data 59. For example, in the example shown in FIG. 10, among all 2×M pieces of extended data 59, the M-th extended data 59 related to vehicle type A corresponds to the classification error data 59'. By excluding this, all N - 1 pieces of extended data 59 are stored in the RAM 7 or the like. Not limited to the example of FIG. 10, the CPU 3 excludes all data determined as the classification error data 59' from the extended data 59.
[0109] In step S48 following step S47, the CPU 3 a plurality of influence coefficients β pAmong them, the influence coefficient β determined to be less than a predetermined reference value through Lasso regression p indicating a small contribution coefficient β p ’ is determined.
[0110] By performing Lasso regression, the influence coefficient β with a small contribution to the determination of sound quality p becomes substantially zero. The CPU 3 compares the magnitude of the value of the influence coefficient β p with, for example, a reference value set to zero, and extracts the small contribution coefficient β p ’ that is presumed to have a small influence on the determination by the golden ear. Then, in step S49 following step S48, the CPU 3 excludes the term related to the influence coefficient β p from the function f represented by the above formula (3) and the like. For example, in the example shown in FIG. 10, among all P influence coefficients β p , the influence coefficient β 2 corresponding to the second acoustic characteristic is equivalent to the small contribution coefficient β p ’.
[0111] After performing step S49, the control process returns from the classification test phase S4 to the machine learning phase S3, and re-executes the machine learning phase S3 in a state reflecting the processing related to the classification test phase S4.
[0112] The machine learning phase S3 re-executed as described above is performed in a state where the processing related to step S46 is reflected. That is, the CPU 3 updates the classifier (specifically, the functional form of the function f) by re-executing the machine learning phase S3 in a state where the classification error data 59’ is excluded from the plurality of extended data 59. Then, the control process transitions from the machine learning phase S3 to the classification test phase S4, and the determination related to step S44 described above is re-executed. If the determination related to step S44 is No, the classification error data 59’ will be further excluded.
[0113] In this way, the CPU 3 repeatedly executes the classification test step and the update of the classifier until the classification error data 59’ is no longer extracted from the plurality of extended data 59.
[0114] Also, the machine learning phase S3 that is executed again as described above will be performed in a state reflecting the processing according to step S49. That is, the CPU 3 excludes the classification error data 59' from the plurality of extended data 59 and the plurality of influence coefficients β p and excludes the small contribution coefficient β p ' from the plurality of influence coefficients β, and then updates the linear classifier by executing the machine learning phase S3 again. As a result, the function f used in the Lasso regression will gradually change into a lower-dimensional function.
[0115] <Example of the sound quality analysis method> FIG. 11 is a diagram for explaining the influence of weight coefficients and the like in the classification test phase S4, and FIG. 12 is a diagram for explaining the ranking of influence coefficients. In this embodiment, audio acoustic data recorded in two types of vehicle cabins was used as the analysis target. That is, the number of the acoustic data 49 is two (n = 2). Also, at the time of recording, the sound collection microphones were arranged on the driver's seat. Further, the sound collection microphones were arranged side by side along the left-right direction (vehicle width direction) so as to correspond to both ears of the driver.
[0116] Also, the number of acoustic characteristics analyzed for each vehicle type, that is, the number of feature amounts and influence coefficients, is 96 (P = 96) at the first execution of the machine learning phase S3.
[0117] Also, in this embodiment, the data is expanded 10,000 times for each of vehicle type A and vehicle type B (M = 10,000). Therefore, the number of the extended data 59 is 20,000 (N = M × n = 20,000) at the first execution of the machine learning phase S3.
[0118] In this embodiment, linear separation as shown in the above formula (3) was executed. At that time, the number of divisions k in the cross-validation was set to 10, and the minimization of the error function S λ (β) considering the regularization parameter λ was performed using the coordinate descent method.
[0119] In this embodiment, the exclusion of classification error data 59' from a plurality of extension data 59 and the small contribution coefficient β p ' from β p were excluded twice (that is, the machine learning phase S3 was executed three times in total). As shown in FIG. 11, each time the machine learning phase S3 is repeated, the classification accuracy (the output result y of the function f i and the label Y i is improved. In this embodiment, out of the initially set 20,000 pieces of extension data 59, 9,777 pieces of classification error data 59' were thinned out. The number of extension data 59 finally used (the number of extension data 59 where the output result y of the function f i and the label Y i match) is 10,223 pieces.
[0120] And the magnitude of each influence coefficient β p was visualized as shown in FIG. 12. Here, the horizontal axis in FIG. 12 is the ID number (value of p) for distinguishing each influence coefficient β p , and the vertical axis is the value of each influence coefficient β p . Here, the circled marks attached with the numbers from "1" to "10" respectively indicate the marks of the 1st to 10th largest influence coefficients β p . By visualizing the graph as shown in FIG. 12, it becomes possible to clarify the influence coefficient β p with a large contribution degree to the determination of the sound quality by the golden ear, and thus the acoustic characteristics corresponding to the influence coefficient β p . By performing lasso regression while intentionally causing overfitting, the difference between each influence coefficient β p increases.
[0121] <Quantification and objective analysis of acoustic characteristics> As described above, in the sound quality analysis method according to the present embodiment, as exemplified in step S12 of FIG. 4, data augmentation is performed on a plurality of acoustic data 49 to generate a plurality of augmented data 59. According to conventionally known findings, excessive data augmentation has been considered undesirable because it causes overfitting.
[0122] However, according to the present embodiment, data augmentation is intentionally performed to create a situation of overfitting or a situation close thereto. As a result, while the classification performance for unknown acoustic data is suppressed, for the known acoustic data 49 used as teacher data, the influence coefficient β p contributing to the determination of the sound quality of the known acoustic data 49, and another influence coefficient β p have a significantly enlarged difference.
[0123] In addition, by intentionally causing overfitting as described above and combining it with machine learning by Lasso regression (a regression method in which weighting coefficients with small contributions become zero), an influence coefficient β p that significantly contributes to the determination of the sound quality of the known acoustic data 49, and an influence coefficient β p with a smaller contribution can be clearly distinguished. As a result, in the sound quality of labeling based on human subjective evaluation or the like, it is possible to objectively and quantitatively clarify which acoustic characteristics contribute to the evaluation.
[0124] Also, as shown in the above formula (1), by performing standardization based on the average value μ p and the standard deviation σ p of the augmented data 59 generated by data augmentation, even when acoustic data 49 is newly added afterwards, it becomes possible to smoothly perform standardization reflecting the added acoustic data 49. As a result, each step in the sound quality analysis method can be performed as it is without changing the content.
[0125] Also, as shown in steps S44, S46, S47, etc. of FIG. 7, the CPU 3 re-executes the machine learning phase S3 with the classification error data 59' excluded from the extended data 59. The classification error data 59' corresponds to, for example, extended data 59 that is determined to have good sound quality by a human judgment but is determined to have poor sound quality by the classifier. More specifically, when using a linear classifier that performs linear separation, such classification error data 59' corresponds to extended data that crosses the boundary line (a straight line defined by the function f) that characterizes the linear separation.
[0126] According to the present embodiment, the CPU 3 excludes such classification error data 59' from the teacher data. By excluding the data that crosses the boundary line as described above, overfitting is further promoted. This is advantageous in objectively and quantitatively clarifying which acoustic characteristics contribute to the evaluation when evaluating the sound quality.
[0127] Also, as described above, by repeatedly executing the classification test phase S4 and the update of the linear classifier until the classification error data 59' is no longer extracted, overfitting is further promoted. This is advantageous in objectively and quantitatively clarifying which acoustic characteristics contribute to the evaluation when evaluating the sound quality.
[0128] Also, as described above, by performing Lasso regression, the influence coefficient β that contributes little to the determination of good or bad p becomes zero. As a result of intensive studies by the inventors of the present application, according to the obtained findings, by performing Lasso regression again with the small contribution coefficient β p such as the influence coefficient β that has become zero p ' excluded, it is possible to suppress the appearance of the classification error data 59' as described above. As a result, a classifier with higher classification accuracy can be obtained.
[0129] In general, in order to obtain audio data for each vehicle model, it is necessary to prepare a plurality of different vehicle bodies. Although it is required to record acoustic data on a large number of vehicle bodies in order to improve the accuracy of machine learning, it is not easy to prepare a plurality of different vehicle bodies.
[0130] On the other hand, by performing data augmentation on such audio data, as shown in FIGS. 8 and 9, even when there are few vehicle bodies to be analyzed, high-precision machine learning can be performed. Further, in such audio, by clarifying the acoustic characteristics with a large contribution degree, when improving the acoustic performance of the audio, it becomes possible to focus on creating specific acoustic characteristics. As a result, it is possible to efficiently improve the performance of audio products.
[0131] 《Other Embodiments》 In the above embodiment, after data-augmenting a plurality of acoustic data 49 into augmented data 59, label Y is associated with each augmented data 59. i However, the configuration is not limited to such a configuration. That is, the learning preparation phase S2 as the learning preparation step may be configured to label the determination result of the quality of the sound quality for each acoustic data 49 in units of the acoustic data 49 before data augmentation. When configured in this way, the learning preparation phase S2 will be executed before the sample amplification phase S1. The label Y to be attached to each acoustic data 49 i may be stored in advance in the HDD 9 or the like in a state associated with the acoustic data 49.
[0132] Also, the configuration of the flowchart shown in each figure is merely illustrative and can be changed as appropriate. For example, the processes according to steps S48 and S49 in FIG. 7 may be performed at the timing immediately after performing Lasso regression (specifically, the timing immediately before starting the classification test phase S4 after finishing the machine learning phase S3).
[0133] In the above embodiment, as an example of the computer 1, one having a single CPU 3 was illustrated. However, the present disclosure is not limited to this example. The computer 1 includes, in addition to a personal computer, parallel computers such as supercomputers and PC clusters. For example, only some processes such as the sample amplification phase S1 and the learning preparation phase S2 in FIG. 3 may be executed by a personal computer, and only processes that require calculation time such as the machine learning phase S3 may be executed by a parallel computer.
[0134] That is, the "computation unit" in the present disclosure may be configured by combining a computation unit in a specific computer and a computation unit in other computers. In that case, the term "computer 1" will mean a system consisting of a plurality of computers.
[0135] In the above embodiment, acoustic data 49 recorded in a specific vehicle interior was taken as the analysis target. However, the present disclosure is not limited to such analysis targets. The technology disclosed herein can be applied generally to audio used indoors.
Industrial Applicability
[0136] As described above, the present disclosure is useful for analyzing the sound quality of various acoustic data such as audio acoustic data used in a vehicle interior, and has industrial applicability.
Explanation of Signs
[0137] 1 Computer 3 CPU (Computation Unit) 7 RAM (Storage Unit) 11 Display Unit 18 Storage Medium 29 Sound Quality Analysis Program 49 Acoustic Data 59 Extended Data S1 Sample Amplification Phase (Sample Amplification Step) S2 Learning Preparation Phase (Learning Preparation Step) S3 Machine Learning Phase (Machine Learning Step) S4 Analysis Test Phase (Analysis Test Step)
Claims
1. A method for analyzing sound quality that determines a linear classifier for classifying the quality of each of a plurality of acoustic data and analyzes acoustic characteristics based on the linear classifier by using a computer including an arithmetic unit that executes a program and a storage unit that reads data, a sample amplification step in which the storage unit reads the plurality of acoustic data in a quantified state for each acoustic characteristic, and the arithmetic unit executes data expansion on the plurality of acoustic data to generate a plurality of expanded data subjected to data expansion; a learning preparation step in which the arithmetic unit labels, for each acoustic data before data expansion or each expanded data after data expansion, a determination result of the quality of sound in units of acoustic data; a machine learning step of learning the plurality of influence coefficients by setting, as weighting coefficients constituting the linear classifier, a plurality of influence coefficients corresponding to the respective acoustic characteristics, and causing the arithmetic unit to perform Lasso regression in a state where a pair of the plurality of expanded data and the quality of sound labeled in the learning preparation step is set as teacher data; and the plurality of acoustic data correspond to each audio sound recorded in a plurality of types of vehicle interiors, the acoustic characteristics are determined based on the recording data of a sound collection microphone set in the vehicle interior and the vehicle interior environment, and include at least one of a time centroid, an interaural level difference, an interaural time difference, an initial decay time, an initial lateral energy ratio, a speech transmission index, a C value, a D value, and an interaural correlation function. A sound quality analysis method characterized by the above.
2. In the sound quality analysis method according to Claim 1, in the sample amplification step, the arithmetic unit standardizes the plurality of expanded data for each of the influence coefficients based on the average value and standard deviation of the plurality of expanded data. A sound quality analysis method characterized by the above.
3. In the sound quality analysis method according to Claim 1 or 2, a classification test step in which, by inputting the plurality of expanded data into the linear classifier, the arithmetic unit collates, for each of the plurality of expanded data, the quality of sound labeled in the learning preparation step and the quality of sound classified by the linear classifier. In the classification test step, the arithmetic unit determines classification error data indicating extended data among the plurality of extended data, in which the quality of the labeled sound quality and the quality of the sound quality classified by the linear classifier are different. The arithmetic unit updates the linear classifier by re-executing the machine learning step with the classification error data excluded from the plurality of extended data. A method for analyzing sound quality, characterized by the above.
4. A method for analyzing sound quality, which determines a linear classifier for classifying the quality of each of a plurality of acoustic data and analyzes acoustic characteristics based on the linear classifier by using a computer including an arithmetic unit that executes a program and a storage unit that reads data. A sample amplification step in which the storage unit reads the plurality of acoustic data in a state where each acoustic characteristic is quantified, and the arithmetic unit performs data expansion on the plurality of acoustic data to generate a plurality of extended data subjected to data expansion. A learning preparation step in which the arithmetic unit labels the determination result of the quality of the sound quality for each acoustic data before data expansion or each extended data after data expansion in units of acoustic data. A machine learning step in which a plurality of influence coefficients corresponding to the respective acoustic characteristics are set as weighting coefficients constituting the linear classifier, and the arithmetic unit performs Lasso regression in a state where a pair of the plurality of extended data and the quality of the sound quality labeled in the learning preparation step is set as teacher data to learn the plurality of influence coefficients. A classification test step in which the arithmetic unit collates the quality of the sound quality labeled in the learning preparation step and the quality of the sound quality classified by the linear classifier for each of the plurality of extended data by inputting the plurality of extended data into the linear classifier. In the classification test step, the arithmetic unit determines classification error data indicating extended data among the plurality of extended data, in which the quality of the labeled sound quality and the quality of the sound quality classified by the linear classifier are different. The arithmetic unit updates the linear classifier by re-executing the machine learning step with the classification error data excluded from the plurality of extended data. A method for analyzing sound quality, characterized by the above.
5. In the method for analyzing sound quality according to claim 3 or 4, The arithmetic unit repeatedly executes the classification test step and the update of the linear classifier until the classification error data is no longer extracted. A sound quality analysis method characterized by the above.
6. In the sound quality analysis method according to any one of claims 3 to 5, after the arithmetic unit executes the machine learning step, among the plurality of influence coefficients, the arithmetic unit determines a small contribution coefficient indicating an influence coefficient determined to be less than a predetermined reference value through the Lasso regression, the arithmetic unit updates the linear classifier by excluding the classification error data from the plurality of extended data and excluding the small contribution coefficient from the plurality of influence coefficients, and then executing the machine learning step again. A sound quality analysis method characterized by the above.
7. A sound quality analysis program that, by causing a computer including an arithmetic unit that executes a program and a storage unit that reads data, determines a linear classifier for classifying the quality of each of a plurality of acoustic data, and analyzes acoustic characteristics based on the linear classifier, wherein the computer a sample amplification step in which the storage unit reads the plurality of acoustic data in a quantified state for each acoustic characteristic, and the arithmetic unit performs data expansion on the plurality of acoustic data to generate a plurality of extended data subjected to data expansion; a learning preparation step in which the arithmetic unit labels the determination result of the quality of each acoustic data before data expansion or each extended data after data expansion in units of acoustic data; a machine learning step in which the arithmetic unit performs Lasso regression in a state where a plurality of influence coefficients corresponding to the respective acoustic characteristics are set as the weighting coefficients constituting the linear classifier, and a combination of the plurality of extended data and the determination results labeled for each of the plurality of extended data is set as teacher data, thereby determining the plurality of influence coefficients; the plurality of acoustic data correspond to audio sounds recorded in a plurality of types of vehicle interiors, the acoustic characteristics are determined based on the recording data of a sound collection microphone set in the vehicle interior and the vehicle interior environment in the vehicle interior, and include at least one of a time centroid, an interaural level difference, an interaural time difference, an initial decay time, an initial lateral energy ratio, a speech transmission index, a C value, a D value, and an interaural correlation function. A sound quality analysis program characterized by the above. A sound quality analysis program that determines a linear classifier for classifying the quality of each of a plurality of acoustic data and analyzes acoustic characteristics based on the linear classifier by causing a computer including an arithmetic unit that executes a program and a storage unit that reads data to execute the program, wherein: the computer:[[]] a sample amplification step of generating a plurality of augmented data to which data augmentation has been applied by causing the storage unit to read the plurality of acoustic data in a state where each acoustic characteristic is quantified and causing the arithmetic unit to perform data augmentation on the plurality of acoustic data; a learning preparation step of labeling, for each of the acoustic data before data augmentation or each of the augmented data after data augmentation, a determination result of whether the sound quality is good or bad in units of acoustic data; a machine learning step of determining the plurality of influence coefficients by performing Lasso regression by the arithmetic unit in a state where a plurality of influence coefficients corresponding to the respective acoustic characteristics are set as the weighting coefficients constituting the linear classifier and a combination of the plurality of augmented data and the determination results labeled for each of the plurality of augmented data is set as teacher data; a classification test step of causing the arithmetic unit to collate, for each of the plurality of augmented data, the sound quality determination result labeled in the learning preparation step and the sound quality determination result classified by the linear classifier by inputting the plurality of augmented data into the linear classifier; in the classification test step, the arithmetic unit determines classification error data indicating augmented data in which the labeled sound quality determination result and the sound quality determination result classified by the linear classifier are different among the plurality of augmented data; the arithmetic unit updates the linear classifier by re-executing the machine learning step with the classification error data excluded from the plurality of augmented data; characterized in that it is a sound quality analysis program.
9. A computer-readable storage medium storing the sound quality analysis program according to claim 7 or 8, characterized in that it is a computer-readable storage medium.
Citation Information
Patent Citations
Sound quality subjective evaluation system based on local area network architecture
CN110208002A
Neural network
JP1996096084A
JP2004-428776A
Live atmosphere estimation device and program for the same
JP2011250049A
Sound data learning system, sound data learning method and sound data learning device
JP2020056918A