Learning device, learning method, and learning program
The learning device and method improve speaker age estimation accuracy by utilizing semi-supervised learning with coarse-grained age labels, enhancing model performance and resource utilization.
Patent Information
- Application Number
- JP2023532904
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-07-05
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2041-07-05
AI Technical Summary
Existing speaker age estimation technologies face challenges in accurately estimating age from voice data due to the use of low-granularity age information and the inability to effectively utilize data with coarse-grained age labels.
A learning device and method that acquire semi-supervised learning data with voice data and age range labels, estimate speaker age using an LSTM network, calculate a loss using a specific loss function, and update model parameters to minimize the loss, thereby effectively utilizing data with coarse-grained age information.
The proposed solution enables more accurate speaker age estimation from voice data by effectively utilizing data with coarse-grained age labels, improving model accuracy and resource utilization.
Smart Images

Figure 0007694662000014 
Figure 0007694662000015 
Figure 0007694662000016
Abstract
Description
Technical Field
[0001] The present invention relates to a learning device, a learning method, and a learning program.
Background Art
[0002] Conventionally, technologies for automatically estimating age from human voices (hereinafter referred to as "speaker age estimation technologies") have been studied. For example, in the case of use in a call center, it is expected to optimize the synthetic voice playback speed of an automatic response system according to the age of a customer. Here, the speaker age estimation technology is defined as supervised learning for estimating age from voice (or acoustic feature amounts extracted therefrom). In order to train a machine learning model, a large amount of voice associated with age information is required, and the accuracy of the model depends on how many speakers can be collected in a well-balanced manner across a wide range of ages.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the above-described prior art, the age of the speaker cannot be estimated with higher accuracy from the voice. This is because the above-described prior art has the following problems.
[0005] First, age is highly private information, and depending on the dataset, not specific ages such as 25 years old but only low-granularity information (e.g., in their 20s, 18 - 35 years old) may be provided (see, for example, Non-Patent Document 1). Also, in tasks such as directly estimating the speaker's age that have been actively studied in recent years (see, for example, Non-Patent Documents 2 and 3), a dataset with real-valued speaker age information is required, and learning data with low-granularity age information as in the above example cannot be used as it is. To effectively utilize such data, as a general method, a method of treating it as data without age labels within the framework of semi-supervised learning can be considered, but in that case, the low-granularity but provided age information cannot be effectively utilized.
Means for Solving the Problem
[0006] To solve the above-described problems and achieve the object, the learning device according to the present invention includes an acquisition unit that acquires semi-supervised learning data including the voice data of the speaker and the correct label indicating the age range of the speaker, an estimation unit that estimates the age of the speaker of the voice data included in the semi-supervised learning data, a calculation unit that calculates a first loss using a first loss function from the correct label indicating the age range and the estimated age, and an update unit that updates the model parameters so as to minimize the first loss.
[0007] Also, the learning method according to the present invention is a learning method executed by a learning device, including: an acquisition step of acquiring learning data with a weak teacher including a speaker's voice data and a correct label indicating the age range of the speaker; an estimation step of estimating the age of the speaker of the voice data included in the learning data with a weak teacher; a calculation step of calculating a first loss using a first loss function from the correct label indicating the age range and the estimated age; and an update step of updating model parameters so as to minimize the first loss.
[0008] Also, the learning program according to the present invention causes a computer to execute: an acquisition step of acquiring learning data with a weak teacher including a speaker's voice data and a correct label indicating the age range of the speaker; an estimation step of estimating the age of the speaker of the voice data included in the learning data with a weak teacher; a calculation step of calculating a first loss using a first loss function from the correct label indicating the age range and the estimated age; and an update step of updating model parameters so as to minimize the first loss.
Advantages of the Invention
[0009] In the present invention, the age of the speaker can be estimated from the voice with higher accuracy.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
[0011] Hereinafter, embodiments of a learning device, a learning method, and a learning program according to the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to the embodiments described below.
[0012] 〔Embodiment〕 Hereinafter, the processing of the learning system 100, the configuration of the learning device 10, the details of the processing, and the flow of the processing according to the embodiment (appropriately, the present embodiment) will be described in order, and finally the effects of the present embodiment will be described.
[0013] [Processing of Learning System 100] Hereinafter, the processing of the learning system (appropriately, the present system) 100 according to the present embodiment will be described. FIG. 1 is a diagram showing an example of a learning system according to an embodiment. Hereinafter, a configuration example of the present system 100, the processing of the present system 100, the problems of the prior art, and the effects of the present system 100 will be described in order.
[0014] (1. Configuration Example of System 100) The present system 100 shown in FIG. 1 has a learning device 10. Note that the present system 100 may include a plurality of learning devices 10. In the present system 100, data involved in learning with a teacher 20, learning with a weak teacher 30, and voice data 40 are acquired by the learning device 10. Here, the data for learning with a teacher 20 is learning data including the voice data of a speaker and the correct label indicating the age of the speaker. The data for learning with a weak teacher 30 is learning data including the voice data of a speaker and the correct label indicating the age range of the speaker. The voice data 40 is the voice data of a speaker whose age is estimated by the learning device 10.
[0015] (2. Processing of System 100) In the present system 100, first, the learning device 10 acquires learning data (step S1). Here, the learning data acquired by the learning device 10 may be learning data in which the data for learning with a teacher 20 and the data for learning with a weak teacher 30 are mixed, or the data for learning with a teacher 20 and the data for learning with a weak teacher 30 may be acquired individually.
[0016] Next, the learning device 10 extracts features from the voice data included in the acquired learning data and estimates the age using a Long Short-Term Memory (LSTM) network or the like (step S2). At this time, the speaker age estimation technique executed by the learning device 10 is not particularly limited.
[0017] Subsequently, the learning device 10 calculates the loss using the loss function from the estimated age (step S3), updates the model parameters so as to minimize the calculated loss (step S4), and evaluates using the updated model parameters (step S5). The processing of steps S1 to S5 above corresponds to the learning process and evaluation process of the learning device 10.
[0018] Then, the learning device 10 acquires voice data 40 (step S6), and estimates the age of the speaker of the voice data 40 using the updated model parameters (step S7). The processes of steps S6 and S7 correspond to the age estimation process of the learning device 10.
[0019] (3. Problems of the Prior Art) Here, the problems of the conventional speaker age estimation technology will be described. Among the learning data used in the speaker age estimation technology, there is voice data to which only age information within a certain range is given due to privacy issues. At this time, for model learning that can directly estimate age with higher accuracy, it would be good if data with such coarse-grained age information could also be utilized. However, to refine the coarse-grained age labels, (1) it is highly likely to be impossible in the case of publicly available datasets, and (2) even in the case of a dataset recorded independently, it is highly likely to be difficult from the viewpoints of privacy, experimental ethics, cost, etc. In the prior art, since data as described above cannot be utilized, there is a need for a technology that can utilize data with coarse-grained age information and directly estimate age with high accuracy.
[0020] (4. Effects of System 100) As described above, in this system 100, the learning device 10 acquires supervised learning data 20 and semi-supervised learning data 30, estimates the age of the speaker from the voice data included in the acquired learning data, calculates a loss function from the estimated age, and updates the model parameters so as to minimize the calculated loss. For this reason, this system 100 realizes higher model accuracy by enabling the use of data with coarse-grained age information that could not be used for learning to directly estimate age until now as learning data. That is, this system 100 realizes a learning method that effectively uses resources that could not be used until now, and in particular, can effectively utilize data that has only been given ambiguous labels due to the highly private information of age.
[0021] [Configuration of Learning Device 10] Using FIG. 2, the configuration of the learning device 10 according to the present embodiment will be described in detail. FIG. 2 is a block diagram showing a configuration example of the learning device according to the embodiment. The learning device 10 includes an input unit 11, an output unit 12, a communication unit 13, a storage unit 14, and a control unit 15.
[0022] (1. Input Unit 11) The input unit 11 is responsible for inputting various types of information into the learning device 10. For example, the input unit 11 is implemented by a mouse, a keyboard, etc., and receives input of setting information and the like into the learning device 10.
[0023] (2. Output Unit 12) The output unit 12 is responsible for outputting various types of information from the learning device 10. For example, the output unit 12 is implemented by a display or the like, and outputs setting information and the like stored in the learning device 10.
[0024] (3. Communication Unit 13) The communication unit 13 is responsible for data communication with other devices. For example, the communication unit 13 performs data communication with each communication device. Also, the communication unit 13 can perform data communication with an operator's terminal (not shown).
[0025] (4. Storage Unit 14) The storage unit 14 stores various types of information referred to when the control unit 15 operates, and various types of information acquired when the control unit 15 operates. The storage unit 14 includes a supervised learning data storage unit 14a, a semi-supervised learning data storage unit 14b, and a model parameter storage unit 14c. Here, the storage unit 14 can be implemented by, for example, a semiconductor memory element such as a RAM (Random Access Memory), a flash memory, or a storage device such as a hard disk, an optical disk, etc. In the example of FIG. 2, the storage unit 14 is installed inside the learning device 10, but it may be installed outside the learning device 10, or a plurality of storage units may be installed.
[0026] (4-1. Supervised Learning Data Storage Unit 14a) Using FIG. 3, an example of the learning data stored in the supervised learning data storage unit 14a will be described. FIG. 3 is a diagram showing an example of the data stored in the supervised learning data storage unit according to the embodiment. In FIG. 3, the supervised learning data storage unit 14a stores S x persons and N voices.
[0027] The supervised learning data storage unit 14a stores supervised learning data 20 which is learning voice data used for model learning and to which a speaker age (real value) is assigned as a correct label. For example, the supervised learning data storage unit 14a stores, as learning data, "voice ID", "speaker ID", "speaker age" (real value), "speaker gender", and "voice". In FIG. 3, the supervised learning data storage unit 14a stores a voice to which an age for supervised learning is assigned as a real value (e.g., 25 years old).
[0028] (4-2. Semi-supervised learning data storage unit 14b) Using FIG. 4, an example of the learning data stored in the semi-supervised learning data storage unit 14b will be described. FIG. 4 is a diagram showing an example of the data stored in the semi-supervised learning data storage unit according to the embodiment. In FIG. 4, the semi-supervised learning data storage unit 14b stores S w persons and M voices.
[0029] The semi-supervised learning data storage unit 14b stores semi-supervised learning data 30 which is learning voice data used for model learning and to which a speaker age (age range) is assigned as a correct label. For example, the semi-supervised learning data storage unit 14b stores, as learning data, "voice ID", "speaker ID", "speaker age" (age range), "speaker gender", and "voice". In FIG. 4, the semi-supervised learning data storage unit 14b stores a voice to which an age for semi-supervised learning is assigned within an arbitrary range (e.g., twenties).
[0030] Here, in the learning data storage unit 14b with weak supervision, all the correct labels assigned as the above age range may have the same granularity, or voices with different granularities may be mixed. Also, the above learning data 30 with weak supervision may be subjected to data augmentation and / or split into learning, development, and evaluation sets as necessary.
[0031] (4-3. Data Standard) In the above teacher-assisted learning data storage unit 14a and the learning data storage unit 14b with weak supervision, the standard of the stored audio data is not particularly limited. For example, the stored audio data may be in the form of 16 kHz·16-bit signed integer, 1-channel linear PCM (Pulse Code Modulation) in the frequency band, or 8 kHz·16-bit signed integer, 1-channel G.711 compressed format in the frequency band. Note that the audio data shown in FIGS. 3 and 4 represents the audio waveform as a relationship between the passage of time and the audio signal intensity and is expressed as a 16-bit signed integer.
[0032] (4-4. Model Parameter Storage Unit 14c) The model parameter storage unit 14c stores a set of parameters Θ' (learned parameters) that have been learned and optimized by the estimation unit 15b, calculation unit 15c, and update unit 15d of the control unit 15 described later. Here, the parameter set Θ' is used to estimate the age from the input audio during evaluation data or actual use.
[0033] (5. Control Unit 15) The control unit 15 controls the entire learning device 10. As shown in FIG. 2, the control unit 15 includes an acquisition unit 15a, an estimation unit 15b, a calculation unit 15c, and an update unit 15d. Here, the control unit 15 is, for example, an electronic circuit such as a CPU (Central Processing Unit) or MPU (Micro Processing Unit), or an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array).
[0034] (5-1. Acquisition Unit 15a) The acquisition unit 15a acquires the weakly supervised learning data 30 including the speaker's voice data and the correct label indicating the age range of the speaker. For example, the acquisition unit 15a acquires the weakly supervised learning data 30 including the speaker's voice data, the age range of the speaker, the gender of the speaker, etc. Also, the acquisition unit 15a acquires the supervised learning data 20 including the speaker's voice data and the correct label indicating the age of the speaker. For example, the acquisition unit 15a acquires the supervised learning data 20 including the speaker's voice data, the real value of the speaker's age, the gender of the speaker, etc.
[0035] On the other hand, the acquisition unit 15a acquires the weakly supervised learning data 30 from the weakly supervised learning data storage unit 14b. Also, the acquisition unit 15a acquires the supervised learning data 20 from the supervised learning data storage unit 14a. Furthermore, the acquisition unit 15a may acquire the learning data via the input unit 11, or may acquire the learning data from other terminals or databases via the communication unit 13.
[0036] (5-2. Estimation Unit 15b) The estimation unit 15b estimates the age of the speaker of the voice data included in the weakly supervised learning data 30. Also, the estimation unit 15b estimates the age of the speaker of the voice data included in the supervised learning data 20. For example, the estimation unit 15b estimates using a model such as SVR (Support Vector Regression) or a neural network that can projectively transform the speaker representation vector into age, with the speaker representation vector learned using a separately prepared dataset with a large number of speakers as the input.
[0037] In addition, the estimation unit 15b estimates the speaker's age by using any time-series acoustic feature such as FBANK (Log Mel-filter bank channel output) or MFCC (Mel-Frequency Cepstral Coefficient) as a model input to a neural network capable of handling time-series feature quantities such as RNN or transformer. Further, the estimation unit 15b improves the accuracy by using any technique such as normalization of feature quantities, batch normalization, L1 / L2 regularization, etc., and estimates the speaker's age.
[0038] The estimation unit 15b estimates the speaker's age by defining a learning model and model parameters for estimating the numerical value of age as a regression problem. Also, the estimation unit 15b estimates the speaker's age by defining a learning model and model parameters for classifying into classes corresponding to age as a classification problem. Note that the details of the age estimation process for the learning data will be described later in [Details of the process] (1. Learning data age estimation process).
[0039] On the other hand, the estimation unit 15b outputs the estimation result to the calculation unit 15c. Note that the estimation unit 15b may store the estimation result in the storage unit 14.
[0040] (5-3. Calculation unit 15c) The calculation unit 15c calculates the first loss using the first loss function from the correct label indicating the age range and the estimated age. Also, the calculation unit 15c calculates the second loss using the second loss function from the correct label indicating the age and the estimated age. Further, the calculation unit 15c determines the data including the correct label indicating the age range of the speaker as the weakly supervised learning data 30 and calculates the first loss using the first loss function, and determines the data including the correct label indicating the age of the speaker as the supervised learning data 20 and calculates the second loss using the second loss function. Note that the details of the age estimation process for the learning data will be described later in [Details of the process] (2. Loss function calculation process).
[0041] On the one hand, the calculation unit 15c outputs the calculation result to the update unit 15d. Note that the calculation unit 15d may store the calculation result in the storage unit 14.
[0042] (5-4. Update Unit 15d) The update unit 15d updates the model parameters so as to minimize the first loss. Also, the update unit 15d updates the model parameters so as to minimize the second loss. For example, the update unit 15d updates the model parameters using the stochastic gradient descent method so as to minimize the loss calculated by the calculation unit 15c. Also, the update unit 15d stores the updated model parameters in the model parameter storage unit 14c of the storage unit 14.
[0043] [Details of the Process] The details of the process according to this embodiment will be described with reference to FIGS. 5 to 9, mathematical formulas, etc. Hereinafter, the learning data age estimation process, the loss function calculation process, and the learning application process will be described in detail.
[0044] (1. Learning Data Age Estimation Process) Hereinafter, the learning data age estimation process will be described in detail. The estimation unit 15b of the learning device 10 estimates the speaker age using the following speaker age estimation techniques. For example, the estimation unit 15b estimates using a model that can projectively transform a speaker expression vector into age, such as SVR or a neural network, with the speaker expression vector learned using a separately prepared dataset of a large number of speakers as input.
[0045] Also, the estimation unit 15b may estimate the speaker age by using any time-series acoustic feature amount such as FBANK or MFCC as a model input to a neural network that can handle time-series feature amounts such as RNN or transformer. Further, the estimation unit 15b may introduce efforts to improve accuracy using any technique such as normalization of feature amounts, batch normalization, L1 / L2 regularization, etc., and estimate the speaker age.
[0046] During learning, the estimation unit 15b estimates the age from the input speech using, for example, appropriate random numbers or a set of model parameters Θ pre-trained in another task. At this time, the estimation unit 15b does not specify a combination of the supervised learning data 20 and the weakly supervised learning data 30 either. For example, the estimation unit 15b may take a form such as multi-task learning by mixing the supervised learning data 20 and the weakly supervised learning data 30 at an arbitrary ratio within the same batch, or may switch the learning between the supervised learning data 20 and the weakly supervised learning data 30 for each arbitrary iteration / epoch, or may take any other method. Also, during evaluation, the estimation unit 15b calculates the estimated age from the input speech using the learned parameter set Θ'.
[0047] (2. Loss function calculation process) Using FIGS. 5 to 7, equations, etc., the details of the loss function calculation process according to the present embodiment will be described. FIGS. 5 to 7 are diagrams showing an example of the loss function calculation process according to the embodiment. Hereinafter, the outline of the speaker age estimation process, the loss function calculation process for the classification problem, and the loss function calculation process for the regression problem will be described in this order.
[0048] (2-1. Outline) The calculation unit 15c of the learning device 10 calculates a loss for updating the model parameters from the estimated speaker age. Hereinafter, the outline of the speaker age estimation process that is a prerequisite for the above calculation process will be described. Here, the direct estimation of the speaker age can be defined as either a regression problem or a classification problem.
[0049] In the case of a regression problem, the model f and the parameter set Θ are defined to directly estimate the age value, and the estimated age ŷ is defined as in the following equation (1). Here, x in the following equation (1) represents the input speech.
[0050]
Equation
[0051] Regarding the classification problem, the model f and the parameter set Θ are defined as a classification problem with one class per age, and the posterior probability shown in the following equation (2) is estimated for each age y from the input speech x. n For, the posterior probability shown in the following equation (2) is estimated.
[0052]
Number
[0053] Here, the estimated age ŷ may be the age indicated by the maximum value of the posterior probability for each age y as the estimation result as shown in the following equation (3), or the expected value obtained from the posterior probability may be used as the estimation result as shown in the following equation (4). n For, the posterior probability shown in the following equation (2) is estimated.
[0054]
Number
[0055]
Number
[0056] (2-2. Classification Problem) Hereinafter, the loss function calculation process for the classification problem will be described in the order of the calculation process of the loss (second loss) by the loss function (second loss function) of the supervised learning data 20 and the calculation process of the loss (first loss) by the loss function (first loss function) of the semi-supervised learning data 30.
[0057] (2-2-1. Supervised Learning Data 20) During model learning using the supervised speech x, the calculation unit 15c calculates using, for example, the cross-entropy loss shown in the following equation (5) as the loss function. The calculation unit 15c may also use other loss functions such as the KL (Kullback-Leibler) divergence loss.
[0058]
Number
[0059] At this time, the calculation unit 15c may use, as the correct target, a hard target that assumes only the correct age shown in the following formula (6) as correct, or a soft target that approximates a normal distribution with the correct age shown in the following formula (7) as the average as the correct target.
[0060]
Equation
[0061]
Equation
[0062] Note that in the above formulas (5) to (7), T(y n ) represents the target value for each age y n , N represents the set of all ages defined as correct, and σ represents the standard deviation of the normal distribution set in advance as a hyperparameter.
[0063] (2-2-2. Data 30 for learning with a weak teacher) When performing model learning using the voice w for learning with a weak teacher, the calculation unit 15c defines a loss function such that, for example, if the estimation result falls within the age range indicated by the correct label, it is regarded as correct. At this time, the calculation unit 15c may define a soft target shown in the following formula (8) in which the age range indicated by the correct label is equally correct as the correct target such as cross-entropy loss, or may calculate the loss function from the total value of the posterior probabilities without assuming a distribution.
[0064]
Equation
[0065] Further, the calculation unit 15c may define all of its range as correct, such as in multi-label learning. Also, when using the total value of posterior probabilities, the calculation unit 15c may be trained to provide a margin to allow for some estimation error. Here, the margin of the calculation unit 15c may be uniquely determined, or may be determined fluidly according to the width of the correct range. The following formula (9) shows an example of a loss function in which the total value of the posterior probability for the age indicated by the correct label is set to "1.0".
[0066] [Number]
[0067] Here, in the above formula (9), MSE (Mean Square Error) is used as the loss, but MAE (Mean Absolute Error) or binary cross-entropy loss may also be used. Also, weights may be applied to the loss function as necessary.
[0068] In the above formulas (8) and (9), Y[w] represents the set of ages indicated by the correct label of the voice w for learning with a weak teacher. For example, in the case of the correct label "in their 20s", Y[w] ∈ (20, 21, 22, 23, 24, 25, 26, 27, 28, 29).
[0069] Here, with reference to FIGS. 5 and 6, the losses given by the above formulas (8) and (9) will be described for the case where the correct label is "in their 30s". FIGS. 5 and 6 are diagrams showing an example of loss function calculation processing according to the embodiment.
[0070] FIG. 5(1) shows the predicted posterior probability for each age in the case of "w: correct age label = in their 30s". Also, FIG. 5(2) shows a probability distribution in which ages from 30 to 39 are equally considered correct. Then, based on the above formula (8), the cross-entropy loss is calculated between the posterior probability and the correct probability distribution, and L = 4.41 is calculated.
[0071] Fig. 6(1) shows the predicted posterior probabilities for each age when "w: correct age label = 30s". Then, based on the above equation (9), the loss is calculated such that the sum of the posterior probabilities enclosed by the dashed line in Fig. 6(2) and the margin equals 1.0. At this time, if m = 0.2, then L = 0.41 is calculated.
[0072] (2-3. Regression problem) Below, regarding the loss function calculation process for the regression problem, the calculation process using the loss function of the supervised learning data 20 and the calculation process using the loss function of the semi-supervised learning data 30 will be explained in this order.
[0073] (2-3-1. Supervised learning data) For data with the speaker age label as a real value, the calculation unit 15c calculates the MSE shown in the following equation (10) and the MAE shown in the following equation (11) as losses. Also, the calculation unit 15c may calculate using a method that regards the estimation error below a certain value (ε) as the correct answer using the ε-insensitive loss shown in the following equation (12).
[0074] [Number]
[0075] [Number]
[0076] [Number]
[0077] (2-3-2. Semi-supervised learning data) For data with a speaker age label as a range, if the estimation result falls within that range, the calculation unit 15c shall consider it equally correct. For example, the calculation unit 15c may calculate using a loss function such that the estimation result falls within the correct range as in the following formula (13), or may calculate using any other arbitrary loss function. Further, the calculation unit 15c may weight the loss function as necessary.
[0078]
Number
[0079] Here, with reference to FIG. 7, when the correct label is "in their 30s", the loss given by the above formula (13) will be explained. FIG. 7 is a diagram showing an example of the loss function calculation process according to the embodiment. FIG. 7 is a graph showing the numerical value of the estimated age on the horizontal axis and the numerical value of the loss on the vertical axis, and the loss given by the above formula (13) is minimized at the estimated age "in their 30s" (30 to 39 years old).
[0080] (3. Processing of the learning application) With reference to FIGS. 8 and 9, the details of the processing of the learning application according to this embodiment will be described. FIGS. 8 and 9 are diagrams showing an example of the setting screen of the learning application according to the embodiment. Hereinafter, the outline of the setting screen of the learning application and the details of the setting screen of the learning application will be described in order.
[0081] (3-1. Outline of the setting screen) With reference to FIG. 8, the outline of the setting screen of the learning application will be described. In the input field shown in FIG. 8(1), for example, an arbitrary value (character, numerical value) is entered by the operation of the user. Also, the input field may contain a default value. In the input field shown in FIG. 8(2), the user selects and inputs from predetermined options by a pull-down or the like by the operation of the user. The button shown in FIG. 8(3) is used when searching for a file, and the user searches using a file management application or the like with the "Browse" button by the operation of the user.
[0082] (3-2. Details of the Setting Screen) Using Fig. 9, the details of the setting screen of the learning application will be described. Below, the specification of the supervised learning data 20, the specification of the model parameters, the specification of the semi-supervised learning data 30, the specification of the mini-batch method, the specification of the loss function, the specification of the margin, and the specification of the learning timing will be explained in this order.
[0083] (3-2-1. Specification of the Supervised Learning Data 20) In Fig. 9(1), the normal supervised learning data 20 is specified by the user's operation. Here, the data can be specified by any method, and the format does not matter as long as the learning program can read it. For example, you can read a text or file with the data path entered, or you can specify the directory where the data is placed and read all the data under it. Also, although it is assumed that the correct answer for each file is entered in "Label data", the correct answer may be described together in the file with the data path described, or the correct answer may be included in the data name, without having to read it as a separate file.
[0084] (3-2-2. Specification of the Model Parameters) In Fig. 9(2), the parameters (batch size, optimization method, model structure, loss function setting, etc.) necessary for the learning of a general neural network are specified by the user's operation. Here, the parameters (design of the loss function) necessary for learning with normal supervised labels are also specified. In addition to this, for example, you may specify whether to incorporate techniques such as L1 / L2 regularization, batch normalization, and normalization of features.
[0085] (3-2-3. Specification of the Semi-Supervised Learning Data 30) In Fig. 9(3), the semi-supervised teacher data 30 is specified by the user's operation. Here, for the learning / evaluation / correct answer information, the program is specified so that it can read the data in the same way as the above-mentioned supervised learning data 20.
[0086] (3-2-4. Specification of Mini-batch Method) In Fig. 9(4), select the method for mini-batching the weakly supervised learning data 30 by the user's operation. For example, in the case of the pull-down "value", directly specify the batch size. In Fig. 9(4), "64" is input by default. Also, in the case of the pull-down "rate", specify the ratio to the "Mini-Batch size". For example, if the above ratio is set to "0.5", it is equal to "value:32" in the default setting. Furthermore, the batch size can also be specified in a free format.
[0087] (3-2-5. Specification of Loss Function) In Fig. 9(5), specify the loss function and its weight coefficient for weakly supervised learning by the user's operation. For example, {CE, MSE, BCE, MAE, ML} can be selected. Note that the above "CE" represents cross-entropy loss, "BCE" represents binary cross-entropy loss, and "ML" represents BCE for multi-label learning. Also, in the loss function calculation process, the weight coefficient specified here is multiplied by the loss function and then the error backpropagation is performed. In the example of Fig. 9(5), MSE is used as the loss and the error backpropagation is performed with the weight coefficient = 1.0.
[0088] (3-2-6. Specification of Margin) In Fig. 9(6), specify the margin during loss calculation by the user's operation. For example, in the case of "order", specify it as a constant regardless of the age range of the data. Also, in the case of "chance rate", vary it according to the age range and specify the magnification.
[0089] (3-2-7. Specification of Learning Timing) In Fig. 9(7), the timing during the learning of the weakly supervised learning data 30 is specified by the user's operation. For example, in the case of "same", the mini-batches of the normal supervised learning data 20 and the weakly supervised learning data 30 are combined. In the example of Fig. 9(7), it becomes 64 + 64 = 128. Also, in the case of "iter", every x iterations, supervised learning → weakly supervised learning → supervised learning... are alternately performed. Further, in the case of "epoch", every x epochs, supervised learning → weakly supervised learning → supervised learning... are alternately performed. Furthermore, in the cases of "iter" and "epoch", "x" may be specified as an arbitrary number. Note that in the example of Fig. 9(7), the default is set to 1.
[0090] [Flow of processing] Using Fig. 10, the overall flow of the learning process will be described. Fig. 10 is a flowchart showing the overall flow of the learning process according to the present embodiment. Hereinafter, after explaining the overall flow of the process, the outline of each process will be described.
[0091] (1. Overall flow of processing) First, the acquisition unit 15a of the learning device 10 executes a learning data acquisition process (step S101). Next, the estimation unit 15b of the learning device 10 executes a learning data age estimation process (step S102). Then, the calculation unit 15c of the learning device 10 executes a loss function calculation process (step S103). Finally, the update unit 15d of the learning device 10 executes a model parameter update process (step S104) and ends the process. Note that the above steps S101 to S104 can also be executed in a different order. Also, among the above steps S101 to S104, some processes may be omitted.
[0092] (2. Flow of each process) First, the learning data acquisition process by the acquisition unit 15a will be described. In the learning data acquisition process, the acquisition unit 15a acquires the supervised learning data 20 and the semi-supervised learning data 30 from the storage unit 14. At this time, the acquisition unit 15a may acquire the supervised learning data 20 and the semi-supervised learning data 30 from the storage unit 14 via the input unit 11 or the communication unit 13 without referring to the storage unit 14.
[0093] Second, the learning data age estimation process by the estimation unit 15b will be described. In the learning data age estimation process, the estimation unit 15b extracts feature quantities from the voice data included in the acquired learning data, and estimates the age using a learning model such as SVR or a neural network. Details of the learning data age estimation process are described in [Details of the Process] (1. Learning Data Age Estimation Process) above.
[0094] Third, the loss function calculation process by the calculation unit 15c will be described. In the loss function calculation process, the calculation unit 15c determines whether the supervised learning data 20 or the semi-supervised learning data 30 is the learning data, determines whether it is a classification problem or a regression problem as the estimation method, and calculates the loss using a loss function suitable for each learning data and estimation method. Details of the learning data age estimation process are described in [Details of the Process] (2. Loss Function Calculation Process) above.
[0095] Fourth, the model parameter update process by the update unit 15d will be described. In the model parameter update process, the update unit 15d updates the model parameters using the stochastic gradient descent method or the like so as to minimize the loss calculated by the calculation unit 15c.
[0096] [Effects of the Embodiment] First, in the learning process according to the above-described embodiment, learning data 30 with weak labels including the speaker's voice data and the correct label indicating the range of the speaker's age is obtained, the age of the speaker in the voice data included in the learning data 30 with weak labels is estimated, and from the correct label indicating the age range and the estimated age, a first loss is calculated using a first loss function, and the model parameters are updated so as to minimize the first loss. Therefore, in this process, the age of the speaker can be estimated from the voice with higher accuracy.
[0097] Second, in the learning process according to the above-described embodiment, learning data 20 with labels including the speaker's voice data and the correct label indicating the speaker's age is obtained, the age of the speaker in the voice data included in the learning data 20 with labels is estimated, and from the correct label indicating the age and the estimated age, a second loss is calculated using a second loss function, and the model parameters are updated so as to minimize the second loss. Therefore, in this process, by using more learning data, the age of the speaker can be estimated from the voice with higher accuracy.
[0098] Third, in the learning process according to the above-described embodiment, data including the correct label indicating the range of the speaker's age is determined as learning data 30 with weak labels and the first loss is calculated using the first loss function, and data including the correct label indicating the speaker's age is determined as learning data 20 with labels and the second loss is calculated using the second loss function. Therefore, in this process, by effectively using more learning data, the age of the speaker can be estimated from the voice with higher accuracy.
[0099] Fourth, in the learning process according to the above-described embodiment, the age of the speaker is estimated by defining a learning model and model parameters for estimating the numerical value of the age as a regression problem. Therefore, in this process, by effectively using more learning data, the age of the speaker can be estimated from the voice with higher accuracy based on the regression model.
[0100] Fifthly, in the learning process according to the above-described embodiment, the age of the speaker is estimated by defining a learning model and model parameters for classifying into classes corresponding to ages as a classification problem. Therefore, in this process, by effectively using more learning data, the age of the speaker can be estimated from the voice with higher accuracy based on the classification model.
[0101] 〔System Configuration, etc.〕 Each component of each device illustrated according to the above embodiment is a functional concept, and does not necessarily have to be physically configured as illustrated. That is, the specific form of the distribution and integration of each device is not limited to that illustrated, and all or part of it can be functionally or physically distributed and integrated in any unit according to various loads, usage situations, etc. Furthermore, each processing function performed by each device can be realized in whole or in any part by a CPU and a program analyzed and executed by the CPU, or can be realized as hardware by wired logic.
[0102] Also, among the processes described in the above embodiment, all or part of the processes described as being automatically performed can be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, regarding the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above document and drawings, they can be arbitrarily changed unless otherwise specified.
[0103] 〔Program〕 Also, a program can be created that describes the processing executed by the learning device 10 described in the above embodiment in a computer-executable language. In this case, by the computer executing the program, the same effects as the above embodiment can be obtained. Furthermore, such a program can be recorded on a computer-readable recording medium, and the same processing as the above embodiment can be realized by causing the computer to read and execute the program recorded on this recording medium.
[0104] FIG. 11 is a diagram showing a computer that executes a program. As illustrated in FIG. 11, the computer 1000 has, for example, a memory 1010, a CPU 1020, a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070, and these components are connected by a bus 1080.
[0105] The memory 1010 includes, as illustrated in FIG. 11, a ROM (Read Only Memory) 1011 and a RAM 1012. The ROM 1011 stores, for example, a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090 as illustrated in FIG. 11. The disk drive interface 1040 is connected to a disk drive 1100 as illustrated in FIG. 11. For example, a removable storage medium such as a magnetic disk or an optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120 as illustrated in FIG. 11. The video adapter 1060 is connected to, for example, a display 1130 as illustrated in FIG. 11.
[0106] Here, as illustrated in FIG. 11, the hard disk drive 1090 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. That is, the above programs are stored, for example, in the hard disk drive 1090 as program modules in which instructions to be executed by the computer 1000 are described.
[0107] Also, the various data described in the above embodiments are stored as program data, for example, in the memory 1010 or the hard disk drive 1090. Then, the CPU 1020 reads out the program modules 1093 and program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as necessary, and executes various processing procedures.
[0108] Note that the program modules 1093 and program data 1094 related to the program are not limited to being stored in the hard disk drive 1090. For example, they may be stored in a removable storage medium and read out by the CPU 1020 via a disk drive or the like. Alternatively, the program modules 1093 and program data 1094 related to the program may be stored in another computer connected via a network (such as a LAN (Local Area Network) or a WAN (Wide Area Network)) and read out by the CPU 1020 via the network interface 1070.
[0109] The above embodiments and their modifications are included in the invention described in the claims and the scope equivalent thereto, in the same way as the technology disclosed in the present application.
Description of Reference Numerals
[0110] 10 Learning device 11 Input unit 12 Output unit 13 Communication unit 14 Storage unit 14a Teacher-based learning data storage unit 14b Weak-teacher-based learning data storage unit 15 Control unit 15a Acquisition unit 15b Estimation unit 15c Calculation unit 15d Update unit 20 Teacher-based learning data 30 Weak-teacher-based learning data 40 Audio data 100 Learning System
Claims
1. An acquisition unit that acquires learning data with weak supervision including the speaker's voice data and the correct label indicating the age range of the speaker; An estimation unit that estimates the age of the speaker of the voice data included in the learning data with weak supervision using a learning model that outputs the age in response to the input of the voice data; A calculation unit that calculates a first loss using a first loss function from the correct label indicating the age range and the estimated age; An update unit that updates the model parameters, which are the learned parameters of the learning model, so as to minimize the first loss; A learning device characterized by comprising the above.
2. The acquisition unit further acquires supervised learning data including the speaker's voice data and the correct label indicating the age of the speaker, The estimation unit further estimates the age of the speaker of the voice data included in the supervised learning data, The calculation unit further calculates a second loss using a second loss function from the correct label indicating the age and the estimated age, The update unit further updates the model parameters so as to minimize the second loss, The learning device according to claim 1, characterized by the above.
3. The calculation unit determines data including the correct label indicating the age range of the speaker as learning data with weak supervision and calculates the first loss using the first loss function, and determines data including the correct label indicating the age of the speaker as supervised learning data and calculates the second loss using the second loss function, The learning device according to claim 2, characterized by the above.
4. The estimation unit estimates the age of the speaker by defining the learning model and the model parameters for estimating the numerical value of the age as a regression problem. The learning device according to claim 3, characterized by the above.
5. The estimation unit estimates the age of the speaker by defining the learning model and the model parameters that classify the speaker into a class corresponding to the age as a classification problem. The learning device according to claim 3, characterized in that.
6. A learning method executed by a learning device, An acquisition step of acquiring learning data with weak supervision including the voice data of the speaker and the correct label indicating the age range of the speaker; An estimation step of estimating the age of the speaker of the voice data included in the learning data with weak supervision using a learning model that outputs the age in response to the input of the voice data; A calculation step of calculating a first loss using a first loss function from the correct label indicating the age range and the estimated age; An update step of updating the model parameters, which are the learned parameters of the learning model, so as to minimize the first loss; A learning method characterized by including.
7. An acquisition step of acquiring learning data with weak supervision including the voice data of the speaker and the correct label indicating the age range of the speaker; An estimation step of estimating the age of the speaker of the voice data included in the learning data with weak supervision using a learning model that outputs the age in response to the input of the voice data; A calculation step of calculating a first loss using a first loss function from the correct label indicating the age range and the estimated age; An update step of updating the model parameters, which are the learned parameters of the learning model, so as to minimize the first loss; A learning program characterized by causing a computer to execute.
Citation Information
Patent Citations
Voice-based age prediction method, device and equipment
CN111128235A
Age prediction method, device and equipment
CN111210840A
Age estimation method, device and equipment
CN111261196A
Face age automatic estimation method based on weak supervision
CN111274882A
Method and device for training VAD by using weak supervision data
CN112786029A