Estimation program, estimation device, and estimation method
Patent Information
- Application Number
- JP2023122265
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-07-27
- Publication Date
- 2026-09-14
- Estimated Expiration
- 2043-07-27
Smart Images

Figure 0007920104000005 
Figure 0007920104000006 
Figure 0007920104000007
Abstract
Description
[Technical Field]
[0001] Embodiments of the present invention include an estimation program , recommendation fixed device , and Sadakata In the law To relate to. [Background technology]
[0002] The degree of engagement or engagement when two or more people are conversing or exchanging opinions is collectively referred to as communication activity. Traditionally, methods have been used to estimate communication activity.
[0003] For example, Patent Document 1 proposes a meeting function calculation device that evaluates meeting activity by considering the relationships between meeting participants, such as the number of times one participant elicits comments from other meeting participants. This device acquires time-series data of audio and images representing the actions or phenomena between meeting participants and calculates correlation data between two or more meeting participants. Then, it generates and displays meeting support information from the time-series data and correlation data according to a predetermined evaluation scale for the meeting.
[0004] Patent Document 2 proposes a meeting support system that quantitatively evaluates the atmosphere of a meeting and assists in facilitating smooth discussions. This system acquires the voices of meeting participants, generates meeting support information such as the degree of discussion engagement based on frequently occurring keywords extracted from the voice-text data, and provides this meeting support information to the meeting participants. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] Japanese Patent Publication No. 2009-163431 [Patent Document 2] Japanese Patent Publication No. 2017-112545 [Overview of the Initiative] [Problems that the invention aims to solve]
[0006] However, conventional technology had a problem in that the accuracy of estimating communication activity decreased significantly when a person could not wear a microphone or when the person's face was not facing the camera. In other words, conventional technology sometimes resulted in a decrease in the accuracy of estimating communication activity.
[0007] The problem that this invention aims to solve is an estimation program that can improve the accuracy of estimating communication activity. , recommendation fixed device , and Sadakata Law It is about providing. [Means for solving the problem]
[0008] The estimation program according to the embodiment includes an acquisition step of acquiring an acoustic signal relating to a sound source and a non-acoustic signal which is an image signal that is temporally synchronized with the acoustic signal and correlates with the morphology of the facial region or the movement of a body part when the speaker is speaking; an acoustic feature calculation step of calculating acoustic features based on the acoustic signal; and a non-acoustic feature calculation step of calculating non-acoustic features that correlate with the speaker's communication features in accordance with the non-acoustic changes represented by the non-acoustic signal. A speech presence calculation step that calculates the speech presence based on the acoustic characteristics and the non-acoustic characteristics, An activity feature calculation step that calculates an activity feature score based on the acoustic features and the non-acoustic features, and an activity estimation step that estimates the communication activity score based on a comparison of the activity feature score with an evaluation scale, A situation information generation step that generates situation information representing the communication status based on the voice presence level, and an output control step that outputs the communication activity level and situation information representing the communication status, This is a program designed to instruct a computer to execute a command. [Brief explanation of the drawing]
[0009] [Figure 1] Schematic diagram of the estimation device. [Figure 2] Diagram illustrating input signals, acoustic signals, and non-acoustic signals. [Figure 3] A schematic diagram illustrating the flow of the communication activity level estimation process. [Figure 4] A flowchart illustrating the process for estimating communication activity. [Figure 5]Schematic diagram of a learning device. [Figure 6] An explanatory diagram illustrating an example of an integrated neural network configuration. [Figure 7] A flowchart illustrating the learning process flow. [Figure 8] A table showing the results of the verification. [Figure 9] Hardware configuration diagram. [Modes for carrying out the invention]
[0010] The estimation program, learning program, estimation device, learning device, estimation method, learning method, and learning model according to this embodiment will be described in detail below with reference to the drawings.
[0011] (First Embodiment) Figure 1 is a schematic diagram of an example of the estimation device 10 of this embodiment. The estimation device 10 is a computer that estimates the communication activity level.
[0012] The estimation device 10 includes a processing unit 11, a storage unit 12, an input unit 13, a communication unit 14, a display unit 15, and an acoustic output unit 16. The processing unit 11, storage unit 12, input unit 13, communication unit 14, display unit 15, and acoustic output unit 16 are communicated together via a bus 17 or the like.
[0013] The memory unit 12 stores various calculation results from the processing unit 11 and estimated programs executed by the processing unit 11. The processing unit 11 will be described in detail later. The memory unit 12 is an example of a computer-readable recording medium. The memory unit 12 is composed of, for example, ROM (Read Only Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), integrated circuit memory, etc.
[0014] The input unit 13 receives various commands from the user. Examples of input units 13 include a keyboard, mouse, various switches, a touchpad, and a touch panel display. Signals corresponding to the commands input by the user through the operation of the input unit 13 are supplied to the processing unit 11. The input unit 13 may also be a computer connected to the processing unit 11 via a wired or wireless connection.
[0015] The communication unit 14 is an interface for communicating information with external devices connected to the estimation device 10 via a network. For example, the communication unit 14 receives acoustic and non-acoustic signals from a device that collects acoustic and non-acoustic signals, and receives the first trained model 21, the second trained model 22, the third trained model 23, the fourth trained model 24, the fifth trained model 25, and the sixth trained model 26 from the learning device described later. Acoustic signals, non-acoustic signals, the first trained model 21, the second trained model 22, the third trained model 23, the fourth trained model 24, the fifth trained model 25, and the sixth trained model 26 will be described later.
[0016] The display unit 15 displays various information. The display unit 15 can be any display known in the art, such as a CRT (Cathode-Ray Tube) display, a liquid crystal display, an organic EL (Electro-Luminescence) display, an LED (Light-Emitting Diode) display, a plasma display, or any other display. The display unit 15 may also be a projector.
[0017] The sound output unit 16 converts electrical signals into sound and radiates them. Any speaker known in the art, such as a magnetic speaker, dynamic speaker, condenser speaker, or other speaker, can be used as the sound output unit 16.
[0018] The processing unit 11 performs information processing in the estimation device 10.
[0019] The processing unit 11 includes an acquisition unit 11A, an acoustic feature calculation unit 11B, a non-acoustic feature calculation unit 11C, an integrated feature calculation unit 11D, a voice presence calculation unit 11E, a laughter presence calculation unit 11F, an emotion feature calculation unit 11G, an activity feature calculation unit 11H, an activity estimation unit 11I, a situation information generation unit 11J, and an output control unit 11K.
[0020] The processing unit 11 includes a processor such as a CPU (Central Processing Unit) and memory such as RAM (Random Access Memory). The processing unit 11 performs communication activity estimation processing, such as estimating the communication activity level of an input signal, by executing an estimation program stored in the storage unit 12. The estimation program is recorded on a non-temporary computer-readable recording medium. The processing unit 11 reads the estimation program from the recording medium and executes it to realize the acquisition unit 11A, acoustic feature calculation unit 11B, non-acoustic feature calculation unit 11C, integrated feature calculation unit 11D, voice presence calculation unit 11E, laughter presence calculation unit 11F, emotion feature calculation unit 11G, activity feature calculation unit 11H, activity estimation unit 11I, situation information generation unit 11J, and output control unit 11K. The estimation program may have multiple modules, each of which implements a divided function of the respective unit (acquisition unit 11A to output control unit 11K).
[0021] The hardware implementation of the processing unit 11 is not limited to the above-described embodiment. For example, it may be configured by a circuit such as an Application Specific Integrated Circuit (ASIC) that implements at least one of the following: acquisition unit 11A, acoustic feature calculation unit 11B, non-acoustic feature calculation unit 11C, integrated feature calculation unit 11D, voice presence calculation unit 11E, laughter presence calculation unit 11F, emotion feature calculation unit 11G, activity feature calculation unit 11H, activity estimation unit 11I, situation information generation unit 11J, and output control unit 11K. At least one of the acquisition unit 11A, acoustic feature calculation unit 11B, non-acoustic feature calculation unit 11C, integrated feature calculation unit 11D, voice presence calculation unit 11E, laughter presence calculation unit 11F, emotion feature calculation unit 11G, activity feature calculation unit 11H, activity estimation unit 11I, situation information generation unit 11J, and output control unit 11K may be implemented on a single integrated circuit or individually on multiple integrated circuits.
[0022] The acquisition unit 11A acquires an input signal that includes an acoustic signal and a non-acoustic signal related to a sound source. The input signal is a video signal that includes an acoustic signal and an image signal related to the same sound source. The acoustic signal and the non-acoustic signal are time-series signals. More specifically, the acoustic signal and the non-acoustic signal are time-series signals consisting of multiple frames that are consecutive in time. The acoustic signal and the non-acoustic signal are time-synchronized on a frame-by-frame basis.
[0023] In this embodiment, the acoustic signal and the non-acoustic signal are assumed to originate from the same sound source. However, it is sufficient that there is a correlation between the acoustic signal and the non-acoustic signal; the sound sources of the acoustic signal and the non-acoustic signal do not need to be strictly identical.
[0024] Figure 2 is an explanatory diagram illustrating an example of an input signal, an acoustic signal, and a non-acoustic signal.
[0025] Figure 2 shows an example where the input signal is a video signal. A video signal is a time-series signal that includes a time-series acoustic signal and a time-series image signal that are synchronized in time. The length of the time interval of the video signal is not particularly limited, but it is assumed to be, for example, about 30 seconds. The video signal is collected by a video camera device that includes, for example, a microphone and an imaging device.
[0026] An acoustic signal is a signal relating to speech from a speech source. The speech source is, for example, the speaker. More specifically, an acoustic signal includes both a speech signal originating from the speaker's utterance and a noise signal originating from noise. In other words, an acoustic signal is a signal relating to both speech and noise from the speaker, who is the speech source.
[0027] The acoustic signal is collected by a microphone. The microphone collects the sound of the speaker's utterance, converts the sound pressure of the collected sound into an analog electrical signal (acoustic signal), and converts this acoustic signal into a digital time-domain electrical signal (acoustic signal) through A / D conversion. The time-domain acoustic signal is acquired by the acquisition unit 11A and converted into a frequency-domain acoustic signal by short-time Fourier transform or the like.
[0028] Non-acoustic signals are signals other than the acoustic signal related to the speaker that are collected approximately simultaneously with the acoustic signal. More specifically, non-acoustic signals are signals other than the voice of the speaker, which is the source of the sound, such as the speaker's image or sensor signals of the speaker detected by sensors. Specifically, for example, non-acoustic signals are image signals related to the speaker who is speaking, or sensor signals related to the physiological responses of the speaker's lips and facial muscles and brain waves caused by speech. In this embodiment, we will explain assuming that the non-acoustic signal is an image signal related to the speaker. That is, in this embodiment, we will explain as an example the form in which the non-acoustic signal is an image signal that is temporally synchronized with the acoustic signal.
[0029] Image signals are collected by an imaging device that includes multiple image sensors, such as a CCD (Charge Coupled Device). The imaging device optically photographs the speaker and generates digital spatial image signals (image data) related to the speaker on a frame-by-frame basis. The image signals are required to correlate with the speaker's speech. The image frame only needs to include at least the lip region, whose shape changes in response to speech, as the subject of the image, and may also include the entire face of the speaker. Image signals are acquired frame by frame by the acquisition unit 11A.
[0030] Time-series acoustic signals and time-series image signals are defined according to equation (1) below.
[0031]
number
[0032] In equation (1), A represents a time-series acoustic signal. In the equation for A representing the acoustic signal, T represents the number of dimensions in the time domain of the frame to be processed. F represents the number of dimensions in the frequency domain of the frame to be processed. V represents a time-series image signal. In the equation for V representing the image signal, T represents time, H represents the height of the image, W represents the width of the image, and C represents the number of dimensions of the color channels. That is, the time-series acoustic signal A is an acoustic signal with T dimensions in the time domain and F dimensions in the frequency domain of the frame to be processed. The image signal V is an image signal with dimensions of time T, height H, width W, and color channels C.
[0033] Returning to Figure 1, we continue the explanation.
[0034] The acoustic feature calculation unit 11B calculates acoustic features based on the acoustic signal. Acoustic features are characteristic quantities of the acoustic signal. Acoustic features have values based on the acoustic signal and have values that correlate with the speaker's voice.
[0035] The acoustic feature calculation unit 11B calculates acoustic features for each frame of the acoustic signal. Therefore, the acoustic feature calculation unit 11B calculates time-series data consisting of continuous acoustic features by calculating acoustic features for each frame from a time-series acoustic signal consisting of multiple frames.
[0036] In detail, the acoustic feature calculation unit 11B calculates acoustic features corresponding to the changes in sound represented by the acoustic signal. The changes in sound represented by the acoustic signal are represented, for example, by pitch, volume, peak value, etc.
[0037] In this embodiment, the acoustic feature calculation unit 11B calculates acoustic features from the acoustic signal using the first trained model 21.
[0038] The first trained model 21 is a neural network trained to take an acoustic signal as input and output acoustic features. For example, an encoder network trained to convert an acoustic signal into acoustic features can be used as this neural network. The first trained model 21 is stored in the memory unit 12, etc., and used by the acoustic feature calculation unit 11B to calculate acoustic features. The first trained model 21 is generated by the learning device described later.
[0039] The relationship between acoustic signals and acoustic features is as follows: An acoustic signal is the time-series waveform data of the sound pressure value of the speech emitted by a speaker. The acoustic signal correlates with the speech emitted by the speaker. For example, the peak value of the acoustic signal is relatively high when the speaker is speaking and relatively low when the speaker is not speaking. Acoustic features are designed so that their value correlates with the peak value of the acoustic signal; in other words, they are designed to distinguish between the speech and silence components contained in the acoustic signal. For example, the higher the peak value of the acoustic signal, the higher the value of the acoustic feature, and the lower the peak value of the acoustic signal, the lower the value of the acoustic feature.
[0040] The non-acoustic feature calculation unit 11C calculates non-acoustic features based on the non-acoustic signal. Non-acoustic features are feature quantities of the non-acoustic signal. Non-acoustic features have values based on the non-acoustic signal and have special values that correlate with the speaker's speech.
[0041] The non-acoustic feature calculation unit 11C calculates non-acoustic features for each frame. Therefore, the non-acoustic feature calculation unit 11C calculates non-acoustic features for each frame from a time-series non-acoustic signal consisting of multiple frames, thereby calculating time-series data consisting of non-acoustic features that are continuous in time.
[0042] In detail, the non-acoustic feature calculation unit 11C calculates non-acoustic features corresponding to non-acoustic changes represented by non-acoustic signals. Non-acoustic changes are represented, for example, by the movement of at least a part of the speaker.
[0043] In this embodiment, the non-acoustic feature calculation unit 11C calculates non-acoustic features from non-acoustic signals using the second trained model 22.
[0044] The second pre-trained model 22 is a neural network trained to take a non-acoustic signal as input and output a non-acoustic feature. For example, this neural network takes an image signal V as input and outputs an image feature E. V An encoder network trained to convert to is used. The second trained model 22 is stored in the memory unit 12, etc., and used by the non-acoustic feature calculation unit 11C to calculate non-acoustic features. The second trained model 22 is generated by the learning device described later.
[0045] The relationship between image signals and image features is as follows: Image signals correlate with the morphology of the facial region and the movement of body parts when the speaker is speaking. Image features are designed to calculate the communication features contained in the image signals. Specifically, image signals have different morphologies when the speaker is making noise and when they are not. Image features are designed so that their values correlate with the speaker's cocommunication features. For example, the more the speaker's body parts move, the higher the value of the image features, and the less the speaker's body parts move, the lower the value of the image features.
[0046] The integrated feature calculation unit 11D calculates an integrated feature obtained by integrating an acoustic feature and a non-acoustic feature. The integrated feature calculation unit 11D calculates the integrated feature based on the acoustic feature and the non-acoustic feature. The integrated feature calculation unit 11D calculates the integrated feature for each frame.
[0047] The integrated feature is an integrated feature obtained by combining an acoustic feature and a non-acoustic feature. The integrated feature is represented, for example, by a multi-dimensional vector obtained by combining an acoustic feature and a non-acoustic feature.
[0048] Specifically, the integrated feature is represented by the following formula (2).
[0049]
Math.
[0050] In formula (2), Z AV represents an integrated feature. E A represents an acoustic feature. E V represents an image feature. t represents a frame time. That is, the integrated feature Z AV is an integrated feature of the acoustic feature E A and the image feature E V , and is calculated for each frame time t.
[0051] In formula (2), E A (t、d) is an acoustic feature vector at frame time t∈{1, 2, …, T} and coordinate d∈{1, 2, …, D} compressed into D dimensions, and is an example of an acoustic feature. E V (t、e) is an image feature vector at frame time t∈{1, 2, …, T} and coordinate e∈{1, 2, …, E} compressed into D dimensions, and is an example of an image feature. Z AV (t、d+e) is an integrated feature vector at frame time t∈{1, 2, …, T} and coordinate d∈{1, 2, …, d+e} compressed into D dimensions, and is an example of an integrated feature.
[0052] The speech presence calculation unit 11E calculates speech presence based on acoustic and non-acoustic features. Speech presence is a feature quantity related to the presence or absence of a speech signal. A speech signal refers to an acoustic signal that originates from a speaker's utterance among acoustic signals. Speech presence is a feature value that represents the presence or absence of a speech signal that originates from a speaker's utterance among acoustic signals that contain noise.
[0053] The voice presence calculation unit 11E calculates the voice presence for each frame. In this embodiment, the voice presence calculation unit 11E calculates the voice presence based on integrated features. Therefore, the voice presence calculation unit 11E calculates the voice presence for each frame from the integrated features of a time series consisting of multiple frames, thereby calculating time-series data consisting of voice presences that are continuous in the time series.
[0054] In detail, the speech presence calculation unit 11E calculates speech presence according to the speech characteristics derived from the acoustic signal represented by the acoustic characteristics and the non-acoustic signal represented by the non-acoustic characteristics.
[0055] In this embodiment, the voice presence calculation unit 11E calculates the voice presence from integrated features using the third trained model 23.
[0056] The third pre-trained model 23 is a neural network trained to take integrated features as input and output speech presence. More specifically, the third pre-trained model 23 is a neural network trained to reduce a third loss related to the difference between the correct speech presence label and the speech presence. As such a neural network, for example, an estimation network trained to estimate speech presence from integrated features is used. The third pre-trained model 23 is stored in the memory unit 12, etc., and used by the speech presence calculation unit 11E to calculate speech presence. The third pre-trained model 23 is generated by the learning device described later.
[0057] The laughter presence calculation unit 11F calculates the laughter presence based on acoustic and non-acoustic features. The laughter presence is a feature quantity related to the presence or absence of a laughter signal. A laughter signal refers to an acoustic signal that originates from a speaker's laughter. The laughter presence has a feature value that represents the presence or absence of a speaker's laughter signal from an acoustic signal that contains noise.
[0058] The laughter presence calculation unit 11F calculates the laughter presence for each frame. In this embodiment, the laughter presence calculation unit 11F calculates the laughter presence based on integrated features. Therefore, the laughter presence calculation unit 11F calculates the laughter presence for each frame from the integrated features of a time series consisting of multiple frames, thereby calculating time-series data consisting of laughter presences that are continuous in the time series.
[0059] In detail, the laughter presence calculation unit 11F calculates the laughter presence according to the characteristics of laughter derived from acoustic signals represented by acoustic characteristics and non-acoustic signals represented by non-acoustic characteristics.
[0060] In this embodiment, the laughter presence calculation unit 11F calculates the laughter presence using the fourth trained model 24.
[0061] The fourth pre-trained model 24 is a neural network trained to take integrated features as input and output laughter presence. More specifically, the fourth pre-trained model 24 is a neural network trained to reduce a fourth loss related to the difference between the ground truth laughter presence label and the actual laughter presence. As this neural network, for example, an estimation network trained to estimate laughter presence from integrated features is used. The fourth pre-trained model 24 is stored in the memory unit 12, etc., and used by the laughter presence calculation unit 11F to calculate laughter presence. The fourth pre-trained model 24 is generated by the learning device described later.
[0062] The emotion feature calculation unit 11G calculates the emotion feature score based on acoustic and non-acoustic features. The emotion feature score is a feature quantity related to the emotion signal. The emotion signal refers to the acoustic signal that originates from the emotion of the speaker's speech within the acoustic signal. The emotion feature score has a value that represents the emotional features of the speaker included in the acoustic signal. In other words, the emotion feature score has a feature value that represents the emotion of the speech signal that originates from the speaker's utterance within the acoustic signal.
[0063] The emotion feature calculation unit 11G calculates emotion features for each frame. In this embodiment, the emotion feature calculation unit 11G calculates emotion features based on integrated features. Therefore, the emotion feature calculation unit 11G calculates emotion features for each frame from the integrated features of a time series consisting of multiple frames, thereby calculating time series data consisting of emotion features that are continuous in the time series.
[0064] In detail, the emotion feature calculation unit 11G calculates an emotion feature score that corresponds to the emotion-related features contained in the acoustic signal and the non-acoustic signal, which is derived from the acoustic signal represented by acoustic features and the non-acoustic signal represented by non-acoustic features.
[0065] In this embodiment, the emotion feature calculation unit 11G calculates the emotion feature score from the integrated features using the fifth trained model 25.
[0066] The fifth pre-trained model 25 is a neural network trained to take integrated features as input and output sentiment feature scores. The fifth pre-trained model 25 is a neural network trained to reduce the fifth loss related to the difference between the ground truth sentiment feature score label and the sentiment feature score for the sentiment features. As this neural network, for example, an estimation network trained to estimate sentiment feature scores from integrated features is used. The fifth pre-trained model 25 is stored in the memory unit 12, etc., and used by the sentiment feature score calculation unit 11G to calculate sentiment feature scores. The fifth pre-trained model 25 is generated by the learning device described later.
[0067] The activation feature calculation unit 11H calculates the activation feature based on acoustic and non-acoustic features. The activation feature is a feature quantity related to the communication activation signal. The communication activation signal refers to the acoustic signal that originates from the heightened communication of the speaker. The activation feature has a feature value that represents the communication activation features of the speaker from an acoustic signal that contains noise.
[0068] The activity feature calculation unit 11H calculates the activity feature for each frame. In this embodiment, the activity feature calculation unit 11H calculates the activity feature based on integrated features. Therefore, the activity feature calculation unit 11H calculates the activity feature for each frame from the integrated features of a time series consisting of multiple frames, thereby calculating time series data consisting of activity features that are continuous in the time series.
[0069] In detail, the activity feature calculation unit 11H calculates an activity feature that represents the activity characteristics derived from the acoustic signal represented by the acoustic characteristics and the non-acoustic signal represented by the non-acoustic characteristics.
[0070] In this embodiment, the activity feature calculation unit 11H calculates the activity feature from the integrated features using the sixth trained model 26.
[0071] The sixth pre-trained model 26 is a neural network trained to take integrated features as input and output sentiment features. More specifically, the sixth pre-trained model 26 is a neural network trained to reduce the sixth loss related to the difference between the ground truth activation feature label and the activation feature for an activation feature. As such a neural network, for example, an estimation network trained to estimate activation features from integrated features is used. The sixth pre-trained model 26 is generated by the learning device described later. The sixth pre-trained model 26 is stored in the memory unit 12, etc., and used by the activation feature calculation unit 11H to calculate activation features. The sixth pre-trained model 26 is generated by the learning device described later.
[0072] The activity level estimation unit 11I estimates the communication activity level based on a comparison of the activity feature level with an evaluation scale. The evaluation scale is set as the boundary between the values corresponding to the communication activity level. For example, consider a case where the activity feature level takes values from "-2" to "3". In this case, the evaluation scale is pre-set as follows: "-2": no conversation, "-1": thinking, "0": stagnant, "1": somewhat active, "2": active, "3": very active. Through the estimation process by the activity level estimation unit 11I, the communication activity level for the input signal, including acoustic and non-acoustic signals, is estimated for each frame.
[0073] The situation information generation unit 11J generates situation information representing the communication situation based at least on the presence of voices. In this embodiment, the situation information generation unit 11J generates situation information based on the presence of voices, the presence of laughter, and the emotional characteristic score.
[0074] Situational information refers to information about the communication situation. For example, situational information includes information about the communication situation through the speaker's utterances, such as the amount of speech, the number of laughs, and the rate of emotional change.
[0075] In other words, the situation information generation unit 11J generates situation information that includes the amount of speech from a speech source represented by the speech presence level, the number of times laughter occurs represented by the laughter presence level, and the rate of change in emotion represented by the emotion characteristic level. For example, the situation information generation unit 11J stores in advance in the storage unit 12 a table that associates the speech presence level with the amount of speech, a table that associates the laughter presence level with the number of times laughter occurs, and a table that associates the emotion characteristic level with the rate of change in emotion. Then, the situation information generation unit 11J uses these tables to identify the amount of speech, the number of times laughter occurs, and the rate of change in emotion that correspond to the speech presence level, laughter presence level, and emotion characteristic level calculated by the speech presence level calculation unit 11E, the laughter presence level calculation unit 11F, and the emotion characteristic level calculation unit 11G, respectively, and calculates situation information that includes these.
[0076] The output control unit 11K outputs various information to the display unit 15 and the sound output unit 16. In this embodiment, the output control unit 11K outputs the communication activity level estimated by the activity level estimation unit 11I. Alternatively, the output control unit 11K may output the communication activity level estimated by the activity level estimation unit 11I and the situation information generated by the situation information generation unit 11J.
[0077] For example, the output control unit 11K displays communication activity level and status information on the display unit 15. In this case, it is preferable that the output control unit 11K displays the communication activity level and status information on the display unit 15 in a visual association with the acoustic signal and / or image signal. Alternatively, the output control unit 11K may output the communication activity level and status information to the acoustic output unit 16.
[0078] In this embodiment, a configuration in which the processing unit 11 of the estimation device 10 includes an integrated feature calculation unit 11D will be described as an example. However, the processing unit 11 may also be configured without an integrated feature calculation unit 11D.
[0079] In other words, the speech presence calculation unit 11E does not necessarily have to calculate speech presence from an integrated feature based on acoustic and non-acoustic features. The speech presence calculation unit 11E may calculate speech presence directly from acoustic and non-acoustic features, or from other intermediate outputs based on acoustic and non-acoustic features. As an example, the speech presence calculation unit 11E may calculate speech presence from the acoustic and non-acoustic features to be processed using a neural network that has been trained to take acoustic and non-acoustic features as input and output speech presence.
[0080] Similarly, the laughter presence calculation unit 11F does not necessarily have to calculate the laughter presence from an integrated feature based on acoustic and non-acoustic features. The laughter presence calculation unit 11F may calculate the laughter presence directly from acoustic and non-acoustic features, or from other intermediate outputs based on acoustic and non-acoustic features. As an example, the laughter presence calculation unit 11F may calculate the laughter presence from the acoustic and non-acoustic features being processed using a neural network trained to take acoustic and non-acoustic features as input and output the laughter presence.
[0081] Similarly, the emotion feature calculation unit 11G does not necessarily have to calculate the emotion feature score from an integrated feature based on acoustic and non-acoustic features. The emotion feature calculation unit 11G may calculate the emotion feature score directly from acoustic and non-acoustic features, or from other intermediate outputs based on acoustic and non-acoustic features. As an example, the emotion feature calculation unit 11G may calculate the emotion feature score from the acoustic and non-acoustic features of the processing target using a neural network that has been trained to take acoustic and non-acoustic features as input and output an emotion feature score.
[0082] Similarly, the activity feature calculation unit 11H does not necessarily have to calculate the activity feature score from an integrated feature based on acoustic and non-acoustic features. The activity feature calculation unit 11H may calculate the activity feature score directly from acoustic and non-acoustic features, or from other intermediate outputs based on acoustic and non-acoustic features. As an example, the activity feature calculation unit 11H may calculate the activity feature score from the acoustic and non-acoustic features of the processing target using a neural network that has been trained to take acoustic and non-acoustic features as input and output the activity feature score.
[0083] Furthermore, the processing unit 11 of the estimation device 10 may be configured without at least one of the laughter presence calculation unit 11F and the emotion feature calculation unit 11G. If the estimation device 10 does not utilize the laughter presence and / or emotion feature, the processing unit 11 may be configured without the laughter presence calculation unit 11F and / or the emotion feature calculation unit 11G.
[0084] Next, an example of the flow of the communication activity level estimation process executed by the processing unit 11 of the estimation device 10 in this embodiment will be described.
[0085] Figures 3 and 4 are explanatory diagrams illustrating an example of the flow of the communication activity estimation process performed by the processing unit 11. Figure 3 is a schematic diagram showing the flow of the communication activity estimation process by the processing unit 11. Figure 4 is a flowchart illustrating an example of the flow of the communication activity estimation process by the processing unit 11. For the purpose of explaining Figures 3 and 4, it is assumed that the non-acoustic signal is an image signal. The communication activity estimation process is performed by the processing unit 11 operating according to the estimation program stored in the memory unit 12, etc.
[0086] The acquisition unit 11A acquires an input signal that includes an acoustic signal and an image signal (step S1, step S100).
[0087] The acoustic feature calculation unit 11B uses the first trained model 21 to calculate acoustic features E from the acoustic signal A acquired in steps S1 and S100. A Calculate (Steps S2, S102).
[0088] The non-acoustic feature calculation unit 11C uses the second trained model 22 to calculate image features E from the image signal V acquired in steps S1 and S100. V Calculate (step S3, step S103).
[0089] The processing order of steps S2 and S102, and steps S3 and S103 is not particularly limited. For example, steps S2 and S102 may be executed after steps S3 and S103. Alternatively, steps S2 and S102 may be executed in parallel with steps S3 and S103.
[0090] The integrated feature calculation unit 11D calculates the integrated feature Z from the acoustic features calculated in steps S2 and S102 and the image features calculated in steps S3 and S103. AV Calculate (step S4, step S104).
[0091] The voice presence calculation unit 11E uses the third trained model 23 to calculate the integrated feature Z calculated in steps S4 and S104. AV From the voice existence calculation Y VAD Calculate (step S5, step S105).
[0092] The laughter presence calculation unit 11F uses the fourth trained model 24 to calculate the integrated feature Z calculated in steps S4 and S104. AV Laughter presence level Y^ LD Calculate (step S6, step S106).
[0093] The emotion feature calculation unit 11G uses the fifth trained model 25 to calculate the integrated feature Z calculated in steps S4 and S104. AV From emotional characteristic score Y^ ER Calculate (step S7, step S107).
[0094] The activity feature calculation unit 11H uses the sixth trained model 26 to calculate the integrated feature Z calculated in steps S4 and S104. AV From the activity feature Y^ COMM Calculate (step S8, step S108).
[0095] Note that the processing order of steps S5 to S8 and steps S105 to S108 is not limited to the order described above. For example, steps S7 to S5 and steps S107 to S105 may be executed after steps S8 and S108. Alternatively, steps S5 to S8 and steps S105 to S108 may be executed in parallel.
[0096] The activity estimation unit 11I calculates the activity characteristic Y^ in steps S8 and S108. COMM Based on comparison with the evaluation scale, the level of communication activity is estimated (Step S9, Step S109).
[0097] The situation information generation unit 11J calculates the voice presence Y^ calculated in steps S5 and S105. VAD And the laughter presence calculation Y^ calculated by steps S6 and S106 LD And the emotional characteristic score Y^ calculated by steps S7 and S107. ER Based on this, situation information is generated (steps S10, S110).
[0098] The output control unit 11K outputs the communication activity level estimated in steps S9 and S109, and the status information generated in steps S10 and S110 (steps S11 and S110). Then, it terminates the communication activity level estimation process.
[0099] Note that the communication activity estimation process shown in Figures 3 and 4 is illustrative, and the communication activity estimation process according to this embodiment is not limited to the procedure shown in Figures 3 and 4. As described above, the processing unit 11 may be configured without an integrated feature calculation unit 11D. In this case, steps S4 and S104 may not be executed. Also, the processing unit 11 may be configured without a laughter presence calculation unit 11F and / or an emotion feature calculation unit 11G. In this case, steps S6, S106, and / or steps S7 and S107 may not be executed.
[0100] As described above, the estimation device 10 of this embodiment comprises an acquisition unit 11A, an acoustic feature calculation unit 11B, a non-acoustic feature calculation unit 11C, an activity feature calculation unit 11H, and an activity estimation unit 11I. The acquisition unit 11A acquires acoustic signals and non-acoustic signals related to the sound source. The acoustic feature calculation unit 11B calculates acoustic features based on the acoustic signals. The non-acoustic feature calculation unit 11C calculates non-acoustic features based on the non-acoustic signals. The activity feature calculation unit 11H calculates activity features based on the acoustic features and non-acoustic features. The activity estimation unit 11I estimates communication activity based on a comparison of the activity features with an evaluation scale.
[0101] In contrast to conventional technologies such as those described in Patent Documents 1 and 2, which estimate communication activity using audio and images, there is a problem in that the accuracy of communication activity estimation drops significantly when a person cannot wear a microphone or when a person's face is not facing the camera. Furthermore, there are still limitations for practical applications, such as estimating communication activity from ceiling-mounted microphones and cameras, and there is a need to improve the accuracy of communication activity estimation.
[0102] On the other hand, the estimation device 10 of this embodiment uses acoustic signals and non-acoustic signals to calculate the activity feature score based on the acoustic features of the voice signal and the non-acoustic features of the non-acoustic signal, and estimates the communication activity score from the activity feature score.
[0103] Thus, the estimation device 10 of this embodiment estimates the communication activity level based on acoustic and non-acoustic signals, making it possible to estimate the communication activity level with high accuracy even when it is not possible to collect human voices at close range or when it is not possible to identify a person's face from an image.
[0104] Therefore, the estimation device 10 of this embodiment can improve the accuracy of estimating communication activity.
[0105] Furthermore, the estimation device 10 of this embodiment calculates the activity feature score using acoustic features based on acoustic signals and non-acoustic features based on non-acoustic signals. Therefore, in addition to the above effects, the estimation device 10 of this embodiment can improve the efficiency of mutual feature extraction for the activity feature score.
[0106] Furthermore, the estimation device 10 of this embodiment generates situational information based on the presence of voice, the presence of laughter, and the emotional characteristic score. Therefore, by generating situational information in addition to the communication activity level, the estimation device 10 of this embodiment can provide information that enables the detection of communication activity level with higher accuracy, in addition to the effects described above.
[0107] (Second embodiment) Figure 5 is a schematic diagram of an example of the learning device 30 in this embodiment. The learning device 30 is a computer that learns the trained model used in the estimation device 10.
[0108] The learning device 30 includes a processing unit 31, a storage unit 32, an input unit 33, a communication unit 34, a display unit 35, and an audio output unit 36. The processing unit 31, storage unit 32, input unit 33, communication unit 34, display unit 35, and audio output unit 36 are connected to each other via a bus 37 or the like.
[0109] The memory unit 32, input unit 33, communication unit 34, display unit 35, and sound output unit 36 are the same as the memory unit 12, input unit 13, communication unit 14, display unit 15, and sound output unit 16 of the above embodiment.
[0110] The processing unit 31 performs information processing in the learning device 30.
[0111] The processing unit 31 includes an acquisition unit 31A, an acoustic feature calculation unit 31B, a non-acoustic feature calculation unit 31C, an integrated feature calculation unit 31D, a voice presence calculation unit 31E, a laughter presence calculation unit 31F, an emotion feature calculation unit 31G, an activity feature calculation unit 31H, an update unit 31I, an update completion determination unit 31J, and an output control unit 31K.
[0112] The processing unit 31 includes a processor such as a CPU and memory such as an RA. The processing unit 31 executes a learning process, such as generating the first trained model 21 to the sixth trained model 26 by learning, by executing a learning program stored in the memory unit 32. The learning program is recorded on a non-temporary computer-readable recording medium. The processing unit 31 reads the learning program from the recording medium and executes it to realize the acquisition unit 31A, acoustic feature calculation unit 31B, non-acoustic feature calculation unit 31C, integrated feature calculation unit 31D, voice presence calculation unit 31E, laughter presence calculation unit 31F, emotion feature calculation unit 31G, activity feature calculation unit 31H, update unit 31I, update completion determination unit 31J, and output control unit 31K. The learning program may have multiple modules, each of which implements a divided function of the respective unit (acquisition unit 31A to output control unit 31K).
[0113] The hardware implementation of the processing unit 31 is not limited to the above-described embodiment. For example, it may be configured by a circuit such as an application-specific integrated circuit (ASIC) that implements at least one of the following: acquisition unit 31A, acoustic feature calculation unit 31B, non-acoustic feature calculation unit 31C, integrated feature calculation unit 31D, voice presence calculation unit 31E, laughter presence calculation unit 31F, emotion feature calculation unit 31G, activity feature calculation unit 31H, update unit 31I, update completion determination unit 31J, and output control unit 31K. At least one of the acquisition unit 31A, acoustic feature calculation unit 31B, non-acoustic feature calculation unit 31C, integrated feature calculation unit 31D, voice presence calculation unit 31E, laughter presence calculation unit 31F, emotion feature calculation unit 31G, activity feature calculation unit 31H, update unit 31I, update completion determination unit 31J, and output control unit 31K may be implemented on a single integrated circuit or individually on multiple integrated circuits.
[0114] The acquisition unit 31A acquires an input signal, which includes acoustic and non-acoustic signals related to the speech source, as training data. The input signal is the same as described above. That is, the input signal is a time-series signal, which includes a time-series acoustic signal and a time-series non-acoustic signal. As described above, the non-acoustic signals include image signals related to the speaker, sensor signals related to the physiological responses of the speaker's lips and facial muscles due to speech, and brain waves, etc.
[0115] The acoustic feature calculation unit 31B calculates acoustic features from an acoustic signal using the first neural network 41. The acoustic features calculated by the acoustic feature calculation unit 31B are the same as those calculated by the acoustic feature calculation unit 11B. The first neural network 41 is a neural network that takes an acoustic signal as input and outputs acoustic features. The first trained model 21 is generated by training the first neural network 41.
[0116] The non-acoustic feature calculation unit 31C calculates non-acoustic features from non-acoustic signals using the second neural network 42. The non-acoustic features calculated by the non-acoustic feature calculation unit 31C are the same as the non-acoustic features calculated by the non-acoustic feature calculation unit 11C. The second neural network 42 is a neural network that takes non-acoustic signals as input and outputs non-acoustic features. A second trained model 22 is generated by training the second neural network 42.
[0117] The integrated feature calculation unit 31D calculates integrated features based on acoustic and non-acoustic features. The integrated features calculated by the integrated feature calculation unit 31D are the same as the integrated features calculated by the integrated feature calculation unit 11D.
[0118] The speech presence calculation unit 31E calculates the speech presence from integrated features using the third neural network 43. The speech presence calculated by the speech presence calculation unit 31E is the same as the speech presence calculated by the speech presence calculation unit 11E. A third trained model 23 is generated by training the third neural network 43.
[0119] The third neural network 43 is a neural network trained to reduce a third loss related to the difference between the ground truth speech presence label and the actual speech presence. The ground truth speech presence label is a ground truth label indicating that each frame of the acoustic signal is an audio segment or a non-audio segment. The ground truth speech presence label may be created manually based on the acoustic signal, automatically using speech recognition technology, or a combination of these manual and automatic methods.
[0120] The laughter presence calculation unit 31F calculates the laughter presence from integrated features using the fourth neural network 44. The laughter presence calculated by the laughter presence calculation unit 31F is the same as the laughter presence calculated by the laughter presence calculation unit 11F. A fourth trained model 24 is generated by training the fourth neural network 44.
[0121] The fourth neural network 44 is a neural network trained to reduce a fourth loss related to the difference between the ground truth laughter presence label and the actual laughter presence. The ground truth laughter presence label is a ground truth label indicating that each frame of the acoustic signal is a laughter interval or a non-laughter interval. The ground truth laughter presence label may be created manually based on the acoustic signal, automatically using speech recognition technology, or a combination of these manual and automatic methods.
[0122] The emotion feature calculation unit 31G calculates emotion features from integrated features using the fifth neural network 45. The emotion features calculated by the emotion feature calculation unit 31G are the same as the emotion features calculated by the emotion feature calculation unit 11G. A fifth trained model 25 is generated by training the fifth neural network 45.
[0123] The fifth neural network 45 is a neural network trained to reduce a fifth loss related to the difference between the ground truth sentiment feature label and the sentiment feature for an emotion feature. The ground truth sentiment feature label is a ground truth label indicating an emotion state, assigned to each frame of an acoustic signal. The ground truth sentiment feature label may be created manually based on an acoustic or non-acoustic signal, automatically using emotion recognition technology, or a combination of these manual and automatic methods.
[0124] The activation feature calculation unit 31H calculates the activation feature score from the integrated features using the sixth neural network 46. The activation feature score calculated by the activation feature calculation unit 31H is the same as the activation feature score calculated by the activation feature calculation unit 11H. By training the sixth neural network 46, the sixth trained model 26 is generated.
[0125] The sixth neural network 46 is a neural network trained to reduce a sixth loss relating to the difference between the ground truth activation feature label and the activation feature calculation for activation features. The ground truth activation feature label is a ground truth label that indicates the activation state of communication, assigned to each frame of an acoustic signal. The ground truth activation feature label may be created manually based on an acoustic or non-acoustic signal, or it may be created automatically using emotion recognition technology, or it may be created by combining these manual and automatic methods.
[0126] The update unit 31I updates the first neural network 41, the second neural network 42, the third neural network 43, the fourth neural network 44, the fifth neural network 45, and the sixth neural network 46 by learning that reduces the combined loss of the third loss, the fourth loss, the fifth loss, and the sixth loss.
[0127] As described above, the third loss is the loss related to the difference between the correct label for voice presence (correct voice presence label) and the actual voice presence. The fourth loss is the loss related to the difference between the correct label for laughter presence (correct laughter presence label) and the actual laughter presence. The fifth loss is the loss related to the difference between the correct label for sentiment feature (correct sentiment feature label) and the actual sentiment feature. The sixth loss is the loss related to the difference between the correct label for activity feature (correct activity feature) and the actual activity feature.
[0128] Updating a neural network means updating the learning parameters of the neural network.
[0129] The update completion determination unit 31J determines whether the conditions for stopping the learning process have been met. These conditions include, for example, that the number of learning parameter updates has reached a predetermined number, or that the amount of learning parameter updates is below a threshold.
[0130] Until the update completion determination unit 31J determines that the stop condition has been met, a series of processes are repeatedly executed, including the calculation of acoustic features by the acoustic feature calculation unit 31B, the calculation of non-acoustic features by the non-acoustic feature calculation unit 31C, the calculation of integrated features by the integrated feature calculation unit 31D, the calculation of voice presence by the voice presence calculation unit 31E, the calculation of laughter presence by the laughter presence calculation unit 31F, the calculation of emotion features by the emotion feature calculation unit 31G, the calculation of activity features by the activity feature calculation unit 31H, and the updating of the first neural network 41, the second neural network 42, the third neural network 43, the fourth neural network 44, the fifth neural network 45, and the sixth neural network 46 by the update unit 31I.
[0131] If the update completion determination unit 31J determines that the stop condition has been met, the output control unit 31K outputs the first neural network 41 as the first trained model 21, the second neural network 42 as the second trained model 22, the third neural network 43 as the third trained model 23, the fourth neural network 44 as the fourth trained model 24, the fifth neural network 45 as the fifth trained model 25, and the sixth neural network 46 as the sixth trained model 26, at the time the determination was made.
[0132] For example, the output control unit 31K outputs the first trained model 21, the second trained model 22, the third trained model 23, the fourth trained model 24, the fifth trained model 25, and the sixth trained model 26 to the estimation device 10. The estimation device 10 can then perform the above processing using these trained models learned by the learning device 30.
[0133] In other words, the learning model that integrates the third pre-trained model 23, the fourth pre-trained model 24, the fifth pre-trained model 25, and the sixth pre-trained model 26 is a learning model that includes a neural network trained to reduce the integrated losses of the third loss related to the difference between the ground truth label for speech presence and speech presence, the fourth loss related to the difference between the ground truth label for laughter presence and laughter presence, the fifth loss related to the difference between the ground truth label for emotion feature and emotion feature, and the sixth loss related to the difference between the ground truth label for activation feature and activation feature. The learning model causes the computer to function to perform calculations by the neural network and output speech presence, laughter presence, emotion feature, and activation feature when integrated features of acoustic and non-acoustic signals are input.
[0134] Next, we will describe an example of an integrated neural network configuration.
[0135] Figure 6 is an explanatory diagram illustrating an example of an integrated neural network configuration.
[0136] The integrated neural network includes a first neural network 41, a second neural network 42, an integrated feature computation module 51D, a third neural network 43, a fourth neural network 44, a fifth neural network 45, and a sixth neural network 46.
[0137] The first neural network 41 takes acoustic signal A as input and processes acoustic feature E A The second neural network 42 takes the image signal V as input and outputs the image features E V The integrated feature computation module 51D outputs the acoustic features E output from the first neural network 41. A and the image features E output from the second neural network 42 V Enter the following to integrate feature Z AV It outputs the following. The integrated feature calculation module 51D is a module that corresponds to the functions of the integrated feature calculation unit 31D.
[0138] The third neural network 43 uses the integrated feature Z output from the integrated feature computation module 51D. AV Enter the voice presence value Y^ VAD The fourth neural network 44 outputs the integrated feature Z output from the integrated feature computation module 51D. AV Enter laughter presence level Y^ LD The fifth neural network 45 outputs the integrated feature Z output from the integrated feature computation module 51D. AV Enter the sentiment feature Y^ ER The sixth neural network 46 outputs the integrated feature Z output from the integrated feature computation module 51D. AV Enter the activity feature Y^ COMM Outputs.
[0139] The first neural network 41, the second neural network 42, the third neural network 43, the fourth neural network 44, the fifth neural network 45, and the sixth neural network 46 are assumed to have initial values assigned to their learning parameters. These learning parameters include weights, biases, etc. Note that the learning parameters may also include arbitrary hyperparameters.
[0140] The first neural network 41 processes acoustic features E from acoustic signal A. A It has an encoder network architecture capable of calculating E. For example, the first neural network 41 includes four one-dimensional convolutional layers. The second neural network 42 processes image features E from the image signal V. V It has an encoder network architecture capable of calculating the following. For example, the second neural network 42 includes three 3D convolutional layers. At least one of the three 3D convolutional layers may be connected to a maximum value pooling layer and / or a global mean value pooling layer.
[0141] The third neural network 43 integrates feature Z AV From the voice presence Y^ VAD It has a detection network architecture capable of calculating Z. For example, the third neural network 43 includes two dense layers. The fourth neural network 44 has an integrated feature Z. AV Laughter presence level Y^ LD It has a detection network architecture capable of calculating [the value]. For example, the fourth neural network 44 includes two dense layers.
[0142] The fifth neural network 45 integrates feature Z AV From emotional characteristic score Y^ ERIt has a detection network architecture capable of calculating the integrated feature Z. For example, the fifth neural network 45 includes two dense layers. The sixth neural network 46 has an integrated feature Z. AV From the activity feature Y^ COMM It has a detection network architecture capable of calculating [the result]. For example, the sixth neural network 46 includes two dense layers.
[0143] All convolutional and dense layers of the first neural network 41, the second neural network 42, the third neural network 43, the fourth neural network 44, the fifth neural network 45, and the sixth neural network 46 are followed by a normalized linear function unit, and a normalized linear function is applied to the output of that layer. However, the final layers of the fifth neural network 45 and the sixth neural network 46 are followed by a sigmoid activation function unit or a softmax activation function unit instead of a normalized linear function unit, and a sigmoid function or softmax function is applied to the output of that final layer.
[0144] The integrated neural network takes an acoustic signal A and an image signal V as inputs and calculates the audio presence Y^ VAD and laughter presence level Y^ LD and emotional characteristic score Y^ ER and activity feature Y^ COMM It outputs the following.
[0145] Voice presence Y^ VAD And the correct answer is the presence level of the audio (Ground Truth) Y VAD A third loss function relating to the difference with, and the laughter presence Y^ LD The correct answer is Laughter Presence (Grand Truth) Y LD The fourth loss function relating to the difference with, and the sentiment feature Y^ ER And the correct emotional characteristic score (Grand Truth) Y ER The fifth loss function relating to the difference with, and the activation feature Y^ COMM and the ground truth activity feature YCOMM The learning parameters of the first neural network 41, the second neural network 42, the third neural network 43, the fourth neural network 44, the fifth neural network 45, and the sixth neural network 46 are updated so as to minimize the sixth loss function relating to the difference with and the integrated loss function including . In this embodiment, loss and loss function have the same meaning.
[0146] Next, an example of the learning process flow by the processing unit 31 of the learning device 30 in this embodiment will be described.
[0147] Figure 7 is a flowchart showing an example of the learning process flow executed by the processing unit 31 of the learning device 30 in this embodiment. The learning process flow will be explained using Figures 6 and 7. For the purpose of the following explanation, non-acoustic signals will be assumed to be image signals.
[0148] The learning process is performed by the processing unit 31 operating in accordance with a learning program stored in the memory unit 32, etc. In this learning process, the processing unit 31 learns the first neural network 41, the second neural network 42, the third neural network 43, the fourth neural network 44, the fifth neural network 45, and the sixth neural network 46 in parallel. More specifically, the processing unit 31 performs supervised learning on training data for an integrated neural network including the first neural network 41, the second neural network 42, the third neural network 43, the fourth neural network 44, the fifth neural network 45, and the sixth neural network 46.
[0149] The acquisition unit 31A acquires an input signal that includes an acoustic signal A and an image signal V (steps S11 and S201). The acquisition unit 31A acquires an input signal that is one training sample. The frame length of the time interval of the input signal is not particularly limited, but for example, it is about 30 seconds. The acoustic signal A and the image signal V included in the input signal are time-synchronized.
[0150] The acoustic feature calculation unit 31B uses the first neural network 41 to calculate acoustic features E from acoustic signals A included in the input signals acquired in steps S11 and S201. A The following is calculated (steps S12 and S202). The first neural network 41 used in steps S12 and S202 is assumed not to have completed training until it is judged positively in step S210, which will be described later. The acoustic signal A input to the first neural network 41 is assumed to have been converted from the time domain to the frequency domain.
[0151] The non-acoustic feature calculation unit 31C uses the second neural network 42 to calculate image features Ev, which are non-acoustic features, from the image signal V, which is a non-acoustic signal included in the input signal acquired in steps S1 and S201 (steps S13 and S203). The second neural network 42 used in steps S203 and S13 is assumed not to have completed training until it is determined to be positive in step S210, which will be described later.
[0152] The processing order of steps S12 and S202 and steps S13 and S203 is not particularly limited. For example, steps S12 and S202 may be executed after steps S13 and S203. Alternatively, steps S12 and S202 may be executed in parallel with steps S13 and S203.
[0153] The integrated feature calculation unit 31D uses the integrated feature calculation module 51D to calculate the acoustic feature E A and image feature E V Integration feature Z AVCalculate the integrated feature Z (steps S14, S204). AV This is calculated for each frame time t.
[0154] The voice presence calculation unit 31E uses the third neural network 43 to calculate the integrated feature Z AV From the voice presence Y^ VAD The following is calculated (steps S15 and S205). The third neural network 43 used in steps S15 and S205 is assumed not to have completed training until it is judged positively in step S210, which will be described later.
[0155] The laughter presence calculation unit 31F uses the fourth neural network 44 to calculate the integrated feature Z AV Laughter presence level Y^ LD The following is calculated (steps S16 and S206). The fourth neural network 44 used in steps S16 and S206 is assumed not to have completed training until it is judged positively in step S210, which will be described later.
[0156] The emotion feature calculation unit 215 uses the third neural network 43 to calculate the integrated feature Z AV From emotional characteristic score Y^ ER The following is calculated (steps S17 and S207). The third neural network 43 used in steps S17 and S207 is assumed not to have completed training until it is judged positively in step S210, which will be described later.
[0157] The activity feature calculation unit 216 uses the fourth neural network 44 to calculate the integrated feature Z AV From the activity feature Y^ COMM The following is calculated (steps S18 and S208). The fourth neural network 44 used in steps S18 and S208 is assumed not to have completed training until it is judged positively in step S210, which will be described later.
[0158] The processing order of steps S205 to S208 and steps S15 to S18 is not particularly limited. The processing may be executed in the order of steps S208 to S205 and steps S18 to S15. Alternatively, the processing of steps S205 to S208 and steps S15 to S18 may be executed in parallel.
[0159] The update unit 31I updates the first neural network 41, the second neural network 42, the third neural network 43, the fourth neural network 44, the fifth neural network 45, and the sixth neural network 46 using an integrated loss that includes the third loss, the fourth loss, the fifth loss, and the sixth loss (steps S19, S209).
[0160] As shown in equation (3) below, the integrated loss L at frame time t and coordinate d TOTAL This is the third loss L VAD and the fourth loss L LD and the fifth loss L ER and the sixth loss L COMM It is determined by the sum of the following.
[0161]
number
[0162] Third loss L VAD This is a loss used to penalize the difference between the correct speech presence label and the actual speech presence regarding the presence of speech. As shown in equation (4) below, the third loss L VAD This is defined by Binary Cross Entropy (BCE). Binary Cross Entropy is defined as the presence label Y of the correct audio. VAD and the voice presence Y^ VAD It is used as a measure to evaluate the difference between [the two /
[0163] The fourth loss L LDis a loss for imposing a penalty on the difference between the ground-truth laughter presence label related to laughter presence and the laughter presence. As shown in the following formula (5), the fourth loss L LD is defined by binary cross entropy. Binary cross entropy is the ground-truth laughter presence label Y LD and the laughter presence Y^ LD is used as a measure for evaluating the difference between
[0164] The fifth loss L ER is a loss for imposing a penalty on the difference between the ground-truth emotional feature label related to emotional features and the emotional feature degree. As shown in the following formula (6), the fifth loss L ER is defined by SoftMax Cross Entropy (SCE). SoftMax cross entropy is the ground-truth emotional feature label Y ER and the emotional feature degree Y^ ER is used as a measure for evaluating the difference between
[0165] The sixth loss L COMM is a loss for imposing a penalty on the difference between the ground-truth activation feature label related to activation features and the activation feature degree. As shown in the following formula (7), the sixth loss L COMM is defined by Mean Squared Error (MSE). Mean squared error is the ground-truth activation feature label Y COMM and the activation feature degree Y^ COMM is used as a measure for evaluating the difference between
[0166]
Numerical Formula
[0167] The updating unit 31I, according to any optimization method, integrates the loss L TOTALThe learning parameters of the first neural network 41, the second neural network 42, the third neural network 43, the fourth neural network 44, the fifth neural network 45, and the sixth neural network 46 are updated to minimize the value. This results in the ground truth speech presence label Y. VAD and the voice presence Y^ VAD The difference between and the correct laughter presence label Y LD and laughter presence level Y^ LD The difference between and the correct sentiment feature label Y ER and emotional characteristic score Y^ ER The difference between and the correct activity feature label Y COMM and activity feature Y^ COMM The learning parameters of the first neural network 41, the second neural network 42, the third neural network 43, the fourth neural network 44, the fifth neural network 45, and the sixth neural network 46 are updated to minimize the overall difference between them. Any optimization method such as stochastic gradient descent or Adam (adaptive moment estimation) can be used.
[0168] The update completion determination unit 31J determines whether or not the stop conditions are met (steps S210, S20). The stop conditions may be set to, for example, when the number of learning parameter updates reaches a predetermined number or when the amount of learning parameter updates is less than a threshold. If it is determined that the stop conditions are not met (step S210: No), the process returns to steps S201 and S11.
[0169] The update completion determination unit 31J continues to make negative judgments until it determines that the stop condition has been met (step S210: Yes). For this reason, the series of processes in steps S11, S201 to S19, and S209 are repeatedly executed until it is determined that the stop condition has been met.
[0170] Furthermore, the sequence of processes from steps S11 to S19 and steps S201 to S209 may be repeatedly executed for a single training sample (batch learning). Alternatively, the sequence of processes from steps S11 to S19 and steps S201 to S209 may be repeated for multiple training samples (mini-batch learning).
[0171] Then, if it is determined in step S210 that the stop condition is met (step S10: Yes), the update termination determination unit 31J outputs the first neural network 41 as the first trained model 21, the second neural network 42 as the second trained model 22, the third neural network 43 as the third trained model 23, the fourth neural network 44 as the fourth trained model 24, the fifth neural network 45 as the fifth trained model 25, and the sixth neural network 46 as the sixth trained model 26 (steps S21, S211).
[0172] As a result of the processing in steps S21 and S210, the first trained model 21, the second trained model 22, the third trained model 23, the fourth trained model 24, the fifth trained model 25, and the sixth trained model 26 are transmitted to the estimation device 10 via the communication unit 34, etc., and stored in the storage unit 12 of the estimation device 10. The update completion determination unit 31J may output an integrated neural network including the first trained model 21, the second trained model 22, the third trained model 23, the fourth trained model 24, the fifth trained model 25, the sixth trained model 26, and the integrated feature calculation module 51D. Then, this routine is terminated.
[0173] The learning process by the processing unit 31 is completed by the processing of steps S11 to S21 and steps S201 to S211.
[0174] As described above, the estimation device 10 estimates the communication activity level from acoustic signals and image signals using the first trained model 21, the second trained model 22, the third trained model 23, the fourth trained model 24, and the sixth trained model 26 generated by the learning device 30. The estimation device 10 may also estimate the communication activity level by inputting the acoustic signals and image signals into the integrated neural network described above.
[0175] As described above, the learning device 30 of this embodiment includes an acquisition unit 31A, an acoustic feature calculation unit 31B, a non-acoustic feature calculation unit 31C, an integrated feature calculation unit 31D, a speech presence calculation unit 31E, a laughter presence calculation unit 31F, an emotion feature calculation unit 31G, an activity feature calculation unit 31H, and an update unit 31I. The acquisition unit 31A acquires acoustic signals and non-acoustic signals related to the speech source. The acoustic feature calculation unit 31B calculates acoustic features from the acoustic signals using a first neural network 41. The non-acoustic feature calculation unit 31C calculates non-acoustic features from the non-acoustic signals using a second neural network 42. The integrated feature calculation unit 31D calculates integrated features based on the acoustic and non-acoustic features. The speech presence calculation unit 31E calculates speech presence from integrated features using a third neural network 43. The laughter presence calculation unit 31F calculates the laughter presence from integrated features using the fourth neural network 44. The emotion feature calculation unit 31G calculates the emotion feature from integrated features using the fifth neural network 45. The activity feature calculation unit 31H calculates the activity feature from integrated features using the sixth neural network 46. The update unit 31I updates the first neural network 41, the second neural network 42, the third neural network 43, the fourth neural network 44, the fifth neural network 45, and the sixth neural network 46 by learning to reduce the integrated losses of the third loss related to the difference between the correct label for voice presence and voice presence, the fourth loss related to the difference between the correct label for laughter presence and laughter presence, the fifth loss related to the difference between the correct label for emotion feature and emotion feature, and the sixth loss related to the difference between the correct label for activity feature and activity feature.
[0176] Therefore, in the learning device 30 of this embodiment, multimodal learning of acoustic signals and non-acoustic signals, and multitask learning of speech presence estimation tasks, laughter presence estimation tasks, emotion feature estimation tasks and activity feature estimation tasks, makes it possible to generate an integrated neural network of a first trained model 21, a second trained model 22, a third trained model 23, a fourth trained model 24, a fifth trained model 25, and a sixth trained model 26, which are highly accurate communication activity estimation models.
[0177] In the estimation device 10, the activity feature calculated using the sixth trained model 26 can be made clearer and more accurate by combining the speech features calculated using the first trained model 21, the non-speech features calculated using the second trained model 22, the speech presence calculated using the third trained model 23, the laughter presence calculated using the fourth trained model 24, and the emotion feature score calculated using the fifth trained model 25. As a result, it becomes possible to detect communication activity even when it is not possible to collect human voices at close range or when it is not possible to identify a person's face from an image. In other words, the learning device 30 of this embodiment can learn a learning model that is capable of improving the accuracy of estimating communication activity.
[0178] (Examples) The performance of the integrated neural network according to this embodiment (hereinafter referred to as Configuration Example 3 of the present invention), which includes the first neural network 41, the second neural network 42, the third neural network 43, the fourth neural network 44, the fifth neural network 45, and the sixth neural network 46 generated according to the above embodiment, was verified using simulated dialogue video data.
[0179] The simulated dialogue video data was collected from one-on-one simulated dialogues held in an office conference room using a ceiling-mounted 360-degree omnidirectional fisheye camera. A total of 12 pairs of subjects, each consisting of two people who consented to data collection, were filmed for 10 minutes for each of the three dialogue tasks (Task B-1, Task B-2, Task C) (total 6 hours = 12 pairs × 3 tasks × 10 minutes). In anticipation of multimodal processing of audio and image information, the data was re-encoded as a monaural audio source with a resolution of 960 × 960 pixels, a frame rate of 25, and a sampling rate of 16 kHz.
[0180] The topics of the simulated dialogues assigned to each dialogue task were as follows: Task B-1 simulated a very active communication situation from the participants by having them freely converse about what they wanted to do after the end of the COVID-19 pandemic. Task B-2 simulated a moderately active communication situation from the participants by having them freely converse about plans to invite and entertain friends or guests from overseas to Japan. Task C simulated a moderately inactive communication situation from the participants by having them jointly discuss test questions from the SPI (Synthetic Personality Inventory) aptitude test.
[0181] The simulated dialogue video data was allocated as training, validation, and test datasets, respectively, and used for deep learning and accuracy validation in Configuration Example 3 of the present invention. The split ratio was 7:1:2, and was adjusted to ensure that features were as unbiased as possible across datasets based on the ground truth label distribution of communication activity. The validation dataset was used to identify hyperparameters suitable for each model and experimental condition.
[0182] Figure 8 is a table showing the verification results using a test dataset of simulated dialogue video data. Verification results are shown for each model. Configuration Example 3 of the present invention significantly surpasses the accuracy of the prior art configuration examples under all experimental conditions, showing a relative improvement of up to approximately 1.32 in average error value and approximately 26.37 percentage points in estimation processing accuracy.
[0183] The above embodiment is merely an example, and various modifications are possible. For example, although the acoustic signal is described as a speech signal, it may also be a vocal cord signal or vocal tract signal obtained by decomposing the speech signal. Also, although the acoustic signal is described as waveform data of sound pressure values in the time domain or frequency domain, it may also be data obtained by converting said waveform data into any space.
[0184] Thus, according to the above embodiments, it becomes possible to detect communication activity even when it is not possible to collect human voices at close range or when it is not possible to identify a person's face from an image.
[0185] Next, an example of the hardware configuration of the estimation device 10 and learning device 30 of the above embodiment will be described.
[0186] Figure 9 is a hardware configuration diagram of an example of the estimation device 10 and learning device 30 of the above embodiment.
[0187] The estimation device 10 and learning device 30 of the above embodiment include a control device such as a CPU (Central Processing Unit) 90D, a storage device such as a ROM (Read Only Memory) 90E, a RAM (Random Access Memory) 90F, and an HDD (Hard Disk Drive) 90G, an I / F unit 90B which is an interface to various devices, an output unit 90A which outputs various information such as output information, an input unit 90C which accepts user operations, and a bus 90H which connects each unit, and thus has a hardware configuration that uses a normal computer.
[0188] In the estimation device 10 and learning device 30 of the above embodiment, each of the above parts is realized on a computer by the CPU 90D reading an information processing program from ROM 90E onto RAM 90F and executing it.
[0189] Furthermore, the information processing program for executing each of the above processes performed by the estimation device 10 and learning device 30 in the above embodiment may be stored in the HDD 90G. Alternatively, the information processing program for executing each of the above processes performed by the estimation device 10 and learning device 30 in the above embodiment may be pre-installed and provided in the ROM 90E.
[0190] Furthermore, the information processing program for executing the above-mentioned processes performed by the estimation device 10 and learning device 30 of the above embodiment may be provided as a computer program product by being stored in an installable or executable file format on a computer-readable storage medium such as a CD-ROM, CD-R, memory card, DVD (Digital Versatile Disc), or flexible disk (FD). Alternatively, the program for executing the above-mentioned processes performed by the estimation device 10 and learning device 30 of the above embodiment may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. Alternatively, the program for executing the above-mentioned processes performed by the estimation device 10 and learning device 30 of the above embodiment may be provided or distributed via a network such as the Internet.
[0191] Although embodiments of the present invention have been described above, these embodiments are presented as examples only and are not intended to limit the scope of the invention. This novel embodiment can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. This embodiment and its variations are included in the scope and spirit of the invention, as well as in the claims of the invention and its equivalents. [Explanation of Symbols]
[0192] 10 Estimation device 11A Acquisition Department 11B Acoustic Feature Calculation Unit 11C Non-acoustic feature calculation unit 11D Integrated Feature Calculation Unit 11E Voice Presence Calculation Unit 11F Laughter Presence Calculation Unit 11G Emotional Feature Calculation Unit 11H Activity characteristic calculation unit 11I Activity Estimator 11J Situation Information Generation Unit 11K Output Control Unit 21. First pre-trained model 22. Second pre-trained model 23. Third pre-trained model 24. Fourth pre-trained model 25. The fifth pre-trained model 26. The sixth pre-trained model 30 Learning device 31A Acquisition Department 31B Acoustic Feature Calculation Unit 31C Non-acoustic feature calculation unit 31D Integrated Feature Calculation Unit 31E Voice Presence Calculation Unit 31F Laughter Presence Calculation Unit 31G Emotional Feature Calculation Unit 31H Activity characteristic calculation unit 31I Update Section 41 The First Neural Network 42. The Second Neural Network 43. The Third Neural Network 44. The Fourth Neural Network 45. The Fifth Neural Network 46. The Sixth Neural Network
Claims
1. An acquisition step of acquiring an acoustic signal relating to a sound source and a non-acoustic signal which is an image signal that is temporally synchronized with the acoustic signal and correlates with the morphology of the facial region or the movement of a body part when the speaker is pronouncing, A step of calculating acoustic features based on the aforementioned acoustic signal, A non-acoustic feature calculation step that calculates non-acoustic features that correlate with the speaker's communication characteristics in response to non-acoustic changes represented by the aforementioned non-acoustic signal, A speech presence calculation step that calculates the speech presence based on the acoustic characteristics and the non-acoustic characteristics, An activity feature calculation step of calculating the activity feature score based on the acoustic features and the non-acoustic features, An activity estimation step in which communication activity is estimated based on a comparison with the aforementioned activity characteristic evaluation scale, A situation information generation step that generates situation information representing the communication situation based on the aforementioned voice presence level, An output control step that outputs status information representing the communication activity level and communication status, A pre-programmed command to get a computer to execute something.
2. The acoustic feature calculation step described above is: The acoustic characteristics corresponding to the change in sound represented by the aforementioned acoustic signal are calculated. The estimation program according to claim 1.
3. The aforementioned non-acoustic changes are represented by the movement of at least a part of the speaker's body. The estimation program according to claim 1.
4. The aforementioned activity characteristic calculation step is: The activity characteristic degree is calculated, which represents the activity characteristic derived from the acoustic signal represented by the acoustic characteristics and the non-acoustic signal represented by the non-acoustic characteristics. The estimation program according to claim 1.
5. A step of calculating the presence of laughter based on the acoustic characteristics and the non-acoustic characteristics, A step of calculating emotional characteristic scores based on the acoustic characteristics and non-acoustic characteristics, It further includes, The aforementioned status information generation step is: The situational information is generated based on the presence level of the voice, the presence level of the laughter, and the emotional characteristic level. The estimation program according to claim 1.
6. The aforementioned step of calculating the presence of sound is: The presence of sound is calculated according to the sound characteristics derived from the acoustic signal represented by the acoustic characteristics and the non-acoustic signal represented by the non-acoustic characteristics. The estimation program according to claim 1.
7. The aforementioned step of calculating the presence of laughter is: The degree of laughter presence is calculated according to the characteristics of laughter derived from the acoustic signal represented by the acoustic characteristics and the non-acoustic signal represented by the non-acoustic characteristics. The estimation program according to claim 5.
8. The aforementioned step of calculating emotional characteristics is: The emotional characteristic score is calculated according to the emotional characteristics derived from the acoustic signal represented by the acoustic characteristics and the non-acoustic signal represented by the non-acoustic characteristics. The estimation program according to claim 5.
9. The aforementioned status information generation step is: The amount of speech produced by the sound source, as represented by the aforementioned sound presence, The number of times laughter occurs, which is represented by the aforementioned laughter presence level, The rate of change in emotion, expressed by the aforementioned emotional characteristic score, The generation of the aforementioned situation information includes, The estimation program according to claim 5.
10. The method further includes an integrated feature calculation step of calculating an integrated feature obtained by integrating the acoustic features and the non-acoustic features, The voice presence calculation step involves calculating the voice presence based on the integrated features, The laughter presence calculation step involves calculating the laughter presence based on the integrated features, The emotional characteristic calculation step involves calculating the emotional characteristic based on the integrated characteristics, The activity characteristic calculation step involves calculating the activity characteristic based on the integrated characteristics. The estimation program according to claim 5.
11. The acoustic feature calculation step involves calculating the acoustic features from the acoustic signal using the first trained model, The non-acoustic feature calculation step involves calculating the non-acoustic features from the non-acoustic signal using a second trained model. The activation feature calculation step involves calculating the activation feature from the integrated features of the acoustic signal and the non-acoustic signal using a sixth trained model. The estimation program according to claim 1.
12. The aforementioned speech presence calculation step involves calculating the speech presence from the integrated features of the acoustic signal and the non-acoustic signal using a third trained model, The laughter presence calculation step involves calculating the laughter presence from the integrated features using the fourth trained model, The emotion feature calculation step involves calculating the emotion feature from the integrated features using a fifth trained model. The estimation program according to claim 5.
13. The estimation program according to claim 1, wherein the non-acoustic signal is an image signal that is time-synchronized with the acoustic signal.
14. The aforementioned non-acoustic features have higher values the more the speaker's body parts move, and lower values the less the speaker's body parts move. The estimation program according to claim 1.
15. An acquisition unit that acquires an acoustic signal relating to a sound source and a non-acoustic signal which is an image signal that is temporally synchronized with the acoustic signal and correlates with the shape of the facial region or the movement of a body part when the speaker is speaking. An acoustic feature calculation unit that calculates acoustic features based on the aforementioned acoustic signal, A non-acoustic feature calculation unit calculates non-acoustic features that correlate with the speaker's communication characteristics in response to non-acoustic changes represented by the aforementioned non-acoustic signal, A voice presence calculation unit that calculates the voice presence based on the acoustic characteristics and the non-acoustic characteristics, An activity feature calculation unit that calculates the activity feature score based on the acoustic features and the non-acoustic features, An activity estimation unit that estimates the communication activity level based on a comparison with the aforementioned activity characteristic evaluation scale, A situation information generation unit that generates situation information representing the communication situation based on the aforementioned voice presence level, An output control unit that outputs status information representing the communication activity level and communication status, An estimation device equipped with the following features.
16. Executed by a computer, An acquisition step of acquiring an acoustic signal relating to a sound source and a non-acoustic signal which is an image signal that is temporally synchronized with the acoustic signal and correlates with the morphology of the facial region or the movement of a body part when the speaker is pronouncing, A step of calculating acoustic features based on the aforementioned acoustic signal, A non-acoustic feature calculation step that calculates non-acoustic features that correlate with the speaker's communication characteristics in response to non-acoustic changes represented by the aforementioned non-acoustic signal, A speech presence calculation step that calculates the speech presence based on the acoustic characteristics and the non-acoustic characteristics, An activity feature calculation step of calculating the activity feature score based on the acoustic features and the non-acoustic features, An activity estimation step in which communication activity is estimated based on a comparison with the aforementioned activity characteristic evaluation scale, A situation information generation step that generates situation information representing the communication situation based on the aforementioned voice presence level, An output control step that outputs status information representing the communication activity level and communication status, An estimation method that includes [this].
Citation Information
Patent Citations
Voice processor, moving picture processor, voice and moving picture processor, and recording medium with voice and moving picture processing program recorded
JP2002006874A
Conference visualizing system and method, conference summary processing server
JP2008262046A
Communication calculation device, function calculation device for meeting, and function calculation method and program for meeting
JP2009163431A
Congress analysis device, method and program
JP2016012216A
Conference support system
JP2017112545A