Device and method
The apparatus and method enhance emotion estimation by generating evidence information based on speech features and reference amounts, addressing the mismatch between subjective and estimated emotions, thereby improving user satisfaction with the results.
Patent Information
- Application Number
- PCT/JP2024/018725
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-21
- Publication Date
- 2025-11-27
Smart Images

Figure JP2024018725_27112025_PF_FP_ABST
Abstract
Description
Apparatus and method
[0001] One aspect of the present disclosure relates to an apparatus and a method.
[0002] Patent Literature 1 discloses a technology for generating an emotion estimator that uses speech data classified into patterns of change in features corresponding to each emotion as training data to estimate the emotion of a speaker when he or she speaks.
[0003] JP 2020-154332 A
[0004] It is known that a discrepancy occurs between a user's subjective non-verbal information, such as emotions, and non-verbal information estimated based on objective data obtained from the user's speech. Even if a user understands the discrepancy, they may not be convinced by the estimated result if it differs from the subjective non-verbal information. Therefore, there is a need for a technology that outputs more appropriate estimation grounds.
[0005] Therefore, an object of the present disclosure is to provide an apparatus and method that can output more appropriate estimation grounds.
[0006] The device disclosed herein includes an acquisition unit that acquires features related to speech, an estimation unit that estimates non-verbal information of a speaker uttering a speech based on predetermined reference amounts related to the features and non-verbal information, a generation unit that generates evidence information related to the basis for the estimation of the non-verbal information by the estimation unit based on the features, the reference amounts, and priorities corresponding to the features, and an output unit that outputs the evidence information.
[0007] According to one aspect of the present disclosure, it is possible to output more appropriate estimation grounds.
[0008] FIG. 1 is a block diagram showing a configuration of a processing system including an apparatus according to a first embodiment of the present disclosure. FIG. 2 is a diagram showing an example of the configuration of information related to emotions, which is an example of non-verbal information. FIG. 3 is a flowchart showing the procedure of an example of a processing method performed by the processing system according to the first embodiment of the present disclosure. FIG. 4 is a flowchart showing an example of a step of generating ground information in the processing method performed by the processing system according to the first embodiment of the present disclosure. FIG. 5 is a block diagram showing a configuration of a processing system including an apparatus according to a second embodiment of the present disclosure. FIG. 6 is a flowchart showing the procedure of an example of a processing method performed by the processing system according to the second embodiment of the present disclosure. FIG. 7 is a block diagram showing a configuration of a processing system including an apparatus according to a third embodiment of the present disclosure. FIG. 8 is a flowchart showing the procedure of an example of a processing method performed by the processing system according to the third embodiment of the present disclosure. FIG. 9 is a flowchart showing an example of a step of generating ground information in the processing method performed by the processing system according to the third embodiment of the present disclosure. FIG. 10 is a diagram showing an example of a hardware configuration of an apparatus according to an embodiment of the present disclosure.
[0009] The present disclosure will be described with reference to the accompanying drawings. Whenever possible, the same parts are designated by the same reference numerals and redundant description will be omitted.
[0010] [First Embodiment] FIG. 1 is a block diagram showing the configuration of a processing system including an apparatus according to a first embodiment of the present disclosure. The processing system 1 shown in FIG. 1 includes a terminal 5, a non-verbal information database 7, and an apparatus 10, all of which are configured to communicate with each other via a network including a wireless communication network and a fixed communication network. The processing system 1 is used to present the basis for estimating non-verbal information of a speaker using the terminal 5, which is an example of a voice terminal. The speaker may be, for example, a user who wishes to objectively view their own non-verbal information. For example, the speaker can appropriately perform stress self-care by using the processing system to detect emotional changes that the speaker himself or herself is not aware of. For example, the speaker can check whether he or she is speaking in an unintentionally overbearing voice and reconsider how he or she interacts with others.
[0011] The device 10 estimates a speaker's non-verbal information based on features related to the speaker's voice and predetermined reference quantities related to non-verbal information. Non-verbal information is information indicating at least one of emotions and mental states. Emotions include, for example, at least one of joy, anger, sadness, excitement, positive intention, negative intention, affirmative intention, and negative intention. Mental states include at least one of psychological stress, depressive reaction, impatience, shock, and irritation. Non-verbal information estimated in this embodiment includes, for example, joy, anger, and sadness. The device 10 of this embodiment receives the speaker's voice from the terminal 5 and outputs to the terminal 5 the speaker's non-verbal information estimated from the voice and evidence information that is the basis for estimating the non-verbal information. The device 10 does not necessarily have to output the speaker's non-verbal information. Each component will be described in detail below.
[0012] The terminal 5 is an audio terminal used by a speaker to input speech. The terminal 5 may also be an audio terminal used by a speaker to converse with another speaker. The terminal 5 is, for example, a personal computer, a smartphone, a tablet terminal, a feature phone, a server device, a game console, or other device. Note that while FIG. 1 illustrates only one terminal as each of the terminals 5, the processing system 1 may include any number of terminals 5, two or more.
[0013] The terminal 5 acquires the speaker's voice. The terminal 5 outputs the acquired voice to the device 10. The terminal 5 accepts an input operation from the speaker regarding the priority for an index indicating voice quality. The terminal 5 outputs the input index indicating voice quality to the device 10. The index indicating voice quality and the priority will be described in detail later.
[0014] Furthermore, the terminal 5 acquires the non-language information estimated by the device 10 and basis information regarding the basis for estimating the non-language information. Note that the terminal 5 may acquire the non-language information and the basis information from the device 10 every time speech is input.
[0015] The non-language information database 7 stores predetermined reference quantities related to non-language information. The reference quantities are data that can be compared with the feature quantities described below. The reference quantities may be numerical data (vector data). The reference quantities are, for example, data that express at least one of each emotion and each state of mind included in the non-language information. The reference quantities may also be statistical quantities of multiple speakers. The statistical quantities are, for example, the average value, median value, etc. of the feature quantities in the non-language information of multiple speakers. The feature quantities will be described in detail below.
[0016] FIG. 2 is a diagram showing an example of the configuration of reference quantities related to non-verbal information. In the example shown in FIG. 2, the reference quantities stored in the non-verbal information database 7 are represented by a table showing the correspondence between non-verbal information and voice characteristics (indicators of voice quality). The reference quantities may be conditions used when determining non-verbal information based on voice. The reference quantities may indicate a range of values corresponding to indices of voice quality. Emotions such as "joy," "anger," and "sadness" are correlated with each indices of voice quality. In the example shown in FIG. 2, reference quantities for "volume," "pitch," and "clarity" are shown as examples of indices of voice quality. The reference quantities are numerical data showing the correspondence between emotions (an example of non-verbal information) and indices of voice quality. The numerical values in the table in FIG. 2 indicate boundary values of the reference quantities (%) when the maximum value of each indices is 100 and the minimum value is 0.
[0017] Note that the indices indicating voice quality do not necessarily have to include any of the indices of volume, pitch, and clarity. The indices indicating voice quality may include, for example, at least one of speaking rate, timbre, resonance, warmth, and nasality. Furthermore, the reference amount of each non-verbal information may include a value for at least one indices. The reference amount of each indices indicating voice quality may include a value for at least one non-verbal information.
[0018] The device 10 shown in Fig. 1 is configured to include, as functional components, an acquisition unit 20, an estimation unit 30, a generation unit 40, and an output unit 50. The device 10 estimates non-verbal information based on speech acquired from the terminal 5. The device 10 generates basis information regarding the basis for estimating the non-verbal information. The device 10 outputs at least the basis information to the terminal 5. The functions of each functional unit of the device 10 will be described in detail below.
[0019] The acquisition unit 20 acquires the speaker's voice and priority, and acquires features based on the voice. The acquisition unit 20 has a voice acquisition unit 21, a detection unit 22, and a priority acquisition unit 23. The voice acquisition unit 21 acquires voice input to the terminal 5 by the speaker from the terminal 5. The voice acquisition unit 21 acquires the voice by accepting the voice transmitted from the terminal 5. Note that the voice acquisition unit 21 may acquire the voice from a database (not shown) in which voices acquired by the terminal 5 are stored.
[0020] The detection unit 22 detects a feature quantity related to the acquired speech based on the speech. The feature quantity is a value corresponding to an index indicating the voice quality indicated by the speech. The feature quantity may be numerical data (vector data). The detection unit 22 analyzes, as a feature quantity, a value corresponding to at least one index of the volume, pitch, and clarity indicated by the speech, as an example of an index indicating the voice quality. The detection unit 22 of this embodiment analyzes, for example, the volume, pitch, and clarity indicated by the speech.
[0021] For example, if a speaker speaks in a loud, low voice and in a rapid-fire manner, the detection unit 22 analyzes the characteristics of the speaker's voice as 70, 20, and 40 for "volume," "pitch," and "clarity," respectively.
[0022] The priority acquisition unit 23 acquires a priority corresponding to the feature. The priority indicates the degree of importance to be attached to each index indicating the voice quality corresponding to the feature when estimating non-linguistic information. The priority is a value that weights the numerical value of each index included in the feature. The priority is, for example, the degree to which the speaker (user) predicts it, and may be set higher in advance the higher the degree to which the device (first machine learning model 31) attaches importance when making estimation. The priority may be set higher, for example, the higher the degree of ease of understanding for the speaker (user). The priority acquisition unit 23 acquires, from the terminal 5, the priority input to the terminal 5 by the speaker. The priority acquisition unit 23 acquires the priority by accepting the priority transmitted from the terminal 5.
[0023] Note that the priority acquisition unit 23 does not necessarily have to acquire priorities input by the speaker whose non-language information is to be estimated. For example, the priority acquisition unit 23 may acquire the priorities from a database (not shown) in which predetermined priorities are stored. For example, the priority acquisition unit 23 may acquire, as priorities, priority statistics obtained from multiple users who may use the device 10 to estimate non-language information. The priority statistics may be the average or median of multiple priorities obtained from multiple users. In this case, the priority acquisition unit 23 may calculate the priority statistics based on priorities acquired from multiple terminals. The priority acquisition unit 23 may determine priorities based on features that are likely to be associated with non-language information that have been previously examined.
[0024] The estimation unit 30 estimates non-verbal information of a speaker uttering a voice based on predetermined reference amounts for the feature amounts and non-verbal information. The estimation unit 30 acquires data indicating reference amounts for the non-verbal information from the non-verbal information database 7. The estimation unit 30 estimates the emotion indicated by the voice as the speaker's non-verbal information based on the feature amounts indicated by the voice and the reference amounts acquired from the non-verbal information database 7.
[0025] The estimation unit 30 estimates non-language information using, for example, a first machine learning model 31. The estimation unit 30 has learned a reference amount in advance, and estimates the non-language information by inputting a feature into the first machine learning model 31, which outputs non-language information corresponding to speech in response to the input feature. The estimation unit 30 inputs the feature into a machine learning model within the device 10 or external to the device 10, and causes the non-language information to be output. The estimation unit 30 of this embodiment causes the non-language information to be output using the first machine learning model 31 within the device 10.
[0026] The first machine learning model 31 may include, for example, at least one of a generative artificial intelligence (AI) model, a discriminative / determinative AI model, and a model that combines these. Hereinafter, an example in which a generative AI model is used as the first machine learning model 31 will be described.
[0027] A generative AI model is a model that can generate content in response to a prompt containing input information, according to the instructions and output format indicated by the prompt, and return the content as response information. The prompt can also include input information, in which case the generative AI model generates response information targeted at the input information. The generative AI model may be, for example, an interactive AI that includes a large language model (LLM) and a user interface (UI) for interacting with the user, enabling text or voice chat with the user. Examples of such generative AI include ChatGPT, GPT (registered trademark)-3.5, GPT-4V, PaLM2, etc.
[0028] The first machine learning model 31 may be stored within the device 10, or may be stored in another device connected to the device 10 via a network and configured to enable information exchange with a user via the device 10. Note that while only one device 10 is illustrated in FIG. 1 , the processing system 1 may include multiple devices 10 and a server device (not shown). The second machine learning model 41 described below also has the same configuration and function as the first machine learning model 31. The first machine learning model 31 and the second machine learning model 41 described below may be a single common machine learning model, each may be a single machine learning model, or each may be configured by combining multiple machine learning models.
[0029] The first machine learning model 31 outputs an emotion estimated to be indicated by the speech based on the input feature quantities as an example of non-linguistic information. The estimation unit 30 inputs the feature quantities acquired by the detection unit 22 to the first machine learning model 31 as input information. For example, the estimation unit 30 inputs feature quantities of 70, 20, and 40 for "volume," "pitch," and "clarity" as input information to the first machine learning model 31. The first machine learning model 31 estimates that the emotion of "anger" is the speaker's non-linguistic information based on pre-trained reference quantities and the input feature quantities, and outputs the emotion. The estimation unit 30 acquires the emotion output by the first machine learning model 31 as non-linguistic information. The first machine learning model 31 may output the likelihood (probability) of the estimated emotion.
[0030] It should be noted that the detection unit 22 of the acquisition unit 20 described above does not need to analyze the features. In this case, for example, the detection unit 22 may input a voice to the first machine learning model 31 and cause it to detect an index indicating the voice quality of the voice. The first machine learning model 31 may directly output, as non-verbal information, an emotion estimated to be expressed by the voice based on the voice input from the detection unit 22. The detection unit 22 of the acquisition unit 20 acquires the features output from the first machine learning model 31. It should be noted that the detection unit 22 may acquire features stored in a database or a server device (not shown).
[0031] Furthermore, the estimation unit 30 does not have to estimate the non-language information using, for example, the first machine learning model 31. The estimation unit 30 may compare the feature amount with a reference amount and estimate the non-language information on a rule-based basis. The estimation unit 30 searches a table showing the correspondence between non-language information and speech features obtained from the non-language information database 7 to determine which emotion corresponds when "volume," "pitch," and "clarity" are 70, 20, and 40. The estimation unit 30 may estimate that the speaker's non-language information includes the emotion of anger based on the table.
[0032] The generation unit 40 generates basis information regarding the basis for estimating the non-language information in the estimation unit 30 based on the feature amount, the reference amount, and the priority. The basis information includes information regarding the feature amount with high priority and the estimated non-language information. The basis information includes at least one of comparison information regarding the result of comparing the feature amount with the reference amount and impression information regarding an impression estimated based on the comparison information and the non-language information.
[0033] The comparison information may include a result of comparison between a value of an index indicating a voice quality with a high priority among the feature quantities and a value of a reference quantity for the index. The comparison information may include a result of comparison between a value of an index indicating at least one voice quality among the feature quantities corresponding to a priority higher than a predetermined threshold and a value of a reference quantity for the at least one index associated with the estimated non-language information. The generation unit 40 may generate, for example, a value of an index indicating a voice quality with a high priority among the feature quantities and a value of an index indicating a voice quality with a high priority among the reference quantities using a machine learning model.
[0034] The generation unit 40 generates, as the comparison information, a result of comparing a feature quantity indicating an index of a voice quality with a high priority and a reference quantity. Furthermore, the generation unit 40 of this embodiment generates impression information based on the non-language information and the comparison information. The generation unit 40 may generate the comparison information using, for example, a machine learning model. The generation unit 40 may also generate basis information including the comparison information and the impression information using, for example, the machine learning model. The generation unit 40 may also generate notification information including non-language information and basis information including the comparison information and the impression information using, for example, the machine learning model.
[0035] The generation unit 40 inputs non-verbal information, feature quantities, reference quantities, and priorities into a machine learning model within the device 10 or external to the device 10, and causes the model to output comparison information and impression information. The generation unit 40 of this embodiment outputs comparison information using a second machine learning model 41 within the device 10. The generation unit 40 estimates notification information by inputting the extracted feature quantities and reference quantities into the second machine learning model 41, which outputs notification information by inputting non-verbal information, feature quantities, and reference quantities. Note that the second machine learning model 41 may output at least one of comparison information and impression information.
[0036] The second machine learning model 41 may include, for example, at least one of a generative AI model, a discriminative / determinative AI model, and a model that combines these. Below, an example will be described in which a generative AI model is used as the second machine learning model 41. In this embodiment, the device 10 is capable of providing a content provision function using the second machine learning model 41 as a large-scale language model.
[0037] For example, a case will be described in which the indices indicating voice quality corresponding to a priority higher than a predetermined threshold are "volume" and "pitch." Based on the input features and priorities, the second machine learning model 41 outputs a value for the indices indicating voice quality with high priority among the features. Based on the priorities, the second machine learning model 41 extracts 70 and 20, which indicate "volume" and "pitch," as values for the indices indicating voice quality with high priority from the features of 70, 20, and 40 for "volume," "pitch," and "clarity."
[0038] The second machine learning model 41 outputs a value for an index indicating a voice quality with a high priority from the reference amounts based on the input non-verbal information, reference amounts, and priorities. The second machine learning model 41 extracts "50 or more" and "less than 30" indicating "volume" and "pitch" as values for an index indicating a voice quality with a high priority from the reference amounts corresponding to the emotion (non-verbal information) estimated by the estimation unit 30.
[0039] The second machine learning model 41 outputs comparison information based on the input feature quantities, reference quantities, and priorities. The second machine learning model 41 outputs comparison information based on the feature quantities and reference quantities extracted for the indices indicating voice quality. The second machine learning model 41 generates, as the comparison information, a result of comparing the feature quantities indicating "volume" and "pitch" with the reference quantities. For example, when a feature quantity of 70 for "volume" and a reference quantity of "50 or more" are input, the second machine learning model 41 outputs comparison information indicating "loud voice" as the volume of the speaker's voice is greater than the boundary value of the reference quantity of "50" (corresponding to the reference quantity). Furthermore, when a feature quantity of 20 for "pitch" and a reference quantity of "less than 30" are input, the second machine learning model 41 outputs comparison information indicating "low voice" as the pitch of the speaker's voice is lower than the boundary value of the reference quantity of "30" (corresponding to the reference quantity).
[0040] The second machine learning model 41 outputs impression information based on the input non-verbal information, feature quantities, reference quantities, and priorities. The second machine learning model 41 generates evidence information including impression information regarding a speaker's impression based on the non-verbal information and the comparison information. The second machine learning model 41 outputs, as impression information, information regarding the speaker's impression estimated based on the comparison information, which is information regarding the impression related to the non-verbal information. The second machine learning model 41 may have learned information regarding impressions in advance. For example, the second machine learning model 41 outputs impression information indicating a "rough impression" based on non-verbal information indicating the emotion of "anger" and comparison information indicating a "loud voice." Furthermore, for example, the second machine learning model 41 outputs impression information indicating a "sulky impression" based on non-verbal information indicating the emotion of "anger" and comparison information indicating a "low voice." Note that the generation unit 40 does not need to input non-verbal information to the second machine learning model 41. In this case, the second machine learning model 41 may generate impression information based on the comparison information even without inputting non-verbal information.
[0041] The second machine learning model 41 outputs notification information including evidence information based on the non-verbal information, the comparison information, and the impression information. For example, the second machine learning model 41 outputs notification information such as "You seem a little angry. You have a loud voice and give a rough impression. You also have a low voice and give a bad impression" based on non-verbal information such as the emotion of anger, comparison information such as "loud voice" and "low voice," and impression information such as "rough impression" and "sulky impression." The generation unit 40 acquires the notification information output by the second machine learning model 41 as information including the non-verbal information and evidence information.
[0042] The output unit 50 may output the basis information to the terminal 5. The output unit 50 may output non-verbal information to the terminal 5. In the present embodiment, the output unit 50 outputs notification information including the non-verbal information and the basis information to the terminal 5. The terminal 5 presents the notification information to the user. For example, the terminal 5 displays the notification information.
[0043] The processing procedure by the processing system 1 and device 10 configured as described above, i.e., the flow of the processing method according to this embodiment, will be described. FIG. 3 is a flowchart showing the procedure of an example of a processing method by the processing system according to the first embodiment of the present disclosure. The processing method MT shown in FIG. 3 (hereinafter, may be simply referred to as "method MT") is started, for example, when a speaker inputs voice into the terminal 5, by the device 10 receiving a signal from the terminal 5 indicating that the voice has been input. Note that the method may also be started when the device 10 receives a signal from the terminal 5 indicating that the speaker's voice is ready to be transmitted from the terminal 5 at any timing other than the timing when the speaker inputs voice.
[0044] In the method MT, first, in step S1, the speech acquisition unit 21 of the acquisition unit 20 acquires speech. The speech acquisition unit 21 receives the speech of the speaker transmitted from the terminal 5, thereby acquiring the speech.
[0045] Next, the detection unit 22 of the acquisition unit 20 acquires a feature amount in step S2. The detection unit 22 detects a feature amount related to the speech acquired by the speech acquisition unit 21 based on the speech.
[0046] Next, the estimation unit 30 acquires a reference amount in step S3. The estimation unit 30 acquires data indicating a reference amount for a feature amount related to non-language information from the non-language information database 7.
[0047] Next, the priority acquisition unit 23 of the acquisition unit 20 acquires a predetermined priority corresponding to the feature amount at step S4. The priority acquisition unit 23 acquires the priority by receiving the priority transmitted from the terminal 5.
[0048] Next, in step S5, the estimation unit 30 estimates the speaker's non-verbal information. Based on the feature quantities detected by the detection unit 22 and the acquired reference quantities, the estimation unit 30 estimates the emotion estimated to be indicated by the speech as an example of the speaker's non-verbal information. The estimation unit 30 inputs the feature quantities acquired by the detection unit 22 as input information to the first machine learning model 31. The first machine learning model 31 outputs the emotion estimated to be indicated by the speech as an example of non-verbal information based on the input feature quantities.
[0049] Next, the generating unit 40 generates basis information in step S6. The generating unit 40 acquires (generates) notification information including the basis information. Details of step S6 for generating the basis information will be described later.
[0050] Subsequently, the output unit 50 outputs the basis information in step S7. The output unit 50 outputs notification information including the basis information generated by the generation unit 40 to the terminal 5. The terminal 5 presents the notification information including the non-verbal information and the basis information to the speaker. When the basis information (or notification information including the basis information) is output and step S7 is completed, the method MT ends.
[0051] Next, step S6 of method MT will be described in detail. Fig. 4 is a flowchart showing an example of a step of generating basis information in the processing method by the processing system according to the first embodiment of the present disclosure. Step S6 shown in Fig. 4 includes steps S61 to S65.
[0052] In step S6, first, in step S61, the generation unit 40 extracts feature quantities corresponding to indices indicating high-priority voice qualities. Based on the feature quantities and priorities, the generation unit 40 extracts feature quantities corresponding to high-priority indices indicating voice qualities. For example, the generation unit 40 inputs the feature quantities and priorities as input information to the second machine learning model 41. Based on the input feature quantities and priorities, the second machine learning model 41 outputs values for indices indicating high-priority voice qualities among the feature quantities.
[0053] Next, in step S62, the generation unit 40 extracts a reference quantity corresponding to an index indicating a voice quality with high priority. The generation unit 40 extracts the reference quantity corresponding to an index indicating a voice quality with high priority from the non-language information database 7. Note that the generation unit 40 may extract the reference quantity corresponding to an index indicating a voice quality with high priority from the reference quantities acquired in step S3. The generation unit 40 inputs, for example, non-language information, the reference quantity, and the priority as input information to the second machine learning model 41. The second machine learning model 41 outputs a value for an index indicating a voice quality with high priority from the reference quantities corresponding to the estimated non-language information based on the input non-language information, the reference quantity, and the priority.
[0054] Next, in step S63, the generation unit 40 generates comparison information. The generation unit 40 compares the extracted feature quantities with the reference quantities and generates the result as comparison information. For example, the generation unit 40 inputs the feature quantities extracted in step S61 and the reference quantities extracted in step S62 as input information to the second machine learning model 41. The second machine learning model 41 outputs comparison information based on the feature quantities and the reference quantities.
[0055] Next, in step S64, the generation unit 40 generates impression information. The generation unit 40 generates the impression information based on the non-language information and the generated comparison information. The generation unit 40 inputs, for example, the estimated non-language information and the comparison information generated in step S63 as input information to the second machine learning model 41. The second machine learning model 41 outputs impression information based on the non-language information and the comparison information.
[0056] Next, in step S65, the generation unit 40 generates notification information. The generation unit 40 generates notification information having non-verbal information and basis information including at least one of comparison information and impression information, and outputs the notification information to the terminal 5. For example, the generation unit 40 inputs the estimated non-verbal information and basis information including at least one of the comparison information generated in step S63 and the comparison information generated in step S64 as input information to the second machine learning model 41. The second machine learning model 41 outputs notification information based on the non-verbal information and the basis information. The generation unit 40 acquires the notification information output from the second machine learning model 41. When the notification information is output and step S65 is completed, step S6 is completed.
[0057] Next, the effects of the device and method of the present disclosure will be described with reference to an example of a conventional problem. For example, a user may use a device (circuit) such as a machine learning model that classifies non-verbal information, including emotions or mental states, using speech as input information, and obtain non-verbal information from the machine learning model. It is known that a discrepancy occurs between the user's subjective non-verbal information and the non-verbal information estimated by the machine learning model based on objective data obtained from the user's speech. Therefore, even if a user understands the discrepancy, they may be dissatisfied with the estimated result when it differs from the subjective non-verbal information presented.
[0058] In this case, it is possible to present to the user the basis for the inference result output by the machine learning model. As the basis for the inference result output by the machine learning model, it is possible to extract the judgment used in the machine learning model in the process of deriving the inference result from the input information. However, even if the extracted judgment can be presented to the user as is, the judgment is merely optimized for processing by the machine learning model, and therefore the user may not be able to fully understand the presented judgment. Furthermore, if the user is presented with only the judgment, the user may not be able to interpret it by focusing on the degree of difference between the feature amount used in the judgment and the reference amount, or specific changes in the feature amount, etc. Therefore, there is a need for a technology that contributes to improving the user's sense of acceptance of non-verbal information estimated based on speech.
[0059] The device 10 of the present disclosure includes an acquisition unit 20 that acquires features related to speech, an estimation unit 30 that estimates non-verbal information of a speaker uttering a speech based on predetermined reference amounts related to the features and non-verbal information, a generation unit 40 that generates evidence information related to the basis for the estimation of the non-verbal information by the estimation unit 30 based on the features, the reference amounts, and predetermined priorities corresponding to the features, and an output unit 50 that outputs the evidence information.
[0060] In addition, the method MT disclosed herein includes a step S2 of acquiring features related to speech, a step S5 of estimating non-verbal information of the speaker uttering the speech based on predetermined reference amounts related to the features and non-verbal information, a step S6 of generating evidence information regarding the basis for estimating the non-verbal information in the estimation step S5 based on the features, the reference amounts, and predetermined priorities corresponding to the features, and a step S7 of outputting the evidence information.
[0061] The device 10 and method MT of the present disclosure output basis information regarding the basis for estimating non-verbal information. Because the basis information is generated based on features, reference values, and priorities, the speaker (an example of a user) can easily understand the basis for estimating the non-verbal information through parameter comparison, etc. For example, the basis information may include the degree of difference between the features and the reference values. Furthermore, for example, the basis information may include features that are easy for the speaker to understand depending on the priorities. In this way, the device 10 and method MT can output more appropriate estimation basis. Specifically, the device 10 and method MT can contribute to improving the speaker's (user's) sense of satisfaction with the estimated non-verbal information.
[0062] Furthermore, in the device 10 of the present disclosure, the estimation unit 30 estimates non-linguistic information by inputting the feature quantities to a first machine learning model 31 that has learned reference quantities in advance and outputs non-linguistic information in response to the input feature quantities. In this case, the generation unit 40 generates evidence information regarding the basis of the non-linguistic information estimated by the first machine learning model 31. Therefore, even if the speaker cannot understand the process of estimating non-linguistic information in the first machine learning model 31, the evidence information obtained by the generation unit 40 and the output unit 50 can be presented to the speaker, thereby promoting the speaker's understanding of the estimated non-linguistic information.
[0063] Furthermore, in the device 10 of the present disclosure, the generation unit 40 extracts values for indices indicating high-priority voice qualities from among the feature quantities and values for indices indicating high-priority voice qualities from among the reference quantities, and generates basis information including comparison information comparing the extracted feature quantities with the extracted reference quantities. In this case, the comparison results for the indices indicating high-priority voice qualities can be output to the speaker, which can encourage the speaker to understand the basis for the estimation of non-verbal information and contribute to improving the speaker's satisfaction with the estimated non-verbal information.
[0064] Furthermore, in the device 10 of the present disclosure, the generation unit 40 generates basis information including impression information regarding the impression of the speaker based on the non-verbal information and the comparison information. In this case, by generating the impression information of the speaker, the relationship between the feature amount and the reference amount can be presented to the speaker in a more easily understandable format, which can encourage the speaker to understand the basis for the estimation of the non-verbal information and contribute to improving the speaker's sense of satisfaction with the estimated non-verbal information.
[0065] Second Embodiment FIG. 5 is a block diagram showing the configuration of a processing system including an apparatus according to a second embodiment of the present disclosure. The processing system 1A shown in FIG. 5 includes a terminal 5A, a non-verbal information database 7, a past information database 8, and an apparatus 10A, all of which are configured to communicate with each other via a network including a wireless communication network and a fixed communication network. The processing system 1A of the second embodiment differs from the processing system 1 of the first embodiment in that it includes the terminal 5A, the past information database 8, and the apparatus 10A. In the processing system 1A of the second embodiment, entities having the same configurations and functions as entities included in the processing system 1 of the first embodiment are given the same names and reference numerals, and descriptions thereof will be omitted. Furthermore, in the processing system 1A of the second embodiment, entities having some of the same configurations and functions as entities included in the processing system 1 of the first embodiment are given the same names, and descriptions of the same configurations and functions will be omitted.
[0066] The processing system 1A is used to present estimation grounds including a result of comparing first non-verbal information of the speaker at a first time point with second non-verbal information of the speaker at a second time point after the first time point when estimating non-verbal information of the speaker using the terminal 5A. The speaker is, for example, a user who wishes to objectively view his or her own non-verbal information and to reflect on his or her own non-verbal information. In the second embodiment, a case will be described in which the speaker inputs speech into the terminal 5A at the second time point and obtains estimation grounds for the second non-verbal information. The second time point is a date and time point later than the first time point. The second time point is, for example, two days after the first time point.
[0067] The device 10A of the second embodiment receives the speaker's speech from the terminal 5A at the first and second time points, and outputs, to the terminal 5A, second non-verbal information that is non-verbal information of the speaker estimated from the speech, and basis information that is the basis for estimating the second non-verbal information. Note that the device 10A does not necessarily have to output the speaker's non-verbal information.
[0068] Terminal 5A acquires the speaker's voice. Terminal 5A outputs the acquired voice to device 10A. Terminal 5A outputs to device 10A the voice that was uttered by the speaker at a first time point and input to terminal 5A. Hereinafter, this voice may be referred to as the first voice. Terminal 5A outputs to device 10A the voice that was uttered by the speaker at a second time point after the first time point and input to terminal 5A. Hereinafter, this voice may be referred to as the second voice.
[0069] Furthermore, terminal 5A acquires the non-language information estimated by device 10A and basis information regarding the basis for estimating the non-language information. Specifically, terminal 5A may acquire first non-language information of a speaker uttering a first voice and second non-language information of a speaker uttering a second voice. Terminal 5A does not need to acquire the first non-language information. Terminal 5A acquires basis information regarding second non-language information including a comparison result between the first feature amount and the second feature amount. Details of the first feature amount and the second feature amount will be described later. Terminal 5A may acquire basis information including a comparison result between the first non-language information and the second non-language information.
[0070] The past information database 8 may store a first speech of a speaker at a first time point acquired by terminal 5A. The past information database 8 may store at least one of a first feature amount, first non-language information, and evidence information related to the first non-language information output by device 10A. The past information database 8 may store at least one of a second feature amount, second non-language information, and evidence information related to the second non-language information output by device 10A.
[0071] The device 10A shown in Fig. 5 is configured to include, as functional components, an acquisition unit 20A, an estimation unit 30, a generation unit 40, and an output unit 50. The estimation unit 30 estimates first non-verbal information and second non-verbal information based on the first speech and second speech acquired from the terminal 5A by the acquisition unit 20A. The generation unit 40 generates basis information regarding the basis for estimating the second non-verbal information. The output unit 50 outputs at least the basis information to the terminal 5A. The functions of each functional unit of the device 10A will be described in detail below.
[0072] The acquisition unit 20A acquires a first speech and a second speech of a speaker, as well as a priority. The acquisition unit 20A acquires a first feature based on the first speech. The acquisition unit 20A acquires a second feature based on the second speech. The acquisition unit 20A has a speech acquisition unit 21A, a detection unit 22A, and a priority acquisition unit 23.
[0073] The speech acquisition unit 21A acquires a first speech input by a speaker to the terminal 5 at a first time point from the terminal 5. The speech acquisition unit 21A acquires the first speech by receiving the first speech transmitted from the terminal 5 at the first time point. Note that the speech acquisition unit 21A may acquire the first speech from a database (e.g., the past information database 8) in which speech acquired by the terminal 5 is stored.
[0074] The speech acquisition unit 21A acquires, from the terminal 5, a second speech input to the terminal 5 by a speaker at a second time point. The speech acquisition unit 21A acquires the second speech by receiving the second speech transmitted from the terminal 5 at a second time point after the first time point. Note that the speech acquisition unit 21A may acquire the second speech from a database (e.g., the past information database 8) in which speech acquired by the terminal 5 is stored.
[0075] The voice acquisition unit 21A may acquire a first voice at a first time point and may acquire a second voice at a second time point. The voice acquisition unit 21A may acquire the first voice and the second voice at the second time point.
[0076] The detection unit 22A detects a first feature quantity related to the acquired first speech based on the acquired first speech. The detection unit 22A detects a second feature quantity related to the acquired second speech based on the acquired second speech. The first feature quantity and the second feature quantity have, for example, the same properties as the feature quantities in the first embodiment described above.
[0077] At a first time point in the second embodiment, for example, if a speaker speaks rapidly in a slightly louder, lower voice, the detection unit 22A analyzes the feature quantities indicated by the speaker's voice, "volume," "pitch," and "clarity," as 60, 20, and 40, respectively. At a second time point in the second embodiment, for example, if a speaker speaks rapidly in a voice louder than at the first time point, and in the same lower voice as at the first time point, the detection unit 22A analyzes the feature quantities indicated by the speaker's voice, "volume," "pitch," and "clarity," as 70, 20, and 40, respectively.
[0078] The detection unit 22A may acquire a first feature amount at a first time point and acquire a second feature amount at a second time point. The speech acquisition unit 21A may acquire the first feature amount and the second feature amount at a second time point.
[0079] The estimation unit 30 estimates first non-language information and second non-language information of a speaker who utters a voice, based on predetermined reference amounts related to the feature amounts and the non-language information. The estimation unit 30 acquires data indicating reference amounts related to the non-language information from the non-language information database 7. The estimation unit 30 estimates an emotion indicated by the first voice as the speaker's first non-language information, based on the first feature amount indicated by the first voice and the reference amount acquired from the non-language information database 7. The estimation unit 30 estimates an emotion indicated by the second voice as the speaker's second non-language information, based on the second feature amount indicated by the second voice and the reference amount acquired from the non-language information database 7.
[0080] The estimation unit 30 estimates the first non-language information and the second non-language information using, for example, a first machine learning model 31. The estimation unit 30 inputs the first feature quantities acquired by the detection unit 22A to the first machine learning model 31 as input information. For example, the estimation unit 30 inputs first feature quantities of 60, 20, and 40 for "volume," "pitch," and "clarity" as input information to the first machine learning model 31. The first machine learning model 31 estimates that the speaker's first non-language information is the emotion of "anger" based on pre-trained reference quantities and the input feature quantities, and outputs the emotion. The estimation unit 30 acquires the emotion output by the first machine learning model 31 as the first non-language information.
[0081] The estimation unit 30 inputs the second feature quantities acquired by the detection unit 22A as input information to the first machine learning model 31. For example, the estimation unit 30 inputs second feature quantities of 70, 20, and 40 for "volume," "pitch," and "clarity" as input information to the first machine learning model 31. The first machine learning model 31 estimates that the speaker's second non-verbal information is the emotion of "anger" based on pre-trained reference quantities and the input feature quantities, and outputs the emotion. The estimation unit 30 acquires the emotion output by the first machine learning model 31 as second non-verbal information.
[0082] The generation unit 40 generates basis information regarding the basis for estimating the second non-language information in the estimation unit 30 based on the first feature amount, the second feature amount, the reference amount, and the priority. The basis information regarding the second non-language information includes a comparison result between the first feature amount and the second feature amount. The comparison result between the first feature amount and the second feature amount includes information indicating a transition from the first speech feature of the speaker at the first time point to the second speech feature of the speaker at the second time point. Hereinafter, information including the comparison result between the first feature amount and the second feature amount may be referred to as transition information.
[0083] The transition information (comparison result) may include an expression evaluating the relationship between the first feature and the second feature. The transition information (comparison result) may include, for example, at least one of the magnitude relationship and the magnitude of the difference between the first feature and the second feature. The transition information includes a comparison result between a value of an index indicating a voice quality with a high priority among the first feature and a value of an index indicating a voice quality with a high priority among the second feature. The generation unit 40 may, for example, use the second machine learning model 41 to generate evidence information including at least the transition information or notification information including the transition information. Note that the notification information may include at least one of the first non-language information and the second non-language information.
[0084] For example, a case will be described in which the indices indicating voice quality corresponding to a priority higher than a predetermined threshold are "volume" and "pitch." Based on the input first feature and priority, the second machine learning model 41 outputs a value for the indices indicating high-priority voice quality among the first feature. Based on the priority, the second machine learning model 41 extracts 60 and 20, which indicate "volume" and "pitch," as values for the indices indicating high-priority voice quality among the first feature values of 60, 20, and 40 for "volume," "pitch," and "clarity."
[0085] The second machine learning model 41 outputs a value for an index indicating a voice quality with a high priority among the second features based on the input second features and priorities. Based on the priorities, the second machine learning model 41 extracts 70 and 20, which indicate "volume" and "pitch," as values for an index indicating a voice quality with a high priority among the second features of 70, 20, and 40 for "volume," "pitch," and "clarity."
[0086] The second machine learning model 41 outputs first comparison information regarding the first non-language information based on the input first feature, reference value, and priority. For example, when a first feature value of 60 for "volume" and a reference value of 50 are input, the second machine learning model 41 determines that the volume of the speaker's first voice is greater than the reference value and outputs first comparison information indicating "loud voice." Furthermore, when a first feature value of 20 for "pitch" and a reference value of 30 are input, the second machine learning model 41 determines that the pitch of the speaker's first voice is lower than the reference value and outputs first comparison information indicating "low voice." Note that the second machine learning model 41 does not need to estimate the first comparison information at the second time point, and may instead acquire first comparison information stored in the past information database 8.
[0087] The second machine learning model 41 outputs second comparison information regarding the second non-language information based on the input first feature amount, second feature amount, reference amount, and priority. For example, when a second feature amount of 70 for "volume" and a reference amount of "50 or more" are input, the second machine learning model 41 determines that the volume of the speaker's second voice is greater than the boundary value "50" of the reference amount (corresponding to the reference amount), and outputs second comparison information indicating "loud voice." Furthermore, when a second feature amount of 20 for "pitch" and a reference amount of "less than 30" are input, the second machine learning model 41 determines that the pitch of the speaker's second voice is lower than the boundary value "30" of the reference amount (corresponding to the reference amount), and outputs second comparison information indicating "low voice."
[0088] The second machine learning model 41 outputs transition information based on at least one of the input first non-verbal information, second non-verbal information, first feature amount, second feature amount, reference amount, and priority. For example, the second machine learning model 41 outputs transition information based on the input first non-verbal information and second non-verbal information. For example, when the first non-verbal information and the second non-verbal information are input, the emotion of "anger" does not change, so the second machine learning model 41 outputs transition information such as "I've been feeling a bit angrier recently."
[0089] For example, the second machine learning model 41 outputs transition information based on the input first feature amount, second feature amount, first comparison information, and second comparison information. For example, when the second machine learning model 41 receives the first feature amount of 60 for "volume," the second feature amount of 70, the first comparison information, and the second comparison information, the second machine learning model 41 outputs transition information that "the voice is even louder today" because the second feature amount is larger than the first feature amount, although the state of "loud voice" remains unchanged.
[0090] Furthermore, for example, when the second machine learning model 41 receives the first feature value of 20 for "pitch," the second feature value of 20, the first comparison information, and the second comparison information, the state of "low voice" remains unchanged and the first feature value and the second feature value are the same, so it outputs transition information that "the voice remains low."
[0091] The second machine learning model 41 outputs notification information including evidence information based on at least one of the second non-verbal information, the first comparison information, the second comparison information, the transition information, and the impression information. For example, the second machine learning model 41 outputs notification information such as "You seem a little angry lately. Today, your voice is even louder and you seem rougher. Also, your voice has remained low, giving you the impression of being in a bad mood" based on the second non-verbal information indicating an emotion of anger, the transition information indicating "your voice is even louder today" and "your voice has remained low," and the impression information indicating "your impression of being rough" and "your impression of being grumpy." The generation unit 40 acquires the notification information output by the second machine learning model 41 as information including the transition information (evidence information).
[0092] The generation unit 40 may output basis information related to the first non-language information. Here, the basis information related to the first non-language information may include the same information as the basis information related to the non-language information according to the first embodiment, or may include basis information including a result of comparing non-language information at a time point prior to the first time point with the basis information related to the non-language information.
[0093] The output unit 50 may output the basis information regarding the first non-language information at the first time point. The output unit 50 may store the basis information regarding the first non-language information in the past information database 8.
[0094] The output unit 50 outputs the basis information regarding the second non-language information at the second time point. The output unit 50 outputs the basis information including at least the transition information to the terminal 5A. The output unit 50 may store the basis information regarding the second non-language information in the past information database 8.
[0095] The processing procedure by the processing system 1A and device 10A configured as described above, i.e., the flow of the processing method according to this embodiment, will be described. FIG. 6 is a flowchart showing the procedure of an example of a processing method by the processing system according to the second embodiment of the present disclosure. The processing method MTA shown in FIG. 6 (hereinafter, simply referred to as "method MTA") is started, for example, when a speaker inputs voice into terminal 5A, and device 10 receives a signal from terminal 5A indicating that the voice has been input. Note that method MTA shown in FIG. 6 may be started when the speaker's voice at a first time point prior to a second time point at which the speaker inputs voice is available from the past information database 8.
[0096] In the method MTA, first, the speech acquisition unit 21A of the acquisition unit 20A acquires a first speech, which is a speech at a first time point, from the past information database 8 in step SA11.
[0097] Next, in Step SA21, the detection unit 22A of the acquisition unit 20A acquires a first feature amount, which is a feature amount at a first time point related to the first voice, based on the first voice acquired by the voice acquisition unit 21A.
[0098] Next, in step SA12, the speech acquisition unit 21A of the acquisition unit 20A acquires second speech, which is speech at a second time point, from the terminal 5A. Note that, at least in step SA12, when the speech at the second time point input by the speaker is in a state where it can be acquired from the past information database 8, the speech acquisition unit 21A may acquire the second speech from the past information database 8.
[0099] Next, in Step SA22, the detection unit 22A of the acquisition unit 20A acquires, based on the second voice acquired by the voice acquisition unit 21A, a second feature amount that is a feature amount related to the second voice at a second time point.
[0100] Next, the estimation unit 30 acquires a reference amount in step SA3. Step SA3 is the same process as step S3 in the first embodiment.
[0101] Next, the estimation unit 30 acquires the priority in step SA4, which is the same process as step S4 in the first embodiment.
[0102] Next, in step SA5, the estimation unit 30 estimates first non-language information based on the first feature amount and the reference amount. Furthermore, in step SA5, the estimation unit 30 estimates second non-language information based on the second feature amount and the reference amount.
[0103] Next, in step SA6, the generation unit 40 generates evidence information regarding the second non-language information. The generation unit 40 generates evidence information including at least transition information, for example, by inputting input information to the second machine learning model 41 as described above. Note that the generation unit 40 may also generate evidence information regarding the first non-language information in step SA6.
[0104] Subsequently, the device 10 outputs the basis information in step SA7. Step SA7 is the same process as step S7 in the first embodiment. When the basis information (or notification information including the basis information) is output and step SA7 is completed, the method MTA ends.
[0105] The device 10A of the present disclosure includes an acquisition unit 20A that acquires a first feature as a feature corresponding to speech at a first time point and a second feature as a feature corresponding to speech at a second time point after the first time point; an estimation unit 30 that estimates first non-language information of a speaker uttering the speech at the first time point and second non-language information of a speaker uttering the speech at the second time point; a generation unit 40 that generates evidence information regarding the second non-language information based on the first feature, the second feature, a reference amount, and a priority; and an output unit 50 that outputs the evidence information regarding the second non-language information, wherein the evidence information regarding the second non-language information includes a comparison result between the first feature and the second feature.
[0106] The device 10A and method MTA of the present disclosure output evidence information regarding the basis for estimating the second non-verbal information. Because the evidence information includes a comparison result between the first feature and the second feature, the speaker (an example of a user) can easily understand the basis for estimating the second non-verbal information by, for example, comparing each parameter of the input data with the speaker's own past data. In this way, the device 10A and method MTA can contribute to improving the speaker's (user's) sense of satisfaction with the estimated second non-verbal information. Furthermore, the device 10A and method MTA can provide the speaker with an opportunity to review the transition of the speaker's own feature from the first time point to the second time point.
[0107] Furthermore, in the device 10A of the present disclosure, the comparison result includes content indicating a transition from the first speech feature of the speaker at the first time point to the first speech feature of the speaker at the second time point. In this case, the device 10A and the method MTA can provide the speaker with an opportunity to reflect on the transition of their own non-verbal information from the first time point to the second time point.
[0108] [Third Embodiment] Figure 7 is a block diagram showing the configuration of a processing system including an apparatus according to a third embodiment of the present disclosure. The processing system 1B shown in Figure 7 includes a terminal 5B, a non-verbal information database 7, a subjective information database 9, and an apparatus 10B, which are configured to communicate with each other via a network including a wireless communication network and a fixed communication network. The processing system 1B of the third embodiment differs from the processing system 1 of the first embodiment in that it includes a terminal 5B, a subjective information database 9, and an apparatus 10B. In the processing system 1B of the third embodiment, entities having the same configurations and functions as entities included in the processing system 1 of the first embodiment are given the same names and reference numerals, and descriptions thereof will be omitted. Furthermore, in the processing system 1B of the third embodiment, entities having some of the same configurations and functions as entities included in the processing system 1 of the first embodiment are given the same names, and descriptions of the same configurations and functions will be omitted.
[0109] When estimating non-verbal information of a speaker using terminal 5B, processing system 1B is used to estimate objective estimated information, which is non-verbal information estimated based on feature amounts and reference amounts that are statistics of multiple speakers, and subjective estimated information, which is non-verbal information estimated based on feature amounts and threshold values set for each speaker. Furthermore, processing system 1B is used to present objective ground information as ground information for the objective estimated information and subjective ground information as ground information for the subjective estimated information.
[0110] The speaker is, for example, a user who wishes to objectively view his / her own non-verbal information and also wishes to visualize the difference between non-verbal information estimated using objective indicators (objective estimated information) and non-verbal information estimated using his / her own subjective indicators (subjective estimated information). In the third embodiment, a case will be described in which a speaker inputs speech to terminal 5B and acquires objective ground information and subjective ground information.
[0111] The device 10B of the second embodiment receives the speech of a speaker from the terminal 5B, and outputs to the terminal 5B objective estimation information and subjective estimation information of the speaker estimated based on the speech, as well as objective ground information that is the basis for estimating the objective estimation information and subjective ground information that is the basis for estimating the subjective estimation information. Note that the device 10B does not have to output at least one of the objective estimation information and the subjective estimation information of the speaker.
[0112] Terminal 5B acquires the speaker's voice. Terminal 5B outputs the acquired voice to device 10B. Terminal 5B also acquires objective estimation information and subjective estimation information estimated by device 10B, as well as objective ground information and subjective ground information. Terminal 5B does not need to acquire at least one of the objective estimation information and the subjective estimation information. Terminal 5B may acquire only information on parts where the objective ground information and the subjective ground information differ. Furthermore, terminal 5B accepts input of subjective information of the speaker. Terminal 5B outputs the acquired subjective information to device 10B. Details of the subjective information will be omitted.
[0113] The subjective information database 9 stores the subjective information described below. The subjective information database 9 stores a threshold value related to non-verbal information set for each speaker. The threshold value is data that can be compared with feature quantities. The threshold value may be numerical data (vector data). The threshold value is, for example, data that expresses at least one of each emotion and each state of mind included in the non-verbal information. The subjective information database 9 may store the subjective information described below.
[0114] The device 10B shown in Fig. 7 includes, as functional components, an acquisition unit 20B, an estimation unit 30, a generation unit 40, an output unit 50, a determination unit 60, and an update unit 70. The estimation unit 30 estimates objective estimation information and subjective estimation information based on the speech acquired from the terminal 5B. The generation unit 40 generates objective ground information and subjective ground information relating to the grounds for estimating the objective estimation information and the subjective estimation information, respectively. The output unit 50 outputs at least the objective ground information and the subjective ground information to the terminal 5B. The functions of each functional unit of the device 10B will be described in detail below.
[0115] The acquisition unit 20B acquires the speaker's voice and priority. The acquisition unit 20B acquires features based on the voice. The acquisition unit 20B receives subjective information, which is non-verbal information expected by the speaker. The acquisition unit 20B has a voice acquisition unit 21, a detection unit 22, a priority acquisition unit 23, and a reception unit 24. The voice acquisition unit 21, the detection unit 22, and the priority acquisition unit 23 of the third embodiment have the same configurations and functions as the voice acquisition unit 21, the detection unit 22, and the priority acquisition unit 23 of the first embodiment. Note that the voice and priority acquired in the third embodiment, and the features detected will be described below using the same example as that shown in the first embodiment.
[0116] The receiving unit 24 receives subjective information including an index indicating the voice quality of the voice expected by the speaker. The subjective information may be information expected by the speaker, regardless of the non-verbal information output by the device 10B. The subjective information is information that the speaker himself / herself subjectively derives, indicating how the speaker evaluates at least one of the indices related to the non-verbal information and the voice quality. The receiving unit 24 acquires, from the terminal 5B, the subjective information input by the speaker to the terminal 5B. The receiving unit 24 acquires the subjective information by accepting the subjective information transmitted from the terminal 5B. Note that the receiving unit 24 may acquire the subjective information from a database (not shown) in which the subjective information acquired by the terminal 5B is stored. For example, the receiving unit 24 acquires subjective information such as, "I feel angry, my voice is loud, and my voice is much lower than usual."
[0117] The estimation unit 30 estimates non-language information of a speaker uttering a voice based on the feature amounts and reference amounts related to the non-language information. The estimation unit 30 of the third embodiment estimates objective estimation information as non-language information based on the feature amounts and reference amounts that are statistics of multiple speakers. The estimation unit 30 acquires data indicating the reference amounts that are statistics of multiple speakers related to the non-language information from the non-language information database 7. The reference amounts shown in FIG. 2 may be data indicating the reference amounts that are statistics of multiple speakers in the third embodiment.
[0118] The estimation unit 30 estimates the objective estimation information using, for example, a first machine learning model 31. The estimation unit 30 of the third embodiment executes the same process as the estimation unit 30 of the first embodiment. In the example of the third embodiment, the first machine learning model 31 also estimates that the speaker's objective estimation information is the emotion of "anger" and acquires the emotion as objective estimation information.
[0119] The estimation unit 30 of the third embodiment estimates subjective estimation information as non-language information based on the feature amount and a threshold value related to non-language information set for each speaker. The estimation unit 30 acquires data indicating the threshold value set by the speaker who inputs the speech related to the non-language information from the subjective information database 9.
[0120] The estimation unit 30 estimates the subjective estimation information using, for example, a first machine learning model 31. The estimation unit 30 of the third embodiment replaces the reference amount with a threshold value and performs processing similar to that of the estimation unit 30 of the first embodiment. In the example of the third embodiment, the first machine learning model 31 also estimates that the speaker's subjective estimation information is the emotion of "anger" and acquires this emotion as objective estimation information. Note that the thresholds for the emotion of "anger" used here are, for example, "50 or more" and "20 or less" for "volume" and "pitch."
[0121] The generating unit 40 generates objective grounds information regarding the grounds for estimating the objective estimation information in the estimating unit 30 based on the feature amount, the reference amount, and the priority. The generating unit 40 of the third embodiment executes the same processing as the generating unit 40 of the first embodiment. In the example of the third embodiment, the second machine learning model 41 also generates at least one of comparison information and impression information as objective grounds information, and may generate notification information including objective grounds information such as "He seems a little angry. He has a loud voice and gives the impression of being rough. He also has a low voice and gives the impression of being in a bad mood."
[0122] The generation unit 40 generates subjective basis information regarding the basis for estimating the subjective estimation information in the estimation unit 30 based on the feature, threshold, and priority. The generation unit 40 of the third embodiment replaces the reference quantity with a threshold and performs processing similar to that of the generation unit 40 of the first embodiment. In the example of the third embodiment, the second machine learning model 41 receives the feature of 20 for "pitch" and the threshold of "20 or less," and determines that the pitch of the speaker's voice is as low as the reference quantity, thereby outputting comparison information indicating "a slightly low voice." The second machine learning model 41 may generate at least one of comparison information and impression information as the subjective basis information, and generate notification information including subjective basis information such as "He seems a little angry. His voice is loud and gives a rough impression. Also, his voice is a little low and gives the impression of being in a bad mood."
[0123] For example, the generating unit 40 generates notification information including objective ground information and subjective ground information. The generating unit 40 generates notification information further including at least one of objective estimation information and subjective estimation information. The generating unit 40 generates notification information that presents the objective ground information and the subjective ground information in a format that allows distinction between them. The generating unit 40 generates, for example, notification information such as "Evaluation based on objective indicators: You seem a little angry. Your voice is loud and gives the impression of being rough. Also, your voice is low and gives the impression of being in a bad mood. Evaluation based on subjective indicators: You seem a little angry. Your voice is loud and gives the impression of being rough. Also, your voice is slightly low and gives the impression of being in a bad mood."
[0124] The output unit 50 outputs objective grounds information and subjective grounds information. The output unit 50 outputs notification information including at least the objective grounds information and the subjective grounds information to the terminal 5B. For example, the output unit 50 displays the notification information in a manner that allows distinction between objective grounds information using an objective reference amount and subjective grounds information using a subjective threshold. The output unit 50 generates notification information that further includes at least one of objective estimation information and subjective estimation information.
[0125] The determination unit 60 determines whether the subjective basis information is different from the subjective information. If the determination unit 60 determines that the subjective basis information is different from the subjective information, it determines that the threshold value acquired from the subjective information database 9 is not set appropriately, and causes the update unit 70 to update the subjective basis information. If the determination unit 60 determines that the subjective basis information is the same as the subjective information, it determines that the threshold value acquired from the subjective information database 9 is set appropriately, and therefore the update unit 70 does not perform the update.
[0126] In the above example, based on the notification information generated by the generation unit 40, which includes subjective grounds information such as "You seem angry. Your voice is loud and rough. Also, your voice is a little low and you seem to be in a bad mood," and the subjective information acquired by the reception unit 24, "I feel angry, my voice is loud, and it's much lower than usual," the judgment unit 60 judges that the subjective grounds information regarding the lowness of the voice differs from the subjective information.
[0127] The determination unit 60 may determine whether the subjective estimation information is different from the subjective information. If the determination unit 60 determines that the subjective estimation information is different from the subjective information, it determines that the thresholds acquired from the subjective information database 9 are not set appropriately overall, and causes the update unit 70 to update the thresholds. If the determination unit 60 determines that the subjective estimation information is identical to the subjective information, it determines that the thresholds acquired from the subjective information database 9 are set appropriately overall, and does not perform the update by the update unit 70.
[0128] The determination unit 60 may determine whether the objective ground information is different from the subjective ground information. If the determination unit 60 determines that the objective ground information is different from the subjective ground information, it may determine that the threshold value acquired from the subjective information database 9 is not appropriately set, and may cause the updating unit 70 to update the threshold value. If the determination unit 60 determines that the objective ground information is identical to the subjective ground information, it may determine that the threshold value acquired from the subjective information database 9 is appropriately set overall, and may not perform updating by the updating unit 70. The determination unit 60 may determine whether the objective estimation information is different from the subjective estimation information. If the determination unit 60 determines that the objective estimation information is different from the subjective estimation information, it may determine that the threshold value acquired from the subjective information database 9 is not appropriately set, and may cause the updating unit 70 to update the threshold value. If the determination unit 60 determines that the objective estimation information is identical to the subjective estimation information, it may determine that the threshold value acquired from the subjective information database 9 is appropriately set overall, and may not perform updating by the updating unit 70.
[0129] The updating unit 70 updates the threshold set for each speaker in accordance with the feature. For example, the updating unit 70 sets the threshold so that the difference between the subjective information and the subjective basis information becomes small. In the above example, since the determining unit 60 determines that the subjective basis information for the low voice level differs from the subjective information, the updating unit 70 sets the threshold to a slightly higher value of "25 or less" so that the feature for the low voice level is significantly different from the threshold.
[0130] The update unit 70 may set the threshold to approach the reference amount based on, for example, the difference between the feature amount and the reference amount and the difference between the feature amount and the threshold. The update unit 70 may notify the terminal 5B of the index for which the threshold should be updated and the value already set for that index, and prompt the speaker to update the threshold. In this case, the update unit 70 may set the value input by the terminal 5B as the new threshold. The update unit 70 may store the threshold in the subjective information database 9 so as to update it to the newly set threshold.
[0131] The processing procedure by the processing system 1B and device 10B configured as described above, i.e., the flow of the processing method according to this embodiment, will be described. Fig. 8 is a flowchart showing the procedure of an example of a processing method by the processing system according to the third embodiment of the present disclosure. The processing method MTB shown in Fig. 8 (hereinafter, may be simply referred to as "method MTB") is started, for example, when a speaker inputs voice into terminal 5B, and device 10B receives a signal from terminal 5B indicating that the voice has been input.
[0132] In method MTB, first, in step SB1, the speech acquisition unit 21 of the acquisition unit 20 acquires the speaker's speech from the terminal 5B. Step SB1 is the same process as step S1 in the first embodiment described above.
[0133] Next, in step SB2, the detection unit 22 of the acquisition unit 20B acquires features related to the speech based on the speech acquired by the speech acquisition unit 21. Step SB2 is the same process as step S2 in the first embodiment described above.
[0134] Next, the estimation unit 30 acquires a reference amount in step SB31. Step SB31 is the same process as step S3 in the first embodiment.
[0135] Next, in step SB32, the estimation unit 30 acquires a threshold value. The estimation unit 30 acquires a threshold value set by the speaker for the feature amount stored in the subjective information database 9. Note that the estimation unit 30 may acquire the threshold value directly from the terminal 5B.
[0136] Next, the estimation unit 30 acquires the priority at step SB4. Step SB4 is the same process as step S4 in the first embodiment.
[0137] Next, in step SB5, the estimation unit 30 estimates objective estimation information based on the feature amount and the reference amount. Furthermore, in step SB5, the estimation unit 30 estimates subjective estimation information based on the feature amount and a threshold value.
[0138] Next, in Step SB6, the generating unit 40 generates objective ground information related to the objective estimation information. In Step SB6, the generating unit 40 generates subjective ground information related to the subjective estimation information.
[0139] Subsequently, the device 10B outputs the objective grounds information and the subjective grounds information in step SB7. When the objective grounds information and the subjective grounds information are output and step SB7 is completed, the method MTB ends.
[0140] 9 is a flowchart showing the procedure of an example of a processing method by the processing system according to the third embodiment of the present disclosure. The processing method MTC shown in FIG. 9 (hereinafter, may be simply referred to as "method MTC") is executed after the objective ground information and the subjective ground information are output by method MTB, for example.
[0141] In method MTC, first, the receiving unit 24 receives subjective information in step SC1. The receiving unit 24 acquires the subjective information of the speaker by receiving the subjective information transmitted from the terminal 5. The receiving unit 24 may acquire the subjective information from the subjective information database 9. Note that in method MTC, only step SC1 may be executed before execution of method MTB.
[0142] Next, in step SC2, the determination unit 60 determines whether the subjective basis information is different from the subjective information. If the determination unit 60 determines that the subjective basis information is different from the subjective information (step SC2: YES), the process proceeds to step SC3. If the determination unit 60 determines that the subjective basis information is the same as the subjective information (step SC2: NO), the method MTC ends.
[0143] If the determination unit 60 determines that the subjective basis information is different from the subjective information (step SC2: YES), the update unit 70 updates the threshold value in step SC3. The update unit 70 updates the threshold value stored in the subjective information database 9 by any of the above-described predetermined methods. The threshold value is updated to complete step SC3, and the method MTC ends.
[0144] In the device 10B of the present disclosure, the estimation unit 30 estimates objective estimation information as non-verbal information based on features and a reference quantity which is a statistical quantity of multiple speakers, and estimates subjective estimation information as non-verbal information based on the features and a threshold value for non-verbal information set for each speaker, the generation unit 40 generates objective ground information as ground information for the objective estimation information and subjective ground information as ground information for the subjective estimation information based on the features, reference quantity, threshold value and priority, and the output unit 50 outputs the objective ground information and the subjective ground information.
[0145] The device 10B and method MTB of the present disclosure estimate objective estimation information based on feature quantities and reference quantities, which are statistics of multiple speakers, and estimate subjective grounds information based on feature quantities and thresholds set for each speaker. The objective grounds information and subjective grounds information make it easier for a speaker (an example of a user) to understand the grounds for estimating the objective estimation information through, for example, a comparison between objective data and subjective data. In this way, the device 10B and method MTB can contribute to improving the speaker's (user's) sense of satisfaction with the estimated non-verbal information. Furthermore, the device 10B and method MTB can provide the speaker with an opportunity to understand the discrepancy between the objective grounds information and the subjective grounds information that arises due to the difference between the reference quantities, which are statistics of multiple speakers, and the thresholds set for each speaker, i.e., an opportunity to view themselves objectively.
[0146] Furthermore, in the device 10B of the present disclosure, the generation unit 40 generates notification information that presents the objective ground information and the subjective ground information in a format that allows the speaker to distinguish between the objective ground information and the subjective ground information. In this case, the device 10B and the method MTB can present the information to the speaker in a format that allows the speaker to easily refer to each of the objective ground information and the subjective ground information, thereby further contributing to an improvement in the speaker's sense of satisfaction.
[0147] In addition, the device 10B of the present disclosure further includes a receiving unit that receives subjective information, which is non-verbal information assumed by the speaker, a determination unit 60 that determines whether the subjective estimation information differs from the subjective information, and an update unit 70 that updates the threshold value when the determination unit 60 determines that the subjective grounds information differs from the subjective information.
[0148] The device 10B and method MTC of the present disclosure update the threshold value when the determination unit 60 determines that the subjective basis information differs from the subjective information. This reduces the discrepancy between the subjectively estimated non-verbal information and the basis information regarding the basis of the non-verbal information, and the non-verbal information estimated via the device 10B and the basis information regarding the basis of the non-verbal information. In this way, the device 10B and method MTC can contribute to improving the speaker's (user's) sense of satisfaction with the estimated non-verbal information (subjective estimation information).
[0149] 1, 5, and 7. In the processing systems 1, 1A, and 1B, the terminals 5, 5A, and 5B may include at least one of the first machine learning model 31 and the second machine learning model 41. This configuration can be realized, for example, by installing an application that executes the functions of at least one of the first machine learning model 31 and the second machine learning model 41, and at least one of the first machine learning model 31 and the second machine learning model 41, on the terminal 5.
[0150] Alternatively, in the processing systems 1, 1A, and 1B, the terminals 5, 5A, and 5B may include the devices 10, 10A, and 10B. This configuration can be realized, for example, by installing applications that execute the functions of the devices 10, 10A, and 10B and applications that execute the functions of the first machine learning model 31 and the second machine learning model 41 on the terminals 5, 5A, and 5B.
[0151] 1, 5 and 7, the non-verbal information database 7, the past information database 8 and the subjective information database 9 are shown as external servers of the terminals 5, 5A and 5B and the devices 10, 10A and 10B, but the data (information) stored in at least one of the non-verbal information database 7, the past information database 8 and the subjective information database 9 may be included in any of the terminals 5, 5A and 5B and the devices 10, 10A and 10B. In this case, at least one of the non-verbal information database 7, the past information database 8 and the subjective information database 9 may not be provided in the processing systems 1, 1A and 1B.
[0152] The device and method of the present disclosure have the following configuration.
[0153] [1] An apparatus comprising: an acquisition unit that acquires features related to speech; an estimation unit that estimates the non-language information of a speaker uttering the speech based on predetermined reference amounts related to the features and non-language information; a generation unit that generates basis information related to the basis for the estimation of the non-language information by the estimation unit based on the features, the reference amounts, and priorities corresponding to the features; and an output unit that outputs the basis information.
[0154] [2] The device according to [1], wherein the estimation unit estimates the non-language information by inputting the feature amounts to a machine learning model that has learned the reference amount in advance and outputs the non-language information by inputting the feature amounts.
[0155] [3] The device according to [1] or [2], wherein the generation unit extracts, from the feature quantities, a value for an index indicating a voice quality with a high priority and, from the reference quantities, a value for an index indicating a voice quality with a high priority, and generates the basis information including comparison information comparing the extracted feature quantities with the extracted reference quantities.
[0156] [4] The device according to [3], wherein the generation unit generates the basis information including impression information regarding an impression of the speaker based on the non-verbal information and the comparison information.
[0157] [5] The device according to any of the above [1] to [4], wherein the acquisition unit acquires a first feature as the feature corresponding to speech at a first time point and a second feature as the feature corresponding to speech at a second time point after the first time point; the estimation unit estimates first non-language information of a speaker who uttered the speech at the first time point and second non-language information of a speaker who uttered the speech at the second time point; the generation unit generates the basis information regarding the second non-language information based on the first feature, the second feature, the reference amount, and the priority; the output unit outputs the basis information regarding the second non-language information; and the basis information regarding the second non-language information includes a comparison result between the first feature and the second feature.
[0158] [6] The device according to [5], wherein the comparison result includes content indicating a transition from the voice characteristics of the speaker at the first time point to the voice characteristics of the speaker at the second time point.
[0159] [7] The device according to any of the above [1] to [6], wherein the estimation unit estimates objective estimation information as the non-language information based on the feature amount and the reference amount which is a statistical amount of a plurality of the speakers, and estimates subjective estimation information as the non-language information based on the feature amount and a threshold value related to the non-language information set for each of the speakers; the generation unit generates objective ground information as the ground information related to the objective estimation information and subjective ground information as the ground information related to the subjective estimation information based on the feature amount, the reference amount, the threshold value and the priority; and the output unit outputs the objective ground information and the subjective ground information.
[0160] [8] The device according to [7], wherein the generation unit generates notification information that presents the objective ground information and the subjective ground information in a format that allows the objective ground information and the subjective ground information to be distinguished from each other.
[0161] [9] The device described in [7] or [8] above, further comprising: a receiving unit that receives subjective information including an index indicating the voice quality of the voice expected by the speaker; a determining unit that determines whether the subjective basis information differs from the subjective information; and an updating unit that updates the threshold when the determining unit determines that the subjective basis information differs from the subjective information.
[0162]
[10] A method comprising: a step of acquiring features related to speech; a step of estimating the non-language information of a speaker uttering the speech based on the features and predetermined reference amounts related to non-language information; a step of generating basis information regarding the basis for estimating the non-language information in the estimating step based on the features, the reference amounts, and priorities corresponding to the features; and a step of outputting the basis information.
[0163] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are connected directly or indirectly (e.g., via wire, wirelessly, etc.) and these multiple devices. The functional block may also be realized by combining the single device or multiple devices with software.
[0164] Functions include, but are not limited to, judgment, determination, assessment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, consideration, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.
[0165] For example, the device 10 (including the device 10A and the device 10B) constituting the conversion system according to an embodiment of the present disclosure may function as a computer that performs processing of the control method of the present disclosure. FIG. 10 is a diagram illustrating an example of the hardware configuration of the device 10 according to an embodiment of the present disclosure. The device 10 described above may be physically configured as a computer including a processor 1001, a memory 1002, a storage device 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, and the like. Note that the device 10 may be configured as a computer including at least one processor such as a CPU or a GPU, or may be configured as a computer including multiple processors or may include multiple computer devices. The terminal 5 and the like may also have a similar hardware configuration.
[0166] In the following description, the term "apparatus" can be interpreted as a circuit, a device, a unit, etc. The hardware configuration of apparatus 10 may be configured to include one or more of the apparatuses shown in the drawings, or may be configured to exclude some of the apparatuses.
[0167] Each function of the device 10 is realized by loading specified software (programs) onto hardware such as the processor 1001 and memory 1002, causing the processor 1001 to perform calculations, control communication via the communication device 1004, and control at least one of reading and writing data in the memory 1002 and storage 1003.
[0168] The processor 1001 controls the entire computer by running, for example, an operating system. The processor 1001 may be configured by a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. For example, the above-mentioned acquisition unit 20, estimation unit 30, generation unit 40, output unit 50, determination unit 60, update unit 70, etc. may be realized by the processor 1001.
[0169] The processor 1001 also reads programs (program code), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes in accordance with the programs. The programs used are programs that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, the acquisition unit 20, the estimation unit 30, the generation unit 40, the output unit 50, the determination unit 60, and the update unit 70 may be implemented by a control program stored in the memory 1002 and running on the processor 1001, and similar implementations may be used for other functional blocks. While the above-described various processes have been described as being executed by a single processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The programs may also be transmitted from a network via a telecommunications line.
[0170] The memory 1002 is a computer-readable recording medium and may be configured, for example, by at least one of a read-only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a random access memory (RAM), etc. The memory 1002 may also be called a register, a cache, a main memory (primary storage device), etc. The memory 1002 can store executable programs (program codes), software modules, etc. for implementing a control method according to an embodiment of the present disclosure.
[0171] Storage 1003 is a computer-readable recording medium, and may be composed of at least one of, for example, an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate medium including at least one of memory 1002 and storage 1003.
[0172] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as, for example, a network device, a network controller, a network card, a communication module, etc. The communication device 1004 may be configured to include a high-frequency switch, a duplexer, a filter, a frequency synthesizer, etc. to realize at least one of frequency division duplex (FDD) and time division duplex (TDD). For example, the above-mentioned acquisition unit 20, estimation unit 30, generation unit 40, output unit 50, determination unit 60, update unit 70, etc. may be realized by the communication device 1004.
[0173] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that accepts input from the outside. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that outputs to the outside. Note that the input device 1005 and the output device 1006 may be integrated into one device (e.g., a touch panel).
[0174] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses between each device.
[0175] The device 10 may also be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware.
[0176] The notification of information is not limited to the aspects / embodiments described in the present disclosure and may be performed using other methods. For example, the notification of information may be performed by physical layer signaling (e.g., Downlink Control Information (DCI) and Uplink Control Information (UCI)), higher layer signaling (e.g., Radio Resource Control (RRC) signaling, Medium Access Control (MAC) signaling, broadcast information (Master Information Block (MIB) and System Information Block (SIB))), other signals, or a combination thereof. Furthermore, the RRC signaling may be referred to as an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.
[0177] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.
[0178] Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.
[0179] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).
[0180] The aspects / embodiments described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to explicit notification, but may be implicit (e.g., not notifying the predetermined information).
[0181] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.
[0182] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0183] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.
[0184] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0185] Note that terms described in this disclosure and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings. For example, at least one of a channel and a symbol may be a signal (signaling). Furthermore, a signal may be a message. Furthermore, a component carrier (CC) may be called a carrier frequency, a cell, a frequency carrier, etc.
[0186] Furthermore, the information, parameters, etc. described in the present disclosure may be expressed using absolute values, may be expressed using relative values from a predetermined value, or may be expressed using other corresponding information. For example, a radio resource may be indicated by an index.
[0187] The names used for the above-described parameters are not intended to be limiting in any way. Furthermore, the mathematical expressions using these parameters may differ from those explicitly disclosed in this disclosure. The various channels (e.g., PUCCH, PDCCH, etc.) and information elements may be identified by any suitable names, and therefore the various names assigned to these various channels and information elements are not intended to be limiting in any way.
[0188] In this disclosure, the terms "Mobile Station (MS)," "user terminal," "User Equipment (UE)," "terminal," and the like may be used interchangeably.
[0189] A mobile station may also be referred to by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or some other suitable terminology.
[0190] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.
[0191] The terms "connected," "coupled," or any variation thereof, refer to any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, "connected" may be read as "access." As used in this disclosure, two elements may be considered to be "connected" or "coupled" to each other using one or more wires, cables, and / or printed electrical connections, as well as electromagnetic energy having wavelengths in the radio frequency range, microwave range, and optical (both visible and invisible) range, as some non-limiting and non-exhaustive examples.
[0192] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."
[0193] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.
[0194] When the terms "include," "including," and variations thereof are used in this disclosure, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, when the term "or" is used in this disclosure, it is not intended to be an exclusive or.
[0195] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.
[0196] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."
[0197] 1, 1A, 1B... processing system, 5, 5A, 5B... terminal, 7... non-verbal information database, 8... past information database, 9... subjective information database, 10, 10A, 10B... device, 20, 20A, 20B... acquisition unit, 21, 21A... speech acquisition unit, 22, 22A... detection unit, 23... priority acquisition unit, 24... reception unit, 30... estimation unit, 31... first machine learning model, 40... generation unit, 41... second machine learning model, 50... output unit, 60... determination unit, 70... update unit, 1001... processor, 1002... memory, 1003... storage, 1004... communication device, 1005... input device, 1006... output device, 1007... bus, MT, MTA, MTB, MTC... processing method.
Claims
1. An apparatus comprising: an acquisition unit that acquires features related to speech; an estimation unit that estimates the non-language information of a speaker uttering the speech based on predetermined reference amounts related to the features and non-language information; a generation unit that generates basis information regarding the basis for the estimation of the non-language information by the estimation unit based on the features, the reference amounts, and priorities corresponding to the features; and an output unit that outputs the basis information.
2. The device described in claim 1, wherein the estimation unit has learned the reference amount in advance and estimates the non-language information by inputting the feature amount into a machine learning model that outputs the non-language information when the feature amount is input.
3. The device described in claim 1, wherein the generation unit extracts values from the features for an index indicating the voice quality with the highest priority and values from the reference quantities for an index indicating the voice quality with the highest priority, and generates the basis information including comparison information comparing the extracted features with the extracted reference quantities.
4. The device according to claim 3, wherein the generation unit generates the basis information including impression information regarding an impression of the speaker based on the non-verbal information and the comparison information.
5. The device described in claim 1, wherein the acquisition unit acquires a first feature as the feature corresponding to the speech at a first time point and a second feature as the feature corresponding to the speech at a second time point after the first time point; the estimation unit estimates first non-language information of a speaker uttering the speech at the first time point and second non-language information of a speaker uttering the speech at the second time point; the generation unit generates the evidence information regarding the second non-language information based on the first feature, the second feature, the reference amount and the priority; the output unit outputs the evidence information regarding the second non-language information; and the evidence information regarding the second non-language information includes a comparison result between the first feature and the second feature.
6. The device of claim 5, wherein the comparison result includes content indicative of a transition from the speaker's voice characteristics at the first time point to the speaker's voice characteristics at the second time point.
7. The device described in claim 1, wherein the estimation unit estimates objective estimation information as the non-language information based on the features and the reference quantity, which is a statistical quantity of the multiple speakers, and estimates subjective estimation information as the non-language information based on the features and a threshold value for the non-language information set for each speaker; the generation unit generates objective ground information as the ground information for the objective estimation information and subjective ground information as the ground information for the subjective estimation information based on the features, the reference quantity, the threshold value, and the priority; and the output unit outputs the objective ground information and the subjective ground information.
8. The device according to claim 7, wherein the generating unit generates notification information that presents the objective ground information and the subjective ground information in a format that allows them to be distinguished from each other.
9. The device according to claim 7, further comprising: a receiving unit that receives subjective information including an index indicating the voice quality of the voice expected by the speaker; a determining unit that determines whether the subjective basis information differs from the subjective information; and an updating unit that updates the threshold value when the determining unit determines that the subjective basis information differs from the subjective information.
10. A method comprising: a step of acquiring features related to speech; a step of estimating the non-language information of a speaker uttering the speech based on the features and predetermined reference amounts related to non-language information; a step of generating basis information regarding the basis for estimating the non-language information in the estimating step based on the features, the reference amounts, and priorities corresponding to the features; and a step of outputting the basis information.
Citation Information
Patent Citations
Interpretable emotion recognition method and system based on global workspace
CN114005468A
Karaoke device
JP2002162978A
Voice evaluation device for evaluating singing with shout technique
JP2014092550A
Singing training device, singing training method and singing training program
JP2016173540A
Emotion analyzer, emotion analysis method and emotion analysis program
JP2021110781A