Impression formation control device, method and program

The impression formation control device addresses the issue of altered speaker intention by determining listener biases and providing external stimuli, ensuring accurate impression control without changing the speaker's message.

JP7722559B2Active Publication Date: 2025-08-13NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024509704
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-25
Publication Date
2025-08-13
Estimated Expiration
2042-03-25

AI Technical Summary

Technical Problem

Existing methods for controlling listener impressions through synthetic speech alter the speaker's intention, leading to miscommunication.

Method used

An impression formation control device that acquires a speaker's speech audio signal, extracts audio features, determines listener biases, generates bias control signals, and provides external stimuli to influence the listener's impression without altering the speaker's intention.

Benefits of technology

Controls listener impressions accurately by using external stimuli, ensuring the speaker's intention is conveyed without alteration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007722559000001
    Figure 0007722559000001
  • Figure 0007722559000002
    Figure 0007722559000002
  • Figure 0007722559000003
    Figure 0007722559000003
Patent Text Reader

Abstract

One aspect of the present invention is configured to: acquire, in control of impression formation in a receiver with respect to a speaker, a spoken voice signal of the speaker to extract a voice feature amount from the spoken voice signal; determine, on the basis of the extracted voice feature amount, a bias to an impression generated in the receiver by the spoken voice signal; generate, on the basis of a determination result of the bias and information indicating a predetermined control direction of the bias, a bias control signal for controlling the bias; and generate and output a stimulus control signal for giving an external stimulus to the receiver in accordance with the bias control signal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] One aspect of the present invention relates to an impression formation control device, method, and program for controlling, for example, the impression formed by a listener with respect to a speaker. [Background technology]

[0002] It is known that a speaker's voice influences the impression formed by the listener, such as the speaker's trustworthiness and likability. For example, Non-Patent Document 1 reports that the pitch of a politician's voice is related to the impression formed of the speaker, the politician. Specifically, it has been reported that the lower the fundamental frequency of the politician's voice, the higher the listener's rating of the politician's likability and trustworthiness (a positive bias occurs in the evaluation), and the higher the fundamental frequency of the politician's voice, the lower the listener's rating of the politician's likability and trustworthiness (a negative bias occurs in the evaluation). Non-Patent Document 1 also describes that it is possible to control the listener's impression of a politician by manipulating the voice using speech synthesis technology. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Yosuke Okada, “The influence of voice pitch on forming impressions of politicians: An experimental study using female voices using voice synthesis software,” Applied Sociology Research, vol. 58, pp. 53-66, 2016. Summary of the Invention [Problem to be solved by the invention]

[0004] However, if synthetic speech is used to control the listener's impression of the speaker, as in Non-Patent Document 1, the speech features are altered, which may result in the speaker's intention not being conveyed correctly to the listener.

[0005] The present invention has been made in light of the above circumstances, and aims to provide a technique that makes it possible to control the impression formed by a listener toward a speaker without altering the speaker's intention. [Means for solving the problem]

[0006] In order to solve the above-mentioned problems, one aspect of the impression formation control device or method of the present invention, when controlling the impression formation of a listener toward a speaker, acquires a speech audio signal of the speaker, extracts audio features from the acquired speech audio signal, determines a bias in the impression created in the listener by the speech audio signal based on the extracted audio features, generates a bias control signal for controlling the bias based on the bias determination result and information indicating a predetermined control direction of the bias, and generates and outputs a stimulus control signal for giving an external stimulus to the listener in accordance with the bias control signal. [Effects of the Invention]

[0007] According to one aspect of the present invention, it is possible to provide a technique that can control the impression formed by a listener regarding a speaker without changing the speaker's intention. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of a system including an impression formation control device according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing an example of a hardware configuration of an impression formation control device according to an embodiment of the present invention. [Figure 3] FIG. 3 is a block diagram showing an example of the software configuration of the impression formation control device according to an embodiment of the present invention. [Figure 4] FIG. 4 is a flowchart showing an example of the processing procedure and processing content of the impression formation control processing executed by the control unit of the impression formation control device shown in FIG. [Figure 5]FIG. 5 is a flowchart showing an example of the processing procedure and processing contents of a first embodiment of the bias determination processing of the processing procedures shown in FIG. [Figure 6] FIG. 6 is a flowchart showing an example of the processing procedure and processing contents of a second embodiment of the bias determination processing of the processing procedures shown in FIG. [Figure 7] FIG. 7 is a flowchart showing an example of the procedure and content of the bias control signal generation process from the procedure shown in FIG. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0010] [One embodiment] (Configuration example) In one embodiment of the present invention, a lecture, seminar, or the like is held via a network.

[0011] (1) System FIG. 1 is a diagram showing an example of the configuration of a system including an impression formation control device SV according to an embodiment of the present invention.

[0012] In one embodiment of the system, for example, a lecturer US1 (hereinafter also referred to as a speaker) at a lecture or seminar uses a lecture terminal TM1 equipped with a microphone MC to transmit a speech voice signal to a participant terminal TM2 used by a participant US2 (hereinafter also referred to as a listener) via a network NW and an impression formation control device SV. The network NW is a wide area network equipped with a public IP network such as the Internet.

[0013] The lecture terminal TM1 and the attendance terminal TM2 are both made up of, for example, personal computers, and these terminals TM1 and TM2 are connected to a network NW via an access network such as a LAN (Local Area Network).

[0014] The terminals TM1 and TM2 may be mobile terminals such as smartphones or tablet terminals, and a wireless LAN or public mobile communication network may be used as the access network. The microphone MC may be an external type or a built-in type relative to the terminal TM1.

[0015] (2) Equipment (2-1) Impression formation control device SV 2 and 3 are block diagrams showing examples of the hardware and software configurations of the impression formation control device SV, respectively.

[0016] The impression formation control device SV is composed of a server computer located on the cloud or the web, for example, and includes a control unit 1 that uses a hardware processor such as a central processing unit (CPU). A storage unit having a program storage unit 2 and a data storage unit 3, and a communication interface (hereinafter, interface will be abbreviated as I / F) unit 4 are connected to this control unit 1 via a bus 5.

[0017] Under the control of the control unit 1, the communication I / F unit 4 uses a communication protocol defined by the network NW to transmit and receive audio data and the like between the lecture terminal TM1 and the attendance terminal TM2.

[0018] The program storage unit 2 is configured, for example, by combining a non-volatile memory such as an HDD (Hard Disk Drive) or SSD (Solid State Drive) as a storage medium that can be written to and read at any time, with a non-volatile memory such as a ROM (Read Only Memory), and stores application programs necessary to execute various control processes according to one embodiment of the present invention, in addition to middleware such as an OS (Operating System).

[0019] The data storage unit 3 is, for example, a combination of a non-volatile memory such as an HDD or SSD as a storage medium that can be written to and read from at any time, and a volatile memory such as RAM (Random Access Memory), and is provided with an audio signal storage unit 31 and a control direction setting information storage unit 32 as storage areas in one embodiment.

[0020] The voice signal storage unit 31 temporarily stores the voice signal of the speaker transmitted from the lecture terminal TM1 for the purpose of impression formation control processing.

[0021] The control direction setting information storage unit 32 stores information for setting a bias control direction when controlling impression formation. Bias represents a physical quantity of the impression that a listener US2 feels toward a speaker US1 giving a lecture, and the control direction setting information is information that defines the control direction of the bias.

[0022] The control unit 1 includes, as processing functions according to an embodiment of the present invention, a speech voice signal acquisition processing unit 11, a voice feature extraction processing unit 12, a bias determination processing unit 13, a bias control signal generation processing unit 14, and a presentation content determination processing unit 15. These processing units 11 to 15 are all realized by causing a hardware processor of the control unit 1 to execute an application program stored in the program storage unit 2.

[0023] Note that part or all of the processing units 11 to 15 may be realized using hardware such as an LSI (Large Scale Integration) or an ASIC (Application Specific Integrated Circuit).

[0024] The speech voice signal acquisition processing unit 11 receives the speech voice signal of the speaker US1 transmitted from the lecture terminal TM1 via the communication I / F unit 4, and temporarily stores the received speech voice signal in the voice signal storage unit 31.

[0025] The speech feature extraction processing unit 12 reads the speech signal from the speech signal storage unit 31 as an input, extracts speech features from the read speech signal, and outputs the extracted speech features. For example, at least one of fundamental frequency, speaking rate, and intonation is extracted as the speech feature.

[0026] The bias determination processing unit 13 receives the speech features of the speech speech signal extracted by the speech feature extraction processing unit 12, determines a bias estimated to occur in the listener US2 based on the speech features of the speech speech signal, and outputs a bias determination result. An example of the determination process will be described in detail in an operation example.

[0027] The bias control signal generation processor 14 receives the bias determination result from the bias determination processor 13 and also reads and receives control direction setting information from the control direction setting information storage unit 32. Based on the bias determination result and the control direction setting information, the bias control signal generation processor 14 generates and outputs a bias control signal for controlling the bias occurring in the listener US2.

[0028] The presentation content determination processor 15 receives the bias control signal generated by the bias control signal generation processor 14 and determines the content of the physical external stimulus to be provided to the listener US2 based on the bias control signal. For example, a tactile stimulus that changes in temperature or hardness is used as the physical external stimulus. An example of the bias control signal generation process will be described in detail in the operation example. Then, the presentation content determination processing unit 15 generates a stimulus control signal corresponding to the content of the external stimulus, and transmits the generated stimulus control signal from the communication I / F unit 4 to the terminal TM2 of the listener US2.

[0029] (2-2) Student's terminal TM2 A presentation device VB for applying a physical external stimulus to the listener US2 is connected to the listening terminal TM2. Examples of the presentation device VB include a mouse with a built-in Peltier element that can display temperature, or an elastic body that can display hardness by expanding and contracting. The listening terminal TM2 receives the stimulus control signal and drives the presentation device VB in response to the stimulus control signal, thereby changing, for example, the temperature or hardness. Changing the temperature or hardness can be expected to have the effect of changing the bias generated in the listener US2.

[0030] In addition to the tactile stimuli such as the changes in temperature or hardness, the physical external stimuli to be given to the listener US2 may include tactile stimuli such as wind pressure or vibration, visual stimuli such as the presence or absence, intensity, or change in the color of light emitted, olfactory stimuli such as the presence or absence of a scent, etc. These physical external stimuli can be given to the listener US2 by using, for example, an electric fan, a vibrator, a display, an aroma diffuser, etc. as the presentation device VB.

[0031] (Example of operation) Next, an example of the operation of the impression formation control device SV configured as above will be described. FIG. 4 is a flowchart showing an example of the processing procedure and processing contents of the impression formation control processing executed by the control unit 1 of the impression formation control device SV.

[0032] (1) Control direction setting Prior to the start of impression formation control, the control unit 1 of the impression formation control device SV performs a process of setting the control direction in response to, for example, an input from a system administrator. The control direction sets the control direction of the bias to be presented to the listener, who is the student, and for example, three types are set: "positive," "negative," and "suppression." The control unit 1 of the impression formation control device SV stores information indicating the set control direction in the control direction setting information storage unit 32. Examples of bias include trust, intimacy, and liking. A positive bias indicates, for example, a direction in which trust deepens in the case of trust, a direction in which intimacy deepens in the case of intimacy, and a direction in which liking increases in the case of liking, while a negative bias indicates the opposite direction. The direction of suppressing bias indicates preventing the bias itself from changing. Note that, although examples of bias will be explained below using trust, intimacy, and liking, these are merely examples of bias. Any bias that is influenced by physical external stimuli is acceptable, and is not limited to trust, intimacy, and liking.

[0033] The control direction may be set by the speaker who is the lecturer, or may be set by the listener who is the student.

[0034] (2) Acquisition of speech signals When attending a lecture or seminar, a listener US2 accesses a URL (Uniform Resource Locator) of a site notified in advance by, for example, the organizer, and a line is then established between, for example, a terminal TM1 for the lecture and a terminal TM2 for the audience via an impression formation control device SV.

[0035] In this state, in step S1, the control unit 1 of the impression formation control device SV determines that the speaker US1, who is the lecturer, has started speaking based on the speech voice signal transmitted from the lecture terminal TM1. Then, when the speech starts, in step S2, under the control of the speech voice signal acquisition processing unit 11, the control unit 1 of the impression formation control device SV receives the speech voice signal of the speaker US1 transmitted from the lecture terminal TM1 thereafter via the communication I / F unit 4, and temporarily stores the received speech voice signal in the voice signal storage unit 31.

[0036] The timing of acquiring the speech voice signal may be set arbitrarily, and the acquisition time length may be set to any length as long as speech features can be extracted. For example, the timing of acquiring the speech voice signal may be a predetermined time determined based on the time it takes for a speaker to switch or for a listener to form an impression of a speaker, or the acquisition time length of the speech voice signal may be set to a time determined based on the time required for a listener to estimate an impression of a speaker, for example, about 10 seconds.

[0037] Furthermore, the speech audio signal may be acquired once, or may be acquired multiple times for a predetermined length during a lecture. If the speech audio signal is acquired multiple times periodically, even if the audio feature of the speech audio signal of speaker US1 changes during the lecture and the impression given to the listener changes, it becomes possible to readjust the bias presented to listener US2 in accordance with this change in impression.

[0038] (3) Extraction of speech features When the speech voice signal is acquired, in step S3, the control unit 1 of the impression formation control device SV reads the speech voice signal from the voice signal storage unit 31 under the control of the voice feature extraction processing unit 12, and extracts voice features from the read speech voice signal. As the voice feature, for example, at least one of fundamental frequency, speaking rate, and intonation is extracted.

[0039] The method for extracting speech features may be, for example, a well-known method described in Reference 1 below, but is not limited to the method described in Reference 1. [References 1] F. Eyben, M. Wo¨llmer, and B. Schuller, “OpenSMILE - The Munich versatile and fast open-source audio feature extractor,” MM'10-Proc. ACM Multimed. 2010 Int. Conf., pp. 1459-1462, 2010.

[0040] (4) Bias judgment Next, in step S4, the control unit 1 of the impression formation control device SV, under the control of the bias determination processing unit 13, determines the bias that is estimated to occur in the listener US2 due to the speech signal of the speaker US1 based on the above speech features.

[0041] The bias determination method used depends on the type of speech feature, as follows: (4-1) In the case of fundamental frequency FIG. 5 is a flowchart showing an example of the procedure and content of the bias determination process executed by the bias determination processor 13 when the audio feature is the "fundamental frequency."

[0042] First, in step S411, the bias determination processing unit 13 receives the "fundamental frequency" extracted as a speech feature from the speech feature extraction processing unit 12, and determines the bias that is estimated to occur in the listener US2 based on the level of this fundamental frequency.

[0043] Here, the relationship between the high and low fundamental frequency and bias is known to be such that, as exemplified in Non-Patent Document 1, low voices tend to receive higher ratings in terms of "trustworthiness," "intimacy," and "likability," while high voices tend to receive lower ratings.

[0044] Therefore, in step S412, the bias determination processing unit 13 determines whether the fundamental frequency f B is, for example, 300 Hz or less, and in step S414, the fundamental frequency f B is, for example, 600 Hz or more.

[0045] As a result of the above judgment, the fundamental frequency f B If the fundamental frequency f is 300 Hz or less, the bias determination processing unit 13 determines in step S413 that the bias estimated to occur in the listener US2 is "positive." B If the fundamental frequency f is 600 Hz or more, the bias determination processing unit 13 determines in step S415 that the bias estimated to occur in the listener US2 is "negative." Bis higher than 300 Hz and lower than 600 Hz, the bias determination processing unit 13 determines in step S416 that no bias occurs in the listener US2, that is, that there is "no bias."

[0046] The evaluation of "trustworthiness," "intimacy," and "likability" differs depending on the relationship between the speaker US1 and the listener US2. B It is desirable that the threshold value for determining whether the frequency is high or low be set arbitrarily, not limited to the above 300 Hz or 600 Hz.

[0047] Finally, in step S417, the bias determination processing unit 13 outputs the determination result obtained in step S413, S415, or S416 to the bias control signal generation processing unit .

[0048] (4-2) In the case of speaking speed FIG. 6 is a flowchart showing an example of the procedure and content of the bias determination process executed by the bias determination processor 13 when the speech feature is "speaking rate."

[0049] First, in step S421, the bias determination processing unit 13 receives the "speech rate" extracted as a speech feature from the speech feature extraction processing unit 12, and determines the bias that is estimated to occur in the listener US2 based on whether the speech rate is fast or slow. It is well known that there is a relationship between speaking speed and bias. For example, findings such as a tendency for relatively fast speaking speeds to be rated as more "extroverted," and conversely, a tendency for slow speaking speeds to be rated as less "extroverted" are described in Reference 2 below.

[0050] [Reference 2] Teruhisa Uchida, “The effect of speech rate on speaker personality impressions,” Psychological Research, vol. 73, no. 2, pp. 131-139, 2002.

[0051] Here, for example, if the bias is "trustworthiness," "intimacy," or "likability," then if the speaking speed is relatively fast, the ratings of "trustworthiness," "intimacy," and "likability" will be high, and conversely, if the speaking speed is slow, the ratings of "trustworthiness," "intimacy," and "likability" will be low.

[0052] Therefore, in step S422, the bias determination processing unit 13 determines whether the speech rate is, for example, 10.8 mora / sec or more, and in step S424, determines whether the speech rate is, for example, 6.96 mora / sec or less. Note that a mora is a unit representing the number of kana, long consonants, double consonants, and nasal consonants in the Japanese syllabary.

[0053] If the result of the above determination is that the speaking rate is 10.8 mora / sec or higher, bias determination processor 13 determines in step S423 that the bias estimated to occur in listener US2 is "positive." On the other hand, if the speaking rate is 6.96 mora / sec or lower, bias determination processor 13 determines in step S425 that the bias estimated to occur in listener US2 is "negative." Note that if the speaking rate is higher than 6.96 mora / sec but lower than 10.8 mora / sec, bias determination processor 13 determines in step S426 that there is "no bias."

[0054] In this case, too, the "trustworthiness," "intimacy," and "likability" will differ depending on the relationship between the speaker US1 and the listener US2. For this reason, it is desirable to be able to set the threshold value for determining the speech rate arbitrarily, without being limited to the above 10.8 mora / sec and 6.96 mora / sec.

[0055] Finally, in step S427, the bias determination processing unit 13 outputs the determination result obtained in step S423, S425 or S426 to the bias control signal generation processing unit .

[0056] (4-3) Inflection It is generally known that there is a relationship between the level of intonation in speech and bias. For example, findings such as a tendency for a stronger intonation to be rated as more "extroverted," and conversely, a weaker intonation to be rated as less "extroverted" are described in Reference 3 below.

[0057] [Reference 3] Teruhisa Uchida, “The influence of intonation and its change pattern on speaker personality impressions,” Psychological Research, vol. 76, no. 4, pp. 382-390, 2005.

[0058] If the above biases are "trustworthiness," "intimacy," and "likability," then a strong intonation will result in a high evaluation of "trustworthiness," "intimacy," and "likability," and conversely, a weak intonation will result in a low evaluation of "trustworthiness," "intimacy," and "likability."

[0059] Therefore, similar to the fundamental frequency determination process procedure shown in Fig. 5, the bias determination processor 13 determines whether the standard deviation of the fundamental frequency, which represents "inflection," is, for example, 40 Hz or more, and also determines whether the standard deviation of the fundamental frequency is, for example, 20 Hz or less. If the result of this determination is that the standard deviation of the fundamental frequency is, for example, 40 Hz or more, the bias determination processor 13 determines that the bias estimated to occur in the listener US2 is "positive." On the other hand, if the standard deviation of the fundamental frequency is, for example, 20 Hz or less, the bias determination processor 13 determines that the bias estimated to occur in the listener US2 is "negative." Note that if the standard deviation of the fundamental frequency is higher than the above-mentioned 20 Hz but lower than the above-mentioned 40 Hz, the bias determination processor 13 determines that the listener has "no bias."

[0060] In this case, too, the evaluations of "trustworthiness," "intimacy," and "likability" differ depending on the relationship between speaker US1 and listener US2, so the threshold value for determining the standard deviation of the fundamental frequency may be set arbitrarily, not limited to the above 40 Hz or 20 Hz.

[0061] Finally, the bias determination processing unit 13 outputs the determination result to the bias control signal generation processing unit 14 .

[0062] (5) Generation of bias control signals Next, in step S5, the control unit 1 of the impression formation control device SV, under the control of the bias control signal generation processing unit 14, executes the process of generating a bias control signal for presenting a physical external stimulus to the listener US2 as follows.

[0063] FIG. 7 is a flowchart showing an example of the procedure and content of the bias control signal generation process executed by the bias control signal generation processor 14.

[0064] The bias control signal generation processing unit 14 first reads the control direction setting information from the control direction setting information storage unit 32 in step S51, and receives the bias determination result from the bias determination processing unit 13 in step S52.

[0065] Next, in step S53, the bias control signal generation processing unit 14 determines whether the control direction setting information read is "positive," "negative," or "inhibit." If the result of this determination is "positive," in step S54, the bias control signal "positive" is generated to generate a "positive" bias in the listener US2.

[0066] The bias control signal "positive" has the function of generating an external stimulus to further enhance the "positive" bias when the determination result of the bias determined from the speech features is "positive." This is expected to have the effect of amplifying the "positive" bias occurring in the listener US2. Furthermore, the bias control signal "positive" has the function of generating an external stimulus to counteract the "negative" bias when the determination result of the bias determined from the speech features is "negative." This is expected to have the effect of causing a "positive" bias in the listener US2. Furthermore, the bias control signal "positive" has the function of generating an external stimulus to cause a "positive" bias in the listener US2 when the determination result of the bias determined from the speech features is "absent."

[0067] On the other hand, if the control direction is determined to be negative in step S53, the bias control signal generation processor 14 generates a bias control signal "negative" in step S55 to apply a negative bias to the listener US2.

[0068] The bias control signal "negative" has a function of generating an external stimulus to counteract the "positive" bias when the bias determination result from the speech features is "positive." This is expected to have the effect of generating a "negative" bias in the listener US2. Furthermore, the bias control signal "negative" has a function of generating an external stimulus to further enhance the "negative" bias when the bias determination result from the speech features is "negative." This is expected to have the effect of amplifying the "negative" bias in the listener US2. Furthermore, the bias control signal "negative" has a function of generating an external stimulus to generate a "negative" bias in the listener US2 when the bias determination result from the speech features is "absent."

[0069] Finally, suppose that the control direction determined in step S53 is "suppression." In this case, in step S56, the bias control signal generation processor 14 determines whether the bias determined from the audio features is "positive," "negative," or "none."

[0070] As a result of this determination, it is assumed that the bias determined from the speech features is "positive." In this case, in step S57, the bias control signal generation processor 14 generates a bias control signal "n-negative" for canceling the "positive" bias for the listener US2. This bias control signal "n-negative" is a signal for generating an external stimulus for changing the bias of the listener US2 in the "neutral" direction.

[0071] On the other hand, suppose that the bias determined from the speech features in step S56 is negative. In this case, in step S58, the bias control signal generation processor 14 generates a bias control signal "n-positive" for changing the negative bias determined from the speech features to a positive direction. This bias control signal "n-positive" is a signal for generating an external stimulus to bias the listener US2 in a positive direction. By generating an external stimulus to bias the listener US2 in a positive direction, it is expected that the bias of the listener US2 will be changed to a neutral direction.

[0072] Also, suppose that the result of the determination in step S56 above is that the bias determined from the audio feature is “none.” In this case, the bias control signal generation processing unit 14 does not generate a bias control signal and ends the bias control signal generation process.

[0073] (6) Deciding what to present and sending stimulus control signals Finally, in step S6, the control unit 1 of the impression formation control device SV, under the control of the presentation content determination processing unit 15, determines the content of the external stimulus to be presented to the listener US2 and executes the process of transmitting a stimulus control signal as follows.

[0074] That is, the presentation content determination processing unit 15 receives the bias control signal from the bias control signal generation processing unit 14, and determines the content of the external stimulus to be given to the listener US2 based on the received bias control signal. Then, in accordance with the determined content of the external stimulus, it generates a stimulus control signal for operating the presentation device VB.

[0075] (6-1) When using "temperature" as an external stimulus In general, people tend to perceive their relationship with an acquaintance or experimenter as "closer" when they hold a warm object in their hands or when the room is warm, compared to when they hold a cold object or when the room is cold. This finding is reported, for example, in Reference 4.

[0076] [Reference 4] H. Ijzerman and GR Semin, “The thermometer of social relations: Mapping social proximity on temperature: Research article,” Psychol. Sci., vol. 20, no. 10, pp. 1214-1220, 2009.

[0077] Therefore, for example, a mouse with a built-in Peltier element capable of presenting temperature is used as presentation device VB. In this case, presentation content determination processing unit 15 determines the content of the external stimulus to be presented as follows.

[0078] (1) When the bias control signal is "positive," the content to be presented is set to "40 degrees," a temperature that people generally perceive as warm. (2) When the bias control signal is “n-positive,” the displayed content is set to “35 degrees,” which is a temperature lower than the “positive” case and which humans perceive as warm. (3) When the bias control signal is "n-negative," the displayed content is set to "30 degrees," which is a temperature lower than that in the "n-positive" case and which humans perceive as cold. (4) When the bias control signal is "negative," the content to be presented is set to "25 degrees," which is a temperature lower than the above "n-negative" and which humans perceive as cold.

[0079] The temperature presentation content is not limited to the above example, and may be set arbitrarily according to individual differences in temperature sensation among the listener US2.

[0080] The presentation content determination processing unit 15 generates a stimulus control signal for causing the presentation device VB to generate the temperature of the presentation content that has been determined. Then, the presentation content determination processing unit 15 transmits the generated stimulus control signal from the communication I / F unit 4 to the attendance terminal TM2 used by the listener.

[0081] When the lecture terminal TM2 receives the stimulus control signal, it drives the presentation device VB in accordance with the received stimulus control signal to generate the temperature specified by the stimulus control signal. Therefore, if the listener US2 is holding a mouse as the presentation device VB at this time, the listener US2 can be given the external stimulus of "temperature," which is expected to have the effect of controlling the listener US2's impression of the speaker US1, who is the lecturer.

[0082] (6-2) When using "hardness" as an external stimulus In general, people tend to evaluate others as harsher and less emotional when they touch a hard object than when they touch a soft object. This finding is reported, for example, in Reference 5 below.

[0083] [Reference 5] JM Ackerman, CC Nocera, and JA Bargh, “Incidental Haptic Sensations Influence Social Judgments and Decisions,” Science (80-. )., vol. 328, no. 5986, pp. 1712-1715, Jun. 2010.

[0084] Therefore, for example, a balloon whose hardness changes when pressure is applied is used as the presentation device VB. A device that uses this balloon to present hardness is shown, for example, in Reference 6 below. Note that instead of a balloon, it is also possible to use an elastic body that can present hardness by expanding and contracting, for example.

[0085] [Reference 6] Mana Sasagawa, et al. “Proposal of a texture presentation system capable of displaying hardness and shape by jamming transition,” Transactions of the Information Processing Society of Japan, vol. 60, no. 2, pp. 376-384, 2019.

[0086] When the balloon is used as the presentation device VB, the presentation content determination processing unit 15 determines the presentation content of the external stimulus as follows.

[0087] (1) When the bias control signal is "positive," the presentation content is set to "-10 kPa," which is a hardness that people generally perceive as soft. (2) When the bias control signal is "n-positive," the presentation content is set to "-30 kPa," which is a hardness that is lower than that in the "positive" case and that humans perceive as hard. (3) When the bias control signal is "n-negative," the presentation content is set to "-50 kPa," which is a hardness that is lower than that in the "n-positive" case and that humans perceive as hard. (4) When the bias control signal is "negative," the presentation content is set to "-70 kPa," which is a hardness that is lower than the above-mentioned "n-negative" and that humans perceive as hard.

[0088] The content of the hardness presentation is not limited to the above example, and may be set arbitrarily according to individual differences in how the listener US2 feels about hardness.

[0089] The presentation content determination processing unit 15 generates a stimulus control signal for causing the presentation device VB to generate the determined hardness of the presentation content. Then, the presentation content determination processing unit 15 transmits the generated stimulus control signal from the communication I / F unit 4 to the attendance terminal TM2 used by the listener.

[0090] When the student terminal TM2 receives the stimulus control signal, it drives the presentation device VB in accordance with the received stimulus control signal to generate the hardness specified by the stimulus control signal. Therefore, if the listener US2 is holding a balloon as the presentation device VB at this time, the listener US2 can be given an external stimulus of the "hardness," which is expected to have the effect of controlling the listener US2's impression of the speaker US1, who is the lecturer.

[0091] (Actions and Effects) As described above, in one embodiment, the impression formation control device SV first acquires a speech signal from a speaker US1 (lecturer) and extracts speech features from the speech signal, and then determines a bias that is likely to occur in a listener US2 (a student) based on the extracted speech features. Next, based on the bias determination result and information indicating a preset bias control direction, a bias control signal for presenting a physical external stimulus to the listener US2 is generated. The content of the external stimulus to be presented to the listener US2 is determined based on the generated bias control signal, and a stimulus control signal corresponding to the content of the external stimulus is transmitted to the terminal TM2 of the listener US2. The stimulus control signal then drives a presentation device VB to present the physical external stimulus to the listener US2 as a bias, thereby changing the listener US2's impression of the speaker US1.

[0092] Therefore, even if a listener PS2, who is a student, has a negative impression of a speaker PS1, who is a lecturer, due to the speech signal of the speaker, PS1, it is possible to cancel or alleviate the negative impression by applying a bias to the listener PS2 by an external stimulus such as temperature or hardness. Furthermore, since the speech features of the speech signal emitted by the speaker PS1 are not altered, it is possible to accurately convey the intention of the speaker PS1 to the listener PS2.

[0093] [Other embodiments] (1) In the above embodiment, a case where a listener PS2, who is a student, attends a lecture by a speaker PS1, who is a lecturer, via a network has been described as an example. However, this is not limited to this case, and the present invention can also be applied to a case where a listener PS2 attends a lecture by a speaker PS1 in person. This case can also be implemented using a configuration similar to that of the above embodiment.

[0094] For example, a speech signal of speaker PS1 is transmitted from lecture terminal TM1 to impression formation control device SV. Then, the impression formation control device SV determines a bias that is estimated to occur in listener PS2 from the speech feature of the speech signal, generates a control signal for controlling the bias of listener PS2 based on the determination result, and transmits it to attendance terminal TM2. Terminal TM2 drives presentation device VB in accordance with the control signal to provide listener PS2 with an external stimulus, thereby expected to control the bias occurring in listener PS2.

[0095] (2) In the above embodiment, the processing functions of the impression formation control device SV are provided in a server computer located on the cloud or the web. However, this is not limiting, and the processing functions of the impression formation control device SV may be provided in, for example, the terminal TM1 for lectures or the terminal TM2 for attendance. Furthermore, the processing functions of the impression formation control device SV may be distributed among the terminal TM1 for lectures, the terminal TM2 for attendance, and a server computer located on the cloud or the web.

[0096] (3) In addition, various modifications can be made to the configuration of the impression formation control device SV, the processing procedure and processing content, the timing of the generation of external stimuli, the types of external stimuli and the means for presenting them, the usage scenarios of the impression formation control device SV, etc., without departing from the gist of this invention.

[0097] Although the embodiments of the present invention have been described in detail above, the above description is merely an example of the present invention in every respect. It goes without saying that various improvements and modifications can be made without departing from the scope of the present invention. In other words, when implementing the present invention, specific configurations according to the embodiments may be appropriately adopted.

[0098] In short, this invention is not limited to the above-described embodiments, and the components can be modified and embodied in practice without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, some components may be omitted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined. [Explanation of symbols]

[0099] SV...impression formation control device US1...Speaker US2…Receiver TM1: Terminal for lectures TM2: Device used for taking classes NW...Network MC: Microphone VB…presentation device 1...Control unit 2...Program memory section 3...Data storage unit 4...Communication I / F section 5...Bus 11... Speech signal acquisition processing unit 12...Speech feature extraction processing unit 13...Bias determination processing unit 14...Bias control signal generation processing section 15... Presentation content determination processing unit 31...Audio signal storage unit 32...Control direction setting information storage unit

Claims

1. An impression formation control device that controls the impression formation of a listener with respect to a speaker, a first processing unit that acquires a speech signal of the speaker; a second processing unit that extracts speech features from the speech signal; a third processing unit that determines a bias in an impression that the speech signal causes to the listener based on the speech feature; a fourth processing unit that generates a bias control signal for controlling the bias based on the bias determination result and information indicating a preset control direction of the bias; a fifth processing unit that generates a stimulus control signal for providing an external stimulus to the listener according to the bias control signal and outputs the generated stimulus control signal; An impression formation control device comprising:

2. the second processing unit extracts at least one of a fundamental frequency, a speaking rate, and an intonation from the speech audio signal as the audio feature; the third processing unit compares the extracted speech feature with a predetermined judgment condition, and judges the bias occurring in the listener based on the comparison result. The impression formation control device according to claim 1 .

3. the fourth processing unit generates the bias control signal that specifies a control direction and a control amount of the temperature based on a result of the bias determination and information that indicates a preset control direction of the bias when the external stimulus is temperature; the fifth processing unit generates the stimulus control signal for applying the external stimulus caused by the change in temperature to the listener in accordance with the bias control signal, and outputs the generated stimulus control signal. The impression formation control device according to claim 1 .

4. The fourth processing unit generates the bias control signal that specifies a control direction and a control amount of the hardness based on the determination result of the bias and information indicating a control direction of the bias that is set in advance when the external stimulus is hardness; The fifth processing unit generates the stimulus control signal for giving the listener the external stimulus by changing the hardness according to the bias control signal, and outputs the generated stimulus control signal. The impression formation control device according to claim 1 .

5. An impression formation control method executed by an information processing device for controlling a listener's impression formation regarding a speaker, comprising: acquiring a speech signal from the speaker; extracting speech features from the speech signal; determining a bias in an impression caused to the listener by the speech signal based on the speech feature; generating a bias control signal for controlling the bias based on the bias determination result and information indicating a preset control direction of the bias; generating a stimulus control signal for providing an external stimulus to the listener according to the bias control signal, and outputting the generated stimulus control signal; An impression formation control method comprising:

6. 5. The impression formation control device according to claim 1, wherein the program causes a processor included in the impression formation control device to execute at least one of the processes of the first processing unit to the fifth processing unit.

Citation Information

Patent Citations

  • Voice quality difference evaluation table generating device, voice quality difference evaluation table generation system for speech corpus, and speech synthesis system

    JP2005070214A

  • Speech synthesizer, speech processor, and program

    JP2006330060A

  • Biofeedback system for correction of nasality

    US20100235170A1

  • Information processing device, information processing method, and program

    WO2019087646A1