Paralanguage information recognition device, paralanguage information recognition method, and paralanguage information recognition program

The paralinguistic information recognition device extracts and utilizes predetermined keywords in media data to estimate user emotions and states without prior neutral speech data, enhancing recognition accuracy in call centers and other applications.

JP2025175816APending Publication Date: 2025-12-03KK TOSHIBA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024082094
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-20
Publication Date
2025-12-03

AI Technical Summary

Technical Problem

Existing paralinguistic information recognition technologies require advance acquisition of neutral speech data of the same speaker, which is difficult to obtain in practical applications like call centers.

Method used

A paralinguistic information recognition device and method that acquires media data of a target user, extracts portions containing predetermined keywords as reference data, and estimates paralinguistic information based on these references without prior neutral speech data.

Benefits of technology

Enables accurate recognition of paralinguistic information without needing advance preparation of neutral speech data, improving recognition performance by using keywords like greetings and proper nouns as neutral speech criteria.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025175816000001_ABST
    Figure 2025175816000001_ABST
Patent Text Reader

Abstract

To provide technology enabling recognition of a user's paralanguage information without requiring prior preparation of the user's media data used as a standard for paralanguage information recognition.SOLUTION: A paralanguage information recognition program according to one embodiment functions a computer as: acquisition means for acquiring media data of a target user; extraction means for extracting portions containing predetermined keywords from the media data as standard media data; and estimation means for estimating the target user's paralanguage information from the media data with reference to the standard media data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] An embodiment of the present invention relates to a paralinguistic information recognition device, a paralinguistic information recognition method, and a paralinguistic information recognition program. [Background technology]

[0002] Paralinguistic information recognition is a technology that recognizes a user's paralinguistic information (e.g., a user's emotions or mental state) from media data related to the user (e.g., audio data or video data). Paralinguistic information recognition is applied to estimating customer satisfaction in call centers, recognizing customer emotions, and estimating worker fatigue levels.

[0003] Emotion recognition, an example of paralinguistic information recognition, is susceptible to the influence of the user's speaking style, and there is a problem that recognition performance is easily degraded. Therefore, in recent years, emotion recognition methods have been devised that take into account the characteristics of the user's speaking style by using the user's neutral voice (specifically, the voice uttered by the user in a neutral state) as a reference.

[0004] However, the above-mentioned emotion recognition requires the acquisition of neutral speech data including the user's neutral speech in advance. In practical applications such as call centers, it is difficult to acquire neutral speech data of the same speaker as the speaker whose speech data is to be recognized in advance. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Patent No. 7420211 Summary of the Invention [Problem to be solved by the invention]

[0006] The problem to be solved by the present invention is to provide a paralinguistic information recognition device, a paralinguistic information recognition method, and a paralinguistic information recognition program that enable recognition of a user's paralinguistic information without the need to prepare in advance the user's media data to be used as a basis for paralinguistic information recognition. [Means for solving the problem]

[0007] A paralinguistic information recognition program according to one embodiment causes a computer to function as an acquisition means for acquiring media data of a target user, an extraction means for extracting a portion of the media data containing a predetermined keyword as reference media data, and an estimation means for estimating the paralinguistic information of the target user from the media data based on the reference media data. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram showing the functional configuration of a paralinguistic information recognition device according to an embodiment. [Figure 2] 2 is a flowchart showing the operation of a keyword detection unit shown in FIG. 1; [Figure 3] 2 is a flowchart showing the operation of the data extraction unit shown in FIG. [Figure 4] 2 is a flowchart showing the operation of the estimation unit shown in FIG. 1; [Figure 5] FIG. 2 is a block diagram showing the hardware configuration of a computer that can realize the paralinguistic information recognition device shown in FIG. [Figure 6] 10 is a flowchart showing a paralinguistic information recognition process in offline operation according to the embodiment. [Figure 7] 10 is a flowchart showing a paralinguistic information recognition process in real time according to an embodiment. [Figure 8] FIG. 10 is a diagram showing an output of the estimation result of paralinguistic information according to the embodiment. [Figure 9] 10A and 10B are diagrams showing the results of an evaluation experiment performed on the method according to the embodiment and a baseline method. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, an embodiment will be described with reference to the drawings.

[0010] The embodiments relate to a technology for estimating a user's paralinguistic information from media data related to the user. In this specification, paralinguistic information indicates a person's mental and / or physical state. Examples of mental states may include, but are not limited to, emotions, interests, concerns, and mental fatigue. Examples of physical states may include, but are not limited to, physical fatigue. The media data may be audio data, video data, or text data. The media data used to estimate the paralinguistic information may include one type of media data or multiple types of media data, such as a combination of audio data and video data. Hereinafter, audio data may be simply referred to as audio, and video data may be simply referred to as video. Furthermore, a user whose paralinguistic information is to be estimated is referred to as a target user.

[0011] First Embodiment In the first embodiment, a case where the media data is audio data will be described.

[0012] Fig. 1 schematically illustrates a paralinguistic information recognition device 100 according to a first embodiment. As illustrated in Fig. 1, the paralinguistic information recognition device 100 includes an acquisition unit 110, an extraction unit 120, an estimation unit 130, and an output unit 140. The paralinguistic information recognition device 100 can be implemented in a computer such as a server or a personal computer.

[0013] The acquisition unit 110 acquires voice data of the target user. The voice data of the target user is voice data containing the voice of the target user. For example, the voice data is acquired by an external media data acquisition device and input to the paralinguistic information recognition device 100. When the input voice data contains voices of multiple users, the acquisition unit 110 obtains the voice data of the target user by extracting portions containing the voice of the target user from the input voice data. In other words, the acquisition unit 110 obtains the voice data of the target user by removing portions containing the voices of users other than the target user from the input voice data.

[0014] The extraction unit 120 receives the speech data of the target user from the acquisition unit 110 and extracts a portion including a predetermined keyword from the received speech data. Typically, a plurality of predetermined keywords are set, and the extraction unit 120 extracts one or more portions including any of the predetermined keywords from the speech data of the target user.

[0015] The predetermined keywords are expressions (e.g., words or phrases) corresponding to greetings, backchannels, replies, or proper nouns. Examples of expressions corresponding to greetings include, but are not limited to, "Good morning" and "Hello." Examples of expressions corresponding to backchannels include, but are not limited to, "I see," "Of course," and the like. Examples of expressions corresponding to replies include, but are not limited to, "Yes" and "No." Examples of proper nouns include, but are not limited to, personal names, product names, and identification information. Identification information may be, for example, a customer's identification number (ID) or a product model number. Generally, when a user utters expressions corresponding to greetings, backchannels, replies, or proper nouns, they tend to speak in a relatively readable tone. Therefore, it is believed that utterances of expressions corresponding to greetings, backchannels, replies, or proper nouns are less likely to contain specific emotions. Therefore, portions containing the predetermined keywords can be considered as neutral speech data. The neutral speech data includes speech uttered by a user in a calm state (hereinafter referred to as "calm speech") and can be used as a criterion for recognizing the user's paralinguistic information.

[0016] The extraction unit 120 extracts a portion including a predetermined keyword from the speech data of the target user as reference speech data.

[0017] The estimation unit 130 estimates the paralinguistic information of the target user from the voice data of the target user acquired by the acquisition unit 110, using the reference voice data extracted by the extraction unit 120 as a reference.

[0018] The output unit 140 outputs the estimation result of the paralinguistic information generated by the estimation unit 130. For example, the output unit 140 displays the estimation result of the paralinguistic information on a display device.

[0019] Next, the extraction unit 120 will be described in more detail.

[0020] The extraction unit 120 includes a keyword detection unit 121 and a data extraction unit 122. The keyword detection unit 121 performs keyword detection on the target user's voice data and generates keyword information indicating a voice section corresponding to a predetermined keyword. The voice section indicates a time section in which voice exists. The keyword information includes, for example, time information for identifying a voice section corresponding to the predetermined keyword. The data extraction unit 122 extracts a portion including the predetermined keyword from the target user's voice data as reference voice data based on the keyword information generated by the keyword detection unit 121.

[0021] FIG. 2 schematically illustrates an example of the operation of the keyword detection unit 121. In step S201 of FIG. 2, the keyword detection unit 121 divides the target user's voice data into utterance units. Here, an utterance refers to a voice section sandwiched between silent sections of a certain time period (typically 0.2 seconds) or more. Specifically, the keyword detection unit 121 first detects a voice section by voice section detection. Voice section detection can be achieved by known techniques such as those described in Reference 1. Next, the keyword detection unit 121 divides the target user's voice data into utterance units based on the detected voice sections. Here, it is assumed that the target user's voice data is divided into multiple utterances. (Reference 1) J. Sohn, NS Kim and W. Sung, "A statistical model-based voice activity detection," in IEEE Signal Processing Letters, vol. 6, no. 1, pp. 1-3, Jan. 1999, doi: 10.1109 / 97.736233.

[0022] In step S202, the keyword detection unit 121 determines whether or not a predetermined keyword is included in each utterance.

[0023] A method for detecting a predetermined keyword can be, for example, a method in which specific words are registered in advance as keywords in a keyword database. In this method, the predetermined keyword is a word registered in the keyword database. A method for determining whether a predetermined keyword is included in an utterance can be, for example, a method in which the utterance is converted into text using speech recognition and whether the text contains the predetermined keyword using text matching. The text obtained by converting an utterance using speech recognition is sometimes called a speech transcript. Alternatively, a keyword spotting model configured to detect an utterance containing a predetermined keyword from speech can be used. The keyword spotting model can be realized using a known keyword detection technology such as that described in Reference 2. Alternatively, an event detection model configured to detect words or utterances expressing a predetermined keyword from speech can be used. The event detection model can be realized using a known event detection technology such as that described in Reference 3. (Reference 2) Zhang, Ao, et al. "U2-KWS: Unified Two-Pass Open-Vocabulary Keyword Spotting with Keyword Bias." 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023. (Reference 3) Keisuke Imoto. "Analysis of Acoustic Events and Acoustic Scenes." Journal of the Acoustical Society of Japan 74.4 (2018): 198-207.

[0024] It is also possible to use a method that does not use a keyword database. For example, a model constructed to extract expressions corresponding to predetermined keywords, such as that disclosed in Reference 4, may be used. For example, the model may be constructed to extract expressions corresponding to greetings ("Good morning" and "Hello" in the above example) from text such as "Good morning. It's nice weather, isn't it?" or "Hello. It's been a while." It may also be constructed to extract proper nouns such as customer IDs and product names. By using such a model, it is possible to detect words or utterances that express predetermined keywords without specifying specific words. However, the method of detecting keywords is not limited to these. (Reference 4) Pakhale, Kalyani. "Comprehensive overview of named Entity Recognition: Models, Domain-Specific applications and challenges." arXiv preprint arXiv:2309.14084 (2023).

[0025] If the predetermined keyword is included in the utterance (step S202; Yes), the keyword detection unit 121 estimates the speech section corresponding to the predetermined keyword and generates keyword information indicating the speech section corresponding to the predetermined keyword (step S203). The estimation of the speech section corresponding to the predetermined keyword can be realized by known techniques such as those described in Reference 5. (Reference 5) Bain, Max, et al. "WhisperX: Time-accurate speech transcription of long-form audio." arXiv preprint arXiv:2303.00747 (2023).

[0026] If the predetermined keyword is not included in the utterance (step S202; No), the keyword detection unit 121 acquires the utterance as speech data to be recognized (step S204).

[0027] FIG. 3 schematically illustrates an example of the operation of the data extraction unit 122. In step S301 of FIG. 3, the data extraction unit 122 acquires keyword information from the keyword detection unit 121. In step S302, the data extraction unit 122 extracts a speech interval including a predetermined keyword based on the keyword information acquired in step S301, and acquires the extracted speech interval as reference speech data for the target user. In one example, the data extraction unit 122 acquires the speech interval itself corresponding to the predetermined keyword as the reference speech data. In another example, the data extraction unit 122 acquires a speech interval including the speech interval corresponding to the predetermined keyword and speech intervals of a predetermined time before and after the speech interval corresponding to the predetermined keyword as the reference speech data. Specifically, the data extraction unit 122 acquires the speech interval from a time point a predetermined time before the speech interval corresponding to the predetermined keyword to a time point a predetermined time after the speech interval corresponding to the predetermined keyword as the reference speech data. In yet another example, the data extraction unit 122 acquires the entire utterance including the predetermined keyword as the reference speech data.

[0028] Next, the estimation unit 130 will be described in more detail.

[0029] 4 shows an outline of an example of the operation of the estimation unit 130. In step S401 of FIG.

[0030] If there is no reference speech data (step S402; No), the flow proceeds to step S404. In step S404, the estimation unit 130 estimates the paralinguistic information of the target user from the speech data to be recognized. For example, the estimation unit 130 estimates the paralinguistic information of the target user for the speech data to be recognized using a general-purpose paralinguistic information recognition model prepared in advance. The general-purpose paralinguistic information recognition model receives speech data as input and is trained in advance to estimate the paralinguistic information of the target user when uttering the speech (utterance) included in the speech data. The estimation unit 130 inputs the speech data to be recognized into the general-purpose paralinguistic information recognition model and obtains information indicating the paralinguistic information output from the general-purpose paralinguistic information recognition model.

[0031] If there is one or more pieces of reference speech data (step S402; Yes), the flow proceeds to step S403. In step S403, the estimation unit 130 uses the reference speech data as a reference to estimate the paralinguistic information of the target user for the speech data to be recognized.

[0032] For example, the following methods can be used to estimate paralinguistic information. Each method uses as input the speech data to be recognized and reference speech data, or a transcript of these two types of speech data. Multiple reference speech data may also be used as input.

[0033] One example of an estimation method is a technique that uses a paralinguistic information recognition model such as an emotion recognition model. The paralinguistic information recognition model is trained to estimate paralinguistic information using speech to be recognized obtained from a user and neutral speech obtained from the same user as input. The paralinguistic information recognition model can be, but is not limited to, a neural network, a support vector machine, or a large-scale language model.

[0034] When the extraction unit 120 uses a model (also called a keyword detection model) constructed to extract expressions corresponding to predetermined keywords from speech data, the output from the keyword detection model (corresponding to the reference speech data) becomes the input to the paralinguistic information recognition model. Therefore, the keyword detection model and the paralinguistic information recognition model may be connected and regarded as one model, and the model may be trained to estimate paralinguistic information from the speech data of the target user.

[0035] Another example of the estimation method is a method that uses the similarity of features that represent speech data to be recognized and reference speech data, or a transcribed text of these two types of speech data.

[0036] When speech data is used as input, the estimation unit 130 uses a pre-prepared speech representation model to obtain features representing the input speech. Examples of the speech representation model include a neural network-based speech-based model (see, for example, Reference 6) and a neural network-based speaker recognition model pre-trained with a specific data set to estimate speakers. The features representing the speech are intermediate layer features (intermediate features) obtained when speech data is input to the speech representation model. Estimation of paralinguistic information can use the cosine similarity between intermediate features obtained from speech data to be recognized and intermediate features obtained from reference speech data. For example, if the cosine similarity exceeds a predetermined threshold, the estimation unit 130 determines that the speech data to be recognized is neutral speech, and that the target user is in a neutral state. If the cosine similarity is equal to or less than the threshold, the estimation unit 130 determines that the speech data to be recognized is non-neutral speech, and that the target user is in a non-neutral state. (Reference 6) Baevski, Alexei, et al. "wav2vec 2.0: A framework for self-supervised learning of speech representations." Advances in neural information processing systems 33 (2020): 12449-12460.

[0037] When using transcribed text as input, similar to when using speech data as input, the estimation unit 130 uses a pre-prepared text expression model to obtain features that represent the text. Examples of text expression models include a text-based model using a neural network (see, for example, Reference 7) and an emotion estimation model that is trained to estimate emotions by inputting text from a specific corpus in advance. The method of estimating paralinguistic information is also similar to when using speech data as input. (Reference 7) Devlin, Jacob, et al. "Bert: Pre-training of deep bidirectional transformers for language understanding." arXiv preprint arXiv:1810.04805 (2018).

[0038] Fig. 5 shows a schematic diagram of an example of the hardware configuration of a computer 500 that can realize the paralinguistic information recognition device 100. As shown in Fig. 5, the computer 500 includes, as hardware elements, a CPU (Central Processing Unit) 501, a RAM (Random Access Memory) 502, an auxiliary storage device 503, and an input / output interface 504. The CPU 501 is connected to the RAM 502, the auxiliary storage device 503, and the input / output interface 504 so as to be able to communicate with them.

[0039] The CPU 501 is an example of a general-purpose processor capable of executing programs. The RAM 502 includes a volatile memory such as a Synchronous Dynamic Random Access Memory (SDRAM) and is used as a work area for the CPU 501. The auxiliary storage device 503 includes a non-volatile memory such as a Hard Disk Drive (HDD) or a Solid State Drive (SSD) and stores a group of programs including a paralinguistic information recognition program, data, and the like.

[0040] The CPU 501 operates in accordance with a program stored in the auxiliary storage device 503. When executed by the CPU 501, the paralinguistic information recognition program causes the CPU 501 to perform the processes described with respect to the paralinguistic information recognition device 100. For example, the CPU 501 operates as an acquisition unit 110, an extraction unit 120, an estimation unit 130, and an output unit 140 in accordance with the paralinguistic information recognition program.

[0041] The input / output interface 504 is an interface for connecting an input device and an output device. The input device is a device that allows a system user (a human operator who operates the paralinguistic information recognition device 100) to input information. Examples of the input device include a keyboard and a mouse. The output device is a device that outputs information to the system user. Examples of the output device include a display device 510 and a speaker. The input device and the output device may be provided in the computer 500.

[0042] When a client-server model is adopted, the computer 500 may be provided with a communication interface instead of or in addition to the input / output interface 504. The CPU 501 communicates with a client device used by a system user via the communication interface. The CPU 501 receives input by the system user from the client device via the communication interface and transmits information to be presented to the system user to the client device via the communication interface.

[0043] Note that the computer 500 may include a dedicated processor such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit) instead of or in addition to the general-purpose processor. The processing circuit refers to a general-purpose processor, a dedicated processor, or a combination of a general-purpose processor and a dedicated processor.

[0044] A program such as a paralinguistic information recognition program may be provided to the computer 500 in a state where it is stored on a computer-readable recording medium. In this case, the computer 500 is equipped with a drive that reads data from the recording medium and acquires the program from the recording medium. Examples of recording media include magnetic disks, optical disks (CD-ROM, CD-R, DVD-ROM, DVD-R, etc.), magneto-optical disks (MO, etc.), and semiconductor memories. The program may also be distributed via a communications network. Specifically, the program may be stored on a server on the communications network, and the computer 500 may download the program from the server.

[0045] The paralinguistic information recognition device 100 can be operated in either offline or real-time mode (also called online mode). Offline mode is a mode in which the paralinguistic information recognition device 100 is operated after audio recording has finished, such as after a conference or a call at a call center. Real-time mode is a mode in which the paralinguistic information recognition device 100 is operated while audio is being recorded, such as during a conference or while a call at a call center is being handled.

[0046] Fig. 6 shows an outline of the paralinguistic information recognition process in offline operation, which is executed by the paralinguistic information recognition device 100. The paralinguistic information recognition process shown in Fig. 6 starts after the end of voice recording.

[0047] In step S601, the keyword detection unit 121 extracts the target user's utterance from the voice data.

[0048] In step S602, the keyword detection unit 121 determines whether or not a predetermined keyword is included in each utterance. If the keyword detection unit 121 detects a predetermined keyword in any utterance, it generates keyword information indicating a speech section corresponding to the predetermined keyword.

[0049] In step S603, the data extraction unit 122 extracts, as reference voice data, a voice section including a predetermined keyword from the voice data of the target user, based on the keyword information generated by the keyword detection unit 121. For example, the data extraction unit 122 obtains an utterance including the predetermined keyword as the reference voice data.

[0050] In step S604, the data extraction unit 122 obtains utterances that do not contain the predetermined keyword as speech data to be recognized.

[0051] In step S605, the estimation unit 130 estimates the paralinguistic information of the target user for each utterance (speech data to be recognized) obtained in step S604 based on the reference speech data obtained in step S603. The estimation unit 130 generates an estimation result indicating the paralinguistic information of the target user for each utterance.

[0052] In step S606, the output unit 140 outputs the estimation result obtained in step S605. For example, the output unit 140 displays the estimation result on a display device (for example, the display device 510 shown in FIG. 5).

[0053] Fig. 7 shows an outline of a real-time paralinguistic information recognition process executed by the paralinguistic information recognition device 100. The paralinguistic information recognition process shown in Fig. 7 starts when audio recording starts. For example, a media data acquisition device provides audio data obtained by audio recording to the paralinguistic information recognition device 100 in real time.

[0054] In step S701, the acquisition unit 110 extracts voice data corresponding to one utterance of the target user from the voice data received from the media data acquisition device. When the one utterance of the target user ends, the flow proceeds to step S702.

[0055] In step S702, the keyword detection unit 121 determines whether or not a predetermined keyword is included in the utterance obtained in step S701.

[0056] If the utterance contains a predetermined keyword (step S702; Yes), the flow proceeds to step S703. In step S703, the data extraction unit 122 obtains a speech section containing the predetermined keyword as reference speech data. For example, the data extraction unit 122 obtains the utterance obtained in step S701 as reference speech data. If the conference or call is continuing (step S709; Yes), the flow returns to step S701.

[0057] If the utterance does not contain a predetermined keyword (step S702; No), the flow proceeds to step S704. In step S704, the data extraction unit 122 obtains the utterance obtained in step S701 as speech data to be recognized.

[0058] In step S705, the estimation unit 130 determines whether or not there is reference voice data for the target user.

[0059] If there is reference speech data for the target user (step S705; Yes), the flow proceeds to step S706. In step S706, the estimation unit 130 estimates the paralinguistic information of the target user for the speech data to be recognized obtained in step S704 based on the reference speech data. For example, the estimation unit 130 inputs the reference speech data and the speech data to be recognized into a paralinguistic information recognition model, and obtains information indicating the paralinguistic information output from the paralinguistic information recognition model.

[0060] If there is no reference speech data for the target user (step S705; No), the flow proceeds to step S707. In step S707, the estimation unit 130 estimates the paralinguistic information of the target user for the speech data to be recognized obtained in step S704, using a general-purpose paralinguistic information recognition model prepared in advance. The general-purpose paralinguistic information recognition model receives speech as input and is trained in advance to estimate the paralinguistic information of the target user when the speech is uttered. The estimation unit 130 inputs the utterance obtained in step S701 into the general-purpose paralinguistic information recognition model, and obtains information indicating the paralinguistic information output from the general-purpose paralinguistic information recognition model.

[0061] In step S708, the output unit 140 displays the result of estimating the paralinguistic information of the target user obtained in step S706 or step S707. For example, if the estimation result of the paralinguistic information indicates that the target user is unsettled, the output unit 140 highlights the transcribed text of the utterance obtained in step S701. For example, if the estimation result of the paralinguistic information indicates that the target user is calm, the output unit 140 displays the transcribed text of the utterance obtained in step S701 without highlighting it.

[0062] If the conference or call is continuing (step S709; No), the flow returns to step S701. If the conference or call has ended (step S709; Yes), the flow ends.

[0063] In this way, the paralinguistic information recognition device 100 performs paralinguistic information recognition using a general-purpose paralinguistic information recognition model until reference speech data corresponding to the target user's calm speech data is obtained. Once the reference speech data of the target user is obtained, the paralinguistic information recognition device 100 recognizes the target user's paralinguistic information using the reference speech data of the target user as a reference for subsequent utterances.

[0064] An example of offline operation will be described. Consider a case where, as a review of telephone responses at a call center, it is confirmed whether a customer was feeling angry (negative). The paralinguistic information recognition device 100 divides the customer's voice data into utterance units, performs a keyword search for each utterance, and acquires the customer's reference voice data. The paralinguistic information recognition device 100 then uses the customer's reference voice data as a reference to perform anger detection for each utterance. If anger is detected, the paralinguistic information recognition device 100 displays an alert indicating that anger has been detected. The paralinguistic information recognition device 100 also highlights the transcript of the reference voice data and the transcript of the utterance in which anger was detected.

[0065] FIG. 8 is a schematic diagram illustrating a display example. As illustrated in FIG. 8, the display device displays the content of the conversation between the customer and the operator. Specifically, a transcript of the customer's speech and a transcript of the operator's speech are displayed in chronological order. In FIG. 8, "xxx" is the customer's name, "yyy" is the company name or the operator's name, and "zzz" is the product name. In the example illustrated in FIG. 8, voice data of "Hello, this is xxx" and voice data of "zzz" are acquired as reference voice data. Then, it is detected that the user is feeling angry in response to the customer's utterance, "What do you mean? If I hadn't called, would that mean I would never have received the call?" As a highlighting display, the portion extracted as the reference voice data is surrounded by a dashed line, and the utterance in which anger is detected is surrounded by a solid line. In another example, the portion extracted as the reference voice data may be highlighted in a first color (e.g., green), and the utterance in which anger is detected may be highlighted in a second color (e.g., red) different from the first color.

[0066] An example of real-time operation will be described. In a call center, the voices of customers and operators are collected and simultaneously divided into utterance units and a transcript is displayed. First, the paralinguistic information recognition device 100 detects predetermined keywords from the customer's utterance and acquires reference voice data. Based on the acquired reference voice data, the paralinguistic information recognition device 100 detects anger in the customer's subsequent utterance and displays an alert. If no reference voice data has been acquired, the paralinguistic information recognition device 100 uses a general-purpose emotion recognition model. The transcribed text is highlighted and an alert is displayed when a predetermined keyword or a specific emotion is detected.

[0067] Another example of real-time operation will be described. Consider a case in which a negotiator detects in real time from a conversation which products a business partner is interested in in order to smoothly advance the negotiation. The model used by the estimation unit 130 is configured to estimate the degree of interest by inputting two types of speech: a calm speech and a speech to be recognized. First, the paralinguistic information recognition device 100 detects predetermined keywords from the speech of the business partner and acquires reference speech data. Then, based on the acquired reference speech data, the paralinguistic information recognition device 100 detects interest from the speech of the business partner when the negotiator introduces each product. The paralinguistic information recognition device 100 displays the products that the negotiator introduced when the interest of the business partner was detected.

[0068] As described above, the paralinguistic information recognition device 100 acquires the speech data of the target user, extracts a portion of the speech data of the target user that includes predetermined keywords as reference speech data, and estimates the paralinguistic information of the target user from the speech data of the target user based on the reference speech data.

[0069] In the above configuration, reference speech data to be used as neutral speech data of the target user is acquired from the speech data of the target user, thereby making it possible to recognize paralinguistic information based on the neutral speech of the target user without having to prepare neutral speech data of the target user in advance.

[0070] The inventors conducted an evaluation experiment of the proposed method (the method according to this embodiment). For the experiment, a dataset simulating conversations between customers and operators at a call center (hereinafter referred to as the CC dataset) was used. In the CC dataset, each customer utterance during each conversation is labeled as either calm or non-calm.

[0071] The model used was an emotion estimation model trained on the online gaming voice chat corpus OGVC (online gaming voice chat corpus with emotional label) disclosed in Reference 8. (Reference 8) Yasuko Arimoto, Hiromi Kawazu, "Online Game Emotional Speech Corpus Using Voice Chat," Proceedings of the 2013 Autumn Meeting of the Acoustical Society of Japan, vol. 1-P-46a, pp. 385-388, (2013)

[0072] In the experiment, the baseline method and the proposed method were compared. The baseline method takes speech data as input, estimates nine emotions using an emotion recognition model, and maps the emotion with the highest probability to either neutral (neutral) or non-neutral (other). The proposed method obtains multiple utterances containing specified keywords as reference speech data, calculates the similarity between the average of the model's intermediate features for this reference speech data and the intermediate features of the emotion recognition model for the speech data to be recognized, and classifies the emotion as neutral if the similarity is above a set threshold, and non-neutral if the similarity is below the set threshold.

[0073] Figure 9 shows the results of the evaluation experiment for the baseline method and the proposed method. As shown in Figure 9, the accuracy rate of the baseline method was 48.5%, while that of the proposed method was 59.2%, demonstrating that the proposed method performed better than the baseline method. This demonstrates that the proposed method contributes to improving paralinguistic information recognition performance.

[0074] <Second embodiment> In the second embodiment, a case will be described in which the media data is moving image data, which refers to a collection of temporally consecutive frame images and does not include audio data.

[0075] The paralinguistic information recognition device according to the second embodiment has the same functional configuration as the paralinguistic information recognition device 100 according to the first embodiment. As shown in FIG. 1 , the paralinguistic information recognition device according to the second embodiment has an acquisition unit 110, an extraction unit 120, an estimation unit 130, and an output unit 140. The extraction unit 120 has a keyword detection unit 121 and a data extraction unit 122. Regarding each component, differences from the first embodiment will be described, and descriptions of similarities to the first embodiment will be omitted.

[0076] The acquisition unit 110 acquires video data of the target user. For example, the acquisition unit 110 receives the video data from a media data acquisition device such as a camera.

[0077] The keyword detection unit 121 divides the video data of the target user into utterance units. Hereinafter, the data portion obtained by dividing the video data of the target user into utterance units is referred to as a video fragment. Dividing the video data into utterance units can be achieved using lip recognition technology, such as that described in Reference 9. The keyword detection unit 121 performs keyword detection for each video fragment and generates keyword information indicating a time period containing a predetermined keyword. For example, the keyword detection unit 121 converts the video fragment into text using lip recognition and determines whether the text contains a predetermined keyword using text matching. Transcription using lip recognition can be achieved using known technology, such as that described in Reference 10. Video fragments in which a predetermined keyword is not detected are provided to the estimation unit 130 as video fragments to be recognized. (Reference 9) Hironori Kai, Daisuke Miyazaki, Ryo Furukawa, Masato Aoyama, Shinsaku Hiura, and Naoki Asada, "Speech detection by lip region extraction and recognition," Research Report on Computer Vision and Image Media (CVIM), vol. 2011, no. 13, pp. 1-8, 2011. [Online]. Available: https: / / cir.nii.ac.jp / crid / 1573105976846729216 (Reference 10) Ma, Pingchuan, et al. "Training strategies for improved lip-reading." ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022.

[0078] The data extraction unit 122 receives keyword information from the keyword detection unit 121 and, based on the keyword information, extracts a time interval including a predetermined keyword from the video data of the target user as reference video data. In one example, the data extraction unit 122 obtains a time interval corresponding to the predetermined keyword as reference video data. In another example, the data extraction unit 122 obtains a time interval from a point in time a predetermined time before the time interval corresponding to the predetermined keyword to a point in time a predetermined time after the time interval corresponding to the predetermined keyword as reference video data. In yet another example, the data extraction unit 122 obtains an entire video fragment including the predetermined keyword as reference video data.

[0079] The estimation unit 130 estimates the target user's paralinguistic information from the target user's video data using reference video data as a reference. For example, the estimation unit 130 receives the reference video data and each video fragment as input, and estimates the target user's paralinguistic information for each video fragment in a manner similar to that of the first embodiment. Alternatively, the estimation unit 130 may receive a transcript of the reference video data and a transcript of each video fragment as input, and estimate the target user's paralinguistic information for each video fragment in a manner similar to that of the first embodiment.

[0080] The output unit 140 outputs the estimation result of the paralinguistic information.

[0081] One example of an application using video data as media data is estimating a driver's fatigue level using video data captured by a security camera installed inside the vehicle. The extraction unit 120 performs lip recognition on the driver's video data and obtains video fragments containing a predetermined keyword as reference video data. The estimation unit 130 then estimates the driver's fatigue level for the video fragment to be recognized based on the reference video data.

[0082] For example, this can be used to prevent taxi accidents. Consider a system that estimates a driver's fatigue level and recommends appropriate rest if it detects driver fatigue in real time. In this case, the estimation unit 130 uses a model that estimates fatigue level by inputting calm video data and video data to be recognized. First, lip recognition is performed on video data captured by an in-vehicle security camera showing a conversation between the driver and a taxi passenger, and video fragments containing predetermined keywords are obtained as the driver's reference video data. The predetermined keywords can be expressions corresponding to greetings used by passengers when boarding a taxi or the amount the driver charges. The estimation unit 130 then estimates the driver's fatigue level based on the reference video data and displays an alert recommending the driver to take a rest if the fatigue level exceeds a predetermined threshold. If no reference video data is available, the estimation unit 130 uses a general-purpose model for estimating fatigue level.

[0083] In the second embodiment, it is possible to recognize the paralinguistic information of a target user without the need to prepare reference video data of the target user in advance. Also, it is possible to recognize the paralinguistic information of a target user from video data that does not include audio data.

[0084] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, and are also included in the scope of the invention and its equivalents as defined in the claims. [Explanation of symbols]

[0085] 100...paralinguistic information recognition device, 110...acquisition unit, 120...extraction unit, 121...keyword detection unit, 122...data extraction unit, 130...estimation unit, 140...output unit, 500...computer, 501...CPU, 502...RAM, 503...auxiliary storage device, 504...input / output interface, 510...display device.

Claims

1. an acquisition means for acquiring media data of the target user; an extraction means for extracting a portion including a predetermined keyword from the media data as reference media data; and an estimation means for estimating the paralinguistic information of the target user from the media data based on the reference media data; A paralinguistic information recognition program that allows a computer to function as a

2. The extraction means a keyword detection means for detecting a keyword portion corresponding to the predetermined keyword in the media data and generating a detection result; a data extraction means for extracting, based on the detection result, the portion of the media data containing the predetermined keyword as the reference media data; Equipped with The paralinguistic information recognition program according to claim 1.

3. the keyword detection means converts the media data into text and detects the keyword portion, which is a portion corresponding to the predetermined keyword, based on the text; 3. The paralinguistic information recognition program according to claim 2.

4. the keyword detection means uses a keyword spotting model or an event detection model to detect the keyword portion in the media data, which is a portion corresponding to the predetermined keyword; 3. The paralinguistic information recognition program according to claim 2.

5. the data extraction means extracts a portion including the keyword portion from the media data, and obtains the extracted portion as reference media data for the target user; 3. The paralinguistic information recognition program according to claim 2.

6. The predetermined keyword is an expression corresponding to a greeting, a backchannel, a reply, or a proper noun. The paralinguistic information recognition program according to claim 1.

7. The media data is audio data or video data. The paralinguistic information recognition program according to claim 1.

8. the extraction means uses a first model that extracts expressions corresponding to predetermined keywords from the speech data; the estimation means estimates the target user's paralinguistic information using a second model connected to an output of the first model; The second model is trained in conjunction with the first model. The paralinguistic information recognition program according to claim 1.

9. the media data is audio data, the reference media data is audio data representing a neutral state of the target user, and the paralinguistic information is an emotion; The paralinguistic information recognition program according to claim 1.

10. causing the computer to further function as a display means for displaying transcribed text obtained from the media data; the display means highlights, in the transcribed text, a portion corresponding to the reference media data and a portion estimated as a predetermined emotion by the estimation means, and further displays an alert indicating that the predetermined emotion has been estimated. The paralinguistic information recognition program according to claim 1.

11. an acquisition unit for acquiring media data of a target user; an extracting unit that extracts a portion including a predetermined keyword from the media data as reference media data; an estimation unit that estimates paralinguistic information of the target user from the media data based on the reference media data; A paralinguistic information recognition device comprising:

12. 1. A computer-implemented method for recognizing paralinguistic information, comprising: Obtaining media data of a target user; extracting a portion including a predetermined keyword from the media data as reference media data; estimating paralinguistic information of the target user from the media data relative to the reference media data; A paralinguistic information recognition method comprising:

Citation Information

Patent Citations

  • Emotion recognition device, emotion recognition model learning device, and methods and programs thereof

    JP7420211B2