Support device
The auxiliary device and method address the challenge of determining speaking timing for hearing-impaired individuals in remote meetings by analyzing audio data for context and providing non-visual notifications, ensuring timely and appropriate speech opportunities.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2022-03-16
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies struggle for individuals with hearing impairments to determine appropriate speaking timing in remote meetings, as visualizing speech content alone is insufficient for assessing the atmosphere, leading to uncertainty about when to speak.
An auxiliary device and method that utilize voice data to determine suitable speaking timing by analyzing audio data for context and atmosphere, and notify users through non-visual means such as device vibration, ensuring appropriate timing for speech.
Enables users to grasp suitable speaking times effectively, reducing confusion and enhancing participation in remote meetings without relying on visual cues.
Smart Images

Figure 0007848528000001 
Figure 0007848528000002 
Figure 0007848528000003
Abstract
Description
Technical Field
[0001] The present invention relates to an auxiliary device, an auxiliary method, and a program.
Background Art
[0002] When a person with a hearing impairment participates in a remote meeting or the like, the spoken content may be visualized by speech recognition. However, simply visualizing the content makes it difficult to grasp whether anyone is trying to speak, and there is a risk that it will be difficult to determine the timing to speak. Therefore, techniques for detecting the speaking timing are known.
[0003] As a technique used for detecting the speaking timing, for example, there is Patent Document 1. Patent Document 1 describes a technique for determining the speaking timing by using an acquired speaking section and user information by using a microphone, a camera, or other sensors.
[0004] Also, as a related technique, for example, there is Patent Document 2. Patent Document 2 describes a determination device aimed at making an appropriate response to a user's speech. For example, Patent Document 2 describes a determination device having an acquisition unit that acquires context information regarding a user's speech and a determination unit that determines an output mode of a response to the user's speech based on the context information acquired by the acquisition unit.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0006] In the technology described in Patent Document 1, while the timing for speaking can be determined by detecting situations where no one is about to speak, it is difficult for the user to judge whether the atmosphere is suitable for them to speak. Therefore, there was a problem in that it was difficult to confirm whether the timing was truly appropriate for speaking. In particular, for people with hearing impairments, it is difficult to grasp whether the atmosphere is good or bad from the speaker's tone, so the above problem was even more problematic. Furthermore, this problem could not be solved even by using the technology described in Patent Document 2.
[0007] Therefore, the objective of the present invention is to solve the above-mentioned problems. [Means for solving the problem]
[0008] An auxiliary device, which is one embodiment of the present invention, A situation determination unit that determines whether or not the situation is such that speech can be uttered based on voice data, A timing determination unit determines an appropriate timing for uttering based on the result of the determination by the aforementioned situation determination unit, A notification unit that notifies in a manner appropriate to the situation determined by the situation determination unit, based on the result of the determination by the timing determination unit, has This is the structure it takes.
[0009] Furthermore, an auxiliary method, which is one embodiment of the present invention, Information processing device, Based on the audio data, it is determined whether or not the person is in a state where they can speak. Based on the results of the assessment, the appropriate timing for speaking is determined. Depending on the timing determination result, notifications will be sent in a manner appropriate to the determined situation. This is the structure it takes.
[0010] Furthermore, a program, which is one embodiment of the present invention, In an information processing device, Based on the voice data, determine whether it is a situation where speaking is possible, Based on the result of the determination, determine the timing suitable for speaking, According to the determination result of the timing, notify in a method corresponding to the determined situation It is a program for realizing the processing.
Effect of the Invention
[0011] According to each configuration as described above, it is possible to appropriately notify the timing suitable for speaking.
Brief Description of the Drawings
[0012] [Figure 1] It is a diagram showing a configuration example of a remote conference system in the first embodiment of the present disclosure. [Figure 2] It is a block diagram showing a configuration example of a processing device. [Figure 3] It is a diagram showing an example of information included in the determination information. [Figure 4] It is a diagram showing an example of information included in the determination information. [Figure 5] It is a diagram for explaining an example during the determination model learning. [Figure 6] It is a diagram showing an example of information included in the notification means information. [Figure 7] It is a flowchart showing an operation example of a speaking timing detection unit. [Figure 8] It is a flowchart showing an operation example of a context determination unit. [Figure 9] It is a flowchart showing another operation example of the context determination unit. [Figure 10] It is a flowchart showing an operation example of a timing determination unit. [Figure 11] It is a flowchart showing an operation example of a notification unit. [Figure 12] It is a diagram showing a hardware configuration example of an auxiliary device in the second embodiment of the present disclosure. [Figure 13] It is a block diagram showing a configuration example of the auxiliary device [Modes for carrying out the invention]
[0013] [First Embodiment] A first embodiment of this disclosure will be described with reference to Figures 1 to 11. Figure 1 is a diagram showing an example configuration of the remote conferencing system 100 in the first embodiment of this disclosure. Figure 2 is a block diagram showing an example configuration of the processing unit 300. Figure 3 is a diagram showing an example of the information contained in the determination information 352. Figure 4 is a diagram showing an example of the information contained in the determination information 352. Figure 5 is a diagram illustrating an example of the learning process of the determination model 353. Figure 6 is a diagram showing an example of the information contained in the notification means information 354. Figure 7 is a flowchart showing an example of the operation of the speech timing detection unit 365. Figure 8 is a flowchart showing an example of the operation of the context determination unit 366. Figure 9 is a flowchart showing another example of the operation of the context determination unit 366. Figure 10 is a flowchart showing an example of the operation of the timing determination unit 367. Figure 11 is a flowchart showing an example of the operation of the notification unit 368.
[0014] In the first embodiment of this disclosure, a remote conferencing system 100 capable of assisting the user of the processing unit 300 by determining and notifying the timing of speech utterance is described. As will be described later, the remote conferencing system 100 detects a timing at which speech can be uttered based on voice data, etc. The remote conferencing system 100 also performs a context determination based on the voice data, etc. to determine whether or not it is a situation in which speech can be uttered. Then, the remote conferencing system 100 determines the timing of speech utterance based on the timing detection result and the result of the context determination, and notifies the user of the processing unit 300 in a manner appropriate to the determination result. For example, the remote conferencing system 100 can notify the result of the determination of the timing of speech utterance in a manner appropriate to the determination result that does not rely on visual perception, such as changing the strength or duration of vibration according to the result of the context determination.
[0015] Figure 1 shows an example of the overall configuration of the remote conferencing system 100. Referring to Figure 1, the remote conferencing system 100 includes, for example, one or more information processing devices 200 and a processing device 300. As shown in Figure 1, the information processing devices 200 and the processing device 300 are connected to each other so that they can communicate with each other, for example, via a network. Note that the remote conferencing system 100 may have multiple processing devices 300.
[0016] The information processing device 200 is a device that conducts remote conferences with other information processing devices 200 and processing devices 300. For example, the information processing device 200 may be a general personal computer or tablet equipped with sound data acquisition functions such as a microphone and imaging functions such as a camera.
[0017] For example, when the information processing device 200 acquires sound data or image data, it can transmit the acquired sound data or image data to other information processing devices 200 or processing devices 300. The information processing device 200 may use a general remote conferencing system or the like to transmit sound data or image data.
[0018] The processing unit 300, like the information processing unit 200, is a device that conducts remote conferences with the information processing unit 200 and other processing units 300. For example, the processing unit 300 may be a general personal computer or tablet equipped with sound data acquisition functions such as a microphone and imaging functions such as a camera. Similar to the information processing unit 200, the processing unit 300 may transmit sound data and image data using a general remote conferencing system.
[0019] Furthermore, the processing unit 300 can assist the user of the processing unit 300 by determining and notifying the timing of speech utterance. Figure 2 shows an example of the configuration of the processing unit 300 characteristic of this embodiment. Referring to Figure 3, the processing unit 300 has, as its main components, an operation input unit 310, an imaging unit 320, a screen display unit 330, a communication I / F unit 340, a storage unit 350, and an arithmetic processing unit 360.
[0020] Figure 3 illustrates a case where the functions of the processing unit 300 are realized using a single information processing unit. However, the processing unit 300 may be realized using multiple information processing units, for example, by being implemented on the cloud. For example, the processing unit 300 may consist of a processing unit having the functions of a non-speech interval detection unit 362, a transcription unit 363, and an atmosphere determination unit 364, and an auxiliary device having the functions of a speech timing detection unit 365, a context determination unit 366, a timing determination unit 367, and a notification unit 368. Furthermore, the processing unit 300 may not include some of the configurations exemplified above, such as not having an operation input unit 310 or a screen display unit 330, and may have configurations other than those exemplified above. For example, the processing unit 300 may have a sensor unit such as a heart rate sensor that acquires vital data.
[0021] The operation input unit 310 consists of an operation input device such as a keyboard or mouse. The operation input unit 310 detects the operation of the user operating the processing unit 300 and outputs it to the arithmetic processing unit 360. The operation input unit 310 may also include a microphone or the like.
[0022] The imaging unit 320 consists of an imaging device such as a camera. The imaging unit 320 acquires image data of users or other subjects using the processing unit 300 and outputs it to the arithmetic processing unit 360.
[0023] The screen display unit 330 consists of a screen display device such as an LCD (Liquid Crystal Display). The screen display unit 330 can display various information stored in the storage unit 350 on the screen in response to instructions from the arithmetic processing unit 360.
[0024] The communication interface unit 340 consists of data communication circuits and the like. The communication interface unit 340 performs data communication with external devices such as the information processing device 200 and other processing devices 300 that are connected via a communication line.
[0025] The storage unit 350 is a storage device such as a hard disk or memory. The storage unit 350 stores processing information and programs 358 necessary for various processes in the arithmetic processing unit 360. The programs 358 are read into the arithmetic processing unit 360 and executed to realize various processing processes. The programs 358 are pre-read from external devices or recording media via data input / output functions such as the communication I / F unit 340 and stored in the storage unit 350. The main information stored in the storage unit 350 includes, for example, a trained model 351 for atmosphere determination, determination information 352, determination model 353, notification means information 354, image data information 355, voice data information 356, and character information 357. Note that the storage unit 350 may not have some of the information exemplified above, such as not having a determination model 353.
[0026] The pre-trained model 351 for atmosphere determination is a model that quantifies and outputs the atmosphere of a place based on audio data acquired from the information processing device 200, the processing device 300, etc. For example, the pre-trained model 351 for atmosphere determination is generated in advance by machine learning using training data in which labels that quantify the atmosphere of the place are attached to audio data, and is stored in the memory unit 350. For example, the pre-trained model 351 for atmosphere determination may be generated using a technique such as that described in Patent Document 3.
[0027] The pre-trained model 351 for atmosphere determination may be, for example, a model that quantifies and outputs the atmosphere of a place based on information that leads to the estimation of participants' emotions, such as vital data, and audio data. For example, the pre-trained model 351 for atmosphere determination may be generated in advance by performing machine learning using training data in which labels that quantify the atmosphere of a place are attached to audio data and vital data, etc.
[0028] The determination information 352 includes information used when performing context determination. For example, the determination information 352 is acquired in advance by methods such as obtaining it from an external device via the communication I / F unit 340 or inputting it using the operation input unit 310, and is stored in the storage unit 350.
[0029] Figures 3 and 4 show examples of the information included in the judgment information 352. For example, as illustrated in Figure 3, the judgment information 352 associates a conditional expression with an atmosphere. Here, the conditional expression includes information indicating words and conditions, such as the inclusion of multiple words or the inclusion of any of the words. For example, "K1,1 and K1,2" as illustrated in Figure 3 indicates the condition that both the words "K1,1" and "K1,2" are included. "K1,1" and "K1,2" can be any words. The atmosphere also includes information indicating the mood, such as "positive" or "negative." The information included in the atmosphere may be further subdivided, or it may be a numerical representation of the degree of positivity or negativity, for example.
[0030] Furthermore, as shown in Figure 4, the judgment information 352 may include information that associates the atmosphere with information indicating whether or not it is a context in which one wishes to speak. For example, Figure 4 indicates that an atmosphere of "positive" is a context in which one wishes to speak. As illustrated in Figure 3, the information indicating the atmosphere and whether or not it is a context in which one wishes to speak may be further subdivided or quantified.
[0031] Furthermore, the judgment information 352 may include information other than those exemplified above. For example, the judgment information 352 may include only words as keywords.
[0032] The judgment model 353 includes a pre-trained model used when performing context determination. In other words, the judgment model 353 is a pre-trained model that determines whether or not a situation is one in which speech is possible. For example, the judgment model 353 is pre-trained in an external device or processing unit 300 and stored in the memory unit 350.
[0033] For example, when training the classification model 353, training data is generated by assigning labels such as "want to speak" or "do not want to speak" to training conference data, which includes audio and image data. Then, the classification model 353 is trained by performing machine learning using the generated training data. Note that the labels may be further subdivided.
[0034] Figure 5 is a diagram illustrating in more detail an example of training the judgment model 353. As illustrated in Figure 5, labeling can be performed while displaying, for example, a video display unit 401 showing the proceedings of a meeting, a label input unit 402 used for label input, a graph display unit 403 representing the atmosphere of the situation quantified using the trained atmosphere judgment model 351, and a transcription result display unit 404. Specifically, for example, a user performing labeling inputs labels for sections where they "want to speak" or "don't want to speak," based on the video displayed in the video display unit 401, the graph displayed in the graph display unit 403, and the text displayed in the transcription result display unit 404. In other words, the user performing labeling inputs labels according to the degree to which they "want to speak" or "don't want to speak." For example, the user inputs a higher value as a label the more they "want to speak." Furthermore, the transcription results displayed in the transcription result display unit 404 can be highlighted based on the intonation of the speech, making it easier for people with hearing impairments to grasp the emotional information contained in the speech.
[0035] Furthermore, after generating training data as described above, a judgment model 353 can be generated by performing supervised learning using the atmosphere of the place as a feature. For example, to ensure a sufficient amount of data for training, the labeled intervals may be divided into fixed intervals, and then supervised learning may be performed using the atmosphere of the place at each timing as a feature. For example, by processing in this way, a judgment model 353 can be trained, which outputs a value indicating the degree to which one "wants to speak" or "does not want to speak," by inputting the output of the trained atmosphere judgment model 351.
[0036] The notification means information 354 includes information indicating the notification means used when making a notification. For example, the notification means information 354 is acquired in advance by methods such as obtaining it from an external device via the communication I / F unit 340 or inputting it using the operation input unit 310, and is stored in the storage unit 350.
[0037] Figure 6 shows an example of the information included in notification means information 354. Referring to Figure 6, the notification means information 354 associates likelihood with vibration pattern. Here, likelihood is a value corresponding to the value output by the determination model 353. The vibration pattern indicates the length and intensity of vibration when vibrating an arbitrary device. For example, through this association, the notification means information 354 shows a vibration pattern corresponding to the output of the determination model 353.
[0038] Image data information 355 contains image data. For example, image data information 355 is updated when the receiving unit 361 acquires image data from the information processing device 200 or other processing devices 300, or when the receiving unit 361 acquires image data acquired by the imaging unit 320.
[0039] The audio data information 356 contains audio data. For example, the audio data information 356 is updated when the receiving unit 361 acquires audio data from the information processing device 200 or other processing devices 300, or when the receiving unit 361 acquires audio data acquired through a microphone or the like provided by the processing device 300.
[0040] The character information 357 contains character data. For example, the character information 357 is updated when the receiving unit 361 receives character data, or when the transcription unit 363 converts the audio data into character data based on the audio data information 356.
[0041] The arithmetic processing unit 360 has an arithmetic unit such as a CPU (Central Processing Unit) and its peripheral circuits. The arithmetic processing unit 360 reads and executes a program 358 from the storage unit 350, thereby realizing various processing functions by having the hardware and the program 358 cooperate. The main processing functions realized by the arithmetic processing unit 360 include, for example, a receiving unit 361, a non-speech interval detection unit 362, a transcription unit 363, an atmosphere determination unit 364, a speech timing detection unit 365, a context determination unit 366, a timing determination unit 367, and a notification unit 368.
[0042] The receiving unit 361 receives image data, audio data, text data, etc. The receiving unit 361 also stores the received image data, audio data, and text data in the storage unit 350 as image data information 355, audio data information 356, and text information 357. For example, the receiving unit 361 may receive image data, audio data, text data, etc. by using a known remote conferencing system.
[0043] For example, the receiving unit 361 receives image data, audio data, character data, etc., from the information processing device 200 and other processing devices 300 via the communication I / F unit 340. The receiving unit 361 can also acquire image data and audio data acquired by the microphone or imaging unit 320 of its own device. The receiving unit 361 may also acquire character data generated by the transcription unit 363.
[0044] The receiving unit 361 may also receive information other than those exemplified above from the information processing device 200 or other processing devices 300, etc. For example, the receiving unit 361 may receive information indicating whether the microphone is on or off, information on whether the mute function is on or off, etc., by using a known remote conferencing system.
[0045] The non-speech section detection unit 362 detects non-speech sections, which are sections where no speech is present, based on the audio data. The non-speech section detection unit 362 may detect non-speech sections using any method. For example, the non-speech section detection unit 362 may detect speech sections from sections that are not speech sections by detecting speech sections using VAD (Voice Activity Detection) on the audio data.
[0046] The transcription unit 363 converts the audio data into text data. The transcription unit 363 may perform the above conversion process using known techniques. The transcription unit 363 also stores the converted text data as text information 357 in the storage unit 350.
[0047] The atmosphere determination unit 364 quantifies and outputs the atmosphere of the place based on the audio data. For example, the atmosphere determination unit 364 outputs a value corresponding to the atmosphere of the place by inputting the audio data into a trained model 351 for atmosphere determination. For example, the atmosphere determination unit 364 may convert the audio data into a value corresponding to the atmosphere of the place using a technique such as that described in Patent Document 3.
[0048] The speech timing detection unit 365 detects the speech timing based on the non-speech interval detection unit 362 detected. For example, the speech timing detection unit 365 detects a time when there is no participant attempting to speak as the speech timing.
[0049] Specifically, for example, the speech timing detection unit 365 checks whether the microphones of the information processing device 200 and other processing devices 300 are off (i.e., muted). If all the microphones of the information processing device 200 and other processing devices 300 are off, the speech timing detection unit 365 detects the speech timing. On the other hand, if even one microphone of the information processing device 200 and other processing devices 300 is on, the speech timing detection unit 365 checks the detection result by the non-speech interval detection unit 362. For example, if the non-speech interval detection unit 362 is detecting a non-speech interval, the speech timing detection unit 365 detects the speech timing. On the other hand, if the non-speech interval detection unit 362 is not detecting a non-speech interval, the speech timing detection unit 365 does not detect the speech timing.
[0050] For example, as described above, the speech timing detection unit 365 detects a speech timing when there is no participant attempting to speak, based on microphone on / off information (information about muting) and the detection results from the non-speech interval detection unit 362.
[0051] The context determination unit 366 performs a context determination to determine whether or not it is a situation in which speech is possible, based on the speech data or the text data obtained by converting the speech data. For example, the context determination unit 366 can perform a context determination based on the determination information 352. The context determination unit 366 may also perform a context determination based on the determination information 352 and the determination model 353.
[0052] For example, the context determination unit 366, triggered by the reception of character data by the receiving unit 361, checks whether the character data satisfies the conditional expression exemplified in Figure 3, thereby determining whether the atmosphere is positive or negative. The context determination unit 366 also compares the result of the positive / negative determination with the information exemplified in Figure 4 to determine whether the context is one in which the user wants to speak. For example, if the context determination unit 366 determines, based on Figures 3 and 4, that the situation is one in which the user wants to speak, it notifies the timing determination unit 367 or the like of a speakable context. On the other hand, if it determines that the situation is not one in which the user wants to speak, the context determination unit 366 notifies the timing determination unit 367 or the like of a speakless context.
[0053] The context determination unit 366 may, for example, perform a positive / negative determination based on the output of the atmosphere determination trained model 351, instead of, or in conjunction with, a positive / negative determination based on the character data and the determination information 352.
[0054] Furthermore, as described above, the context determination unit 366 may perform context determination based on the determination information 352 and the determination model 353. For example, the context determination unit 366 checks whether the keywords included in the determination information 352 are included in the character data, triggered by the reception of character data by the receiving unit 361. If the keywords are included in the character data, the context determination unit 366 checks the atmosphere of the situation. For example, the context determination unit 366 inputs the output from the atmosphere determination unit 364 to the determination model 353 to obtain output from the determination model 353. For example, if the determination model 353 produces an output classified as "want to speak", the context determination unit 366 notifies the timing determination unit 367 or the like of a context in which speaking is possible. On the other hand, if the determination model 353 produces an output classified as "do not want to speak", the context determination unit 366 notifies the timing determination unit 367 or the like of a context in which speaking is not possible.
[0055] Furthermore, if the keyword is not included in the character data, the context determination unit 366 can notify the timing determination unit 367 or the like of a non-speakable context. The context determination unit 366 may also be configured to check whether the character data contains characters identical or of the same type as the keyword included in the determination information 352.
[0056] For example, the context determination unit 366 determines the situation by processing one or a combination of the methods exemplified above, and notifies the user of whether the context allows for speaking or not, based on the voice data. The context determination unit 366 can not only output information indicating whether the user "wants to speak" or "does not want to speak," but also output a value indicating the degree to which the user "wants to speak" or "does not want to speak," by using the determination model 353, etc. In other words, the context determination unit 366 can output information indicating the degree of ease of speaking. The degree of ease of speaking indicates, for example, that the higher the value, the easier it is for the user of the processing unit 300 to speak.
[0057] Based on the detection results from the speech timing detection unit 365 and the determination results from the context determination unit 366, the timing determination unit 367 notifies the notification unit 368 or the like of a speech-appropriate timing, indicating that it is a suitable time to speak if it is a suitable time to speak and the context is one in which the user wants to speak.
[0058] For example, the timing determination unit 367, triggered by the detection of speech timing by the speech timing detection unit 365, checks the result of the determination made by the context determination unit 366. For example, if the latest determination result from the context determination unit 366 is a "speech-ready context," the timing determination unit 367 notifies the notification unit 368 or the like of the speech-appropriate timing. For example, the notification of the speech-appropriate timing may include information indicating that the timing is suitable for speaking, as well as information indicating the degree to which the user "wants to speak." On the other hand, if the latest determination result from the context determination unit 366 is a "non-speech-ready context," the timing determination unit 367 does not make the above notification.
[0059] The notification unit 368 provides notifications to the user. For example, the notification unit 368 can notify a predetermined device, such as a smartphone or other mobile terminal owned by the user, in a manner that corresponds to the degree to which the user "wants to speak". In other words, the notification unit 368 notifies the device to vibrate in a manner that corresponds to the degree to which the user "wants to speak". For example, upon receiving a speech-suitability timing, the notification unit 368 refers to the notification means information 354 to confirm the strength and duration of the vibration according to the degree to which the user "wants to speak". Then, the notification unit 368 notifies the predetermined device to vibrate according to the degree to which the user "wants to speak".
[0060] The above is an example of the configuration of the processing unit 300, which functions as an auxiliary device to assist the user.
[0061] Next, we will describe an example of the operation of the processing unit 300 with reference to Figures 7 to 11. First, we will describe an example of the operation of the speech timing detection unit 365 with reference to Figure 7.
[0062] Figure 7 is a flowchart showing an example of the operation of the speech timing detection unit 365. Referring to Figure 7, the speech timing detection unit 365 checks whether the microphones of the information processing device 200 and other processing devices 300 are turned off (i.e., muted) or not (step S101).
[0063] , If all microphones on the information processing device 200 and other processing devices 300 are turned off (step S101, Yes), the speech timing detection unit 365 detects the speech timing (step S103). On the other hand, if even one microphone on the information processing device 200 and other processing devices 300 is turned on (step S101, No), the speech timing detection unit 365 checks the detection result by the non-speech interval detection unit 362 (step S102). Then, for example, if the non-speech interval detection unit 362 is detecting a non-speech interval (step S102, Yes), the speech timing detection unit 365 detects the speech timing (step S103). On the other hand, if the non-speech interval detection unit 362 is not detecting a non-speech interval (step S102, No), the speech timing detection unit 365 does not detect the speech timing.
[0064] The above is an example of the operation of the speech timing detection unit 365. Next, an example of the operation of the context determination unit 366 will be explained with reference to Figures 8 and 9.
[0065] Figure 8 is a flowchart illustrating an example of the operation of the context determination unit 366. Referring to Figure 8, the context determination unit 366, triggered by the reception of character data by the receiving unit 361, checks whether the character data satisfies the conditional expression exemplified in Figure 3, thereby determining whether the atmosphere is positive or negative (step S201). The context determination unit 366 also compares the result of the positive / negative determination with the information exemplified in Figure 4 to determine whether the context is one in which the user wants to speak (step S202). For example, if the context determination unit 366 determines, based on Figures 3 and 4, that the situation is one in which the user wants to speak (step S202, Yes), it notifies the timing determination unit 367, etc., of the speakable context (step S203). On the other hand, if it determines that the situation is not one in which the user wants to speak (step S202, No), the context determination unit 366 notifies the timing determination unit 367, etc., of the speakable context (step S204).
[0066] Furthermore, the context determination unit 366 may perform operations as illustrated in Figure 9 instead of the operations shown in Figure 8. Figure 9 is a flowchart showing other examples of operations of the context determination unit 366. Referring to Figure 9, the context determination unit 366 checks whether the keyword included in the determination information 352 is included in the character data, triggered by the reception of character data by the receiving unit 361 (step S211). If the keyword is included in the character data (step S211, Yes), the context determination unit 366 checks the atmosphere of the situation (step S212). For example, the context determination unit 366 inputs the output from the atmosphere determination unit 364 to the determination model 353 to obtain an output from the determination model 353. Then, for example, if the determination model 353 produces an output classified as "want to speak" (step S212, Yes), the context determination unit 366 notifies the timing determination unit 367 or the like of the utterable context (step S213). On the other hand, if the judgment model 353 produces an output that is classified as "not wanting to speak" (step S212, No), the context determination unit 366 notifies the timing determination unit 367 or the like of the non-speakable context (step S214). In addition, if the keyword is not included in the character data (step S211, No), the context determination unit 366 can also notify the timing determination unit 367 or the like of the non-speakable context (step S214).
[0067] The above is an example of the operation of the context determination unit 366. Next, an example of the operation of the timing determination unit 367 will be explained with reference to Figure 10.
[0068] Figure 10 is a flowchart illustrating an example of the operation of the timing determination unit 367. Referring to Figure 10, the timing determination unit 367 is triggered by the detection of a speech timing by the speech timing detection unit 365 (step S301) and confirms the result of the determination by the context determination unit 366 (step S302). For example, if the latest determination result from the context determination unit 366 is a "speech-ready context" (step S302, Yes), the timing determination unit 367 notifies the notification unit 368 or the like of the speech-ready timing (step S303). For example, the notification of the speech-ready timing may include information indicating that the timing is suitable for speaking, as well as information indicating the degree to which the user "wants to speak". On the other hand, if the latest determination result from the context determination unit 366 is a "non-speech-ready context" (step S302, No), the timing determination unit 367 does not make the above notification.
[0069] The above is an example of the operation of the timing determination unit 367. Next, an example of the operation of the notification unit 368 will be explained with reference to Figure 11.
[0070] Figure 11 is a flowchart illustrating an example of the operation of the notification unit 368. Referring to Figure 11, for example, the notification unit 368 receives the speech matching timing notified in the process of step S303 described above (step S401). Then, the notification unit 368 refers to the notification means information 354. The notification unit 368 then notifies a predetermined device or the like to perform vibrations according to the degree to which the user wants to speak (step S402).
[0071] The above is an example of the operation of the notification unit 368.
[0072] Thus, the processing unit 300 includes a context determination unit 366, a timing determination unit 367, and a notification unit 368. With this configuration, the timing determination unit 367 can notify the notification unit 368 or the like of a suitable timing for speaking if the latest determination result from the context determination unit 366 is a "speakable context". As a result, the notification unit 368 can notify a predetermined device or the like to vibrate according to the degree to which the user "wants to speak". In this way, with the above configuration, it is possible to determine whether or not it is a situation in which one can speak, and if it is a situation in which one can speak, notify the user in a way that corresponds to the degree to which the user "wants to speak". This makes it possible to notify the user of a suitable timing for speaking in an appropriate manner.
[0073] Furthermore, the aforementioned processing unit 300 can notify the user of an appropriate timing for speaking without using visual information such as device vibration. Additionally, the processing unit 300 can provide notifications based on the ease of speaking, without using visual information. Therefore, even for users with hearing impairments, confusion due to excessive visual information can be prevented.
[0074] The configuration of the processing unit 300 is not limited to the case described above. For example, the notification unit 368 may notify the user of the processing unit 300 and also transmit information indicating that a notification has been sent to the information processing unit 200 and other processing units 300. By making such a notification, for example, the information processing unit 200 and other processing units 300 can provide the user of the processing unit 300 that sent the notification with an opportunity to speak. Furthermore, if the above mechanism is performed by multiple processing units 300, for example, the information processing unit 200 acting as a meeting facilitator can list the processing units 300 that are at a suitable timing for speaking. This allows the facilitator to be notified of participants who are suitable to speak, for example, when there is no speaking and the discussion is not progressing. This listing may be done, for example, by color-coding the list of participants.
[0075] In this embodiment, the remote conferencing system 100 has been described. However, the present invention is not limited to cases where meetings are held remotely. For example, the processing device 300 may be used in various systems that transmit and receive audio data.
[0076] [Second Embodiment] Next, a second embodiment of the present disclosure will be described with reference to Figures 12 and 13. Figure 12 is a diagram showing an example of the hardware configuration of the auxiliary device 500. Figure 13 is a block diagram showing an example of the configuration of the auxiliary device 500.
[0077] In a second embodiment of this disclosure, an example configuration of an auxiliary device 500, which is an information processing device that assists in the transmission and reception of voice data over a network, such as in remote conferencing, will be described. Figure 12 shows an example of the hardware configuration of the auxiliary device 500. Referring to Figure 12, the auxiliary device 500 has, as an example, the following hardware configuration. ·CPU (Central Processing Unit) 501 (computing unit) ROM (Read Only Memory) 502 (Storage Device) • RAM (Random Access Memory) 503 (storage device) • Program group 504 loaded into RAM 503 • Storage device 505 for storing the program group 504 • Drive device 506 for reading and writing to the recording medium 510 outside the information processing device. • Communication interface 507 connecting to the communication network 511 outside the information processing device. • Input / output interface 508 for data input and output. • Bus 509 connecting each component
[0078] Furthermore, the auxiliary device 500 can realize the functions of the status determination unit 521, timing determination unit 522, and notification unit 523 shown in Figure 13 by having the CPU 501 acquire the program group 504 and execute it. The program group 504 is, for example, stored in advance in a storage device 505 or ROM 502, and the CPU 501 loads it into RAM 503 or the like and executes it as needed. Alternatively, the program group 504 may be supplied to the CPU 501 via a communication network 511, or it may be stored in advance in a recording medium 510, and the drive device 506 may read the program and supply it to the CPU 501.
[0079] Figure 12 shows an example of the hardware configuration of the auxiliary device 500. The hardware configuration of the auxiliary device 500 is not limited to the case described above. For example, the auxiliary device 500 may consist of only a part of the configuration described above, such as not having the drive device 506.
[0080] The situation determination unit 521 determines whether or not the situation is such that speech can be uttered, based on the audio data. The situation determination unit 521 may also determine whether or not the situation is such that speech can be uttered, based on the text data obtained by converting the audio data.
[0081] The timing determination unit 522 determines the appropriate timing for uttering based on the result of the determination by the situation determination unit 521.
[0082] The notification unit 523 notifies in a manner appropriate to the situation determined by the situation determination unit 521, based on the result of the determination by the timing determination unit 522. For example, the notification unit 523 can notify a predetermined device to vibrate in a manner appropriate to the situation.
[0083] Thus, the auxiliary device 500 includes a situation determination unit 521, a timing determination unit 522, and a notification unit 523. With this configuration, the notification unit 523 can notify in a manner appropriate to the situation determined by the situation determination unit 521, according to the result of the determination by the timing determination unit 522. As a result, the timing suitable for speaking can be notified in an appropriate manner.
[0084] The auxiliary device 500 described above can be realized by incorporating a predetermined program into an information processing device such as the auxiliary device 500. Specifically, another form of the present invention is a program for an information processing device such as the auxiliary device 500 that performs the following processes: determining whether or not it is possible to speak based on voice data; determining an appropriate timing for speaking based on the result of the determination; and notifying the user in a manner appropriate to the determined situation.
[0085] Furthermore, the auxiliary method performed by the information processing device such as the auxiliary device 500 described above involves the information processing device such as the auxiliary device 500 determining, based on voice data, whether or not it is possible to speak, determining an appropriate timing for speaking based on the result of the determination, and notifying the user in a manner appropriate to the determined situation.
[0086] Even if the invention is a program, or a recording medium readable by a computer on which the program is recorded, or an auxiliary method having the above-described configuration, it can achieve the same effects and functions as the auxiliary device 500 described above, and thus achieve the objectives of the present disclosure described above.
[0087] <Note> Some or all of the above embodiments may also be described as follows. The following outlines the auxiliary devices and other components of the present invention. However, the present invention is not limited to the following configurations.
[0088] (Note 1) A situation determination unit that determines whether or not the situation is such that speech can be uttered based on voice data, A timing determination unit determines an appropriate timing for uttering based on the result of the determination by the aforementioned situation determination unit, A notification unit that notifies in a manner appropriate to the situation determined by the situation determination unit, based on the result of the determination by the timing determination unit, has Auxiliary equipment. (Note 2) The auxiliary device described in Appendix 1, The status determination unit determines whether or not the situation is such that speech is possible by checking whether or not predetermined keywords exist in the text data obtained by converting the voice data. Auxiliary equipment. (Note 3) An auxiliary device as described in Appendix 1 or Appendix 2, The situation determination unit makes a positive or negative determination based on the character data obtained by converting the voice data, and determines whether or not the situation is such that speech is possible based on the positive or negative determination result. Auxiliary equipment. (Note 4) An auxiliary device described in any one of the items from Appendix 1 to Appendix 3, The situation determination unit uses a pre-trained model for determining whether or not a situation is one in which speech is possible to determine whether or not a situation is one in which speech is possible. Auxiliary equipment. (Note 5) An auxiliary device described in any one of the items from Appendix 1 to Appendix 4, It has a detection unit that detects the timing at which speech can be uttered based on voice data, The timing determination unit determines the appropriate timing for speech based on the detection result by the detection unit and the determination result by the situation determination unit. Auxiliary equipment. (Note 6) The auxiliary device described in Appendix 5, The timing determination unit determines that the timing is suitable for speaking when the detection unit detects a suitable speaking timing and the situation determination unit determines that the situation is suitable for speaking. Auxiliary equipment. (Note 7) An auxiliary device described in any one of the items from Appendix 1 to Appendix 6, The aforementioned situation determination unit is configured to output information indicating the degree of ease of speaking by determining whether or not the situation is such that speech is possible. The notification unit notifies a predetermined device in a manner that corresponds to the degree of ease of speaking. Auxiliary equipment. (Note 8) An auxiliary device described in any one of the items from Appendix 1 to Appendix 7, The notification unit notifies the device to vibrate in a manner appropriate to the degree of ease of speaking. Auxiliary equipment. (Note 9) Information processing device, Based on the audio data, it is determined whether or not the person is in a state where they can speak. Based on the results of the assessment, the appropriate timing for speaking is determined. Depending on the result of the determination, the situation determination unit will notify in a manner appropriate to the determined situation. Auxiliary methods. (Note 10) In an information processing device, Based on the audio data, it is determined whether or not the person is in a state where they can speak. Based on the results of the assessment, the appropriate timing for speaking is determined. Depending on the result of the determination, the situation determination unit will notify in a manner appropriate to the determined situation. A program to perform the processing.
[0089] Although the present invention has been described above with reference to the embodiments described above, the present invention is not limited to the embodiments described above. Various modifications to the structure and details of the present invention can be made within the scope of the present invention as can be understood by those skilled in the art. [Explanation of symbols]
[0090] 100 Remote conferencing systems 200 Information Processing Devices 300 Processing Units 310 Operation Input Section 320 Imaging Unit 330 Screen display section 340 Communication I / F section 350 Storage section 351 Pre-trained model for atmosphere determination 352 Judgment information 353 Judgment Model 354 Notification method information 355 Image data information 356 Audio data information 357 characters 358 programs 360 arithmetic processing unit 361 Receiver 362 Non-speech interval detection unit 363 Transcription Section 364 Atmosphere determination unit 365 Speech Timing Detection Unit 366 Context determination unit 367 Timing determination unit 368 Notification Department 401 Video Display Unit 402 Label Input Section 403 Graph Display Section 404 Transcription Result Display Section 500 Auxiliary equipment 501 CPU 502 ROM 503 RAM 504 Program Groups 505 Storage device 506 Drive unit 507 Communication Interface 508 Input / Output Interfaces Bus 509 510 Recording media 511 Communication Network 521 Situation Determination Unit 522 Timing determination unit 523 Notification Department
Claims
1. A situation determination unit that determines whether or not the situation is such that speech can be uttered based on voice data, A timing determination unit determines an appropriate timing for uttering based on the result of the determination by the aforementioned situation determination unit, A notification unit that notifies in a manner appropriate to the situation determined by the situation determination unit, based on the result of the determination by the timing determination unit, It has, The status determination unit determines whether or not the situation is such that speech is possible by checking whether or not predetermined keywords exist in the text data obtained by converting the voice data. Auxiliary equipment.
2. An auxiliary device according to claim 1, The situation determination unit makes a positive or negative determination based on the character data obtained by converting the voice data, and determines whether or not the situation is such that speech is possible based on the positive or negative determination result. Auxiliary equipment.
3. An auxiliary device according to claim 1 or claim 2, The situation determination unit uses a pre-trained model for determining whether or not a situation is one in which speech is possible to determine whether or not a situation is one in which speech is possible. Auxiliary equipment.
4. An auxiliary device according to any one of claims 1 to 3, It has a detection unit that detects the timing at which speech can be uttered based on voice data, The timing determination unit determines the appropriate timing for speech based on the detection result by the detection unit and the determination result by the situation determination unit. Auxiliary equipment.
5. The auxiliary device according to claim 4, The timing determination unit determines that the timing is suitable for speaking when the detection unit detects a suitable speaking timing and the situation determination unit determines that the situation is suitable for speaking. Auxiliary equipment.
6. An auxiliary device according to any one of claims 1 to 5, The aforementioned situation determination unit is configured to output information indicating the degree of ease of speaking by determining whether or not the situation is such that speech is possible. The notification unit notifies a predetermined device in a manner that corresponds to the degree of ease of speaking. Auxiliary equipment.
7. An auxiliary device according to any one of claims 1 to 6, The notification unit notifies the device to vibrate in a manner appropriate to the degree of ease of speaking. Auxiliary equipment.
8. A situation determination unit that determines whether or not the situation is such that speech can be uttered based on voice data, A timing determination unit determines an appropriate timing for uttering based on the result of the determination by the aforementioned situation determination unit, A notification unit that notifies in a manner appropriate to the situation determined by the situation determination unit, based on the result of the determination by the timing determination unit, It has, The situation determination unit makes a positive or negative determination based on the character data obtained by converting the voice data, and determines whether or not the situation is such that speech is possible based on the positive or negative determination result. Auxiliary equipment.
9. Information processing device, Based on the audio data, it is determined whether or not the person is in a state where they can speak. Based on the results of the assessment, the appropriate timing for speaking is determined. Depending on the result of the assessment, notification will be given in a manner appropriate to the situation assessed. When determining the situation, the presence or absence of predetermined keywords in the text data obtained by converting the aforementioned audio data is checked to determine whether or not the situation allows for speech. Auxiliary methods.
10. In an information processing device, Based on the audio data, it is determined whether or not the person is in a state where they can speak. Based on the results of the assessment, the appropriate timing for speaking is determined. Depending on the result of the assessment, notification will be given in a manner appropriate to the situation assessed. When determining the situation, the presence or absence of predetermined keywords in the text data obtained by converting the aforementioned audio data is checked to determine whether or not the situation allows for speech. A program to perform the processing.
Citation Information
Patent Citations
Voice input support program, voice input support device and voice input support method
JP2008102384A
Speech recognition device and navigation device using the same
JP2011027905A
Estimation device, estimation method, and estimation program
JP2017146279A
Conversation support system, conversation support device and conversation support program
JP2019049733A
Voice analysis system, voice analysis method, and voice analysis program
JP2019056879A