Voice processing system, voice processing method, and voice processing program

The audio processing system addresses the issue of multiple voice inputs by determining and outputting the most relevant voice, improving accuracy in voice recognition and synthesis in multi-device conference settings.

JP2025113611APending Publication Date: 2025-08-04SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024007857
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-08-04

AI Technical Summary

Technical Problem

In a conference setting with multiple voice devices, the speech of one user is input to multiple devices, leading to the generation of multiple text information for the same spoken voice and reduced accuracy of voice processing.

Method used

An audio processing system that includes an acquisition processing unit to acquire voices from multiple devices, a determination processing unit to determine voice similarity, and an output processing unit to output a specific voice based on similarity thresholds, ensuring accurate voice recognition and synthesis.

Benefits of technology

Improves the accuracy of voice processing by filtering and outputting the most relevant voice, reducing redundant text generation and enhancing the clarity of voice recognition and synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025113611000001_ABST
    Figure 2025113611000001_ABST
Patent Text Reader

Abstract

To provide a voice processing system, a voice processing method, and a voice processing program which can improve the accuracy of voice processing when conversation is conducted by using a plurality of voice devices in the same space.SOLUTION: A conference support device 1 includes: an acquisition processing section 112 for acquiring a voice produced by a user, which is input into each microphone of a plurality of voice devices 2 arranged in the same space; a determination processing section 117 for determining the degree of mutual similarity of multiple voices acquired from each of the plurality of voice devices 2; and an output processing section 118 for outputting a specific first voice among the multiple voices to a voice recognition processing section 113 and a voice synthesis processing section 115 when the degree of the mutual similarity of the multiple voices is a threshold or more.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technique for controlling voices when multiple users have conversations using voice devices individually.

Background Art

[0002] Conventionally, a technique for converting a user's spoken voice into text information and displaying it is known. For example, there is known a conference system capable of summarizing conference text information including text information obtained from the spoken content of conference participants for each predetermined section and sequentially displaying the summary results (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Here, in one conference room, multiple users may hold a conference using voice devices respectively. In this way, when multiple voice devices are used in the same space, the speech of one user is input to the microphones of multiple voice devices, resulting in problems such as the generation of multiple text information for the same spoken voice and the synthesis processing of the same spoken voice, which reduces the accuracy of voice processing.

[0005] An object of the present disclosure is to provide a voice processing system, a voice processing method, and a voice processing program capable of improving the accuracy of voice processing when conversations are held using multiple voice devices in the same space.

Means for Solving the Problems

[0006] An audio processing system according to one aspect of the present disclosure includes an acquisition processing unit, a determination processing unit, an output processing unit, and a display processing unit. The acquisition processing unit acquires voice uttered by a user input to each microphone of a plurality of audio devices arranged in the same space. The determination processing unit determines the similarity between a plurality of voices acquired from each of the plurality of audio devices. The output processing unit outputs a specific first voice among the plurality of voices to an audio processing unit when the similarity between the plurality of voices is equal to or greater than a threshold value.

[0007] An audio processing method according to another aspect of the present disclosure includes acquiring voice uttered by a user input to each microphone of a plurality of audio devices arranged in the same space, determining the similarity between a plurality of voices acquired from each of the plurality of audio devices, and outputting a specific first voice among the plurality of voices to an audio processing unit when the similarity between the plurality of voices is equal to or greater than a threshold value, which is an audio processing method executed by one or more processors.

[0008] An audio processing program according to another aspect of the present disclosure includes acquiring voice uttered by a user input to each microphone of a plurality of audio devices arranged in the same space, determining the similarity between a plurality of voices acquired from each of the plurality of audio devices, and outputting a specific first voice among the plurality of voices to an audio processing unit when the similarity between the plurality of voices is equal to or greater than a threshold value, which is an audio processing program for causing one or more processors to execute.

Advantages of the Invention

[0009] According to the present disclosure, it is possible to provide an audio processing system, an audio processing method, and an audio processing program capable of improving the accuracy of audio processing when having a conversation using a plurality of audio devices in the same space.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7A

Figure 7B

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

[0011] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that the following embodiments are examples of embodying the present disclosure and do not have the character of limiting the technical scope of the present disclosure.

[0012] The information processing system according to the present disclosure can be applied, for example, to a case where a plurality of users use voice devices each having a microphone and a speaker in the same space (for example, a conference room) to have a conversation (conference) with users in another space. Note that the information processing system can also be applied to a case where a plurality of users have a conversation using voice devices in one space. Further, the information processing system can also be applied to a case where one user uses a voice device in one space to have a conversation with a user in another space.

[0013] FIG. 1 shows an application example of the conference support system 100 according to this embodiment. As shown in FIG. 1, users A to D participate in a meeting in conference room R1, and other users (not shown) participate in a meeting in conference room R2. Users A to D each use a neckband-type voice device 2A to 2D that can be worn around the neck to have conversations. The users in conference room R2 may use the voice device 2 or may use one microphone speaker device installed in conference room R2. Note that FIG. 1 shows an example in which each of users A to D uses the voice device 2, but the present invention is not limited to this, and only some users may use the voice device 2. The conference support system 100 is an example of the information processing system and the voice processing system of the present disclosure.

[0014] Each voice device 2 in conference room R1 is wirelessly connected (connected by Bluetooth (registered trademark)) to the conference support device 1, and the voice input to the microphone of each voice device 2 is input to the conference support device 1 and transmitted from the conference support device 1 to the conference terminal 3. The conference application of the conference server 4 transmits the voice received by the conference terminal 3 to conference room R2. As a result, the voice in conference room R1 is output (played back) from the speaker of the voice device 2 (or the microphone speaker device) of the users in conference room R2. Similarly, the voice input to the microphone of the voice device 2 (or the microphone speaker device) in conference room R2 is played back from the speaker of the voice device 2 of each user in conference room R1 by the conference application of the conference server 4.

[0015] In this way, the conference support system 100 is a system that enables a plurality of users to have conversations using the voice device 2 individually in the same space (conference room R1 in FIG. 1). The conference support system 100 may include a display device 5 that can be used in the conference. On the display device 5, conference information such as the camera images of the conference participants and conference materials is displayed by the conference application, or the recognition result (text information) obtained by converting the voice into text by voice recognition processing is displayed.

[0016] As shown in FIG. 1, the conference support system 100 includes a conference support device 1, an audio device 2, a conference terminal 3, and a conference server 4. The audio device 2 is a wirelessly connected audio device equipped with a microphone and a speaker. The audio device 2 may also have functions such as an AI speaker or a smart speaker. The conference support system 100 is a system that includes multiple audio devices 2 and transmits and receives audio data of a user's speech between the multiple audio devices 2. The audio devices 2 may be the same type of audio device or different types of audio devices. For example, the multiple audio devices 2 may include wirelessly connected audio devices and wired connected audio devices. The multiple audio devices 2 may also include neckband-type audio devices, headset-type audio devices, and stationary audio devices.

[0017] The conference supporting device 1 controls the audio (input audio, output audio, etc.) of the audio device 2, and when a conference starts in a conference room, for example, it executes a process of transmitting and receiving audio between multiple audio devices 2. For example, the conference supporting device 1 controls multiple audio devices 2 arranged in the same space. The conference supporting device 1 also stores the audio acquired from the audio devices 2 as audio for recording, and executes a process of converting the acquired audio into text (speech recognition process). Note that the conference supporting device 1 alone may constitute the information processing system and audio processing system of the present disclosure.

[0018] The information processing system of the present disclosure may also have functions for providing various services, such as a conference service, a subtitling (transcription) service using voice recognition, a translation service, and a minutes service. In this embodiment, the conference support system 100 includes a conference server 4 that provides conference services. The conference server 4 provides an online meeting service using a conference application, which is a type of general-purpose software. For example, the conference application is installed in a conference terminal 3. By starting up the conference terminal 3 and logging in, it becomes possible to hold an online meeting using the conference application (for example, an online meeting in conference rooms R1 and R2).

[0019] The conference terminal 3 may be, for example, a user terminal used by a representative user (host) who hosts the conference among the users participating in the conference. For example, the user terminal 3A of user A who is the host functions as the conference terminal 3. In this case, users B to D can start the conference application on the user terminals 3B to 3D and view the conference screen P2 (Fig. 8, etc.).

[0020] [Conference support device 1] As shown in Fig. 2, the conference support device 1 is a device including a control unit 11, a storage unit 12, a communication unit 13, etc. For example, the conference support device 1 is connected to a plurality of audio devices 2 and has functions such as mixing or splitting the audio input from the plurality of audio devices 2 or the conference terminal 3, and an audio recognition function for converting the input audio into text information.

[0021] The communication unit 13 is a communication unit for connecting the conference support device 1 to a communication network by wire or wirelessly and performing data communication according to a predetermined communication protocol with external devices such as the audio device 2, the conference terminal 3, and the display device 5 via the communication network. For example, the communication unit 13 executes pairing processing by the Bluetooth method and wirelessly connects to each audio device 2.

[0022] The storage unit 12 is a non-volatile storage unit such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), or a flash memory for storing various kinds of information. Data such as device information D1 regarding the audio device 2, speech information D2 regarding the speech content, summary information D3 regarding the summary text, and topic information D4 regarding the topic are stored in the storage unit 12.

[0023] FIG. 3 shows an example of device information D1. Information such as a connection ID, a device name, and voice data is registered in the device information D1. The connection ID is identification information (device information) used when connecting the voice device 2, and is, for example, a Bluetooth address. The device name is the device name of the voice device 2. The voice data is data of voice (spoken voice) acquired from the voice device 2. Thus, voice data is stored for each device in the device information D1. Each voice data is stored with the identification information of the device (voice device 2) attached thereto.

[0024] FIG. 4 shows an example of utterance information D2. Information such as a conversation ID, an utterance time, an utterer, an utterance content, and a topic ID is registered in the utterance information D2. The conversation ID is identification information of the utterance content of the user (utterer). The utterance time is the time when the user uttered. The utterer is the name or identification information (user ID) of the user. The utterance content is character information obtained by converting the voice uttered by the user into text information. The topic ID is identification information of the topic of the meeting. When a meeting regarding a predetermined topic is started, the control unit 11 registers each piece of information in the utterance information D2 based on the spoken voice of the user acquired from the voice device 2.

[0025] FIG. 5 shows an example of summary information D3. Information such as an interval ID, a summary ID, a conversation ID, a summary content, a summary creation time, and a topic ID is registered in the summary information D3. The interval ID is identification information capable of identifying an interval (details will be described later) divided according to a predetermined time or a predetermined number of characters. The summary ID is identification information capable of identifying a summary sentence (details will be described later) generated based on the text information converted from the voice. The summary content is the content of the summary sentence, and the summary creation time is the time when the summary sentence was generated. The control unit 11 generates a summary sentence based on the spoken voice of the user during the meeting and registers each piece of information in the summary information D3.

[0026] FIG. 6 shows an example of the issue information D4. In the issue information D4, information such as an issue ID, an issue name, a current issue, a summary ID, and post-meeting summary data is registered. The issue name is the name of the issue. The current issue is information indicating whether the issue is currently under discussion (during the meeting). FIG. 6 shows that the issue regarding the "next-generation office" is currently under discussion. The post-meeting summary data indicates data (minutes data) generated after the meeting ends, and includes content related to the minutes, such as a summary, conclusions, and action items (future directions, etc.). When a meeting starts for an issue designated by, for example, the organizer of the meeting, the control unit 11 registers each piece of information in the issue information D4 based on the speech audio of the user acquired from the audio device 2, and registers the post-meeting summary data in the issue information D4 when the meeting ends.

[0027] In addition, the storage unit 12 stores control programs such as a meeting support program (an example of the information processing program of the present disclosure) for causing the control unit 11 to execute the meeting support process (see FIG. 18) described later. For example, the meeting support program may be non-temporarily recorded on a computer-readable recording medium such as a CD or DVD, and read by a reading device (not shown) such as a CD drive or DVD drive provided in the meeting support device 1 and stored in the storage unit 12.

[0028] The control unit 11 includes control devices such as a CPU, a ROM, and a RAM. The CPU is a processor that executes various arithmetic processes. The ROM is a non-volatile storage unit in which control programs such as BIOS and OS for causing the CPU to execute various arithmetic processes are pre-stored. The RAM is a volatile or non-volatile storage unit that stores various pieces of information, and is used as a temporary storage memory (working area) for various processes executed by the CPU. Then, the control unit 11 controls the meeting support device 1 by causing the CPU to execute various control programs pre-stored in the ROM or the storage unit 12.

[0029] Specifically, as shown in FIG. 2, the control unit 11 includes various processing units such as a setting processing unit 111, an acquisition processing unit 112, a speech recognition processing unit 113, a summary generation processing unit 114, a speech synthesis processing unit 115, and a display processing unit 116. The control unit 11 functions as the various processing units by executing various processes according to the control program using the CPU. Further, some or all of the processing units may be configured by electronic circuits. The control program may be a program for causing a plurality of processors to function as the processing units.

[0030] The setting processing unit 111 sets the topic of a conversation (meeting). Specifically, the setting processing unit 111 accepts operations such as registering a new topic, selecting a topic, and changing the topic currently being discussed from a user (such as the meeting organizer). For example, user A, who is the meeting organizer, starts a meeting application on the user terminal 3A (meeting terminal 3) to display a setting screen P1 (see FIG. 7A). User A can select the "Topic" column on the setting screen P1 to register one or more topics. Also, user A can select one topic from the multiple registered topics on the setting screen P1. FIG. 7B shows an example of the setting screen P11 after the meeting has started (during the meeting). FIG. 7B displays identification information ("During conversation") indicating the topic currently being discussed (e.g., "Next-generation office"). When user A wants to change the topic, user A selects another topic on the setting screen P1. Also, user A can select the "Settings" column on the setting screen P1 to select the users participating in the meeting or select the audio device 2 used in the meeting.

[0031] When the setting processing unit 111 sets a topic in response to a user operation, it registers the identification information (topic ID) and topic name of the topic in the topic information D4 (see FIG. 6). Also, when the meeting starts, the setting processing unit 111 registers information ("Yes") that can identify the topic of the meeting (here, "Next-generation office") in the topic information D4.

[0032] Note that the setting processing unit 111 may permit only a user having a predetermined authority to perform an operation of setting an agenda and an operation of starting a meeting. As another embodiment, the setting processing unit 111 may analyze the speech content of the user to automatically set or change an agenda.

[0033] The acquisition processing unit 112 acquires the voice spoken by the user. Specifically, when the meeting is started and the user speaks, the acquisition processing unit 112 acquires the speech voice input to the microphone of the voice device 2 of the user. Further, the acquisition processing unit 112 acquires time information corresponding to the speech time of the speech voice of the user. For example, the acquisition processing unit 112 acquires the time when the speech voice of the user is input to the microphone of the voice device 2, or the time when the speech voice is acquired.

[0034] When the acquisition processing unit 112 acquires voice from the voice device 2, the voice data is stored in the device information D1 (see FIG. 3) in association with the identification information (connection ID, device name, etc.) of the voice device 2. In the example shown in FIG. 1, the acquisition processing unit 112 acquires the respective speech voices of the users A to D in the conference room R1 from the voice device 2, and stores each voice data in the device information D1 in association with the connection ID.

[0035] The speech recognition processing unit 113 performs a speech recognition process of converting speech into text based on the speech acquired by the acquisition processing unit 112. The speech recognition processing unit 113 uses a predetermined speech recognition engine (trained model) to convert speech into text information. Note that the speech recognition engine is generated by learning various speech data (teacher data), for example. The speech recognition engine is installed in the conference support device 1. Note that the speech recognition processing unit 113 may perform speech processing such as echo cancellation, noise cancellation, and gain adjustment on the speech acquired by the acquisition processing unit 112.

[0036] In addition, the speech recognition processing unit 113 registers the text information as the utterance content in the utterance information D2 (see FIG. 4). Further, the speech recognition processing unit 113 registers the identification information (conversation ID) of the utterance content, the time information (utterance time), and the identification information (name) of the speaker in association with the utterance content. Further, the speech recognition processing unit 113 registers the identification information (topic ID) of the topic corresponding to the utterance content in association with the utterance content. Further, the speech recognition processing unit 113 registers the topic ID of the topic set by the setting processing unit 111 in association with the utterance content. The speech recognition processing unit 113 is an example of the conversion processing unit of the present disclosure.

[0037] The summary generation processing unit 114 generates a summary sentence in a short text format that summarizes the user's utterance content based on the text information converted by the speech recognition processing unit 113. The summary generation processing unit 114 uses a predetermined summary generation engine (trained model) to generate a summary sentence from the text information. Note that the summary generation engine is generated by learning various conversation data and summary data (teacher data), for example. The summary generation engine is installed in the conference support device 1. For example, the summary generation processing unit 114 generates one summary sentence from a plurality of text information corresponding to the speech voices of one or more users.

[0038] In addition, the summary generation processing unit 114 generates a summary sentence for each predetermined section. Specifically, the summary generation processing unit 114 generates one or more summary sentences every predetermined time (for example, 5 minutes) or every predetermined number of characters (for example, 1500 characters) of the text information.

[0039] In addition, the summary generation processing unit 114 registers the summary sentence as summary content in summary information D3 (see FIG. 5). Further, the summary generation processing unit 114 associates and registers the identification information (section ID) of the section, the identification information (summary ID) of the summary sentence, and the identification information (conversation ID) of the conversation corresponding to the text information with the summary content. Further, the summary generation processing unit 114 associates and registers the time when the summary sentence was generated with the summary content. Further, the summary generation processing unit 114 associates and registers the topic ID of the topic set by the setting processing unit 111 with the summary content. In the example shown in FIG. 5, conversations corresponding to a plurality of text information with conversation IDs "101", "102", "103",... are held in the section with section ID "K01", indicating that one summary (summary ID "Y001") is generated from these multiple conversation contents. The number of summary sentences generated in one section may be one or more than one.

[0040] In addition, the summary generation processing unit 114 registers the summary ID in the topic information D4 (see FIG. 6). Note that when the summary generation processing unit 114 generates a plurality of summary sentences for one topic, the plurality of summary IDs corresponding to each summary sentence are associated with the topic (topic ID) and registered in the topic information D4.

[0041] The voice synthesis processing unit 115 synthesizes the voice acquired by the acquisition processing unit 112. Further, the voice synthesis processing unit 115 outputs the synthesized voice to the conference terminal 3 (for example, user terminal 3A) in the conference room R1. The conference application of the conference server 4 transmits the voice received by the conference terminal 3 in the conference room R1 to the conference terminal 3 in the conference room R2.

[0042] The display processing unit 116 causes various types of information corresponding to the conference application to be displayed on the display screen. Specifically, in the conference screen P2, the display processing unit 116 displays the text information converted by the speech recognition processing unit 113 and the summary text corresponding to the topic generated by the summary generation processing unit 114 side by side. For example, as shown in FIG. 8, on the user terminal 3, the display processing unit 116 displays the text information during the conversation in the conversation display area P22 of the conference screen P2, and displays the summary text generated for each predetermined section in the summary display area P21 of the conference screen P2. Note that the display processing unit 116 sequentially adds and displays the text information in real time while scrolling the screen in the conversation display area P22.

[0043] In addition, in the conversation display area P22, the display processing unit 116 associates and displays the identification information (such as the user name and user icon) of the user of the speech corresponding to the text information and the time information (speech time) with the text information (speech content).

[0044] In addition, in the summary display area P21, the display processing unit 116 associates and displays the identification information (for example, the summary creation time T1) that can identify the section with the summary text. For example, when the summary generation processing unit 114 generates a summary text every five minutes, the display processing unit 116 displays the times at five-minute intervals in the summary display area P21 (see FIG. 8).

[0045] Note that in the summary display area P21, the multiple bullet-pointed sentences included in one section (for example, the section from 10:06 to 10:11) are each a summary text. As another embodiment, the summary generation processing unit 114 may generate each of the bullet-pointed sentences as a key point and generate one summary text by summarizing the multiple key points.

[0046] In addition, the display processing unit 116 displays a plurality of topics in the summary display area P21 so that they can be selected, and also displays the topics selected by the user in an identifiable manner. When no topic is selected in the summary display area P21, the display processing unit 116 displays a summary text corresponding to the current topic of conversation. As a result, on the conference screen P2, the text information of the conversation regarding the current topic of conversation and the summary text regarding the topic are displayed side by side. The user can collectively check the current content and summary of the conversation on one conference screen P2.

[0047] Here, when the conference organizer (user A) changes the topic from topic 1 ("Next-generation office") to topic 2 ("Cost reduction plan") on the setting screen P1 (see Fig. 7A), as shown in Fig. 9, the display processing unit 116 displays the text information of the current conversation regarding topic 2 in the conversation display area P22, and also displays the summary text corresponding to topic 2 in the summary display area P21.

[0048] Furthermore, for example, when the current topic of conversation is "topic 2" (see Fig. 9) and a user (for example, users B to D) selects "topic 1" in the summary display area P21, as shown in Fig. 10, the display processing unit 116 displays the summary text corresponding to topic 1 in the summary display area P21. As a result, on the conference screen P2, the text information regarding the current topic of conversation 2 and the summary text regarding topic 1 selected by the user are displayed side by side. Also, in this case, the display processing unit 116 associates an identification mark M1 that can identify the current topic of conversation 2 with the topic 2 in the summary display area P21 so that the user can recognize the current topic of conversation 2. As another embodiment, the display processing unit 116 may display the icon of the current topic of conversation 2 and the icon of topic 1 selected by the user in different display modes in the summary display area P21.

[0049] In this way, when the user wants to check the content of the previous "Agenda 1" after the agenda of the meeting switches from "Agenda 1" (see Figure 8) to "Agenda 2" (see Figure 9), during the conversation of Agenda 2, the user can select "Agenda 1" in the summary display area P21 (see Figure 10) to check the summary text of Agenda 1. The user can check the summary texts of past agendas and the conversation content during the current conversation together on one meeting screen P2.

[0050] In addition, the display processing unit 116 may display user identification information that can identify the user who spoke within a predetermined period, in association with the summary text corresponding to the predetermined period. For example, as shown in Figure 11, the display processing unit 116 displays user icons U1 that can identify the users ("Suzuki", "Takahashi", "Sato") who spoke in the period from 10:01 to 10:06, in association with the summary text corresponding to the period. Thereby, the user can grasp the users who spoke the content (text information) that is the basis of the summary text on the meeting screen P2.

[0051] Here, the display processing unit 116 may execute processing to change the display content of each display area according to user operations on the summary display area P21 and the conversation display area P22 of the meeting screen P2. For example, as shown in Figure 12, when the user presses "10:11", which is the creation time of the summary text in the period from 10:06 to 10:11, in the summary display area P21, or when the user presses the summary text of the period, the display processing unit 116 preferentially displays the text information that is the basis of the summary text in the conversation display area P22. In this case, the display processing unit 116 temporarily stops the display processing of the text information of the speech voice acquired in real time, and fixedly displays the text information that is the basis of the summary text selected by the user. Thereby, the user can select a summary text that he / she is concerned about during the conversation and check the speech content that is the basis of the summary text on the meeting screen P2 where the summary text is displayed.

[0052] In addition, the display processing unit 116 may display a release button K3 for releasing the pause state on the conference screen P2. When the user presses the release button K3, the display processing unit 116 resumes the display processing of the text information of the currently ongoing speech obtained in real time in the conversation display area P22.

[0053] As another embodiment, for example, as shown in FIG. 13, when the user presses the text information (conversation content) in the conversation display area P22, the display processing unit 116 may additionally display the text information in the summary display area P21. For example, when the user determines that there is important information among the plurality of text information displayed in the conversation display area P22, the user selects the text information. Thereby, the display processing unit 116 additionally displays the selected text information as a summary sentence Y1 in the summary display area P21. As another embodiment, the summary generation processing unit 114 may add the text information selected by the user and regenerate the summary sentence. In this case, the summary generation processing unit 114 may assign a weight to the text information selected by the user and regenerate the summary sentence. Thereby, the user can modify the summary sentence as intended. Note that the selection operation of the text information may be, for example, an operation (drag-and-drop operation) in which the user moves the text information from the conversation display area P22 to the summary display area P21.

[0054] In addition, the display processing unit 116 may make the display modes of the summary sentence and the text information different for each topic. For example, as shown in FIG. 14, the display processing unit 116 causes the background image of the topic icon G1 and the background image C11 of the summary sentence to be displayed in a manner corresponding to the topic in the summary display area P21, and causes the background image C12 of the text information to be displayed in a manner corresponding to the topic in the conversation display area P22. For example, as shown in FIG. 14, the display processing unit 116 displays the topic icon G1 of topic 1, the summary sentence of topic 1, and the text information that is the source of the summary sentence of topic 1 in the same display mode (for example, the same background image).

[0055] In addition, the display processing unit 116 causes the scroll bars B1 to B3 corresponding to each of the topics 1 to 3 to be displayed in the conversation display area P22 in a manner corresponding to the topic, and causes each scroll bar to be displayed with a length corresponding to the ratio of the conversation time regarding the topic. Further, the display processing unit 116 causes a position mark B10 indicating the position of the text information being displayed to be displayed on the scroll bar in the conversation display area P22. As a result, the user can grasp the length of the conversation corresponding to each topic, the position of the conversation content currently being displayed, and the like.

[0056] In addition, as shown in FIG. 14, the display processing unit 116 may cause an identification mark M2 capable of identifying the topic to be displayed in the conversation display area P22.

[0057] In addition, in the configuration shown in FIG. 14, for example, when the user selects the topic icon G1 (“Topic 2”), the display processing unit 116 causes a summary text of Topic 2 to be displayed in the summary display area P21 as shown in FIG. 15, and may cause the text information corresponding to the conversation of Topic 2 to be displayed prominently in the conversation display area P22. Here, the position mark B10 is displayed at the head of the scroll bar B2 corresponding to Topic 2.

[0058] In addition, in the configuration shown in FIG. 14, for example, when the user selects the conversation content (text information) of Topic 2 in the conversation display area P22, the display processing unit 116 may cause a summary text corresponding to the selected conversation content of Topic 2 to be prominently displayed in the summary display area P21 as shown in FIG. 16. Note that the display processing unit 116 can extract a summary text (summary ID) corresponding to the conversation content (conversation ID) by referring to the summary information D3 shown in FIG. 5.

[0059] As described above, the control unit 11 converts the voice uttered by the user into text information, generates a summary text based on the text information, and arranges and displays the text information and the summary text on the conference screen P2. In addition, the control unit 11 sets the topic of the conversation, generates the summary text for each topic, and arranges and displays the text information and the summary text corresponding to the topic on the conference screen P2.

[0060] [Ambient voice control] Incidentally, when a plurality of voice devices 2 are used in the same space (same room), the speech voice of one user is input to the microphones of the plurality of voice devices 2, resulting in problems such as the generation of multiple text information for the same speech voice or the synthesis processing of the same speech voice, which reduces the accuracy of voice processing. For example, when user B speaks in conference room R1 (see FIG. 1), the speech voice of user B is not only input to the microphone of voice device 2B of user B, but may also be input (ambient input) to the microphones of voice devices 2A, 2C, and 2D of other users A, C, and D. In this case, problems such as the generation of multiple text information of the same content corresponding to each voice obtained from each voice device 2 or the synthesis processing of multiple voices of the same content occur. Therefore, the conference support device 1 of the present disclosure may be configured to solve the above problems.

[0061] Specifically, as shown in FIG. 17, in addition to each processing unit shown in FIG. 2, the control unit 11 further includes a determination processing unit 117 that determines whether a plurality of voices input from a plurality of voice devices 2 substantially simultaneously are similar to each other, and a voice output processing unit 118 that outputs a voice based on the determination result of the determination processing unit 117.

[0062] The determination processing unit 117 determines the similarity between a plurality of voices input to the respective microphones of a plurality of voice devices 2 arranged in the same space substantially simultaneously. For example, the determination processing unit 117 compares the waveforms of the plurality of voices acquired by the acquisition processing unit 112 to determine the similarity. Further, the determination processing unit 117 determines whether the similarity between the plurality of voices is equal to or greater than a threshold value. The threshold value is set, for example, to a reference value that can identify whether voices heard by a person with the ear are the same voice.

[0063] The voice output processing unit 118 outputs the voice to the voice processing unit (voice recognition processing unit 113, voice synthesis processing unit 115) based on the determination result of the determination processing unit 117. Specifically, when the similarity between the plurality of voices is equal to or greater than a threshold value, the voice output processing unit 118 outputs a specific voice (first voice) among the plurality of voices to the voice recognition processing unit 113 and the voice synthesis processing unit 115. That is, when the plurality of voices acquired by the acquisition processing unit 112 are substantially the same voice, the voice output processing unit 118 outputs one of the plurality of voices to the voice recognition processing unit 113 and the voice synthesis processing unit 115.

[0064] Also, when the similarity between the plurality of voices is equal to or greater than the threshold value, the voice output processing unit 118 outputs the voice with the maximum sound pressure (first voice) among the plurality of voices to the voice recognition processing unit 113 and the voice synthesis processing unit 115. For example, when user B speaks, the sound pressure of the voice input to the microphone of the voice device 2B of user B is the maximum, and the sound pressure of the voice input to each microphone of the voice devices 2A, 2C, and 2D located at a position away from the voice device 2B is lower than the sound pressure of the voice input to the microphone of the voice device 2B. Therefore, the voice output processing unit 118 determines that the voice with the maximum sound pressure among the plurality of similar voices is the normal voice, excludes the other voices, and outputs only the voice with the maximum sound pressure to the voice recognition processing unit 113 and the voice synthesis processing unit 115. As a result, the voice recognition processing unit 113 can execute voice recognition processing based on an appropriate voice, and thus one appropriate summary sentence can be generated by the subsequent processing of the summary generation processing unit 114. Also, since the voice synthesis processing unit 115 can execute synthesis processing based on an appropriate voice, an appropriate voice can be played back in the conference room R2.

[0065] As another embodiment, when the similarity between the plurality of voices is equal to or greater than the threshold value, the voice output processing unit 118 may output the voice (first voice) with the shortest delay time among the plurality of voices to the voice recognition processing unit 113 and the voice synthesis processing unit 115. For example, when user B speaks, compared with the voice input to the microphone of the voice device 2B of user B, the voices input to the microphones of the voice devices 2A, 2C, and 2D located at positions away from the voice device 2B have longer delay times. For this reason, the acquisition processing unit 112 first receives the voice input to the microphone of the voice device 2B of user B, and then, after a predetermined time has elapsed, the voices input to the microphones of the voice devices 2A, 2C, and 2D are sequentially input according to the distance. Therefore, the voice output processing unit 118 determines that the voice with the shortest delay time (the voice that reaches the conference support device 1 the earliest) among the plurality of similar voices is the normal voice, excludes the other voices, and outputs only the voice with the shortest delay time to the voice recognition processing unit 113 and the voice synthesis processing unit 115.

[0066] Note that the control unit 11 may determine the voice to be output based on both the sound pressure and the delay time. For example, the control unit 11 extracts a plurality of voices in descending order of the sound pressure from the plurality of similar voices, and determines the voice with the shortest delay time among the extracted plurality of voices as the voice to be output.

[0067] Also, when the similarity between the plurality of voices is less than the threshold value, the voice output processing unit 118 outputs the plurality of voices to the voice recognition processing unit 113 and the voice synthesis processing unit 115. For example, when a plurality of users speak almost simultaneously, voices with different waveforms are input to the acquisition processing unit 112. In this case, the determination processing unit 117 determines that the plurality of voices are not similar to each other, and the voice output processing unit 118 outputs each of the plurality of voices to the voice recognition processing unit 113 and the voice synthesis processing unit 115. In this case, the voice recognition processing unit 113 sequentially performs voice recognition processing on each voice. Also, the voice synthesis processing unit 115 synthesizes each voice into one voice and outputs it to the conference terminal 3.

[0068] As another embodiment, when the similarity between the plurality of voices is less than the threshold value, the voice output processing unit 118 may output a predetermined number of voices among the plurality of voices to the voice recognition processing unit 113 and output the plurality of voices to the voice synthesis processing unit 115. For example, when 10 voice devices 2 are connected to the conference support device 1, if voice recognition processing and summary text generation processing are performed on the voices input from each of the 10 voice devices 2 substantially simultaneously, it may take a long time to process and there may be a delay in the display of text information and summary text. Therefore, the voice output processing unit 118 selects the top 3 voices in descending order of sound pressure among the 10 voices acquired from each of the 10 voice devices 2 and outputs them to the voice recognition processing unit 113. Note that 10 voices are input to the voice synthesis processing unit 115. Thereby, the voice recognition processing unit 113 executes voice recognition processing based on the specific 3 voices, and the voice synthesis processing unit 115 executes processing for synthesizing 10 voices into one voice.

[0069] According to the above configuration, it is possible to prevent problems such as generation of a plurality of summary texts corresponding to each voice acquired from each voice device 2 or synthesis processing of a plurality of voices with the same content.

[0070] [Conference Support Processing] FIG. 18 shows an example of the procedure of the conference support processing executed by the control unit 11 of the conference support device 1.

[0071] Note that the present disclosure can be regarded as a conference support method (the information processing method and voice processing method of the present disclosure) for executing one or a plurality of steps included in the conference support processing. Also, one or a plurality of steps included in the conference support processing described here may be appropriately omitted. Also, the execution order of each step in the conference support processing may be different within a range that produces the same operational effects. Further, here, the case where the control unit 11 executes each step in the conference support processing is described as an example, but in other embodiments, one or a plurality of processors may execute each step in the conference support processing in a distributed manner.

[0072] Here, as shown in FIG. 1, a case where a meeting is held using a plurality of audio devices 2 arranged in the meeting room R1 will be described as an example. Also, it is assumed that the topic has been registered in advance before the meeting starts. Note that the user can register a topic even during the meeting after it has started.

[0073] First, in step S1, the control unit 11 determines whether an operation to start the meeting has been received. For example, user A, who is the organizer of the meeting, starts the meeting application on the meeting terminal 3 (user terminal 3A) and presses the start button K1 on the setting screen P1 (see FIG. 7A). When the control unit 11 receives the pressing operation of the start button K1, it determines that the meeting start operation has been received. If the control unit 11 determines that the meeting start operation has been received (S1: Yes), the process proceeds to step S2. The control unit 11 waits until the meeting start operation is received (S1: No).

[0074] Next, in step S2, the control unit 11 starts the process of acquiring the voice spoken by the user from the audio device 2. For example, when the meeting starts and the voice spoken by user B is input to the microphone of the audio device 2B, the control unit 11 acquires the voice from the audio device 2B. If the voice of user B is also input to the microphones of the audio devices 2 of other users, the control unit 11 also acquires the voice from the corresponding audio device 2.

[0075] Next, in step S3, the control unit 11 determines whether an operation to set the topic of the meeting has been received. For example, when user A selects the topic of the meeting on the setting screen P1, the control unit 11 receives the selection operation. If the control unit 11 receives the selection operation (S3: Yes), the process proceeds to step S4. The control unit 11 waits until the selection operation is received from the user (S3: No).

[0076] In step S4, the control unit 11 sets the topic selected by the user. When the control unit 11 sets the topic, it registers information related to the topic (topic ID, topic name, etc.) in the topic information D4 (see FIG. 6).

[0077] Next, in step S5, the control unit 11 executes voice output processing. The control unit 11 outputs the voice acquired from one voice device 2 or a plurality of voices acquired from a plurality of voice devices 2 substantially simultaneously to the voice recognition processing unit and the voice synthesis processing unit according to a predetermined condition. FIG. 19 shows a specific example of the voice output processing.

[0078] In step S51 of FIG. 19, the control unit 11 determines whether a plurality of voices are input from each of the plurality of voice devices 2 substantially simultaneously. When the control unit 11 determines that a plurality of voices are input substantially simultaneously (S51: Yes), the process proceeds to step S52. On the other hand, when the control unit 11 determines that a plurality of voices are not input substantially simultaneously (S51: No), the process proceeds to step S55.

[0079] In step S52, the control unit 11 compares the waveforms of the plurality of voices. Next, in step S53, the control unit 11 determines whether the similarity between the plurality of voices is equal to or greater than a threshold value. Specifically, the control unit 11 compares the waveforms of each voice to calculate the similarity between each voice, and determines whether the calculated similarity is equal to or greater than the threshold value. When the control unit 11 determines that the similarity is equal to or greater than the threshold value (S53: Yes), the process proceeds to step S54. On the other hand, when the control unit 11 determines that the similarity is less than the threshold value (S53: No), the process proceeds to step S55.

[0080] In step S54, the control unit 11 outputs the voice with the maximum sound pressure among the plurality of similar voices to the voice recognition processing unit and the voice synthesis processing unit. As another embodiment, the control unit 11 may output the voice with the shortest delay time (the fastest arrival time) among the plurality of similar voices to the voice recognition processing unit and the voice synthesis processing unit.

[0081] In contrast, in step S55, the control unit 11 outputs the input one or more voices to the voice recognition processing unit and the voice synthesis processing unit. For example, when a plurality of voices are not input substantially simultaneously (S51: No), the control unit 11 outputs each voice to the voice recognition processing unit and the voice synthesis processing unit in the order of input. Also, for example, when a plurality of voices are not similar (S53: No), the control unit 11 outputs each voice to the voice recognition processing unit and the voice synthesis processing unit.

[0082] After the voice output process (S5), the control unit 11 causes the process to proceed to steps S6 and S61 (see FIG. 18).

[0083] In step S6, the control unit 11 executes voice recognition processing. Specifically, the control unit 11 uses a predetermined voice recognition engine (trained model) to acquire the voice output in step S5 and convert it into text information. Also, the control unit 11 associates the text information (utterance content) with time information (utterance time), the speaker, and the topic (topic ID) and registers it in the utterance information D2 (see FIG. 4).

[0084] Next, in step S7, the control unit 11 executes a process of generating a summary sentence. Specifically, the control unit 11 uses a predetermined summary generation engine (trained model) to generate a summary sentence from the text information. Also, the control unit 11 generates a summary sentence for each predetermined section. Specifically, the control unit 11 generates a summary sentence every predetermined time (for example, 5 minutes) or every predetermined number of characters (for example, 1500 characters) of the text information.

[0085] The control unit 11 associates the generated summary sentence (summary content) with the section (section ID), the summary ID, the conversation (conversation ID) corresponding to the text information, and the topic ID and registers it in the summary information D3 (see FIG. 5). Also, the control unit 11 registers the summary ID in the topic information D4 (see FIG. 6).

[0086] Next, in step S8, the control unit 11 causes the text information and the summary text to be displayed on the display screen. Specifically, in the conference screen P2 (see FIG. 8), the control unit 11 causes the text information during the conversation to be displayed in the conversation display area P22, and causes the summary text generated for each predetermined section to be displayed in the summary display area P21. Further, the control unit 11 sequentially adds and displays the text information in real time while scrolling the screen in the conversation display area P22. In addition, the control unit 11 causes the identification information (user name, user icon) of the user of the uttered voice corresponding to the text information and the time information (utterance time) to be displayed in association with the text information in the conversation display area P22.

[0087] Further, in the summary display area P21, the control unit 11 causes the identification information (summary creation time) that can identify the section and the user icon U1 that can identify the user who uttered in the section to be displayed in association with the summary text corresponding to the section (see FIG. 11).

[0088] Next, in step S9, the control unit 11 determines whether or not a user operation has been received on the conference screen P2. When the control unit 11 determines that a user operation has been received on the conference screen P2 (S9: Yes), the process proceeds to step S10. On the other hand, when the control unit 11 determines that no user operation has been received on the conference screen P2 (S9: No), the process proceeds to step S11. For example, the user can perform an operation of selecting a topic, an operation of selecting a summary text, an operation of selecting text information (conversation content), etc. on the conference screen P2.

[0089] In step S10, the control unit 11 executes display change processing according to the user operation. For example, when the topic currently in conversation is "Topic 2" (see FIG. 9), if the user selects "Topic 1" in the summary display area P21, the control unit 11 causes the summary sentence corresponding to Topic 1 to be displayed in the summary display area P21 as shown in FIG. 10. In this case, the control unit 11 displays an identification mark M1 that can identify "Topic 2" currently in conversation in the vicinity of "Topic 2" in the summary display area P21.

[0090] Also, for example, as shown in FIG. 12, when the user presses "10:11", which is the creation time of the summary sentence in a specific section (the section from 10:06 to 10:11) in the summary display area P21, or when the user presses the summary sentence of that section, the control unit 11 causes the conversation content (text information) that is the source of the summary sentence to be displayed in the conversation display area P22. When the user presses the cancel button K3, the control unit 11 resumes the display processing of the real-time conversation content in the conversation display area P22.

[0091] Also, for example, as shown in FIG. 13, when the user presses the conversation content (text information) in the conversation display area P22, the control unit 11 additionally displays the text information (summary sentence Y1) in the summary display area P21.

[0092] After step S10, the control unit 11 transfers the process to step S11.

[0093] On the other hand, in step S61, the control unit 11 executes voice synthesis processing. Specifically, when a plurality of similar voices are input almost simultaneously, the control unit 11 executes synthesis processing for the voice with the maximum sound pressure. Also, when a plurality of dissimilar voices are input almost simultaneously, the control unit 11 executes processing to synthesize the plurality of voices into one voice.

[0094] Next, in step S62, the control unit 11 outputs the synthesized voice to the conference terminal 3 (for example, the user terminal 3A). The conference application of the conference server 4 transmits the voice received by the conference terminal 3 in the conference room R1 to the conference terminal 3 in the conference room R2. After step S62, the control unit 11 transfers the process to step S11.

[0095] In step S11, the control unit 11 determines whether an operation to end the conference has been received. For example, when user A ends the conference, the user ends the conference application. When the control unit 11 receives an end operation of the conference application, it determines that an end operation of the conference has been received. When the control unit 11 determines that an end operation of the conference has been received (S11: Yes), the conference support process ends. On the other hand, when the control unit 11 does not receive an end operation of the conference (S11: No), the process returns to step S1.

[0096] Returning to step S1, for example, when user A performs an operation to change the topic, the control unit 11 sets a new topic (S4) and starts a conversation on the topic. For example, the user presses the end button K2 on the setting screen P11 (see FIG. 7B) and selects another topic on the setting screen P1 (see FIG. 7A). When the control unit 11 sets the selected topic (S4), it executes the above-described process (S5 to S10) for the voice of the utterance content regarding the topic. The control unit 11 repeatedly executes the above-described process according to the topic until the conference ends.

[0097] Here, for example, if the user sets "Topic 1" and then changes it to "Topic 2" after a conversation on Topic 1 has taken place, summary texts for each of Topic 1 and Topic 2 are generated. Also, if the user returns to "Topic 1" and a conversation takes place, a summary text of the previous conversation and a summary text of the subsequent conversation regarding Topic 1 are generated respectively. As another embodiment, the control unit 11 may regenerate a summary text that combines the summary text of the previous conversation and the summary text of the subsequent conversation for "Topic 1". In this way, even when the topic is changed to a new topic or returned to the original topic, an appropriate summary text can be generated for each topic. Also, the user can confirm the summary text and conversation content of the previous conversation and the summary text and conversation content of the subsequent conversation for the same topic on the conference screen P2.

[0098] As another embodiment, when the meeting ends, the control unit 11 may generate minutes of the meeting. For example, when a meeting is held on a plurality of topics, the control unit 11 generates minutes that summarize the summary texts for each topic. Also, based on a plurality of summary texts of one topic, the control unit 11 may generate the gist, conclusion, action items, etc. of the topic and compile them into the minutes. The control unit 11 may store the minutes in the storage unit 12 or upload them to a shared folder of a data server (not shown). Each user may be able to access the shared folder on the user terminal 3 and view the minutes.

[0099] As described above, the conference support system 100 according to the present disclosure acquires the voice spoken by the user, converts the voice into text information, generates a summary text based on the text information, and displays the text information and the summary text side by side on the display screen (conference screen P2). Thereby, since the speech content that is the basis of the summary text can be confirmed together with the summary result, it becomes possible to improve the convenience of the function of displaying the summary based on the speech content.

[0100] As another embodiment, the conference support system 100 sets the topic of conversation, acquires the voice spoken by the user, converts the voice into text information, generates a summary text for each topic based on the text information, and displays the text information and the summary text corresponding to the topic side by side on the display screen (conference screen P2). Thereby, it is possible to grasp the summary text for each topic, and also to confirm the utterance content that is the basis of the summary text together with the summary result, so that the convenience can be further improved.

[0101] In each of the above-described embodiments, the conference support system 100 further acquires the voice spoken by the user input to each microphone of the plurality of voice devices 2 arranged in the same space, determines the similarity between the plurality of voices acquired from each of the plurality of voice devices, and when the similarity between the plurality of voices is equal to or greater than a threshold value, a specific first voice among the plurality of voices may be output to the voice processing unit. Thereby, it is possible to prevent a problem that the accuracy of voice processing such as generation of a plurality of summary texts corresponding to each voice input from each voice device 2 substantially simultaneously or synthesis processing of a plurality of voices having the same content decreases.

[0102] Note that in the above-described embodiment, an example of an online conference by connecting the conference rooms R1 and R2 via a network is shown. However, the conference support system 100 of the present disclosure may be configured by only one conference room R1. In this case, for example, in the conference room R1, the conference support device 1 executes a speech-to-text process of displaying the text information obtained by converting the voice input to the microphone of the voice device 2 on the display device 5. Also, one voice device 2 (for example, a stationary microphone speaker device) is installed in the conference room R1, and the conference support device 1 may convert the voice of one or more users input to the microphone of the voice device 2 into text information. That is, when the conference support system 100 has a speech-to-text function, a plurality of voice devices 2 may be arranged for each user, or one voice device 2 may be arranged for a conference room or a plurality of users.

[0103] As another embodiment, the speech recognition processing may include a translation function that converts speech in a first language (e.g., Japanese) into text in a second language (e.g., English). Further, the control unit 11 may, for example, translate and display each of the text information and the summary text, or may display the text information in the conversation display area P22 in the language of the conversation without translation and translate and display only the summary text in the summary display area P21. By adopting a configuration of translating only the summary text, it is possible to reduce the time and cost spent on translation. Further, the control unit 11 may display a translation button on the conference screen P2 that can switch the ON / OFF of the translation function.

[0104] [Appendix 1 of the Disclosure] The following is an appendix on the summary of the disclosure extracted from the above-described embodiments. Note that each configuration and each processing function described in the following appendix can be arbitrarily combined by selection.

[0105] <Appendix 1> An acquisition processing unit that acquires the speech uttered by the user, A conversion processing unit that converts the speech acquired by the acquisition processing unit into text information, A generation processing unit that generates a summary text summarizing the utterance content of the user based on the text information converted by the conversion processing unit, A display processing unit that arranges and displays the text information converted by the conversion processing unit and the summary text generated by the generation processing unit on a display screen, An information processing system comprising the above.

[0106] <Appendix 2> The display processing unit displays the time information corresponding to the utterance time of the speech in association with the text information. The information processing system according to Appendix 1.

[0107] <Appendix 3> The generation processing unit generates the summary text for each predetermined interval, The display processing unit displays the summary text for each predetermined interval. The information processing system according to Supplementary Note 1 or 2.

[0108] <Supplementary Note 4> The generation processing unit generates the summary sentence at predetermined time intervals or for every predetermined number of characters of the text information. The information processing system according to Supplementary Note 3.

[0109] <Supplementary Note 5> The display processing unit causes the predetermined section to be displayed in an identifiable manner. The information processing system according to Supplementary Note 3 or 4.

[0110] <Supplementary Note 6> The display processing unit displays, in association with the summary sentence corresponding to the predetermined section, user identification information that can identify the user who spoke within the predetermined section. The information processing system according to any one of Supplementary Notes 3 to 5.

[0111] <Supplementary Note 7> The display processing unit preferentially displays the text information that is the source of the summary sentence selected by the user among the plurality of displayed summary sentences. The information processing system according to any one of Supplementary Notes 1 to 6.

[0112] <Supplementary Note 8> The display processing unit displays the translated sentence of the summary sentence in association with the summary sentence. The information processing system according to any one of Supplementary Notes 1 to 7.

[0113] [Supplementary Note 2 of the Disclosure] The following is a supplementary note on the outline of the disclosure extracted from the above-described embodiment. Note that each configuration and each processing function described in the following supplementary note can be arbitrarily combined by making selections.

[0114] <Supplementary Note 1> A setting processing unit that sets the topic of conversation, An acquisition processing unit that acquires the voice spoken by the user, A conversion processing unit that converts the voice acquired by the acquisition processing unit into text information; A generation processing unit that generates a summary sentence summarizing the user's speech content for each of the topics set by the setting processing unit based on the text information converted by the conversion processing unit; A display processing unit that displays the text information and the summary sentence corresponding to the topic generated by the generation processing unit side by side on a display screen; An information processing system comprising:

[0115] <Appendix 2> The display processing unit displays the text information currently in conversation in a first area of the display screen, and displays the summary sentence corresponding to the topic selected by the user in a second area of the display screen. The information processing system according to Appendix 1.

[0116] <Appendix 3> The display processing unit displays a plurality of topics set by the setting processing unit on the display screen so that they can be selected, and displays the summary sentence corresponding to the topic selected by the user among the plurality of topics. The information processing system according to Appendix 1 or 2.

[0117] <Appendix 4> The display processing unit preferentially displays the text information that is the source of the summary sentence corresponding to the topic selected by the user among the text information converted by the conversion processing unit. The information processing system according to Appendix 3.

[0118] <Appendix 5> The display processing unit displays the text information that is the source of the summary sentence corresponding to the topic selected by the user and the text information that is the source of the summary sentence corresponding to the topic not selected by the user in different display modes among the text information converted by the conversion processing unit. The information processing system described in Supplementary Note 3 or 4.

[0119] <Supplementary Note 6> The display processing unit causes the summary text generated based on the text information selected by the user from among the text information displayed on the display screen to be displayed on the display screen. The information processing system according to any one of Supplementary Notes 3 to 5.

[0120] <Supplementary Note 7> The display processing unit causes the current topic of conversation and the topic selected by the user to be displayed on the display screen in a distinguishable manner. The information processing system according to any one of Supplementary Notes 3 to 6.

[0121] <Supplementary Note 8> The display processing unit causes time information corresponding to the speaking time of the voice to be displayed in association with the text information. The information processing system according to any one of Supplementary Notes 1 to 7.

[0122] <Supplementary Note 9> The generation processing unit generates the summary text at predetermined time intervals or for every predetermined number of characters of the text information. The display processing unit causes the summary text to be displayed at predetermined time intervals or for every predetermined number of characters. The information processing system according to any one of Supplementary Notes 1 to 8.

[0123] <Supplementary Note 10> The setting processing unit accepts, from the user, at least one of an operation to add a new topic, an operation to start a conversation by specifying a topic, and an operation to change the current topic of conversation, on a setting screen different from the display screen. The information processing system according to any one of Supplementary Notes 1 to 9.

[0124] <Supplementary Note 11> The display processing unit causes the current topic of conversation to be displayed in a distinguishable manner on the setting screen. 11. The information processing system of claim 10.

[0125] [Disclosure Note 3] The following will provide an outline of the disclosure extracted from the above-described embodiment. Note that the configurations and processing functions described in the following supplementary notes can be selected and combined as desired.

[0126] <Appendix 1> an acquisition processing unit that acquires voices uttered by a user input to respective microphones of a plurality of audio devices arranged in the same space; a determination processing unit that determines a degree of similarity between a plurality of sounds acquired from each of the plurality of audio devices; an output processing unit that outputs a specific first voice from among the plurality of voices to a voice processing unit when a similarity between the plurality of voices is equal to or greater than a threshold; A voice processing system comprising:

[0127] <Appendix 2> the output processing unit outputs the first voice having the largest sound pressure among the plurality of voices to the voice processing unit when the similarity between the plurality of voices is equal to or greater than the threshold value; 10. The speech processing system of claim 1.

[0128] <Appendix 3> the output processing unit outputs the first voice having the shortest delay time among the plurality of voices to the voice processing unit when the similarity between the plurality of voices is equal to or greater than the threshold value. 3. The speech processing system according to claim 1 or 2.

[0129] <Appendix 4> the output processing unit outputs the plurality of sounds to the sound processing unit when the similarity between the plurality of sounds is less than the threshold value; 4. A speech processing system according to any one of Supplementary notes 1 to 3.

[0130] <Appendix 5> the determination processing unit compares the waveforms of the plurality of sounds to determine the degree of similarity; The voice processing system according to any one of Supplementary Notes 1 to 4.

[0131] <Supplementary Note 6> The voice processing unit executes at least one of a voice conversion process for converting the voice acquired by the acquisition processing unit into text information and a voice synthesis process for synthesizing the voice acquired by the acquisition processing unit. The voice processing system according to any one of Supplementary Notes 1 to 5.

[0132] <Supplementary Note 7> The voice processing unit When the similarity between the plurality of voices is equal to or greater than the threshold value, converts the first voice into text information. When the similarity between the plurality of voices is less than the threshold value, converts each of the plurality of voices into text information. The voice processing system according to any one of Supplementary Notes 1 to 6.

[0133] <Supplementary Note 8> When the similarity between the plurality of voices is less than the threshold value, the output processing unit outputs a predetermined number of voices among the plurality of voices to the voice processing unit that executes the voice conversion process, and outputs the plurality of voices to the voice processing unit that executes the voice synthesis process. The voice processing system according to Supplementary Note 6.

Explanation of Reference Numerals

[0134] 100: Conference support system 1: Conference support device 2: Voice device 3: User terminal, conference terminal 4: Conference server 5: Display device 11: Control unit 12: Storage unit 13: Communication unit 111: Setting processing unit 112: Acquisition processing unit 113: Voice recognition processing unit 114: Summary generation processing unit 115: Voice synthesis processing unit 116: Display processing unit 117: Judgment processing unit 118: Output processing unit D1: Equipment information D2: Utterance information D3: Summary information D4: Topic information G1: Topic icon K1: Start button K2: End button K3: Cancel button M1: Identification mark M2: Identification mark P1: Setting screen P11: Setting screen P2: Conference screen P21: Summary display area P22: Conversation display area

Claims

1. An acquisition processing unit that acquires voice spoken by a user input to each microphone of a plurality of voice devices arranged in the same space; A determination processing unit that determines the similarity between a plurality of voices acquired from each of the plurality of voice devices; An output processing unit that outputs a specific first voice among the plurality of voices to a voice processing unit when the similarity between the plurality of voices is equal to or greater than a threshold value; A voice processing system comprising:

2. When the similarity between the plurality of voices is equal to or greater than the threshold value, the output processing unit outputs the first voice having the maximum sound pressure among the plurality of voices to the voice processing unit. The voice processing system according to claim 1.

3. When the similarity between the plurality of voices is equal to or greater than the threshold value, the output processing unit outputs the first voice having the shortest delay time among the plurality of voices to the voice processing unit. The voice processing system according to claim 1 or 2.

4. When the similarity between the plurality of voices is less than the threshold value, the output processing unit outputs the plurality of voices to the voice processing unit. The voice processing system according to claim 1.

5. The determination processing unit determines the similarity by comparing the waveforms of the plurality of voices. The voice processing system according to claim 1.

6. The voice processing unit executes at least one of a voice conversion process of converting the voice acquired by the acquisition processing unit into text information and a voice synthesis process of synthesizing the voice acquired by the acquisition processing unit. The voice processing system according to claim 1.

7. The voice processing unit When the similarity between the plurality of voices is equal to or greater than the threshold value, converts the first voice into text information; When the similarity between the plurality of voices is less than the threshold value, converts each of the plurality of voices into text information. The voice processing system according to claim 1.

8. When the similarity between the plurality of voices is less than the threshold value, the output processing unit outputs a predetermined number of voices among the plurality of voices to the voice processing unit that executes the voice conversion process, and outputs the plurality of voices to the voice processing unit that executes the voice synthesis process. The voice processing system according to claim 6.

9. Acquiring voice spoken by a user input to each microphone of a plurality of voice devices arranged in the same space; Determining the similarity between a plurality of voices acquired from each of the plurality of voice devices; When the similarity between the plurality of voices is equal to or greater than a threshold value, output a specific first voice among the plurality of voices to a voice processing unit; A voice processing method executed by one or more processors. **Claim 10** Obtain voices spoken by a user input to microphones of respective ones of a plurality of voice devices arranged in the same space; Determine the similarity between the plurality of voices obtained from respective ones of the plurality of voice devices; When the similarity between the plurality of voices is equal to or greater than a threshold value, output a specific first voice among the plurality of voices to a voice processing unit; A voice processing program for causing one or more processors to execute.

Citation Information

Patent Citations

  • Conference system, summary device, control method of conference system, control method of summary device, and program

    JP2019105740A