Voice recognition result display system, voice recognition result display device, voice recognition result display method and program
The speech recognition system addresses the challenge of long text displays by segmenting text into manageable blocks, improving readability and comprehension in conference settings.
Patent Information
- Application Number
- JP2022056877
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-30
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2042-03-30
AI Technical Summary
In conference settings, particularly when hearing-impaired individuals participate, long uninterrupted speech segments lead to prolonged text displays that become difficult to read, hindering understanding.
A speech recognition system that generates image data to display text information as a single block for short segments and divides it into multiple blocks when the character count exceeds a threshold, ensuring each block remains below the threshold, with additional features like text framing, numbering, and color unification to enhance readability.
Prevents text information from becoming difficult to read by segmenting long utterances, making it easier to comprehend, especially for hearing-impaired participants.
Smart Images

Figure 0007770229000001 
Figure 0007770229000002 
Figure 0007770229000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a speech recognition result display system, a speech recognition result display device, a speech recognition result display method, and a program. [Background technology]
[0002] Conventionally, when multiple participants are in a conference, there is known a device that converts the content of speech (speech) of each participant into text and displays it (see, for example, Patent Document 1). Such a device is sometimes used to support hearing-impaired people in participating in conferences. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-179480 Summary of the Invention [Problem to be solved by the invention]
[0004] For example, in the device described in Patent Document 1, if a conference participant continues speaking for a long period of time without interruption, the text display may also continue for a long period of time without interruption. In this case, it may be difficult for other participants to read the text. This problem is particularly pronounced when hearing-impaired people participate in a conference, because hearing-impaired people understand the content of the conference by reading the text.
[0005] The present invention has been made in consideration of the above-mentioned problems, and aims to provide a speech recognition result display system, a speech recognition result display device, a speech recognition result display method, and a program that can prevent text information from becoming difficult to read when a single speech segment becomes long. [Means for solving the problem]
[0006] (1) In order to achieve the above object, a speech recognition result display system (1) according to one embodiment of the present invention includes an acquisition unit (222) that acquires a user's speech, a speech recognition unit (a speech recognition unit 223, a text conversion unit 224, a dependency analysis unit 225) that performs speech recognition on the speech acquired by the acquisition unit and outputs text information, an image generation unit (a minutes creation unit 226, an image generation unit 241) that generates image data based on the text information, and a display unit (a display unit 203, a display unit 303) that displays the image data, wherein the image generation unit generates the image data that displays the text information included in one speech section of the user as a single block when the number of characters of the text information included in the one speech section is less than a predetermined value, and performs an image data generation process that generates the image data that displays the text information included in the one speech section divided into multiple blocks when the number of characters of the text information included in the one speech section is equal to or greater than the predetermined value, and the number of characters of the text information included in each of the multiple blocks is less than the predetermined value.
[0007] (2) In the speech recognition result display system according to one aspect of the present invention, the image generation unit may perform control so that the single speech section is not divided into the multiple segments midway through a phrase.
[0008] (3) In the speech recognition result display system according to one aspect of the present invention, the image generating unit may add a text frame to each of the plurality of blocks when generating the image data.
[0009] (4) In addition, in a speech recognition result display system according to one aspect of the present invention, when the text information included in one speech section spans multiple text frames, the image generation unit may assign consecutive numbers to the multiple text frames.
[0010] (5) In addition, in a speech recognition result display system according to one aspect of the present invention, when the text information included in the single speech section spans multiple text frames, the image generation unit may unify the colors of the multiple text frames.
[0011] (6) In the speech recognition result display system according to an aspect of the present invention, the image generation unit may generate the image data that distinguishes each of the plurality of groups by inserting a line break.
[0012] (7) In the speech recognition result display system according to an aspect of the present invention, the image generation unit may perform the image data generation process every time the text information is output from the speech recognition unit.
[0013] (8) In order to achieve the above object, a speech recognition result display device according to one aspect of the present invention includes an acquisition unit that acquires a user's speech, a speech recognition unit that performs speech recognition on the speech acquired by the acquisition unit and outputs text information, an image generation unit that generates image data based on the text information, and a display unit that displays the image data, wherein the image generation unit generates the image data that displays the text information included in one speech section of the user as a single block when the number of characters of the text information included in the one speech section is less than a predetermined value, and performs an image data generation process that generates the image data that displays the text information included in the one speech section divided into multiple blocks when the number of characters of the text information included in the one speech section is equal to or greater than the predetermined value, and the number of characters of the text information included in each of the multiple blocks is less than the predetermined value.
[0014] (9) In order to achieve the above-mentioned object, a speech recognition result display method according to one aspect of the present invention is a speech recognition result display method in a speech recognition result display system, comprising: an acquisition step in which an acquisition unit acquires a user's spoken speech; a speech recognition step in which a speech recognition unit performs speech recognition on the spoken speech acquired by the acquisition unit and outputs text information; an image generation step in which an image generation unit generates image data based on the text information; and a display step in which a display unit displays the image data, wherein in the image generation step, if the number of characters of the text information included in one utterance section of the user is less than a predetermined value, the image generation unit generates the image data that displays the text information included in the one utterance section as a single block; and if the number of characters of the text information included in the one utterance section is equal to or greater than the predetermined value, the image generation unit performs image data generation processing that generates the image data that displays the text information included in the one utterance section divided into multiple blocks, and the number of characters of the text information included in each of the multiple blocks is less than the predetermined value.
[0015] (10) In order to achieve the above object, a program according to one aspect of the present invention causes a speech recognition result display system to execute an acquisition step of acquiring a user's spoken speech, a speech recognition step of performing speech recognition on the spoken speech and outputting text information, an image generation step of generating image data based on the text information, and a display step of displaying the image data, wherein in the image generation step, if the number of characters of the text information included in one speech section of the user is less than a predetermined value, the image data is generated to display the text information included in the one speech section as a single block, and if the number of characters of the text information included in the one speech section is equal to or greater than the predetermined value, an image data generation process is performed to generate the image data to display the text information included in the one speech section divided into multiple blocks, and the number of characters of the text information included in each of the multiple blocks is less than the predetermined value. [Effects of the Invention]
[0016] According to the above (1), (8), (9) or (10), it is possible to prevent the text information from becoming difficult to read when the length of one utterance segment becomes long.
[0017] According to (2) above, text information divided into multiple chunks can be made easier to read. According to (3) above, each group can be easily visually recognized, making the text information even easier to read. According to the above (4) or (5), it becomes easier to visually recognize a group corresponding to one speech section, and it is possible to make the text information easier to read more effectively. According to (6) above, each group can be easily visually recognized, making the text information even easier to read. According to the above (7), the content of speech by each participant can be reflected more quickly in the display on the display unit. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a block diagram showing an example of the configuration of a speech recognition result display device (speech recognition result display system) according to an embodiment of the present invention. [Figure 2] 10A and 10B are diagrams illustrating an example of an image displayed on a display unit of the parent device according to the embodiment. [Figure 3] 10A and 10B are diagrams showing other example images displayed on the display unit of the parent device according to the embodiment. [Figure 4] 10A and 10B are diagrams illustrating an example of an image displayed on a display unit of the slave device according to the embodiment. [Figure 5] 10 is a flowchart showing an example of a processing procedure performed by a speech recognition result display device (speech recognition result display system) according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0019] First, an example of a situation in which the speech recognition result display device (speech recognition result display system) of this embodiment is used will be described. The speech recognition result display device (speech recognition result display system) of this embodiment is used in a conference attended by two or more participants. Some of the participants may have difficulty speaking. Participants who can speak wear a microphone or use a terminal equipped with a microphone (smartphone, tablet terminal, personal computer, etc.). Participants who are hearing impaired use a terminal into which text can be input. The speech recognition result display device recognizes and converts the speech uttered by the participants into text, and displays the text on each participant's terminal. The speech recognition result display device also displays text information input by the hearing impaired participant as text on each participant's terminal.
[0020] FIG. 1 is a block diagram showing an example of the configuration of a speech recognition result display device (speech recognition result display system) 1 according to this embodiment. 1, a voice recognition result display device (voice recognition result display system) 1 includes a parent device 2 and child devices 3a, 3b, etc. When one of the child devices 3a, 3b, etc. is not specified, it will be simply referred to as a child device 3. The master unit 2 and the slave unit 3 are connected via a wired or wireless network 4 .
[0021] The base unit 2 includes a sound pickup unit 201, an operation unit 202, a display unit 203, a communication unit 204, an authentication unit 211, an acoustic model / dictionary storage unit 221, an acquisition unit 222, a voice recognition unit 223, a text conversion unit 224, a dependency analysis unit 225, a minutes creation unit 226, a minutes storage unit 227, a text acquisition unit 231, and an image generation unit 241.
[0022] The slave unit 3 includes a sound pickup unit 301, an operation unit 302, a display unit 303, a communication unit 304, and a processing unit 305. The sound pickup unit 301, the operation unit 302, the display unit 303, the communication unit 304, and the processing unit 305 are connected via a bus 306.
[0023] <Handset 3> First, the slave unit 3 will be described. The slave device 3 is, for example, a smartphone, a tablet terminal, a personal computer, etc. The slave device 3 may include an audio output unit, a motion sensor, a GPS (Global Positioning System), etc.
[0024] The sound collection unit 301 is a microphone. The sound collection unit 301 collects the user's speech, converts the collected speech from an analog signal to a digital signal, and outputs the converted digital signal to the processing unit 305.
[0025] The operation unit 302 detects a user operation and outputs the detection result to the processing unit 305. The operation unit 302 is, for example, a touch panel sensor provided on the display unit 303, or a wired or wirelessly connected keyboard.
[0026] The processing unit 305 generates setting information based on the operation result of the operation unit 302, and outputs the generated setting information to the communication unit 304. Here, the setting information includes identification information of the participants. The setting information may include information indicating whether the sound collection unit is being used and information indicating whether the operation unit is being used. The processing unit 305 generates a login instruction based on the operation result of the operation unit 302, and outputs the generated login instruction to the communication unit 304. Here, the login instruction includes identification information of the participants and identification information of the remote unit 3. The processing unit 305 adds identification information to text information based on the operation result of the operation unit 302, and outputs the text information to the communication unit 304. The processing unit 305 adds identification information to the speech output by the sound collection unit 301, and outputs the text information to the communication unit 304. The processing unit 305 acquires image data output by the communication unit 304, and outputs the acquired image data to the display unit 303. The processing unit 305 establishes communication with the parent unit 2 based on the login permission information output by the communication unit 304. When the processing unit 305 receives a speech restriction instruction (input restriction instruction) from the parent unit 2, it may impose a restriction on text input. Furthermore, when the processing unit 305 receives a speech restriction instruction from the parent unit 2, it may also impose a restriction on voice input.
[0027] The display unit 303 displays the image data output by the processing unit 305. The display unit 303 is, for example, a liquid crystal display device, an organic EL (electroluminescence) display device, an electronic ink display device, etc. The image displayed on the display unit 303 will be described later.
[0028] The communication unit 304 transmits the setting information output by the processing unit 305 to the base unit 2 via the network 4. The communication unit 304 transmits the login instruction output by the processing unit 305 to the base unit 2 via the network 4. The communication unit 304 transmits the text information or spoken voice output by the processing unit 305 to the base unit 2 via the network 4. The transmitted text information or spoken voice includes the user's identification information and the identification information of the handset 3. The communication unit 304 receives image data transmitted by the base unit 2 and outputs the received image data to the processing unit 305. When the communication unit 304 receives information permitting login transmitted by the base unit 2, it outputs the received information permitting login to the processing unit 305.
[0029] <Base unit 2> Next, the base unit 2 will be described. The base unit 2 is, for example, a notebook computer.
[0030] The sound collection unit 201 is a microphone. The sound collection unit 201 collects the user's speech, converts the collected speech from an analog signal to a digital signal, and outputs the converted digital speech to the acquisition unit 222.
[0031] The operation unit 202 detects a user operation and outputs the detection result to the text acquisition unit 231. The operation unit 202 is, for example, a touch panel sensor or a keyboard provided on the display unit 203. During login processing, the operation unit 202 outputs the operation detection result to the authentication unit 211.
[0032] The display unit 203 is, for example, a liquid crystal display device, an organic EL display device, an electronic ink display device, etc. The display unit 203 displays image data output by the image generation unit 241. The image displayed on the display unit 203 will be described later.
[0033] The communication unit 204 receives the speech transmitted by the handset 3 and outputs the received speech to the acquisition unit 222. The communication unit 204 receives the text information transmitted by the handset 3 and outputs the received text information to the text acquisition unit 231. The communication unit 204 receives the login instruction transmitted by the handset 3 and outputs the received login instruction to the authentication unit 211. The communication unit 204 transmits the image data output by the image generation unit 241 to the handset 3 via the network 4. The communication unit 204 transmits the login permission information output by the authentication unit 211 to the handset 3 via the network 4.
[0034] The authentication unit 211 determines whether to permit login based on the identification information of the participant and the identification information of the slave unit 3 included in the login instruction output by the communication unit 204. If the login is permitted, the authentication unit 211 outputs information permitting login to the communication unit 204. The authentication unit 211 determines whether to permit login of the user of the master unit 2 based on the result of operation of the operation unit 202. If the login is permitted, the authentication unit 211 outputs information permitting login to each functional unit and permits operation of each functional unit. The functional units are the communication unit 204, the authentication unit 211, the acoustic model and dictionary storage unit 221, the acquisition unit 222, the speech recognition unit 223, the text conversion unit 224, the dependency analysis unit 225, the minutes creation unit 226, the minutes storage unit 227, the text acquisition unit 231, and the image generation unit 241.
[0035] The acoustic model and dictionary storage unit 221 stores, for example, an acoustic model, a language model, a word dictionary, etc. An acoustic model is a model based on sound features, and a language model is a model of information about words and their arrangement. A word dictionary is a dictionary with a large vocabulary, such as a large vocabulary word dictionary. The base unit 2 may store and update words not stored in the acoustic model and dictionary storage unit 221. The acoustic model and dictionary storage unit 221 may have a database (DB) for each conference, for example. For example, a first DB may be for general conferences, a second DB for presentations, and a third DB for international conferences. Using a DB tailored to the conference in this way makes it easier to appropriately convert homonyms, etc.
[0036] The acquisition unit 222 acquires the speech output by the sound collection unit 201 or the speech output by the communication unit 204 , and outputs the acquired speech to the speech recognition unit 223 .
[0037] The speech recognition unit 223 acquires the speech output by the acquisition unit 222. The speech recognition unit 223 detects a speech section from the speech. For example, the speech section is detected by detecting a speech signal equal to or greater than a predetermined threshold. The speech recognition unit 223 may detect the speech section using another well-known method. The speech recognition unit 223 performs speech recognition on the detected speech section speech signal by referring to the acoustic model and dictionary storage unit 221 using a well-known method. The speech recognition unit 223 performs speech recognition using, for example, the method disclosed in Japanese Patent Application Laid-Open No. 2015-64554. The speech recognition unit 223 outputs the recognized recognition result and the speech signal to the text conversion unit 224. The speech recognition unit 223 outputs the recognition result and the speech signal in association with each other, for example, for each sentence, for each speech section, or for each utterance. When speech sounds are simultaneously input, the speech recognition unit 223 performs speech recognition for each sound collection unit (201 or 301) by, for example, time-division processing. When the microphone is a microphone array, the speech recognition unit 223 also performs well-known speech recognition processing such as sound source separation processing, sound source localization processing, and sound source identification processing.
[0038] The text conversion unit 224 converts the recognition result output by the speech recognition unit 223 into text by referring to the acoustic model and dictionary storage unit 221. The text information includes information of at least one character. The text conversion unit 224 outputs the converted text information and the acquired voice signal to the dependency analysis unit 225. The text conversion unit 224 may convert the recognition result of the utterance information into text by deleting interjections such as "ah," "um," "eh," and "well," etc.
[0039] The dependency analysis unit 225 performs morphological analysis and dependency analysis on the text information output by the text conversion unit 224 or the text information output by the communication unit 204, by referring to the acoustic model and dictionary storage unit 221. For the dependency analysis, for example, a shift-reduce method, a spanning tree method, or an SVM (Support Vector Machines) in a stepwise application method for chunk identification is used. The dependency analysis unit 225 outputs the text information on which the morphological analysis and dependency analysis have been performed, and the analysis results, to the minutes creation unit 226. When the dependency analysis unit 225 acquires an audio signal output by the text conversion unit 224, it also outputs the acquired audio signal to the minutes creation unit 226.
[0040] The minutes creating unit 226 creates minutes for each speaker based on the text information output by the dependency analysis unit 225 or the text acquisition unit 231. The minutes creating unit 226 stores the created minutes and an audio signal corresponding to them in the minutes storage unit 227. The minutes creating unit 226 also outputs the created minutes to the image generation unit 241. Note that the minutes creating unit 226 may create the minutes by deleting interjections such as "ah," "er," "uh," and "well," etc.
[0041] The minutes storage unit 227 stores the minutes and the audio signals in association with each other.
[0042] The text acquisition unit 231 acquires the operation result output by the operation unit 202 or the operation result of the operation unit 302 output by the communication unit 204, and generates text information based on the acquired result. The text acquisition unit 231 outputs the generated text information to the minutes creation unit 226.
[0043] The image generating unit 241 acquires the minutes information output by the minutes creating unit 226. The image generating unit 241 generates image data based on the minutes information. The image generation unit 241 generates image data corresponding to the number of characters of text information included in one utterance section of the user for text information (minutes information) based on the user's speech. In other words, if the number of characters included in one utterance section is less than a predetermined value, the image generation unit 241 generates image data that displays the text information included in one utterance section as a single block. On the other hand, if the number of characters of the text information included in one utterance section is equal to or greater than the predetermined value, the image generation unit 241 generates image data that displays the text information included in one utterance section divided into multiple blocks. In this case, the image generation unit 241 generates image data so that the number of characters of the text information included in each of the multiple blocks is less than the predetermined value. Note that a "single utterance section" refers to a section (time) in which a user continues speaking without interrupting the utterance for a predetermined period of time or more. The "predetermined value" may be, for example, approximately 500 characters. The image generation unit 241 outputs the image data generated based on the above to the display unit 203 and the communication unit 204.
[0044] <Display image of base unit 2> Next, an example of an image displayed on the display unit 203 of the base unit 2 will be described. FIG. 2 is a diagram showing an example of an image displayed on the display unit 203 of the parent device 2 according to this embodiment. An image g10 is the image displayed on the display unit 203 of the parent device 2.
[0045] Area g100 is an area where participant information is edited. Area g101 is an area for participant information. Reference symbol g102 is the name of the participant. Reference symbol g103 is an icon indicating that the participant will input text using the operation unit 202 of the parent unit 2 or the operation unit 302 of the child unit 3. Reference symbol g104 is an icon indicating that the participant will speak using the sound pickup unit 201 of the parent unit 2 or the sound pickup unit 301 of the child unit 3. Reference symbol g105 is the number (or identification information) of the microphone used by the participant.
[0046] Area g200 is an area where the minutes are displayed. Note that FIG. 2 shows the state after login. Reference symbol g201 is a login / logout button image. Reference symbol g202 is a start / end button image of the voice recognition result display device (voice recognition result display system) 1. Reference symbol g203 is an indication that lights up while the voice recognition result display device (voice recognition result display system) 1 is in use. Reference symbol g204 is a button image that displays the minutes stored in the minutes storage unit 227 and plays back audio signals. Reference symbol g205 is a button image that allows the user of the base unit 2 to select whether or not to use the sound pickup unit 201.
[0047] Reference symbol g211 is text information input by the first participant by operating the operation unit (202 or 302). As shown in the example of FIG. 2, in the image g10 displayed by the display unit 203, a rectangular text frame surrounding the text information is attached to the text information. A text frame is also attached to the text information described below. Reference symbol g212 is a pictogram input by the first participant by operating the operation unit (202 or 302). Reference symbol g213 is information indicating the date and time when the first participant input the text information and pictogram. Reference symbol g214 is the name of the first participant.
[0048] Reference symbol g221 is text information input by the second participant by operating the operation unit (202 or 302). Reference symbols g222 and g223 are text information obtained by voice recognition of what was spoken by the second participant. Reference symbol g224 is an icon indicating that text was input by operating the operation unit (202 or 302). Reference symbol g225 is an icon indicating that speech was input using the sound collection unit (201 or 301). Reference symbol g231 is text information obtained by voice recognition of what was spoken by the third participant.
[0049] Here, the text information related to the symbol g222, "Today's agenda is about the upcoming exhibition," is included in one speech section of the user (Fujisawa), and the number of characters is less than a predetermined value, so it is displayed as a single block (text frame). On the other hand, the text information related to the symbol g223, "At the last meeting, please tell me...," is included in one speech section of the user (Fujisawa), but the number of characters is greater than or equal to the predetermined value, so it is displayed as multiple blocks (text frames) indicated by the symbols g223a to g223d. Furthermore, the number of characters included in each of the text frames indicated by the symbols g223a to g223d is less than the predetermined value.
[0050] 2, the display unit 203 may display image data in which consecutive numbers indicated by reference symbols g241 to g244 are assigned to the text frames indicated by reference symbols g223a to g223d. The display unit 203 may also display image data in which the colors of the text frames indicated by reference symbols g223a to g223d are the same.
[0051] In order for the display unit 203 to perform the above display, the image generation unit 241 generates image data in which a text frame is assigned to each group. The image generation unit 241 may also generate image data in which a sequential number is assigned to each text frame or the colors of the text frames are unified.
[0052] Furthermore, the image generating unit 241 may control the generation of image data so that text information corresponding to one utterance section is not divided into multiple blocks (text frames) in the middle of a phrase. Alternatively, when dividing text information corresponding to one utterance section into multiple blocks (text frames), the image generating unit 241 may control the generation of image data so that the text information is divided at a position where a period or a comma exists.
[0053] 2 is an example, and images displayed on the display unit 203 are not limited to this. For example, the display unit 203 may display image data that distinguishes each of the multiple groups (reference symbols g223A to g223D) by inserting line breaks, as in image g10a shown in Fig. 3. In this case, the image generation unit 241 generates image data that distinguishes each of the multiple groups (reference symbols g223A to g223D) by inserting line breaks.
[0054] Furthermore, the shape of the text frame is not limited to a rectangle, but may be, for example, a polygon, a circle, an ellipse, etc. Alternatively, a text frame does not have to be attached to each group.
[0055] <Display screen of handset 3> Next, an example of an image displayed on the display unit 303 of the slave device 3 will be described. FIG. 4 is a diagram showing an example of an image displayed on the display unit 303 of the slave unit 3 according to this embodiment. An image g30 is an image displayed on the display unit 303 of the child device 3.
[0056] Area g300 is an area where the minutes are displayed. Reference symbol g311 is text information input by the first participant by operating the operation unit (202 or 302). Reference symbol g321 is text information input by the second participant by operating the operation unit (202 or 302). Reference symbols g322 and g323 are text information obtained by voice recognition of what was spoken by the second participant. Reference symbol g331 is text information obtained by voice recognition of what was spoken by the third participant. Area g301 is an area for text input. Note that the operation unit 302 may be a software keyboard displayed on the display unit 303, and may be connected to the handset 3 by wire or wirelessly.
[0057] Here, the text information related to the symbol g322, "Today's agenda is about the upcoming exhibition," is included in one speech section of the user (Fujisawa), and the number of characters is less than a predetermined value, so it is displayed as a single block (text frame). On the other hand, the text information related to the symbol g323, "At the last meeting, please tell me...," is included in one speech section of the user (Fujisawa), but the number of characters is greater than or equal to the predetermined value, so it is displayed as multiple blocks (text frames) indicated by the symbols g323a to g323c. Furthermore, the number of characters included in each of the text frames indicated by the symbols g323a to g323c is less than the predetermined value.
[0058] 4, the display unit 303 may display image data in which consecutive numbers indicated by reference symbols g341 to g343 are assigned to the text frames indicated by reference symbols g323a to g323c. The display unit 303 may also display image data in which the colors of the text frames indicated by reference symbols g323a to g323c are the same.
[0059] In order for the display unit 303 to perform the above display, the image generation unit 241 generates image data in which a text frame is assigned to each group. The image generation unit 241 may also generate image data in which a sequential number is assigned to each text frame or the colors of the text frames are unified.
[0060] Furthermore, the image generating unit 241 may control the generation of image data so that text information corresponding to one utterance section is not divided into multiple blocks (text frames) in the middle of a phrase. Alternatively, when dividing text information corresponding to one utterance section into multiple blocks (text frames), the image generating unit 241 may control the generation of image data so that the text information is divided at a position where a period or a comma exists.
[0061] 2 and 4, the position at which text information corresponding to one utterance section is divided into multiple blocks (text frames) may be different between the image displayed on the display unit 203 of the parent unit 2 and the image displayed on the display unit 303 of the child unit 3. For example, the "predetermined value" (described above) that is the criterion for whether the image generating unit 241 divides the text information corresponding to one utterance section into multiple blocks may be different between the parent unit 2 and the child unit 3. Similarly, the position at which the text information is divided into multiple blocks and the "predetermined value" may be different for each of the child units 3a, 3b, etc.
[0062] 4 is an example, and the image displayed on the display unit 303 is not limited to this. For example, the display unit 303 may display image data that distinguishes each of the multiple groups by inserting a line break. In this case, the image generation unit 241 generates image data that distinguishes each of the multiple groups by inserting a line break.
[0063] Furthermore, the shape of the text frame is not limited to a rectangle, but may be, for example, a polygon, a circle, an ellipse, etc. Alternatively, a text frame does not have to be attached to each group.
[0064] <Processing performed by the speech recognition result display device (speech recognition result display system) 1> Next, an example of a processing procedure performed by the speech recognition result display device (speech recognition result display system) 1 will be described. Fig. 5 is a flowchart showing an example of a processing procedure performed by the speech recognition result display device (speech recognition result display system) 1 according to this embodiment. Note that the following describes a case where a user inputs text information using the sound collection unit (201 or 301).
[0065] (Step S1) The authentication unit 211 performs login processing based on the operation content of the operation unit (202 or 302). For example, when each user operates the operation unit (202 or 302) to input identification information (user ID) that identifies the user and a password, the authentication unit 211 performs login processing based on the input identification information and password.
[0066] (Step S2) The acquisition unit 222 acquires the user's utterance voice output by the sound collection unit 201 or the communication unit 204, and outputs the acquired utterance voice to the voice recognition unit 223.
[0067] (Step S3) The voice recognition unit 223 acquires the uttered voice output by the acquisition unit 222, and performs voice recognition processing on the acquired uttered voice.
[0068] (Step S4) The text conversion unit 224 performs text conversion processing on the speech recognition result.
[0069] (Step S5) The dependency analysis unit 225 performs dependency analysis and morphological analysis on the converted text information for each speaker.
[0070] (Step S6) The dependency analysis unit 225 outputs the text information that has been subjected to the dependency analysis and morphological analysis processes to the minutes creation unit 226.
[0071] (Step S7) The minutes creating unit 226 creates minutes based on the text information output by the dependency analyzing unit 225, and outputs them to the image generating unit 241.
[0072] (Step S8) The image generating unit 241 determines whether the number of characters in the text information included in one utterance section is equal to or greater than a predetermined value. If the image generating unit 241 determines that the number of characters in the text information included in one utterance section is equal to or greater than the predetermined value (Step S8; YES), the process of Step S9 is performed, and if the image generating unit 241 determines that the number of characters in the text information included in one utterance section is less than the predetermined value (Step S8; NO), the process of Step S10 is performed. Note that the above determination may be made by the minutes creating unit 226.
[0073] (Step S9) The image generation unit 241 generates image data for displaying text information included in one speech section divided into multiple groups, and outputs the image data to the display unit 203 or the communication unit 204. In step S9, the image generation unit 241 generates the image data so that the number of characters of the text information included in each of the multiple groups is less than a predetermined value.
[0074] (Step S10) The image generating unit 241 generates image data that displays text information included in one utterance section as a single unit, and outputs the image data to the display unit 203 or the communication unit 204.
[0075] (Step S11) The display unit (203 or 303) displays the image output by the image generation unit 241.
[0076] The speech recognition result display device (speech recognition result display system) 1 thereafter repeats the processes of steps S2 to S11. Note that the processes of steps S7 to S9 (image data generation processes) may be performed every time the dependency analysis unit 225 outputs text information. In other words, the processes of steps S2 to S11 may be repeated in real time during one speech segment. That is, the generation of images by the image generation unit 241 may be repeated in real time, and the images displayed by the display unit (203 or 303) may be continuously updated in real time. The process in FIG. 5 is an example, and the present invention is not limited to this.
[0077] As described above, in this embodiment, when one utterance section becomes long, the text information included in the utterance section is divided into a plurality of chunks and displayed. As a result, according to this embodiment, it is possible to prevent the text information from becoming difficult to read when one utterance section becomes long.
[0078] In the above example, the voice recognition result display device (voice recognition result display system) 1 allows both text input via the operation unit (202 or 302) and text input using voice recognition by the sound collection unit (201 or 301), but this is not limiting. For example, the voice recognition result display device (voice recognition result display system) 1 may only allow text input using voice recognition by the sound collection unit (201 or 301).
[0079] In the above example, the voice recognition result display device 1 includes the master unit 2 and multiple slave units 3, but this is not limiting. For example, the voice recognition result display device 1 may include only one slave unit 3, or the voice recognition result display device 1 may not include any slave units 3.
[0080] Furthermore, each of the authentication unit 211, the acoustic model and dictionary storage unit 221, the acquisition unit 222, the speech recognition unit 223, the text conversion unit 224, the dependency analysis unit 225, the minutes creation unit 226, the minutes storage unit 227, the text acquisition unit 231, and the image generation unit 241 may be included in the slave unit 3. Similarly, the processing unit 305 may be included in the master unit 2.
[0081] Furthermore, each functional unit of the speech recognition result display device 1 may be provided in a device other than the parent device 2 and the child device 3. Alternatively, each functional unit of the speech recognition result display system 1 does not have to be provided in the parent device 2, the child device 3, or other physical device, but may be provided on one or more servers or clouds. The functional units are the communication unit 204, the authentication unit 211, the acoustic model and dictionary storage unit 221, the acquisition unit 222, the speech recognition unit 223, the text conversion unit 224, the dependency analysis unit 225, the minutes creation unit 226, the minutes storage unit 227, the text acquisition unit 231, the image generation unit 241, the communication unit 304, and the processing unit 305.
[0082] Note that a program for implementing all or part of the functions of the speech recognition result display device (speech recognition result display system) 1 of the present invention may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be loaded into a computer system and executed to perform all or part of the processing performed by the speech recognition result display device (speech recognition result display system) 1. Note that the term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer system" also includes a WWW system equipped with a homepage providing environment (or display environment). The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. The term "computer-readable recording medium" also refers to devices that retain a program for a certain period of time, such as volatile memory (RAM) within a computer system that acts as a server or client when the program is transmitted via a network such as the Internet or a communication line such as a telephone line.
[0083] The program may also be transmitted from a computer system storing the program in a storage device or the like to another computer system via a transmission medium or by transmission waves in the transmission medium. Here, the "transmission medium" that transmits the program refers to a medium that has the function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line. The program may also be a program that realizes part of the above-mentioned functions. Furthermore, the program may be a so-called differential file (differential program) that can realize the above-mentioned functions in combination with a program already recorded in the computer system.
[0084] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]
[0085] 1... Speech recognition result display device (speech recognition result display system) 203... Display unit 222... Acquisition unit 223... Speech recognition unit 224... Text conversion unit 225... Dependency analysis unit 226... Minutes creation unit 241... Image generation unit 303... Display unit
Claims
1. an acquisition unit that acquires a user's speech; a speech recognition unit that performs speech recognition on the speech acquired by the acquisition unit and outputs text information; an image generation unit that generates image data based on the text information; a display unit that displays the image data, the image generation unit generates the image data to display the text information included in one utterance section as a single block when the number of characters of the text information included in one utterance section of the user is less than a predetermined value, and generates the image data to display the text information included in the one utterance section as a single block when the number of characters of the text information included in the one utterance section is equal to or greater than the predetermined value, the number of characters of the text information included in each of the plurality of chunks is less than the predetermined value; the image generation unit adds a text frame to each of the plurality of blocks when generating the image data; when the text information included in the single speech section spans a plurality of the text frames, the image generation unit assigns consecutive number displays to the plurality of text frames; the image generation unit does not assign the sequential number indication to the text frame corresponding to the single utterance section in which the number of characters of the text information is less than the predetermined value and which is displayed as the single unit. Speech recognition result display system.
2. the image generation unit controls the single utterance section so that it is not divided into the plurality of segments midway through a phrase.
2. The speech recognition result display system according to claim 1.
3. When the text information included in the single speech section spans a plurality of the text frames, the image generation unit unifies the colors of the plurality of text frames.
3. The speech recognition result display system according to claim 1.
4. the image generation unit generates the image data that distinguishes each of the plurality of groups by performing line breaks.
3. The speech recognition result display system according to claim 1.
5. the image generation unit performs the image data generation process every time the text information is output from the voice recognition unit. The speech recognition result display system according to any one of claims 1 to 4.
6. an acquisition unit that acquires a user's speech; a speech recognition unit that performs speech recognition on the speech acquired by the acquisition unit and outputs text information; an image generation unit that generates image data based on the text information; a display unit that displays the image data, the image generation unit generates the image data to display the text information included in one utterance section as a single block when the number of characters of the text information included in one utterance section of the user is less than a predetermined value, and generates the image data to display the text information included in the one utterance section as a single block when the number of characters of the text information included in the one utterance section is equal to or greater than the predetermined value, the number of characters of the text information included in each of the plurality of chunks is less than the predetermined value; the image generation unit adds a text frame to each of the plurality of blocks when generating the image data; when the text information included in the single speech section spans a plurality of the text frames, the image generation unit assigns consecutive number displays to the plurality of text frames; the image generation unit does not assign the sequential number indication to the text frame corresponding to the single utterance section in which the number of characters of the text information is less than the predetermined value and which is displayed as the single unit. Voice recognition result display device.
7. A speech recognition result display method in a speech recognition result display system, comprising: an acquisition step in which an acquisition unit acquires a speech voice of a user; a speech recognition step in which a speech recognition unit performs speech recognition on the uttered speech acquired by the acquisition unit and outputs text information; an image generating step in which an image generating unit generates image data based on the text information; a display step in which a display unit displays the image data; In the image generating step, when the number of characters of the text information included in one utterance section of the user is less than a predetermined value, the image generating unit generates the image data that displays the text information included in the one utterance section as a single block, and when the number of characters of the text information included in the one utterance section is equal to or greater than the predetermined value, the image generating unit performs an image data generating process that generates the image data that displays the text information included in the one utterance section divided into a plurality of blocks; the number of characters of the text information included in each of the plurality of chunks is less than the predetermined value; In the image generating step, the image generating unit adds a text frame to each of the plurality of blocks when generating the image data, In the image generating step, when the text information included in the single speech section spans a plurality of the text frames, the image generating unit assigns consecutive number displays to the plurality of text frames; In the image generating step, the image generating unit does not assign the sequential number indication to the text frame corresponding to the single utterance section in which the number of characters of the text information is less than the predetermined value and which is displayed as the single unit. How to display voice recognition results.
8. For the voice recognition result display system, an acquisition step of acquiring a user's speech; a speech recognition step of performing speech recognition on the uttered speech and outputting text information; an image generation step of generating image data based on the text information; a display step of displaying the image data, In the image generating step, when the number of characters of the text information included in one utterance section of the user is less than a predetermined value, the image data is generated to display the text information included in the one utterance section as a single block, and when the number of characters of the text information included in the one utterance section is equal to or greater than the predetermined value, the image data is generated to display the text information included in the one utterance section divided into a plurality of blocks, the number of characters of the text information included in each of the plurality of chunks is less than the predetermined value; In the image generating step, when generating the image data, a text frame is added to each of the plurality of blocks; In the image generating step, when the text information included in the one speech section spans a plurality of the text frames, consecutive numbers are assigned to the plurality of text frames; In the image generating step, the consecutive number indication is not given to the text frame corresponding to the single speech section in which the number of characters of the text information is less than the predetermined value and which is displayed as one unit. program.
Citation Information
Patent Citations
Setting device, program, recording medium, and setting method
JP2005258198A
Image and sound processing apparatus
JP2010154259A
Conference system, control method therefor, and program
JP2019179480A
Display support device, method and program
JP2020057401A
Sound recognition subtitle production system and control method thereof
KR102185387B1