Information processing system, information processing method and information processing program

The system enhances speech recognition accuracy by correcting misrecognized characters in short sentences or those with high misrecognition rates by comparing them with adjacent text sentences, addressing conventional accuracy issues.

JP2025113610APending Publication Date: 2025-08-04SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024007856
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-08-04

AI Technical Summary

Technical Problem

Conventional speech-to-text systems face accuracy issues when input text sentences are short or contain many masked portions, leading to decreased prediction accuracy for masked words.

Method used

An information processing system that includes a conversion processing unit, first and second extraction processing units, and a correction processing unit to extract and correct misrecognized characters by comparing a target text sentence with a correction text sentence, ensuring the total number of characters or misrecognition ratio meets specific conditions.

Benefits of technology

Improves speech recognition accuracy by correcting misrecognized characters using additional text context, especially when input sentences are short or have high misrecognition rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025113610000001_ABST
    Figure 2025113610000001_ABST
Patent Text Reader

Abstract

To provide an information processing system, an information processing method and an information processing program, which can improve voice recognition accuracy to convert voice into a text.SOLUTION: An information processor 1 includes: a voice recognition processing unit 113 for converting uttered voice of a user into text information; a target text extraction processing unit for extracting a target text sentence in prescribed length from text information; a recognition determination processing unit 115 for determining whether the total number of characters in the target text sentence is less than a prescribed number of characters or not when characters of an erroneous recognition candidate are included in the target text sentence or whether a ratio of the number of characters of the erroneous recognition candidate with respect to the total number of characters of the target text sentence is more than a prescribed ratio or not; a corrected text extraction processing unit for extracting a corrected text sentence when the total number of characters is determined to be less than the prescribed number of characters or when the ratio is determined to be equal to or grater than the prescribed ratio; and a correction processing unit 116 for correcting characters of erroneous recognition based on the target text sentence or the corrected text sentence.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technology for converting and displaying a user's spoken voice into text.

Background Art

[0002] Conventionally, a technology for converting a user's spoken voice into text information and displaying it is known. For example, a system is known that replaces a word of a proper name included in a sentence (text sentence) obtained by converting a user's spoken voice into text with a mask symbol, predicts a word corresponding to the mask symbol based on a part excluding the mask symbol in the replaced sentence with the mask symbol, and corrects the mask symbol to the word and outputs it (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the conventional technology, for example, when the input text sentence is short or when there are many masked portions in the text sentence, there arises a problem that the accuracy of predicting an appropriate word for the masked portion decreases.

[0005] An object of the present disclosure is to provide an information processing system, an information processing method, and an information processing program capable of improving the speech recognition accuracy for converting speech into text.

Means for Solving the Problems

[0006] An information processing system according to one aspect of the present disclosure includes a conversion processing unit, a first extraction processing unit, a determination processing unit, a second extraction processing unit, and a correction processing unit. The conversion processing unit recognizes speech uttered by a user and converts it into text information. The first extraction processing unit extracts a first text sentence of a predetermined length from the text information converted by the conversion processing unit. The determination processing unit determines whether the total number of characters in the first text sentence is less than a predetermined number of characters, or whether the ratio of the number of characters of the misrecognition candidate to the total number of characters in the first text sentence is equal to or greater than a predetermined ratio, when the first text sentence extracted by the first extraction processing unit includes characters of a misrecognition candidate. The second extraction processing unit extracts a second text sentence different from the first text sentence from the text information when it is determined by the determination processing unit that the total number of characters is less than the predetermined number of characters, or when it is determined that the ratio is equal to or greater than the predetermined ratio. The correction processing unit corrects misrecognized characters among the misrecognition candidate characters based on the first text sentence and the second text sentence.

[0007] An information processing method according to another aspect of the present disclosure includes recognizing speech uttered by a user and converting it into text information, extracting a first text sentence of a predetermined length from the text information, determining whether the total number of characters in the first text sentence is less than a predetermined number of characters, or whether the ratio of the number of characters of the misrecognition candidate to the total number of characters in the first text sentence is equal to or greater than a predetermined ratio, when the first text sentence includes characters of a misrecognition candidate, extracting a second text sentence different from the first text sentence from the text information when it is determined that the total number of characters is less than the predetermined number of characters, or when it is determined that the ratio is equal to or greater than the predetermined ratio, and correcting misrecognized characters among the misrecognition candidate characters based on the first text sentence and the second text sentence, which is an information processing method executed by one or more processors.

[0008] An information processing program according to another aspect of the present disclosure includes: recognizing voice spoken by a user and converting it into text information; extracting a first text sentence of a predetermined length from the text information; when the first text sentence includes characters of misrecognition candidates, determining whether the total number of characters in the first text sentence is less than a predetermined number of characters, or whether the ratio of the number of characters of the misrecognition candidates to the total number of characters in the first text sentence is equal to or greater than a predetermined ratio; when it is determined that the total number of characters is less than the predetermined number of characters, or when it is determined that the ratio is equal to or greater than the predetermined ratio, extracting a second text sentence different from the first text sentence from the text information; and correcting misrecognized characters among the characters of the misrecognition candidates based on the first text sentence and the second text sentence, and causing one or more processors to execute the above.

Advantages of the Invention

[0009] According to the present disclosure, it is possible to provide an information processing system, an information processing method, and an information processing program capable of improving the voice recognition accuracy for converting voice into text.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 6A

Figure 6B

Figure 6C

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

[0011] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that the following embodiments are an example of embodying the present disclosure and do not have the character of limiting the technical scope of the present disclosure.

[0012] The information processing system according to the present disclosure can be applied, for example, to a case where a plurality of users each use a voice device equipped with a microphone and a speaker in one space (for example, a meeting room) to have a conversation. Note that the information processing system can also be applied to a case where a plurality of users each use a voice device in one space to have a conversation (meeting) with users in another space. Furthermore, the information processing system can also be applied to a case where one user uses a voice device in one space to have a conversation with users in another space.

[0013] FIG. 1 shows an application example of the meeting support system 100 according to the present embodiment. As shown in FIG. 1, users A to D participate in a meeting in the meeting room R1. Users A to D each use a neckband-type voice device 2A to 2D that can be worn around the neck to have a conversation. Note that FIG. 1 shows an example in which each of users A to D uses the voice device 2, but the present invention is not limited thereto, and only some users may use the voice device 2. The meeting support system 100 is an example of the information processing system of the present disclosure.

[0014] Each voice device 2 in the meeting room R1 is wirelessly connected (connected by Bluetooth (registered trademark)) to the information processing device 1, and the voice input to the microphone of each voice device 2 is input to the information processing device 1. The information processing device 1 executes a voice recognition process for converting the voice input from each voice device 2 into text information, and outputs the voice recognition result to the user terminal 3, the display device 4, and the like. For example, the information processing device 1 performs real-time speech-to-text conversion of the speech content while the user is speaking and displays it on the user terminal 3 and the display device 4. In addition, the information processing device 1 can output (playback) the voice input to the microphone of each voice device 2 from the speaker of each voice device 2. In addition, the information processing device 1 can display meeting information such as meeting materials on the user terminal 3, the display device 4, and the like. Each user can view text information corresponding to the conversation content, meeting information, and the like on the user terminal 3 and the display device 4. In addition, each user can have a conversation with other users via the voice device 2.

[0015] As shown in FIG. 1, the conference support system 100 includes an information processing device 1, an audio device 2, a user terminal 3, and a display device 4. The audio device 2 is a wireless connection type audio device equipped with a microphone and a speaker. Note that the audio device 2 may have functions such as an AI speaker or a smart speaker. The conference support system 100 includes a plurality of audio devices 2 and is a system capable of transmitting and receiving audio data of the user's speech between the plurality of audio devices 2. Each audio device 2 may be of the same type of audio device or different types of audio devices.

[0016] The information processing device 1 controls the audio (input audio, output audio, etc.) of the audio device 2 and, for example, when a conference starts in a conference room, executes a process of transmitting and receiving audio between the plurality of audio devices 2. For example, the information processing device 1 controls a plurality of audio devices 2 arranged in the same space. In addition, the information processing device 1 accumulates the audio obtained from the audio device 2 as recording audio or executes a process (speech recognition process) of converting the obtained audio into text. Note that the information processing device 1 alone may constitute the information processing system of the present disclosure.

[0017] In addition, the information processing system of the present disclosure may have a function of providing various services such as a conference service, a subtitle (captioning) service by speech recognition, a translation service, and a minutes service. For example, a general-purpose conference application is installed on each user terminal 3. By starting and logging in to the user terminal 3, the user can hold a conference using the conference application.

[0018] [Information Processing Device 1] As shown in FIG. 2, the information processing device 1 is a device including a control unit 11, a storage unit 12, a communication unit 13, etc. For example, the information processing device 1 is connected to a plurality of audio devices 2 and has functions such as mixing or splitting the audio input from the plurality of audio devices 2 and converting the input audio into text information.

[0019] The communication unit 13 is a communication unit for connecting the information processing device 1 to a communication network by wire or wirelessly and performing data communication according to a predetermined communication protocol with external devices such as the voice device 2, the user terminal 3, and the display device 4 via the communication network. For example, the communication unit 13 executes pairing processing by the Bluetooth method to wirelessly connect to each voice device 2.

[0020] The storage unit 12 is a non-volatile storage unit such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), or a flash memory that stores various types of information. Data such as device information D1 regarding the voice device 2 and speech information D2 regarding the speech content is stored in the storage unit 12.

[0021] FIG. 3 shows an example of the device information D1. Information such as a connection ID, a device name, and voice data is registered in the device information D1. The connection ID is identification information (device information) used when connecting the voice device 2 and is, for example, a Bluetooth address. The device name is the device name of the voice device 2. The voice data is data of voice (spoken voice) acquired from the voice device 2. In this way, voice data is stored for each device in the device information D1. Each voice data is stored with the identification information of the device (voice device 2) attached.

[0022] FIG. 4 shows an example of the speech information D2. In the speech information D2, information such as a conversation ID, a speech time, a speaker, and a speech content is registered. The conversation ID is identification information for the speech content of the user (speaker). The speech time is the time when the user spoke. The speaker is the name or identification information (user ID) of the user. The speech content is character information obtained by converting the speech of the user into text information. One speech content is a text sentence obtained by dividing the text information of the speech voice by a predetermined length (one sentence). For example, when a silent state continues for several seconds in the speech voice, it is extracted as one text sentence. Also, when the text information includes a period “.”, it is extracted as one text sentence. The control unit 11 registers in the speech information D2 every time it extracts a text sentence of a predetermined length (one sentence). The text sentence of the predetermined length is an example of the first text sentence of the present disclosure.

[0023] Further, in the storage unit 12, control programs such as a speech recognition program (an example of the information processing program of the present disclosure) for causing the control unit 11 to execute the speech recognition process (see FIG. 10) described later are stored. For example, the speech recognition program may be non-temporarily recorded on a computer-readable recording medium such as a CD or a DVD, and read by a reading device (not shown) such as a CD drive or a DVD drive provided in the information processing apparatus 1 and stored in the storage unit 12.

[0024] The control unit 11 includes control devices such as a CPU, a ROM, and a RAM. The CPU is a processor that executes various arithmetic processes. The ROM is a non-volatile storage unit in which control programs such as a BIOS and an OS for causing the CPU to execute various arithmetic processes are stored in advance. The RAM is a volatile or non-volatile storage unit that stores various information, and is used as a temporary storage memory (working area) for various processes executed by the CPU. Then, the control unit 11 controls the information processing apparatus 1 by causing the CPU to execute various control programs stored in advance in the ROM or the storage unit 12.

[0025] Specifically, as shown in FIG. 2, the control unit 11 includes various processing units such as a setting processing unit 111, an acquisition processing unit 112, a voice recognition processing unit 113, a target text extraction processing unit 114, a recognition determination processing unit 115, a correction processing unit 116, a character determination processing unit 117, a corrected text extraction processing unit 118, and a display processing unit 119. The control unit 11 functions as the various processing units by executing various processes according to the control program using the CPU. Also, some or all of the processing units may be configured by electronic circuits. The control program may be a program for causing a plurality of processors to function as the processing units.

[0026] The setting processing unit 111 sets the topic of the conversation (meeting). Specifically, the setting processing unit 111 accepts operations such as registering a new topic, selecting a topic, and changing the topic currently being discussed from the user (such as the meeting organizer). For example, user A, who is the meeting organizer, starts the meeting application on the user terminal 3A (meeting terminal 3) to display the setting screen P1 (see FIG. 5A). User A can select the "Topic" column on the setting screen P1 to register one or more topics. Also, user A can select one topic from the multiple registered topics on the setting screen P1. Further, user A can start the meeting by pressing the start button K1 on the setting screen P1. FIG. 5B shows an example of the setting screen P11 after the meeting has started (during the meeting). FIG. 5B displays identification information ("During conversation") indicating the topic currently being discussed (such as "Next-generation office"). When user A wants to change the topic, user A presses the end button K2 on the setting screen P11 and selects another topic on the setting screen P1 (see FIG. 5A). Also, user A can select the "Settings" column on the setting screen P1 to select the users participating in the meeting or select the voice device 2 used in the meeting.

[0027] The acquisition processing unit 112 acquires the voice spoken by the user. Specifically, when the meeting starts and the user speaks, the acquisition processing unit 112 acquires the spoken voice input to the microphone of the voice device 2 of the user. Further, the acquisition processing unit 112 acquires time information corresponding to the speaking time of the user's spoken voice. For example, the acquisition processing unit 112 acquires the time when the user's spoken voice is input to the microphone of the voice device 2, or the time when the spoken voice is acquired.

[0028] When the acquisition processing unit 112 acquires voice from the voice device 2, it stores the voice data in the device information D1 (see FIG. 3) in association with the identification information (connection ID, device name, etc.) of the voice device 2. In the example shown in FIG. 1, the acquisition processing unit 112 acquires the respective spoken voices of users A to D in the conference room R1 from the voice device 2, and stores each voice data in the device information D1 in association with the connection ID.

[0029] The voice recognition processing unit 113 executes a process of converting voice into text (voice recognition process) based on the voice acquired by the acquisition processing unit 112. The voice recognition processing unit 113 uses a predetermined voice recognition engine (trained model) to convert the voice into text information. Note that the voice recognition engine is generated, for example, by learning various voice data (teacher data). The information processing apparatus 1 is equipped with the voice recognition engine. Note that the voice recognition processing unit 113 may perform voice processing such as echo cancellation, noise cancellation, and gain adjustment on the voice acquired by the acquisition processing unit 112.

[0030] Further, the voice recognition processing unit 113 registers the text information as the speech content in the speech information D2 (see FIG. 4). The voice recognition processing unit 113 also registers the identification information (conversation ID) of the speech content, the time information (speaking time), and the identification information (name) of the speaker in association with the speech content. The voice recognition processing unit 113 is an example of the conversion processing unit of the present disclosure.

[0031] The target text extraction processing unit 114 extracts a target text sentence (first text sentence) of a predetermined length from the text information converted by the speech recognition processing unit 113. For example, the target text extraction processing unit 114 extracts one sentence delimited by a silent time or a period "。" as the target text sentence.

[0032] The recognition determination processing unit 115 determines whether a character (character or word of misrecognition candidate) that may be misrecognized is included in the target text sentence of a predetermined length extracted by the target text extraction processing unit 114. Specifically, the recognition determination processing unit 115 refers to a pre-registered word list and determines for each of a plurality of characters included in the target text sentence whether it is appropriately converted, whether there is a possibility of error, etc. For example, the recognition determination processing unit 115 refers to the word list and determines whether the converted character may be misrecognized from the words, context, etc. included in the target text sentence. The word list may have a correct word and an incorrect word pre-registered.

[0033] FIG. 6A shows the content (correct answer) spoken by the user, and FIG. 6B shows the speech recognition result (target text sentence) for the spoken voice. The recognition determination processing unit 115 refers to the word list for the recognition result shown in FIG. 6B and determines the four words "current price", "reproduction", "sales depreciation rate", and "current utility cost" as characters of misrecognition candidates. Note that at this stage, the correct character may be included in the characters of the misrecognition candidates.

[0034] When the recognition determination processing unit 115 determines that a character of misrecognition candidate is included in the target text sentence, it executes a process of masking the character. For example, as shown in FIG. 6C, the recognition determination processing unit 115 replaces each of "current price", "reproduction", "sales depreciation rate", and "current utility cost", which are characters of misrecognition candidates, with a mask symbol ([MASK]). Note that the recognition determination processing unit 115 may replace each character constituting the word with a mask symbol.

[0035] When the target text sentence contains misrecognized characters, the correction processing unit 116 executes correction processing to correct the characters to correct characters. Specifically, the correction processing unit 116 can execute the correction processing of correcting the misrecognized characters among the misrecognition candidate characters included in the target text sentence by using a well-known language prediction model (MLM: Masked Language Model). When the language prediction model determines that the misrecognition candidate character is a correct character (not a misrecognition), the correction processing unit 116 displays the recognized character before masking without performing the correction processing.

[0036] Here, when the input target text sentence is short, or when there are many misrecognition candidate characters (masked parts) included in the target text sentence, the number of correctly recognized characters (correct characters) decreases, resulting in a problem that the accuracy of predicting the correct characters for the masked parts decreases. In contrast, the control unit 11 has the following configuration that can prevent the above problem.

[0037] Specifically, when the target text sentence extracted by the target text extraction processing unit 114 contains misrecognition candidate characters, the character determination processing unit 117 determines whether the total number of characters in the target text sentence is less than a predetermined number of characters (first condition), or whether the ratio of the number of misrecognition candidate characters to the total number of characters in the target text sentence is equal to or greater than a predetermined ratio (second condition).

[0038] The predetermined number of characters is set so that the time difference between the speech timing and the display timing is small, for example, in a usage scenario where the user's speech content is displayed in characters in real time. In this case, the predetermined number of characters is set to 30 characters, for example. The character determination processing unit 117 determines whether the total number of characters in the target text sentence is less than 30 characters (first condition).

[0039] Further, the predetermined ratio is set so that the prediction accuracy of the language prediction model can maintain a predetermined accuracy. For example, when the language prediction model is a model that requires that the ratio of the number of characters of misrecognition candidates to the total number of characters of the target text sentence is less than 15%, the predetermined ratio is set to 15%. The character determination processing unit 117 determines whether the ratio of the number of characters of misrecognition candidates to the total number of characters of the target text sentence is 15% or more (second condition).

[0040] When the correction text extraction processing unit 118 satisfies at least one of the first condition and the second condition, the correction text extraction processing unit 118 extracts a correction text sentence for correction that is different from the target text sentence from the text information converted by the speech recognition processing unit 113. Specifically, when the total number of characters of the target text sentence is less than a predetermined number of characters, or when the ratio of the number of characters of misrecognition candidates to the total number of characters of the target text sentence is equal to or more than a predetermined ratio, the correction text extraction processing unit 118 extracts the correction text sentence. The correction text sentence is an example of the second text sentence of the present disclosure.

[0041] For example, the correction text extraction processing unit 118 extracts the text sentence immediately before the target text sentence (the previous sentence) in time series from the text information as the correction text sentence. FIG. 7 shows the speech recognition results arranged in time series. It shows that the conversation was carried out in the order of text sentences T1, T2, and T3. In this case, the correction text extraction processing unit 118 extracts the text sentence T2 immediately before the target text sentence T3 as the correction text sentence. Each text sentence before the target text sentence (text sentences T1 and T2 in FIG. 7) is a corrected sentence in which misrecognized characters have been corrected to correct characters by correction processing described later.

[0042] When the correction text extraction processing unit 118 extracts the correction text sentence, the character determination processing unit 117 executes the determination processing corresponding to the first condition and the second condition again. Specifically, the character determination processing unit 117 determines whether the total number of characters in the target text sentence T3 and the correction text sentence T2 is less than a predetermined number of characters (the first condition), or whether the ratio of the number of characters of the misrecognition candidates to the total number of characters in the target text sentence T3 and the correction text sentence T2 is equal to or greater than a predetermined ratio (the second condition). When the total number of characters in the target text sentence T3 and the correction text sentence T2 is less than the predetermined number of characters, or when the ratio of the number of characters of the misrecognition candidates to the total number of characters in the target text sentence T3 and the correction text sentence T2 is equal to or greater than the predetermined ratio, the correction text extraction processing unit 118 further extracts the correction text sentence. For example, the correction text extraction processing unit 118 extracts the text sentence T1 immediately before the correction text sentence T2 as the correction text sentence.

[0043] In this way, the correction text extraction processing unit 118 adds the correction text sentences in chronological order until the total number of characters becomes equal to or greater than the predetermined number of characters and the ratio of the number of characters of the misrecognition candidates becomes less than the predetermined ratio.

[0044] The correction processing unit 116 executes a correction process for correcting the misrecognized characters of the target text sentence. Specifically, when the total number of characters in the target text sentence is equal to or greater than a predetermined number of characters and the ratio of the number of characters of the misrecognition candidates to the total number of characters in the target text sentence is less than a predetermined ratio, the correction processing unit 116 predicts the correct characters corresponding to the misrecognition candidate characters for the target text sentence by using a well-known language prediction model, and replaces the misrecognized characters among the misrecognition candidate characters with the correct characters. For example, the correction processing unit 116 lists up the words of the correction candidates for the misrecognized characters from the word list based on the context of the target text sentence, extracts the word with the highest prediction rate (correct rate) from the listed correction candidates, and replaces it with the mask symbol. Note that the correction processing unit 116 omits the replacement process when the masked misrecognition candidate character (the recognized character before masking) matches the correction candidate.

[0045] On the other hand, when the total number of characters in the target text sentence is less than a predetermined number of characters, or when the ratio of the number of characters of the misrecognition candidates to the total number of characters in the target text sentence is equal to or greater than a predetermined ratio, the correction processing unit 116 executes a correction process for correcting the misrecognized characters based on the target text sentence and the correction text sentence. Specifically, the correction processing unit 116 predicts the correct characters corresponding to the misrecognition candidate characters based on the target text sentence with the misrecognized characters masked and the correction text sentence extracted by the correction text extraction processing unit 118, and replaces the misrecognized characters among the misrecognition candidate characters with the correct characters.

[0046] For example, when the total number of characters in the target text sentence T3 is less than a predetermined number of characters, or when the ratio of the number of characters of the misrecognition candidates to the total number of characters in the target text sentence T3 is equal to or greater than a predetermined ratio, the correction processing unit 116 predicts the correct characters based on the target text sentence T3 with the misrecognition candidate characters masked and the correction text sentence before the target text sentence T3, and replaces the masked misrecognized characters with the correct characters. For example, the correction processing unit 116 lists up the words of the correction candidates for the misrecognized characters from the word list based on the context between the target text sentence and one or more correction text sentences, extracts the word with the highest prediction rate (correct rate) from the listed correction candidates, and replaces it with the mask symbol. FIG. 8 shows the target text sentence T3 in which the misrecognized characters have been corrected.

[0047] Thus, when the total number of characters in the target text sentence T3 is less than a predetermined number of characters, or when the ratio of the number of characters of the misrecognition candidates to the total number of characters in the target text sentence T3 is equal to or greater than a predetermined ratio, by increasing the number of characters (text sentence) input to the language prediction model, the context can be grasped more accurately, so that the prediction accuracy of the words of the correction candidates is improved. Therefore, the speech recognition accuracy of the target text sentence can be improved.

[0048] The display processing unit 119 causes various types of information corresponding to the conference application to be displayed on the display screen. Specifically, the display processing unit 119 causes the voice recognition result to be displayed on the conference screen P2. For example, as shown in FIG. 9, the display processing unit 119 causes the text sentence (corrected text sentence) during the conversation to be displayed in real time in the conversation display area P22 of the conference screen P2 on the user terminal 3, and causes the summary sentence generated for each predetermined section to be displayed in the summary display area P21 of the conference screen P2. Note that the display processing unit 119 sequentially adds and displays text sentences in real time while scrolling the screen in the conversation display area P22.

[0049] In addition, the display processing unit 119 may display the identification information (such as the user name and user icon) of the user of the uttered voice corresponding to the text sentence and the time information (utterance time) in association with the text sentence (utterance content) in the conversation display area P22.

[0050] In addition, the display processing unit 119 may display, in the summary display area P21, identification information (for example, summary creation time) that can identify the section in association with the summary sentence. When the display processing unit 119 generates a summary sentence every five minutes, the display processing unit 119 displays the times at five-minute intervals in the summary display area P21.

[0051] In addition, the display processing unit 119 may display a plurality of topics in a selectable manner in the summary display area P21 and display the topics selected by the user in an identifiable manner. When no topic is selected in the summary display area P21, the display processing unit 119 displays the summary sentence corresponding to the topic currently being discussed. As a result, on the conference screen P2, the text sentences of the conversation regarding the topic currently being discussed and the summary sentence regarding the topic are displayed side by side. The user can collectively check the content (text sentences) and summary of the conversation currently in progress on one conference screen P2.

[0052] As described above, the control unit 11 converts the voice spoken by the user into a text sentence, and when the text sentence includes characters of misrecognition candidates, corrects and outputs the misrecognized characters among the characters of the misrecognition candidates. Further, when the text sentence including the characters of the misrecognition candidates is short, or when the ratio of the characters of the misrecognition candidates is high, another text sentence is added to the text sentence and the correction process is executed.

[0053] [Speech recognition processing] FIG. 10 shows an example of the procedure of the speech recognition processing executed by the control unit 11 of the information processing apparatus 1.

[0054] Note that the present disclosure can be regarded as a speech recognition method (the information processing method of the present disclosure) for executing one or more steps included in the speech recognition processing. Further, one or more steps included in the speech recognition processing described herein may be appropriately omitted. Further, the execution order of each step in the speech recognition processing may be different as long as the same operational effects are produced. Furthermore, although the case where the control unit 11 executes each step in the speech recognition processing is described herein as an example, in other embodiments, one or a plurality of processors may execute each step in the speech recognition processing in a distributed manner.

[0055] Here, as shown in FIG. 1, the case where a plurality of users hold a meeting using the voice apparatuses 2 in the conference room R1 will be described as an example. When the meeting is started, the control unit 11 executes the following processing.

[0056] In step S1, the control unit 11 acquires the voice spoken by the user from the voice apparatus 2. For example, when the voice spoken by the user after the start of the meeting is input to the microphone of the voice apparatus 2, the control unit 11 acquires the voice from the voice apparatus 2.

[0057] Next, in step S2, the control unit 11 executes speech recognition processing. Specifically, the control unit 11 uses a predetermined speech recognition engine (trained model) to convert the speech into text information. The control unit 11 associates the text information (utterance content) with time information (utterance time) and the speaker, and registers it in the utterance information D2 (see FIG. 4).

[0058] Next, in step S3, the control unit 11 extracts a target text sentence of a predetermined length from the converted text information. For example, the control unit 11 extracts one sentence delimited by a silent time or a period “.” as the target text sentence.

[0059] Next, in step S4, the control unit 11 determines whether the extracted target text sentence contains misrecognition candidate characters (words). For example, the control unit 11 refers to a pre-registered word list and determines whether the converted characters are correct based on the words and context included in the target text sentence. For example, in the example shown in FIG. 6B, the control unit 11 determines “current price”, “reproduction”, “sales depreciation rate”, and “current utility cost” in the target text sentence as misrecognition candidate characters. When the control unit 11 determines that the target text sentence contains the misrecognition candidate characters (S4: Yes), the process proceeds to step S5. On the other hand, when the control unit 11 determines that the target text sentence does not contain the misrecognition candidate characters (S4: No), the process proceeds to step S10.

[0060] In step S5, the control unit 11 executes a process of masking the misrecognition candidate characters. For example, as shown in FIG. 6C, the control unit 11 replaces each of “current price”, “reproduction”, “sales depreciation rate”, and “current utility cost”, which are misrecognition candidate characters, with a mask symbol ([MASK]).

[0061] Next, in step S6, the control unit 11 determines whether the total number of characters in the target text sentence is less than a predetermined number of characters (first condition). For example, the control unit 11 determines whether the total number of characters in the target text sentence is less than 30 characters. If the control unit 11 determines that the total number of characters in the target text sentence is less than the predetermined number of characters (S6: Yes), the process proceeds to step S8. On the other hand, if the control unit 11 determines that the total number of characters in the target text sentence is equal to or more than the predetermined number of characters (S6: No), the process proceeds to step S7.

[0062] In step S7, the control unit 11 determines whether the ratio of the number of characters of the misrecognition candidates to the total number of characters in the target text sentence is equal to or more than a predetermined ratio (second condition). For example, the control unit 11 determines whether the ratio of the number of characters of the misrecognition candidates to the total number of characters in the target text sentence is 15% or more. If the control unit 11 determines that the ratio of the number of characters of the misrecognition candidates to the total number of characters in the target text sentence is equal to or more than the predetermined ratio (S7: Yes), the process proceeds to step S8. On the other hand, if the control unit 11 determines that the ratio of the number of characters of the misrecognition candidates to the total number of characters in the target text sentence is less than the predetermined ratio (S7: No), the process proceeds to step S9.

[0063] In step S8, the control unit 11 extracts a correction text sentence to be used for correcting the misrecognized characters included in the target text sentence. Specifically, the control unit 11 extracts, from the text information converted in step S2, the text sentence immediately preceding the target text sentence in time series as the correction text sentence. In the example shown in FIG. 7, the control unit 11 extracts the text sentence T2 immediately preceding the target text sentence T3 as the correction text sentence. Note that the control unit 11 adds the correction text sentence in time series order (in the order closer to the current time) until the total number of characters becomes equal to or more than the predetermined number of characters and the ratio of the number of characters of the misrecognition candidates becomes less than the predetermined ratio. After the process of step S8, the control unit 11 proceeds to step S9.

[0064] Next, in step S9, the control unit 11 executes a correction process for correcting the misrecognized characters in the target text sentence. Specifically, when the control unit 11 extracts one or more of the correction text sentences (step S8), based on the target text sentence and the correction text sentence, it predicts the correct character corresponding to the misrecognition candidate character, and replaces the misrecognized character among the masked misrecognition candidate characters with the correct character. For example, the control unit 11 uses a well-known language prediction model (MLM) to list up the correction candidate words for the misrecognized character from a pre-registered word list based on the context of the target text sentence and the correction text sentence, and extracts the word with the highest prediction rate from the listed correction candidates and replaces it with the mask symbol.

[0065] On the other hand, when the total number of characters in the target text sentence is equal to or more than a predetermined number of characters (S6: No), and the ratio of the number of misrecognition candidate characters to the total number of characters in the target text sentence is less than a predetermined ratio (S7: No), the control unit 11 executes the correction process based on the context of only the target text sentence. After the process of step S9, the control unit 11 shifts the process to step S10.

[0066] In step S10, the control unit 11 causes the recognition result for the target text sentence to be displayed. Specifically, when the target text sentence does not include misrecognition candidate characters (S4: No), the control unit 11 causes the target text sentence to be displayed on the conference screen P2 (conversation display area P22) of the user terminal 3 without executing the correction process (see FIG. 9). On the other hand, when the target text sentence includes misrecognition candidate characters (S4: Yes), the control unit 11 causes the corrected target text sentence obtained by performing the correction process on the target text sentence to be displayed on the conference screen P2 (conversation display area P22) of the user terminal 3 (see FIG. 9).

[0067] Next, in step S11, the control unit 11 determines whether an end operation of the meeting has been received. For example, when the user ends the meeting, the user ends the meeting application. When the control unit 11 receives an end operation of the meeting application, it determines that an end operation of the meeting has been received. When the control unit 11 determines that an end operation of the meeting has been received (S11: Yes), it ends the voice recognition process. On the other hand, when the control unit 11 does not receive an end operation of the meeting (S11: No), it returns the process to step S1.

[0068] When returning to step S1, the control unit 11 acquires the user's uttered voice and executes each of the above processes on the voice. The control unit 11 repeatedly executes the above-described process until the meeting ends.

[0069] As described above, the conference support system 100 according to the present disclosure recognizes the voice uttered by the user, converts it into text information, extracts a target text sentence (first text sentence) of a predetermined length from the converted text information, and when the target text sentence includes a misrecognition candidate character, determines whether the total number of characters in the target text sentence is less than a predetermined number of characters, or whether the ratio of the number of misrecognition candidate characters to the total number of characters in the target text sentence is equal to or greater than a predetermined ratio. Further, when the total number of characters is less than the predetermined number of characters, or when the ratio is equal to or greater than the predetermined ratio, the conference support system 100 extracts a correction text sentence different from the target text sentence from the text information, and corrects the misrecognized characters among the misrecognition candidate characters based on the target text sentence and the correction text sentence. Further, the conference support system 100 extracts, as the correction text sentence, the text sentence immediately before the target text sentence in time series from the text information.

[0070] According to the above configuration, when the input target text sentence is short, or when there are many misrecognition candidate characters included in the target text sentence, another text sentence (for example, the previous sentence) can be added to the target text sentence to predict the correct characters for the misrecognized characters, so the prediction accuracy can be improved. Therefore, it becomes possible to improve the speech recognition accuracy for converting speech into text.

[0071] [Other Embodiments] As another embodiment of the present disclosure, when the total number of characters in the target text sentence is less than a predetermined number of characters, or when the ratio of the number of misrecognition candidate characters to the total number of characters in the target text sentence is equal to or greater than a predetermined ratio, the control unit 11 may extract, from the text information, a text sentence after the target text sentence in time series as the correction text sentence. For example, in the example shown in FIG. 11, for the target text sentence T3, when the total number of characters is less than a predetermined number of characters, or when the ratio of the number of misrecognition candidate characters is equal to or greater than a predetermined ratio, the control unit 11 extracts the text sentence T4 immediately after the target text sentence T3 as the correction text sentence.

[0072] In this case, the control unit 11 predicts the correct characters based on the target text sentence T3 with the misrecognition candidate characters masked and the correction text sentence T4 immediately after the target text sentence T3, and replaces the masked misrecognized characters with the correct characters.

[0073] Note that when using the text sentence T4 after the target text sentence T3, there may be a possibility that the text sentence T4 contains misrecognized characters and the correction process of the text sentence T4 has not been performed at the time when the correction process for the target text sentence T3 is performed. For this reason, the control unit 11 may further add the text sentence immediately after the text sentence T4 to the target text sentence.

[0074] Also, as another embodiment, the control unit 11 may extract the text sentences T2 and T4 before and after the target text sentence T3 as the target text sentence.

[0075] Thus, the control unit 11 may extract one or more text sentences before the target text sentence T3 as the correction text sentence, or may extract one or more text sentences after the target text sentence T3 as the correction text sentence, or may extract one or more text sentences before and after the target text sentence T3 as the correction text sentence.

[0076] Note that when the control unit 11 emphasizes real-time speech recognition, it is desirable to extract the text sentence before the target text sentence as the correction text sentence.

[0077] Also, as another embodiment, when the control unit 11 creates minutes of a meeting or the like after the meeting ends, it may use all the text sentences of the text information as the correction text sentence and execute correction processing on the target text sentence.

[0078] Also, as another embodiment, whether to use the text sentence before the target text sentence as the correction text sentence, whether to use the text sentence after the target text sentence as the correction text sentence, or whether to use the text sentences before and after the target text sentence as the correction text sentence may be set by the user in advance. Also, when selecting a text sentence after the target text sentence, an upper limit value of the number of selectable text sentences may be set so as not to impair the real-time nature of speech recognition.

[0079] As another embodiment, when the masking rate (the ratio of the number of characters of the misrecognition candidate) of the text sentence immediately before the target text sentence is high, the control unit 11 may further extract the text sentence before that as the correction text sentence. Note that the relevance of the content (topic) may decrease as the distance from the target text sentence increases. Therefore, when the masking rate of the text sentence immediately before the target text sentence is high, the control unit 11 may extract the text sentence immediately after the target text sentence as the correction text sentence.

[0080] Incidentally, the topic of conversation may change before and after the target text sentence. For example, the target text sentence may be a text sentence for the speech of a new topic, and the topic may be different from the text sentence immediately before. In this case, for example, if the correction process is performed based on the target text sentence and the immediately preceding text sentence (corrected text sentence), the prediction accuracy of the correct characters may decrease. Therefore, the control unit 11 may extract, as the corrected text sentence, a text sentence relevant to the content of the target text sentence from the converted text information. For example, the control unit 11 determines the topic of each converted text sentence based on the set topic (see FIG. 5), and extracts, as the corrected text sentence, a text sentence having the same topic as the target text sentence. Note that the control unit 11 may analyze the content of each text sentence regardless of the topic and determine the degree of relevance (similarity) to the target text sentence.

[0081] Further, the control unit 11 may total the frequency of appearance of words from the beginning of the conversation, and determine that there is no relevance to other text sentences when the number of words that have not appeared in the past in the target text sentence is equal to or more than a threshold value (for example, 5).

[0082] As another embodiment, the control unit 11 may be able to correct the recognition result displayed on the conference screen P2. For example, the recognition result is displayed in the conversation display area P22 of the conference screen P2 shown in FIG. 12. When the user selects (clicks) a predetermined text sentence (target text sentence) on the conference screen P2, the control unit 11 displays a reception screen P3 for receiving a selection operation of a corrected text sentence (corrected sentence). On the reception screen P3, selection buttons for selecting the additional quantity of the text sentence before the target text sentence and the additional quantity of the text sentence after the target text sentence are displayed. For example, when the user sets the quantity of the "subsequent sentence" to "1" and presses the "execute" button on the reception screen P3, the control unit 11 executes the correction process for the target text sentence based on the target text sentence and one text sentence (corrected text sentence) immediately after the target text sentence. The control unit 11 replaces the target text sentence displayed in the conversation display area P22 with the corrected text sentence.

[0083] In this way, the user can check and correct the recognition result displayed on the conference screen P2.

[0084] Note that the reception screen P3 is not limited to the configuration shown in FIG. 12 and may be composed of various operation buttons. For example, an "automatic correction" button may be displayed on the reception screen P3. When the user presses the "automatic correction" button, the control unit 11 automatically determines the text sentence before and after the target text sentence and extracts it as a corrected text sentence.

[0085] Note that in the above-described embodiment, an example in which only one conference room R1 is configured is shown. However, the conference support system 100 of the present disclosure may be configured to perform an online conference by connecting the conference rooms R1 and R2 via a network. In this case, for example, the voice spoken by the user in the conference room R1 is transferred to the conference room R2 and reproduced, and the voice spoken by the user in the conference room R2 is transferred to the conference room R1 and reproduced.

[0086] Note that the information processing system of the present disclosure may include only a speech-to-text function that converts the voice input from the outside into text and corrects and outputs the misrecognized characters. The speech-to-text function may be applied to a conference as shown in the above-described embodiment, or may be applied to fields other than a conference.

[0087] [Supplementary Note of the Disclosure] The following describes the outline of the disclosure extracted from the above-described embodiment. Note that each configuration and each processing function described in the following supplementary note can be arbitrarily selected and combined.

[0088] <Supplementary Note 1> A conversion processing unit that recognizes the voice spoken by the user and converts it into text information, A first extraction processing unit that extracts a first text sentence of a predetermined length from the text information converted by the conversion processing unit, When the first text sentence extracted by the first extraction processing unit includes characters with misrecognition candidates, a determination processing unit that determines whether the total number of characters in the first text sentence is less than a predetermined number of characters, or whether the ratio of the number of characters of the misrecognition candidates to the total number of characters in the first text sentence is equal to or greater than a predetermined ratio, When it is determined by the determination processing unit that the total number of characters is less than the predetermined number of characters, or when it is determined that the ratio is equal to or greater than the predetermined ratio, a second extraction processing unit that extracts a second text sentence different from the first text sentence from the text information, Based on the first text sentence and the second text sentence, a correction processing unit that corrects misrecognized characters among the misrecognition candidate characters, An information processing system comprising the same.

[0089] <Appendix 2> The correction processing unit predicts the correct characters corresponding to the misrecognized characters based on the first text sentence with the misrecognition candidate characters masked and the second text sentence, and replaces the masked misrecognized characters with the correct characters. The information processing system according to Appendix 1.

[0090] <Appendix 3> The second extraction processing unit extracts, as the second text sentence, the text sentence immediately before the first text sentence in time series from the text information. The information processing system according to Appendix 1 or 2.

[0091] <Appendix 4> The second extraction processing unit extracts, as the second text sentence, the text sentence immediately after the first text sentence in time series from the text information. The information processing system according to any one of claims 1 to 3.

[0092] <Appendix 5> The second extraction processing unit extracts, as the second text sentence, a text sentence related to the content of the first text sentence from the text information. The information processing system according to any one of Supplementary Notes 1 to 4.

[0093] <Supplementary Note 6> The second extraction processing unit extracts the second text sentence until the total number of characters of the first text sentence and the second text sentence becomes equal to or more than the predetermined number of characters, or until the ratio of the number of characters of the misrecognition candidate to the total number of characters of the first text sentence and the second text sentence becomes less than the predetermined ratio. The information processing system according to any one of Supplementary Notes 1 to 5.

[0094] <Supplementary Note 7> The second extraction processing unit receives an operation of selecting one or a plurality of text sentences from among a plurality of text sentences included in the recognition result on a display screen that displays the recognition result corresponding to the text information, and extracts the selected one or a plurality of text sentences as the second text sentence. The information processing system according to any one of Supplementary Notes 1 to 6.

[0095] <Supplementary Note 8> Causing the display device to display the first text sentence in which the misrecognized characters have been corrected by the correction processing unit. The information processing system according to any one of Supplementary Notes 1 to 7.

Explanation of Signs

[0096] 100: Conference Support System 1: Information Processing Apparatus 2: Audio Device 3: User Terminal 4: Display Device 11: Control Unit 12: Storage Unit 13: Communication Unit 111: Setting Processing Unit 112: Acquisition Processing Unit 113: Speech Recognition Processing Unit 114: Target Text Extraction Processing Unit 115: Recognition Judgment Processing Unit 116: Correction Processing Unit 117: Character determination processing unit 118: Correction text extraction processing unit 119: Display processing unit P1: Setting screen P2: Conference screen P21: Summary display area P22: Conversation display area P3: Reception screen T3: Target text sentence T2: Correction text sentence

Claims

1. A conversion processing unit that recognizes the voice spoken by a user and converts it into text information; A first extraction processing unit that extracts a first text sentence of a predetermined length from the text information converted by the conversion processing unit; When the first text sentence extracted by the first extraction processing unit contains characters of misrecognition candidates, it is determined whether the total number of characters in the first text sentence is less than a predetermined number of characters, or whether the ratio of the number of characters of the misrecognition candidates to the total number of characters in the first text sentence is equal to or greater than a predetermined ratio. A determination processing unit; When it is determined by the determination processing unit that the total number of characters is less than the predetermined number of characters, or when it is determined that the ratio is equal to or greater than the predetermined ratio, a second extraction processing unit that extracts a second text sentence different from the first text sentence from the text information; A correction processing unit that corrects misrecognized characters among the misrecognition candidate characters based on the first text sentence and the second text sentence; An information processing system comprising:

2. Based on the first text sentence in which the characters of the misrecognition candidates are masked and the second text sentence, the correction processing unit predicts the correct characters corresponding to the misrecognized characters and replaces the masked misrecognized characters with the correct characters. The information processing system according to Claim 1.

3. The second extraction processing unit extracts, as the second text sentence, the text sentence immediately before the first text sentence in time series from the text information. The information processing system according to Claim 1.

4. The second extraction processing unit extracts, as the second text sentence, the text sentence immediately after the first text sentence in time series from the text information. The information processing system according to Claim 1.

5. The second extraction processing unit extracts, as the second text sentence, a text sentence related to the content of the first text sentence from the text information. The information processing system according to Claim 1.

6. The second extraction processing unit extracts the second text sentence until the total number of characters of the first text sentence and the second text sentence becomes equal to or greater than the predetermined number of characters, or until the ratio of the number of characters of the misrecognition candidates to the total number of characters of the first text sentence and the second text sentence becomes less than the predetermined ratio. The information processing system according to Claim 1.

7. In a display screen that displays a recognition result corresponding to the text information, the second extraction processing unit receives an operation of selecting one or more text sentences from among a plurality of text sentences included in the recognition result, and extracts the selected one or more text sentences as the second text sentence. The information processing system according to claim 1.

8. Causing the display device to display the first text sentence in which the misrecognized characters have been corrected by the correction processing unit. The information processing system according to any one of claims 1 to 7.

9. Recognizing the voice spoken by the user and converting it into text information, extracting a first text sentence of a predetermined length from the text information, when the first text sentence includes misrecognition candidate characters, determining whether the total number of characters in the first text sentence is less than a predetermined number of characters, or whether the ratio of the number of misrecognition candidate characters to the total number of characters in the first text sentence is equal to or greater than a predetermined ratio, when it is determined that the total number of characters is less than the predetermined number of characters, or when it is determined that the ratio is equal to or greater than the predetermined ratio, extracting a second text sentence different from the first text sentence from the text information, correcting the misrecognized characters among the misrecognition candidate characters based on the first text sentence and the second text sentence, An information processing method executed by one or more processors.

10. Recognizing the voice spoken by the user and converting it into text information, extracting a first text sentence of a predetermined length from the text information, when the first text sentence includes misrecognition candidate characters, determining whether the total number of characters in the first text sentence is less than a predetermined number of characters, or whether the ratio of the number of misrecognition candidate characters to the total number of characters in the first text sentence is equal to or greater than a predetermined ratio, when it is determined that the total number of characters is less than the predetermined number of characters, or when it is determined that the ratio is equal to or greater than the predetermined ratio, extracting a second text sentence different from the first text sentence from the text information, correcting the misrecognized characters among the misrecognition candidate characters based on the first text sentence and the second text sentence, An information processing program for causing one or more processors to execute.

Citation Information

Patent Citations

  • Information processing device, information processing method, and program

    JP7216863B1