Information processing device, information processing method and information processing system

The system addresses inaccuracies in meeting minutes by integrating voice and operation data correction, ensuring accurate reflection of participant operations in speech recognition results.

JP2025135904APending Publication Date: 2025-09-19SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024033974
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-06
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing meeting systems fail to accurately reflect participant operations in automatically generated meeting minutes, leading to inaccuracies in speech recognition results.

Method used

An information processing device and system that incorporates voice recognition and operation information acquisition to correct character data based on participant operations, using a voice information acquisition unit, voice recognition unit, and character data correction unit to generate corrected character data.

Benefits of technology

The system accurately reflects participant operations, enhancing the accuracy of voice recognition results in meeting minutes by correcting character data based on operation details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025135904000001_ABST
    Figure 2025135904000001_ABST
Patent Text Reader

Abstract

To provide an information processing device, an information processing method and an information processing system that make further corrections on character data converted from speech data of each utterance in a conference based upon details of an operation on a terminal device by a conference participant.SOLUTION: An information processing device 100 comprises a speech information acquisition unit 111, a speech recognition unit 112, an operation information acquisition unit 113, and a character data correction unit 117. The speech information acquisition unit 111 acquires speech information Ix including speech data Dx and utterance time tx for each of utterances X by a plurality of participants P of the conference. The speech recognition unit 112 converts the speech data Dx into character data Dc. The operation information acquisition unit 113 acquires operation information Iy including operation details Cy and operation time ty for each of operations by the participants P. The character data correction unit 117 generates corrected character data Dcc obtained by correcting the character data Dc based upon the operation information Iy.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing device, an information processing method, and an information processing system. [Background technology]

[0002] In the browsing system described in Patent Document 1, the identification information of the speaker and the text data of the speech recognition results are stored in association with each other for each time information of the speech. On the speech list screen of a series of speech recognition results, the speeches can be marked and the speeches can be searched according to specified search conditions. The search results are displayed on the search screen. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2010-238050 Summary of the Invention [Problem to be solved by the invention]

[0004] In the past, even if meeting participants themselves performed operations such as taking notes on a terminal device, the content of those operations was not reflected in the automatically generated minutes, and the results of speech recognition were sometimes inaccurate.

[0005] The present disclosure aims to provide an information processing device, an information processing method, and an information processing system in which further correction is made to text data converted from audio data for each statement made during a meeting based on the operations performed by participants on their terminal devices. [Means for solving the problem]

[0006] The information processing device of the present disclosure includes a voice information acquisition unit, a voice recognition unit, an operation information acquisition unit, and a character data correction unit. The voice information acquisition unit acquires voice information including voice data and a time of speech for each speech made by a plurality of participants in a conference. The voice recognition unit converts the voice data into character data. The operation information acquisition unit acquires operation information including the operation content and the operation time for each operation made by the participants. The character data correction unit generates corrected character data by correcting the character data based on the operation information.

[0007] The information processing method disclosed herein includes a voice information acquisition step, a voice recognition step, an operation information acquisition step, and a character data correction step. The voice information acquisition step acquires voice information including voice data and a speech time for each speech made by a plurality of participants in a conference. The voice recognition step converts the voice data into character data. The operation information acquisition step acquires operation information including operation content and operation time for each operation made by the participants. The character data correction step generates corrected character data by correcting the character data based on the operation information.

[0008] The information processing system of the present disclosure includes an information processing device and a terminal device. The information processing device has a first communication unit, a voice information acquisition unit, a voice recognition unit, an operation information acquisition unit, and a character data correction unit. The first communication unit communicates with the terminal device. The voice information acquisition unit acquires voice information including voice data and a speech time for each speech made by a plurality of participants in a conference. The voice recognition unit converts the voice data into character data. The operation information acquisition unit acquires operation information including an operation content and an operation time for each operation made by the participant. The character data correction unit generates corrected character data by correcting the character data based on the operation information. The terminal device has a display unit, a display control unit, and a second communication unit. The display control unit controls the display unit. The second communication unit communicates with the first communication unit. A predetermined proximity time is set for each piece of character data. When the operation content is a character string including a second predetermined symbol, the character data correction unit of the information processing device generates a corrected character string by correcting the second predetermined symbol included in the character string based on the character data corresponding to the predetermined proximity time including the operation time corresponding to the operation content and the corrected character data. The display control unit of the terminal device displays the corrected character string obtained from the information processing device via the second communication unit on the display unit. [Effects of the Invention]

[0009] According to the present disclosure, the text data converted from the voice data of each utterance made during a conference is further corrected based on the operation details of the participants on their terminal devices. As a result, the operation details of the participants are appropriately reflected, making it possible to generate more accurate voice recognition results. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a block diagram showing a configuration of a minutes generation system 1 including an information processing device 100 according to a first embodiment of the present disclosure. [Figure 2] 10 is a time chart showing the temporal relationship between statements X and operations Y by each participant P. [Figure 3] 10 is an explanatory diagram illustrating an example of an operation by a participant P to create his / her own minutes, memos, etc. on the terminal device 10. FIG. [Figure 4] 10 is an explanatory diagram showing an example in which operation inputs and the like on the terminal device 10 are reflected in the minutes generated by the information processing device 100. FIG. [Figure 5] 1 is a flowchart showing an information processing method by the information processing device 100. DETAILED DESCRIPTION OF THE INVENTION

[0011] Embodiments of the present disclosure will be described with reference to the drawings. In the drawings, the same or corresponding parts are designated by the same reference numerals, and description thereof will not be repeated.

[0012] First Embodiment 1.1 Overall configuration of the minutes generation system 1 First, the overall configuration of a minutes-of-meeting generation system 1 including an information processing device 100 that automatically generates minutes of a meeting will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the minutes-of-meeting generation system 1 including the information processing device 100 according to the first embodiment of the present disclosure.

[0013] As shown in Fig. 1, the minutes generation system 1 includes an information processing device 100 and a terminal device 10. The terminal device 10 is a device operated by a participant P to create his or her own minutes, memos, etc. during or before or after a meeting. The terminal device 10 is connected to the information processing device 100 via a wired or wireless network, but the connection method is not limited to a network connection. For example, the terminal device 10 may be directly connected to the information processing device 100 via a USB cable or a dedicated cable, or via a wireless connection method such as Bluetooth (registered trademark).

[0014] 1.2 Configuration of the information processing device 100 The information processing device 100 includes a CPU (Central Processing Unit) 110, a display unit 120 having a display screen 121, a storage unit 130, a communication unit 140, and a voice input unit 150.

[0015] The voice input unit 150 receives input of the voice of a utterance X by each participant P and outputs voice data Dx to the voice information acquisition unit 111. The voice input unit 150 may output to the voice information acquisition unit 111 not only voices that are directly input, but also voice data that has been transmitted from the outside via a network and received by the communication unit 140, for example.

[0016] The audio input unit 150 may be, for example, a microphone, but is not limited thereto. The audio input unit 150 may be built into the information processing device 100. The audio input unit 150 may be a wired or wireless handheld microphone, or a microphone of a wireless headset. Two or more of these may be used in combination. When a wireless headset is connected to the terminal device 10, audio data from the microphone is received by the information processing device 100 via, for example, the communication units 16 and 140, and transmitted to the audio input unit 150. However, the wireless headset may be directly connected wirelessly to the information processing device 100. When multiple wireless headsets are used, each participant P uses a different wireless headset, so that utterances X of multiple participants P are not mixed in one piece of audio data. This makes it easier to identify the speaker of the audio data.

[0017] The CPU 110 controls each unit of the information processing device 100. The CPU 110 has a voice information acquisition unit 111, a voice recognition unit 112, an operation information acquisition unit 113, a minutes generation unit 114, a display control unit 115, a clock unit 116, and a character data correction unit 117.

[0018] The CPU 110 may be, for example, a processor or an MPU (Micro Processing Unit), but is not limited to these. For example, if an OS (Operating System, sometimes called "basic software") that runs on the CPU 110 and programs for the functions corresponding to each of the above-mentioned units are stored in the storage unit 130, the above-mentioned units are realized by the CPU 110 executing the OS and programs. Examples of OS include, but are not limited to, Microsoft Windows (registered trademark), Android (registered trademark), and Linux (registered trademark). Each of the above-mentioned units may be configured with, for example, an electronic circuit, a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), or the like, but is not limited to these.

[0019] The clock unit 116 has a clock function and outputs at least the current hour, minute, and second data. The clock unit 116 may further have a calendar function and be able to output year, month, and day data.

[0020] The voice information acquisition unit 111 acquires voice information Ix including voice data Dx and speech time tx for each speech X made by multiple participants P of the conference. The voice information acquisition unit 111 acquires the voice data Dx from the voice input unit 150 and acquires the speech time tx from the clock unit 116. For example, if the speaker of each voice data Dx is known in advance because each participant P uses an individual wireless headset or the like, the voice information Ix may include the speaker's name.

[0021] The voice recognition unit 112 converts the voice data Dx into character data Dc. Voice recognition may be, for example, an existing technology in which AI (Artificial Intelligence) analyzes words or conversations uttered by a person and converts them into text data, but is not limited to this.

[0022] The operation information acquisition unit 113 acquires operation information Iy including operation content Cy and operation time ty for each operation Y performed by a participant P. Examples of the operation content Cy include input of characters, symbols, etc., and specific operations. Examples of input of characters, symbols, etc. include, but are not limited to, input of a predetermined character string related to the progress of the conference and input of special symbols (e.g., ? or ☆). Examples of specific operations include, but are not limited to, click operations, double-click operations, tap operations, etc. Note that the operation content Cy is received by the information processing device 100 via, for example, the communication unit 16 and the communication unit 140, and transmitted to the operation information acquisition unit 113.

[0023] The character data correcting unit 117 generates corrected character data Dcc by correcting the character data Dc based on the operation information Iy.

[0024] Therefore, the character data Dc converted from the voice data Dx for each utterance X during the conference is further corrected based on the operation content, such as memo entry, made by the participant P. As a result, the operation content Cy made by the participant P is appropriately reflected, thereby enabling the generation of more accurate voice recognition results. Note that the character data Dc includes, for example, the content of each utterance X (so-called "transcription"), but is not limited to this.

[0025] The character data correction unit 117 may generate corrected character data Dcc by regarding the operation content Cy corresponding to the operation time ty included in the predetermined proximity time tn determined for each character data Dc as corresponding to the character data Dc. The predetermined proximity time tn is the time from the utterance time tx corresponding to the character data Dc until a predetermined time has elapsed. This predetermined time is expected to be, for example, several seconds to several tens of seconds, but is not limited to such a time. Furthermore, if the operation content Cy includes character input data, the corresponding character data Dc may be identified based on its approximation to the input data. By combining a character-based proximity judgment, the predetermined time can be assumed to be longer. The predetermined proximity time tn will be described later with reference to FIG. 2.

[0026] Therefore, corrected character data Dcc is generated by reflecting the operation content Cy that is likely to actually correspond to character data Dc converted from voice data Dx for each utterance X. As a result, the corrected character data Dcc can be displayed.

[0027] The minutes generation unit 114 generates minutes data Dm based on the audio information Ix, the character data Dc, and the corrected character data Dcc. Therefore, the minutes data Dm is generated based not only on the character data Dc converted from the audio data Dx, but also on the corrected character data Dcc. As a result, the minutes data Dm can be generated more accurately. Note that the minutes data Dm includes, for example, the content of each utterance X (so-called "transcription"), a summary, and an excerpt of the utterance content, but is not limited to these.

[0028] The display unit 120 displays at least one of the character data Dc and the corrected character data Dcc. Therefore, the character data Dc and the corrected character data Dcc are displayed on the display unit 120. As a result, if both are displayed side by side, it becomes possible to visually confirm how the original character data Dc has been corrected. However, during normal use, there is no need to display the pre-corrected character data Dc when the corrected character data Dcc is present.

[0029] The display unit 120 includes those with and without memory properties. Examples of display units without memory properties include, but are not limited to, liquid crystal and organic electroluminescence (EL). Examples of display units with memory properties include, but are not limited to, electronic paper.

[0030] The display control unit 115 may further display the operation information Iy on the display unit 120. Therefore, the operation information Iy including the operation content Cy and the operation time ty is displayed on the display unit 120. As a result, it becomes possible to visually confirm what operation was performed for what utterance X.

[0031] The storage unit 130 stores the voice data Dx linked to the utterance time tx, and stores the operation content Cy linked to the operation time ty. Therefore, the voice data Dx for each utterance X is stored linked to the utterance time tx, and the operation content Cy is stored linked to the operation time ty. As a result, it is possible to generate the corrected character data Dcc and the minutes data Dm not only during the conference but also after the conference has ended.

[0032] The storage unit 130 also stores information and data necessary for controlling each unit of the information processing device 100. The storage unit 130 may store or store an OS, programs, etc. executed by the CPU 110. The storage unit 130 includes memory, specifically, volatile memory and nonvolatile memory. Volatile memory includes, for example, dynamic random access memory (DRAM) and static random access memory (SRAM), but is not limited to these. Nonvolatile memory includes, for example, read-only memory (ROM), flash memory, solid state drive (SSD), and hard disk, but is not limited to these.

[0033] The communication unit 140 is a communication interface, and is connected to the communication unit 16 of the terminal device 10 via a wired or wireless connection. This enables two-way communication between the terminal device 10 and the information processing device 100. The communication unit 140 is an example of a "first communication unit" in the present disclosure.

[0034] 1.3 Temporal relationship between statements X and actions Y by each participant P Next, an example of determining whether or not a utterance X and an operation Y by each participant P are temporally close to each other will be described with reference to Fig. 2. Fig. 2 is a time chart showing the temporal relationship between the utterance X and the operation Y by each participant P.

[0035] As shown in Figure 2, assume that the audio data Dx for each utterance X in the order of the utterance time tx is managed with a serial number or time information. For example, the nth audio data Dx is represented here as "Dr(n)". Each audio data Dx is shown as a rectangle, and the width of this rectangle corresponds to the duration of each utterance X. The left edge of the rectangle corresponds to the start time of utterance X, and the right edge of the rectangle corresponds to the end time of utterance X.

[0036] Here, the statement time tx corresponds to the start time (left edge of the rectangle) of each statement X, but is not limited to this correspondence. The statement time tx may be set to, for example, the midpoint between the start time and the end time.

[0037] Each piece of voice data Dx is associated with character data Dc converted by the voice recognition unit 112. This character data Dc is also managed with a serial number or time information, just like the voice data Dx. The nth piece of voice data Dx(n) corresponds to the nth piece of character data Dc(n).

[0038] On the other hand, operation contents Cy performed by participant P during the conference are also managed with a serial number or time information attached. For example, the m-th operation contents Cy is represented here as "Cy(m)".

[0039] Similar to the character data correction unit 117 described above, the minutes generation unit 114 considers the operation content Cy corresponding to the operation time ty included in the predetermined proximity time tn determined for each character data Dc to correspond to the character data Dc. For example, below the character data Dc(n) in the upper center of FIG. 2, the range of the predetermined proximity time tn is illustrated immediately above the time axis t. This predetermined proximity time tn includes the operation time ty of the operation content Cy(m). Therefore, the minutes generation unit 114 considers the operation content Cy(m) to correspond to the character data Dc(n).

[0040] That is, the minutes generating unit 114 determines that the operation content Cy(m) is temporally close to the nth comment X. This means that there is a high possibility that the operation content Cy(m) corresponds to the nth comment X.

[0041] However, when the duration is long, such as in the (n+1)th utterance X, it may be difficult to accurately determine whether any operation Y is temporally close to any part of the utterance X. In such a case, it is preferable to subdivide the utterance X into individual sentences and treat each separately. For example, if the (n+1)th utterance X is divided into four parts and treated, it can be determined that the second part and the operation Cy(m+1) are temporally close to each other.

[0042] 1.4 Configuration of the terminal device 10 As described above, the terminal device 10 is a device operated by a participant P during a conference or before or after the conference to create his or her own minutes, memos, etc. Examples of the terminal device 10 include, but are not limited to, a laptop computer, a smartphone, and a tablet.

[0043] 1, the terminal device 10 includes a CPU 11, an imaging unit 12, an operation unit 13, a display unit 14, a storage unit 15, and a communication unit 16. These hardware configurations are similar to those of general notebook computers, smartphones, tablets, etc., and therefore detailed description thereof will be omitted. The communication unit 16 is an example of the "second communication unit" of the present disclosure.

[0044] The imaging unit 12 reads the face of the participant P who is operating the terminal device 10. It is also possible to identify the participant P by recognizing the facial image. The imaging unit 12 may be, for example, a built-in camera, but is not limited to this.

[0045] The operation unit 13 receives operation inputs. Examples of the operation unit 13 include, but are not limited to, buttons, a keyboard, and a touch panel that allows touch operations.

[0046] Examples of the display unit 14 include, but are not limited to, liquid crystal and organic EL. The display unit 14 may be a touch panel, and may also serve as the operation unit 13.

[0047] 1.5 Display example of minutes by the information processing device 100 Next, an example of displaying minutes by the information processing device 100 will be described with reference to Fig. 3 and Fig. 4. Fig. 3 is an explanatory diagram illustrating an example of an operation by a participant P to create his / her own minutes, memos, etc. on the terminal device 10. Fig. 4 is an explanatory diagram showing an example in which operation inputs, etc. on the terminal device 10 are reflected in the minutes generated by the information processing device 100. Fig. 5 is an explanatory diagram showing an example in which main points and summaries are displayed together with the minutes.

[0048] 1.5.1 Display example on terminal device 10 3, when a participant P performs an operation such as inputting characters from the operation unit 13 of the terminal device 10 connected to the information processing device 100, the input content is displayed on the display unit 14, and additional display is also performed by the information processing device 100. Note that the operation content Cy on the terminal device 10 is stored in the storage unit 130 each time, linked to the operation time ty.

[0049] For example, when a simple sentence such as "The investment scale is large" is input, the simple sentence is displayed on the display unit 14. The operation time ty is displayed right-justified on the same line as the line on which the simple sentence is displayed.

[0050] For example, suppose that the character string "pigupada" is included in the minutes data Dm displayed on the display unit 120 of the information processing device 100. The participant P, realizing that this is a mistranslation, may input "×pigupada → BIGPAD" from the operation unit 13.

[0051] When the operation information acquiring unit 113 acquires such operation content Cy, the character data correcting unit 117 corrects the character data Dc "pigupada" to generate corrected character data Dcc "BIGPAD".

[0052] Storage unit 130 may store character data Dc and corrected character data Dcc generated by character data correction unit 117 in association with each other. Therefore, the correction content of such erroneous conversion is stored in storage unit 130 in association with the original character data Dc and corrected character data Dcc. As a result, it becomes possible to avoid similar erroneous conversions thereafter and perform correct conversions from the beginning.

[0053] If participant P enters a string containing a special symbol (in this case, "?") in the part he or she was unable to hear, such as "Number of employees???", the content will be completed if possible based on the correction character data Dcc and displayed as "Number of employees: 120".

[0054] 1.5.2 Display example on information processing device 100 4, the character data Dc or the corrected character data Dcc is displayed on the display unit 120 of the information processing device 100. Specifically, for example, each transcribed utterance X is displayed in chronological order together with the utterance time tx and the name of the speaker.

[0055] For example, on the display unit 120 of the information processing device 100, for Yamada's statement X at 14:04, the character data Dc "equation book" which is a mistaken conversion of the voice data Dx that sounds like "toushikibo" is displayed with a strikethrough. Immediately afterwards, corrected character data Dcc "investment scale" which is a correction of this character data Dc is displayed. This is an automatic correction in real time that reflects the operation content Cy from the terminal device 10 at the nearby time of 14:20, specifically the input "large investment scale."

[0056] When the operation content Cy includes the input of a first predetermined symbol, the character data correction unit 117 generates corrected character data Dcc according to the function previously assigned to the first predetermined symbol. Therefore, appropriate correction is performed according to the function previously assigned to the first predetermined symbol. As a result, advanced correction that accurately reflects the intention of the operation content Cy of the participant P is possible.

[0057] For example, Kinoshita's utterance X at 14:19 includes the character string "pigupada," which was incorrectly converted without being correctly recognized by the speech recognition unit 112. This character data Dc is later corrected by reflecting the operation Cy entered from the terminal device 10 at the nearby time of 14:25, specifically, "×pigupada → BIGPAD." Here, "×" and "→" are pre-assigned to the erroneous conversion correction function. Participant P is assumed to enter an example of an erroneous conversion after "×," then enter "→" before entering a correct conversion example. However, "×" and "→" are merely examples. The symbols used here are examples of the "first predetermined symbol" of this disclosure.

[0058] 1.5.3 Another display example on the terminal device 10 4, the display unit 14 of the terminal device 10 being used by the participant P displays "Number of employees ???" in the third row from the top. This "?" is set in advance as one of the special symbols to be used when the participant wants the information processing device 100 to complete numbers that they were unable to hear. The symbol used in this case is an example of the "second predetermined symbol" of the present disclosure.

[0059] 2, a predetermined proximity time tn is determined for each character data Dc. When an operation content Cy is a character string C including a second predetermined symbol, the character data correction unit 117 of the information processing device 100 generates a corrected character string Cc by correcting the second predetermined symbol included in the character string C, based on the character data Dc and corrected character data Dcc corresponding to the predetermined proximity time tn including the operation time ty corresponding to the operation content Cy.

[0060] The display control unit 11a of the terminal device 10 causes the display unit 14 to display the corrected character string Cc acquired from the information processing device 100 via the communication unit 16.

[0061] Therefore, in the terminal device 10, the character string C input including the second predetermined symbol is corrected by the correction process in the information processing device 100 and displayed as a corrected character string Cc in which the second predetermined symbol has been correctly corrected. As a result, for example, even numerical values ​​that could not be heard are correctly displayed, so that even the participants themselves can keep a useful record of the conference.

[0062] For example, the operation Cy "Number of employees ???" input to the terminal device 10 at 14:38 is corrected by the information processing device 100 and displayed as "Number of employees 5,000."

[0063] 1.6 Information processing method for generating minutes by the information processing device 100 Next, an outline of an information processing method for generating minutes by the information processing device 100 will be described with reference to Fig. 5. Fig. 5 is a flowchart showing the information processing method by the information processing device 100.

[0064] 5, first, in step S1, the voice information acquisition unit 111 acquires voice information Ix including voice data Dx and speech time tx for each speech X made by multiple participants P of the conference. Note that step S1 is an example of the "voice information acquisition step" of the present disclosure.

[0065] In step S2, the voice recognition unit 112 analyzes the voice data Dx and converts it into character data Dc. Note that step S2 is an example of the "voice recognition step" of the present disclosure.

[0066] In step S3, the operation information acquiring unit 113 acquires operation information Iy including operation content Cy and operation time ty for each operation Y by the participant P. Note that step S3 is an example of the "operation information acquiring step" of the present disclosure.

[0067] In step S4, the character data correcting unit 117 corrects the character data Dc based on the operation information Iy to generate corrected character data Dcc, and then ends the series of processes. Note that step S4 is an example of the "character data correcting step" of the present disclosure.

[0068] Therefore, the character data Dc converted from the voice data Dx for each utterance X during the conference is further corrected based on the operation details such as memo entry by the participant P. As a result, the operation details Cy by the participant P are appropriately reflected, making it possible to generate more accurate voice recognition results.

[0069] The present invention can be embodied in various other forms without departing from its spirit or main features. Therefore, the above-described embodiments are merely illustrative in all respects and should not be interpreted as limiting. The scope of the present invention is defined by the claims and is not limited to the text of the specification. Furthermore, all modifications and variations within the equivalent range of the claims are within the scope of the present invention. [Industrial Applicability]

[0070] The present disclosure is applicable to an information processing device, an information processing method, and an information processing system. [Explanation of symbols]

[0071] 1. Meeting minutes generation system 10 Terminal Equipment 11 CPU 11a Display control unit 12 Imaging unit 13 Control section 14 Display section 15 Storage section 16 Communications Department (Second Communications Department) 100 Information processing device 110 CPU 111 Voice information acquisition unit 112 Voice Recognition Unit 113 Operation information acquisition unit 114 Minutes Generation Department 115 Display control unit 116 Clock Section 117 Character data correction unit 120 Display section 121 Display screen 130 Storage section 140 Communications Department (First Communications Department) 150 Audio input section

Claims

1. a voice information acquisition unit that acquires voice information including voice data and speech time for each speech made by a plurality of participants in a conference; a voice recognition unit that converts the voice data into character data; an operation information acquisition unit that acquires operation information including operation content and operation time for each operation performed by the participant; a character data correction unit that generates corrected character data by correcting the character data based on the operation information; An information processing device comprising:

2. the character data correction unit generates the corrected character data by regarding the operation content corresponding to the operation time included in the predetermined proximity time determined for each of the character data as corresponding to the character data; The information processing device according to claim 1 , wherein the predetermined proximity time is a time from the speech time corresponding to the character data until a predetermined time has elapsed.

3. The information processing apparatus according to claim 1 , further comprising a display unit that displays at least one of the character data and the corrected character data.

4. The information processing apparatus according to claim 3 , further comprising a display control unit that causes the operation information to be further displayed on the display unit.

5. The information processing device according to claim 1 , further comprising a storage unit that stores the voice data in association with the utterance time and stores the operation content in association with the operation time.

6. The information processing apparatus according to claim 5 , wherein the storage unit stores the character data and the corrected character data generated by the character data correction unit in association with each other.

7. 3. The information processing device according to claim 1, wherein, when the operation content includes input of a first predetermined symbol, the character data correction unit generates the corrected character data in accordance with a function previously assigned to the first predetermined symbol.

8. The information processing apparatus according to claim 1 , further comprising: a minutes generating unit that generates minutes data based on the voice information, the character data, and the corrected character data.

9. a voice information acquisition step of acquiring voice information including voice data and a speech time for each speech made by a plurality of participants in the conference; a voice recognition step of converting the voice data into character data; an operation information acquisition step of acquiring operation information including operation content and operation time for each operation performed by the participant; a character data correcting step of generating corrected character data by correcting the character data based on the operation information; An information processing method, including:

10. An information processing system including an information processing device and a terminal device, The information processing device includes: a first communication unit that communicates with the terminal device; a voice information acquisition unit that acquires voice information including voice data and speech time for each speech made by a plurality of participants in a conference; a voice recognition unit that converts the voice data into character data; an operation information acquisition unit that acquires operation information including operation content and operation time for each operation performed by the participant; a character data correction unit that generates corrected character data by correcting the character data based on the operation information; and The terminal device A display unit; a display control unit that controls the display unit; a second communication unit that communicates with the first communication unit; and a predetermined proximity time is determined for each of the character data; the character data correction unit of the information processing device, when the operation content is a character string including a second predetermined symbol, generates a corrected character string in which the second predetermined symbol included in the character string is corrected based on the character data corresponding to the predetermined proximity time including the operation time corresponding to the operation content and the correction character data; The display control unit of the terminal device causes the correction character string acquired from the information processing device via the second communication unit to be displayed on the display unit.

Citation Information

Patent Citations

  • Browsing system and method, and program

    JP2010238050A