Information processing device, information processing method, program, and storage medium
The information processing device addresses the inefficiency of manual text-based searching by linking speech content text data with audio data, facilitating easy access to relevant audio playback in interviews.
Patent Information
- Application Number
- JP2024011641
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-30
- Publication Date
- 2025-08-12
AI Technical Summary
Existing methods for summarizing business negotiations require users to manually search through text data to find relevant audio recordings, which is time-consuming and inefficient.
An information processing device that generates linked utterance content text data by associating speech content text data with audio link text data, allowing direct access to corresponding audio data without manual searching.
Improves the ease of confirming interview details by enabling direct access to relevant audio playback from text data, reducing the time and effort required to locate specific portions of recorded interviews.
Smart Images

Figure 2025117015000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device and method, a program and a recording medium, and in particular to a technology for increasing the ease of checking the details of an interview, which are difficult to check by simply viewing the text data of the utterance content generated based on the recorded data of the interview. [Background technology]
[0002] For example, in the case of business negotiations conducted as part of a company's sales activities, where meetings are conducted through dialogue between multiple speakers, it is common practice to prepare minutes summarizing the content of the meetings.
[0003] Furthermore, as described in Patent Document 1 below, there is a technique for converting the speech of each person in a business negotiation into text based on audio recording data of the business negotiation, or for creating a summary of the content of the business negotiation based on the converted speech data. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent No. 7223469 Summary of the Invention [Problem to be solved by the invention]
[0005] Here, suppose that a user, such as a person who actually conducted an interview, finds a passage of interest in the minutes while reviewing the contents of the interview after the fact by referring to the minutes that were created. In this case, it would be sufficient to be able to understand the details from the text of the passage, but it is possible that the details cannot be understood from the text alone. It is conceivable that audio recordings of the interviews can be saved as in Patent Document 1, and if such audio recordings are available, the user can play back the relevant passages in the audio recordings to check the details. However, it takes a lot of time and effort for the user to search for the relevant part in the recorded data using only the text in the minutes as a clue.
[0006] The present invention has been made in consideration of the above circumstances, and aims to improve the ease with which a user can subsequently check the details of an interview when viewing the text data of the utterance content generated based on the recorded data of the interview. [Means for solving the problem]
[0007] The information processing device of the present invention comprises a generation processing unit that generates linked utterance content text data by associating utterance content text data, which is text data indicating the content of an utterance in an interview generated based on recorded interview data, with audio link text data, which is text data indicating a link to the audio data of the utterance portion in the recorded interview data that corresponds to the utterance content text data, and a transmission processing unit that transmits the linked utterance content text data generated by the generation processing unit to an external server device. By associating the voice link text data with the speech content text data as described above, it is no longer necessary for the user to search for the corresponding voice data portion in the recorded data when realizing audio playback of the speech portion corresponding to the speech content text data. [Effects of the Invention]
[0008] According to the present invention, it is possible to improve the ease of subsequent confirmation of the details of an interview when a user views the text data of the utterance content generated based on the recorded data of the interview. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a block diagram illustrating an example of the configuration of a minutes-taking support system according to an embodiment. [Figure 2] 1 is a perspective view showing an example of a schematic external appearance of a sound collection device configured as a smartphone. [Figure 3] FIG. 10 is a diagram illustrating an example of placement of a sound pickup device when recording an interview. [Figure 4] FIG. 10 is a diagram showing an example of a screen display of a sound collection device when recording an interview. [Figure 5] 1 is a block diagram illustrating an example of a hardware configuration of a sound collection device according to an embodiment. [Figure 6] FIG. 1 is a block diagram illustrating an example of a hardware configuration of an information processing apparatus according to an embodiment. [Figure 7] FIG. 2 is a block diagram illustrating an example of the hardware configuration of a minutes management server according to an embodiment. [Figure 8] FIG. 10 is a diagram showing an example of a registration screen for interview participant information in an embodiment. [Figure 9] FIG. 10 is a diagram showing an example of a recording list screen in the embodiment. [Figure 10] FIG. 10 is a diagram showing an example of a minutes creation screen (before speaker / group setting and analysis processing are performed) in an embodiment. [Figure 11] FIG. 10 is a diagram showing an example of a setting screen for setting speaker groups in the embodiment. [Figure 12] FIG. 10 is a diagram showing an example of information displayed in the speaker / group setting area after information has been set. [Figure 13] FIG. 10 is a diagram showing an example of a minutes creation screen (after setting speakers and groups) in an embodiment. [Figure 14] FIG. 10 is a diagram showing an example of display of analysis result text data on the minutes creation screen. [Figure 15] 10A and 10B are diagrams illustrating a method for displaying gist sentences of small categories in a gist list display area according to an embodiment. [Figure 16] 10 is an explanatory diagram of a method for displaying gist sentences of small categories in a gist list display area according to an embodiment. FIG. [Figure 17] 10A and 10B are explanatory diagrams illustrating an example of an operation for instructing audio playback of an utterance portion corresponding to a gist sentence. [Figure 18] FIG. 10 is a diagram for explaining transcription of a summary sentence into a minutes creation area. [Figure 19]FIG. 10 is a diagram for explaining the transcription of a summary sentence into the minutes creation area. [Figure 20] 10 is an explanatory diagram of uploading of minutes data with an audio link and processing related to the uploaded minutes data. FIG. [Figure 21] 10 is an explanatory diagram of an example of an operation for playing back audio while viewing minutes of a meeting. FIG. [Figure 22] FIG. 10 is an explanatory diagram of an example of an operation for playing back audio while viewing minutes of a meeting. [Figure 23] FIG. 2 is an explanatory diagram of functions of an information processing apparatus according to an embodiment. [Figure 24] 10 is an explanatory diagram of a process for generating gist sentences for each category according to an embodiment, and an audio link associated with the generated gist sentences. FIG. [Figure 25] FIG. 10 is a diagram showing an example of minutes data including audio link text data. [Figure 26] 10 is a flowchart of a process for generating minutes data in an embodiment. [Figure 27] 10 is a flowchart of a process related to access authentication in the embodiment. [Figure 28] FIG. 10 is a diagram illustrating an example of association of a voice link with recorded text data. [Figure 29] FIG. 10 is a diagram illustrating a configuration example of a modified minutes support system. [Figure 30] FIG. 10 is an explanatory diagram of a modified example relating to the display of analysis result information. [Figure 31] FIG. 10 is an explanatory diagram of a modified example of information display of a sound collection device during interview recording. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of the present invention will be described in the following order. <1. System configuration example> <2. Configuration examples of each device> <3. Flow from interview to minutes creation and example of display screen> <4. Example of functional configuration> <5. Processing Procedure> <6. Variations> <7. Summary of embodiments>
[0011] <1. System configuration example> FIG. 1 is a block diagram showing an example of the configuration of a minutes-taking support system 100 including an information processing device according to an embodiment of the present invention. As shown in the figure, the minutes support system 100 according to the embodiment includes a voice management server 1, a sound collection device 20, and a minutes management server 50.
[0012] The sound collection device 20 is equipped with a plurality of microphones arranged at different positions on the device body and a display unit for displaying information, one of the plurality of microphones being a microphone on the other side for collecting the speech of the other party during the interview, and the other being a microphone on the own side for collecting the speech of the own party during the interview. In this example, the sound collection device 20 is configured as a smartphone, and the minutes support system 100 of this example employs a configuration in which the smartphone is used as a sound collection device for recording the interview.
[0013] Figure 2 is an oblique view showing an example of the general appearance of a sound collection device 20 configured as a smartphone, where Figure 2A is an oblique view of the sound collection device 20 seen from the front side, and Figure 2B is an oblique view of the sound collection device 20 seen from the back (rear) side. Here, the "front" of the sound collection device 20 means the side on which the display surface of the display unit is located. 2A, the display screen of the display unit 27 is disposed on the front side of the sound collection device 20, and a call speaker 28a is disposed above the display screen. A first microphone (microphone) 33a used for NC (noise canceling) is disposed near the call speaker 28a. Further, a second microphone 34a used for collecting sound during a call is disposed below the sound collection device 20, specifically, in this example, on the underside of the sound collection device 20. In this example, two microphones are disposed as the second microphone 34a disposed below the sound collection device 20 to enable stereo sound collection. Here, with regard to the sound collection device 20 configured as a smartphone, the upper and lower parts are concepts in which the side on which the call speaker 28a is arranged is the upper side when the sound collection device is viewed from the front.
[0014] 2B, a camera unit 36 is disposed at one of the left and right ends of the upper part on the back side of the sound collection device 20. The camera unit 36 disposed on the back side allows the user to capture an image of a subject while checking the captured image displayed on the display unit 27. In the sound collection device 20 of this example, a first microphone 33b used for NC is also arranged near the camera unit .
[0015] As can be understood from the above explanation, the sound collection device 20 of this example is provided with first microphones 33a, 33b for NC located on the top of the device body, and second microphones 34a, 34b capable of collecting stereo sound located on the bottom of the device body. In this example, at least one of the first microphones 33a, 33b located at the top is defined as the above-mentioned local microphone, and the second microphones 34a, 34b located at the bottom are defined as the above-mentioned remote microphone. In this example, the first microphone 33a is defined as the local microphone.
[0016] Here, the second microphones 34a and 34b can be said to be microphones with higher performance than the first microphone 33a (or the first microphone 33b) because they are capable of collecting stereo sound.
[0017] In this example, it is assumed that directional microphones are used as the first microphones 33a and 33b and the second microphones 34a and 34b.
[0018] FIG. 3 is a diagram for explaining an example of placement of the sound collection device 20 when recording an interview. In this embodiment, an example of an "interview" is a "business negotiation" between a sales representative Hs and a customer Hc. The sales representative Hs is a person who participates in business negotiations, engages in sales activities, and also prepares minutes of the business negotiations. The sales representative Hs can be considered the host in a business meeting, while the customer Hc can be considered the guest in the meeting. The business negotiation here is premised on a face-to-face format where the sales representative Hs and the customer Hc actually meet face to face.
[0019] In many cases, business negotiations are conducted in a room such as a conference room, with the sales representative Hs seated on one side of a table or desk and the customer Hc seated on the other side. Here, the participants in the interview are assumed to be multiple sales representatives Hs and multiple customers Hc (specifically, two of each is used as an example), but it is sufficient for a business negotiation to be conducted with the participation of at least one sales representative Hs and one customer Hc.
[0020] In this case, the sound collection device 20 is a smartphone used by the sales representative Hs. When recording a business negotiation, the sales representative Hs places the sound collection device 20 on a support surface such as the top surface of a table as shown in the figure. As described above, in this example, the second microphones 34a and 34b are defined as the other party's microphones, so when the sound collection device 20 is placed on the support surface for recording, the second microphones 34a and 34b should be directed toward the customer Hc as the other party.
[0021] For this reason, the sound collection device 20 in this example has a function of displaying direction guidance information Id on the display unit 27, which guides the direction in which the other party's microphone (in this example, the second microphones 34a, 34b) can be pointed toward the other party when recording an interview.
[0022] FIG. 4 is a diagram showing an example of how the direction guidance information Id is displayed. As the direction guidance information Id, the display screen of the display unit 27 is configured to display information indicating the side on which at least the microphones on the other side (second microphones 34a, 34b) are located. 4A shows an example in which, as the direction guidance information Id, not only information indicating the side on which the other-side microphone is located (hereinafter referred to as "other-side microphone presentation information") but also information indicating the side on which the own-side microphone (in this example, first microphone 33a) is located (hereinafter referred to as "own-side microphone presentation information") is displayed. Specifically, in FIG. 4A, the other-side microphone presentation information is displayed in an area on the display screen that is closer to the position of second microphones 34a and 34b than the position of first microphone 33a, and the own-side microphone presentation information is displayed in an area on the display screen that is closer to the position of first microphone 33a than the positions of second microphones 34a and 34b, and further, the other-side microphone presentation information includes the characters "other side" and the own-side microphone presentation information includes the characters "own side."
[0023] Here, when both the remote microphone information and the self microphone information are displayed as shown in FIG. 4A, the remote microphone information can be displayed more emphasized than the self microphone information. Various examples of the emphasis can be considered. Examples include emphasis by color, emphasis by display size such as font size, emphasis by adding a specific figure or pattern, etc.
[0024] 4B shows an example in which only the other party's microphone presentation information is displayed as the direction guidance information Id. Specifically, FIG. 4B shows an example in which the other party's microphone presentation information is displayed with the characters "other party" and an arrow indicating the side where the other party's microphone is located (in other words, an arrow indicating the direction to face the other party).
[0025] Although the above example shows only the information presented by the microphone on the other side being displayed as the direction guidance information Id, it is also possible to display only the information presented by the microphone on the own side as the direction guidance information Id. By directing the local microphone toward the user in accordance with the local microphone presentation information, the other user's microphone is correctly directed toward the other user, and therefore the correct placement orientation of the sound collection device 20 can be guided by displaying the local microphone presentation information.
[0026] For confirmation, in the sound collection device 20, the positional relationship of the first microphone 33a and the second microphones 34a and 34b with respect to the display screen of the display unit 27 remains unchanged, so the display process of the direction guidance information Id can be realized as a process of displaying a predetermined image on the display unit 27. The direction guidance information Id may be displayed, for example, in response to the start of an application for recording the interview (a recording application, which will be described later).
[0027] In FIG. 1, the voice management server 1 and the minutes management server 50 are configured as computer devices equipped with a microcomputer having a CPU (Central Processing Unit), ROM (Read Only Memory), and RAM (Random Access Memory). The sound collection device 20 is capable of performing data communication between the voice management server 1 and the minutes management server 50 via a network NT, which is a predetermined communication network such as the Internet. Furthermore, the voice management server 1 and the minutes management server 50 are also capable of performing data communication with each other via the network NT.
[0028] The minutes-taking support system 100 of the embodiment has a function of analyzing the contents of the interview based on the recorded data of the interview, and a function of presenting text data showing the results of the analysis (hereinafter referred to as "analysis result text data") to the user on a screen. The user can use the analysis result text data presented on the screen to create minutes of the interview.
[0029] As can be understood from the above description, the sound collection device 20 has a function of recording a meeting (a business negotiation in this example). Specifically, the sound collection device 20 records the sound signals collected by the first microphone 33a and the second microphones 34a and 34b, thereby recording the meeting. The recorded data of the interview is transmitted to the voice management server 1 via the network NT, and is stored (accumulated) in the voice management server 1.
[0030] In this embodiment, the function of recording the interview is realized by a "recording application." In this embodiment, the recording application is assumed to be a so-called SaaS (Software as a Service) type in which the software (software program) serving as the recording application is stored on the voice management server 1 side and the software is provided as a service. In this case, various processes of the sound collection device 20 related to recording, such as recording the picked-up signal by the first microphone 33a and the second microphones 34a and 34b, transmitting the recorded data to the sound management server 1, and displaying the aforementioned direction guidance information Id, are performed by the sound collection device 20 under control from the sound management server 1.
[0031] The voice management server 1 has a function of converting the recorded data of an interview (a business negotiation in this example) into text, i.e., a function of generating text data by converting (transcribed or written down) the spoken portion of the recorded data of the interview into text, and an analysis function of performing an analysis process related to the content of the interview based on the text data obtained by the text conversion function (hereinafter referred to as "recorded text data"). Specifically, the analysis function in this embodiment includes functions of performing a summary creation process that creates a summary of the interview, a key point creation process that creates a list of key points of the content (topics) of the interview, and a question and answer identification process that identifies the question and answer portions of the interview.
[0032] Here, the text data of the summary sentences generated by the summary creation process described above, the text data of the key points sentences created by the key points list creation process, and the text data of the questions and answers identified by the question and answer identification process can be said to be data that has been converted into text from the elements of the content of the interview. In this specification, data obtained by converting elements of the content of an interview into text will be referred to as "element text data."
[0033] Furthermore, both such "element text data" and the above-mentioned recorded text data can be said to be text data that indicates the content of speech in an interview, generated based on the recorded data of the interview. In this specification, the text data indicating the content of the utterances made in the interview, which is generated based on the recorded data of the interview, will be referred to as "text data of the utterance content."
[0034] In the minutes support system 100, text data showing the analysis results obtained by the analysis function of the voice management server 1 as described above (hereinafter referred to as "analysis result text data") is displayed on a screen for the sales representative Hs as a user. Specifically, the analysis result text data showing the analysis results as a summary, a list of key points, questions and answers, etc. as described above is displayed on a screen and presented to the user. In this example, the screen display of such analysis result text data is performed on the display unit 27 of the sound collection device 20.
[0035] The user can efficiently create minutes using the analysis result text data of the interview presented as described above. Specifically, the user can transcribe (copy) text data such as summaries presented as the analysis result text data and use them to create minutes, which eliminates the need to create minutes from scratch and improves the efficiency of creating minutes.
[0036] In this embodiment, the analysis result text data is presented by displaying a minutes-creation screen (editing screen Ge, described later) that includes an analysis result display area for displaying the analysis result text data and a minutes-creation area for creating minutes. Here, the minutes-creation area is an area in which the analysis result text data (particularly, data on key points, described later) can be transcribed as constituent sentences of the minutes, allowing sentences serving as minutes to be created.
[0037] Since a minutes creation screen including not only the analysis result display area but also the minutes creation area is displayed, it is possible to eliminate the need to open a separate editing screen when creating minutes using the analysis result text data. Therefore, the effort required for the user to create minutes of an interview can be reduced, and the ease of creating minutes can be improved. In particular, when the interview is a business meeting as in this example, the improved ease of creating minutes can improve the efficiency of business negotiation activities and contribute to an improvement in the success rate.
[0038] <2. Configuration examples of each device> 5 to 7, an example of the configuration of each device constituting the minutes-taking support system 100 according to the embodiment will be described.
[0039] FIG. 5 is a block diagram showing an example of the hardware configuration of the sound collection device 20. As shown in FIG. 5, the sound collection device 20 includes a CPU 21, a ROM 22, and a RAM 23. The CPU 21 functions as an arithmetic processing unit that performs various processes, and executes the various processes in accordance with a program stored in the ROM 22 or a program loaded from the storage unit 29 into the RAM 23. The RAM 23 also stores data and the like required for the CPU 21 to execute the various processes.
[0040] The CPU 21, ROM 22, and RAM 23 are connected to one another via a bus 24. An input / output interface (I / F) 25 is also connected to this bus 24.
[0041] The input / output interface 25 is connected to an input unit 26 that includes an operator and an operation device. For example, various types of operators and operation devices such as a keyboard, a mouse, keys, a dial, a touch panel, a touch pad, a remote controller, etc. are assumed as the input unit 26. In this example, the touch panel in the input unit 26 is formed on the display screen of the display unit 27, and the user can perform touch operations on the display screen. The input unit 26 detects a user operation, and the CPU 21 interprets a signal corresponding to the input operation.
[0042] Furthermore, the input / output interface 25 is connected integrally or separately to a display unit 27 made up of an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) panel, etc., and an audio output unit 28 made up of a speaker, etc. The display unit 27 is used to display various types of information, and is configured as a display device provided on the housing of the sound collection device 20, for example.
[0043] The display unit 27 displays images for various image processing, moving images to be processed, etc. on the display screen based on instructions from the CPU 21. Furthermore, the display unit 27 displays various operation menus, icons, messages, etc., that is, GUI (Graphical User Interface), based on instructions from the CPU 21.
[0044] The input / output interface 25 may be connected to a storage unit 29 and a communication unit 30 . The storage unit 29 is configured with a hard disk drive (HDD) or a solid state drive (SSD), and stores various types of information.
[0045] The communication unit 30 performs communication processing via a transmission path such as the Internet, downloads applications, and performs communication with various devices via wired / wireless communication, bus communication, and the like.
[0046] A drive 32 is also connected to the input / output interface 5 as required, and a removable recording medium 31 such as a memory card or optical disk is appropriately attached thereto.
[0047] The drive 32 allows data files such as programs used in various processes to be read from the removable recording medium 31. The read data files are stored in the storage unit 29, and images and sounds contained in the data files are output on the display unit 27 and the audio output unit 28. Furthermore, the computer programs and the like read from the removable recording medium 31 are installed in the storage unit 29 as needed.
[0048] Furthermore, the input / output interface 25 is connected to a first sound collection unit 33 and a second sound collection unit 34 . The first sound collection unit 33 is a comprehensive representation of the sound collection unit for NC, including the first microphones 33a and 33b mentioned above, and is configured to have at least the first microphones 33a and 33b, an amplifier unit that amplifies the collected sound signals, and an A / D conversion unit that A / D converts the amplified collected sound signals. The second sound collection unit 34 is a comprehensive reference to a stereo sound collection unit including the second microphones 34a and 34b described above, and is configured to have at least the second microphones 34a and 34b, an amplifier unit that amplifies the collected sound signals, and an A / D conversion unit that A / D converts the amplified collected sound signals. The audio signals collected by the first audio collection unit 33 and the second audio collection unit 34 are at least temporarily recorded in a storage device such as the storage unit 29, thereby recording the interview.
[0049] A sensor unit 35 is also connected to the input / output interface 25. The sensor unit 35 collectively indicates sensors provided in the sound collection device 20 serving as a smartphone, and the image sensor provided in the camera unit 36 described above is included in this sensor unit 35. The sensor unit 35 also includes an IMU (Inertial Measurement Unit) equipped with an acceleration sensor, an angular velocity sensor, a direction sensor, etc.
[0050] FIG. 6 is a block diagram showing an example of the hardware configuration of the voice management server 1. As shown in FIG. As shown in FIG. 6, the voice management server 1 includes a CPU 1a, a ROM 2, a RAM 3, a bus 4, an analysis unit 1b, an input / output interface 5, an input unit 6, a display unit 7, a voice output unit 8, a memory unit 9, a communication unit 10, and a drive 12. Here, the CPU 1a, ROM 2, RAM 3, bus 4, input / output interface 5, input unit 6, display unit 7, audio output unit 8, storage unit 9, communication unit 10, and drive 12 are similar to the CPU 21, ROM 22, RAM 23, bus 24, input / output interface 25, input unit 26, display unit 27, audio output unit 28, storage unit 29, communication unit 30, and drive 32 in the sound collection device 20, and therefore redundant explanations will be avoided. Similar to the removable recording medium 31 described above, a removable recording medium 11 such as a memory card or optical disk is appropriately attached to the drive 12.
[0051] In the voice management server 1, the analysis unit 1b is connected to the bus 4. The analysis unit 1b is configured as, for example, a signal processor or an AI (Artificial Intelligence) processor, and performs various analytical processes related to the contents of the interview based on the recorded interview data. Specifically, in this example, it performs at least the summary creation process, the main points creation process, and the question and answer identification process described above. Details of the analysis process executed by the analysis unit 1b, including the summary creation process, the gist list creation process, and the question and answer identification process, will be explained again later.
[0052] Here, the analysis processing by the analysis unit 1b may be executed by the CPU 1a, particularly when the analysis processing is performed as a rule-based processing rather than an AI processing.
[0053] The CPU 1a performs processing operations based on various programs stored in the memory unit 9, ROM 2, etc., thereby executing the information processing and communication processing necessary to realize the processing of the voice management server 1 as an embodiment described below.
[0054] The voice management server 1 is not limited to being configured as a single computer device as shown in Fig. 6, but may be configured as a system of multiple computer devices. The multiple computer devices may be systemized using a LAN (Local Area Network) or the like, or may be located in a remote location using a VPN (Virtual Private Network) using the Internet or the like. The multiple computer devices may include computer devices as a server group (cloud) available through a cloud computing service.
[0055] FIG. 7 is a block diagram showing an example of the hardware configuration of the minutes management server 50. As shown in FIG. In the following description, parts that are the same as parts that have already been described will be given the same reference numerals and description thereof will be omitted.
[0056] The difference from the voice management server 1 shown in FIG. 6 is that the voice management server 1 does not include the analysis unit 1b. Hereinafter, for the sake of convenience, the CPU 1a included in the minutes management server 50 will be referred to as "CPU 50a."
[0057] The minutes management server 50 is not limited to being configured as a single computer device as shown in Fig. 7, but may be configured as a system of multiple computers. In this case, the multiple computers may be configured as a system using a LAN or the like, or may be located in a remote location using a VPN or the like over the Internet or the like. The multiple computers may include computers as a server group (cloud) available through a cloud computing service.
[0058] <3. Flow from interview to minutes creation and example of display screen> As described above, the business negotiation in this example is conducted face-to-face in real space between the sales representative Hs and the customer Hc, and the content of the conversation during the business negotiation is recorded by the sound collection device 20 used by the sales representative Hs. In this example, the sales representative Hs conducts a business negotiation with the above-mentioned recording application running.
[0059] FIG. 8 shows an example of a registration screen Gr for participant information displayed by the recording application. As shown in the figure, the registration screen Gr has input boxes b1, b2, and b3 for entering the company name, name, and position of the customer Hc (customer) in the business negotiation, and an input box b4 for entering the name of the sales representative Hs himself. Input boxes b1, b2, and b3 for customer Hc can be added by operating the add button B1 in the figure, and can accommodate input of information for multiple customers Hc. The registration screen Gr also has a start button B2 for issuing a command to start recording.
[0060] On this registration screen Gr, the sales representative Hs inputs the company name, name, and position of the customer Hc, as well as his or her own name, and operates the start button B2 to cause the sound collection device 20 to start recording the business negotiation.
[0061] Although not shown in the figure, after recording starts, a stop recording button is displayed on the screen of the recording app, and sales representative Hs can operate the stop button to instruct the sound collection device 20 to stop recording the business negotiation (end instruction).
[0062] When an instruction to stop recording the business negotiation is given, the CPU 20a of the sound collection device 20 performs processing to store the recording data of the business negotiation and participant information in the voice management server 1. Specifically, the recording data of the business negotiation, the information on the name and position of the customer Hc entered on the registration screen Gr, and the information on the name of the sales representative Hs are sent to the voice management server 1 and stored therein.
[0063] The voice management server 1 associates the recorded data of the business negotiations sent from the sound collection device 20 with participant information (information on the company name, name, position of the customer Hc, and the name of the sales representative Hs) and stores them in a predetermined storage device such as the memory unit 9. At this time, the voice management server 1 assigns a business negotiation ID, which serves as an identifier for the business negotiation, to each piece of recorded data of the business negotiation.
[0064] In the minutes support system 100 of this embodiment, a list of recording data of business negotiations stored in the voice management server 1 can be presented to a user as a sales representative Hs. Specifically, in this example, it is assumed that the recording application is provided with a function for accessing such a list of recorded data. In other words, in this case, the user as the sales representative Hs can be presented with a recording list screen Gi, which is a list screen of recorded data, by operating an access button (not shown) to the recorded data list displayed on the screen of the recording application.
[0065] FIG. 9 shows an example of the recording list screen Gi. As shown in the figure, the recording list screen Gi displays, for each negotiation ID, the date and time information of the negotiation, a play button Bp, participant information, and an input memo field. In this example, the date and time information of the business negotiation includes the start time (recording start time) and end time (recording end time). The playback button Bp is a button for issuing an instruction to play back the recording data of the business negotiation associated with the business negotiation ID. The participant information displayed is the information entered on the registration screen Gr. If there are any items left unentered on the registration screen Gr, a message to that effect is displayed. The input memo field allows the user to enter text memo information after the negotiation. The information entered in the input memo field is stored in association with the negotiation ID.
[0066] In addition, the recording list screen Gi can also display links to the data of the materials used in the business negotiations (see "Main Materials" in the figure).
[0067] On the recording list screen Gi, a minutes creation button B3 is displayed for each negotiation ID, which instructs the launch of software as a minutes editor. The minutes editor is software that displays the aforementioned editing screen Ge (minutes creation screen), i.e., a screen for creating minutes that includes an analysis result display area and a minutes creation area, and performs processes related to creating minutes. The user as the sales representative Hs operates the minutes creation button B3 displayed for the business negotiation for which the minutes are to be created on the recording list screen Gi.
[0068] In this example, the minutes editor software is stored on the voice management server 1 side, and the minutes editor function is provided to users in the form of SaaS. Therefore, in response to the operation of the minutes creation button B3, the sound collection device 20 requests the voice management server 1 to start the minutes editor service.
[0069] In the above example, the recording application is given the ability to access the recording list screen Gi, but access to the recording list screen Gi can be achieved by specifying the URI (Uniform Resource Identifier) of the recording list screen Gi, and the software used does not matter.
[0070] In addition, in the above example, the recorded data of the business negotiation is sent to the voice management server 1 in response to the operation to stop recording, but the recorded data may also be sent to the voice management server 1 in response to an operation by the sales representative Hs after the business negotiation, separate from the operation to stop recording, or may be sent in real time during the business negotiation.
[0071] As will be described later, in this embodiment, the editing screen Ge displayed by the minutes editor is provided with a button that instructs the execution of analysis processing related to the content of the business negotiation, such as creating the summary described above, and the voice management server 1 executes the analysis processing in response to the operation of the button.
[0072] Fig. 10 is a diagram showing an example of the edit screen Ge. Specifically, the edit screen Ge shown in Fig. 10 is the edit screen Ge before an instruction to execute an analysis process is given. In response to a request from the sound collection device 20 to start the minutes editor service in response to the operation of the minutes creation button B3 described above, the audio management server 1 performs a process of displaying the editing screen Ge shown in Figure 10 on the sound collection device 20. As will be described later, in the editing screen Ge of this example, a display area for analysis result text data (analysis result display area Aa, which will be described later) is displayed on the screen in response to an instruction to execute analysis processing.
[0073] The editing screen Ge in Figure 10 has an utterance text display area At, a recording period display area Ap, a speaker / group setting area Ah, and a text editing area Ae, as well as an input box b5, a play button Bp, copy (transcription) buttons B10 and B11, an analysis execution button B13, a speaker setting instruction button B14, an output button B15, a clipboard copy button B16, an upload button B17, and a post button B18.
[0074] As shown in the figure, in the editing screen Ge of this example, the speech text display area At, the speaker / group setting area Ah, and the text editing area Ae are arranged in the following order from left to right on the screen: speech text display area At, speaker / group setting area Ah, text editing area Ae. In this example, the recording period display area Ap is disposed below the spoken text display area At.
[0075] The speech text display area At is an area in which text data of speeches made in an interview (the above-mentioned recorded text data) is displayed for each speaker. In this embodiment, the process of converting recorded data into text is performed by the voice management server 1. The technology for converting recorded voice data into text, i.e., the technology for converting voice data into text data, is a well-known technology, and detailed description thereof will be omitted here.
[0076] In this example, the voice management server 1 divides the recorded text data into predetermined block units and manages them. The block units here are topic units. In other words, the blocking process of the text data (text block process) in this example is a process of detecting topic boundaries from the text data. Such text block process can be performed using a machine-learned AI (artificial intelligence) model. Alternatively, the text block process can be realized as a rule-based process using keyword matching. For details, see, for example, Japanese Patent Application Laid-Open No. 2023-34235. The text blocking process identifies at least the start time (start time of speech) of each block of text data. By identifying the start time of the block, the position (playback position) corresponding to the start time of the block is also identified for the recorded audio data (audio data).
[0077] Furthermore, the voice management server 1 in this embodiment also performs speaker separation processing based on recorded data. Speaker separation processing is a process for identifying which speaker each speech portion is from. Specifically, for example, a voiceprint analysis process is performed on the recorded data, and a speaker ID is associated with each detected voiceprint. This speaker ID is assigned, for example, as Speaker A, Speaker B, Speaker C, etc. Then, a process is performed to identify the speech sections of each speaker in the recorded data based on the voiceprint information of each speaker. By identifying the speech sections of each speaker in the recorded data in this way, it is possible to identify the correspondence between the speech portions and the speakers in the speech text data obtained by converting the recorded data into text.
[0078] The CPU 1a in the voice management server 1 causes the sound collection device 20 to display the text data of the speech part for each speaker in the speech text display area At based on the text data of each speech part obtained by the speaker separation process described above and information indicating the correspondence between each speech part and the speaker ID.
[0079] In this example, the text data of the speech portions of each speaker is displayed in chronological order from the top to the bottom of the screen. In this example, the speech text display area At displays speaker icons as identifiers for the speech portions of each speaker. At the stage shown in Fig. 10, the speaker icons display speaker IDs assigned in the speaker separation process, such as speakers A, B, and C.
[0080] 10, all of the text data of the utterance portion for each speaker is displayed close to either the left or right edge of the utterance text display area At. Specifically, in this example, the text data is displayed close to the left edge of the utterance text display area At as shown in the figure.
[0081] The recording period display area Ap is an area for displaying recording period display information that graphically represents the recording period of the interview. As shown in the figure, in the recording period display area Ap in this example, information indicating the recording period of the interview in the form of a bar extending in the left-right direction is displayed as recording period display information. In the recording period display area Ap, a seek bar Sb for indicating the playback position of the recorded data is displayed in association with the recording period display information. This seek bar Sb can be moved left and right, and the user can move the seek bar Sb and then operate the play button Bp to instruct the sound collection device 20 to play the recorded data from the playback position indicated by the position of the seek bar Sb.
[0082] In this embodiment, the recording period display information displays information showing the distribution of the speech periods of each speaker. Specifically, the recording period display information in this case displays a bar indicating the recording period for each speaker, and displays the speech period information for each speaker in a different display format for each speaker. For example, it is possible to use different colors or patterns for the speech periods for each speaker. In this example, the bar for each speaker indicating the recording period displays the speaker icon of the corresponding speaker.
[0083] By displaying information showing the distribution of speaking periods for each speaker as recording period display information as described above, the user can easily grasp information such as which speaker spoke for how long during what period during the interview, as well as the amount of speech each speaker made, and can easily grasp which speaker's speech should be focused on when creating minutes. Furthermore, by displaying the information on the speech period for each speaker in a different display format for the recording period display information, the user can intuitively identify the speech period for each speaker. This improves the ease with which a user can search for a specific speech portion of a specific speaker when they want to check the recording data for that portion during the process of creating minutes.
[0084] In addition, in the recording period display area Ap of this embodiment, the negotiation process period information Is is displayed in association with the recording period display information. Here, the negotiation process period information Is is information indicating the name and period of each negotiation process when the negotiation is divided into processes according to a predetermined standard.
[0085] The voice management server 1 in this embodiment performs processing to identify the negotiation process within the negotiation based on the recording data of the negotiation. The negotiation process here refers to the stage at which the negotiation progresses, and examples of such information include "icebreaker (small talk)," "needs analysis," "proposal / explanation," and "negotiation." In this example, the analysis unit 1b in the voice management server 1 performs the process of identifying the negotiation process. Specifically, the analysis unit 1b performs a process of identifying the negotiation process based on text data obtained by converting audio recordings of the negotiations into text. This process can be performed by an AI model generated by machine learning such as deep learning, using the text data as learning input data and the correct answer data for the negotiation process as training data. Alternatively, it can be realized as a rule-based process, in which frequently occurring words for each negotiation process are predetermined as process identification keywords, and the negotiation process is identified based on the results of matching with the process identification keywords.
[0086] The CPU 1a of the voice management server 1 performs processing to display the negotiation process period information Is on the sound collection device 20 based on the period information for each negotiation process (information indicating the start time and end period of the process) obtained as a result of the negotiation process identification processing by the analysis unit 1b as described above. The figure shows an example of the negotiation process period information Is when five processes, ST1 to ST5, are identified as the negotiation process.
[0087] As described above, by displaying the negotiation process period information Is, i.e., the chronological information of the negotiation process, in association with the recording period display information, when a user wants to check the recording data of a specific speech portion of a specific speaker during the process of creating minutes, it is possible to improve the ease of searching for that speech portion.
[0088] Although the recording period display information has been described above as an example in which the recording period is represented by a bar, the recording period may also be represented by another shape such as an arc. When the recording period is represented in an arc shape, it is conceivable to use a rotational operator as the playback position indication operator, rather than a sliding operator such as the seek bar Sb.
[0089] The speaker / group setting area Ah is an area for setting the name of each speaker and the group for each speaker. Group setting here means setting which group each speaker belongs to, with respect to two groups: the self-group, which is the group on the sales representative Hs's (user's) side, and the other-side group, which is the group on the sales representative Hs's other side (i.e., the customer Hc's side).
[0090] In the speaker / group setting area Ah, when group setting has not been performed, a display is made to indicate that each speaker belongs to the other party's group, as shown in the figure. Note that the display format when group setting has not been performed is not limited to this, and other display formats can also be used, such as a display indicating that each speaker belongs to the user's group.
[0091] In the editing screen Ge of this example, when the user wishes to set the speaker's name and group, he or she operates the speaker setting instruction button B14. Then, a setting screen Gs for setting speakers and groups as shown in FIG. 11 is displayed on the speaker / group setting area Ah. As shown in the figure, the setting screen Gs is provided with an input box b6 for inputting the speaker's name, a selection operator for selecting your own group / other party group, a save button B20, and a cancel button B21. Here, an example of the setting screen Gs for speaker B is shown, but in this example, in response to the operation of the speaker setting instruction button B14 described above, the setting screen Gs is displayed in order from speaker A, and the user can input name information in the input box b6 in order from speaker A, set a group by operating the selection operator s1, and perform a save operation using the save button B20, thereby inputting the name and setting a group for each speaker. When the cancel button B21 is operated on the setting screen Gs, the CPU 1a terminates the acceptance of the setting operation of the speaker name and group, and performs processing to return to the screen display state of FIG. 10, for example.
[0092] Here, in this embodiment, since the sound collection device 20 has a self-side microphone and a remote-side microphone, it is possible to determine whether the speaker belongs to the self-side or remote-side group by analyzing which of these microphones collected the speech sound. When performing such group determination processing, it is not necessary to have the user perform group setting operations using the setting screen Gs as described above. In this case, it is sufficient to have the user input and set the name information for each speaker on the setting screen Gs.
[0093] FIG. 12 shows an example of display information in the speaker / group setting area Ah when the name and group for each speaker are set using the setting screen Gs described above. As shown in the figure, the set name information is displayed for each speaker. Also, as shown in Figure 10, in the speaker / group setting area Ah, before the name of each speaker is set, information such as A, B, or C indicating the speaker ID is displayed as the speaker icon, but after the name is set, the initial letter of the name is displayed on the speaker icon, as shown in Figure 12. In addition, in this case, in the speaker / group setting area Ah, according to the group settings for each speaker, the speaker set to the other party's group will be displayed in the other party's group area, and the speaker set to your own group will be displayed in your own group area.
[0094] In this example, the name and group settings for each speaker in the speaker / group setting area Ah are also reflected in the utterance text display area At. Specifically, as in the editing screen Ge shown in FIG. 13, the speaker icon for each speaker in the spoken text display area At is changed to a notation showing the initials of the set name.
[0095] Furthermore, in the utterance text display area At, according to the group setting for each speaker, the utterance text data of a speaker belonging to the own group is moved to one of the left and right edges, and the utterance text data of a speaker belonging to the other group is moved to the other left or right edge. Specifically, in this example, the utterance text data of a speaker belonging to the own group is moved to the right edge of the utterance text display area At, and the utterance text data of a speaker belonging to the other group is moved to the left edge of the utterance text display area At.
[0096] This allows the user to intuitively distinguish between the utterances of the user's group and the other group in the interview text data displayed in the utterance text display area At. In particular, when the interview is a business negotiation as in this embodiment, the user as the sales representative Hs can refer to all the utterances of the members of the other group, the customer Hc, which increases convenience.
[0097] In this example, the name setting information for each speaker in the speaker / group setting area Ah is also reflected in the display information in the recording period display area Ap. Specifically, the display of the speaker icon displayed in each bar for each speaker will be changed from the speaker ID to the initials of the speaker's name.
[0098] Furthermore, in this example, the setting information of the name of each speaker in the speaker / group setting area Ah is also reflected in the text editing area Ae. Specifically, in this example, as shown by "X" in the figure, the name information for each speaker set in the speaker / group setting area Ah (i.e., the name information of the interview participants) is automatically transcribed (copied) into the text editing area Ae. In addition to the setting information in the speaker / group setting area Ah, it is possible to automatically input information about the date and time of the interview into the text editing area Ae as shown in the figure. Furthermore, it is also possible to automatically input text information (for example, "Materials used:" as shown in the figure) to prompt the user to enter the URI where the materials used in the interview (business negotiation) are saved. In this example, the interview information including the name information of each speaker and the information on the date and time of the interview is automatically copied to the top of the text editing area Ae.
[0099] On the editing screen Ge, the analysis execution button B13 is a button for instructing the execution of analysis processes related to the contents of the interview, such as the creation of the summary described above. Specifically, in this example, this button is for instructing the execution of at least the summary creation process, the main points list creation process, and the question and answer identification process described above.
[0100] In this example, the analysis processes of the summary creation process, the main point list creation process, and the question and answer identification process are performed by the analysis unit 1b in the speech management server 1. Therefore, when the analysis execution button B13 is operated, the CPU 1a instructs the analysis unit 1b to execute the analysis processes.
[0101] The CPU 1a displays the analysis result text data obtained as a result of the analysis process by the analysis unit 1b on the editing screen Ge.
[0102] FIG. 14 shows an example of the display of analysis result text data on the editing screen Ge. On the editing screen Ge, the analysis result text data is displayed in the analysis result display area Aa. As shown in the figure, in this example, the analysis result display area Aa is positioned between the speaker / group setting area Ah and the text editing area Ae in the left-right direction of the screen.
[0103] The analysis result display area Aa is provided with a summary display area A1, a main point list display area A2, a question and answer display area A3, and copy buttons B25, B26, and B27. The summary display area A1 is an area for displaying text data as a summary created in the summary creation process.
[0104] The main point list display area A2 is an area for displaying the text data of the main point list created in the main point list creation process, i.e., the text data showing each main point. The main points here refer to the main points discussed in the interview. Therefore, the text data showing each main point can be said to be data that converts the main points discussed in the interview into text. Hereafter, the data that is a text version of the main points discussed in such interviews will be referred to as "main points."
[0105] The question and answer display area A3 is an area for displaying text data of the question portion and text data of the answer portion identified in the question and answer identification process.
[0106] A copy button B25 is provided for the summary display area A1. The copy button B25 is a button for instructing copying of the text data displayed in the summary display area A1 to the text editing area Ae. In the present example, the summary display area A1 allows a range of text data to be specified. After specifying the range of text data to be copied in the summary display area A1, the user can operate the copy button B25 to instruct copying of the text data in the specified range to the text editing area Ae.
[0107] In this example, the gist list display area A2 is capable of displaying gist sentences for each topic category. In this example, multiple categories with different levels of abstraction are defined as topic classifications in interviews, and key points to be displayed in the key point list display area A2 are generated for each category. Specifically, in this example, three levels of abstraction for topic classifications are defined: large, medium, and small (large being the most abstract). In other words, there are three topic classifications: "large category," "medium category," and "small category."
[0108] In this example, the major categories are classified based on the concept of the "negotiation process" mentioned above, and the key points of the major categories would be, for example, any of the above-mentioned "icebreaker," "needs analysis," "proposal / explanation," and "negotiation." The main points of the intermediate category are those with a lower level of abstraction than the main points of the major category. For example, the main points of the intermediate category for the major category = "Icebreaker" could be "golf" or "weather" as specific topics in the sales negotiation step of "Icebreaker", and the main points of the intermediate category for the major category = "Needs analysis" could be "current situation confirmation" or "challenges" as specific topics in the sales negotiation step of "Needs analysis". The main points of the minor category are less abstract than the main points of the medium category. For example, the main points of the minor category for the medium category "golf" are "I played last week" and "I want a new club," while the main points of the minor category for the medium category "weather" are "My head hurts on rainy days" and "If it's sunny next week, I'll go golfing."
[0109] Note that an example of a method for identifying a sales negotiation process as a major classification has already been explained, so a duplicate explanation will be avoided. A specific example of a method for generating gist sentences for medium and small categories will be described later.
[0110] 14 shows the initial display state of the analysis result display area Aa, which is displayed in response to analysis processing of the converted text data. The initial display state here refers to the state at the stage when display starts. In this embodiment, in the gist list display area A2 in the initial display state, gist sentences of categories with a certain level of abstraction or higher are displayed, and gist sentences of categories with a level of abstraction lower than the certain level are hidden. Specifically, in this example, as shown in the figure, only gist sentences of major and medium categories are displayed, and gist sentences of minor categories are hidden. If the key points for all categories were displayed from the beginning, the display content in the key points list display area would become cluttered, making it difficult for the user to grasp the outline of the key points of the interview. By displaying only the key points for categories with a certain level of abstraction or higher as described above and hiding the key points for the remaining categories, it is possible to prevent the display content in the key points list display area from becoming cluttered, making it easier for the user to grasp the outline of the key points of the interview.
[0111] In this example, where the service of analyzing text data and supporting the creation of minutes using the editing screen Ge is provided in a SaaS format, display control of the key points list display area A2 as described above is primarily performed by the CPU 1a of the voice management server 1.
[0112] Although an example in which the number of categories of gist sentences is three is given here, the number of categories may be, for example, two (major categories and minor categories), or may be four or more categories.
[0113] In the gist list display area A2 in the initial display state, a check box cb for selecting a gist sentence is displayed for each of the displayed gist sentences of the major and medium categories.
[0114] In the gist list display area A2, an expandable box ex is provided for the display area of each subcategory under each medium category. The user can select an expandable box ex under any medium category to display the gist sentence of the subcategory under that medium category. Figure 15 shows an example of a case where an expandable box ex located under the medium category "Golf" in the major category "Icebreaker" is selected (Figure 15A), and the key points of the subcategory under "Golf" are displayed (Figure 15B). Furthermore, Figure 16 illustrates a case where, in response to the selection of the expandable box ex located under the intermediate category "Weather" (Figure 16A), the main points of the subcategory under "Weather" are displayed (Figure 16B). As shown in FIGS. 15 and 16, check boxes cb are displayed for the main points of the displayed small categories as well as for the main points of the large and medium categories. When the expandable box ex is expanded to display the main points of the subcategory as shown in Figures 15B and 16B, a close button (marked with an x in the figure) is displayed as shown, making it possible to return the main points of the subcategory to a hidden state.
[0115] In the gist list display area A2 of this example, the gist sentences are arranged in the chronological order of the original utterances. Specifically, in the gist list display area A2 of this example, the older the original utterance, the higher the gist sentence is placed in the area.
[0116] In addition, in the gist list display area A2 in this example, when an operation to select a gist sentence is performed, the audio data of the utterance portion corresponding to the selected gist sentence is played back. FIG. 17 shows an example in which the main sentence of the minor category "played last week" is selected. In the gist list display area A2 in this example, it is possible to perform an operation to select each gist sentence in this way, and when the user performs an operation to select one of the gist sentences, the CPU 1a performs a process to play back the audio data of the utterance portion in the recorded data that corresponds to the selected gist sentence.
[0117] In this way, by being able to play back the audio data of the utterance portion corresponding to the gist sentence on the editing screen Ge, when a user wants to clarify the content of a gist sentence whose details are unclear when creating minutes, the user does not need to search for the corresponding audio data portion by himself / herself, thereby improving the efficiency of creating minutes.
[0118] Here, in the minutes support system 100 of this embodiment, in order to enable smooth playback of audio data of the speech points corresponding to such key point sentences, the audio management server 1 performs a pre-link assignment process in which, before the key point sentences are displayed in the key point list display area A2, the audio management server 1 assigns an "audio link" to the audio data, specifically, an "audio link" which is a link to the audio data portion of the speech point corresponding to the key point sentence. This pre-linking process will be explained later.
[0119] In FIG. 14, a copy button B26 is provided for the main point list display area A2. The user can check any of the check boxes cb for each gist sentence and then operate the copy button B26 to instruct copying of the gist sentence corresponding to the checked check box cb into the text editing area Ae.
[0120] A specific example will be described with reference to Figures 18 and 19. For convenience of illustration, Figures 18 and 19 mainly show only the analysis result display area Aa and the text editing area Ae of the editing screen Ge. 18 shows a state in which the main point sentences of the intermediate category "golf" in the main point list display area A2 are selected as the target for copying. When the copy button B26 is operated in this state, at least the selected main point sentences of "golf" are copied to the text editing area Ae, as shown in FIG. In this example, if the key sentence selected to be copied is a key sentence from a medium category, as shown in the figure, the key sentences from the large category corresponding to the key sentence from the medium category are copied, as well as the key sentences from all the small categories corresponding to the key sentence from the medium category. Although not illustrated, when a main sentence of a major category is selected as the copy target, the main sentences of all the medium and small categories under that major category are copied along with that main sentence. For example, when a main sentence of the major category "Icebreaker" is selected, the main sentence of that major category, "Icebreaker," the main sentence of the medium category "Golf" under that major category, the main sentences of the small category under that subordinate category, "I played last week" and "I want a new club," the main sentence of the other medium category under that major category, "Weather," and the main sentences of the small category under that subordinate category, "Rainy days give me a headache" and "If it's sunny next week, I'll go golfing," are copied into the text editing area Ae. Furthermore, when a main point sentence of a small category is selected as the copy target, the main point sentence of the medium category and the large category corresponding to the main point sentence of the small category are copied together with the main point sentence of the small category. For example, when the main point sentence of the small category = "If it's sunny next week, I'll go golfing" is selected, the main point sentence of the small category, "If it's sunny next week, I'll go golfing", the main point sentence of the medium category corresponding to the small category, "Weather", and the main point sentence of the large category corresponding to the small category, "Icebreaker", are copied into the text editing area Ae.
[0121] As mentioned above, in this example, the interview information is automatically copied to the top of the text editing area Ae, but the text data copied from the analysis result display area Aa, including the key points list display area A2, is placed below the event information. At this time, it is conceivable to insert title information indicating that this is the main subject of the minutes, such as the text "Minutes Contents" shown in the figure, between the event information and the copied text data.
[0122] In FIG. 14, text data showing the question portion and text data showing the answer portion are displayed in question and answer display area A3. In the question and answer display area A3 of this example, text data showing the question portion and text data showing the answer portion are arranged in the chronological order of the original utterances.
[0123] The copy button B27 is a button for instructing copying of the text data in the question and answer display area A3 to the text editing area Ae. In the question and answer display area A3, it is possible to perform an operation to specify the text data of the question and answer portions to be copied. After specifying the range of text data of the question and answer portions to be copied in the question and answer display area A3, the user can issue an instruction to copy the specified text data to the text editing area Ae by operating the copy button B27.
[0124] In the question and answer display area A3 of this example, the text data showing the question portion and the text data showing the answer portion are displayed so that it is possible to identify whether the comment came from the user's group or the other party's group. Specifically, in this example, the text data for the question and answer sections of your group is displayed closer to the right edge of the question and answer display area A3, and the text data for the question and answer sections of the other group is displayed closer to the left edge of the question and answer display area A3. This makes it easier for a user to understand specific questions and answers from his or her own or the other party when creating minutes of a meeting.
[0125] The analysis result display area Aa includes a summary display area A1, a main points display area A2, and a question and answer display area A3. In this embodiment, the analysis result display area Aa has the summary display area A1 positioned at the top of these display areas. That is, the CPU 1a performs display processing on the analysis result display area Aa so that the text data representing the summary is positioned at the top of the display area.
[0126] Here, "upper" means the direction toward the top of the screen when the utterance text display area At, the analysis result display area Aa, and the text editing area Ae are arranged horizontally on the screen, and means the direction toward the left of the screen when the utterance text display area At, the analysis result display area Aa, and the text editing area Ae are arranged vertically on the screen. Information displayed on the screen tends to be viewed from the top down. Therefore, by placing the "Summary" at the topmost part of the screen as described above, it is possible to have the user check the contents of the "Summary" before referring to the "Key Points" and "Questions and Answers." In other words, it is possible to have the user grasp the overall outline of the interview from the "Summary" before referring to the "Key Points" and "Questions and Answers." As a result, it is possible to make it easier for the user to understand the contents of the "Key Points" and "Questions and Answers." Therefore, the time required for the user to understand the contents of the interview can be reduced, and the efficiency of creating minutes can be improved.
[0127] In the analysis result display area Aa, the arrangement order of the summary display area A1, the main points list display area A2, and the question and answer display area A3 may be changeable based on a user operation (customizable by the user). On the other hand, it is also possible to make it impossible for the user to customize the arrangement order of the utterance text display area At, the analysis result display area Aa, and the text editing area Ae on the editing screen Ge.
[0128] Next, the text editing area Ae as the minutes creation area will be described. The text editing area Ae is an area where the analysis result text data in the analysis result display area Aa, including the main points, can be transcribed as constituent sentences of the minutes, allowing for the creation of minutes. In this text editing area Ae, the copied analysis result text data can be edited, for example, by deleting or adding parts. Since the analysis result text data in the analysis result display area Aa can be copied and used as key points or summaries in the sentences that make up the minutes, there is no need to input the text for creating the minutes from scratch, which improves the efficiency of creating the minutes.
[0129] In the text editing area Ae of this embodiment, if the transcribed analysis result text data is corrected, the corrected text portion is displayed in a different display mode from the uncorrected text portion. Although not shown, if the analysis result text data copied from the analysis result display area Aa is corrected in the text editing area Ae by a user operation, the corrected text portion is displayed in a different color from the uncorrected text portion, for example. Alternatively, only the corrected text portion may be displayed in bold or underlined.
[0130] In addition, in this embodiment, when the transcribed analysis result text data is corrected in the text editing area Ae, the CPU 1a performs a process of recording the text data before correction and the text data after correction in correspondence with each other in a memory device. The text data before and after the correction is recorded in a memory device within the voice management server 1, such as the storage unit 9 of the voice management server 1. Alternatively, the text data before and after the correction may be recorded in a memory device in the minutes management server 50 or the sound collection device 20.
[0131] Here, the text data before and after correction in the text editing area Ae can be used to train minutes writers. Specifically, for example, by having a low performer who is unfamiliar with minute writing refer to the text data before and after correction by a high performer who is skilled in minute writing, it becomes possible for the low performer to learn the skills related to minute writing, which can contribute to the training of low performer writers. In particular, if the interview is a business negotiation, it becomes possible for the low performer to learn the skills of a high performer sales representative, which can contribute to the improvement of the sales representative's skills. Furthermore, the above-described pre- and post-correction text data can also be used to improve and enhance the analytical process for obtaining analysis result text data from interview recording data. Specifically, for example, when analytical processing is performed using an AI model, the pre- and post-correction text data can be used as learning data for machine learning to improve and enhance the analytical process. Alternatively, even when analytical processing is performed using rule-based processing, the pre- and post-correction text data can be referenced when updating the analytical process, thereby contributing to the improvement and enhancement of the analytical process.
[0132] In addition, there are other benefits to saving the text data before and after the corrections described above, such as the following: For example, if a sales representative Hs1 modifies the string "likes golf" as part of a summary in the minutes to "President xx likes golf," while another sales representative Hs2 modifies the same summary by deleting the phrase "likes golf," then sales representative Hs1 will be recognized as a high performer and sales representative Hs2 as a low performer based on past sales negotiation results, etc. Additionally, we envision an AI that automatically creates meeting minutes. In this case, by saving the text data before and after the correction, the AI can learn the correction tendency of sales representative Hs1 (always recording the customer's hobbies in the minutes, and adding a subject so that it is clear whose hobby it is). Specifically, it is possible to enable the AI to acquire the function of including the name of the customer's hobbies in the minutes.
[0133] On the editing screen Ge, an output button B15 and a clipboard copy button B16 are provided as operators related to the text editing area Ae. The output button B15 is a button for instructing that the text created in the text editing area Ae (text as minutes) be output in a predetermined file format such as Word. Specifically, it is a button for instructing that the text created in the text editing area Ae be saved in the sound collection device 20 (i.e., locally) in a predetermined file format. The clipboard copy button B16 is a button for instructing copying of the text created in the text editing area Ae to the clipboard of the sound collection device 20. By providing the output button B15 and clipboard copy button B16, the user can save the created text data as minutes in a predetermined file format in the sound collection device 20. Therefore, the output button B15 and clipboard copy button B16 can be said to be buttons for instructing local saving of the created minutes.
[0134] Furthermore, in the editing screen Ge of this example, an upload button B17 and a post button B18 are also provided as operators related to the text editing area Ae. The upload button B17 is a button for instructing uploading of the text created in the text editing area Ae to a cloud server of a predetermined sales information management service (for example, Salesforce: registered trademark). Specifically, the upload button B17 is a button for instructing uploading of the text created in the text editing area Ae to the minutes management server 50. For example, when the upload button B17 is operated, a dialog is displayed, and when the user specifies their company name on the dialog and performs an operation to instruct the upload to be performed, the text created in the text editing area Ae is uploaded to the minutes management server 50.
[0135] The post button B18 is a button for instructing that the text created in the text editing area Ae be posted as shared information in a predetermined sales communication service (e.g., Slack: a registered trademark). For example, when the post button B18 is operated, a dialog box is displayed. When the user specifies a channel for information sharing on the dialog box and issues a posting instruction, the text created in the text editing area Ae is posted as shared information in the specified channel in the sales communication service.
[0136] In addition, in the editing screen Ge of this example, a copy button B10 and a copy button B11 are also provided as operators for issuing a copy instruction to the text editing area Ae. The copy button B10 is a button for instructing copying of the text data in the utterance text display area At to the text editing area Ae. The copy button B11 is a button for instructing copying of the speaker name information and group setting information set in the speaker / group setting area Ah to the text editing area Ae. By providing these copy buttons, the efficiency of creating minutes can be further improved.
[0137] Here, in this embodiment, when an instruction to save the minutes data is given, such as by operating the upload button B17, the audio management server 1 (CPU 1a) performs a process (hereinafter referred to as the "minutes data generation process") to generate minutes data in which the above-mentioned audio link is attached to the main points as the minutes data. By generating minutes data in this manner in which audio links are attached to key sentences, when a key sentence is specified in the minutes, the audio data portion indicated by the audio link can be played back, making it possible to smoothly play back the audio data of the utterance corresponding to the specified key sentence. Therefore, when realizing audio playback of the speech corresponding to the key points in the minutes, it is no longer necessary for the user to search for the relevant audio data portion from the recorded data, which makes it easier to check the details of the interview in the minutes after the fact.
[0138] In this embodiment, the audio link is associated with data in text format that indicates a link to the corresponding audio data portion. Here, some minutes management servers 50 that store minutes data including gist sentences do not support hyperlinks. By providing text data for the audio links associated with the gist sentences as described above, even when minutes data is stored on such a server that does not support hyperlinks, the user can play back the audio data of the corresponding spoken portion using the audio link information in text data. In other words, even when minutes data is stored on a minutes management server 50 that does not support hyperlinks, the user does not need to search for the relevant audio data portion, making it easier to confirm the details of the interview after the fact. For example, a user can copy an audio link in the form of text data and enter it into a search bar in a web browser to access the corresponding audio data and play the audio data.
[0139] As mentioned above, in this example, before the gist sentences are displayed in the gist list display area A2, the pre-linking process is performed to add audio links to the gist sentences. Therefore, when generating the minutes data, the audio links added by the pre-linking process are inherited. Since it is only necessary to take over the audio link that has already been assigned, the minutes data can be generated quickly.
[0140] Furthermore, in this embodiment, when an instruction to save the minutes data (cloud save instruction) is given to a specified server device other than the device displaying the editing screen Ge (minutes creation screen), such as by using the upload button B17 described above, as an instruction to digitize the minutes data, the audio management server 1 generates minutes data in a format in which an audio link is attached to the main points. On the other hand, when an instruction to save the minutes data (local save instruction) is given to the device displaying the editing screen Ge, such as by operating the output button B15 described above, the audio management server 1 generates minutes data in a format in which an audio link is not attached to the main points. Specifically, in this example, the above-mentioned cloud save instruction corresponds to an upload instruction operation on the dialog that is displayed in response to the operation of the upload button B17, and a post instruction operation on the dialog that is displayed in response to the operation of the post button B18. In this example, the local save instruction corresponds to the operation of the output button B15 and the operation of the clipboard copy button B16.
[0141] By switching between adding and not adding audio links as described above, if the user does not think that the audio link function is necessary for the minutes, they can simply instruct the created minutes to be saved locally.Since the user can choose whether to create minutes data with or without an audio link, it is possible to prevent situations where the user is only able to create and save minutes data with an audio link even though they do not think that the audio link function is necessary.
[0142] Here, in this embodiment, it is possible to correct the key points copied to the text editing area Ae, but in this example, the audio management server 1 will not release the audio link when generating the minutes data, even if the transcribed key points are corrected, as long as even one character from before the correction remains.
[0143] FIG. 20 is an explanatory diagram of uploading of minutes data with an audio link and processing related to the uploaded minutes data. First, the minutes data created through the editing screen Ge is <1> As shown in the figure, the minutes are uploaded to the minutes management server 50. In this example, this upload is performed based on the operation of the upload button B17.
[0144] The minutes data uploaded to the minutes management server 50 can be viewed by the sales representative Hs as a user (see in the figure). <2> Here, the minutes data is viewed using the sound collection device 20 (i.e., the device that created the minutes), but the minutes data can also be viewed using a device other than the sound collection device 20.
[0145] While viewing the minutes data, the user performs a voice data access operation using the voice link text data (see the figure). <3> Specifically, you can copy the text data as audio links associated with the key points in the minutes, enter it into the search bar of a web browser, and perform a search.
[0146] FIG. 21 is a diagram for explaining an example of an operation for playing back audio while viewing the minutes. As shown in Figure 21A, transcribed key sentences can be selected in the minutes. Figure 21A shows an example in which the key sentence "The customer played last week," which is a subcategory under "golf" in the minutes, is selected.
[0147] When an operation to select a gist sentence in the minutes is performed, a dialog box showing information on processes that can be executed on the selected gist sentence is displayed, as shown in Fig. 21B. Specifically, this dialog box contains information on the audio link associated with the selected gist sentence (in this example, this is URI information indicating the location of the audio data corresponding to the gist sentence), and an edit button that instructs transition to edit mode for the selected gist sentence. The dialog also has a back button, and the user can close the dialog by operating the back button.
[0148] If the user wishes to hear the audio corresponding to the selected gist sentence, the user performs an operation to select the audio link in the dialog. Then, within the dialog, another dialog like the one shown in Figure 22 is displayed. This separate dialog displays a copy button (the "Copy" part in the figure) for instructing to copy the audio link, and a back button. By operating the copy button, the user can instruct to copy the text data as an audio link. Then, using the copied audio link, the user can access the audio data corresponding to the selected gist sentence (in this example, by entering it into the search bar of the web browser and executing a search). The user can close the separate dialog box by operating the back button.
[0149] In response to the above-described voice data access operation, an access request is generated from the sound collection device 20 (the device displaying the minutes) to the voice management server 1 (see FIG. <4> (See
[0150] In this example, when there is a request to access audio data using an audio link, the audio management server 1 performs access authentication processing (see in the figure). <5> ). In this example, it is assumed that a contract has been made in advance between the administrator of the voice management server 1 and the user (or the company to which the user belongs) to enjoy the service of voice playback of the part corresponding to the gist sentence. In this access authentication process, at least a process of determining whether the user who has made the access is a legitimate user who has signed the contract is performed. For example, in the access authentication process, the voice management server 1 requests the sound collection device 20 to transmit the user account and password information specified at the time of the contract, and determines whether the transmitted information matches the user account and password information registered for the legitimate user. If the information matches, the authentication is successful; if it does not match, an authentication error occurs.
[0151] If the authentication is successful, the voice management server 1 allows the device that made the access request to access the voice data (recorded data). In the event of an authentication error, the voice management server 1 denies access to the voice data by the device that made the access request, and performs appropriate error processing, such as displaying a message on the device indicating that access is not permitted.
[0152] <4. Example of functional configuration> With reference to FIG. 23, a functional configuration for realizing the information processing method according to the embodiment described above will be described. Figure 23A is a functional block diagram showing various functions as an embodiment possessed by the CPU 1a in the voice management server 1, and Figure 23B is a functional block diagram showing various functions as an embodiment possessed by the analysis unit 1b in the voice management server 1.
[0153] As shown in FIG. 23A, the CPU 1a has functions as an audio recording processing unit F11, a screen control processing unit F12, a minutes generation processing unit F13, a transmission processing unit F14, and an access authentication processing unit F15.
[0154] The audio recording processing unit F11 performs a process of recording the audio data of the interview. Specifically, this is a process of storing the audio data of the interview (audio data detected by the microphone) transmitted from the audio collection device 20 in the storage unit 9 of the audio management server 1.
[0155] The screen control processing unit F12 performs control processing for displaying an editing screen Ge as a minutes creation screen on the sound collection device 20, and various processing in response to operations on the editing screen Ge. In particular, the screen control processing unit F12 in this embodiment performs processing such as displaying key sentences of categories with a certain level of abstraction or higher and hiding key sentences of categories with a level of abstraction lower than the certain level in the initial display state of the key point list display area A2, as described above in Figure 14.
[0156] The minutes generation processing unit F13 generates minutes data including a gist sentence. Specifically, the minutes generation processing unit F13 in this embodiment generates minutes data including a gist sentence associated with an audio link. A specific example of the association of audio links with key points in the minutes data will be explained later.
[0157] The transmission processing unit F14 performs processing to transmit the minutes data generated by the minutes generation processing unit F13 to an external server device.
[0158] As can be understood from the above explanation, the minutes generation processing unit F13 in this example generates minutes data with an audio link when a cloud storage instruction is given, and generates minutes data without an audio link when a local storage instruction is given. Furthermore, when a cloud storage instruction is given, the transmission processing unit F14 performs processing to transmit the minutes data with the audio link generated by the minutes generation processing unit F13 to a server device corresponding to the operation. Specifically, when an upload execution instruction operation is given on the dialog displayed in response to the operation of the upload button B17, the transmission processing unit F14 transmits the minutes data with the audio link to the minutes management server 50. On the other hand, when a post execution instruction operation is given on the dialog displayed in response to the operation of the post button B18, the transmission processing unit F14 transmits the minutes data with the audio link to a predetermined server device other than the minutes management server 50.
[0159] The access authentication processing unit F15 performs the above-mentioned access authentication processing when an external device requests access to the recorded data based on the audio link in the minutes data. Note that a specific example of the access authentication processing has already been explained, so a duplicate explanation will be avoided.
[0160] In FIG. 23B, the analysis unit 1b has functions as a text blocking processing unit F21, a summary creation processing unit F22, a gist list creation processing unit F23, and a question and answer identification processing unit F24.
[0161] The text blocking processor F21 performs the above-mentioned text blocking process on the recorded-text data, which is the data obtained by converting the recorded data of the interview into text. Specifically, the text blocking process in this example is a process of detecting topic breaks from the recorded-text data. As mentioned above, this text blocking process can be realized by a machine-learned AI model or rule-based processing using keyword matching. The text blocking process identifies at least the start time (start time of speech) of each block of text data. By identifying the start time of the block, the position (playback position) corresponding to the start time of the block is also identified for the recorded audio data (audio data).
[0162] The summary creation processing unit F22 performs a summary creation process, which is a process for creating a summary of the interview, based on the audio-recorded text data. The main point list creation processing unit F23 performs a main point list creation process, which is a process for creating a list of main points of the interview, based on the audio-text data. The question and answer identification processor F24 performs a question and answer identification process, which is a process for identifying the question portion and the answer portion in the interview, based on the audio-recorded text data.
[0163] Here, it is conceivable that the summary creation process and the question and answer identification process can each be realized as processes using an AI model. For example, for the summary creation process, it is conceivable to use an AI model trained by machine learning using deep learning or the like so that text data can be used as input data and a summary can be obtained as output data. Similarly, for the question and answer identification process, it is conceivable to use an AI model trained by machine learning using deep learning or the like so that text data can be used as input data and a question portion and an answer portion can be obtained as output data. Alternatively, the summary creation process and question and answer identification process may be realized as rule-based processes. For example, the summarization process can be realized by identifying and deleting unnecessary sentences in spoken language, while the question and answer identification process can be realized by identifying the question portion of the utterance based on the recorded data, assuming that the ending of the question portion rises, and identifying the portion of the utterance following the question portion as the answer portion.
[0164] Here, in this embodiment, the gist list creation processing unit F23 generates gist sentences for each of the aforementioned categories (large, medium, and small categories in this example) to create a gist list, and also performs the aforementioned "pre-link assignment process" for each of the generated gist sentences. It is possible that the process of generating the gist sentences and the association of the voice links with the gist sentences may be performed automatically after the business negotiations are completed, for example.
[0165] Referring to FIG. 24, the process of generating gist sentences for each category as an embodiment performed by the gist list creation processor F23 and the audio links associated with the generated gist sentences will be described.
[0166] As described above, in this example, each negotiation process identified in the negotiation process identification process is assigned to the main point sentences of the major classification.
[0167] In this example, the main points of the middle category are generated based on the results of keyword matching using keywords defined corresponding to the topics of the middle category (hereinafter referred to as "middle category keywords"). In this example, a plurality of middle category keywords are defined for each topic of the major category, i.e., for each negotiation process in this example, and the main points of the middle category are generated based on the results of keyword matching using middle category keywords corresponding to each section of the topic of the major category in the text data (i.e., for each negotiation process period in this example). For example, in generating the main points of the middle category in the topic section of the major category = "icebreaker," keywords such as "golf" and "weather" are defined as the middle category keywords corresponding to "icebreaker," and if there is a section in which the word "golf" is detected in the text data for the negotiation process period of "icebreaker," a main point sentence of "golf" is generated as the main point sentence of the middle category under "icebreaker." Furthermore, if the word "weather" is detected in a section other than the section in which the word "golf" was detected in the text data of the "icebreaker" section, a further gist sentence of "weather" is generated as a gist sentence of the intermediate category under "icebreaker." For example, by using this method, it is possible to generate gist sentences for each major category and its subordinate medium categories.
[0168] In this case, the boundaries between intermediate categories, i.e., the start and end of each intermediate category period, follow the boundaries of the blocks detected in the text block segmentation process. Within the sales negotiation process period of each major category, intermediate categories are identified by keyword matching, starting from the first block of the sales negotiation process period. A block in which a word matching a pre-defined keyword is detected (this keyword is designated as the first keyword) is identified as a block of the intermediate category to which the first keyword should be applied (designated as the first block). If a word matching another keyword (designated as the second keyword) is detected in a subsequent block, this block is identified as a block of the intermediate category to which the second keyword should be applied (designated as the second block). In this case, if there is a block between the first block and the second block in which no keyword is detected (designated as the non-detection block), the period from the first block to the non-detection block (or, if there are multiple non-detection blocks, the period up to the last non-detection block) is identified as the period of the intermediate category to which the first keyword should be applied. This process of identifying the period of the medium classification is performed for each major classification.
[0169] The main point sentences in the small category are generated with a lower level of abstraction of the topic than the main point sentences in the medium category. Specifically, in this example, the gist sentences for the small categories are generated by summarizing the text for each block detected in the text block division process. In other words, in this example, the gist sentences for the small categories can be said to be summaries for each block. The generation of the summaries here can be performed using the same method as the generation of summaries in the summary creation process described above.
[0170] In this example, the audio links are associated with each small group of gist sentences as shown in the figure. In other words, in this example, the voice management server 1 generates voice data files for each block of the interview recording data received from the sound collection device 20, and stores these voice data files for each block, and the voice links are generated as URI information indicating the location information of each of these voice data files for each block. Then, for each gist sentence of each minor category, text data of a URI indicating the location of the audio data file of the block from which the gist sentence was generated is associated as an audio link.
[0171] In addition, while Figure 15 and other figures above give two examples of key sentences for the minor category under the medium category = "golf", namely "I played last week" and "I want a new club", Figure 24 gives an example of a case where there are three blocks in the period of the medium category = "golf", i.e., three key sentences should be generated for the minor category under the medium category = "golf", but Figure 24 ignores consistency with Figure 15 and other figures above.
[0172] In the above, it is assumed that one small category is made up of one block, but it is also possible to consider that one small category is made up of a predetermined number of blocks. In this case, the audio link attached to the gist sentence of the small category can be text data indicating the URI of the audio data file of the corresponding predetermined number of blocks.
[0173] In Figure 23B, in generating key point sentences, the key point list creation processing unit F23 in this embodiment generates key point sentences for categories with a certain level of abstraction or above based on the speech content of both the one's own group and the other group, and generates key point sentences for categories with a level of abstraction below a certain level based only on the speech of the one's own group or the other group. Specifically, in this example, the gist list creation processing unit F23 generates gist sentences for the major categories based on the utterances of both the own group and the other group, while generating gist sentences for the medium and small categories based only on the utterances of the other group. For clarity, the group to which each utterance belongs can be identified for each utterance in the converted data by performing the speaker separation process described above and setting the group for each speaker in the speaker / group setting area Ah.
[0174] For gist sentences, in order to properly classify high-level abstractions, it is advantageous to have a large number of utterances as a basis, and it is desirable to use the utterances of both the self-group and the other group when generating gist sentences for high-level abstractions, as described above. On the other hand, for gist sentences for low-level abstractions, if they are to be recorded in the minutes, the content of the other group's utterances becomes important (especially when considering business negotiations, medical interviews, educational interviews, etc.), so it is desirable to use only the utterances of the other group when generating gist sentences for low-level abstractions, as described above. Therefore, by selecting the utterance to be used as the basis for generating the gist sentence from either the user's own utterance or the other person's utterance depending on the level of abstraction of the classification as described above, the utterance to be used to generate the gist sentence can be appropriately selected depending on the level of abstraction of the classification of the gist sentence to be generated, thereby ensuring that the gist sentence is generated appropriately.
[0175] As can be understood from the above explanation, in this example, audio links are associated with the key point sentences through a pre-link assignment process by the key point list creation processing unit F23, and then when the minutes data including the key point sentences is sent (in this example, for cloud storage), the minutes generation processing unit F13 generates the minutes data including the key point sentences to which audio links are associated. In this case, the correspondence between the audio links and the gist sentences in the minutes data can be realized by generating data in a format in which text data of URIs as audio links are added to each gist sentence (in this example, subcategory gist sentences) as shown in Figure 25. The example in Figure 25 illustrates minutes data in the form of text data in a format in which text as audio links is added following the text of the gist sentence for each gist sentence.
[0176] As explained above in Figures 21 and 22, in this example, when a user views the minutes data stored in the minutes management server 50, the text data of the audio links included in the minutes data as described above is hidden and is displayed in response to user operations (such as selecting the target gist sentence). However, instead of this, the minutes data including the text data of the audio links as shown in Figure 25 can be displayed as is, and the user can copy the text data of the audio links displayed in this way.
[0177] Here, the voice management server 1 in this embodiment has a function of converting recorded interview data into text, and in this example, this text conversion function is a function possessed by the analysis unit 1b. Furthermore, the voice management server 1 in this embodiment performs the above-mentioned speaker separation process, and the speaker separation process is also performed by the analysis unit 1b.
[0178] <5. Processing Procedure> An example of a specific processing procedure to be executed to realize the information processing method according to the embodiment will be described with reference to the flowcharts of FIGS. In this example, the processes shown in FIGS. 26 and 27 are executed by the CPU 1a in the voice management server 1 based on a program stored in the ROM 2 or the storage unit 9, for example.
[0179] FIG. 26 is a flowchart of the process for generating the minutes data. First, in step S101, the CPU 1a waits for an instruction to save the minutes data. That is, in this example, the CPU 1a waits for any of the following: an instruction to save the minutes data from the Output button B15, an instruction to copy to clipboard B16, an instruction to upload the minutes data on the dialog box displayed in response to an operation of the Upload button B17, and an instruction to post the minutes data on the dialog box displayed in response to an operation of the Post button B18.
[0180] If an operation to instruct the minutes data to be digitized has been performed, the CPU 1a proceeds to step S102 and determines whether the operation is a cloud storage system. That is, it determines whether the operation detected in step S101 is an upload execution instruction operation on a dialog box displayed in response to the operation of the upload button B17, or a post execution instruction operation on a dialog box displayed in response to the operation of the post button B18.
[0181] If the cloud storage system is selected in step S102, the CPU 1a proceeds to step S103 and generates minutes data that retains the audio link status of each gist sentence. That is, as shown in Fig. 25, the CPU 1a generates minutes data in which audio link text data is associated with each gist sentence.
[0182] Then, in step S104 following step S103, the CPU 1a performs upload processing of the generated minutes data. That is, if the operation detected in step S101 is an operation to instruct upload execution on the dialog displayed in response to operation of the upload button B17, the CPU 1a transmits the generated minutes data to the minutes management server 50, and if the operation detected in step S101 is an operation to instruct posting execution on the dialog displayed in response to operation of the post button B18, the CPU 1a transmits the generated minutes data to a predetermined server device as the posting destination.
[0183] On the other hand, if the operation is not a cloud storage operation in step S103, the CPU 1a proceeds to step S105 and generates minutes data in which the audio link status for the gist sentences of each category is released. For example, the CPU 1a generates minutes data that inherits only the text data of the gist sentences.
[0184] In step S106 following step S105, the CPU 1a performs processing to transmit the generated minutes data to the sound collection device 20. This allows the sound collection device 20 to output the text data of the minutes to Word or copy it to the clipboard.
[0185] The CPU 1a ends the series of processes shown in FIG. 26 in response to execution of the process of either step S104 or S106.
[0186] The minutes editor service may be provided by software stored locally in the sound collection device 20. 26 is executed by the CPU 20a of the sound collection device 20. At this time, it goes without saying that instead of the transmission process of step S106, a process of outputting the text data of the minutes to Word or copying it to the clipboard should be executed.
[0187] In addition, although the above example shows that when saving minutes data, the format of the minutes data can be selected from either a format in which audio link text data is associated with the text or a format in which no audio links are associated with the text, it is also possible to select from three formats for the format of the minutes data to be saved: these two formats plus a format in which audio links are associated as hyperlinks. This makes it possible to also handle cases where minutes data is saved in a server device that supports hyperlinks. It is also possible to select between a format in which voice link text data is associated and a format in which a voice link as a hyperlink is associated.
[0188] FIG. 27 is a flowchart of the process related to access authentication. First, in step S110, the CPU 1a waits for an access request to the audio data, and if such an access request is received, the CPU 1a proceeds to step S111 to perform authentication processing. That is, the access authentication processing described above is executed. Note that an example of the access authentication processing has already been explained, so a duplicate explanation will be avoided.
[0189] In step S112 following step S111, the CPU 1a determines whether or not the authentication is successful. If the authentication is successful, the CPU 1a proceeds to step S113 and performs access permission processing. That is, the access request of step S110 is permitted. This allows the device that has made the access request based on the audio link of the minutes data (the audio pickup device 20 in this example) to reproduce and output the audio data of the portion corresponding to the selected gist sentence.
[0190] On the other hand, if it is determined in step S112 that the authentication was not successful, the CPU 1a proceeds to step S114 and performs error processing, i.e., denies access to the audio data by the device that made the access request, and performs corresponding error processing, such as displaying a message on the device indicating that access is not permitted.
[0191] The CPU 1a ends the series of processes shown in FIG. 27 in response to execution of the process of either step S113 or S114.
[0192] <6. Variations> The embodiment is not limited to the specific example described above, and various modified configurations can be adopted. For example, in the above example, an audio link using text data (hereinafter referred to as "audio link text data") is associated with a key point sentence, but audio link text data can also be associated with other element text data, such as a summary sentence, when the element text data is applied to the minutes.
[0193] Alternatively, when the recorded text data is stored in the minutes management server 50, for example, the voice link text data can be associated with the recorded text data. FIG. 28 shows an example of the association of a voice link with recorded text data. Since the recorded text data is expected to be a relatively long data set from the start to the end of the interview, it is desirable to associate the voice link text data with each part of the recorded text data. In this case, the CPU 1a associates voice link text data with each unit of text data obtained by dividing the recorded text data into predetermined utterance units. Specifically, the predetermined speech unit may be a unit of a predetermined time of speech. In this case, the CPU 1a divides the recorded text data into units of the predetermined time to generate unit text data, and associates each unit text data with voice link text data that indicates a link to the corresponding voice data. Here, the predetermined speech unit may be a unit such as a predetermined number of characters or a unit for each change of speaker, and is not limited to a unit for each predetermined time as described above.
[0194] Furthermore, in the explanation so far, we have given an example of adding audio links to all gist sentences, regardless of whether they are from your own group or the other group, but it is also possible to add audio links only to gist sentences whose original utterance is an utterance from the other group (hereinafter referred to as "gist sentences based on the other group's utterance"). In interviews, what the other party says is often more important than what you say. In particular, in business negotiations, it is important to confirm the other party's latent desires, so what the other party (customer) says is more important than what you say. By assigning audio links only to the key sentences that are based on the other party's utterance as described above, it is possible to smoothly confirm the other party's latent desires, while reducing the processing burden associated with assigning audio links and reducing the required data capacity. When an audio link is assigned only to the gist sentences based on the other party's utterance, it is possible to save only the speech data of the other party's group, without saving the speech data of the own group.
[0195] In the above example, the main categories of the main points sentences are business negotiation process categories, but the main categories can also be categories other than business negotiation process categories. For example, arbitrary categories can be defined as the main categories, and multiple keywords can be prepared for each category. Then, if a matching keyword is detected in the text, the wording of that keyword can be adopted as the wording of the main points sentence for that major category, making it possible to accommodate arbitrary categories. Regarding the major classification, rather than simply adopting the wording of the matched keywords as the main points, it is also possible to adopt a method in which, for example, when keywords such as "what will be done in the future," "who will be the decision maker," and "how to proceed from here" are detected, more abstract wording such as "next action" is applied.
[0196] Furthermore, in the explanation so far, the example has been given of the interview being a business negotiation, but interviews to which the technology of the present invention can be applied include various opportunities for communication between people, such as business negotiations, as well as interviews in medical, counseling, lectures, speeches, education, and entertainment.
[0197] Furthermore, although the above example illustrates a case where the interview is conducted face-to-face, the technology of the present invention can also be suitably applied to a case where the interview is conducted as an online interview. 29 illustrates an example in which an online interview is conducted between an information processing device 60 used by a sales representative Hs and an information processing device 70 such as a PC used by a customer Hc. These information processing devices 60, 70 are assumed to be computer devices with sound collection capabilities, such as personal computers, smartphones, and tablet terminals. The online interview can be realized by exchanging at least audio between the information processing device 60 and the information processing device 70. Of course, a camera can be provided to enable the online interview to include video. By using video data as well as audio recording data as information related to the interview, it becomes possible to perform analysis based on the speaker's gestures and hand movements during the analysis process, thereby improving analysis performance.
[0198] In Figure 29, the dotted arrows indicate the transmission path of the voices of the interview participants, with the voice of sales representative Hs being picked up by the microphone of information processing device 60, and the voice of customer Hc being picked up by the microphone of information processing device 70, and then input into information processing device 60 via network NT. In this case, the information processing device 60 records the interview not only the voice data picked up by its own microphone but also the voice data acquired from the information processing device 70. Then, the information processing device 60 transmits (records) the recorded data to the voice management server 1 as shown by the solid arrow in the figure. After the recorded data is recorded in the voice management server 1, the processing of the voice management server 1 and the information processing device 60 is the same as in the case of a face-to-face interview, so duplicate explanations will be avoided.
[0199] In addition, in the above example, the device (voice management server 1) that performs the analysis processing of the interview content based on the recorded data and the device (sound collection device 20) that performs the display processing of the editing screen Ge are separate devices, but it is also possible to configure these analysis processing and display of the editing screen Ge to be performed by the same information processing device.
[0200] 20 and 27, the authentication process based on the contract is performed for access to the recorded interview data. However, the content of the contract may vary widely. For example, there may be multiple contracts with different restrictions for accessing the recorded interview data (audio playback). For example, there may be a contract with a time limit that allows use for a specified period, such as one month or one year, and a contract without a time limit. In this case, the voice management server 1 performs corresponding processing according to the contract details for each user who has made an access request.
[0201] Here, it is conceivable that the voice management server 1 and the minutes management server 50 are managed by different companies. In this case, it is conceivable that the service for supporting the creation of minutes using the editing screen Ge and the service for storing the recorded interview data are under separate contracts. For example, it is conceivable to develop multiple contract plans, such as a minutes editing plan, a minutes editing + recorded data storage (limited storage) plan, and a minutes editing + recorded data storage (unlimited storage) plan. Alternatively, the same company may manage both the minutes creation support service and the recording data storage service of the interviews. In that case, the functions of the voice management server 1 and the minutes management server 50 may be integrated into a single server device.
[0202] In addition, in the explanation so far, we have given an example in which the analysis of speech based on recorded interview data, such as creating a summary or key points, is performed using operations on the minutes creation screen (editing screen Ge) as a trigger, but it is also possible to have this analysis process be performed using the operation to end the recording of the interview as a trigger. In this case, it is conceivable to make it possible to display the analysis result information (analysis result text data) obtained by the analysis process on a management screen for business negotiations, such as the recording list screen Gi illustrated in Fig. 9. For example, as in the recording list screen Gi illustrated in Fig. 30, a display instruction button B30 for instructing the display of analysis result information such as a summary may be provided on the management screen, and the analysis result information may be displayed on the management screen in response to the operation of the display instruction button B30.
[0203] Furthermore, when a smartphone (a device with a calling function) is used as the sound collection device 20 used to record an interview as in the embodiment, it is conceivable that recording will be given priority even if a call comes in during recording. For example, if an incoming call occurs while an interview is being recorded, the ringtone and vibration upon receiving a call can be turned off, and incoming call notification information Ic, such as that shown in Fig. 31, can be displayed on the display unit 27. As the incoming call notification information Ic, for example, a receiver icon indicating an incoming call, or the caller's phone number information, as shown in the figure, can be displayed. As shown in the figure, the direction guidance information Id can be continuously displayed even during an incoming call.
[0204] Although the above example shows an incoming call, it is also conceivable to perform processing that prioritizes recording of other notifications such as e-mails.
[0205] <7. Summary of embodiments> As described above, the information processing device (voice management server 1 or sound collection device 20) as an embodiment comprises a generation processing unit (minutes generation processing unit F13) that generates linked utterance content text data by associating utterance content text data, which is text data indicating the content of utterances in an interview generated based on recorded data of the interview, with audio link text data, which is text data indicating a link to the audio data of the utterance portion in the recorded data that corresponds to the utterance content text data, and a transmission processing unit (same F14) that transmits the linked utterance content text data generated by the generation processing unit to an external server device. By associating the voice link text data with the speech content text data as described above, it is no longer necessary for the user to search for the corresponding voice data portion in the recorded data when realizing audio playback of the speech portion corresponding to the speech content text data. Therefore, when a user views the text data of the utterance content generated based on the recorded data of the interview, it is possible to improve the ease of subsequent confirmation of the details of the interview. Furthermore, according to the above configuration, the voice link is text data, and the speech content text data with the voice link is transmitted to an external server device. Some server devices that store speech content text data do not support hyperlinks. By using the voice link as text data as described above, even when the speech content text data is stored in such a server device that does not support hyperlinks, the user can play back the voice data of the corresponding speech portion using the voice link information in the text data. In other words, even when the speech content text data is stored in a server device that does not support hyperlinks, the user does not need to search for the corresponding voice data portion, thereby making it easier to confirm the details of the interview after the fact.
[0206] In addition, in the information processing device according to the embodiment, the speech content text data is element text data obtained by converting elements of the speech content in the interview into text. This means that when the text data of the utterance content stored in an external server device is converted into element text data such as a summary or a main point, link information to the corresponding audio data can be added to the element text data.
[0207] Furthermore, in the information processing device of the embodiment, the generation processing unit generates minutes data of the interview including element text data to which audio link text data is associated based on user operation, and the transmission processing unit transmits the minutes data generated by the generation processing unit to an external server device. This makes it easier for a user to check the details of the interview after the fact when viewing the minutes data including the element text data.
[0208] Furthermore, in the information processing device of the embodiment, when a local save instruction is given to the device displaying the minutes creation screen, which is an instruction to save the minutes, the generation processing unit generates minutes data in a format in which audio link text data is not associated with element text data, and when a cloud save instruction is given to a specified server device other than the device displaying the minutes creation screen, which is an instruction to save the minutes, the generation processing unit generates minutes data in a format in which audio link text data is associated with element text data. This allows the user to select whether to create minutes data with or without audio links, preventing situations where the user is only able to create and save minutes data with audio links even though the user does not think the audio link function is necessary.
[0209] In addition, in the information processing device as an embodiment, the speech content text data is recorded text data obtained by converting the spoken portion of the recorded data into text, and the generation processing unit associates voice link text data with each unit of text data obtained by dividing the recorded text data into specified speech units. As a result, when the speech content text data stored in an external server device is used as recorded speech data, link information to the corresponding voice data can be added for each predetermined unit of speech, such as for each predetermined time period, each predetermined number of characters, each speaker, etc. In other words, when a user views the recorded speech data stored in the server device, the user can play back the corresponding voice data for each predetermined unit of speech using the voice link text data.
[0210] Furthermore, in the information processing device according to the embodiment, the interview is assumed to be a business negotiation. This makes it easier to check the details of the text data of the conversations about business negotiations after the fact.
[0211] An information processing method as an embodiment is an information processing method in which an information processing device associates speech content text data, which is text data indicating the content of an utterance in an interview generated based on recorded data of the interview, with audio link text data, which is text data indicating a link to the audio data of the speech portion in the recorded data that corresponds to the speech content text data, to generate linked speech content text data, and processes the generated linked utterance content text data to an external server device. This information processing method can also provide the same functions and effects as the information processing device of the above embodiment.
[0212] The program as an embodiment is a program that causes a computer device such as a CPU 1a to execute all or part of the processing shown in FIG. In other words, the program as an embodiment is a program readable by a computer device, and causes the computer device to execute the following process: generate linked utterance content text data by associating speech content text data, which is text data indicating the content of utterances in an interview generated based on recorded interview data, with audio link text data, which is text data indicating a link to the audio data of the utterance portion in the recorded data that corresponds to the speech content text data; and transmit the generated linked utterance content text data to an external server device.
[0213] Such a program can realize a computer device that executes all or part of the functions of the voice management server 1 described in the embodiment. For example, by downloading the program to a personal computer, a portable information processing device, or the like, the personal computer can function as the voice management server 1.
[0214] Such a program can be pre-recorded on a hard disk drive (HDD) as a recording medium built into a computer or other device, or on a ROM in a microcomputer having a CPU, etc. It can also be provided by being temporarily or permanently recorded on a removable recording medium such as a memory card or optical disk. Such programs can also be downloaded from a download site via a network such as a LAN (Local Area Network) or the Internet. [Explanation of symbols]
[0215] 100 Minutes Support System 1. Audio management server 20 Sound collection device 50 Minutes Management Server NT Network Hs Sales Representative Hc customer 21,1a,50a CPU 1b Analysis Department 2,22 ROM 3,23 RAM 4,24 Bus 5,25 Input / Output Interface 6,26 Input section 7,27 Display section 8,28 Audio output section 9,29 Storage part 10,30 Communications Department 11,31 Removable storage media 12,32 drive 28a Speaker for calls 33 First sound collection section 33a, 33b First microphone 34 Second sound collection section 34a, 34b Second microphone 35 Sensor section 36 Camera Unit F11 Audio data recording processing section F12 Screen control processing part F13 Minutes generation processing section F14 Transmission processing section F15 Access authentication processing section
Claims
1. a generation processing unit that generates linked utterance content text data by associating speech content text data, which is text data indicating the content of an utterance in an interview generated based on audio recording data of the interview, with audio link text data, which is text data indicating a link to audio data of an utterance portion in the audio recording data corresponding to the speech content text data; and a transmission processing unit that transmits the linked utterance content text data generated by the generation processing unit to an external server device. Information processing device.
2. The speech content text data is element text data obtained by converting elements of the speech content in the interview into text. The information processing device according to claim 1 .
3. The generation processing unit generating, based on a user operation, minutes data of the interview including the element text data associated with the voice link text data; The transmission processing unit The minutes data generated by the generation processing unit is transmitted to an external server device. The information processing device according to claim 2 .
4. The generation processing unit When a local save instruction is given to a device displaying the minutes creation screen of the interview, the device generates minutes data in a format in which the voice link text data is not associated with the element text data, and when a cloud save instruction is given to a predetermined server device other than the device displaying the minutes creation screen, the device generates minutes data in a format in which the voice link text data is associated with the element text data. The information processing device according to claim 3 .
5. The speech content text data is recorded text data obtained by converting a spoken portion of the recorded data into text, The generation processing unit The voice link text data is associated with each unit of text data obtained by dividing the recorded text data into predetermined utterance units. The information processing device according to claim 1 .
6. The above interview is a business meeting 6. The information processing device according to claim 1.
7. The information processing device Generate linked utterance content text data by associating speech content text data, which is text data indicating the content of an utterance in the interview generated based on the recorded interview data, with voice link text data, which is text data indicating a link to voice data of the utterance portion in the recorded interview data corresponding to the speech content text data; The generated linked utterance content text data is transmitted to an external server device. Information processing methods.
8. A computer readable program, Generate linked utterance content text data by associating speech content text data, which is text data indicating the content of an utterance in the interview generated based on the recorded interview data, with voice link text data, which is text data indicating a link to voice data of the utterance portion in the recorded interview data corresponding to the speech content text data; and causing the computer device to execute a process of transmitting the generated linked utterance content text data to an external server device. program.
9. A recording medium on which a computer-readable program is recorded, Generate linked utterance content text data by associating speech content text data, which is text data indicating the content of an utterance in the interview generated based on the recorded interview data, with voice link text data, which is text data indicating a link to voice data of the utterance portion in the recorded interview data corresponding to the speech content text data; and transmitting the generated linked utterance content text data to an external server device. Recording medium.
Citation Information
Patent Citations
Speech information documentation device
JP7223469B1