Information processing system, information processing device, information processing method, and program
The described system addresses the character limit in voice-to-text conversion by using a database and AI server prompts to ensure complete transcription of voice data into text, regardless of duration.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- BIGLOBE INC
- Filing Date
- 2024-11-19
- Publication Date
- 2026-05-29
AI Technical Summary
Existing voice-to-text conversion technologies using generative AI are limited by the number of characters they can process, making it difficult to convert voice data longer than a predetermined time into text data efficiently.
An information processing system that includes a database storing prompts for an AI server to generate, determine, and concatenate text data until a termination condition is met, allowing for the conversion of audio data exceeding a predetermined duration into text data.
Enables the easy conversion of voice data of any duration into text data by using a system of prompts and AI server interactions to ensure complete transcription.
Smart Images

Figure 2026088739000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing system, an information processing apparatus, an information processing method, and a program.
Background Art
[0002] Generally, businesses that provide goods and services have call centers that respond to inquiries and receptions from consumers by phone. In many cases, the calls between consumers and call centers are recorded for reasons such as quality assurance. To record calls, there are not only those that record the call content by voice recording, but also those that convert the call content into text data using speech recognition software and record it (for example, see Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] When converting voice data into text data using speech recognition software as in the above-described technology, it costs a large amount of money. In recent years, it is also conceivable to convert voice data into text data using generative AI (Artificial Intelligence) that has been rapidly spreading. However, when generative AI converts voice data into text data, there is a limit to the number of characters to be output, and it is not possible to convert the content of a single call between a general consumer and a call center at once. Thus, there is a problem that it is difficult to easily convert voice data longer than a predetermined time into text data.
[0005] The object of the present invention is to provide an information processing system, information processing device, information processing method, and program that can easily convert audio data exceeding a predetermined duration into text data. [Means for solving the problem]
[0006] The information processing system of the present invention is It has a database, an information processing device, and an AI server. The database stores a first prompt for causing the AI server to generate text data from audio data of a predetermined length, a second prompt for causing the AI server to determine whether the text data satisfies a termination condition, and a third prompt for causing the AI server to generate text data from the audio data after the audio corresponding to the text data if the text data does not satisfy the termination condition. The aforementioned information processing device is The acquisition unit acquires the aforementioned audio data, A reading unit that reads the first prompt, the second prompt, and the third prompt from the database, The AI interface unit inputs the audio data acquired by the acquisition unit and the first prompt read by the reading unit to the AI server, causing the AI server to generate the text data from the audio data, and acquires the generated text data from the AI server. The AI interface unit inputs the text data acquired from the AI server and the second prompt read by the reading unit to the AI server to cause the AI server to determine whether the text data satisfies the termination condition, and obtains the result of the determination from the AI server. If the obtained determination result is that the text data does not satisfy the termination condition, the AI interface unit inputs the audio data acquired by the acquisition unit, the third prompt read by the reading unit, and text data concatenated from the text data acquired from the AI server in the order they were acquired to the AI server to determine whether the audio data follows the audio corresponding to the concatenated text data. The process involves having the AI server generate text data, retrieving the generated text data from the AI server, and then repeatedly having the AI server determine whether the text data satisfies the termination condition until the result of the determination obtained from the AI server indicates that the text data satisfies the termination condition, and having the AI server generate text data from the audio data following the audio corresponding to the concatenated text data. If the result of the determination obtained from the AI server indicates that the text data satisfies the termination condition, the AI server concatenates the text data obtained so far in the order in which it was obtained and outputs it.
[0007] Furthermore, the information processing apparatus of the present invention is An acquisition unit that acquires audio data of a predetermined length, A reading unit reads from a database a first prompt for the AI server to generate text data from audio data, a second prompt for the AI server to determine whether the text data satisfies a termination condition, and a third prompt for the AI server to generate text data from the audio data after the audio corresponding to the text data if the text data does not satisfy the termination condition. An AI interface unit inputs the audio data acquired by the acquisition unit and the first prompt read by the reading unit to the AI server, causing the AI server to generate the text data from the audio data, and acquires the generated text data from the AI server. The AI interface unit inputs the text data acquired from the AI server and the second prompt read by the reading unit to the AI server to cause the AI server to determine whether the text data satisfies the termination condition, and obtains the result of the determination from the AI server. If the obtained determination result is that the text data does not satisfy the termination condition, the AI interface unit inputs the audio data acquired by the acquisition unit, the third prompt read by the reading unit, and text data concatenated from the text data acquired from the AI server in the order they were acquired to the AI server to determine whether the audio data follows the audio corresponding to the concatenated text data. The process involves having the AI server generate text data, retrieving the generated text data from the AI server, and then repeatedly having the AI server determine whether the text data satisfies the termination condition until the result of the determination obtained from the AI server indicates that the text data satisfies the termination condition, and having the AI server generate text data from the audio data following the audio corresponding to the concatenated text data. If the result of the determination obtained from the AI server indicates that the text data satisfies the termination condition, the AI server concatenates the text data obtained so far in the order in which it was obtained and outputs it.
[0008] Furthermore, the information processing method of the present invention is A process for acquiring audio data of a predetermined length, The process involves reading a first prompt from the database to generate text data from audio data for the AI server, The process involves inputting the acquired audio data and the first prompt read from the database to the AI server, causing the AI server to generate the text data from the audio data, The process of obtaining the text data generated by the aforementioned AI server, The process involves reading a second prompt from the database to cause the AI server to determine whether the text data satisfies the termination condition, A process in which text data obtained from the AI server and the second prompt read from the database are input to the AI server, and the AI server determines whether the text data satisfies the termination condition, The process of obtaining the result determined by the AI server from the AI server, The process involves reading a third prompt from the database to cause the AI server to generate text data from the audio after the audio corresponding to the text data in the audio data, If the result of the judgment obtained is that the text data does not satisfy the termination condition, the process involves inputting the obtained audio data, the third prompt read from the database, and text data concatenated from the previously obtained text data in the order they were obtained to the AI server, causing the AI server to generate text data from the audio data after the audio corresponding to the concatenated text data. The process involves repeatedly instructing the AI server to determine whether the text data satisfies the termination condition until the result of the acquired determination is that the text data satisfies the termination condition, and instructing the AI server to generate text data from the audio data after the audio corresponding to the concatenated text data. If the result of the acquired determination is that the text data satisfies the termination condition, the AI server performs the process of concatenating the text data acquired so far in the order in which they were acquired and outputting it.
[0009] Furthermore, the program of the present invention, A program to be executed by a computer, On the computer, Procedure for obtaining audio data of a predetermined length, The procedure for reading the first prompt from the database to generate text data from audio data on the AI server, A procedure for inputting the acquired audio data and the first prompt read from the database to the AI server, causing the AI server to generate the text data from the audio data, The procedure for obtaining the text data generated by the aforementioned AI server, A procedure for reading a second prompt from the database to cause the AI server to determine whether the text data satisfies the termination condition, A procedure for inputting text data obtained from the AI server and the second prompt read from the database into the AI server, causing the AI server to determine whether the text data satisfies the termination condition, The process of obtaining the result determined by the AI server from the AI server, A procedure for reading a third prompt from the database for causing the AI server to generate text data from the audio after the audio corresponding to the text data in the audio data, If the result of the judgment obtained is that the text data does not satisfy the termination condition, the procedure involves inputting the obtained audio data, the third prompt read from the database, and text data obtained by concatenating the text data obtained from the AI server in the order they were obtained, to the AI server, causing the AI server to generate text data from the audio data after the audio corresponding to the concatenated text data. The procedure of causing the AI server to determine whether the acquired text data satisfies the end condition until the result of the acquired determination indicates that the text data satisfies the end condition, and the procedure of causing the AI server to generate text data from the voice after the voice corresponding to the concatenated text data in the voice data are repeated. When the result of the acquired determination indicates that the text data satisfies the end condition, the procedure of concatenating and outputting the text data acquired from the AI server in the order in which it was acquired is executed.
Effect of the Invention
[0010] In the present invention, voice data of a predetermined time or longer can be easily converted into text data.
Brief Description of the Drawings
[0011] [Figure 1] It is a diagram showing an information processing system according to the first embodiment. [Figure 2] It is a diagram showing an example of a prompt stored in the database shown in FIG. 1. [Figure 3] It is a diagram showing an example of components included in the information processing apparatus shown in FIG. 1. [Figure 4A] It is a sequence diagram for explaining an information processing method in the information processing system shown in FIG. 1. [Figure 4B] It is a sequence diagram for explaining an information processing method in the information processing system shown in FIG. 1. [Figure 5] It is a diagram showing an information processing system according to the fifth embodiment. [Figure 6] It is a diagram showing an example of a prompt stored in the database shown in FIG. 5. [Figure 7] It is a diagram showing an example of components included in the information processing apparatus shown in FIG. 5. [Figure 8] It is a diagram showing an information processing system according to the third embodiment. [Figure 9]Figure 8 shows an example of a prompt stored in the database. [Figure 10] This figure shows an example of the components of the information processing device shown in Figure 8. [Figure 11] Figure 8 is a sequence diagram illustrating the information processing method in the information processing system shown. [Figure 12] This figure shows an information processing system according to the fourth embodiment. [Figure 13] Figure 12 shows an example of a prompt stored in the database. [Figure 14] This figure shows an example of the components of the information processing device shown in Figure 12. [Figure 15] Figure 12 is a sequence diagram illustrating the information processing method in the information processing system shown. [Modes for carrying out the invention]
[0012] Embodiments of the present invention will be described below with reference to the drawings. (First Embodiment)
[0013] Figure 1 shows an information processing system according to the first embodiment. As shown in Figure 1, the information processing system in this embodiment includes an information processing device 100, a database 200, and an AI server 300. The information processing device 100, the database 200, and the AI server 300 may be connected to each other via a communication network 400, or the information processing device 100 may be directly connected to the database 200 or the AI server 300. These connections may be made wirelessly or via wired connections.
[0014] The AI server 300 is a communication device on which a generative AI is implemented. The AI server 300 can be a physical server or a virtual server. The generative AI implemented in the AI server 300 can be a general-purpose generative AI, and when a prompt (an instruction) is input, it generates and outputs data according to the instructions indicated by the prompt.
[0015] The database 200 is a storage device that pre-stores prompts that the information processing device 100 will input to the AI server 300. Figure 2 shows an example of prompts stored in the database 200 shown in Figure 1. As shown in Figure 2, the database 200 shown in Figure 1 has prompts pre-stored to be input to the AI server 300. In the example shown in Figure 2, the prompts are "Convert the audio data to text data and output it." (the first prompt, described later), "Determine whether the text data satisfies the termination condition and output the determination result." (the second prompt, described later), and "End of input text data (for example, an identifier indicating the end)."<Text_end> The following is stored: "Conversion to text data has been completed up to this point, so please convert the remaining audio data to text data and output it." (the third prompt, described later). The first and third prompts may also contain additional instructions regarding text conversion or instructions to add control identifiers when displaying the text data. For example, these instructions may include: "Distinguish between speakers and add an identifier corresponding to each speaker for each utterance," "Separate utterances into a predetermined maximum number of characters (e.g., 200 characters) and add identifiers to indicate the start and end of each utterance," "Add a start time and an end time to each utterance," and "Add identifiers to indicate the start and end of the text data to be output (e.g.,<Text_end> This includes adding ) and, and not converting words that do not make sense between statements, such as 'uh' and 'um', into text. This information is prewritten using communication devices and input devices of administrators connected to database 200. Furthermore, prompts stored in database 200 can be added, deleted, and edited using communication devices and input devices connected to database 200.
[0016] The information processing device 100 is a device operated by the operators who manage this information processing system or by the administrators who manage this information processing system. The information processing device 100 may be, for example, a computer of the size of a server, a computer of the size of a PC (Personal Computer), or a virtual computer on the cloud.
[0017] Figure 3 shows an example of the components of the information processing device 100 shown in Figure 1. As shown in Figure 3, the information processing device 100 shown in Figure 1 has an acquisition unit 110, a reading unit 120, and an AI interface unit 130. Figure 3 shows only the main components of the information processing device 100 shown in Figure 1 that are relevant to this embodiment.
[0018] The acquisition unit 110 acquires audio data to be converted into text data. The acquisition unit 110 connects to the system where the audio data is recorded via the communication network 400 using an API (Application Programming Interface) and acquires the audio data. Alternatively, if the audio data is stored in the database 200, the acquisition unit 110 may connect to the database 200 via the communication network 400 and acquire the audio data by reading it from the database 200. Furthermore, the acquisition unit 110 may connect to the system where the audio data is recorded via the communication network 400 and acquire the audio data from the system via streaming. The audio data acquired by the acquisition unit 110 may be telephone call data exceeding a predetermined time, recordings of meetings, audio data recorded at lectures or plays, or video data recorded in place of audio data. The audio data may contain the voices of multiple speakers. The audio data and video data acquired by the acquisition unit 110 may be handled as computer files. The file format can be MP3, WAV, or MP4, or it can be audio data acquired in real time via streaming; there are no specific restrictions. The acquisition unit 110 outputs the acquired audio data to the AI interface unit 130. The predetermined time is approximately 8,000 characters, or about 10 minutes.
[0019] The reading unit 120 reads the necessary prompts from the database 200. The reading unit 120 outputs the read prompts to the AI interface unit 130.
[0020] The AI interface unit 130 inputs the voice data acquired by the acquisition unit 110 and a first prompt read by the reading unit 120 from the database 200, which causes the AI server 300 to generate text data from the voice data, to the AI server 300. When the AI interface unit 130 acquires the text data generated by the AI server 300, it inputs the acquired text data and a second prompt read by the reading unit 120 from the database 200, which causes the AI server 300 to determine whether the acquired text data satisfies the termination condition, to the AI server 300. The AI interface unit 130 acquires the result determined by the AI server 300. If the result of the acquired judgment indicates that the text data does not satisfy the termination condition, the AI interface unit 130 inputs to the AI server 300 a third prompt to cause the AI server 300 to generate text data from the audio data acquired by the acquisition unit 110, the audio data read from the database 200 by the reading unit 120 after the audio corresponding to the converted text data, and text data concatenated from the text data acquired so far in the order it was acquired. On the other hand, if the result of the acquired judgment indicates that the text data satisfies the termination condition, the AI interface unit 130 outputs the text data acquired so far from the AI server 300, concatenated in the order it was acquired. This termination condition is a condition that the generation AI implemented in the AI server 300 has learned in advance using existing call content stored in audio data, text data, phrases commonly used at the end of a call, or situations that the generation AI often outputs at the end of a call, such as the repetition of the same phrase.
[0021] The information processing method in the information processing system shown in Figure 1 will be described below. Figures 4A and 4B are sequence diagrams illustrating the information processing method in the information processing system shown in Figure 1.
[0022] First, the acquisition unit 110 acquires the target audio data (step S1). The reading unit 120 reads the first prompt from the database 200 (step S2). The order in which the process of step S1 and the process of step S2 are performed does not matter. Next, the AI interface unit 130 inputs the audio data acquired by the acquisition unit 110 and the first prompt read by the reading unit 120 to the AI server 300 (step S3). The AI server 300 then generates text data for the input audio data according to the input first prompt (step S4) and outputs the generated text data (step S5).
[0023] When the AI interface unit 130 acquires the text data output by the AI server 300, the reading unit 120 reads a second prompt from the database 200 (step S6). Alternatively, the reading unit 120 may read the second prompt from the database 200 before the AI interface unit 130 acquires the text data output by the AI server 300. The AI interface unit 130 inputs the text data acquired from the AI server 300 and the second prompt read by the reading unit 120 from the database 200 to the AI server 300 (step S7). The AI server 300 then determines whether the input text data satisfies the termination condition according to the input second prompt (step S8). Finally, the AI server 300 outputs the result of the determination (step S9). The AI server 300 may also indicate the determination result using a flag. For example, the AI server 300 may output a flag of "1" if the text data satisfies the termination condition, and output a flag of "0" if the text data does not satisfy the termination condition.
[0024] The AI interface unit 130 acquires the judgment result output by the AI server 300 and determines whether the judgment result satisfies the termination condition (step S10). For example, the AI interface unit 130 determines whether the AI server 300 output a flag of "1" or "0". If the judgment result satisfies the termination condition, the AI interface unit 130 outputs the text data (step S11). The output concatenated text data is stored in the storage unit or database 200 of the information processing device 100. Subsequently, the output text data is displayed on a communication device such as an administrator or caller connected to the information processing device 100 or database 200. If the text data contains a control identifier, the text is displayed on the communication device according to the control identifier.
[0025] On the other hand, if the judgment result in step S10 is one that does not satisfy the termination condition, the reading unit 120 reads a third prompt from the database 200 (step S12). The AI interface unit 130 also concatenates the text data acquired from the AI server 300 in the order in which it was acquired (step S13). The order in which the processing in step S12 and the processing in step S13 are performed does not matter. The AI interface unit 130 inputs the audio data acquired by the acquisition unit 110, the text data concatenated in step S13, and the third prompt read by the reading unit 120 to the AI server 300 (step S14). The AI server 300 then generates text data following the concatenated text data for the input audio data, according to the input third prompt (step S15). The AI server 300 outputs the generated text data (step S16).
[0026] When the AI interface unit 130 acquires the text data output from the AI server 300 in step S16, it inputs the acquired text data and the second prompt read by the reading unit 120 from the database 200 to the AI server 300 (step S17). Here, the text data input to the AI server 300 along with the second prompt may be a concatenation of the text data acquired from the AI server 300 in the order in which they were acquired, as long as it includes at least the text data output from the AI server 300 in step S16. The AI server 300 then determines whether the input text data satisfies the termination condition according to the input second prompt (step S18). The AI server 300 then outputs the result of the determination (step S19).
[0027] The AI interface unit 130 acquires the judgment result output by the AI server 300 and determines whether the judgment result satisfies the termination condition (step S20). If the judgment result satisfies the termination condition, the AI interface unit 130 concatenates the text data with the text data concatenated in step S13 (step S21). Then, the processing in step S11 is performed.
[0028] On the other hand, if the result of the judgment in step S20 is that the termination condition is not met, the process in step S13 is repeated.
[0029] Furthermore, the first and third prompts stored in the database 200 may include generating emotion information from the audio data and associating it with text data. In that case, when the AI server 300 generates text data from the input audio data according to the input first or third prompt, it may perform emotion analysis on the audio data and output the result as emotion information associated with the text data. For example, if the AI server 300 detects a negative emotion from the audio data, it may add emotion information "-1" to the part of the text data corresponding to that part of the audio (e.g., a word or sentence), and if it detects a positive emotion from the audio data, it may add emotion information "+1" to the part of the text data corresponding to that part of the audio data (e.g., a word or sentence) and output it. The AI interface unit 130 acquires the emotion information output by the AI server 300. Furthermore, negative and positive emotional information may represent multiple levels corresponding to the intensity of the emotion, including not only "-1" and "+1," but also "-2" and "+2," etc., and a neutral value of "0" may also be included, representing neither negative nor positive.
[0030] Thus, in this configuration, when converting audio data to text data using a generation AI, there are limitations on the number of characters and other quantities of text data that the generation AI can convert and generate. Therefore, the generation AI is made to determine whether the converted text data satisfies the termination condition. If the termination condition is not met, the subsequent audio corresponding to the converted text data is further converted into text data, and this process is repeated until the converted text data satisfies the termination condition. Once the termination condition is met, the information processing device 100 concatenates and outputs the text data that the generation AI has generated up to that point. As a result, audio data exceeding a predetermined time can be easily converted into text data. (Second Embodiment)
[0031] Figure 5 shows an information processing system according to the second embodiment. As shown in Figure 5, the information processing system in this embodiment includes an information processing device 101, a database 201, and an AI server 300. The information processing device 101, the database 201, and the AI server 300 may be connected to each other via a communication network 400, or the information processing device 101 may be directly connected to the database 201 or the AI server 300. These connections may be made wirelessly or via wired connections. The AI server 300 is the same as that in the first embodiment. Note that explanations of points that are the same as in the first embodiment may be omitted, and the explanation will focus on the points that differ from the first embodiment.
[0032] Database 201 is a storage device that pre-stores prompts that the information processing device 101 inputs to the AI server 300. Figure 6 shows an example of prompts stored in the database 201 shown in Figure 5. As shown in Figure 6, the database 201 shown in Figure 5 pre-stores prompts to be input to the AI server 300. In the example shown in Figure 6, the first prompt and the third prompt are the same as in the first embodiment. The second prompt differs from that of the first embodiment in that it is a prompt with termination conditions set. The second prompt shown in Figure 6 is an instruction to determine that the call has ended if the text data contains phrases that are commonly used when ending a call, or if the same phrases are repeated. Note that phrases that are commonly used when ending a call may be included in the second prompt, or they may be stored in one of the storage devices and their storage location may be included in the second prompt. These are examples of termination conditions. The termination conditions can be the words contained in the text data, or they can be quantities such as the number of characters in the text data that the generation AI can generate, or the file size of the text data. These are pieces of information that have been prewritten using communication devices and input devices connected to database 201. In addition, prompts stored in database 201 can be added, deleted, and edited using communication devices and input devices connected to database 201.
[0033] The information processing device 101 is a device operated by the operators who run this information processing system or by the administrators who manage this information processing system. The information processing device 101 may be a computer of the size of a server, a computer of the size of a PC, or a virtual computer on the cloud. The information processing device 101 differs from the first embodiment in that it connects to database 201 instead of database 200 in the first embodiment.
[0034] Figure 7 shows an example of the components of the information processing device 101 shown in Figure 5. As shown in Figure 7, the information processing device 101 shown in Figure 5 has an acquisition unit 110, a reading unit 120, and an AI interface unit 131. The acquisition unit 110 and the reading unit 120 are the same as those in the first embodiment. Figure 7 shows only the main components of the information processing device 101 shown in Figure 5 that are relevant to this embodiment.
[0035] In the first embodiment, the AI interface unit 130 inputs a second prompt stored in the database 200 to the AI server 300 when it acquires text data generated by the AI server 300. However, the AI interface unit 131 has the function of inputting a second prompt containing a termination condition stored in the database 201 to the AI server 300, causing the AI server 300 to determine whether the text data acquired from the AI server 300 satisfies the termination condition. Other functions of the AI interface unit 131 are the same as those of the AI interface unit 130 in the first embodiment.
[0036] The following describes the information processing method in the information processing system in this embodiment shown in Figure 5. In the sequence diagrams of Figures 4A and 4B used to describe the first embodiment, the database 200 is replaced by the database 201, the information processing device 100 is replaced by the information processing device 101, and the AI interface unit 130 is replaced by the AI interface unit 131.
[0037] In particular, in step S6 of the first embodiment, this embodiment differs from the first embodiment in that when the AI interface unit 131 acquires the text data output by the AI server 300, the reading unit 120 reads a second prompt from the database 201. Similarly, in step S7 of the first embodiment, this embodiment differs from the first embodiment in that the AI interface unit 131 inputs the text data acquired from the AI server 300 and the second prompt read by the reading unit 120 from the database 201 to the AI server 300. Furthermore, in step S17 of the first embodiment, this embodiment differs from the first embodiment in that when the AI interface unit 131 acquires the text data output from the AI server 300 in step S16, it inputs the acquired text data and the second prompt read by the reading unit 120 from the database 201 to the AI server 300.
[0038] Thus, in this embodiment, unlike the first embodiment, the database 201 stores the termination conditions for text data in advance as prompts and inputs them to the generating AI. Therefore, the termination of a call can be determined based on desired conditions. Alternatively, the information processing device 101 may determine the termination of a call based on the text data acquired from the AI server 300 and the termination conditions stored in the database 201. Furthermore, this embodiment and the first embodiment may be combined so that the AI interface unit 131 inputs a second prompt to the AI server 300, which uses both the termination conditions previously learned by the generating AI and the termination conditions previously stored in advance as prompts by the database 201, to prompt the AI server 300 for determination. (Third embodiment)
[0039] Figure 8 shows an information processing system according to the third embodiment. As shown in Figure 3, the information processing system in this embodiment includes an information processing device 102, a database 202, and an AI server 300. The information processing device 102, the database 202, and the AI server 300 may be connected to each other via a communication network 400, or the information processing device 102 may be directly connected to the database 202 or the AI server 300. These connections may be made wirelessly or via wired connections. The AI server 300 is the same as that in the first embodiment. Note that explanations of points that are the same as in the first embodiment may be omitted, and the explanation will focus on the points that differ from the first embodiment.
[0040] Database 202 is a storage device that pre-stores prompts that the information processing device 102 inputs to the AI server 300. Figure 9 shows an example of prompts stored in the database 202 shown in Figure 8. As shown in Figure 9, the database 202 shown in Figure 8 pre-stores prompts to be input to the AI server 300. In the example shown in Figure 9, the first, second, and third prompts are the same as in the first embodiment. In the example shown in Figure 9, the fourth prompt, "Create a summary of the entered text data," is added. When instructing the AI server 300 to create a summary, instructions such as adding a control identifier for displaying the summary data may be added to the fourth prompt. For example, this instruction may be, "Add a summary start identifier and a summary end identifier to the beginning and end of the summary, respectively." This information is pre-written using a communication device or input device connected to the database 202. Furthermore, the prompts stored in the database 202 can be added, deleted, or edited using a communication device or input device connected to the database 202.
[0041] The information processing device 102 is a device operated by the operator who runs this information processing system or by the administrator who manages this information processing system. The information processing device 102 may be a computer of the size of a server, a computer of the size of a PC, or a virtual computer on the cloud. The information processing device 102 differs from the first embodiment in that it connects to database 202 instead of database 200 in the first embodiment.
[0042] Figure 10 shows an example of the components of the information processing device 102 shown in Figure 8. As shown in Figure 10, the information processing device 102 shown in Figure 8 has an acquisition unit 110, a reading unit 120, and an AI interface unit 132. The acquisition unit 110 and the reading unit 120 are the same as those in the first embodiment. Figure 10 shows only the main components of the information processing device 102 shown in Figure 8 that are relevant to this embodiment.
[0043] In addition to the functions of the AI interface unit 130 in the first embodiment, the AI interface unit 132 inputs the concatenated text data, which is determined to have ended the call, and a prompt (the fourth prompt) read by the reading unit 120 from the database 202, to the AI server 300 to generate a summary, causing the AI server 300 to generate a summary. The AI interface unit 132 retrieves the summary generated by the AI server 300 from the AI server 300. Other functions of the AI interface unit 132 are the same as those of the AI interface unit 130 in the first embodiment.
[0044] The information processing method in the information processing system shown in Figure 8 will be described below. Figure 11 is a sequence diagram illustrating the information processing method in the information processing system shown in Figure 8. Here, the processing after the completion of the same processing as in the first embodiment (transcription of voice data until the completion of the call) (before or after step S11) will be described.
[0045] First, the reading unit 120 reads the fourth prompt from the database 202 (step S31). Next, the AI interface unit 132 inputs the converted text data and the fourth prompt read by the reading unit 120 to the AI server 300 (step S32). The AI server 300 then generates a summary of the input text data according to the input fourth prompt (step S33) and outputs the generated summary as summary data (step S34). The AI interface unit 132 then acquires the summary data output by the AI server 300 and outputs the acquired summary data (step S35). The output summary data is stored in the storage unit of the information processing device 102 or in the database 202. Subsequently, the output summary data is displayed on communication devices such as administrators and callers connected to the information processing device 102 or database 202. If the summary data contains a control identifier, the summary data is displayed on the communication device according to the control identifier.
[0046] Thus, in this embodiment, in addition to the first embodiment, a generation AI is used to generate a summary of the text data. This makes it easier to analyze the content of the call. Furthermore, this embodiment may be combined with the second embodiment instead of the first embodiment, and the processing may be performed after the same processing as the second embodiment (transcription of audio data until the end of the call) is completed, or it may be combined with a combination of the first embodiment and the second embodiment. (Fourth embodiment)
[0047] Figure 12 shows an information processing system according to the fourth embodiment. As shown in Figure 12, the information processing system in this embodiment includes an information processing device 103, a database 203, and an AI server 300. The information processing device 103, the database 203, and the AI server 300 may be connected to each other via a communication network 400, or the information processing device 103 may be directly connected to the database 203 or the AI server 300. These connections may be made wirelessly or via wired connections. The AI server 300 is the same as that in the first embodiment. Note that explanations of points that are the same as in the first embodiment may be omitted, and the explanation will focus on the points that differ from the first embodiment.
[0048] Database 203 is a storage device that pre-stores prompts that the information processing device 103 inputs to the AI server 300. Figure 13 shows an example of prompts stored in database 203 shown in Figure 12. Database 202 shown in Figure 8 pre-stores prompts to be input to the AI server 300, as shown in Figure 9. In the example shown in Figure 13, the first, second, and third prompts are the same as in the first embodiment. In the example shown in Figure 13, a fifth prompt, "Analyze the cause of the emotion and output it," has been added. This fifth prompt is used when emotion information is associated with text data generated from call data and output, causing the AI server 300 to analyze and output the cause of the emotion indicated by the emotion information based on the associated emotion information. When the AI server 300 analyzes the cause and outputs it, instructions to add control identifiers when displaying the cause data may be added to the prompt. For example, such instructions include "Add identifiers indicating the start and end of the phrase to which the emotion information is attached and the phrase to which the cause is attached, and further add identifiers to associate the phrase with the cause." These prompts are pre-written information using communication and input devices connected to database 203. Furthermore, prompts stored in database 203 can be added, deleted, and edited using communication and input devices connected to database 203.
[0049] The information processing device 103 is a device operated by the operators who run this information processing system or by the administrators who manage this information processing system. The information processing device 103 may be a computer of the size of a server, a computer of the size of a PC, or a virtual computer on the cloud. The information processing device 103 connects to database 203 instead of database 200 in the first embodiment.
[0050] Figure 14 shows an example of the components of the information processing device 103 shown in Figure 12. As shown in Figure 14, the information processing device 103 shown in Figure 12 has an acquisition unit 110, a reading unit 120, and an AI interface unit 133. The acquisition unit 110 and the reading unit 120 are the same as those in the first embodiment. Figure 14 shows only the main components of the information processing device 103 shown in Figure 12 that are relevant to this embodiment.
[0051] In addition to the functions of the AI interface unit 130 in the first embodiment, the AI interface unit 133 inputs text data that has been linked when it is determined that the call has ended, emotion information output in association with that text data, and a prompt (fifth prompt) read by the reading unit 120 from the database 203 to the AI server 300, causing the AI server 300 to analyze and output the cause of the emotion indicated by the emotion information, thereby causing the AI server 300 to analyze the cause of the emotion indicated by the emotion information. The AI interface unit 133 obtains the cause of the emotion generated by the AI server 300 from the AI server 300. Other functions of the AI interface unit 133 are the same as those of the AI interface unit 130 in the first embodiment.
[0052] The information processing method in the information processing system shown in Figure 12 will be described below. Figure 15 is a sequence diagram illustrating the information processing method in the information processing system shown in Figure 12. Here, the processing after the completion of the same processing as in the first embodiment (transcription of voice data until the completion of the call) (before or after step S11) will be described.
[0053] First, the reading unit 120 reads the fifth prompt from the database 202 (step S41). Next, the AI interface unit 133 inputs the transcribed text data, the emotion information associated with the text data, and the fifth prompt read by the reading unit 120 to the AI server 300 (step S42). The AI server 300 then analyzes the cause of the emotion based on the input text data and emotion information according to the input fifth prompt (step S43), and outputs the analyzed cause as cause data (step S44). For example, the AI server 300 may extract the portion of the text data associated with positive emotion information (for example, a statement made by one speaker in a call), and further extract the word (statement or phrase) that caused the positive emotion from the text data immediately preceding that portion (for example, a statement made by the other speaker in a call), and then associate the portion of the text data associated with positive emotion information with the word that caused the positive emotion to create cause data. Furthermore, the AI server 300 may extract a portion of the text data associated with negative emotional information (for example, a statement made by one speaker in a call), and then extract a word (statement or phrase) from the text data immediately preceding that portion (for example, a statement made by the other speaker in a call) that caused the negative emotion. The AI server 300 may then associate the portion of the text data associated with the negative emotional information with the word that caused the negative emotion to create cause data. The AI interface unit 133 then acquires the cause data output by the AI server 300 and outputs the acquired cause data (step S45). The output cause data is stored in the storage unit or database 203 of the information processing device 103. Subsequently, the output cause data is displayed to communication devices such as administrators or callers connected to the information processing device 103 or database 203. If the cause data includes a control identifier, the statement to which the emotional information is attached and the word that caused it are displayed to the communication device according to the control identifier.
[0054] Thus, in this embodiment, in addition to the first embodiment, emotion information is generated from call data, and the cause of that emotion is analyzed and output based on the generated emotion information. Therefore, it is possible to easily analyze what kind of statements reassure the other party and what kind of statements cause the other party to feel uneasy. In this embodiment, instead of processing after the completion of the same processing as the first embodiment, the cause analysis instruction may also be added to the prompt in addition to the instruction to perform emotion analysis at the first or third prompt. Furthermore, this embodiment may be combined with the second embodiment instead of the first embodiment, and processed after the completion of the same processing as the second embodiment (transcription of voice data until the completion of the call), or it may be combined with a combination of the first and second embodiments. Moreover, this embodiment may be combined with the third embodiment, and further combined with at least one of the first and second embodiments.
[0055] The above explanation describes how each component is assigned a specific function (process), but this assignment is not limited to those described above. Similarly, the configurations of the components described above are merely examples and are not limited to them. Multiple components can be combined into a single component, a single component can be divided into multiple components, or one component can perform the function (process) of another component. Furthermore, combinations of various embodiments are also possible.
[0056] The processing performed by each of the above-mentioned components may be carried out by logic circuits created according to their respective purposes. Alternatively, a computer program (hereinafter referred to as "program") describing the processing content as a procedure may be recorded on a recording medium readable by each of the information processing devices 100 to 103, and the program recorded on this recording medium may be read by each of the information processing devices 100 to 103 and executed. Recording media readable by each of the information processing devices 100 to 103 refer to portable recording media such as floppy disks, magneto-optical disks, DVDs (Digital Versatile Discs), CDs (Compact Discs), Blu-ray Discs (Registered Trademarks), USB (Universal Serial Bus) memory, and SD cards, as well as ROM (Read Only Memory), RAM, and HDDs (Hard Disc Drives) built into each of the information processing devices 100 to 103. The program recorded on this recording medium is read by the CPU provided in each of the information processing devices 100 to 103, and the same processing as described above is performed under the control of the CPU. Here, the CPU acts as a computer that executes programs read from a recording medium on which those programs are stored. [Explanation of symbols]
[0057] 100-103 Information Processing Equipment 110 Acquisition Department 120 Reading section 130-133 AI Interface Section Databases 200-203 300 AI servers 400 Communication Networks
Claims
1. It has a database, an information processing device, and an AI server. The database stores a first prompt for causing the AI server to generate text data from audio data of a predetermined length, a second prompt for causing the AI server to determine whether the text data satisfies a termination condition, and a third prompt for causing the AI server to generate text data from the audio data after the audio corresponding to the text data if the text data does not satisfy the termination condition. The aforementioned information processing device is The acquisition unit acquires the aforementioned audio data, A reading unit that reads the first prompt, the second prompt, and the third prompt from the database, The system includes an AI interface unit that inputs the audio data acquired by the acquisition unit and the first prompt read by the reading unit to the AI server, causes the AI server to generate the text data from the audio data, and retrieves the generated text data from the AI server. The AI interface unit inputs the text data acquired from the AI server and the second prompt read by the reading unit to the AI server to cause the AI server to determine whether the text data satisfies the termination condition, and obtains the result of the determination from the AI server. If the obtained determination result is that the text data does not satisfy the termination condition, the AI interface unit inputs the audio data acquired by the acquisition unit, the third prompt read by the reading unit, and text data concatenated from the text data acquired from the AI server in the order in which they were acquired to the AI server to input the audio data after the audio corresponding to the concatenated text data. An information processing system that causes the AI server to generate text data, retrieves the generated text data from the AI server, and then repeatedly performs the following processes: causing the AI server to determine whether the text data satisfies the termination condition until the result of the determination obtained from the AI server indicates that the text data satisfies the termination condition; and causing the AI server to generate text data from the audio data after the audio corresponding to the concatenated text data. If the result of the determination obtained from the AI server indicates that the text data satisfies the termination condition, the system outputs the text data obtained from the AI server in the order in which it was obtained.
2. In the information processing system described in claim 1, The termination condition is an information processing system in which the AI server has previously learned the condition using the text data.
3. In the information processing system described in claim 1, The database is an information processing system that stores the termination condition in the second prompt.
4. In the information processing system described in claim 3, The database is an information processing system that stores as a termination condition whether the text data contains text in which the same phrase is repeated.
5. In the information processing system according to any one of claims 1 to 4, The database stores a fourth prompt for the AI server to generate a summary of the concatenated text data. The reading unit reads the fourth prompt from the database after the AI interface unit outputs the concatenated text data. The AI interface unit is an information processing system that inputs the concatenated text data and the fourth prompt read by the reading unit to the AI server to cause the AI server to generate the summary, and retrieves the generated summary from the AI server.
6. In the information processing system according to any one of claims 1 to 4, The database stores, including the generation of emotion information indicating emotion from the audio data and associating it with the text data, for the first prompt and the third prompt. The AI interface unit is an information processing system that causes the AI server to generate the emotion information in association with the text data, and retrieves the generated emotion information from the AI server.
7. In the information processing system described in claim 6, The database stores a fifth prompt that causes the AI server to analyze the cause of the emotion indicated by the emotion information, based on the emotion information associated with the text data. The reading unit further reads the fifth prompt from the database, The AI interface unit is an information processing system that inputs the text data, the emotion information, and the fifth prompt read by the reading unit to the AI server, causes the AI server to analyze the cause of the emotion indicated by the emotion information, and obtains the cause from the AI server.
8. In the information processing system according to any one of claims 1 to 4, The aforementioned audio data is an information processing system consisting of telephone call data.
9. An acquisition unit that acquires audio data of a predetermined length, A reading unit reads from a database a first prompt for the AI server to generate text data from audio data, a second prompt for the AI server to determine whether the text data satisfies a termination condition, and a third prompt for the AI server to generate text data from the audio data after the audio corresponding to the text data if the text data does not satisfy the termination condition. An AI interface unit inputs the audio data acquired by the acquisition unit and the first prompt read by the reading unit to the AI server, causing the AI server to generate the text data from the audio data, and acquires the generated text data from the AI server. The AI interface unit inputs the text data acquired from the AI server and the second prompt read by the reading unit to the AI server to cause the AI server to determine whether the text data satisfies the termination condition, and obtains the result of the determination from the AI server. If the obtained determination result is that the text data does not satisfy the termination condition, the AI interface unit inputs the audio data acquired by the acquisition unit, the third prompt read by the reading unit, and text data concatenated from the text data acquired from the AI server in the order in which they were acquired to the AI server to determine the text data from the audio data after the audio corresponding to the concatenated text data. An information processing device that causes the AI server to generate text data, retrieves the generated text data from the AI server, and then repeatedly performs the following processes: causing the AI server to determine whether the text data satisfies the termination condition until the result of the determination obtained from the AI server is that the text data satisfies the termination condition; and causing the AI server to generate text data from the audio data after the audio corresponding to the concatenated text data, and if the result of the determination obtained from the AI server is that the text data satisfies the termination condition, the information processing device outputs the text data obtained from the AI server in the order in which it was obtained.
10. A process for acquiring audio data of a predetermined length, The process involves reading a first prompt from the database to generate text data from audio data for the AI server, The process involves inputting the acquired audio data and the first prompt read from the database to the AI server, causing the AI server to generate the text data from the audio data, The process of obtaining the text data generated by the aforementioned AI server, The process involves reading a second prompt from the database to cause the AI server to determine whether the text data satisfies the termination condition, A process in which text data obtained from the AI server and the second prompt read from the database are input to the AI server, and the AI server is instructed to determine whether the text data satisfies the termination condition. The process involves obtaining the result determined by the AI server from the AI server, The process involves reading a third prompt from the database for the AI server to generate text data from the audio after the audio corresponding to the text data in the audio data, If the result of the acquired determination is that the text data does not satisfy the termination condition, the acquired audio data, the third prompt read from the database, and the text data obtained by concatenating the text data acquired so far from the AI server in the order they were acquired are input to the AI server, causing the AI server to generate text data from the audio data after the audio corresponding to the concatenated text data. The process involves repeatedly instructing the AI server to determine whether the text data satisfies the termination condition until the result of the acquired determination is that the text data satisfies the termination condition, and instructing the AI server to generate text data from the audio data after the audio corresponding to the concatenated text data. An information processing method that, if the result of the acquired determination is that the text data satisfies the termination condition, performs the process of concatenating the text data acquired so far from the AI server in the order in which it was acquired and outputting it.
11. On the computer, Procedure for obtaining audio data of a predetermined length, The procedure for reading the first prompt from the database to generate text data from audio data on the AI server, A procedure for inputting the acquired audio data and the first prompt read from the database to the AI server, causing the AI server to generate the text data from the audio data, The procedure for obtaining the text data generated by the aforementioned AI server, A procedure for reading a second prompt from the database to cause the AI server to determine whether the text data satisfies the termination condition, A procedure for inputting text data obtained from the AI server and the second prompt read from the database into the AI server, and causing the AI server to determine whether the text data satisfies the termination condition, The process involves obtaining the result determined by the AI server from the AI server, A procedure for reading a third prompt from the database for causing the AI server to generate text data from the audio after the audio corresponding to the text data in the audio data, If the result of the judgment obtained is that the text data does not satisfy the termination condition, the procedure involves inputting the obtained audio data, the third prompt read from the database, and text data concatenated from the previously obtained text data in the order they were obtained to the AI server, causing the AI server to generate text data from the audio data after the audio corresponding to the concatenated text data. The procedure involves repeating the steps of: having the AI server determine whether the text data satisfies the termination condition until the result of the acquired determination is that the text data satisfies the termination condition; and having the AI server generate text data from the audio data after the audio corresponding to the concatenated text data. A program to perform the following steps: if the result of the acquired judgment is that the text data satisfies the termination condition, the AI server concatenates and outputs the text data acquired so far in the order in which it was acquired.