Program, method, information processing device, and system
By executing a program on the computer that includes audio acquisition, QA output, QA designation and FAQ acquisition steps, and using a large-scale language model to generate and organize FAQ, the problem of inappropriate FAQ creation in the prior art is solved, and high-quality Q&A information provision is achieved.
Patent Information
- Application Number
- JP2024101612
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2025-05-07
- Estimated Expiration
- 2044-06-25
AI Technical Summary
The prior art is difficult to create suitable FAQs, resulting in users not being able to obtain valid information.
By executing a program on the computer, the program includes an audio acquisition step, a QA output step, a QA designation step, and a FAQ acquisition step. The program uses a large-scale language model to generate and combine appropriate Q&A pairs based on the content of the call sound and organize them into FAQs.
The creation of suitable FAQ is realized, allowing users to obtain high-quality Q&A information, and solves the problem that FAQ creation in the prior art is not suitable.
Smart Images

Figure 0007672025000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a program, a method, an information processing device, and a system. [Background technology]
[0002] Techniques for creating FAQs are known. Patent Document 1 discloses a technique for providing an FAQ providing device and an FAQ providing program that provide appropriate FAQs corresponding to a user's system configuration. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2021-157569 A Summary of the Invention [Problem to be solved by the invention]
[0004] There is an issue with not being able to create suitable FAQs. Therefore, the present disclosure has been made to solve the above problem, and has an object to provide a technique for creating suitable FAQs. [Means for solving the problem]
[0005] A program to be executed by a computer having a processor and a storage unit, the program executing: a voice acquisition step of acquiring voice data in which a telephone call voice is recorded; a QA output step of outputting a plurality of QA information consisting of a plurality of questions and a plurality of answers to the questions based on the content of the telephone call voice stored in the voice data acquired in the voice acquisition step; a QA identification step of identifying a predetermined QA group consisting of a plurality of similar QA information from the plurality of QA information output in the QA output step; and an FAQ acquisition step of creating a second prompt by including the QA group identified in the QA identification step in a second instruction for outputting FAQ information formatted from the plurality of QA information, and acquiring FAQ information consisting of questions and answers to the questions based on output information output by inputting the second prompt into a large-scale language model. Effect of the Invention
[0006] According to the present disclosure, it is possible to create suitable FAQs. [Brief description of the drawings]
[0007] [Figure 1] FIG. 2 is a block diagram showing the functional configuration of the system 1. [Diagram 2] FIG. 2 is a block diagram showing the functional configuration of the server 10. [Diagram 3] 2 is a block diagram showing the functional configuration of a user terminal 20. FIG. [Figure 4] FIG. 13 is a diagram showing the data structure of a user table 1012. [Diagram 5] FIG. 10 shows the data structure of a call table 1014. [Figure 6] FIG. 13 is a diagram showing the data structure of a QA table 1015. [Figure 7] FIG. 13 is a diagram showing the data structure of a FAQ table 1016. [Figure 8] 13 is a flowchart showing the operation of an FAQ creation process. [Figure 9]FIG. 2 is a block diagram showing the basic hardware configuration of a computer 90. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0008] Hereinafter, an embodiment of the present disclosure will be described with reference to the drawings. In all the drawings explaining the embodiment, the same reference numerals are given to common components, and repeated explanations are omitted. Note that the following embodiment does not unduly limit the contents of the present disclosure described in the claims. In addition, not all of the components shown in the embodiment are essential components of the present disclosure. In addition, each figure is a schematic diagram and is not necessarily illustrated strictly.
[0009] <System 1 Configuration> The system 1 in the present disclosure is an information processing system that provides an information processing service for creating FAQs consisting of pairs of questions and answers. The system 1 includes an information processing device including a server 10, a user terminal 20, and a generating AI 50, which are connected via a network N. FIG. 1 is a block diagram showing the functional configuration of the system 1. FIG. 2 is a block diagram showing the functional configuration of the server 10. As shown in FIG. FIG. 3 is a block diagram showing the functional configuration of the user terminal 20. As shown in FIG.
[0010] Each information processing device is composed of a computer equipped with a calculation device and a storage device. The basic hardware configuration of the computer and the basic functional configuration of the computer realized by the hardware configuration will be described later. For each of the server 10, the user terminal 20, and the generation AI 50, the description overlapping with the basic hardware configuration and the basic functional configuration of the computer described later will be omitted.
[0011] <Server 10 Configuration> The server 10 is an information processing device that provides an information processing service for creating FAQs consisting of pairs of questions and answers. The server 10 includes a storage unit 101 and a control unit 104 .
[0012] <Configuration of the storage unit 101 of the server 10> The storage unit 101 of the server 10 includes an application program 1011 , a user table 1012 , a call table 1014 , a QA table 1015 , and an FAQ table 1016 .
[0013] The application program 1011 is a program for causing the control unit 104 of the server 10 to function as each functional unit. Application programs 1011 include applications such as a web browser application.
[0014] User table 1012 is a table that stores and manages information about member users (hereinafter, users) who use the service. When a user registers to use the service, the user's information is stored in a new record in user table 1012. This allows the user to use the service according to the present disclosure. The user table 1012 is a table having columns of user ID, user name, and user data, with the user ID as the primary key. FIG. 4 is a diagram showing the data structure of the user table 1012. As shown in FIG.
[0015] The user ID is an item for storing user identification information for identifying a user. The user identification information is an item for which a unique value is set for each user. The user name is an item for storing the name of the user. The user name may be set to any character string such as a nickname instead of a name. The user data includes information specific to each individual user and attribute information relating to the characteristics and background of the user. The user's unique information includes information unique to the user, such as the user's date of birth (age), sex, and the like. The user's attribute information includes information such as the user's educational history (highest level of education, major, graduation year), occupation, work history, interests, place of residence, and language.
[0016] The call table 1014 is a table for storing and managing information related to calls (call information). The call table 1014 is a table having a call ID as a primary key, and columns of call ID, user IDs, start time, end time, call duration, and call data. FIG. 5 is a diagram showing the data structure of the call table 1014. As shown in FIG.
[0017] The call ID is an item for storing call identification information for identifying a call. The call identification information is an item in which a unique value is set for each piece of call information. The user IDs is an item for storing user identification information of users who are the parties of the call voice. The user identification information of the customer and the operator is stored. In the case of a call in which three or more users participate in a specified room, the configuration may be such that three or more user identification information are stored. The start time is an item for storing the start time of the voice call. The end time is an item for storing the end time of the voice call. The call duration is an item for storing the call duration of the call voice. The call data is an item for storing voice data of a call.
[0018] The QA table 1015 is a table for storing and managing information relating to QA (QA information). The QA table 1015 is a table having columns for a first call ID, first question data, first answer data, and additional data. FIG. 6 is a diagram showing the data structure of the QA table 1015.
[0019] The first call ID is an item for storing call identification information for identifying a call. The first question data is an item that stores information about a question (first question) in a Q&A consisting of a question and answer pair. The first answer data is an item that stores information about an answer (first answer) in a Q&A consisting of a question and answer pair. The additional data is an item for storing additional information including background information on the background that led one party to make an inquiry to the other party in the call voice and problem information on the problem. Specifically, the additional data stores information on the background that led a customer or the like to make an inquiry to an operator and the problem that led to the question.
[0020] The FAQ table 1016 is a table for storing and managing information relating to FAQs (FAQ information). The FAQ table 1016 is a table having columns of second call IDs, second question data, and second answer data. FIG. 7 is a diagram showing the data structure of the FAQ table 1016.
[0021] The second call IDs is an item for storing call identification information for identifying a call. The second call IDs may be configured to store a plurality of pieces of call identification information. The second question data is an item that stores information about a question (second question) in a Q&A consisting of a question and answer pair. The second answer data is an item that stores information about an answer (second answer) in a Q&A consisting of a question and answer pair.
[0022] <Configuration of the control unit 104 of the server 10> The control unit 104 of the server 10 includes a user registration control unit 1041. The control unit 104 executes an application program 1011 stored in the storage unit 101, thereby realizing each functional unit.
[0023] The user registration control unit 1041 performs processing to store, in the user table 1012, information on users who wish to use the service according to the present disclosure. The information stored in the user table 1012 is generated when a user opens a web page operated by a service provider from any information processing terminal, enters information into a specific input form, and transmits the information to the server 10. The user registration control unit 1041 stores the received information in a new record in the user table 1012, completing the user registration. This allows the user stored in the user table 1012 to use the service. Before the user registration control unit 1041 registers user information in the user table 1012, the service provider may carry out a predetermined examination to restrict whether or not the user is permitted to use the service. The user ID may be any character string or number that can identify the user, any character string or number desired by the user, or an arbitrary character string or number may be automatically set by the user registration control unit 1041.
[0024] <Configuration of User Terminal 20> The user terminal 20 is an information processing device operated by a user who uses a service. The user terminal 20 may be, for example, a mobile terminal such as a smartphone or a tablet, a stationary PC (Personal Computer), or a laptop PC. It may also be a wearable terminal such as an HMD (Head Mount Display) or a wristwatch-type terminal. The user terminal 20 includes a storage unit 201 , a control unit 204 , an input device 206 , and an output device 208 .
[0025] <Configuration of the storage unit 201 of the user terminal 20> The storage unit 201 of the user terminal 20 includes a user ID 2011 and an application program 2012 .
[0026] The user ID 2011 is an account ID of the user. The user transmits the user ID 2011 from the user terminal 20 to the server 10. The server 10 identifies the user based on the user ID 2011 and provides the user with the service according to the present disclosure. The user ID 2011 includes information such as a session ID temporarily assigned by the server 10 when identifying the user who is using the user terminal 20.
[0027] The application program 2012 may be stored in advance in the storage unit 201, or may be configured to be downloaded from a web server operated by a service provider via a communication IF. The application programs 2012 include applications such as a web browser application. The application program 2012 includes an interpreted programming language such as JavaScript™ that runs on a web browser application stored on the user terminal 20 .
[0028] <Configuration of the control unit 204 of the user terminal 20> The control unit 204 of the user terminal 20 includes an input control unit 2041 and an output control unit 2042. The control unit 204 executes an application program 2012 stored in the storage unit 201, thereby realizing each functional unit.
[0029] <Configuration of the input device 206 of the user terminal 20> The input device 206 of the user terminal 20 includes a camera 2061 , a microphone 2062 , a position information sensor 2063 , a motion sensor 2064 , and a touch device 2065 .
[0030] <Configuration of the output device 208 of the user terminal 20> The output device 208 of the user terminal 20 includes a display 2081 and a speaker 2082 .
[0031] <Configuration of generated AI50> Generative AI50 refers to large-scale artificial intelligence models used in the field of natural language processing (NLP). These models learn from large amounts of text data (such as web pages, books, articles, etc.) to understand the patterns of the language used by humans and can effectively perform natural language generation (NLG) tasks. Generative AI50 is used in many NLP tasks, such as generating responses to specific questions, automatically generating articles, summarizing text, translation, sentiment analysis, etc. It can also be applied to various uses, such as education, entertainment, customer service, product development, etc. There are the following types of Generative AI50. In this disclosure, a large language model that mainly outputs text information as output information will be described as a type of Generative AI50. ·OpenAI ChatGPT ·Google Gemini ·Stable Diffusion ·midjourney Note that Generative AI50 may also be configured to be realized as a part of the functions of server 10.
[0032] <Operation of System 1> The following describes each process of System 1. FIG. 8 is a flowchart showing the operation of the FAQ creation process.
[0033] <FAQ Creation Process> The FAQ creation process is a process of creating FAQs (Frequently Asked Questions) based on call voice data regarding inquiries from operators at inquiry windows such as call centers by customers. An FAQ, also called "frequently asked questions," is information regarding pairs of questions frequently asked by customers or internal employees, etc., regarding products, services, or business content, and their answers. In this disclosure, QA refers to information consisting of pairs of questions and answers, and an FAQ refers to a collection of frequently asked QAs among QAs consisting of multiple pairs of questions and answers.
[0034] <Overview of FAQ Creation Process> The FAQ creation process is a series of processes that stores call data related to call voice, creates a plurality of QAs based on the call data, identifies a group consisting of a plurality of similar QAs among the created plurality of QAs, and creates an FAQ based on the group consisting of the plurality of QAs. The FAQ creation process may be configured to be executed each time a new record is registered in the call table 1014 or periodically at predetermined intervals.
[0035] <Details of the FAQ creation process> The details of the FAQ creation process will be described below.
[0036] <Call data storage step> In step S101, the control unit 104 of the server 10 stores the call voice data in the call table 1014 of the server 10. Specifically, in the present disclosure, an example will be described in which a telephone inquiry regarding a product, service, etc. is received from a customer at a corporate call center or the like, and data when an operator responds to the telephone inquiry is stored in the call data item of the call table 1014. Note that the call data does not necessarily need to be stored in the call table 1014 of the server 10, and may be configured to be stored in a server such as another external call service. Note that the call data may be stored via an electronic conference system or the like that enables a call between two or more parties other than a telephone by video, voice, or the like. Note that the storage of the call data does not necessarily need to be executed as part of the FAQ creation process. A configuration in which it is automatically stored in the call table 1014 according to the customer response service at the corporate call center is preferable.
[0037] In step S101, the control unit 104 of the server 10 executes a voice acquisition step of acquiring voice data in which the call voice is recorded. Specifically, the control unit 104 of the server 10 acquires a record of call information by searching the call table 1014. Note that the record of call information for which the FAQ creation process has been executed in the past may be excluded and acquired. Further, the control unit 104 of the server 10 may acquire all the records of call information stored in the call table 1014. Further, the control unit 104 of the server 10 may acquire some of the records of call information stored in the call table 1014. The control unit 104 of the server 10 may acquire the call information of the user related to the party of the call voice related to a specific user ID by searching the user IDs item of the call table 1014, or may acquire the call information of the users related to a plurality of different parties from the call table 1014.
[0038] <QA output step> In step S102, the control unit 104 of the server 10 executes a QA output step of outputting a plurality of QA information including a plurality of question sentences and a plurality of answer sentences for the question sentences based on the content of the call voice stored in the voice data acquired in the voice acquisition step. The QA output step may output a plurality of QA information based on the voice data related to one call voice. The QA output step may output a plurality of QA information based on a plurality of voice data related to a plurality of call voices. The QA output step may output a plurality of QA information based on a plurality of voice data related to a plurality of call voices with different parties of the call voice. Specifically, the control unit 104 of the server 10 executes a process of acquiring a plurality of QA information from one or more call data acquired in step S101.
[0039] The QA output step includes a step of creating a first prompt by including the content of the call voice stored in the voice data in a first instruction to output one or more QA information based on the content of the call voice, and a step of inputting the first prompt into a large language model to output a plurality of QA information. Specifically, the control unit 104 of the server 10 inputs the call data acquired in step S101 into a voice recognition system or the like to acquire a transcription text of the voice spoken in the call voice. The control unit 104 of the server 10 may be equipped with a voice recognition system or the like, or may use a voice recognition system or the like operated by an external information service provider. The control unit 104 of the server 10 creates a first prompt by inserting the transcribed text into the {transcribed text} portion of the first instruction shown below.
[0040] The QAA output step may generate a first prompt by including, in the first instruction, an eleventh instruction for deleting unnecessary speech content included in the content of the call voice. The QA output step may create a first prompt by including, in the first instruction, a twelfth instruction that does not delete any letters or numbers in the product model number. The QA output step may include a step of generating a first prompt by including in the first instruction a thirteenth instruction for outputting background information that led one party to make a query to the other party in the conversation voice, and a step of inputting the first prompt to a large-scale language model to output a plurality of pieces of QA information and background information related to the QA information in association with each other.
[0041] The QA output step may create a first prompt by including information about the parties of the call voice in the first instruction. Specifically, the control unit 104 of the server 10 acquires user data by searching the user ID field of the user table 1012 based on the user IDs included in the call information. The control unit 104 of the server 10 creates a first prompt by inserting the fields (metadata) of operator name, customer company name, and customer representative name into the {operator name}, {company_name}, and {customer_name} portions of the first instruction, respectively.
[0042] [1st instruction] #background You are a customer support expert. You will perform the following tasks to process the transcribed text of the conversation between the customer and the customer support representative: #task Please note the following: Remove unnecessary noise from the text and output it. (11th instruction) Do not delete letters or numbers such as part numbers. (Instruction No. 12) Each QA should include the background and the problem, and should be complete in itself. (13th Instruction) #Output specifications Please output all QAs contained in the transcript in the json format shown below. [ { q :$question1, #Question asked a :$answer1, #Answer to question bg:$background1, # Background information that led to the question (what you were trying to do, what caused the problem, and why you are asking the question) pb:$problem1, # The problem that led to the question (what is bothering you) }, { q :$question2, a :$answer2, bg:$background2, pb:$problem2, }, ... ] #Metadata · Operator name: {operator name} ·Customer name :{company_name} · Customer contact name: {customer_name} {transcript text}
[0043] In the present disclosure, the inclusion of the transcribed text in the first instruction has been disclosed as an example, but is not limited thereto. For example, if the generation AI 50 is a multimodal artificial intelligence system capable of accepting voice and video data, the call data may be sent to the generation AI 50 together with the first instruction without transcribing it using a voice recognition system. Also, in such a case, the control unit 104 of the server 10 includes in the creation of the first prompt by including the call data in the first instruction. In other words, the first prompt includes a data structure consisting of the first instruction and the call data.
[0044] The control unit 104 of the server 10 sends a request including the created first prompt to the API endpoint (URL) of the generation AI 50. The generation AI 50 outputs a response to the first prompt to the server 10. The response output by the generation AI 50 includes QA information. Specifically, the generation AI 50 outputs the first response described below. The control unit 104 of the server 10 extracts and accepts the QA from the response received from the generated AI 50. In this disclosure, the objects each consisting of one each of q, a, bg, and pb included in the first response are the QA. In other words, one or more QAs can be obtained from one call data.
[0045] [First response] [ { Q: Can I change my payment method after placing an order? A:You cannot change the payment method after placing an order. bg:I placed an order but want to change the payment method, PB:You cannot change the payment method after placing an order, }, { Q: Is it possible to cancel an order? A:You can cancel your order before it is shipped. If you wish to cancel, you will need your order number. bg:I placed an order but want to change the payment method, PB:You cannot change the payment method after placing an order, }, ... ]
[0046] The control unit 104 of the server 10 stores q in the first question data, a in the first answer data, and bg and p in the additional data field of the QA table 1015 of the server 10 for each of q, a, bg, and pb of the QA acquired from the generation AI 50. In addition, the call ID of the call information from which the QA is extracted is stored in the first call ID field. As a result, the first question data, first answer data, and first additional data are stored in the QA table in association with each other for each piece of call information (first call ID). In the present disclosure, the QA creation step in step S102 may be executed collectively for one or more (part or all) call information records. After execution of this process, the grouping step in step S103 may be executed.
[0047] <Grouping step> In step S103, the control unit 104 of the server 10 executes a QA specifying step of specifying a predetermined QA group consisting of multiple similar QA information from the multiple QA information output in the QA output step. The QA specifying step executes a step of specifying a QA group consisting of multiple QA information whose QA information and background information associated with the QA information are similar. Specifically, the control unit 104 of the server 10 performs clustering processing based on the similarity of at least one of the first question data, the first answer data, the background information contained in the additional data, and the problem information contained in the multiple QA information acquired in step S102. The control unit 104 of the server 10 groups multiple pieces of QA information stored in the QA table 1015 based on the similarity. The control unit 104 of the server 10 may calculate the similarity for a text combining the first question data and the first answer data. The control unit 104 of the server 10 preferably calculates the similarity of the text combining the first question data, the first answer data, and the background information included in the additional data. This makes it possible to group similar Q&A information together, taking into consideration the background that led a customer or the like to make an inquiry to an operator.
[0048] In step S103, the QA identification step calculates the similarity between a specified QA information and each of the multiple QA information output in the QA output step, and executes a step of identifying multiple QA information included in a specified QA group based on the similarity. Specifically, the control unit 104 of the server 10 calculates the similarity between text information included in multiple pieces of QA information stored in the QA table 1015. Specifically, the control unit 104 of the server 10 calculates the similarity between texts based on any algorithm such as cosine similarity, Jaccard similarity, inverse Euclidean distance, inverse Mahalanobis distance, Dice coefficient, inverse word movement distance, inverse edit distance, BERT similarity, etc. The control unit 104 of the server 10 can cluster (group) multiple pieces of QA information by applying any unsupervised clustering technique such as k-means, supervised clustering technique, etc. to the calculated similarity. Furthermore, the control unit 104 of the server 10 may verify, using a large-scale language model, whether questions and answers are similar for the clustering result (grouping result) based on the calculated similarity. Specifically, the control unit 104 may determine whether groups based on a plurality of pieces of QA information are similar based on the output contents output by inputting a prompt generated by including a plurality of pieces of QA information included in a predetermined clustering in an instruction for determining whether a plurality of pieces of QA information are similar or not. If it is determined that the images are not similar, the clustering process may be executed again, or a grouping process using a large-scale language model, which will be described later, may be executed.
[0049] In step S103, the QA identification step creates a third prompt by including predetermined QA information in an instruction to extract QA information similar to the predetermined QA information from a plurality of QA information, and based on the output information output by inputting the third prompt into a large language model, executes a step of identifying a plurality of QA information. Specifically, the control unit 104 of the server 10 may create a group of similar QA information by including at least one or more of the first question data, the first answer data, the background information included in the additional data, and the question information of the plurality of QA information in an instruction to group the plurality of QA information. For example, the QA information can be grouped by the following instruction text. In the locations of {QA information 1}, {QA information 2}, and {QA information 3} in the following instruction text, a prompt can be created by inserting at least one or more of the first question data, the first answer data, the background information included in the additional data, and the question information of the QA information, respectively. The control unit 104 of the server 10 sends a request including the created prompt to the API endpoint (URL) of the generation AI 50. The generation AI 50 outputs a response to the prompt to the server 10. The control unit 104 of the server 10 identifies the group based on the ID of the QA information included in each group included in the acquired response.
[0050] 〔Instruction text〕 Please group the following texts that are similar to each other. Specifically, for each group, output the IDs included in that group. ID1: {QA information 1} ID2: {QA information 2} ID3: {QA information 3} #Output format { Group 1: [IDA, IDB ···], Group 2: [IDC, IDD ···], }
[0051] <FAQ acquisition step> In step S104, the control unit 104 of the server 10 creates a second prompt by including the QA group identified in the QA identification step in a second instruction for outputting FAQ information formatted from multiple pieces of QA information, and executes an FAQ acquisition step for acquiring FAQ information consisting of a question sentence and an answer sentence to the question sentence, based on the output information output by inputting the second prompt into a large-scale language model. Specifically, the control unit 104 of the server 10 creates a second prompt by including at least one of the first question data, the first answer data, the background information contained in the additional data, and the problem information of the multiple QA information for each group created in step S103 in the {QA group} section of the second instruction shown below.
[0052] The FAQ acquisition step may create a second prompt by including in the second instructions a 21st instruction for outputting an FAQ formatted taking into account background information, and a plurality of pieces of background information associated with each of the plurality of QA data included in the QA group.
[0053] [Second instruction] The contents of the following Q&A will be reorganized and output to avoid duplication. Also, please note the following points: -Create FAQs for each question included in the data -Content that can be considered individual circumstances regarding accounts or product purchases will not be included in the FAQ. - Individual contracting party names, contracting organization names, dates, etc. will be generalized. Please make the content complete so that the reader can understand the background by reading one Q&A. -Please create FAQs taking into consideration the background information contained in the QA. (21st Instruction) #Output specifications Please output all FAQs contained in the transcript in the json format shown below. [ { q :$question1, #Question asked a :$answer1, #Answer to question }, { q :$question2, a :$answer2, }, ... ] #QA List {QA Group}
[0054] The following character strings related to multiple QA data are inserted in the {QA group} section. It is not necessary to include all of the questions, answers, backgrounds, and problems; it is sufficient to include some of them. It is preferable to include at least the questions and answers. ##QA1 Question: $question1 Answer:$answer1 Background: $background1 Problem: $problem1 ##QA2 Question: $question2 Answer:$answer2 Background: $background2 Problem: $problem2
[0055] The second instruction in the FAQ acquisition step may be an instruction to output one formatted FAQ from a plurality of pieces of QA information. Specifically, the second instruction may include an instruction to output one FAQ data for the predetermined QA group identified in step S103. This makes it possible to create one FAQ of a predetermined high quality from multiple similar QAs.
[0056] The control unit 104 of the server 10 stores q as the second question data and a as the second answer data of a new record in the FAQ table 1016 of the server 10 for each of q and a of one or more FAQs acquired from the generation AI 50. This allows the creation of suitable FAQs based on the call voice, without the burden of manual work.
[0057] <Basic computer hardware configuration> 9 is a block diagram showing the basic hardware configuration of a computer 90. The computer 90 includes at least a processor 901, a main memory device 902, an auxiliary memory device 903, and a communication IF 991 (interface). These are electrically connected to each other by a communication bus 921.
[0058] The processor 901 is hardware for executing an instruction set described in a program, and is composed of an arithmetic unit, a register, a peripheral circuit, and the like.
[0059] The main memory device 902 is for temporarily storing programs, data to be processed by the programs, etc. For example, it is a volatile memory such as a DRAM (Dynamic Random Access Memory).
[0060] The auxiliary storage device 903 is a storage device for saving data and programs, such as a flash memory, a hard disk drive (HDD), a magneto-optical disk, a CD-ROM, a DVD-ROM, or a semiconductor memory.
[0061] The communication IF 991 is an interface for inputting and outputting signals for communicating with other computers via a network using a wired or wireless communication standard. The network is composed of the Internet, a LAN, various mobile communication systems constructed by wireless base stations, etc. For example, the network includes 3G, 4G, 5G mobile communication systems, LTE (Long Term Evolution), wireless networks that can connect to the Internet via a specified access point (e.g., Wi-Fi (registered trademark)), etc. In the case of wireless connection, communication protocols include, for example, Z-Wave (registered trademark), ZigBee (registered trademark), Bluetooth (registered trademark), etc. In the case of wired connection, the network also includes a network that is directly connected by a USB (Universal Serial Bus) cable or the like.
[0062] It should be noted that the computer 90 can be virtually realized by distributing all or part of each hardware configuration among multiple computers 90 and connecting them together via a network. In this way, the computer 90 is a concept that includes not only a computer 90 housed in a single housing or case, but also a virtualized computer system.
[0063] <Basic functional configuration of computer 90> A description will now be given of the functional configuration of a computer realized by the basic hardware configuration (FIG. 9) of a computer 90. The computer comprises at least the functional units of a control unit, a storage unit, and a communication unit.
[0064] The functional units of the computer 90 can also be realized by distributing all or part of the functional units among multiple computers 90 connected to each other via a network. The computer 90 is a concept that includes not only a single computer 90 but also a virtualized computer system.
[0065] The control unit is realized by the processor 901 reading out various programs stored in the auxiliary storage device 903, expanding the programs in the main storage device 902, and executing processes according to the programs. The control unit can realize functional units that perform various information processing depending on the type of program. In this way, the computer is realized as an information processing device that performs information processing.
[0066] The functions performed by the components described herein may be implemented in circuitry or processing circuitry, including general purpose processors, application specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), a CPU (a Central Processing Unit), conventional circuits, and / or combinations thereof, programmed to perform the functions described. Processors include transistors and other circuits and are considered to be circuitry or processing circuitry. A processor may be a programmed processor that executes a program stored in a memory. In this specification, a circuitry, unit, or means is hardware that is programmed to realize or performs the described functions, which may be any hardware disclosed in this specification or any hardware known to be programmed to realize or perform the described functions. If the hardware is a processor considered to be a type of circuitry, the circuitry, means, or unit is a combination of the hardware and software used to configure the hardware and / or processor.
[0067] The storage unit is realized by a main storage device 902 and an auxiliary storage device 903. The storage unit stores data, various programs, and various databases. Furthermore, the processor 901 can secure a storage area corresponding to the storage unit in the main storage device 902 or the auxiliary storage device 903 in accordance with a program. Furthermore, the control unit can cause the processor 901 to execute processes of adding, updating, and deleting data stored in the storage unit in accordance with the various programs.
[0068] The term database refers to a relational database, which is used to manage sets of data called masters and tables in a tabular format structurally defined by rows and columns, by associating them with each other. In a database, a table is called a table or master, a column in a table is called a column, and a row in a table is called a record. In a relational database, relationships between tables and masters can be set and associated. Usually, a column that serves as a primary key for uniquely identifying a record is set in each table and each master, but setting a primary key in a column is not essential. The control unit can cause the processor 901 to add, delete, or update records in a specific table or master stored in the storage unit according to various programs. Furthermore, by storing data, various programs, and various databases in the storage unit, it can be considered that the information processing device and information processing system according to the present disclosure have been manufactured.
[0069] In addition, the database and master in this disclosure may include any data structure (such as a list, a dictionary, an associative array, or an object) in which information is structurally defined. The data structure also includes data that can be considered as a data structure by combining data with a function, class, method, or the like written in any programming language.
[0070] The communication unit is realized by the communication IF 991. The communication unit realizes a function of communicating with other computers 90 via a network. The communication unit can receive information transmitted from other computers 90 and input the information to the control unit. The control unit can cause the processor 901 to execute information processing on the received information in accordance with various programs. In addition, the communication unit can transmit information output from the control unit to other computers 90.
[0071] <Additional Notes> The matters described in the above embodiments will be supplemented below.
[0072] (Appendix 1) A program to be executed by a computer having a processor and a storage unit, the program executing: a voice acquisition step (S101) of acquiring voice data in which a telephone call voice is recorded; a QA output step (S102) of outputting a plurality of QA information items consisting of a plurality of questions and a plurality of answers to the questions based on the content of the telephone call voice stored in the voice data acquired in the voice acquisition step; a QA identification step (S103) of identifying a specific QA group consisting of a plurality of similar QA information items from the plurality of QA information items output in the QA output step; and an FAQ acquisition step (S104) of creating a second prompt by including the QA group identified in the QA identification step in a second instruction for outputting FAQ information formatted from the plurality of QA information items, and acquiring FAQ information consisting of questions and answers to the questions based on output information output by inputting the second prompt into a large-scale language model. This allows the creation of suitable FAQs based on the call voice, without the burden of manual work.
[0073] (Appendix 2) The program according to claim 1, wherein the QA output step (S102) is a step of outputting a plurality of pieces of QA information based on voice data relating to one call voice. This makes it possible to extract multiple QAs from a single voice call and create suitable FAQs based on QA groups consisting of similar QAs. Multifaceted QAs can be extracted from a single voice call, and suitable FAQs can be created by integrating QA groups consisting of similar QAs from the multiple QAs.
[0074] (Appendix 3) The program according to claim 1, wherein the QA output step (S102) is a step of outputting a plurality of pieces of QA information based on a plurality of pieces of voice data relating to a plurality of call voices. This makes it possible to extract multiple QAs from multiple call voices and create suitable FAQs based on QA groups consisting of similar QAs. It is possible to create suitable FAQs by integrating QA groups consisting of similar QAs from multiple QAs extracted from each of the multiple call voices. For example, a higher quality FAQ can be created by taking into account the partial content of a given call.
[0075] (Appendix 4) The program according to appendix 3, wherein the QA output step (S102) is a step of outputting a plurality of pieces of QA information based on a plurality of pieces of voice data relating to a plurality of call voices of different parties of the call voice. This makes it possible to extract multiple QAs from multiple call audio with different parties (speakers) and create suitable FAQs based on QA groups consisting of similar QAs. It is possible to create FAQs that integrate information across the contents of calls between multiple parties. By integrating inquiries from various customers and answers from various operators, it is possible to create high-quality, suitable FAQs.
[0076] (Appendix 5) The program according to claim 1, wherein the second instruction in the FAQ acquisition step is an instruction to output one formatted FAQ from multiple pieces of QA information. This allows you to combine QA groups consisting of multiple similar QAs to create a single, high-quality, and optimal FAQ.
[0077] (Appendix 6) A program described in any one of Appendices 1 to 5, wherein the QA output step (S102) includes a step of creating a first prompt by including the content of the call voice stored in the voice data in a first instruction to output one or more pieces of QA information based on the content of the call voice, and a step of outputting multiple pieces of QA information by inputting the first prompt into a large-scale language model. This makes it possible to output multiple pieces of QA information from audio data in which the voice of a call is recorded.
[0078] (Appendix 7) The program described in Appendix 6, wherein the QA output step (S102) includes a step of creating a first prompt by including, in the first instruction, an eleventh instruction for deleting unnecessary speech content contained in the content of the call voice. This makes it possible to output higher quality QA information.
[0079] (Appendix 8) The program of claim 6, wherein the QA output step (S102) includes a step of creating a first prompt by including in the first instruction a twelfth instruction that does not delete any letters or numbers in the product model number. This makes it possible to output higher quality QA information.
[0080] (Appendix 9) The program according to claim 6, wherein the QA output step (S102) includes a step of creating a first prompt by including information about the parties of the call voice in the first instruction. This makes it possible to output higher quality QA information by taking into consideration the parties in the call voice (the names and titles of the customer, operator, etc.).
[0081] (Appendix 10) The program described in Appendix 6, wherein the QA output step (S102) includes a step of creating a first prompt by including in the first instruction a thirteenth instruction for outputting background information that led one party to make a query to the other party in the conversation voice, and a step of inputting the first prompt into a large-scale language model to associate and output multiple pieces of QA information and background information related to the QA information. This makes it possible to output the background that led a customer or other person to make an inquiry to an operator together with the QA information.
[0082] (Appendix 11) The program according to claim 10, wherein the QA identification step (S103) is a step of identifying a QA group consisting of multiple QA information pieces having similar QA information and background information associated with the QA information pieces. This allows the system to create suitable FAQs based on QA groups consisting of QAs with similar backgrounds, taking into consideration the background that led a customer or the like to contact an operator.
[0083] (Appendix 12) The program according to appendix 10 or 11, wherein the FAQ acquisition step (S104) includes a step of creating a second prompt by including in the second instruction a 21st instruction for outputting an FAQ formatted taking into account background information, and a plurality of pieces of background information associated with each of the plurality of QA data included in the QA group. This makes it possible to output higher quality FAQ information that takes into account the circumstances that led a customer or other person to make an inquiry to an operator. For example, it is possible to create FAQs that are made up of questions and answers that take into account the background, making it possible to create FAQs that are more convenient for users by taking into account the background circumstances.
[0084] (Appendix 13) A program described in any one of Appendices 1 to 12, wherein the QA identification step (S103) is a step of calculating a similarity between a specified QA information and each of multiple QA information output in the QA output step, and identifying multiple QA information included in a specified QA group based on the similarity. This makes it possible to identify similar QA information at low cost without using a large-scale language model.
[0085] (Appendix 14) A program described in any one of Appendices 1 to 12, wherein the QA identification step (S103) is a step of creating a third prompt by including specified QA information in instructions for extracting QA information similar to the specified QA information from multiple QA information, and identifying multiple QA information based on output information output by inputting the third prompt into a large-scale language model. This allows us to use large-scale language models to identify similar QA information of higher quality.
[0086] (Appendix 15) A method implemented on a computer having a processor and a memory, the processor performing all the steps performed in the invention according to any one of claims 1 to 14. This allows the creation of suitable FAQs based on the call voice, without the burden of manual work.
[0087] (Appendix 16) An information processing device comprising a control unit and a memory unit, wherein the control unit executes all of the steps executed in the invention according to any one of Supplementary Note 1 to Supplementary Note 14. This allows the creation of suitable FAQs based on the call voice, without the burden of manual work.
[0088] (Appendix 17) A system comprising means for performing all the steps performed in any one of claims 1 to 14. This allows the creation of suitable FAQs based on the call voice, without the burden of manual work. [Explanation of symbols]
[0089] 1 system, 10 server, 101 memory unit, 104 control unit, 106 input device, 108 output device, 20 user terminal, 201 memory unit, 204 control unit, 206 input device, 208 output device, 50 generation AI, 501 memory unit, 504 control unit, 506 input device, 508 output device
Claims
1. A program to be executed by a computer having a processor and a storage unit, a voice acquisition step of acquiring voice data in which a call voice is recorded and attribute information of a party of the call voice; a QA output step of outputting a plurality of QA information pieces each including a plurality of questions and a plurality of answers to the questions, for each of the plurality of voice data related to the call voice acquired in the voice acquisition step; a QA specifying step of specifying a predetermined QA group including a plurality of QA information items having similar QA information items and background information associated with the QA information items among the plurality of QA information items output in the QA output step; a FAQ acquisition step of generating a second prompt by including the QA group identified in the QA identification step in a second instruction for outputting FAQ information shaped from a plurality of QA information, and acquiring FAQ information consisting of a question sentence and an answer sentence to the question sentence based on output information outputted by inputting the second prompt into a large-scale language model; Run the command, The QA output step includes: A step of creating a first prompt by including in a first instruction for outputting one or more pieces of QA information based on the content of the telephone call voice a thirteenth instruction for outputting background information that led one party to make an inquiry to the other party in the telephone call voice, the content of the telephone call voice stored in the voice data, and attribute information of the parties of the telephone call voice; a step of inputting the first prompt into a large-scale language model to output the plurality of QA information and background information related to the QA information in association with each other; Including, program.
2. The QA output step is a step of outputting a plurality of pieces of QA information based on a plurality of pieces of voice data related to a plurality of call voices of different parties of the call voice. The program according to claim 1.
3. A program to be executed by a computer having a processor and a storage unit, A voice acquisition step of acquiring voice data in which a call voice is recorded; a QA output step of outputting a plurality of QA information pieces each including a plurality of questions and a plurality of answers to the questions, for each of the plurality of voice data related to the call voice acquired in the voice acquisition step; a QA specifying step of specifying a predetermined QA group including a plurality of QA information items having similar QA information items and background information associated with the QA information items among the plurality of QA information items output in the QA output step; a FAQ acquisition step of generating a second prompt by including the QA group identified in the QA identification step in a second instruction for outputting FAQ information shaped from a plurality of QA information, and acquiring FAQ information consisting of a question sentence and an answer sentence to the question sentence based on output information outputted by inputting the second prompt into a large-scale language model; Run the command, The QA output step includes: creating a first prompt by including in a first instruction for outputting one or more pieces of QA information based on the content of the call voice a thirteenth instruction for outputting the content of the call voice stored in the voice data and background information that led one party to make a query to the other party in the call voice; a step of inputting the first prompt into a large-scale language model to output the plurality of QA information and background information related to the QA information in association with each other; Including, program.
4. The second instruction in the FAQ acquisition step is an instruction to output one formatted FAQ from a plurality of pieces of QA information.
3. The program according to claim 1 or 2.
5. the QA output step includes a step of generating the first prompt by including, in the first instruction, an eleventh instruction for deleting unnecessary speech content included in the content of the call voice; 3. The program according to claim 1 or 2.
6. The QA output step includes a step of creating the first prompt by including, in the first instruction, a twelfth instruction not to delete alphabets or numbers in a model number of a product.
3. The program according to claim 1 or 2.
7. A program for execution by a computer having a processor and a storage unit, a voice acquisition step of acquiring voice data in which a call voice is recorded and attribute information of a party of the call voice; a QA output step of outputting a plurality of QA information pieces each including a plurality of questions and a plurality of answers to the questions, for each of the plurality of voice data related to the call voice acquired in the voice acquisition step; a QA specifying step of specifying a predetermined QA group consisting of a plurality of similar QA information from among the plurality of QA information output in the QA output step; a 21st instruction to output an FAQ formatted in consideration of background information in a second instruction to output FAQ information formatted from a plurality of QA information, and a FAQ acquisition step to create a second prompt by including the QA group identified in the QA identification step and a plurality of pieces of background information associated with each of the plurality of QA data included in the QA group, and to acquire FAQ information consisting of a question sentence and an answer sentence to the question sentence based on output information outputted by inputting the second prompt into a large-scale language model; Run the command, The QA output step includes: A step of creating a first prompt by including in a first instruction for outputting one or more pieces of QA information based on the content of the telephone call voice a thirteenth instruction for outputting background information that led one party to make an inquiry to the other party in the telephone call voice, the content of the telephone call voice stored in the voice data, and attribute information of the parties of the telephone call voice; a step of inputting the first prompt into a large-scale language model to output the plurality of QA information and background information related to the QA information in association with each other; Including, program.
8. The QA specifying step is a step of calculating a similarity between each of the plurality of QA information outputted in the QA output step and a predetermined QA information, and specifying the plurality of QA information included in the predetermined QA group based on the similarity.
3. The program according to claim 1 or 2.
9. The QA specifying step is a step of creating a third prompt by including predetermined QA information in an instruction for extracting QA information similar to the predetermined QA information from a plurality of pieces of QA information, and specifying the plurality of pieces of QA information based on output information output by inputting the third prompt into a large-scale language model.
3. The program according to claim 1 or 2.
10. A method implemented on a computer having a processor and a memory, the processor executing all of the steps performed in any one of claims 1 to 3.
11. 4. An information processing apparatus comprising a control unit and a storage unit, the control unit executing all of the steps executed in the invention according to any one of claims 1 to 3.
12. A system comprising means for executing all the steps performed in any one of claims 1 to 3.
Citation Information
Patent Citations
Information processing device, information processing method, and information processing program
JP2022126428A
FAQ providing apparatus and FAQ providing program
JP2021157569A