Program, information processing method, and information processing device
A program using a language model to convert audio into structured text data addresses the challenge of quick report and record creation, enhancing accuracy and efficiency.
Patent Information
- Application Number
- JP2025094105
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-11-12
AI Technical Summary
Reports and records often need to be completed quickly, leading to errors and reduced efficiency due to the time pressure.
A program that uses a language model to generate output text data from audio input, including a date and situation, facilitating easy and accurate creation of reports and records.
Enables faster and more accurate generation of reports and records by converting audio into structured text data, reducing errors and improving efficiency.
Smart Images

Figure 0007756274000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a program, an information processing method, and an information processing device. [Background technology]
[0002] In various business operations, reports and minutes are prepared for the purpose of managing progress, schedules, etc. Patent Document 1 discloses a program for generating a summary text that summarizes text information contained in one or more sections of speech data. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2023-169093 Summary of the Invention [Problem to be solved by the invention]
[0004] Reports, minutes, and other record-keeping documents for various tasks (hereinafter referred to as "records, etc.") often need to be completed and submitted in a short time between tasks, placing a burden on the person in charge of preparing them. In addition, the need to submit them in a short time tends to make mistakes, and the correction work involved can further reduce work efficiency.
[0005] One aspect of the present disclosure provides a program or the like that enables reports, minutes, and other records to be created more easily and accurately. [Means for solving the problem]
[0006] A program according to one aspect of the present disclosure causes a computer to execute a process of receiving audio related to an event after the event has ended, providing audio information based on the audio to a language model, and generating output text data consisting of multiple items including a date and a situation. [Effects of the Invention]
[0007] According to a program or the like according to one aspect of the present disclosure, reports, minutes, and other records can be created more easily and accurately. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a schematic diagram of a record creation system according to a first embodiment. [Figure 2] FIG. 2 is a block diagram illustrating an example of the configuration of a server device. [Figure 3] FIG. 2 is a block diagram showing an example of the configuration of a generation server device. [Figure 4] FIG. 2 is a block diagram illustrating an example of the configuration of a terminal device. [Figure 5] 10 is a flowchart illustrating an example of a record document creation process. [Figure 6] FIG. 10 is a diagram illustrating an example of an item definition portion included in a prompt. [Figure 7] FIG. 10 is a diagram showing an example of output text data consisting of a plurality of items. [Figure 8] FIG. 10 is a diagram illustrating an example of a registration information type selection screen. [Figure 9] FIG. 10 is a diagram illustrating an example of an input screen. [Figure 10] FIG. 10 is a diagram showing an example of a converted text display screen. [Figure 11] FIG. 10 is a diagram illustrating an example of an output text data list screen. [Figure 12] FIG. 10 is a diagram illustrating an example of an output result screen. [Figure 13] FIG. 10 is a diagram showing an example of an editing screen on which all text corresponding to one item is displayed. [Figure 14]FIG. 10 is a diagram illustrating an example of a system configuration without a generation server device. [Figure 15] FIG. 10 is a schematic diagram of a record creation system according to a second embodiment. [Figure 16] FIG. 10 is a block diagram showing an example of the configuration of a terminal device according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, examples of a program, an information processing method, and an information processing device according to one aspect of the present disclosure will be described in detail with reference to the drawings. In the description, like elements will be given like reference numerals, and duplicated descriptions will be omitted as appropriate.
[0010] [Embodiment 1] Fig. 1 is a schematic diagram of a record creation system according to embodiment 1. As shown in Fig. 1, the record creation system includes an information processing device 1, an information processing device 2, and a generation server device 4, and each device transmits and receives information via a network 3 such as the Internet. The generation server device 4 may be communicatively connected to a language model 5, such as a large language model (LLM), stored inside or outside the generation server device 4, so that the language model 5 can be used.
[0011] The information processing device 1 is an information processing device that processes, stores, and transmits / receives various types of information. The information processing device 1 is, for example, a server device, a personal computer, a tablet terminal, a smartphone, or other information processing device. The information processing device 1 may be configured as one of a plurality of virtual devices (virtual machines) configured within one information processing device. In this embodiment, to avoid complicating the explanation, the information processing device 1 will be described as a server device 1. However, the information processing device 1 is not limited to a server device, and may be configured as any of the above-mentioned types of information processing devices.
[0012] The information processing device 2 receives information output from the server device 1 via the network 3, accepts operations by a user of the information processing device 2, and transmits information related to the operations to the server device 1 via the network 3. The information processing device 2 is an information processing device such as a personal computer, a smartphone, a mobile phone, a wearable device, or a tablet. In this embodiment, to avoid complicating the explanation, the information processing device 2 will be referred to as a terminal device 2.
[0013] 2 is a block diagram showing an example of the configuration of the server device 1. The server device 1 may be an information processing device configured mainly with electronic circuits using semiconductor circuit elements. The server device 1 has a control unit 11, a communication unit 12, a reading unit 13, and a storage unit 14. The control unit 11, the communication unit 12, the reading unit 13, and the storage unit 14 are connected to each other so as to be able to communicate with each other via a bus 19 or the like. Note that the server device 1 may have another configuration, such as one that does not have the reading unit 13.
[0014] The control unit 11 includes a CPU (Central Processing Unit), an MPU (Micro-Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array The control unit 11 may be configured to have one or more of a processing device such as a digital signal processor (DSP), a digital signal processor (DSP), a quantum processor, etc. The control unit 11 may be configured to read out and execute a program (or program product) P1 stored in the storage unit 14.
[0015] The communication unit 12 is a communication module for performing communication-related processing, and can send and receive information to and from the terminal device 2, etc. via the network 3. The reading unit 13 can read portable storage media 1a such as CD (Compact Disc)-ROMs, DVD (Digital Versatile Disc)-ROMs, and USB (registered trademark) memories. The control unit 11 can read the program P1 and / or data from the portable storage medium 1a via the reading unit 13 and store them in the storage unit 14. The control unit 11 can also download the program P1 from another computer via the network 3, etc., and store them in the storage unit 14.
[0016] The storage unit 14 may include a volatile storage unit such as a random access memory (RAM), and a non-volatile storage unit such as a read only memory (ROM), a hard disk drive (HDD), and a flash memory. The control unit 11 can temporarily store in the volatile storage unit programs and / or data read from the non-volatile storage unit or received from the communication unit 12 for use. The storage unit 14 can store data for a prompt P2 (described later) in addition to a program (or program product) P1 executed by the control unit 11. The storage unit 14 can store multiple prompts P2.
[0017] FIG. 3 is a block diagram showing an example configuration of the generation server device 4. The generation server device 4 can be an information processing device mainly configured with electronic circuits using semiconductor circuit elements. The generation server device 4 has a control unit 41, a communication unit 42, a reading unit 43, and a storage unit 44. The control unit 41, the communication unit 42, the reading unit 43, and the storage unit 44 are communicatively connected to each other via a bus 49 or the like. These components can be configured similarly to the information processing device (server device) 1 of FIG. 2, except that the program P1 and the prompt P2 are not stored in the storage unit 44. Therefore, redundant explanations will be omitted. Note that the generation server device 4 may have a different configuration, such as not having a reading unit 43 that reads the portable storage medium 4a, as in the case of the server device 1.
[0018] The generation server device 4 is an information processing device that handles a language model 5 such as a Generative Pre-trained Transformer (GPT, registered trademark), a Bidirectional Encoder Representations from Transformer (BERT), a Large Language Model (LLM), or a similar large-scale language model such as Gemini (registered trademark). The language model 5 may be stored in a storage unit 44 within the generation server device 4, or may be stored in a storage or the like that is communicatively connected to the generation server device 4.
[0019] The generation server device 4 inputs input data such as images, audio, and text into a language model 5 to generate a response sentence. The language model 5 is a trained machine learning model, and can be, for example, a large-scale language model such as GPT (Generative Pre-trained Transformer) or BERT (Bidirectional Encoder Representations from Transformer), but may also be another language model. The server device 1 may be communicatively connected to the generation server device 4 via a network 3, or may be communicatively connected to the generation server device 4 directly or via a local network or the like without via the network 3.
[0020] 4 is a block diagram showing an example configuration of the terminal device 2. Here, the terminal device 2 may be an information processing device configured mainly with electronic circuits using semiconductor circuit elements. The terminal device 2 may have a control unit 21, a communication unit 22, a reading unit 23, a storage unit 24, a display unit 26, an input unit 27, and a microphone 28. The control unit 21, the communication unit 22, the reading unit 23, the storage unit 24, the display unit 26, the input unit 27, and the microphone 28 may be connected to each other via a bus 29 or the like so as to be able to communicate with each other.
[0021] The control unit 21 includes a CPU (Central Processing Unit), an MPU (Micro-Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array The control unit 21 may be configured to have one or more of a processing device such as a digital signal processor (DSP), a digital signal processor (DSP), a quantum processor, etc. The control unit 21 may be configured to read and execute a program (or a program product) stored in a storage unit 24 described later.
[0022] The communication unit 22 is a communication module for performing communication-related processing, and can send and receive information to and from the server device 1, etc., via the network 3. The reading unit 23 can read portable storage media 2a, such as CD (Compact Disc)-ROMs, DVD (Digital Versatile Disc)-ROMs, and USB (registered trademark) memories. The control unit 21 can read programs and / or data from the portable storage media 2a via the reading unit 23 and store them in the storage unit 24. The control unit 21 can also download programs from other computers via the network 3, etc., and store them in the storage unit 24.
[0023] The storage unit 24 may include a volatile storage device such as a random access memory (RAM), and a non-volatile storage device such as a read only memory (ROM), a hard disk drive (HDD), and a flash memory. The control unit 21 may temporarily store the programs and / or data read from the non-volatile storage device or received from the communication unit 22 in the volatile storage device so that the programs and / or data can be read and written at high speed by the control unit 21 for use.
[0024] The display unit 26 may be a liquid crystal display, an organic EL (Electro Luminescence) display, or the like. The input unit 27 may be an input device such as a keyboard, a mouse, a touch panel, or a camera. The microphone 28 converts sounds around the terminal device 2 into digital data based on instructions from the control unit 21. The control unit 21 stores the converted digital data in the storage unit 24.
[0025] 5 is a flowchart showing an example of a record document creation process. The server device 1 executes the record document creation process based on the program P1, etc. As shown in this flowchart, the server device 1 displays a registration information type selection screen 100 (FIG. 8) for selecting a registration information type (step S101). The registration information type is the type of content of the record document that the user of the terminal device 2 intends to record after the event ends.
[0026] Here, an "event" can be a task such as a meeting, an interview, or equipment maintenance. The registered information type can be, for example, a report on maintenance work, construction work, or other tasks, or a type of record document such as minutes of a weekly or monthly meeting or other meeting. An "event" can also be the need to order parts, reserve equipment, or make an inquiry. Each registered information type can include common items such as date and time and status, and can also include specific items such as a selectively determined department name, a meter value that requires a numerical value to be entered, and the presence or absence of a malfunction.
[0027] Prompt P2, which specifies such items, will be described in detail later with reference to FIG. 6. As will be described later, the input for recording can be "event-related voice" from the user of terminal device 2. "Event-related voice" can be voice with content equivalent to minutes of a meeting or interview, voice with content equivalent to a report on maintenance or other work, voice with content equivalent to ordering parts, reserving equipment, and inquiries, or other voice. Here, "voice" can include not only the voice itself, but also voice data obtained by digitally converting voice.
[0028] The server device 1 executing the program P1 can, for example, display on the display unit 26 of the terminal device 2 a plurality of selectable registration information types from a plurality of registration information types previously stored in the storage unit 14. Furthermore, the user of the terminal device 2 may select one of the displayed registration information types, and the server device 1 may then accept the selected registration information type.
[0029] The server device 1 determines whether or not a selection of a registration information type has been accepted on the registration information type selection screen 100 (step S102). If a registration information type has not been selected in step S102 (step S102: NO), the process of step S102 is repeated. If a registration information type has been selected (step S102: YES), a prompt P2 corresponding to the registration information type is selected (step S103).
[0030] The prompt P2 may be selected from a plurality of registration information types, such as "visit memo," "maintenance record," "order data registration," "tentative reservation request," and "inquiry" shown in Fig. 8. Here, one or more prompts P2 corresponding to each registration information type may be stored in the storage unit 14 of the server device 1, as shown in Fig. 2. The server device 1 can obtain the prompt P2 corresponding to the selected registration information type from the storage unit 14 based on a command from the program P1.
[0031] When prompt P2 is selected, the server device 1, based on the instructions of program P1, sends input screen 200 (form) (Figure 9) based on the content of the selected prompt P2 to the terminal device 2 and displays it on the display unit 26 of the terminal device 2 (step S104).
[0032] The terminal device 2 on which the input screen 200 is displayed determines whether or not an instruction to start voice input has been given (step S105). If an instruction to start voice input is not detected (step S105: NO), the process of step S105 is repeated. If an instruction to start voice input is detected (step S105: YES), the terminal device 2 detects voice via the microphone 28 and converts it into voice data (step S106). The server device 1 receives the voice data via the network 3. Here, if an operation has been performed on the terminal device 2 by voice input at the stage of selecting the registration information type in step S102 or before that, the process of step S105 may not be performed.
[0033] The input voice may be, for example, a reading of the following sentence: "Today I had an appointment with Mr. X, so I visited ABC Co., Ltd. and addressed President Y, where I was able to meet with President Y and Executive Director Z. Since it had only been about a week since the last proposal, I was worried about whether it would be easy to go ahead with the contract, but the president decided to give his consent. The contract has already been handed over, so we plan to receive the documents during my next visit on February 28th."
[0034] The server device 1 can store the received voice data in the storage unit 14. The server device 1 can also convert the received voice data into voice information (e.g., voice text data) that is information based on voice. In this case, the "voice information" can be data of the received voice itself, voice text data converted from voice into text, and information on translated or other converted data thereof.
[0035] Speech-to-text conversion can be performed while the speech is being input. In this case, the server device 1 may transmit the converted text information to the terminal device 2 and display a converted text display screen 300 (FIG. 10) on the display unit 26 of the terminal device 2. In this embodiment, speech-to-text conversion is performed while the speech is being input, but it may also be performed after the speech input is completed. Also, in this embodiment, speech data is converted (e.g., converted to text), but text conversion etc. need not be performed, and the speech data may be treated as speech information as is.
[0036] When the voice input is finished, the user issues an instruction to finish the voice input using the input unit 27 or microphone 28. The control unit 21 of the terminal device 2 determines whether the voice input is finished (step S107). If it is determined that the voice input is not finished (step S107: NO), the process of step S106 is repeated. If it is determined that the voice input is finished (step S107: YES), the server device 1 (receives a notification of the end of the voice input from the terminal device 2) provides the voice information, which is information based on the received voice, to the language model 5 based on an instruction from the program P1, and acquires and generates "output text data P3 consisting of multiple items including a date and a situation" (step S108). The output text data P3 will be described in detail later with reference to FIG. 7.
[0037] Here, the server device 1 may provide speech information to the language model 5 via a generation server device 4 connected to the network 3. The "output text data P3 consisting of multiple items including a date and a situation" may be output by specifying a prompt P2 provided to the language model 5 together with the speech information so as to output text consisting of multiple items including a date and a situation. Alternatively, the language model 5 itself may be configured to output the "output text data P3 consisting of multiple items including a date and a situation" without using the prompt P2. Alternatively, for example, the generation server device 4 may be configured to provide a prompt P2 specified so as to output the "output text data P3 consisting of multiple items including a date and a situation."
[0038] In determining whether the voice input has ended in step S107, the terminal device 2 may end the voice input process if it detects in the voice information that the voice input has ended. This allows the user to end the voice reception without touching the terminal device 2. Furthermore, following the voice reception end process, the server device 1 provides voice information such as voice data or voice text data to the language model 5 to generate output text data P3, allowing the user to check the output result screen 500 (FIG. 12), which will be described later, without touching the terminal device 2. This allows the user to easily perform voice input, for example, even when the user's hands are dirty at a work site.
[0039] Here, the process of generating output text data P3 by providing speech information to the language model 5 in step S108 (hereinafter referred to as "analysis process") can be performed as background processing by the server device 1 or the generation server device 4. In this case, the terminal device 2 can perform other operations, such as inputting speech of other registered information types, without waiting for the output text data P3 to be generated. Note that instead of performing background processing, the terminal device 2 may wait until it receives the output text data P3.
[0040] In this way, the information processing device 1 can receive voice related to an event after the event has ended based on instructions from the program P1, and provide voice information, which is information based on the voice, to the language model 5 to generate output text data P3 consisting of multiple items including the date and situation. This makes it possible to generate output text data P3 for multiple items included in a predetermined format such as a report, minutes, or other record, making it easier and more accurate to create documents such as records.
[0041] Furthermore, the speech input time may be, for example, within 5 minutes, 3 minutes, or 1 minute. This allows the input speech to be compiled in a short period of time, thereby shortening the input time. Furthermore, the content can be compiled without overlapping content, and when the language model 5 converts it into output text data P3, the content can be more accurate and error-free.
[0042] If the analysis process of step S108 is set to background processing, the server device 1 determines whether or not a request has been made to display the output text data list screen 400 (FIG. 11) (step S109). If a request has not been made to display the output text data list screen 400 (step S109: NO), the server device 1 repeats the process of step S109. If a request has been made to display the output text data list screen 400 (step S109: YES), the server device 1 displays the output text data list screen 400 (FIG. 11) on the terminal device 2 (S110).
[0043] The output text data list screen 400 may display a list of speech information for which a request for analysis processing has been made in step S108. Furthermore, the server device 1 may display a list of speech information for which output text data P3 has been acquired in the output text data list screen 400. When displaying the list of speech information for which a request for analysis processing has been made in step S107, the server device 1 may also display whether the output text data P3 has been acquired (analyzed) or whether the output text data P3 has not been acquired (analysis in progress). Furthermore, when the analysis processing is completed, the server device 1 may notify the terminal device 2 that the analysis has been completed.
[0044] Next, it is determined whether or not the (analyzed) speech information from which the output text data P3 was obtained has been selected on the output text data list screen 400 (step S111). If the speech information from which the output text data P3 was obtained has not been selected (step S111: NO), the process of step S111 is repeated. If the speech information from which the output text data was obtained has been selected (step S111: YES), the server device 1 displays the output result screen 500 on the terminal device 2 (step S112).
[0045] The output result screen 500 can be a screen in which the contents of the output text data P3 are reflected on the input screen 200. For example, the server device 1 may transmit a form in which the contents of the output text data P3 are reflected on the input screen 200 to the terminal device 2, and display the form on the display unit 26 of the terminal device 2 as the output result screen 500. Alternatively, the server device 1 may transmit the output text data P3 as is to the terminal device 2, and display the form on the display unit 26 of the terminal device 2 as the output result screen 500 in which the output text data P3 is reflected on the input screen 200.
[0046] The output result screen 500 will be described in detail later with reference to Fig. 12. Here, when an instruction to end the display of the output result screen 500 is received, the record document creation process ends. Note that the server device 1 may re-display the output result screen 500 by reading out the output text data P3 or data for displaying the output result screen 500 stored in the storage unit 14 in response to a request based on the operation of the user's terminal device 2.
[0047] If there is only one type of registration information, that is, if there is only one type of prompt P2 to select, or if the server device 1 uses only one fixed prompt P2 in the processing executed based on the program P1, the processing of steps S101 to S103 may be skipped and the processing may start from step S104. Alternatively, the processing of steps S101 to S104 may be skipped and the processing may start from step S105.
[0048] Figure 6 is a diagram showing an example of an item definition section P21 included in the prompt P2. Here, an example will be described in which a log relating to "visit notes" is selected as the registration information type. Here, Figure 6 uses the JSON (JavaScript (registered trademark) Object Notation) format, which allows combinations of items and their values to be saved in an enumerated or nested state, but formats other than JSON can also be used.
[0049] As shown in FIG. 6, the item definition section P21 of the prompt P2 has multiple items as item names ("fieldName"), including department name, visit date, activity time (from), activity time (to), purpose, progress status, interviewee, and detailed content. Here, the multiple items in the prompt P2 may include an item (first item) that requests content to be selected from multiple listed input candidates. In FIG. 6, the first item corresponds to the item with "list" written in the "dataType" column of each item, and specifically, the item names "purpose" and "progress status" each correspond to the first item.
[0050] For the item name "Purpose", candidates are listed in the "dataList" column, separated by commas, for example, "Visit, Visit (leaving business card), Telephone / Email, Internal affairs, Visit to the company, Headquarters meeting / Training, Other, Accompany / Attendance, Receiving documents, Telesales". The content of the item name "Purpose" is selected from these listed contents. The selection criteria are entered in the "hint" column, which specifies the output content, and can be specified, for example, "Please ignore schedules. If you are unsure, select visit".
[0051] For the item name "Progress," candidates are listed in the "dataList" column, separated by commas, such as "Initial, Meeting, In Progress, Suspended, Closed." The content of the item name "Progress" is selected from these listed options. The selection criteria are entered in the "hint" column, which specifies the output content, and may specify, for example, "How much progress has the project made as a result of the visit? Do not consider the project closed unless there is a contract-like content. An informal agreement is not a closed contract." In this way, for items whose content can be selectively determined from a limited set of terms, the limited terms can be listed in advance in prompt P2 to provide multiple input candidates. This allows the content of each item in the output text data P3, described below, to be appropriate.
[0052] Furthermore, the multiple items in prompt P2 may include an item (second item) that requests text content indicating a date and time. Here, "date and time" includes cases where only the date or only the time is entered. In FIG. 6, the items with "date" or "time" entered in the dataType column of each item correspond to the second item, and specifically, the item names "visit date," "activity time (from)," and "activity time (to)" correspond to the second item, respectively. The item definition section P21 may specify that the date and time are output in the format entered in the "dataFormat" column (e.g., "yyyy / MM / dd" or "HH:mm," etc.).
[0053] The criteria for the date and time to be output are entered in the "hint" field, which specifies the output content. For example, it can be specified as "Date of visit, or today's date if not entered. Today is February 16, 2024." By outputting the date and time data as text in a predetermined format, that format can be used as the content to be displayed on the screen. This makes it easier and more accurate to create documents such as reports, minutes, and other records.
[0054] Furthermore, the multiple items in prompt P2 may be classified as either a first item, a second item, or a third item that does not fall under either the first or second items. Here, in FIG. 6, an item with "text" written in the "dataType" column of each item may be considered to be the third item. The item definition section P21 specifies, for example, that the data be output in text format by writing "text" in the "dataType" column. Specifically, the item names are "department name," "interviewee," and "details." The criteria for the content to be written in text format are written in the "hint" column, which specifies the output content. For example, for the "interviewee" column, the following could be specified: "For those who attended the meeting during the visit, please leave a space between the name and title. If there are multiple items, please use an array format."
[0055] The third item may be neither content selected from a list of multiple input candidates nor text content indicating a date and time, nor may it be any other item. In this way, the creator of prompt P2 can create prompt P2 by entering any of the first, second, and third items. This allows the creator of prompt P2 to determine which of the first, second, and third items the item corresponds to, making it easier to create prompt P2.
[0056] Furthermore, the multiple items in prompt P2 may specify the output of specific content contained in the voice information, and may also include an item that specifies the content to be output if the specific content is not present. Specifically, in FIG. 6, for example, the item named "Visit Date" specifies "Date of visit, or today's date if not specified. Today is February 16, 2024." In this case, if the voice information contains the date of the visit, the date of the visit is determined as the content of the "Visit Date" item. However, if the voice information does not contain the date of the visit, today's date, i.e., the date the voice information was acquired, can be used.
[0057] For example, the item "Business" in Figure 6 specifies "If you don't know, please select "Visit." In this case, if the voice information contains content corresponding to the "Business," that content is determined as the content of the "Business." However, if the voice information does not contain content corresponding to the Business, "Visit" can be determined as the content of the "Business."
[0058] Normally, when there is no specific content to be entered in a field from the speech information, the content to be entered in the field is left blank, and the person in charge would enter it manually. However, even when there is no specific content, by pre-specifying other input content in this way, it is possible to output content corresponding to the field. This makes it easier to create documents such as reports, minutes, and other records. Note that "when there is no specific content" may also include a case where the language model 5 is unable to recognize the specific content.
[0059] Furthermore, the specification for the output content of the prompt P2 item may include the current date and time. Similar to the "date and time" described above, the "date and time" may include only the date or only the time. For example, in FIG. 6, "Today is February 16, 2024" in the item "Visit Date" corresponds to the current date and time as a date only, and "The current time is 1:47 PM" in the item "Activity Time (from)" corresponds to the current date and time as a time only. This allows the output of date and time based on the time when prompt P2 was created or sent, even when the output content includes the current date and time. Furthermore, even if, for example, time has passed between the time when the date and time were entered into prompt P2 and the processing of language model 5 due to processing reasons, the correct date and time entered into prompt P2 allows output of output text data P3 containing an accurate date and time.
[0060] In this way, the prompt P2 specifies the content and format of the output text data P3 to be generated, which is made up of multiple items. Such prompt P2 is provided to the language model 5 along with speech information and is reflected in the output text data P3. The prompt P2 may also indicate hints for generating the data, the length of the data, and the data format. Here, the "format" and "data length and data format" may specify, for example, the format of the content of the item, such as the data size, whether it is fixed length or variable length, the number of characters (maximum or fixed) in the case of a string, and the type of characters (numbers, alphabets, etc.).
[0061] In this embodiment, the selected item definition portion P21 is used as the basis for the prompt P2, and the prompt P2 is provided together with the speech information to the language model 5. Here, the server device 1 may add the following preamble portion to the top of the item definition portion P21 of the prompt P2 shown in Figure 6 to create the prompt P2.
[0062] "You will identify specific keywords or phrases in the input string, and select and convert the appropriate information based on that. This conversion involves extracting the content corresponding to the specified fieldName and hint, and processing it based on the rules defined by length, dataType, dataFormat, and dataList. The results will be provided in JSON format. Please strictly adhere to the hint, length, dataType, and dataFormat. Please also output the criteria for each item." Below are the steps to follow: 1. Analyze the input string and identify the content that corresponds to fieldName or hint. Hint must be strictly observed. 2. For each fieldName, apply the appropriate format depending on the dataType (text, date, time, list, etc.). For date and time, convert according to the specific dataFormat (e.g. "yyyy / MM / dd" or "HH:mm"). 3. For fieldName with dataList, choose from the provided candidates and check if they match the specified selection. 4. For each fieldName, apply length-based constraints as needed. 5. Assemble the extracted and converted information based on the above steps in json format.
[0063] In this embodiment, the item definition portion P21 and the preamble portion of the prompt P2 are separated, but the entire prompt P2 may be stored in the memory unit 14, and the program P1 may select the entire prompt P2 in step S102 described above.
[0064] As described above, the server device 1 receives an input of a registration information type, which is the type of information to be registered, based on instructions from the program P1, and can select one prompt P2 from multiple types of pre-stored prompts P2 based on this registration information type. The selected prompt P2 can then be provided to the language model 5 along with speech information. Here, "selecting one prompt P2" includes selecting a portion of multiple types of prompts P2, as described above, and combining the common portion to create one prompt P2. This allows for easier and more accurate creation of documents such as reports, minutes, and other records.
[0065] 7 is a diagram showing an example of output text data P3 consisting of multiple items. The output text data P3 in FIG. 7 is composed of a combination of the name of each item and its content for each of the multiple items. As shown in this diagram, the output text data P3 adds a "value" item to each item in the item definition section P21 of FIG. 6, and the content of the "value" is entered based on the content of the audio information. The format of the output text data P3 can be a format that allows each item and its value to be combined and saved.
[0066] For example, a format such as JSON as shown in FIG. 7 can be used, which allows combinations of items and their values to be saved in an enumerated or nested state. By specifying a predetermined format for a report, minutes, or other record in the prompt P2, output text data P3 can be output in the specified format. This makes it possible to create reports, minutes, and other records more easily and accurately. The output text data P3 may be acquired and generated by the server device 1 via the generation server device 4 and the network 3.
[0067] FIG. 8 is a diagram showing an example of a registration information type selection screen 100. The registration information type selection screen 100 is displayed on the display unit 26 of the terminal device 2. As shown in this figure, the registration information type selection screen 100 has an item display area 101 in which items for selecting events (registration information types) to be recorded are listed. The item display area 101 can include, for example, items such as visit notes, maintenance records, order data registration, tentative reservation requests, and inquiries. Here, "visit notes" is an item for keeping records of visits when visiting a customer. "Maintenance records" is an item for keeping records of when maintenance is performed on equipment, etc.
[0068] "Order data registration" is an item for recording the type and quantity of parts that need to be ordered when it becomes necessary to order parts, etc. "Tentative reservation request" is an item for recording the equipment to be used, the date and time of use, etc. when it becomes necessary to use equipment. "Inquiry" is an item for recording when it becomes necessary to make an inquiry. The user of the terminal device 2 can select one of the items and notify the server device 1.
[0069] Here, the selection operation may be based on voice input or by operating a touch panel, etc. The example of the registration information type selection screen 100 in Fig. 8 shows an example in which voice input is performed, and a message indicating that voice is being acquired and a voice input stop button 116 for stopping voice input by tapping the screen are displayed.
[0070] 9 is a diagram showing an example of an input screen 200. The input screen 200 is displayed on the display unit 26 of the terminal device 2. As shown in this figure, the input screen 200 displays a title 201 indicating that the registration information type is "visit memo." The input screen 200 has an input item display area 203. Here, the input item display area 203 can reflect and display the item names of the prompt P2, namely, department name, visit date, activity time (from), activity time (to), purpose, progress status, interviewee, and detailed content, as input items.
[0071] Furthermore, the input screen 200 may display a voice input start icon 215 that accepts an operation to start voice input. Here, the touch panel serving as the input unit 27 may be arranged to function on the screen of the display unit 26. This allows the user of the terminal device 2 to start voice input by tapping the position where the voice input start icon 215 is displayed. Note that the input unit 27 that accepts an instruction to start voice input is not limited to a touch panel, and may be a mouse, a keyboard, or the like. Note that if an operation is already being performed by voice input, a voice input stop button may be displayed.
[0072] In this way, the input screen 200 can display a screen showing multiple items to be input, as well as a voice input start icon 215 that accepts an operation to start voice input. This allows the user to visually understand that multiple items to be input can be input by voice. Furthermore, it makes it easier and more accurate to create documents such as reports, minutes, and other records.
[0073] 10 is a diagram showing an example of a converted text display screen 300. The converted text display screen 300 is displayed on the display unit 26 of the terminal device 2. The converted text display screen 300 has a text display area 301 that sequentially displays text data that has been converted from the voice being input. The screen may also have a voice input stop button 316 that indicates that voice is being acquired and that can be used to stop voice input by tapping the screen.
[0074] Here, the speech input may be stopped by vocally inputting a message to stop the speech input. In this case, whether to stop the speech input may be determined by a speech recognition process performed by the terminal device 2 or the server device 1 that has received speech information from the terminal device 2. Note that, when the speech data is directly input to the language model 5, the converted text display screen 300 does not need to be displayed.
[0075] 11 is a diagram showing an example of an output text data list screen 400. The output text data list screen 400 is displayed on the display unit 26 of the terminal device 2. As shown in this figure, the output text data list screen 400 has an audio information list area 401 in which a list of audio information to be processed is displayed. The audio information list area 401 may display a list of audio information for which a request for analysis processing has been made in step S108.
[0076] In this case, for example, the analysis process requests or the voice information may be displayed in order of most recent, and the date and time, the type of registered information, and the status of whether or not the analysis has been completed may be displayed. The voice information list area 401 may display a list of the voice information from which the output text data P3 has been acquired. The user of the terminal device 2 may tap or otherwise select analyzed voice information to display the output result screen 500.
[0077] 12 is a diagram showing an example of the output result screen 500. The output result screen 500 is displayed on the display unit 26 of the terminal device 2. The input item display area 503 of the output result screen 500 reflects the contents of the output text data P3 generated in the server device 1. Specifically, the items in the input item display area 503, namely, department name, visit date, activity time (from), activity time (to), purpose, progress status, interviewee, and detailed content, reflect the contents of the item names corresponding to the items in the output text data P3 shown in FIG.
[0078] 12, the output result screen 500 has a title 501 and an input item display area 503, similar to the input screen 200. Also, the output result screen 500 has a voice input start icon 515, similar to the input screen 200. The voice input start icon 515 accepts an operation to start voice input, similar to the input screen 200. Also, the output result screen 500 has a playback icon 517 that accepts an operation to start playing back voice data.
[0079] The playback icon 517 accepts an operation to start playback of audio data stored in the storage unit 14 when audio is input. If a touch panel is used as the input unit 27, the user of the terminal device 2 can instruct the start of a corresponding function by tapping the display position of each icon. Note that the input unit 27 is not limited to a touch panel and may be a mouse, keyboard, etc.
[0080] Here, for example, in the case where the content of the item name "Detailed Content" in the output text data P3 in Fig. 7 has a large number of characters and cannot be displayed in the "Detailed Content" column of the input item display area 203 of the output result screen 500 in Fig. 12, it is possible to display only a portion of the content to indicate that it has been entered. In this way, the output result screen 500 can display at least a portion of the text corresponding to each of the multiple items in the output text data P3 in the corresponding column for each of the multiple items.
[0081] In this case, the server device 1 can display all of the text corresponding to the selected item by accepting an instruction to select one of the multiple items displayed on the output result screen 500. This allows the user to know whether or not there is text even if the number of characters of text to be displayed for an item on the output result screen 500 is large, and also allows the user to visually understand the operation to display the entire text.
[0082] Fig. 13 is a diagram showing an example of an editing screen 600 that displays all of the text corresponding to one item. Fig. 13 shows editing screen 600 that displays all of the content text of the item for the item name "Detailed Content" on output result screen 500 in Fig. 12. As shown in Fig. 13, editing screen 600 has a content display area 610. Editing screen 600 also has an edit button 601, a save button 602, a playback icon 603, and a voice input start icon 604.
[0083] The edit button 601 is used when editing the content of an item displayed in the content display area 610, and tapping the edit button 601 makes the content of the content display area 610 editable. The save button 602 is used when saving the edited content. The functions provided by operating the play icon 603 and the voice input start icon 604 are the same as those of the play icon 517 and the voice input start icon 515 on the output result screen 500 described above, respectively, and therefore redundant explanations will be omitted.
[0084] Furthermore, since the server device 1 stores the received audio as an audio file, it is possible to receive an instruction to save edited content for one of the multiple items and save the edited content while playing the stored audio file based on an instruction from the program P1, such as by tapping the playback icon 603. Here, playing the audio file includes the server device 1 causing the terminal device 2 to perform streaming playback.
[0085] This allows the user to check the audio file even if he or she is unsure whether the content of the output result screen 500 output based on the output text data P3 is appropriate. Furthermore, if any corrections are necessary, the user can correct the content to be appropriate. Furthermore, even if the user forgets the content of the voice input at a later date, the user can play back the audio file and check it.
[0086] The editing screen 600 can be displayed by tapping or other operations on the portion of the output result screen 500 where the corresponding item is displayed. Note that the operation is not limited to operation on a touch panel, and may be operation using a mouse, keyboard, or the like. Note that the editing screen 600 may be displayed for the purpose of editing the content even when content is displayed for the item in the input item display area 503 of the output result screen 500.
[0087] Furthermore, the output result screen 500 may accept an instruction to display speech text data obtained by converting the input speech into text. Specifically, an operation object such as an icon that accepts an instruction to display the speech text data is arranged on the output result screen 500, and the server device 1 that executes the program P1 accepts an instruction to display the speech text data by operating the operation object. This makes it possible to access the speech text data that is the source of the content displayed on the output result screen 500 from the output result screen 500, so that even if the content of the output result screen 500 has blanks or unclear points, the content of the speech text data can be easily confirmed.
[0088] [Variations] Fig. 14 is a diagram showing an example of a system configuration without providing a generation server device 4. In Fig. 1, the generation server device 4 is provided separately from the server device 1, but as shown in Fig. 14, a configuration may also be adopted in which the server device 1 directly inputs speech information to the language model 5 without providing a generation server device 4. Here, a prompt P2 may also be input together with the speech information.
[0089] [Embodiment 2] Fig. 15 is a schematic diagram of a record creation system according to embodiment 2. In embodiment 1, a server device 1, which is an information processing device 1, executes a program P1, but in embodiment 2, a terminal device 2, which is an information processing device 2, executes the program P1. As shown in this figure, the record creation system in Fig. 15 does not have an information processing device (server device) 1, as compared to the record creation system in Fig. 1, and is composed of an information processing device (terminal device) 2, a generation server device 4, a network 3, and a language model 5.
[0090] Fig. 16 is a block diagram showing an example of the configuration of a terminal device 2 according to embodiment 2. The hardware configuration of the terminal device 2 is the same as the hardware configuration shown in Fig. 4, except that a program (or program product) P1 and a prompt P2 executed by the control unit 21 are stored in the storage unit 24. The program P1 and the prompt P2 have the same contents as the program P1 and the prompt P2 stored in the storage unit 14 of the server device 1 in embodiment 1, and can perform the same output as in embodiment 1 on the display unit 26 and output the same screen.
[0091] 15 and 16, the terminal device 2 can operate in the same procedure as the flowchart of Fig. 5 by replacing the operation of the server device 1 based on the program P1 with the operation of the terminal device 2 based on the program P1. In this case, in the second embodiment, the communication between the server device 1 and the terminal device 2 via the network 3 as in the first embodiment is not performed, and the processing is performed within the terminal device 2.
[0092] As a result, similar to the server device 1 of embodiment 1, the terminal device 2 can receive voice related to an event after the event has ended based on instructions from the program P1 stored in the storage unit 24, provide voice information, which is information based on the voice, to the language model 5, and generate output text data P3 consisting of multiple items including the date and situation. Therefore, similar to embodiment 1, output text data P3 can be generated for multiple items included in a predetermined format such as a report, minutes, or other record, making it possible to create documents such as records more easily and accurately.
[0093] Furthermore, the terminal device 2 can provide the language model 5 with a prompt P2, along with speech information, that specifies the content and format of the output text data P3 to be generated based on instructions from the program P1. Furthermore, the prompt P2 can indicate hints for generating the data, the length of the data, and the data format. For example, the terminal device 2 may store in the storage unit 24 a prompt P2 that includes an item definition portion P21 shown in FIG. 6.
[0094] Furthermore, based on the instructions of the program P1, the terminal device 2 can accept input of the registration information type, which is the type of information to be registered (FIG. 5, step S102), select one prompt P2 from multiple types of prompts P2 stored in advance based on the registration information type (FIG. 5, step S103), and provide the selected prompt P2 together with speech information to the language model 5 (FIG. 5, step S108).
[0095] Furthermore, in prompt P2, the multiple items can include, for example, a first item selected from a list of multiple input candidates and a second item consisting of text indicating a date and time, as shown in item definition portion P21 included in prompt P2 in Fig. 6. In this case, the multiple items may be classified into any of a first item, a second item, and a third item that does not fall into either the first or second item.
[0096] Furthermore, in prompt P2, specifically, for example, as shown in the item named "Visit Date" in Figure 6, the multiple items may specify the output of specific content contained in the audio information, and may also include an item that specifies the content to be output when the specific content does not exist.
[0097] In addition, based on the instructions of the program P1, the terminal device 2 can display an input screen 200 (see FIG. 9) that shows multiple items to be input and also shows a voice input start icon 215 (515, 604) that accepts an operation to start voice input.
[0098] An output result screen 500 may be displayed in which at least a portion of the text corresponding to each of the multiple items in the output text data P3 is shown in the corresponding column for each of the multiple items, and upon receiving an instruction to select one of the multiple items shown on the output result screen 500, all of the text corresponding to the selected item may be displayed (see Figures 12 and 13).
[0099] Furthermore, the terminal device 2 can save the received audio in an audio file based on a command from the program P1. Furthermore, while playing back the saved audio file, the terminal device 2 can accept an instruction to save edited content for one of the multiple items and save the edited content (see playback icon 603 in FIG. 13, etc.).
[0100] Furthermore, the terminal device 2 can display an output result screen 500 that shows at least a portion of the text corresponding to each of the multiple items in the output text data P3 in the corresponding fields of the multiple items, based on a command from the program P1. The output result screen 500 can also show a link that allows access to the audio file (see FIG. 12).
[0101] In addition, in the second embodiment, the contents of the output result screen 500 can be stored in an accessible state on a server device or the like that is constantly connected to the network 3, so that the output result screen 500 can be referenced from multiple terminal devices.
[0102] In the above-described embodiments, the cases where the program P1 is executed on the server device 1 and the terminal device 2 are respectively shown, but the program P1 may be stored in a distributed manner on the server device 1 and the terminal device 2, and the server device 1 and the terminal device 2 may cooperate to operate the entire program P1, so that a single information processing device is formed from multiple devices.
[0103] According to the program P1, the information processing method related to the program P1, and the information processing devices 1 and / or 2 of the embodiments of the present disclosure, it is possible to more easily and accurately create reports, minutes, and other records. The program P1 can be referred to as a program product, software, or software product, and these may be provided on a recording medium or in a form distributed via a communication network.
[0104] The embodiments of the present disclosure are illustrative in all respects and are not restrictive. The scope of the present invention is not defined by the above disclosure but is defined by the claims, and it is intended to include all modifications within the meaning and scope of the claims.
[0105] The matters described in each embodiment can be combined with each other. Furthermore, the independent claims and dependent claims described in the claims can be combined with each other in any and all combinations, regardless of the reference format. Furthermore, the claims do not use a multi-claim format in which a multi-claim, which is a claim that references two or more claims, further references a multi-claim format (multi-multi claim), but a combination using a multi-multi claim format in which each claim in the same category references all of its higher-level claims is also possible. [Explanation of symbols]
[0106] 1. Information processing device (server device) 11 Control section 12 Communications Department 13 Reading unit 14 Storage section 19 Bus 1a Portable storage media 2. Information processing equipment (terminal equipment) 21 Control section 22 Communications Department 23 Reading unit 24 Memory section 26 Display section 27 Input section 28. Mike 29 Bus 2a Portable storage media 3 Network 4. Generation server device 41 Control Unit 42 Communications Department 43 Reading unit 44 Storage section 46 Display section 49 Bus 4a Portable storage media 5. Language Model P1 Program (Program Product) P2 prompt P21 Item definition part P3 Output text data 100 Registration information type selection screen 101 Item display area 116 Audio input stop button 200 Input Screen 201 Title 203 Input item display area 215 Voice input start icon 300 Converted text display screen 301 Text display area 316 Audio input stop button 400 Output text data list screen 401 Audio information list area 500 Output result screen 501 Title 503 Input field display area 515 Voice input start icon 517 Playback Icon 600 Editing Screen 601 Edit button 602 Save button 603 Playback Icon 604 Voice input start icon 610 Content display area
Claims
1. After the conference with the client is over, audio relating to the conference with the client is received; By providing the language model with prompts including voice information, which is information based on the voice, output examples including visit, visit to the company, and telesales, and an instruction to treat it as a visit if the user is unclear, output examples including a deal being concluded and an ongoing deal, and an instruction not to determine that a deal is concluded unless there is content that resembles a contract, the language model outputs the matter including a visit, visit to the company, or telesales, and the progress status including a deal being concluded or an ongoing deal. A program that causes a computer to perform a process.
2. After the conference with the client is over, audio relating to the conference with the client is received; By providing the language model with prompts including voice information, which is information based on the voice, output examples including visit, visit to the company, and telesales, and an instruction to treat it as a visit if the user is unclear, output examples including a deal being concluded and an ongoing deal, and an instruction not to determine that a deal is concluded unless there is content that resembles a contract, the language model outputs the matter including a visit, visit to the company, or telesales, and the progress status including a deal being concluded or an ongoing deal. An information processing method in which processing is performed by a computer.
3. A control unit is provided, the control unit After the conference with the client is over, audio relating to the conference with the client is received; By providing the language model with prompts including voice information, which is information based on the voice, output examples including visit, visit to the company, and telesales, and an instruction to treat it as a visit if the user is unclear, output examples including a deal being concluded and an ongoing deal, and an instruction not to determine that a deal is concluded unless there is content that resembles a contract, the language model outputs the matter including a visit, visit to the company, or telesales, and the progress status including a deal being concluded or an ongoing deal. Information processing device.
Citation Information
Patent Citations
Voice recording management system, voice recording management device, voice recording management method, and program
JP2023027001A
Data entry support device and method, data content voice input system, record management device, and allocation method
JP7541412B1
Program, information processing device, information processing system, information processing method, and information processing terminal
JP2023169093A
JPP7541412B