Program, Information Processing Method, and Information Processing Apparatus

A program simplifies and enhances the accuracy of document creation by processing voice information through a language model to generate structured text data, addressing the inefficiencies and errors in quick document completion.

JP7710588B1Active Publication Date: 2025-07-18SUMITOMO MITSUI FINANCE AND LEASING
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2024197591
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-07-18
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

Documents for reports, minutes, and other records often need to be completed quickly, leading to inefficiencies and potential errors due to time constraints.

Method used

A program that utilizes a computer to receive voice information after an event, process it through a language model, and generate output text data including date and situation, simplifying and enhancing accuracy in document creation.

Benefits of technology

Enables the creation of reports, minutes, and other records more simply and accurately, reducing input time and minimizing errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710588000001_ABST
    Figure 0007710588000001_ABST
Patent Text Reader

Abstract

Provide a program or the like that can create reports, minutes, and other records more simply and accurately. 【Solution means】After the event ends, the program receives the voice related to the event, gives the voice information, which is the information based on the voice, to the language model 5, and causes the computer to execute the process of generating output text data composed of a plurality of items including the date and situation. Further, the program can give a prompt that defines the content and format of the output text data of the plurality of generated items to the language model 5 together with the voice information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a program, an information processing method, and an information processing apparatus.

Background Art

[0002] In various operations, reports and minutes are created for the purpose of managing progress, schedules, etc. Patent Document 1 discloses a program for generating a summary text that summarizes text information included in one or a plurality of segment voice data.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Documents for reports, minutes, and other records (hereinafter referred to as "records, etc.") in various operations often need to be completed and submitted within a short time between operations, and the creation work is a burden on the person in charge. In addition, description errors are likely to occur due to the requirement of submission within a short time, and including the correction work may further reduce the work efficiency.

[0005] One aspect of the present disclosure provides a program or the like that can create reports, minutes, and other records, etc. more simply and accurately.

Means for Solving the Problems

[0006] A program according to one aspect of the present disclosure causes a computer to execute a process of receiving, after an event ends, voice related to the event, providing voice information, which is information based on the voice, to a language model, and generating output text data including a plurality of items including a date and a situation.

Advantages of the Invention

[0007] According to a program or the like according to one aspect of the present disclosure, creation of reports, minutes, and other record documents can be performed more simply and accurately.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Modes for Carrying Out the Invention

[0009] Hereinafter, an example of a program, an information processing method, and an information processing apparatus according to one aspect of the present disclosure will be described in detail with reference to the drawings. In the description, the same reference numerals are assigned to the same elements, and redundant descriptions are omitted as appropriate.

[0010] [Embodiment 1] FIG. 1 is a schematic diagram of a recording document creation system according to Embodiment 1. As shown in FIG. 1, the recording document creation system includes an information processing apparatus 1, an information processing apparatus 2, and a generation server apparatus 4, and each apparatus transmits and receives information via a network 3 such as the Internet. The generation server apparatus 4 may be communicatively connected so as to be able to use a language model 5 such as a large language model (LLM) stored inside or outside the generation server apparatus 4.

[0011] The information processing apparatus 1 is an information processing device that performs processing, storage, and transmission / reception of various information. The information processing apparatus 1 is, for example, an information processing device such as a server device, a personal computer, a tablet terminal, and a smartphone. The information processing apparatus 1 may be composed of a plurality of information processing devices, or may be configured as one of a plurality of virtual devices (virtual machines) configured within one information processing device. In this embodiment, the information processing apparatus 1 will be described by substituting the server apparatus 1 in order to avoid complicated explanations. However, the information processing apparatus 1 is not limited to the server apparatus, and may be configured as any of the information processing devices as described above.

[0012] The information processing device 2 receives the information output from the server device 1 via the network 3, accepts the operations of the users of the information processing device 2, and transmits the information related to the operations to the server device 1 via the network 3. The information processing device 2 is an information processing device such as, for example, a personal computer, a smartphone, a mobile phone, a wearable device, and a tablet. In this embodiment, in order to avoid complicated explanations, the information processing device 2 is described by substituting it with the terminal device 2.

[0013] FIG. 2 is a block diagram showing a configuration example of the server device 1. The server device 1 may be an information processing device mainly composed of an electronic circuit using semiconductor circuit elements. The server device 1 includes a control unit 11, a communication unit 12, a reading unit 13, and a storage unit 14. The components of the control unit 11, the communication unit 12, the reading unit 13, and the storage unit 14 are connected to each other so as to be communicable via a bus 19 or the like. Note that the server device 1 may have another configuration such as not having the reading unit 13.

[0014] The control unit 11 can be configured to include any one or more of processing devices such as a CPU (Central Processing Unit), an MPU (Micro-Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), a DSP (Digital Signal Processor), and a quantum processor. The control unit 11 can be configured to read and execute the program (or program product) P1 stored in the storage unit 14.

[0015] The communication unit 12 is a communication module for performing communication-related processing, and can transmit and receive information to and from the terminal device 2 and the like via the network 3. The reading unit 13 can read a portable storage medium 1a such as a CD (Compact Disc)-ROM, a DVD (Digital Versatile Disc)-ROM, or a USB (registered trademark) memory. The control unit 11 can read the program P1 and / or data from the portable storage medium 1a via the reading unit 13 and store them in the storage unit 14. Further, the control unit 11 can download the program P1 from another computer via the network 3 or the like and store it in the storage unit 14.

[0016] The storage unit 14 can include a volatile storage unit such as a RAM (Random Access Memory), and a non-volatile storage unit such as a ROM (Read Only Memory), an HDD (Hard disk drive), and a flash memory. The control unit 11 can temporarily store the program and / or data read from the non-volatile storage unit or received from the communication unit 12 in the volatile storage unit for use. The storage unit 14 can store the data of the prompt P2 to be described later in addition to the program (or program product) P1 executed by the control unit 11. The storage unit 14 can store a plurality of prompts P2.

[0017] FIG. 3 is a block diagram showing a configuration example of the generation server device 4. The generation server device 4 can be an information processing device mainly composed of an electronic circuit using semiconductor circuit elements. The generation server device 4 includes a control unit 41, a communication unit 42, a reading unit 43, and a storage unit 44. Each configuration of the control unit 41, the communication unit 42, the reading unit 43, and the storage unit 44 is communicably connected to each other by a bus 49 or the like. Since these configurations can be the same as those of the information processing device (server device) 1 in FIG. 2 except that the program P1 and the prompt P2 are not stored in the storage unit 44, duplicate descriptions are omitted. Note that the generation server device 4 may have other configurations such as not having a reading unit 43 for reading the portable storage medium 4a as in the case of the server device 1.

[0018] The generation server device 4 is an information processing device that handles a language model 5 such as a large language model (LLM) like GPT (Generative Pre-trained Transformer, registered trademark), BERT (Bidirectional Encoder Representations from Transformer), and Gemini (registered trademark). The language model 5 may be stored in the storage unit 44 within the generation server device 4, or may be stored in a storage or the like that is communicatively connected to the generation server device 4.

[0019] The generation server device 4 inputs input data such as images, voices, and text into the language model 5 and generates a response sentence. The language model 5 is a learned machine learning model and can be, for example, a large language model such as GPT (Generative Pre-trained Transformer) or BERT (Bidirectional Encoder Representations from Transformer), but may also be other language models. The server device 1 may be communicatively connected to the generation server device 4 via the network 3, or may be communicatively connected to the generation server device 4 directly or via a local network or the like without going through the network 3.

[0020] FIG. 4 is a block diagram showing a configuration example of the terminal device 2. Here, the terminal device 2 can be an information processing device mainly composed of an electronic circuit using semiconductor circuit elements. The terminal device 2 may have a control unit 21, a communication unit 22, a reading unit 23, a storage unit 24, a display unit 26, an input unit 27, and a microphone 28. The respective configurations of the control unit 21, the communication unit 22, the reading unit 23, the storage unit 24, the display unit 26, the input unit 27, and the microphone 28 may be communicatively connected to each other by a bus 29 or the like.

[0021] The control unit 21 can be configured to include any one or more of processing devices such as a CPU (Central Processing Unit), MPU (Micro-Processing Unit), GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), DSP (Digital Signal Processor), and quantum processor. The control unit 21 can be configured to read and execute a program (or program product) stored in the storage unit 24 described later.

[0022] The communication unit 22 is a communication module for performing communication-related processing, and can transmit and receive information to and from the server device 1 etc. via the network 3. The reading unit 23 can read a portable storage medium 2a such as a CD (Compact Disc)-ROM, DVD (Digital Versatile Disc)-ROM, or USB (registered trademark) memory. The control unit 21 can read a program and / or data from the portable storage medium 2a via the reading unit 23 and store it in the storage unit 24. Also, the control unit 21 can download a program from another computer via the network 3 etc. and store it in the storage unit 24.

[0023] The storage unit 24 can include a volatile storage device such as a RAM (Random Access Memory), and a non-volatile storage device such as a ROM (Read Only Memory), HDD (Hard disk drive), and flash memory. The control unit 21 may temporarily store in the volatile storage device a program and / or data read from the non-volatile storage device or received from the communication unit 22 for high-speed reading and writing and utilization by the control unit 21.

[0024] The display unit 26 can be a liquid crystal display, an organic EL (Electro Luminescence) display, or the like. The input unit 27 can be an input device such as a keyboard, a mouse, a touch panel, and a camera. The microphone 28 converts the sound around the terminal device 2 into digital data based on an instruction from the control unit 21. The control unit 21 stores the converted digital data in the storage unit 24.

[0025] FIG. 5 is a flowchart showing an example of a record document creation process. The server device 1 executes a record document creation process based on a program P1 or the like. As shown in this flowchart, the server device 1 displays a registration information type selection screen 100 (FIG. 8) for selecting a registration information type (step S101). The registration information type is the type of the content of the record document that the user of the terminal device 2 intends to record after the event ends.

[0026] Here, the "event" can be work such as a meeting, an interview, and equipment maintenance. The registration information type can be, for example, the type of record documents such as reports on maintenance work, construction work, and other work, as well as minutes of weekly or monthly meetings and other meetings. Also, the "event" may be something that requires parts ordering, equipment reservation, and inquiries. In addition, items such as date and time and situation common to each registration information type can be included, and items specific to each, such as a department name determined selectively, a meter value for inputting a numerical value, and the presence or absence of defects, may also be included.

[0027] Regarding the prompt P2 that defines such items, it will be described in detail in the description with reference to FIG. 6 below. Note that, as will be described later, for the input for recording, "voice related to an event" by the user of the terminal device 2 can be used. "Voice related to an event" can be voice corresponding to the content equivalent to the minutes of a meeting or interview, voice corresponding to the content equivalent to the report of work such as maintenance, and voice such as the content of ordering parts, reservation of equipment, and inquiries, and other voices. Here, "voice" can include, in addition to the voice itself, voice data obtained by digitally converting the voice.

[0028] The server device 1 that executes the program P1 can display, for example, on the display unit 26 of the terminal device 2, a plurality of registration information types that can be selected from a plurality of registration information types stored in advance in the storage unit 14. Also, it may be that when the user of the terminal device 2 selects any one of the displayed plurality of registration information types, the server device 1 receives the selected registration information type.

[0029] The server device 1 determines whether or not a selection of a registration information type has been received on the registration information type selection screen 100 (step S102). If no registration information type has been selected in step S102 (step S102: NO), the process of step S102 is repeated. If a registration information type has been selected (step S102: YES), a prompt P2 corresponding to the registration information type is selected (step S103).

[0030] The prompt P2 may be selected from a plurality of registration information types such as "visit memo", "maintenance record", "order data registration", "temporary reservation application", and "inquiry" shown in FIG. 8. Here, one or more prompts P2 corresponding to each registration information type may be stored in the storage unit 14 of the server device 1 as shown in FIG. 2. The server device 1 can acquire the prompt P2 corresponding to the selected registration information type from the storage unit 14 based on the instruction of the program P1.

[0031] When the prompt P2 is selected, the server device 1 transmits an input screen 200 (in the form) (Fig. 9) based on the content of the selected prompt P2 to the terminal device 2 according to the instruction of the program P1 and displays it on the display unit 26 of the terminal device 2 (step S104).

[0032] The terminal device 2 on which the input screen 200 is displayed determines whether there is an instruction to start voice input (step S105). If an instruction to start voice input is not detected (step S105: NO), the process of step S105 is repeated. If an instruction to start voice input is detected (step S105: YES), the terminal device 2 detects voice via the microphone 28 and converts it into voice data (step S106). The server device 1 receives the voice data via the network 3. Here, if the operation is performed by an instruction based on voice input to the terminal device 2 at the stage of selecting the registration information type in step S102 or earlier, the process of step S105 may not be necessary.

[0033] The voice to be input can be, for example, the one that reads the following text. "Today, I had an appointment with Mr. X, visited the president of Company Y at ABC Co., Ltd., and was able to meet with President Y and Managing Director Z. Since it was about one week since the previous proposal, I thought it would be difficult to make a decision on the contract, but thanks to the president's decision, I was able to obtain an internal approval. The contract has already been delivered, and it is planned to receive the documents during the next visit on February 28th."

[0034] The server device 1 can save the received voice data in the storage unit 14. Also, the server device 1 can convert the received voice data into voice information (for example, voice text data) which is information based on the voice. In this case, "voice information" can be the data of the received voice itself, the voice text data obtained by converting the voice into text, and the information of these translated or other converted data.

[0035] The text conversion of the voice can be performed during the voice input. In this case, the server device 1 may transmit the converted text information to the terminal device 2 and display the converted text display screen 300 (FIG. 10) on the display unit 26 of the terminal device 2. In the present embodiment, the text conversion of the voice is assumed to be performed during the voice input, but it may be performed after the voice input is completed. Further, in the present embodiment, the voice data is converted (for example, text conversion), but the text conversion or the like may not be performed, and the voice data may be handled as voice information as it is.

[0036] When the voice input ends, the user gives an instruction to end the voice input using the input unit 27 or the microphone 28. The control unit 21 of the terminal device 2 determines whether the voice input has ended (step S107). If it is determined that the voice input has not ended (step S107: NO), the process of step S106 is repeated. If it is determined that the voice input has ended (step S107: YES), the server device 1 (receives a notification of the end of the voice input from the terminal device 2) gives the voice information, which is information based on the received voice, to the language model 5 based on the instruction of the program P1, and acquires and generates "output text data P3 composed of a plurality of items including the date and the situation" (step S108). The output text data P3 will be described in detail in the description with reference to FIG. 7 below.

[0037] Here, the server device 1 may provide voice information to the language model 5 via the generation server device 4 connected to the network 3. The "output text data P3 consisting of a plurality of items including date and situation" may be output by stipulating that the prompt P2 given to the language model 5 together with the voice information outputs a text consisting of a plurality of items including date and situation. Also, the language model 5 itself may be configured to output the "output text data P3 consisting of a plurality of items including date and situation" without using the prompt P2. Further, for example, the generation server device 4 may be configured to provide a prompt P2 stipulated to output the "output text data P3 consisting of a plurality of items including date and situation".

[0038] In addition, in the determination of the end of the voice input in step S107, the terminal device 2 may end the voice input process when it detects that the voice information contains the content indicating the end of the voice input. Thereby, the user can end the reception of the voice without touching the terminal device 2. Further, continuously with the process of ending the reception of the voice, the server device 1 provides voice information such as voice data or voice text data to the language model 5 to generate the output text data P3, so that the user can confirm the output result screen 500 (FIG. 12) described later without touching the terminal device 2. Thereby, for example, even when the user's hands are dirty at the work site, the user can easily perform voice input.

[0039] Here, the process of providing the voice information in step S108 to the language model 5 to generate the output text data P3 (hereinafter referred to as the "analysis process") can be a background process by the server device 1 or the generation server device 4. In this case, the terminal device 2 can perform other operations, for example, input voices of other registration information types, without waiting for the output text data P3 to be generated. Note that, instead of the background process, the terminal device 2 may wait until it receives the output text data P3.

[0040] In this way, based on the instructions of program P1, the information processing device 1 can receive the voice related to the event after the event ends, provide the voice information, which is the information based on the voice, to the language model 5, and generate the output text data P3 consisting of a plurality of items including the date and situation. As a result, since the output text data P3 can be generated for a plurality of items included in a predetermined format such as a report, minutes of a meeting, or other record documents, the creation of documents such as record documents can be performed more simply and accurately.

[0041] Also, the input time of the voice may be, for example, within 5 minutes, 3 minutes, or 1 minute. Thereby, since the input voice can be made into a summary in a short time, the input time can be shortened. Also, it can be made into a coherent content that does not include overlapping content, and when the language model 5 converts it into the output text data P3, it can be made into more accurate content without mistakes.

[0042] When the analysis process of step S108 is set as a background process, the server device 1 determines whether the display of the output text data list screen 400 (FIG. 11) is requested (step S109). If the display of the output text data list screen 400 is not requested (step S109: NO), the process of step S109 is repeated. If the display of the output text data list screen 400 is requested (step S109: YES), the output text data list screen 400 (FIG. 11) is displayed on the terminal device 2 (S110).

[0043] On the output text data list screen 400, it may be possible to display a list of the voice information for which the analysis process request in step S108 was made. Also, on the output text data list screen 400, the server device 1 may display a list of the voice information for which the output text data P3 was acquired. When displaying a list of the voice information for which the analysis process request in step S107 was made, the server device 1 may also display whether the output text data P3 has been acquired (analyzed) or not acquired (being analyzed). Note that when the analysis process is completed, the server device 1 may notify the terminal device 2 that the analysis is complete.

[0044] Subsequently, on the output text data list screen 400, it is determined whether the voice information for which the output text data P3 has been acquired (analyzed) is selected (step S111). If the voice information for which the output text data P3 has been acquired is not selected (step S111: NO), the process of step S111 is repeated. If the voice information for which the output text data has been acquired is selected (step S111: YES), the server device 1 causes the terminal device 2 to display the output result screen 500 (step S112).

[0045] The output result screen 500 can be a screen in which the content of the output text data P3 is reflected on the input screen 200. For example, the server device 1 may transmit a form in which the content of the output text data P3 is reflected on the input screen 200 to the terminal device 2 and display it as the output result screen 500 on the display unit 26 of the terminal device 2. Also, the server device 1 may transmit the output text data P3 as it is to the terminal device 2, and in the terminal device 2, it may be displayed on the display unit 26 as the output result screen 500 in which the output text data P3 is reflected on the input screen 200.

[0046] The output result screen 500 will be described in detail in the description with reference to FIG. 12 described later. Here, when an instruction to end the display of the output result screen 500 is received, the record document creation process is ended. Note that the server device 1 may redisplay the output result screen 500 by reading out the output text data P3 stored in the storage unit 14 or the data for displaying the output result screen 500 in response to a request based on the operation of the user's terminal device 2.

[0047] Note that when there is only one type of registration information, that is, when there is only one type of prompt P2 to be selected, or when the prompt P2 used in the process executed by the server device 1 based on the program P1 is fixed to one, etc., the processes of steps S101 to S103 are not performed, and the process can start from the process of step S104. It is also possible to start from the process of step S105 without performing the processes of steps S101 to S104.

[0048] FIG. 6 is a diagram showing an example of the item definition part P21 included in the prompt P2. Here, the case where a record document related to "visit memo" is selected as the registration information type will be described as an example. Here, in FIG. 6, the JSON (JavaScript (registered trademark) Object Notation) format in which combinations of each item and values of each item can be saved in an enumerated or nested state is used, but formats other than JSON can also be used.

[0049] As shown in FIG. 6, the item definition part P21 of the prompt P2 has a plurality of items as the item name ("fieldName"), such as department name, visit date, activity time (from), activity time (to), Component , progress status, interviewer, and detailed content. Here, the plurality of items in the prompt P2 may have an item (first item) that requests content to be selected from a plurality of enumerated input candidates. In FIG. 6, the first item corresponds to the item described as "list" in the "dataType" column of each item. Specifically, the item names " Component " and "progress status" each correspond to the first item.

[0050] The item name " Component " has candidates separated by commas in the "dataList" column. For example, they are listed as "Visiting, Visiting (Leaving Business Cards), Phone / Email, In-Company Affairs, Company Visit, Company Meeting / Training, Others, Accompanying / Attending, Document Receiving, Telesales". The content of the item name " Component " is selected from these listed contents. The selection criteria are described in the "hint" column that defines the output content. For example, it is defined as "Please ignore the schedule. If you don't know, please select Visiting", etc.

[0051] The item name "Progress Status" has candidates separated by commas in the "dataList" column. For example, they are listed as "First Time, Meeting, In Progress, Interrupted, Contracted". The content of the item name "Progress Status" is selected from these listed contents. The selection criteria are described in the "hint" column that defines the output content. For example, it is defined as "To what extent has the project progressed as a result of the visit. Unless there is something like a contract, please do not judge it as contracted. Verbal approval is not a contract", etc. In this way, for items whose content is selectively determined from limited terms, the limited terms can be listed in Prompt P2 in advance as multiple input candidates. Thereby, the content of each item in the output text data P3 described later can be made appropriate.

[0052] Also, in Prompt P2, multiple items may have an item (the second item) that requests content in text indicating a date and time. Here, "date and time" includes cases where it is only the date or only the time. In FIG. 6, the items with "date" or "time" described in the "dataType" column of each item correspond to the second item. Specifically, the item names "Visit Date", "Activity Time (from)", and "Activity Time (to)" each correspond to the second item. The item definition part P21 may stipulate that the date and time be output in the format described in the "dataFormat" column (for example, "yyyy / MM / dd" and "HH:mm", etc.).

[0053] The criteria for the date and time to be output are described in the "hint" column that defines the content of the output. For example, it can be defined as "the date of the visit, or if not specified, today's date. Today is February 16, 2024." etc. By outputting the date and time data as text in a predefined format, it can be used as content for on-screen display in that format. As a result, document creation such as reports, meeting minutes, and other records can be performed more simply and accurately.

[0054] Also, the multiple items in Prompt P2 may be classified into any of the first item, the second item, and a third item that does not correspond to either the first item or the second item. Here, in FIG. 6, the items with "text" described in the "dataType" column of each item may correspond to the third item. The item definition part P21 stipulates, for example, to output in text by describing "text" in this "dataType" column. Specifically, the item names are "Department Name", "Interviewee", and "Detailed Content". The criteria for the content described in text are described in the "hint" column that defines the content of the output. For example, in the case of the "Interviewee" column, it can be defined as "Those who attended the meeting during the visit, please leave a space between the name and the honorific. If there are multiple, please present them in an array." etc.

[0055] Note that the third item may be an item other than those indicated by the content selected from a plurality of enumerated input candidates and the content in text indicating the date and time. In this way, the creator of Prompt P2 can create Prompt P2 in any of the ways of describing the first item, the second item, and the third item. As a result, the creator of Prompt P2 can determine which of the first to third items the item corresponds to and create Prompt P2 more simply.

[0056] In addition, the multiple items in Prompt P2 may include items that define the output for specific content included in the voice information and also define the content to be output when the specific content does not exist. Specifically, in FIG. 6, for example, in the item with the item name "Visit Date", it is defined as "The date of the visit, or if not stated, today's date. Today is February 16, 2024." In this case, if the date of the visit exists in the voice information, the date of the visit is determined as the content of the item "Visit Date". However, if the date of the visit does not exist in the voice information, it can be today's date, that is, the date when the voice information was acquired.

[0057] In addition, for example, in the item with the item name " Component " in FIG. 6, it is defined as "If you don't know, please visit." In this case, if the content corresponding to " Component " exists in the voice information, that content is determined as the content of " Component ". However, if the content corresponding to Component does not exist in the voice information, "Visit" can be determined as the content of " Component ".

[0058] Normally, when there is no specific content to be described in the item from the voice information, it is conceivable that the content input into the item is left blank or the like and manually input by the person in charge. However, even in such a case where there is no specific content, by prescribing other input content in advance, the content corresponding to the item can be output. As a result, the creation of documents such as reports, meeting minutes, and other record books can be performed more simply. Note that "when there is no specific content" may include the case where the language model 5 cannot recognize this specific content.

[0059] In addition, the regulations regarding the output content of the items in Prompt P2 may include the current date and time. Similar to the "date and time" mentioned above, the "date and time" can include cases where it is only the date or only the time. For example, in FIG. 6, "Today is February 16, 2024." in the item name "Visit Date" corresponds to the current date and time as only the date, and "The current time is 13:47." in the item name "Activity Time (from)" corresponds to the current date and time as only the time. Thus, when the output content includes the current date and time, the content of the date and time based on the time point when Prompt P2 is created or sent can be output. Also, for example, even if some time has passed from the time when the date and time is input into Prompt P2 until the processing by the language model 5 due to processing convenience, since the correct date and time is input into Prompt P2, the output text data P3 including the error-free date and time can be output.

[0060] In this way, Prompt P2 defines the content and format of the output text data P3 composed of a plurality of generated items. Such a Prompt P2 is provided to the language model 5 together with the voice information and reflected in the output text data P3. Also, Prompt P2 may indicate hints for data generation, the length of the data, and the data format. Here, the "format" as well as the "length and data format of the data" can define, for example, as the format for the content of the item, the data size, the distinction between fixed length or variable length, the (maximum or fixed) number of characters in the case of a character string, the types of characters such as numbers and alphabets, etc.

[0061] In this embodiment, based on the selected item definition part P21, Prompt P2 is formed, and this Prompt P2 is provided to the language model 5 together with the voice information. Here, the server device 1 may add the following preamble part above the item definition part P21 of Prompt P2 shown in FIG. 6 to form Prompt P2.

[0062] You identify specific keywords or phrases within the input string, select appropriate information based on them, and perform conversions. This conversion involves extracting the content corresponding to the specified fieldName and hint, and processing it based on the rules defined by length, dataType, dataFormat, and dataList. The result is provided in JSON format. Please strictly follow hint, length, dataType, and dataFormat. Also output the criteria for judging each item. The following shows the execution steps. 1. Analyze the input string and identify the content corresponding to fieldName or hint. Strictly follow hint. 2. For each fieldName, apply an appropriate format according to the dataType (such as text, date, time, list, etc.). For date and time, convert according to a specific dataFormat (for example, "yyyy / MM / dd" or "HH:mm"). 3. For fieldName with a dataList, select from the provided candidates and check if they match the specified options. 4. For each fieldName, apply restrictions based on length as necessary. 5. Construct the extracted and converted information in JSON format based on the above steps.

[0063] Note that in this form, the item definition part P21 and the preamble part for the prompt P2 are separated, but the entire prompt P2 may be stored in the storage unit 14, and the program P1 may select the entire prompt P2 in the above step S102.

[0064] As described above, the server device 1 receives an input of a registration information type, which is a type of information to be registered, based on a command of the program P1, and can select one prompt P2 from a plurality of types of prompts P2 stored in advance based on this registration information type. Further, the selected prompt P2 can be given to the language model 5 together with the voice information. Here, "selecting one prompt P2" includes selecting a part of the plurality of types of prompts P2 as described above and combining it with the common part to form one prompt P2. Thereby, documents such as reports, minutes, and other record documents can be created more simply and accurately.

[0065] FIG. 7 is a diagram showing an example of output text data P3 composed of a plurality of items. The output text data P3 in FIG. 7 is composed of a combination of the name and content of each item for a plurality of items. As shown in this figure, in the output text data P3, an item of "value" is added to each item of the item definition part P21 in FIG. 6, and the content of "value" is input based on the content of the voice information. As the format of the output text data P3, a format that can store a combination of each item and its value can be used.

[0066] For example, a format such as JSON shown in FIG. 7, in which combinations of each item and its value can be listed or stored in a nested state, can be used. By defining a predetermined format such as a report, minutes, or other record document in the prompt P2, the output text data P3 in the specified format can be output. Thereby, documents such as reports, minutes, and other record documents can be created more simply and accurately. The output text data P3 may be acquired and generated by the server device 1 via the generation server device 4 and the network 3.

[0067] FIG. 8 is a diagram showing an example of a registration information type selection screen 100. The registration information type selection screen 100 is displayed on the display unit 26 of the terminal device 2. As shown in this figure, the registration information type selection screen 100 has an item display area 101 in which items for selecting an event (registration information type) to leave a record are listed. The item display area 101 can include, for example, items such as visit memo, maintenance record, order data registration, provisional reservation application, and inquiry. Here, the "visit memo" is an item for leaving a record of a visit when visiting a customer. The "maintenance record" is an item for leaving a record when performing maintenance on facilities and the like.

[0068] "Order data registration" is an item for recording the types and quantities of parts that require ordering when an order for parts or the like is necessary. "Provisional reservation application" is an item for recording the facilities to be used, the use date and time, etc. when the use of facilities is necessary. "Inquiry" is an item for recording when an inquiry is necessary. The user of the terminal device 2 can perform an operation of selecting one of the items and notify the server device 1.

[0069] Here, the selection operation may be based on voice input or may be by an operation such as that on a touch panel. In the example of the registration information type selection screen 100 in FIG. 8, an example of performing voice input is shown, and a voice input stop button 116 for stopping the voice input is displayed due to the fact that voice is being acquired and a tap operation on the screen.

[0070] FIG. 9 is a diagram showing an example of an input screen 200. The input screen 200 is displayed on the display unit 26 of the terminal device 2. As shown in this figure, the input screen 200 shows a title 201 indicating that the registration information type is "visit memo". The input screen 200 has an input item display area 203. Here, the input item display area 203 can reflect and display the department name, visit date, activity time (from), activity time (to), Component , progress status, interviewer, and detailed content, which are the item names of the prompt P2, as input items respectively.

[0071] Further, the input screen 200 may display a voice input start icon 215 for receiving an operation to start voice input. Here, the touch panel as the input unit 27 may be arranged to function on the screen of the display unit 26. Thereby, the user of the terminal device 2 can start voice input by tapping the position where the voice input start icon 215 is displayed. Note that the input unit 27 for receiving an instruction to start voice input is not limited to a touch panel, and may be a mouse, a keyboard, or the like. When an operation is already being performed by voice input, a voice input stop button may be displayed.

[0072] In this way, the input screen 200 can display a screen that shows a plurality of items to be input and shows the voice input start icon 215 for receiving an operation to start voice input. Thereby, the user can visually grasp that voice input can be performed for the plurality of items to be input. Also, document creation such as reports, minutes, and other records can be performed more simply and accurately.

[0073] FIG. 10 is a diagram showing an example of a converted text display screen 300. The converted text display screen 300 is displayed on the display unit 26 of the terminal device 2. The converted text display screen 300 has a text display area 301 for sequentially displaying text data for which text conversion has been completed with the voice being input. Also, it may have a voice input stop button 316 for stopping voice input by means of an operation of tapping the screen and indicating that voice is being acquired.

[0074] Here, the stop of voice input may be performed by inputting by voice content such as an instruction to stop voice input. In this case, it may be determined whether to stop voice input by voice recognition processing performed by the terminal device 2 or the server device 1 that has received voice information from the terminal device 2. When directly inputting voice data into the language model 5 or the like, the converted text display screen 300 may not be displayed.

[0075] FIG. 11 is a diagram showing an example of an output text data list screen 400. The output text data list screen 400 is displayed on the display unit 26 of the terminal device 2. As shown in this figure, the output text data list screen 400 has a voice information list area 401 in which a list of voice information to be processed is displayed. The voice information list area 401 may display a list of voice information for which the analysis process request in step S108 has been made.

[0076] In this case, for example, it can be displayed in the order of new analysis process requests or in the order of new voice information, and the date and time, registration information type, and status of whether it has been analyzed can be displayed respectively. The voice information list area 401 may display a list of voice information for which the output text data P3 has been acquired. By selecting the analyzed voice information by an operation such as tapping by the user of the terminal device 2, the output result screen 500 can be displayed.

[0077] FIG. 12 is a diagram showing an example of an output result screen 500. The output result screen 500 is displayed on the display unit 26 of the terminal device 2. The content of the output text data P3 generated in the server device 1 is reflected in the input item display area 503 of the output result screen 500. Specifically, in the items of the deployment name, visit date, activity time (from), activity time (to), Component , progress status, interviewer, and detailed content in the input item display area 503, the content of the item names corresponding to the respective items of the output text data P3 shown in FIG. 7 is reflected.

[0078] As shown in FIG. 12, the output result screen 500 has a title 501 and an input item display area 503, similar to the input screen 200. Also, the output result screen 500 has a voice input start icon 515, similar to the input screen 200. The voice input start icon 515 accepts an operation to start voice input, similar to the input screen 200. Further, the output result screen 500 has a playback icon 517 that accepts an operation to start playback of voice data.

[0079] The playback icon 517 accepts an operation to start playback of the voice data stored in the storage unit 14 at the time of voice input. When a touch panel is adopted as the input unit 27, the user of the terminal device 2 can give an instruction to start the corresponding function by tapping on the display position of each icon. Note that the input unit 27 is not limited to a touch panel and may be a mouse, keyboard, or the like.

[0080] Here, for example, as in the content of the item name "Detailed Content" of the output text data P3 in FIG. 7, since there are a large number of characters, when it cannot be displayed in the "Detailed Content" column of the input item display area 203 of the output result screen 500 in FIG. 12, only a part can be displayed to indicate that it has been input. In this way, the output result screen 500 can show at least a part of the text corresponding to each of the plurality of items of the output text data P3 in the corresponding columns of the plurality of items.

[0081] In this case, the server device 1 accepts an instruction to select one of the plurality of items shown on the output result screen 500 and can display all of the text corresponding to the selected one item. Thereby, even when the number of characters of the text to be displayed for an item on the output result screen 500 is large, the user can grasp the presence or absence of the text and can also visually grasp the operation for displaying the entire text.

[0082] FIG. 13 is a diagram showing an example of an edit screen 600 in which all of the text corresponding to one item is displayed. In FIG. 13, an edit screen 600 in which all of the content text of the item name "Detailed Content" of the output result screen 500 in FIG. 12 is displayed is shown. As shown in FIG. 13, the edit screen 600 has a content display area 610. The edit screen 600 also has an edit button 601, a save button 602, a playback icon 603, and a voice input start icon 604.

[0083] The editing button 601 is used when editing the content of an item displayed in the content display area 610. By tapping the editing button 601, the content in the content display area 610 can be set to an editable state. The save button 602 is used when saving the edited content. Since the functions triggered by operating the playback icon 603 and the voice input start icon 604 are the same as those of the playback icon 517 and the voice input start icon 515 on the output result screen 500 described above, duplicate explanations are omitted.

[0084] Also, since the server device 1 saves the received voice as a voice file, by tapping the playback icon 603 or the like, based on the instruction of program P1, while playing the saved voice file, it can receive a save instruction for the edited content of one of the multiple items and save the edited content. Here, playing the voice file includes the server device 1 causing the terminal device 2 to perform streaming playback.

[0085] Thereby, even when the user suspects that the content of the output result screen 500 output based on the output text data P3 may not be appropriate, the user can check the voice file. Also, if correction is necessary, it can be corrected to appropriate content. Further, even if the user forgets the content at the time of voice input at a later date, the user can play and check the voice file.

[0086] The editing screen 600 can be displayed by operating such as tapping on the part where the corresponding item of the output result screen 500 is displayed. Note that the operation is not limited to the operation of the touch panel and may be by operations such as a mouse and a keyboard. Note that the editing screen 600 may be displayed for the purpose of editing the content even when the content of an item is displayed in the input item display area 503 of the output result screen 500.

[0087] Also, on the output result screen 500, for example, it may be configured to receive an instruction and display the voice text data obtained by converting the input voice into text. Specifically, an operation object such as an icon for receiving an instruction to display the voice text data is arranged on the output result screen 500, and the server device 1 that executes the program P1 receives an instruction to display the voice text data by operating the operation object. As a result, since it becomes possible to access the voice text data that is the source of the content displayed on the output result screen 500 from the output result screen 500, even if there are blanks or unclear points in the content of the output result screen 500, the content of the voice text data can be easily confirmed.

[0088] [Modification Example] FIG. 14 is a diagram showing an example of configuring the system without providing the generation server device 4. In FIG. 1, the generation server device 4 is provided separately from the server device 1, but as shown in FIG. 14, a configuration may be adopted in which the server device 1 directly inputs voice information to the language model 5 without providing the generation server device 4. Here, a prompt P2 may be further input together with the voice information.

[0089] [Embodiment 2] FIG. 15 is a schematic diagram of a record creation system according to Embodiment 2. In Embodiment 1, the server device 1 which is the information processing device 1 executes the program P1, but in Embodiment 2, a configuration is adopted in which the terminal device 2 which is the information processing device 2 executes the program P1. As shown in this figure, compared with the record creation system of FIG. 1, the record creation system of FIG. 15 does not include the information processing device (server device) 1, and is composed of the information processing device (terminal device) 2, the generation server device 4, the network 3, and the language model 5.

[0090] FIG. 16 is a block diagram showing a configuration example of the terminal device 2 according to Embodiment 2. The hardware configuration of the terminal device 2 is the same as the hardware configuration shown in FIG. 4, except that the program (or program product) P1 and the prompt P2 executed by the control unit 21 are stored in the storage unit 24. The program P1 and the prompt P2 have the same content as the program P1 and the prompt P2 stored in the storage unit 14 of the server device 1 in Embodiment 1, and can perform the same output to the display unit 26 and output the same screen as in Embodiment 1.

[0091] Also in the recording document creation system of FIGS. 15 and 16, the terminal device 2 can operate in the same procedure as the flowchart of FIG. 5 by replacing the operation of the server device 1 based on the program P1 with the operation of the terminal device 2 based on the program P1. In this case, in Embodiment 2, the communication via the network 3 between the server device 1 and the terminal device 2 in Embodiment 1 is not performed, and the processing is performed within the terminal device 2.

[0092] Thereby, similar to the server device 1 of Embodiment 1, the terminal device 2 can receive the voice related to the event after the event ends based on the command of the program P1 stored in the storage unit 24, and give the voice information, which is the information based on the voice, to the language model 5, and generate the output text data P3 including a plurality of items including the date and the situation. Therefore, similar to Embodiment 1, since the output text data P3 can be generated for a plurality of items included in a predetermined format such as a report, minutes of a meeting, and other recording documents, the creation of documents such as recording documents can be performed more simply and accurately.

[0093] Further, the terminal device 2 can provide the language model 5 with a prompt P2 that defines the content and format of the output text data P3 of a plurality of items to be generated, together with voice information, based on the instruction of the program P1. Further, the prompt P2 can indicate a hint for data generation, the length of the data, and the data format. For example, the terminal device 2 may have a prompt P2 including an item definition part P21 shown in FIG. 6 in the storage unit 24.

[0094] Also, based on the instruction of the program P1, the terminal device 2 receives an input of a registration information type, which is the type of information to be registered (FIG. 5, step S102), and based on the registration information type, selects one prompt P2 from a plurality of types of pre-stored prompts P2 (FIG. 5, step S103), and can provide the selected prompt P2 to the language model 5 together with voice information (FIG. 5, step S108).

[0095] Also, in the prompt P2, the plurality of items can include, for example, a first item selected from a plurality of enumerated input candidates as shown in the item definition part P21 included in the prompt P2 of FIG. 6, and a second item consisting of text indicating the date and time. In this case, the plurality of items may be classified into any one of a first item, a second item, and a third item that does not correspond to either the first item or the second item.

[0096] Also, specifically in the prompt P2, for example, as shown in the item with the item name "visit date" in FIG. 6, the plurality of items may include an item that defines output for specific content included in the voice information and also defines output content when the specific content does not exist.

[0097] Further, based on the instruction of the program P1, the terminal device 2 can display an input screen 200 (see FIG. 9) indicating a plurality of items to be input and showing a voice input start icon 215 (515, 604) that receives an operation to start voice input.

[0098] Display an output result screen 500 that shows at least a part of the text corresponding to each of the multiple items of the output text data P3 in the corresponding columns of the multiple items, and receive an instruction to select one of the multiple items shown on the output result screen 500, and it may display all of the text corresponding to the selected one item (see FIGS. 12 and 13).

[0099] Also, the terminal device 2 can save the received voice in a voice file based on the instruction of the program P1. Furthermore, while playing back the saved voice file, it can receive an instruction to save the edited content for one of the multiple items and save the edited content (see FIG. 13, playback icon 603, etc.).

[0100] Also, the terminal device 2 can display an output result screen 500 that shows at least a part of the text corresponding to each of the multiple items of the output text data P3 in the corresponding columns of the multiple items based on the instruction of the program P1. Also, the output result screen 500 can show a link that can access the voice file (see FIG. 12).

[0101] Note that also in the second embodiment, by storing the content of the output result screen 500 so that it can be accessed by a server device or the like constantly connected to the network 3, the output result screen 500 can be referred to from a plurality of terminal devices.

[0102] Note that in each of the above-described embodiments, the case where the program P1 is executed in the server device 1 and the case where it is executed in the terminal device 2 are shown respectively. However, the program P1 may be stored distributively in the server device 1 and the terminal device 2, and the server device 1 and the terminal device 2 may cooperate to operate the entire program P1, and a plurality of devices may constitute one information processing device.

[0103] According to the program P1 according to the embodiments of the present disclosure, the information processing method according to the program P1, and the information processing apparatuses 1 and / or 2, the creation of reports, minutes, and other records can be performed more simply and accurately. Note that the program of the program P1 can be referred to as a program product, software, or software product, and these may be provided on a recording medium or in a form distributed via a communication network.

[0104] The embodiments of the present disclosure are illustrative in all respects and not restrictive. The scope of the present invention is not shown in the above-described disclosure, but is shown by the claims, and it is intended that all modifications within the meaning and scope equivalent to the claims be included.

[0105] The matters described in each embodiment can be combined with each other. Also, the independent claims and dependent claims described in the claims can be combined with each other in all possible combinations regardless of the citation form. Further, in the claims, although a multi-claim that is a claim that cites two or more claims does not use a form of a further multi-claim (multi-multi-claim) that further cites, each claim in the same category may be in a form of a combination using a multi-multi-claim form in which all of its upper claims are cited.

Description of Reference Numerals

[0106] 1 Information processing apparatus (server apparatus) 11 Control unit 12 Communication unit 13 Reading unit 14 Storage unit 19 Bus 1a Portable storage medium 2 Information processing apparatus (terminal apparatus) 21 Control unit 22 Communication unit 23 Reading unit 24 Storage unit 26 Display unit 27 Input unit 28 microphones 29 buses 2a Portable storage medium 3 Network 4 Generation server device 41 Control unit 42 Communication unit 43 Reading unit 44 Memory unit 46 Display unit 49 Bus 4a Portable storage medium 5 Language model P1 Program (program product) P2 Prompt P21 Item definition part P3 Output text data 100 Registration information type selection screen 101 Item display area 116 Voice input stop button 200 Input screen 201 Title 203 Input item display area 215 Voice input start icon 300 Converted text display screen 301 Text display area 316 Voice input stop button 400 Output text data list screen 401 Voice information list area 500 Output result screen 501 Title 503 Input item display area 515 Voice input start icon 517 Playback icon 600 Edit screen 601 Edit button 602 Save button 603 Playback icon 604 Voice input start icon 610 Content display area

Claims

1. After the event ends, display a screen that shows a plurality of items to be input and a voice input start icon that accepts an operation to start voice input, receive the voice regarding the event, give a prompt to a language model that defines outputting voice information which is information based on the voice and specific content included in the voice information, and outputting predetermined input content when the specific content does not exist, and generate output text data consisting of a plurality of items including a date and a situation, the prompt defines outputting a date and time shifted by a predetermined time from the date and time based on the prompt creation time or the transmission time when the specific content does not exist A program that causes a computer to execute the process.

2. Give a prompt that defines the content and format of the output text data of the plurality of items to be generated, together with the voice information, to the language model The program according to claim 1.

3. The prompt indicates a hint for data generation, the length of the data, and the data format The program according to claim 2.

4. Receive an input of a registration information type which is a type of information to be registered, Based on the registration information type, select one of the plurality of types of prompts stored in advance, Give the selected prompt to the language model together with the voice information The program according to claim 2.

5. The plurality of items of the prompt are a first item that asks for content selected from a plurality of listed input candidates, and a second item that asks for content in text indicating a date and time, and includes The program according to claim 2.

6. The plurality of items are classified into any of the first item, the second item, and a third item that does not correspond to either the first item or the second item The program according to claim 5.

7. Display an output result screen that shows at least a part of the text corresponding to each of the plurality of items of the output text data in the corresponding columns of the plurality of items, Receive an instruction to select one of the plurality of items shown on the output result screen, Display all of the text corresponding to the selected one item The program according to claim 1.

8. Save the received voice in a voice file, While playing the saved audio file, receive a save instruction for the edited content of one of the plurality of items, Save the edited content The program according to claim 1.

9. Display an output result screen that shows at least a part of the text corresponding to each of the plurality of items of the output text data in the corresponding columns of the plurality of items, The output result screen shows a link that can access the audio file The program according to claim 8.

10. After the event ends, display a screen that shows a plurality of items to be input and shows an audio input start icon that receives an operation to start audio input, Receive the audio related to the event, Give a prompt to the language model that defines outputting voice information, which is information based on the voice, and specific content included in the voice information, and outputting predetermined input content when the specific content does not exist, and generate output text data consisting of a plurality of items including date and situation, The prompt defines that when the specific content does not exist, it outputs a date and time shifted by a predetermined time from the date and time based on the prompt creation time or the transmission time Information processing method.

11. Comprising a control unit, the control unit After the event ends, display a screen that shows a plurality of items to be input and shows an audio input start icon that receives an operation to start audio input, Receive the audio related to the event, Give a prompt to the language model that defines outputting voice information, which is information based on the voice, and specific content included in the voice information, and outputting predetermined input content when the specific content does not exist, and generate output text data consisting of a plurality of items including date and situation, The prompt defines that when the specific content does not exist, it outputs a date and time shifted by a predetermined time from the date and time based on the prompt creation time or the transmission time Information processing device.

Citation Information

Patent Citations

  • Electronic medical record system

    JP2001043284A

  • Communication system

    JP2002014894A

  • Operator's operation support system

    JP2009031810A

  • Input support system, input support method, and program

    JP2022030754A

  • Program, information processing device, information processing system, information processing method, and information processing terminal

    JP2023169093A