Information terminal, program, and summary generation system

The information terminal automates voice recording and summary generation based on location changes, addressing the issue of forgetfulness and manual operation in existing systems, and provides efficient meeting data capture.

JP7861247B1Active Publication Date: 2026-05-19UPWARD INC(JP)
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
UPWARD INC(JP)
Filing Date
2025-09-17
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing voice recording systems require manual operation to start and stop recording, leading to potential forgetfulness and missed recordings during meetings.

Method used

An information terminal that automatically starts and stops recording based on location changes within predefined visit areas, with optional delay settings, and integrates with an information processing device to create summary data from recorded audio using AI-generated text.

Benefits of technology

Prevents missed recordings by automating the recording process and generates meeting summaries automatically, reducing user effort and ensuring comprehensive data capture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007861247000001_ABST
    Figure 0007861247000001_ABST
Patent Text Reader

Abstract

This prevents you from forgetting to record the meeting audio. [Solution] The information terminal 1 includes a location identification unit 171 that identifies the current location of the information terminal 1, and a recording processing unit 174 that, by referring to visit area data that indicates a visit area corresponding to a visit location visited by user U of the information terminal 1, detects that the current location of the information terminal 1 has changed from a state in which it is not included in the visit area indicated by the visit area data to a state in which it is included in the visit area, and executes a process to start recording.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information terminal, a summary creation system, an information processing apparatus, a summary creation method, and a program.

Background Art

[0002] Conventionally, a technique for recording the voice during a meeting and transferring the recorded voice is known (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In order to record the voice during a meeting, it is necessary to perform an operation to start the recording. There were cases where recording could not be performed without performing this operation.

[0005] Therefore, the present invention has been made in view of these points, and an object thereof is to prevent forgetting to record.

Means for Solving the Problems

[0006] An information terminal according to a first aspect of the present invention includes a position specifying unit that specifies the current position of the information terminal, and by referring to access area data indicating an access area corresponding to an access location visited by a user of the information terminal, when it is detected that the current position has changed from a state where it is not included in the access area indicated by the access area data to a state where it is included in the access area, a recording processing unit that executes processing for starting recording.

[0007] The recording processing unit may, after referring to the visit area data and performing a process to start recording, terminate the recording when the current location changes from being included in the visit area to being no longer included in the visit area.

[0008] The information terminal further includes a receiving unit that accepts a setting for a delay time from when the current location is included in the visit area until recording begins, and the recording processing unit may perform a process to start recording after the delay time has elapsed from when the current location is included in the visit area.

[0009] The recording processing unit may refer to location data associated with the visited location and the delay time from when the current location is included in the visited area until recording begins, and perform a process to start recording after the delay time has elapsed from when the current location is included in the visited area.

[0010] The recording processing unit may, in response to the current location being included in the visit area corresponding to the visit location, display an operation image on the display unit to accept an operation to start recording, and start recording in response to the operation image being operated.

[0011] The information terminal may further include a data communication unit that transmits the recorded audio data to an information processing device that performs a process to summarize the content of the audio based on the audio data.

[0012] The data communication unit may identify the state of the communication line on which the voice data is transmitted, and the recording processing unit may generate the voice data by sampling the voice at a sampling rate determined based on the state of the communication line on which the data communication unit identified.

[0013] A second aspect of the present invention provides a summary creation system comprising the above-mentioned information terminal and an information processing device that creates summary data based on audio data transmitted from the information terminal, wherein the information processing device includes an audio data acquisition unit that acquires the audio data, a prompt creation unit that creates a prompt indicating a method for creating the summary data based on the audio data, and a summary creation unit that inputs the audio data and the prompt to a generation AI and creates the summary data based on the summary text output from the generation AI.

[0014] The summary creation system further includes a storage unit that stores a plurality of prompt templates indicating a method for creating the summary data, associated with at least one of the user's attributes or the attributes of the visited location. The prompt creation unit may select a prompt template based on the degree of matching between each of the plurality of prompt templates and the user's attributes or the attributes of the visited location, and create the prompt based on the selected prompt template.

[0015] The prompt creation unit may select one or more prompt templates based on the degree of matching between each of the plurality of prompt templates and the user's attributes or the attributes of the visited location, and then create the prompt from the one or more selected prompt templates based on the prompt template selected by the user.

[0016] The prompt creation unit may create a modified prompt by accepting a modification of the prompt by the user, and store the created modified prompt in the storage unit in association with the user. The summary creation unit may, after the modified prompt has been stored in the storage unit, input the modified prompt stored in the storage unit in association with the user to the generation AI when the voice data acquisition unit acquires the voice data transmitted from the information terminal used by the user.

[0017] The information processing device further includes a determination unit that determines the characteristics of the audio data input to the generating AI based on the characteristics of the training audio data used for machine learning by the generating AI that generates summary text based on the audio data, and a notification unit that notifies the information terminal that generates the audio data of the determined characteristics, and the summary creation unit may input the audio data generated by the information terminal having the characteristics determined by the determination unit to the generating AI, and create summary data based on the audio data based on the summary text output from the generating AI.

[0018] The determination unit may refer to the generation AI information, which associates a plurality of generation AIs with the characteristics of the learning audio data, and determine the characteristics of the audio data to be input to the generation AI based on the characteristics of the learning audio data associated with the generation AI used by the summarization unit to create the summary data.

[0019] The information processing device further includes an information acquisition unit that acquires information from the information terminal indicating the state of the communication environment of the information terminal that transmits the voice data to the information processing device, and the determination unit may further determine the characteristics of the voice data to be input to the generating AI based on the state of the communication environment of the information terminal indicated by the information acquired by the information acquisition unit.

[0020] The determination unit may reduce the sampling rate, which is a characteristic of the voice data input to the generating AI, to a predetermined lower limit as the communication speed of the information terminal that transmits the voice data to the information processing device slows down.

[0021] The determination unit may determine the sampling rate of the audio data input to the generating AI to be lower than the sampling rate of the training audio data.

[0022] When the communication speed of the information terminal that transmits the voice data to the information processing apparatus is less than a threshold value, the determination unit may determine the voice compression method, which is a feature of the voice data input to the generation AI, to be a voice compression method having a higher compression rate than the voice compression method of the learning voice data.

[0023] The information processing apparatus further includes an information acquisition unit that acquires information indicating the voice processing performance of the information terminal that transmits the voice data to the information processing apparatus from the information terminal, and the determination unit may further determine the sampling rate of the voice data input to the generation AI based on the voice processing performance indicated by the information acquired by the information acquisition unit.

[0024] The information processing apparatus according to the third aspect of the present invention includes a determination unit that determines the sampling rate of the voice data input to the generation AI based on the sampling rate of the learning voice data used by the generation AI that generates the summary text based on the voice data, a notification unit that notifies the information terminal that generates the voice data of the determined sampling rate, and a summary creation unit that inputs the voice data generated by the information terminal at the sampling rate determined by the determination unit to the generation AI and creates summary data based on the voice data based on the summary text output from the generation AI.

[0025] The summary creation method according to the fourth aspect of the present invention includes steps executed by a computer, the steps of: acquiring the information terminal as described above and the voice data transmitted from the information terminal; creating a prompt indicating a method for creating summary data based on the voice data; and inputting the voice data and the prompt to the generation AI and creating the summary data based on the text output from the generation AI.

[0026] The method for creating an abstract according to the fifth aspect of the present invention includes: a step of determining, by a computer, the characteristics of the voice data input to the generation AI based on the characteristics of the learning voice data used by the generation AI for machine learning to generate an abstract text based on the voice data; a step of notifying the information terminal that generates the voice data of the determined characteristics; and a step of inputting the voice data generated by the information terminal with the determined characteristics into the generation AI and creating abstract data based on the voice data based on the abstract text output from the generation AI.

[0027] The method for creating an abstract according to the sixth aspect of the present invention includes: a step of generating the voice data of the characteristics determined based on the characteristics of the learning voice data used by the generation AI for machine learning to generate an abstract text based on the voice data, which is executed by a computer; and a step of inputting the voice data into the generation AI and creating abstract data based on the voice data based on the abstract text output from the generation AI.

[0028] The program according to the seventh aspect of the present invention causes a computer to execute: a step of specifying the current position of the computer; and a step of referring to access area data indicating an access area corresponding to an access location visited by a user of the computer, and executing a process for starting recording in response to detecting that the current position has changed from a state not included in the access area indicated by the access area data to a state included in the access area.

[0029] The program according to the eighth aspect of the present invention causes a computer to execute: a step of generating the voice data of the characteristics determined based on the characteristics of the learning voice data used by the generation AI for machine learning to generate an abstract text based on the voice data; and a step of inputting the voice data into the generation AI and creating abstract data based on the voice data based on the abstract text output from the generation AI.

Advantages of the Invention

[0030] According to this invention, the effect of preventing forgetting to record is achieved. [Brief explanation of the drawing]

[0031] [Figure 1] This is a diagram illustrating the overview of the summary generation system S. [Figure 2] This diagram shows the configuration of information terminal 1. [Figure 3] This figure shows the relationship between the sampling rate of the audio data input to a speech-to-text generation AI and the character error rate. [Figure 4] This is a diagram showing the configuration of the information processing device 2. [Figure 5] This figure shows an example of data for potential places to visit. [Figure 6] This figure shows an example of geofence data. [Figure 7] This figure shows an example of user location data. [Figure 8] This figure shows an example of data used for template selection. [Figure 9] This is a flowchart showing the processing flow at information terminal 1. [Figure 10] This is a flowchart showing the processing flow in the information processing device 2. [Modes for carrying out the invention]

[0032] [Overview of Summary Generation System S] Figure 1 is a diagram illustrating the overview of the summary creation system S. The summary creation system S is a system for creating summary data based on recorded audio. The recorded audio is, for example, audio recorded at a meeting, and the summary creation system S creates summary data to be used as meeting minutes.

[0033] The summary creation system S comprises an information terminal 1, an information processing device 2, a generation AI server 3, and a customer management server 4. The information processing device 2 may have some or all of the functions of the generation AI server 3 and the customer management server 4.

[0034] Information terminal 1 is a terminal used by user U, who requests the creation of summary data, and is, for example, a smartphone or tablet. User U is, for example, a sales representative who visits customers and records meetings with them, but user U's job responsibilities and visit locations are arbitrary.

[0035] As shown in Figure 1(a), when information terminal 1 detects that user U has entered a visit area G corresponding to the location where the meeting will be held, it executes a recording start process. The visit area G is defined, for example, by a geofence that indicates a virtual boundary. Information terminal 1 has a GPS (Global Positioning System) receiver for detecting its own position, and detects that user U has entered the visit area G by detecting a change in the state from when the location of information terminal 1 is not included in the pre-set visit area G to when the location of information terminal 1 is included in the visit area G.

[0036] The recording start process is, for example, a process that automatically starts recording. The recording start process may also be a process that displays a screen for starting recording on the display of the information terminal 1, or a process that vibrates the information terminal 1 or generates a predetermined sound on the information terminal 1 so that user U is reminded to start recording.

[0037] Subsequently, as shown in Figure 1(b), when the information terminal 1 detects that user U has left the visited area G, it executes a recording termination process. The recording termination process is, for example, a process that automatically terminates the recording. The recording termination process may also be a process that displays a screen for terminating the recording on the display of the information terminal 1, or a process that vibrates the information terminal 1 or generates a predetermined sound on the information terminal 1 to remind user U to terminate the recording.

[0038] While information terminal 1 is recording, or after recording has finished, information terminal 1 transmits the recorded audio data to information processing device 2. Information terminal 1 transmits the audio data to information processing device 2 via, for example, a Wi-Fi® or mobile phone network wireless communication line. Information terminal 1 may transmit the audio data in association with location information, or it may transmit the audio data in association with information such as a customer name indicating a place visited (hereinafter referred to as "place visited information").

[0039] Information processing device 2 is a computer that performs the process of summarizing the content of audio data received from information terminal 1. Information processing device 2 creates summarized data by sending the audio data to generation AI server 3. Specifically, information processing device 2 sends the audio data to generation AI server 3 along with a prompt requesting the creation of summarized data, and receives the summarized data generated by generation AI server 3.

[0040] The information processing device 2 may obtain text data by sending the audio data to the generating AI server 3 along with a first prompt requesting that the audio data be converted into text data, and then obtain summary data by sending a second prompt to the generating AI server 3 along with the text data, for creating summary data based on the text data.

[0041] The generation AI server 3 is a computer equipped with a natural language processing model such as a large-scale language model (LLM) or a small-scale language model (SLM). Based on the content of the instructions contained in the prompt, the generation AI server 3 generates summary data by performing natural language processing on the audio data received from the information processing device 2.

[0042] The information processing device 2 outputs the summary data received from the generation AI server 3. The information processing device 2 may also process the summary data received from the generation AI server 3 and output the processed summary data. As an example, as shown in Figure 1(b), the information processing device 2 sends the summary data to the customer management server 4. The information processing device 2 sends the summary data to the customer management server 4, for example, along with location information or visit location information associated with the voice data. The information processing device 2 may also send the summary data to the customer management server 4 in association with information indicating the visit location received from the information terminal 1. The information processing device 2 may also send the summary data to the information terminal 1 or another terminal.

[0043] The customer management server 4 is, for example, a computer used by the organization to which user U belongs, and stores information about customers associated with customers that user U may visit. When the customer management server 4 receives summary data from the information processing device 2, it stores the received summary data in association with the customer corresponding to the location information or visit location information transmitted in association with the summary data.

[0044] With the summary creation system S configured as described above, user U can prevent forgetting to record conversations during meetings held at the visited location. Furthermore, since summary data is automatically created based on the recorded audio data, user U can avoid the trouble of creating meeting minutes.

[0045] [Configuration of Information Terminal 1] Figure 2 shows the configuration of information terminal 1. Information terminal 1 includes a GPS unit 11, a microphone 12, a terminal communication unit 13, a display unit 14, an operation unit 15, a storage unit 16, and a control unit 17. The control unit 17 includes a location identification unit 171, a reception unit 172, a data communication unit 173, and a recording processing unit 174.

[0046] The GPS unit 11 receives radio waves from GPS satellites. The GPS unit 11 inputs the location information transmitted by the received radio waves into the location identification unit 171. The location information is used to determine the latitude and longitude of the information terminal 1.

[0047] The microphone 12 receives sounds from the information terminal 1 and converts the received sounds into electrical signals. The microphone 12 inputs the generated electrical signals to the recording processing unit 174.

[0048] The terminal communication unit 13 has a communication interface for sending and receiving various types of data with the information processing device 2. For example, the terminal communication unit 13 receives configuration data transmitted from the information processing device 2 and inputs the received data to the data communication unit 173. The terminal communication unit 13 also transmits location data indicating the location of the information terminal 1 and recorded audio data to the information processing device 2.

[0049] The display unit 14 is a display that shows various kinds of information. For example, the display unit 14 displays an operation screen for starting recording and an operation screen for ending recording in response to instructions from the recording processing unit 174.

[0050] The operation unit 15 is a device for receiving operations from user U, and is, for example, a touch panel mounted on top of the display unit 14. When the operation unit 15 detects that user U has performed an operation, it notifies the reception unit 172 of information (for example, coordinates) indicating the location where user U performed the operation.

[0051] The storage unit 16 has storage media such as ROM (Read Only Memory) and RAM (Random Access Memory). The storage unit 16 stores the program executed by the control unit 17. The storage unit 16 may also temporarily store audio data generated by the recording processing unit 174.

[0052] The control unit 17 includes, for example, a CPU (Central Processing Unit). The control unit 17 functions as a location identification unit 171, a reception unit 172, a data communication unit 173, and a recording processing unit 174 by executing a program stored in the storage unit 16.

[0053] The location identification unit 171 identifies the current location of the information terminal 1 based on the location information input from the GPS unit 11. The location identification unit 171 notifies the data communication unit 173 and the recording processing unit 174 of the identified current location.

[0054] The reception unit 172 accepts various settings via the operation unit 15. For example, the reception unit 172 accepts the setting of the delay time from when the current location of the information terminal 1 is included in the visit area until recording starts. The delay time is determined based on the minimum time that is expected to be required from when user U arrives at the visit location until the meeting starts.

[0055] The data communication unit 173 transmits and receives various types of data to and from the information processing device 2 via the terminal communication unit 13. For example, the data communication unit 173 transmits location data notified by the location identification unit 171 to the information processing device 2. After transmitting the location data, the data communication unit 173 receives stay flag data indicating whether or not the location of the information terminal 1 indicated by the location data is included in the visited area. If the location of the information terminal 1 is included in the visited area, the data communication unit 173 may receive a visited location ID to identify the visited area.

[0056] The data communication unit 173 notifies the recording processing unit 174 that the information terminal 1 has changed to a state in which it is included in the visited area, or that the information terminal 1 has changed to a state in which it is not included in the visited area. The data communication unit 173 may also notify the recording processing unit 174 whether or not the information terminal 1 is included in the visited area by notifying the recording processing unit 174 of the received stay flag data.

[0057] The data communication unit 173 transmits the audio data recorded by the recording processing unit 174 to the information processing device 2. The data communication unit 173 may transmit the audio data generated by the recording processing unit 174 to the information processing device 2 in real time, or it may transmit the audio data from the start to the end of the recording to the information processing device 2 in response to receiving notification from the recording processing unit 174 that the recording has ended.

[0058] The data communication unit 173 may identify the state of the communication environment for transmitting voice data. The state of the communication environment is represented by the communication speed or the data error rate. For example, the data communication unit 173 may identify whether or not the information terminal 1 is in a state where it can connect to Wi-Fi (registered trademark) as the state of the communication environment. Before transmitting voice data, the data communication unit 173 may send and receive data with the information processing device 2 to check the state of the communication environment, and may identify the state of the communication environment based on the communication speed or data error rate when the data is sent and received. The data communication unit 173 may also identify the state of the communication environment based on data indicating the relationship between the visited location and the communication environment, which is stored in the storage unit 16 in advance. The data communication unit 173 notifies the recording processing unit 174 of the identified state of the communication environment.

[0059] The recording processing unit 174 samples the electrical signal input from the microphone 12 at a predetermined sampling rate, converts it from analog to digital, and stores the converted digital data as audio data in the storage unit 16. The recording processing unit 174 stores the audio data in the storage unit 16, for example, in association with time information. As will be described in detail later, the predetermined sampling rate is, for example, a sampling rate determined based on the sampling rate of the audio data used for machine learning by the generating AI server 3, which is used by the information processing device 2 when creating summary data based on the audio data.

[0060] The recording processing unit 174, by referring to visit area data that indicates a visit area corresponding to a visit location visited by the user of the information terminal 1, detects that the current location of the information terminal 1 has changed from a state where it is not included in the visit area indicated by the visit area data to a state where it is included in the visit area, and executes a process to start recording. For example, the recording processing unit 174 executes a process to start recording when the stay flag data notified by the data communication unit 173 changes from a value indicating that the user is not in the visit area to a value indicating that the user is in the visit area.

[0061] The recording processing unit 174 starts recording, for example, by beginning to store audio data based on the electrical signal input from the microphone 12 in the storage unit 16 when the current location is included in the area corresponding to the visited location. The recording processing unit 174 may also display an operation image on the display unit 14 to accept the operation to start recording when the current location is included in the area corresponding to the visited location, and start recording when the operation image is operated.

[0062] The recording processing unit 174 may perform processing to start recording after a delay time has elapsed since the current location of the information terminal 1 was included in the visit area. The delay time is, for example, the time when the reception unit 172 received the information from user U before the current location was included in the visit area. This allows user U to pre-set the time required from arrival at the customer's office until the meeting starts, so that the recording processing unit 174 can start recording or receive a recording start operation around the time the meeting is about to begin.

[0063] The recording processing unit 174 may refer to location data that associates the visited location with the delay time from when the current location is included in the visited area until recording begins, and perform processing to start recording after the delay time has elapsed since the current location was included in the visited area. The location data is data created in advance based on instructions from user U and is stored in the storage unit 16. By using such location data, the recording processing unit 174 eliminates the need for user U to set a delay time each time they visit a customer they frequently visit.

[0064] The recording processing unit 174 may, when the current location of the information terminal 1 is included in the visited area, display an operation image on the display unit 14 to accept the operation to start recording, and then automatically start recording after a delay time has elapsed since the operation image was displayed on the display unit 14. This prevents recording from not being made if the user U forgets to perform the operation to start recording.

[0065] The recording processing unit 174 may, after referring to the visit area data and executing the process to start recording, terminate the recording when the current location changes from being included in the visit area to being excluded from the visit area. The recording processing unit 174 may also, upon realizing that the current location is no longer included in the area corresponding to the visit location, display an operation image on the display unit 14 to accept the operation to terminate the recording, and terminate the recording when the operation image is operated. By operating in this manner, the recording processing unit 174 can prevent recording from continuing even after the meeting with the customer has ended, thus preventing the recording of unnecessary audio and the depletion of the battery level of the information terminal 1.

[0066] The recording processing unit 174 may determine the characteristics of the audio data to be generated based on the state of the communication line identified by the data communication unit 173. These characteristics include, for example, the sampling rate or encoding method of the electrical signals used to generate the audio data. Examples of encoding methods include PCM format and AAC format. The recording processing unit 174 generates the audio data by sampling the audio at a sampling rate determined based on the state of the communication line.

[0067] The recording processing unit 174 generates audio data by, for example, converting an electrical signal into digital data by sampling it at, for example, 16 kHz, and then converting it into digital data at a sampling rate determined based on the state of the communication line before transmitting it to the communication line. The recording processing unit 174 may also generate audio data by sampling the electrical signal at a sampling rate determined based on the state of the communication line.

[0068] The recording processing unit 174 lowers the sampling rate as the communication line conditions worsen and the data error rate increases. In this case, instead of reducing the amount of audio data within a predetermined time by lowering the sampling rate, the recording processing unit 174 may transmit the same audio data multiple times. This increases the probability that audio data without data errors reaches the information processing device 2. The recording processing unit 174 may also lower the sampling rate as the communication speed of the communication line slows down.

[0069] The recording processing unit 174 may determine the sampling rate so as to fall within the range of sampling rates notified from the information processing device 2 via the terminal communication unit 13. The inventors of the present invention have found that audio data with a sampling rate close to the sampling rate of the training audio data used for machine learning by the generation AI used by the information processing device 2 when creating a summary based on the audio data is less prone to conversion errors when the generation AI converts audio to text. Therefore, the recording processing unit 174 may generate audio data based on the sampling rate or encoding scheme notified from the information processing device 2, which stores a sampling rate or encoding scheme suitable for the generation AI.

[0070] Figure 3 shows the relationship between the sampling rate of the audio data input to a speech-to-text generation AI and the character error rate. Figure 3 shows the results obtained by inputting audio data with multiple different sampling rates into a generation AI trained using training audio data with a sampling rate of 16 kHz, and measuring the error rate of the characters converted by each audio data. The character error rate is represented by the number of incorrect words among the words included in the summary data generated by the generation AI. In Figure 3, the solid line shows the results when the number of trained models of the generation AI is the largest, the dashed line shows the results when the number of trained models is in the middle, and the dotted line shows the results when the number of trained models is the smallest.

[0071] In all cases, the character error rate is relatively low when the sampling rate is between 10 kHz and 25 kHz. Therefore, the recording processing unit 174 generates audio data at a sampling rate within the range of sampling rates determined based on the sampling rate used by the generation AI for learning (for example, between 10 kHz and 16 kHz), thereby improving the quality of the summary created by the information processing device 2 using the generation AI.

[0072] The recording processing unit 174 may generate audio data using an encoding scheme determined based on the state of the communication line identified by the data communication unit 173. In this case, the recording processing unit 174 may determine the encoding scheme from one or more encoding schemes of audio data used by the generating AI for learning, which are notified by the information processing device 2. This improves the quality of the summaries created by the information processing device 2 using the generating AI.

[0073] [Configuration of Information Processing Device 2] Figure 4 shows the configuration of the information processing device 2. The information processing device 2 includes a device communication unit 21, a storage unit 22, and a control unit 23. The control unit 23 includes an information acquisition unit 231, a decision unit 232, a notification unit 233, an audio data acquisition unit 234, a prompt creation unit 235, and a summary creation unit 236.

[0074] The device communication unit 21 has a communication interface for sending and receiving various types of data between the information terminal 1, the generation AI server 3, and the customer management server 4. The device communication unit 21 receives voice data from the information terminal 1, sends the voice data and a prompt, which is instruction data for instructing the generation AI server 3 to create a summary, and outputs the summary data received from the generation AI server 3 to the customer management server 4.

[0075] The storage unit 22 has storage media such as ROM, RAM, and SSD (Solid State Drive). The storage unit 22 stores programs executed by the control unit 23. The storage unit 22 also stores various data used to determine whether or not the information terminal 1 is staying at a visited location. Specifically, the storage unit 22 stores information indicating the attributes of multiple users U (e.g., gender, age, etc.) associated with information for identifying user U (e.g., name or employee code). The storage unit 22 also stores visit candidate location data indicating potential places that user U may visit, geofence data which is an example of visit area data indicating the area of ​​each visit location, and user location data indicating the location of user U.

[0076] Figure 5 shows an example of potential visit location data. In the potential visit location data, the customer ID, customer attributes, potential visit location, and geofence ID are associated.

[0077] "Customer ID" is information used to identify a customer. "Customer attributes" include industry, location, environment of the location, company size, capital structure, etc. "Potential visit locations" are places that user U may visit, such as the customer's name, the name of a building on the customer's premises, etc. As shown in Figure 5, multiple potential visit locations may be set for a single customer.

[0078] The "Geofence ID" is information used to identify a geofence and is set for each of the multiple potential visit locations. In the potential visit location data, a delay time may be associated with each potential visit location from the time that information terminal 1 enters the visit area until information terminal 1 executes the recording start process.

[0079] Figure 6 shows an example of geofence data. In the geofence data, the geofence ID, the geofence's center position (latitude and longitude), and the geofence's radius are associated. The geofence's center position and radius are set based on the shape and size of the customer's business premises.

[0080] Figure 7 shows an example of user location data. In user location data, the user ID, time, location, and stay flag are associated. The "user ID" is information used to identify the user. The "time" is the time when information terminal 1 notified information processing device 2 of its location. In Figure 7, the time is shown in 1-minute intervals, but this time interval is arbitrary.

[0081] In Figure 7, "Location" indicates the location of information terminal 1, and is represented, for example, by latitude and longitude. The user location data shown in Figure 7 indicates that locations P0, P1, P5, and P6 are not included in the geofence of the destination visited by user U, while locations P2 to P4 are included in the geofence of the destination. The "Stay Flag" is data indicating whether or not user U is staying within the customer's geofence. The user location data shown in Figure 7 indicates that user U stayed within the geofence between 10:02 and 10:57. In the user location data shown in Figure 7, location P3 corresponds to the location of the conference room, illustrating an example where user U arrives at the main gate of the destination at 10:02, moves around the business premises of the destination, arrives at the conference room at 10:03, and stays in the conference room until 10:56.

[0082] The memory unit 22 also stores data used by the summary creation unit 236 when creating summary data. For example, the memory unit 22 stores prompt templates that the summary creation unit 236 sends to the generation AI server 3 when creating summary data. The prompt template is data that includes information to specify the content of the summary data, such as the items to be included in the summary data, the length of the summary data, and the specialized terminology that can be used in the summary data.

[0083] The storage unit 22 may store multiple prompt templates indicating how to create summary data, associated with at least one of the user's attributes or the attributes of the visited location (industry, business type, business scale, etc.). The storage unit 22 may also store template selection data indicating the degree of matching of the multiple templates, associated with the customer's attributes.

[0084] Figure 8 shows an example of template selection data. In the template selection data shown in Figure 8, customer attributes, indicated by the customer's industry, business type, and company size, are associated with the degree of matching for each of the four templates A to D. The higher the numerical value shown in Figure 8, the greater the degree of matching. Since the specialized terminology used in meetings and the items that should be recorded differ depending on the customer's attributes, using different prompt templates depending on the customer's attributes makes it easier to create summary data that is appropriate for the customer's attributes.

[0085] The degree of matching is a probability value inferred, for example, using logistic regression. This probability value is generated by a computer creating template selection data by performing the following steps.

[0086] First, the computer creates training data that associates customer attributes with the template selected by that customer. In this training data, customer attributes are linked to a template type (for example, one of A through D). Attributes can be either categorical or continuous variables.

[0087] Next, the computer uses a classifier such as logistic regression to create a learning model that takes customer attribute values ​​as input features and outputs predicted values ​​for which each template type will be selected. Finally, by inputting the attribute values ​​into the created learning model, the computer can create template selection data, as shown in Figure 8, in which customer attributes are associated with the degree of matching for each template.

[0088] The memory unit 22 may further store generation AI information that shows the characteristics of each generation AI used when creating summary data. In the generation AI information, information for identifying multiple generation AIs (e.g., the names of the generation AIs) and the characteristics of the training audio data are associated. The characteristics of the generation AI are, for example, the encoding scheme or sampling rate of the training audio data used by the generation AI for training. The memory unit 22 may also store data showing the relationship between the sampling rate and the character error rate, as shown in Figure 3.

[0089] Returning to Figure 4, the details of the control unit 23 will be explained. The control unit 23 functions as an information acquisition unit 231, a decision unit 232, a notification unit 233, a voice data acquisition unit 234, a prompt creation unit 235, and a summary creation unit 236 by executing the program stored in the memory unit 22.

[0090] The information acquisition unit 231 acquires various types of information. For example, the information acquisition unit 231 acquires information from the information terminal 1 indicating that the information terminal 1 has entered the visited area G. This information includes information for identifying user U, as well as information for identifying the visited area G (for example, a destination ID).

[0091] The information acquisition unit 231 may acquire information from the information terminal 1 indicating the status of the communication environment of the information terminal 1. The information indicating the status of the communication environment may be, for example, the communication speed or data error rate identified by the information terminal 1. The information indicating the status of the communication environment may also be information indicating a communication environment level, such as "good," "normal," or "bad." The information acquisition unit 231 notifies the determination unit 232 of the information indicating the status of the communication environment.

[0092] The information acquisition unit 231 may acquire information from the information terminal 1 indicating the voice processing performance of the information terminal 1. Voice processing performance is represented, for example, by the voice encoding scheme that the information terminal 1 can use, the model name of the processor installed in the information terminal 1, or the memory capacity. The information acquisition unit 231 notifies the determination unit 232 of the acquired information indicating the voice processing performance.

[0093] The information acquisition unit 231 may acquire information indicating the attributes of user U or the attributes of other people with whom user U is conversing. These attributes are those that can affect the characteristics of the voice, such as gender, age, and voice volume. The information acquisition unit 231 may acquire information indicating attributes set by the information terminal 1 from the information terminal 1, or it may acquire information indicating attributes that have been registered in advance in association with user U or other people from an external device or storage unit 22.

[0094] The determination unit 232 refers to the generation AI information, which associates multiple generation AIs with the characteristics of the training audio data, and determines the characteristics of the audio data to be input to the generation AI based on the characteristics of the training audio data associated with the generation AI used by the summarization unit 236 to create the summary data. The determination unit 232 refers to the characteristics of each of the multiple generation AIs stored in the storage unit 22 and identifies an encoding scheme or sampling rate suitable for the audio data to be input to the generation AI used by the summarization unit 236. The notification unit 233 notifies the information terminal 1, which generates the audio data, of the characteristics determined by the determination unit 232. The notification unit 233 may also notify the information terminal 1 of multiple encoding schemes or multiple sampling rates or a range of usable sampling rates.

[0095] The determination unit 232 may determine the characteristics of the voice data to be input to the generating AI based on the state of the communication environment of the information terminal 1 indicated by the information acquisition unit 231. For example, the determination unit 232 lowers the sampling rate as the communication speed slows down or the data error rate increases. In this case, the determination unit 232 may determine the sampling rate of the voice data to be input to the generating AI to be lower than the sampling rate of the training voice data. If the communication speed of the information terminal 1 is below a threshold, the determination unit 232 may determine the voice compression method, which is a characteristic of the voice data to be input to the generating AI, to be a voice compression method with a higher compression ratio than the voice compression method of the training voice data.

[0096] As explained with reference to Figure 3, even if the sampling rate is lower than the sampling rate of the training audio data (e.g., 16 kHz), as long as it is within a predetermined range (e.g., 10 kHz or higher), the character error rate when the generating AI converts the audio data into text will remain below a certain range. Therefore, the determination unit 232 may reduce the sampling rate to a predetermined lower limit as the communication speed of the information terminal 1 that transmits the audio data to the information processing device slows down. This reduces the amount of audio data, thereby suppressing the number of character errors in the summarized data while reducing data communication costs.

[0097] The determination unit 232 may further determine the sampling rate of the audio data input to the generating AI based on the audio processing performance indicated by the information acquisition unit 231. This makes it possible for the information processing device 2 to stably create summary data based on the audio data recorded by the information terminal 1, even if the processing performance of the information terminal 1 is low.

[0098] Incidentally, the character error rate when the generating AI converts speech to text can vary depending on the quality of the input speech. For example, the clearer the speech, the lower the character error rate is likely to be. Therefore, the determination unit 232 may determine the sampling rate of the speech data input to the generating AI based on the attributes of user U or the attributes of the other person with whom user U is conversing, as indicated by the information acquired by the information acquisition unit 231. As an example, the determination unit 232 may set a smaller sampling rate when the speech data indicates the voice of a young woman, who tends to have a clear voice, compared to when it indicates the voice of an elderly man, who tends to have a unclear voice. This makes it possible to reduce the amount of speech data while reducing the error rate when the generating AI converts speech to text.

[0099] The voice data acquisition unit 234 acquires voice data transmitted from the information terminal 1 via the device communication unit 21. The voice data acquisition unit 234 stores the acquired voice data in the storage unit 22.

[0100] The prompt creation unit 235 creates a prompt indicating how to create summary data based on the audio data stored in the storage unit 22. The method for creating the summary data is indicated by the format of the summary data, the content to be included in the summary data, or the length of the summary data. For example, the prompt creation unit 235 refers to the template selection data shown in Figure 8, selects a prompt template based on the degree of matching between each of the multiple prompt templates and the attributes of user U or the visited location, and creates a prompt based on the selected prompt template. The prompt creation unit 235 inputs the created prompt to the summary creation unit 236.

[0101] The prompt creation unit 235 may select one or more prompt templates based on the degree of matching between each of the multiple prompt templates and the attributes of user U or the visited location, and send information indicating the one or more selected prompt templates (for example, the template name or information indicating the template's overview) to the information terminal 1. The information terminal 1 displays the information indicating the one or more prompt templates, and the prompt creation unit 235 may create a prompt based on the prompt template selected by user U from the one or more selected prompt templates. This ensures that even if multiple prompt templates are available to user U, summary data that matches user U's preferences or purpose is created.

[0102] The prompt creation unit 235 may create a prompt based on a prompt template selected from multiple prompt templates based on the items to be included in the summary data received by the information acquisition unit 231 from the user U. For example, the prompt creation unit 235 selects a prompt template suitable for the summary items by referring to template selection data in which summary items and prompt templates are associated.

[0103] The summary creation unit 236 inputs the audio data received from the audio data acquisition unit 234 and the prompts received from the prompt creation unit 235 to the generation AI, and creates summary data based on the summary text output from the generation AI. Specifically, the summary creation unit 236 transmits the audio data and prompts to the generation AI server 3 via the device communication unit 21. The summary creation unit 236 creates summary data based on the audio data by acquiring the summary data generated by the generation AI server 3 from the audio data based on the summary data creation method indicated by the prompts. The summary creation unit 236 may also create the output summary data by processing the summary data transmitted from the generation AI server 3.

[0104] However, there may be cases where user U is not satisfied with the content of the summary data created by the summary creation unit 236. In this case, the summary creation unit 236 may obtain a request to correct the summary data from the information terminal 1, correct the summary data based on the content of the correction request, and send the corrected summary data to the information terminal 1. The summary creation unit 236 may also communicate the correction content to the prompt creation unit 235, and the prompt creation unit 235 may update the summary data by sending the corrected prompt to the generation AI server 3.

[0105] In such cases, to increase the probability that summary data that suits user U's preferences or purposes will be created in the future, the prompt creation unit 235 may create a revised prompt by transmitting the prompt used to create the summary data to the information terminal 1 via the device communication unit 21 and accepting revisions to the prompt by user U. The summary creation unit 236 creates the revised summary data by transmitting the revised prompt to the generation AI server 3. The prompt creation unit 235 may store the created revised prompt in the storage unit 22 in association with user U. The prompt creation unit 235 may store the revised prompt in the storage unit 22 in association with user U, provided that user U has approved that the summary data created using the revised prompt is acceptable.

[0106] After the correction prompt is stored in the storage unit 22, when the audio data acquisition unit 234 acquires the audio data transmitted from the information terminal 1, the summary creation unit 236 inputs the correction prompt stored in the storage unit 22 in association with user U to the generation AI (i.e., sends it to the generation AI server 3). As a result, the next time user U requests the creation of summary data, the summary creation unit 236 can create the summary data using the correction prompt, increasing the probability that summary data that matches user U's preferences or purpose will be created.

[0107] [Processing flow at information terminal 1] Figure 9 is a flowchart showing the processing flow in information terminal 1. The flowchart shown in Figure 9 starts from the moment the application software for recording audio is launched.

[0108] The location identification unit 171 identifies the current location of the information terminal 1 (S11) and monitors whether the information terminal 1 has entered a pre-set visit area (e.g., a geofence) (S12). If the location identification unit 171 determines that the information terminal 1 has entered a geofence (YES in S12), the data communication unit 173 notifies the generation AI server 3 of the status of the communication environment and obtains the sampling rate to be used when recording voice from the information processing device 2 (S13).

[0109] The recording processing unit 174 displays a button for starting recording on the display unit 14 (S14) and monitors the operation of pressing the displayed button (S15). When the recording start button is operated (YES in S15), the recording processing unit 174 starts recording at the sampling rate acquired by the data communication unit 173 (S16). During this time, the location identification unit 171 monitors whether the information terminal 1 has left the geofence (S17). While the location identification unit 171 determines that the information terminal 1 has not left the geofence (NO in S17), the recording processing unit 174 continues recording and transmits the audio data to the information processing device 2 via the data communication unit 173.

[0110] If the location identification unit 171 determines that information terminal 1 has left the geofence (YES in S17), the recording processing unit 174 terminates recording (S18). The recording processing unit 174, for example, displays a button to terminate recording on the display unit 14, and terminates recording when the displayed button is pressed.

[0111] [Processing flow in Information Processing Device 2] Figure 10 is a flowchart showing the processing flow in the information processing device 2. When the information acquisition unit 231 acquires information from the information terminal 1 that the information terminal 1 has entered a geofence, thereby determining that the information terminal 1 has entered a geofence (S31), the determination unit 232 determines the state of the communication environment of the information terminal 1 based on the information indicating the state of the communication environment transmitted from the information terminal 1 (S32). The determination unit 232 determines the sampling rate based on the determined state of the communication environment and notifies the information terminal 1 of the determined sampling rate (S33).

[0112] Subsequently, the audio data acquisition unit 234 acquires audio data generated by sampling the audio at the sampling rate notified by the determination unit 232 (S34), and stores the acquired audio data in the storage unit 22.

[0113] Next, the prompt creation unit 235 identifies at least one of the user U's attributes or the attributes of the destination by referring to the data stored in the storage unit 22 (S35), and selects a prompt template corresponding to the identified attribute by referring to the template selection data (S36).

[0114] The summary creation unit 236 sends the prompt created based on the prompt template selected by the prompt creation unit 235 to the generation AI server 3 along with the audio data (S37), and retrieves the text data generated from the generation AI server 3 based on the audio data (S38). The summary creation unit 236 outputs the summary data created based on the retrieved text data to the information terminal 1 (S39).

[0115] If, after the summary creation unit 236 outputs summary data to the information terminal 1, the prompt creation unit 235 receives a request from the information terminal 1 to modify the summary data (YES in S40), the summary creation unit 236 updates the summary data by sending a prompt updated by the prompt creation unit 235 based on the content to be modified to the generation AI server 3 (S41).

[0116] [Differentiation] In the above description, summary data was created by the information terminal 1 transmitting recorded audio data to the information processing device 2, and the information processing device 2 transmitting the audio data and prompts to the generation AI server 3. However, some of the functions performed by the information processing device 2 may be performed by the information terminal 1. For example, the information terminal 1 may create a prompt and obtain summary data by directly transmitting the created prompt and audio data to the generation AI server 3.

[0117] Furthermore, in the above explanation, the information processing device 2 created summary data by sending audio data and prompts to the generating AI server 3. However, the information processing device 2 does not need to create summary data by sending audio data and prompts to another generating AI server that has the function of converting audio data into text data, thereby creating text data. If the audio data is generated at a sampling rate corresponding to the sampling rate of the audio data used by the generating AI server for machine learning, the error rate when converting from audio data to text data will be reduced.

[0118] Furthermore, in the above explanation, the information terminal 1 transmits location data indicating its current location to the information processing device 2, and the information processing device 2 receives stay flag data to determine whether or not the current location is included in the visited area. However, the information terminal 1 may also store geofence data indicating the visited area. In this case, the information terminal 1 can determine whether or not it is included in the visited area without transmitting location data to the information processing device 2.

[0119] [Effects of the summary generation system S] As explained above, when information terminal 1 detects that its current location has changed from being outside the visit area to being included in the visit area, it executes a process to start recording. This prevents user U from forgetting to record audio of meetings at visit locations, allowing user U to concentrate on sales activities and other tasks.

[0120] Furthermore, when creating summary data using a generative AI capable of generating summaries based on speech, a problem arises in that the quality of the summary data deteriorates if the characteristics of the speech data used by the generative AI for training do not match the characteristics of the speech data generated by the information terminal 1 through recording. In contrast, the summary creation system S can determine the characteristics of the speech data generated by the information terminal 1 based on the characteristics of the speech data used by the generative AI for training, thereby suppressing the deterioration of the quality of the summary data.

[0121] Furthermore, depending on the communication environment at the location where user U is staying, a problem may arise where information terminal 1 cannot transmit large amounts of audio data to information processing device 2. In response to this, the summarization system S can change the sampling rate or encoding method used by information terminal 1 when generating audio data depending on the communication environment. Therefore, by having information terminal 1 generate audio data using an appropriate sampling rate or encoding method determined based on the characteristics of the audio data used by the generating AI for learning and the state of the communication environment, it becomes possible to ensure the quality of the summary data even when the communication environment is poor.

[0122] Furthermore, the summary data created by the information processing device 2 is associated with the visited location (customer) where the information terminal 1 was located at the time the audio data was generated, and is sent to the customer management server 4, where it is registered in the customer management server 4's database. Therefore, user U does not need to perform any work to save the summary data to the customer management server 4.

[0123] Although the present invention has been described above using embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of its gist. For example, all or part of the apparatus can be configured by functionally or physically distributing and integrating in any unit. Furthermore, new embodiments resulting from any combination of multiple embodiments are also included in the embodiments of the present invention. The effects of the new embodiments resulting from the combinations are combined with the effects of the original embodiments. [Explanation of symbols]

[0124] 1. Information terminal 2. Information Processing Device 3. Generation AI Server 4. Customer Management Server 11 GPS section 12 Microphones 13 Terminal Communication Unit 14 Display section 15 Control section 16 Memory section 17 Control Unit 21. Device Communication Unit 22 Memory section 23 Control Unit 171 Location identification part 172 Reception Department 173 Data Communications Department 174 Recording Processing Unit 231 Information Acquisition Department 232 Decision Section 233 Notification Department 234 Audio data acquisition unit 235 Prompt Creation Section 236 Summary Preparation Department

Claims

1. It is an information terminal, A location identification unit that identifies the current location of the information terminal, A recording processing unit, which, by referring to visit area data indicating a visit area corresponding to a visit location visited by the user of the information terminal, detects that the current location has changed from a state where it is not included in the visit area indicated by the visit area data to a state where it is included in the visit area, executes a process to start recording. It has, The recording processing unit is an information terminal that, by referring to location data associated with the visited location and the delay time from when the current location is included in the visited area until recording begins, executes a process to start recording after the delay time has elapsed from when the current location is included in the visited area.

2. The recording processing unit refers to the visited area data, performs a process to start recording, and then terminates the recording when the current location changes from being included in the visited area to being no longer included in the visited area. The information terminal according to claim 1.

3. The system further includes a receiving unit that accepts a setting for the delay time from when the current location is included in the visit area until recording begins. The recording processing unit executes a process to start recording after the current location is included in the visited area and the delay time has elapsed. The information terminal according to claim 1.

4. The recording processing unit, upon realizing that the current location is included in the visit area corresponding to the visit location, displays an operation image on the display unit to accept an operation to start recording, and starts recording upon realizing that the operation image has been operated. The information terminal according to claim 1.

5. The device further includes a data communication unit that transmits recorded audio data to an information processing device that performs a process to summarize the content of the audio based on the audio data. The information terminal according to claim 1.

6. The data communication unit identifies the state of the communication line that transmits the voice data, The recording processing unit generates the audio data by sampling the audio at a sampling rate determined based on the state of the communication line identified by the data communication unit. The information terminal according to claim 5.

7. On the computer, The steps include determining the current location of the computer, The steps include:

1. Referencing visit area data indicating a visit area corresponding to a visit location visited by the computer user, and detecting that the current location has changed from a state where it is not included in the visit area indicated by the visit area data to a state where it is included in the visit area, and then referencing location data associated with the visit location and the delay time from when the current location becomes included in the visit area until recording begins, to perform a process to start recording after the delay time has elapsed from when the current location becomes included in the visit area; A program to execute.

8. The system comprises an information terminal according to any one of claims 1 to 6, and an information processing device that creates summary data based on voice data transmitted from the information terminal, The aforementioned information processing device is A voice data acquisition unit that acquires the aforementioned voice data, A prompt creation unit that creates a prompt indicating a method for creating the summary data based on the aforementioned audio data, A summarization unit inputs the aforementioned audio data and the aforementioned prompt to a generating AI, and creates the summary data based on the summary text output from the generating AI. A summary generation system having the following features.

9. The system further includes a storage unit that stores a plurality of prompt templates indicating the method for creating the summary data, associated with at least one of the user's attributes or the attributes of the visited location. The prompt creation unit selects a prompt template based on the degree of matching between each of the plurality of prompt templates and the user's attributes or the attributes of the visited location, and creates the prompt based on the selected prompt template. The summary creation system according to claim 8.

10. The prompt creation unit selects one or more prompt templates based on the degree of matching between each of the plurality of prompt templates and the user's attributes or the attributes of the visited location, and creates the prompt from the one or more selected prompt templates based on the prompt template selected by the user. The summary creation system according to claim 9.

11. The prompt creation unit creates a modified prompt by receiving a modification of the prompt by the user, and stores the created modified prompt in the storage unit in association with the user. The summary generation unit, after the correction prompt has been stored in the storage unit, inputs the correction prompt stored in the storage unit in association with the user to the generation AI when the voice data acquisition unit acquires the voice data transmitted from the information terminal used by the user. The summary creation system according to claim 8.

12. The aforementioned information processing device is A determination unit determines the characteristics of the audio data input to a generation AI that generates a summary text based on audio data, based on the characteristics of the training audio data used for machine learning by the generation AI. The information terminal that generates the aforementioned voice data is equipped with a notification unit that notifies the determined characteristics, It further possesses, The summary creation unit inputs the audio data generated by the information terminal, which has the characteristics determined by the determination unit, into the generation AI, and creates the summary data based on the audio data based on the summary text output from the generation AI. The summary creation system according to claim 8.

13. The determination unit refers to the generation AI information, which associates a plurality of generation AIs with the characteristics of the learning audio data, and determines the characteristics of the audio data to be input to the generation AI based on the characteristics of the learning audio data associated with the generation AI used by the summarization unit to create the summary data. The summary creation system according to claim 12.

14. The information processing device further includes an information acquisition unit that acquires from the information terminal information indicating the state of the communication environment of the information terminal that transmits the voice data to the information processing device, The determination unit further determines the characteristics of the voice data to be input to the generating AI based on the state of the communication environment of the information terminal indicated by the information acquisition unit. The summary creation system according to claim 12.

15. The determination unit reduces the sampling rate, which is a characteristic of the voice data input to the generating AI, to a predetermined lower limit as the communication speed of the information terminal that transmits the voice data to the information processing device slows down. The summary creation system according to claim 14.

16. The determination unit determines the sampling rate of the audio data input to the generating AI to be lower than the sampling rate of the training audio data. The summary creation system according to claim 12.

17. The determination unit, when the communication speed of the information terminal that transmits the audio data to the information processing device is below a threshold, determines the audio compression method, which is a characteristic of the audio data input to the generating AI, to be an audio compression method with a higher compression ratio than the audio compression method of the learning audio data. The summary creation system according to claim 12.

18. The information processing device further includes an information acquisition unit that acquires from the information terminal information indicating the voice processing performance of the information terminal that transmits the voice data to the information processing device, The determination unit determines the sampling rate of the audio data to be input to the generating AI, based on the audio processing performance indicated by the information acquired by the information acquisition unit. The summary creation system according to claim 12.