Processing device, processing method, and processing program

The system automates the creation and management of meeting minutes using AI, addressing the inefficiency of manual searching in web conferencing by providing immediate and organized access to meeting data.

JP2025180827APending Publication Date: 2025-12-11NTT DOCOMO BUSINESS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024088431
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-30
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing web conferencing systems require users to manually search for meeting phrases during the conference, disrupting the meeting flow and causing inefficiencies.

Method used

A processing device and method that utilizes a server device and generation AI to automatically convert voice data to text, create meeting minutes, and manage them in a database, allowing for seamless integration and retrieval of meeting information without user intervention.

Benefits of technology

Enables smooth conference progression by reducing user burden and allowing immediate access to meeting minutes, with features like anonymization and easy retrieval of related news, enhancing meeting efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025180827000001_ABST
    Figure 2025180827000001_ABST
Patent Text Reader

Abstract

To support a smooth progress of a conference while alleviating burden on a user.SOLUTION: A server device 310 includes: an acquisition part that acquires information related to a conference; a recognition part 134 that converts voice data of each user participating in the conference into a text data and associates time information with each text; a determination part 3132 that determines, based on the text data, whether or not it is necessary to search for a word uttered at the time of the conference; a search execution part 3133 that searches for a word to be searched for when a search is necessary; and a search result display control part 3134 that displays a search result searched by the search execution unit 3133 on each user terminal used by each user.SELECTED DRAWING: Figure 22
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a processing device, a processing method, and a processing program. [Background technology]

[0002] In recent years, web conferencing services that connect via applications or browsers have become widespread. In these services, the equipment that provides the conferencing service is installed on the network, and users participate in the conference using applications or browsers running on their terminals. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-230532 Summary of the Invention [Problem to be solved by the invention]

[0004] To participate in a web conference, a user installs software that runs on a terminal device, for example, and then starts up the terminal device, launches the installed software, and then participates in the web conference using the registered conference ID.

[0005] During a meeting, a user may wish to search for phrases that were spoken during the meeting, etc. However, if a user performs a cumbersome process such as a search during the meeting, the meeting may be interrupted, and the meeting may not proceed smoothly.

[0006] Therefore, the present invention has been made in consideration of the above, and aims to provide a processing device, processing method, and processing program that enable appropriate minutes to be created while reducing the burden on the user. [Means for solving the problem]

[0007] In order to solve the above-mentioned problems and achieve the object, the processing device of the present invention is characterized by having an acquisition unit that acquires information about a conference, a recognition unit that converts the voice data of each user who participated in the conference into text data and associates time information with each text, a determination unit that determines whether or not it is necessary to search for words spoken during the conference based on the text data, a search execution unit that searches for the words to be searched for if the search is necessary, and a display control unit that displays the search results searched by the search execution unit on each user terminal used by each user. [Effects of the Invention]

[0008] According to the present invention, it is possible to support the smooth progress of a conference while reducing the burden on the user. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a processing system according to the first embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of the configuration of the server device illustrated in FIG. [Figure 3] FIG. 3 is a diagram illustrating an example of a data configuration of the relationship information illustrated in FIG. [Figure 4] FIG. 4 is a diagram illustrating the flow of processing in the processing system shown in FIG. [Figure 5] FIG. 5 is a diagram illustrating the flow of processing in the processing system shown in FIG. [Figure 6] FIG. 6 is a diagram illustrating the flow of processing in the processing system shown in FIG. [Figure 7] FIG. 7 is a sequence diagram showing the processing procedure of the processing method according to the first embodiment. [Figure 8] FIG. 8 is a diagram illustrating an example of a configuration of a processing system according to the first modification of the first embodiment. [Figure 9] FIG. 9 is a diagram illustrating a processing flow of the processing system according to the first modification of the first embodiment. [Figure 10]FIG. 10 is a sequence diagram showing a processing procedure of a processing method according to the first modification of the first embodiment. [Figure 11] FIG. 11 is a diagram illustrating another example of the configuration of the processing system according to the first modification of the first embodiment. [Figure 12] FIG. 12 is a diagram illustrating an example of the configuration of a processing system according to the second embodiment. [Figure 13] FIG. 13 is a diagram illustrating the flow of processing in the processing system according to the second embodiment. [Figure 14] FIG. 14 is a diagram showing an example of a screen of a user terminal. [Figure 15] FIG. 15 is a sequence diagram showing the processing procedure of the processing method according to the second embodiment. [Figure 16] FIG. 16 is a diagram illustrating a processing flow in the first modification of the second embodiment. [Figure 17] FIG. 17 is a sequence diagram showing a processing procedure of a processing method according to the first modification of the second embodiment. [Figure 18] FIG. 18 is a diagram illustrating an example of a configuration of a processing system according to Modification 2 of Embodiment 2. In FIG. [Figure 19] FIG. 19 is a diagram illustrating a processing flow in the first modification of the second embodiment. [Figure 20] FIG. 20 is a sequence diagram showing a processing procedure of a processing method according to the second modification of the second embodiment. [Figure 21] FIG. 21 is a diagram illustrating another example of the configuration of the processing system according to the second modification of the second embodiment. [Figure 22] FIG. 22 is a diagram illustrating an example of the configuration of a processing system according to the third embodiment. [Figure 23] FIG. 23 is a diagram illustrating the flow of processing in the processing system according to the third embodiment. [Figure 24] FIG. 24 is a diagram illustrating the flow of processing in the processing system according to the third embodiment. [Figure 25] FIG. 25 is a diagram showing an example of a screen of a user terminal. [Figure 26]FIG. 26 is a sequence diagram showing the processing procedure of the processing method according to the third embodiment. [Figure 27] FIG. 27 is a diagram illustrating an example of a configuration of a processing system according to the first modification of the third embodiment. [Figure 28] FIG. 28 is a diagram illustrating a processing flow in the first modification of the third embodiment. [Figure 29] FIG. 29 is a sequence diagram showing a processing procedure of a processing method according to the first modification of the third embodiment. [Figure 30] FIG. 30 is a diagram illustrating another example of the configuration of the processing system according to the second modification of the third embodiment. [Figure 31] FIG. 31 is a diagram illustrating an example of the configuration of a processing system according to the fourth embodiment. [Figure 32] FIG. 32 is a diagram illustrating an example of a computer that implements a server device by executing a program. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to this embodiment. In addition, in the description of the drawings, the same parts are designated by the same reference numerals.

[0011] [Embodiment 1] [Processing System] The following describes the configuration of a processing system according to embodiment 1. The processing system according to embodiment 1 automatically creates minutes of a conference (for example, a web conference) between multiple users by using a generative AI (artificial intelligence) (generative model).

[0012] 1 is a diagram illustrating an example of the configuration of a processing system according to embodiment 1. As illustrated in FIG. 1, the processing system 1 includes a server device 10, user terminals 20a to 20d, and a generation AI server 30. The server device 10 communicates with the user terminals 20a to 20d and the generation AI server 30.

[0013] The user terminals 20a to 20d are terminals used by the users Ua to Ud. The user terminals 20a to 20d are, for example, personal computers (PCs), notebook PCs, tablet terminals, smartphones, etc. The user terminals 20a to 20d are terminal devices that can input and output voice data and text data and communicate with the server device 10. The users Ua to Ud participate in, for example, a web conference held between the users Ua to Ud via an application or browser running on the user terminals 20a to 20d. In the following description, the user terminals 20a to 20d will be referred to as user terminals 20 when no particular distinction is made. The number of user terminals 20 may be two or more and is not limited to four.

[0014] The generation AI server 30 is equipped with a generation AI 31 (generation model), which is a natural language processing model. The generation AI 31 performs natural language processing on input text data in accordance with set prompts, creates minutes, and outputs them. The generation AI 31 is, for example, a large-scale natural language processing model.

[0015] The server device 10 communicates with the generation AI server 30. The server device 10 communicates with the user terminals 20a to 20d. The server device 10 converts the voice data of the users Ua to Ud during the conference into text data and transmits it to the generation AI server 30.

[0016] The server device 10 stores the list of meeting agenda items, the minutes output from the generation AI server 30, and the identification information of the meeting in association with each other in a minutes database (DB). In this way, the server device 10 automatically creates the minutes of the meeting and stores them in a database.

[0017] [Server device] Next, a description will be given of the server device 10. Fig. 2 is a diagram showing an example of the configuration of the server device 10 shown in Fig. 1. As shown in Fig. 2, the server device 10 includes a communication unit 11, a storage unit 12, and a control unit 13.

[0018] The communication unit 11 is a communication interface that transmits and receives various information to and from other devices connected via a network, etc. The communication unit 11 is realized by a NIC (Network Interface Card) or the like, and performs communication between the control unit 13 (described later) and other devices (e.g., user terminals 20a to 20d, generation AI server 30) via telecommunication lines such as a LAN (Local Area Network) or the Internet.

[0019] The storage unit 12 is realized by semiconductor memory elements such as RAM (Random Access Memory) and flash memory, and stores processing programs that operate the server device 10, data used during execution of the processing programs, etc. The storage unit 12 has user information 121, schedule information 122, conference information 123, text data 124, a minutes DB 125, and a recorded audio DB 126.

[0020] The user information 121 is information including the ID, name, department, position, and project in charge of each of the users Ua to Ud who use the user terminals 20a to 20d.

[0021] The schedule information 122 is a conference schedule for each of the users Ua to Ud. The conference schedule includes, for example, the conference room (in the case of a Web conference, the Web conference ID, etc.), date and time, agenda, members participating in the conference, and information on related conferences and projects.

[0022] The conference information 123 includes the date and time of the conference, the users participating in the conference, a conference summary, a list of conference topics, etc. The conference summary is, for example, input by the users participating in the conference before the conference. The list of conference topics is a list of topics to be discussed in the conference, and is, for example, input by the users participating in the conference before the conference. The conference information 123 may also include the level of confidentiality of each conference.

[0023] The text data 124 is text data converted from the voice data by a recognition unit 134 (described later). The voice data is voice data of each user Ua to Ud who participated in the conference, and is transmitted from the user terminals 20a to 20d.

[0024] The minutes DB 125 is a DB in which minutes of meetings are stored. The minutes of meetings are generated by the generation AI 31. The minutes DB 125 stores the ID of each meeting and the minutes in association with each other. The minutes DB 125 stores relationship information indicating the correspondence between the minutes along with the minutes. The minutes are grouped by minutes that share the same meeting topic, members, or keywords, or by minutes that have similar content, and each minutes ID is associated with identification information of the group to which the minutes corresponding to this minutes ID belong as relationship information. Furthermore, news related to the meeting (related news) is linked to each minutes as relationship information.

[0025] Fig. 3 is a diagram showing an example of the data configuration of the related information shown in Fig. 2. As shown in Fig. 3, the related information includes the following items: minutes ID, conference ID, group, member, keyword, topic, confidentiality, and related news.

[0026] Minutes that share any of the following are classified into the same group: meeting topic, members, or keywords; or minutes with similar content. Keywords include project identification information, words that appear within a predetermined order of frequency during the meeting, people's names, and the members' affiliations and job titles. Confidentiality may be a level of confidentiality set in advance for the meeting, or a level of confidentiality estimated from the meeting members, the content of the meeting, and / or the meeting topic. Related news may be, for example, a document file containing the news or URL (Uniform Resource Locator) information containing the news. Related news may include internal news related to the meeting, external release information, or general news that was discussed during the meeting.

[0027] For example, the minutes ID "D1" is associated with the meeting ID "M1," as well as related information such as group "G1," members "Ua, Ub, Uc," keyword "Project B," agenda item "Development Progress," confidentiality "High," and related news "N1."

[0028] The recorded audio DB 126 is a DB that stores audio data and / or image data of the speech of each user in a conference. The recorded audio DB 126 stores the ID of each conference in association with the recorded audio and / or recorded image. Each entry in the minutes of the conference and the playback position of the recorded audio or recorded image of the conference corresponding to each entry are associated with each other by the association unit 137 (described later).

[0029] The control unit 13 controls the entire server device 10. The control unit 13 is, for example, an electronic circuit such as a CPU (Central Processing Unit), or an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array). The control unit 13 also has an internal memory for storing programs defining various processing procedures and control data, and executes each process using the internal memory. The control unit 13 also functions as various processing units when various programs are run. The control unit 13 has a schedule acquisition unit 131 (acquisition unit), a conference information acquisition unit 132 (acquisition unit), a prompt creation unit 133, a recognition unit 134, a preprocessing unit 135, a storage unit 136, and an association unit 137.

[0030] The schedule acquisition unit 131 acquires the schedules of the users Ua to Ud and stores them in the schedule information 122. The schedule acquisition unit 131 acquires the schedules of the users Ua to Ud, for example, by referring to the conference schedules of the users Ua to Ud registered in the company network.

[0031] The conference information acquisition unit 132 acquires conference information of conferences held by users Ua to Ud. The conference information acquisition unit 132 acquires the date and time of the conference, the users participating in the conference, a conference summary, and a list of conference agendas. For example, the conference information acquisition unit 132 causes the user terminal 20a to output a voice or text saying, "Please tell me the summary and agenda of today's conference," and acquires a summary and a list of conference agendas of the conferences held by users Ua to Ud through the voice or text input by user Ua in response. The conference information acquisition unit 132 acquires a list of conference agendas before the conference or at the start of the conference.

[0032] The prompt creation unit 133 creates a prompt that commands the creation of minutes of a meeting based on the input text data. The prompt creation unit 133 may add information about the meeting for which minutes are to be created (e.g., the schedule, an outline or agenda of the meeting, and members of the meeting) to the prompt to facilitate the creation of minutes.

[0033] The prompt creation unit 133 then sets the created prompt in the generation AI 31. The prompt creation unit 133 instructs, for example, in the prompt, to create minutes at a predetermined timing. The predetermined timing may be, for example, when the meeting ends. Alternatively, the predetermined timing may be every five minutes, when silence continues for a predetermined period of time, when a frequently used word changes, when the agenda changes, or the like, and may be set as appropriate.

[0034] The recognition unit 134 acquires voice data of each user Ua to Ud who spoke in the conference via communication with the user terminals 20a to 20d. The recognition unit 134 performs voice recognition on the acquired voice data and converts the voice data into text data. The recognition unit 134 performs voice recognition processing using a trained voice recognition model (machine learning model) that outputs text data of the spoken voice from the voice data. The recognition unit 134 also associates time information and user information indicating the speaker with the text data and inputs the text data to the generation AI 31.

[0035] The preprocessing unit 135 acquires the minutes created by the generation AI 31 and performs preprocessing. Specifically, the preprocessing unit 135 provides the minutes to any user who participated in the conference. Then, when the preprocessing unit 135 receives an instruction to amend the minutes from the user terminal 20 operated by the user, it amends the minutes in accordance with the amendment instruction. When the preprocessing unit 135 receives an amendment from the user, it inputs the amendment content to the generation AI 31, causing it to re-create the minutes, and may also cause the generation AI 31 to learn the amendment content of the minutes to improve the accuracy of the minutes generation.

[0036] The preprocessing unit 135 also determines whether there are any sections in the minutes that are subject to anonymization, and performs anonymization processing on the sections that are subject to anonymization. The preprocessing unit 135 determines that sections in the minutes that contain words that are subject to anonymization are sections that are subject to anonymization (e.g., confidential sections). The preprocessing unit 135 then performs anonymization processing, such as blanking out or blacking out, on the sections that are subject to anonymization (see Reference 1). Reference: JP 2020-149628 A

[0037] The storage unit 136 associates the minutes of the meeting output from the generation AI 31 with identification information of the meeting (for example, a meeting ID) and stores them in the minutes DB 125. The storage unit 136 stores the minutes that have been preprocessed by the preprocessing unit 135 in the minutes DB 125. The storage unit 136 stores relationship information indicating the correspondence relationship of the minutes created by the association unit 137 (described later) in the minutes DB 125 together with the minutes.

[0038] The association unit 137 groups minutes of a meeting that share any of the meeting topic, members, and keywords. The association unit 137 groups minutes of a meeting based on the meeting topic, members of the meeting, the number of occurrences of keywords (product, project), and the name of the meeting. Alternatively, the association unit 137 groups minutes of a meeting by associating minutes of a meeting that have similar content. The association unit 137 associates minutes of a meeting that have similar content by determining whether the minutes of a meeting are similar. The association unit 137 may also cause the generation AI 31 to determine whether the minutes of a meeting are similar. The association unit 137 may further classify the minutes of a meeting within a group based on their commonalities, thereby hierarchizing the minutes of a meeting.

[0039] In this way, in the minutes DB 125, minutes of related meetings are grouped so that they belong to the same group, making it easy for the user to search for minutes.

[0040] The association unit 137 may acquire news related to the meeting based on the contents of the minutes, and associate the acquired news with the minutes. The related news may be, for example, a document file containing the news or URL information containing the news. The related news may be internal news related to the meeting, external release information, or general news that was discussed during the meeting.

[0041] The associating unit 137 acquires the confidentiality level of the meeting corresponding to the minutes and associates it with the minutes. The associating unit 137 determines the confidentiality level of the meeting corresponding to the minutes based on the settings of the meeting corresponding to the minutes (presence or absence of a non-public schedule and a tag indicating high confidentiality). If the confidentiality level of another related meeting and minutes is set to be the highest, the associating unit 137 determines that the minutes to be determined also have the highest confidentiality level. Furthermore, if a word that appeared during the meeting is, for example, a word that is contained in the minutes with the highest confidentiality level, the associating unit 137 also determines that the minutes of this meeting also have the highest confidentiality level.

[0042] The association unit 137 creates the relationship information shown in FIG. 3 based on the association process results of each of the minutes, and causes the storage unit 136 to store the relationship information in the minutes DB 125.

[0043] Furthermore, the association unit 137 associates each entry of the minutes in the minutes DB 125 with the playback position of the recorded audio or recorded image of the meeting corresponding to each entry, for each entry of the minutes and the recorded audio or recorded image of the meeting in the recorded audio DB 126. By this association, when the user selects any part of the minutes while viewing the minutes, the display transitions to the corresponding part of the recorded audio or recorded image, and the recorded audio or recorded image can be viewed.

[0044] [Processing flow] Next, a description will be given of the flow of processing in the processing system 1. Figures 4 to 6 are diagrams illustrating the flow of processing in the processing system 1 shown in Figure 1.

[0045] 4, the server device 10 communicates with, for example, the user terminal 20a used by the user Ua to acquire the conference schedule and conference information (step S1). The server device 10 inputs the conference information into a prompt that commands the creation of minutes of the conference (step S2). The server device 10 sets the prompt in the generation AI 31 (steps S3 and S4).

[0046] 5, when the conference begins, the server device 10 acquires voice data of the users Ua to Ud during the conference through communication with the user terminals 20a to 20d (steps S5-1 to S5-4). The server device 10 then performs voice recognition on the acquired voice data of the users Ua to Ud and converts the voice data into text data (step S6). The server device 10 associates time information and user information indicating the speaker with the conference text data and inputs them to the generation AI 31 (step S7).

[0047] The generation AI 31 creates the minutes of the meeting based on the input text data (step S8). The generation AI 31 may associate the user's identification information with the portion of the minutes that corresponds to the user's remarks. Then, the generation AI 31 transmits the created minutes to the server device 10 (step S9).

[0048] As shown in FIG. 6, the server device 10 provides the minutes to the user terminal 20a, for example, for viewing (step S10). When an instruction to amend the minutes is received from the user terminal 20 (step S11), the server device 10 amends the minutes in accordance with the instruction (step S12). The server device 10 also amends the minutes and performs anonymization processing on the portions of the minutes that need to be anonymized (step S12), and stores the minutes in the minutes DB 125 (step S13). The server device 10 also performs grouping of the minutes, inclusion of related news, determination of the confidentiality of the meeting, etc. (step S12), and stores related information for the minutes in the minutes DB 125. The server device 10 may associate each entry in the minutes with a playback position of recorded audio of the meeting, etc., in the recorded audio DB 126.

[0049] [Processing method] Next, a description will be given of a processing method executed by the processing system 1. Fig. 7 is a sequence diagram showing the processing steps of the processing method according to the first embodiment.

[0050] As shown in FIG. 7, for example, when a Web conference is started on the user terminals 20a to 20d via an application or a browser (steps S21-1 to S21-4), conference room information in which users Ua to Ud can participate is output to the user terminals 20a to 20d through communication between the server device 10 and the user terminals 20a to 20d (steps S22-1 to 22-4).

[0051] The server device 10 acquires the conference schedule, such as the date and time, agenda, members participating in the conference, and information on related conferences and projects (step S23). The users Ua to Ud operate the user terminals 20a to 20d to select a conference room to participate in. Then, for example, the user Ua inputs conference information into the user terminal 20a (step S24). The user terminal 20a transmits the input conference information to the server device 10 (step S25). It is sufficient for any one of the users Ua to Ud to input the conference information. Alternatively, the conference information may be registered in advance in the conference schedule, in which case steps S24 and S25 are omitted.

[0052] The server device 10 generates a prompt instructing the creation of minutes of the meeting based on the input text data (step S26). The server device 10 may add information about the meeting for which minutes are to be created to the prompt. The server device 10 sets the generated prompt in the generation AI 31 (steps S27 and S28).

[0053] When the server device 10 receives voice data of users Ua to Ud from the user terminals 20a to 20d (steps S29-1 to S29-4, S30-1 to S30-4), it performs voice recognition on the voice data and converts the voice data into text data (step S31). The server device 10 associates time information and user information indicating the speaker with the converted text data and inputs the text data to the generation AI 31 of the generation AI server 30 (step S32).

[0054] In the generation AI server 30, the generation AI 31 creates the minutes of the meeting based on the input text data (step S33). The generation AI server 30 transmits the minutes created by the generation AI 31 to the server device 10 (step S34).

[0055] The server device 10, for example, transmits the minutes to the user terminal 20a (step S35), and upon receiving an instruction to modify the minutes from the user terminal 20Ua (step S36), executes pre-processing to modify the minutes (step S37). As pre-processing, the server device 10 may perform anonymization processing on the portions of the minutes that are to be anonymized. In the server device 10, the storage unit 136 stores the pre-processed minutes of the conference in the minutes DB 125 in association with identification information of the conference (e.g., the conference ID) (step S38).

[0056] Next, the server device 10 performs a correspondence process on the minutes in the minutes DB 125 (step S39). The correspondence process includes grouping the minutes, associating the minutes with related news, associating the minutes with the confidentiality levels of the meetings corresponding to the minutes, and associating each entry in the minutes with the playback position of the recorded audio or video of the meeting corresponding to each entry.

[0057] [Effects of the First Embodiment] In this way, the server device 10 according to the first embodiment automatically creates minutes of a meeting by using the generation AI 31. Therefore, the user does not need to perform the cumbersome process of creating minutes of a meeting while playing back the recording and audio of the meeting and checking the contents of the meeting and chat. Furthermore, the user can check the minutes immediately after the end of the meeting or even during the meeting, allowing the user to review the meeting and smoothly proceed with the meeting.

[0058] Furthermore, the server device 10 performs pre-processing such as correction and concealment processing on the minutes before storing them in the minutes DB 125. Therefore, the server device 10 can store minutes with appropriate content and with confidential information concealed in a database, thereby enabling the minutes DB 125 to be utilized effectively.

[0059] The server device 10 also groups the minutes in the minutes DB 125 and associates news related to the conference with the minutes. This allows the user to easily search for related minutes and easily check news related to the conference without searching via an external server, etc. The server device 10 may display minutes of conferences related to the conference on the user terminal 20 before the conference so that the user can review the conference.

[0060] Furthermore, the server device 10 associates each entry of the minutes in the minutes DB 125 with the playback position of the recorded audio or recorded image of the meeting corresponding to each entry, for each entry of the minutes and the recorded audio or recorded image of the meeting in the recorded audio DB 126. This association allows a user to select any part of the minutes while viewing the minutes, automatically transition to the corresponding part of the recorded audio or recorded image, and view the recorded audio or recorded image. The user can quickly view the recorded audio or recorded image without having to search for a part of the recorded audio or recorded image that corresponds to the desired part of the minutes.

[0061] [First Modification of First Embodiment] Next, a first modification of the first embodiment will be described. In the first modification of the first embodiment, one of a plurality of generation AIs is selected as the generation AI that creates the minutes of the meeting, thereby enabling more appropriate minutes to be created.

[0062] Fig. 8 is a diagram showing an example of the configuration of a processing system according to Modification 1 of Embodiment 1. As shown in Fig. 8, in a processing system 1A according to Modification 1 of Embodiment 1, a generation AI server 40 is installed within an in-house network 100 in which user terminals 20a to 20d and a server device 10A are installed. In addition, the server device 10A is capable of communicating with a generation AI server 50, which is an external server.

[0063] The generation AI server 40 is equipped with Tsuzumi (registered trademark) 41 (first generation model), which is a generation AI. Tsuzumi 41 is a natural language processing model fine-tuned to a specific field. Examples of specific fields include finance, medicine, semiconductors, IT (Information Technology), academia, factories (plants), law, and office services. Tsuzumi 41 was built with an emphasis on low power consumption, and has a faster processing speed than ChatGPT 51 (described below). Note that the specific fields are not limited to those mentioned above.

[0064] The generation AI server 50 is equipped with ChatGPT (registered trademark) 51 (second generation model), which is a generation AI. ChatGPT 51 is a large-scale natural language processing model that is slower than Tsuzumi 41 but has higher accuracy.

[0065] Tsuzumi41 and ChatGPT51 perform natural language processing on the input text data according to the configured prompts, and create and output meeting minutes based on the input text data.

[0066] In a conference between users Ua to Ud, the server device 10A acquires the minutes of the conference by converting the voice data of each user who participated in the conference into text data and inputting the converted text data into Tsuzumi 41 or ChatGPT 51. Then, the server device 10A stores the minutes created by Tsuzumi 41 or ChatGPT 51 in the minutes DB 125.

[0067] Specifically, the server device 10A selects Tsuzumi41 or ChatGPT51 based on the conference information about the conference the user will be attending, the content of the conference, or predetermined rules. Then, the server device 10A sets a prompt to instruct the selected generation AI to create minutes of the conference based on the input text data. The server device 10A may add information about the conference for which minutes are to be created to the prompt to facilitate the creation of minutes.

[0068] This allows the server device 10 to obtain the most suitable minutes according to the situation from the selected generation AI.

[0069] [Server device] The server device 10A will be described. The server device 10A has a control unit 13A instead of the control unit 13 of the server device 10 shown in Fig. 2. The control unit 13A has a selection unit 138A. Fig. 9 is a diagram illustrating the flow of processing in the processing system 1A according to the first modification of the first embodiment.

[0070] The selection unit 138A selects one of a plurality of generation AIs, which are natural language processing models, based on information about the meeting, the content of the meeting, or predetermined rules. The information about the meeting and the content of the meeting are included, for example, in the meeting schedule or the meeting information input from the user terminal 20a (step S41). For example, the selection unit 138A selects Tsuzumi41 or ChatGPT51 as the generation AI suitable for the meeting based on the schedule and the meeting information (step S42 in FIG. 9). Furthermore, after the meeting starts, the selection unit 138A may determine the content of the meeting from the minutes created by the generation AI and switch the generation AI to create the minutes.

[0071] The selection unit 138A determines, for example, the accuracy of the meeting, the degree of response speed, the status of the meeting, the industry related to the meeting, whether the meeting is in a specific field or an accuracy-oriented meeting, or the level of confidentiality of the meeting based on information about the meeting and the content of the meeting. Furthermore, the selection unit 138A may determine, in accordance with a predetermined rule, the accuracy of the meeting, the degree of response speed, the status of the meeting, the industry related to the meeting, whether the meeting is in a specific field or an accuracy-oriented meeting, or the level of confidentiality of the meeting. The selection unit 138A is not limited to the above, and may change the content of the determination depending on the industry, field, members, and situation.

[0072] The selection unit 138A may also use a generation AI (Tsuzumi41 or ChatGPT51) to determine the accuracy of the meeting, the degree of response speed, the status of the meeting, the industry related to the meeting, whether the meeting is in a specific field or an accuracy-oriented meeting, or the level of confidentiality of the meeting. In this case, the selection unit 138A sets a prompt to the generation AI to instruct it to determine the accuracy of the meeting, the degree of response speed, the status of the meeting, the industry related to the meeting, whether the meeting is in a specific field or an accuracy-oriented meeting, or the level of confidentiality of the meeting, based on the input information about the meeting and / or the content of the meeting.

[0073] Based on the determined content, the selection unit 138A selects Tsuzumi41 or ChatGPT51. For example, when the conference is in a specific field (for example, finance), the response speed is set to a relatively fast level, and speed is emphasized, the selection unit 138A selects Tsuzumi41.

[0074] Furthermore, if the confidentiality level of the meeting is above standard, the selection unit 138A selects Tsuzumi 41 within the in-house network 200. The generation AI server 40 may have multiple Tsuzumi 41s corresponding to the fields of finance, medicine, semiconductors, IT, academia, factories, law, or office services. In this case, the selection unit 138A selects Tsuzumi 41 corresponding to the determined field of the meeting and creates minutes of the meeting.

[0075] Furthermore, for example, the selection unit 138A selects ChatGPT51 when accuracy is important. Furthermore, the selection unit 138A selects ChatGPT51 when the confidentiality level of the conference is below standard. The selection unit 138A may determine the content of the conference after the conference starts and switch the generation AI that determines whether or not annotations need to be displayed. Furthermore, the selection unit 138A may input, for example, information about the conference and the content of the conference to the generation AI and cause the generation AI to determine the generation AI that will create the minutes. The generation AI that performs this determination may be generation AI31, Tsuzumi41, ChatGPT51, or another generation AI.

[0076] The prompt creation unit 133 inputs the information about the conference and / or the content determined based on the content of the conference into the prompt together with the conference information (step S42). The prompt creation unit 133 sets the created prompt to the generation AI (e.g., Tsuzumi41) selected by the selection unit 138A (step S43).

[0077] In this way, the server device 10A selects a generation AI suitable for the conference based on the determination content, adjusts the prompt when instructing the generation AI, and then instructs the selected generation AI to create minutes.

[0078] Then, the selected generation AI (for example, Tsuzumi41) creates the minutes of the meeting based on the input text data of the meeting in accordance with the set prompts (step S44).

[0079] [Processing method] Next, a description will be given of the processing procedure of the processing method according to Modification 1 of Embodiment 1. Fig. 10 is a sequence diagram showing the processing procedure of the processing method according to Modification 1 of Embodiment 1. Steps S51-1 to S55 in Fig. 10 are the same processes as steps S21-1 to S25 in Fig. 7.

[0080] The server device 10A selects Tsuzumi41 or ChatGPT51 as the generating AI for creating the minutes of the meeting based on the information about the meeting, the content of the meeting, or a predetermined rule (step S56). Fig. 10 shows the case where Tsuzumi41 is selected.

[0081] Server device 10A creates a prompt instructing the creation of minutes of the meeting based on the input text data (step S57). Server device 10A sets the created prompt to Tsuzumi 41 selected in step S56 (steps S58 and S59).

[0082] Then, when the server device 10A receives voice data of users Ua to Ud from the user terminals 20a to 20d (steps S60-1 to S60-4, S61-1 to S61-4), it converts the voice data into text data (step S62). The server device 10A associates the time information and user information indicating the speaker with the converted text data and inputs the text data to Tsuzumi 41 selected in step S56 (step S63). In the generation AI server 40, Tsuzumi 41 creates minutes of the meeting based on the input text data (step S64) and transmits the created minutes to the server device 10A (step S65).

[0083] Steps S66 to S70 in FIG. 10 are the same processes as steps S35 to S39 shown in FIG.

[0084] [Effects of Modification 1 of Embodiment 1] In variant example 1 of embodiment 1, minutes that are more suitable for the meeting can be created by selecting one of multiple generation AIs as the generation AI to create the minutes of the meeting based on information about the meeting, the content of the meeting, or predetermined rules.

[0085] Note that the generation AI is just an example, and a server having a plurality of other generation AIs installed therein may be further provided. Fig. 11 is a diagram showing another example of the configuration of the processing system according to the first modification of the first embodiment.

[0086] As shown in Figure 11, in the processing system 1B, multiple generation AI servers 40A and 40B are installed within the in-house network 100. The generation AI server 40A is equipped with Tsuzumi 41A, which is fine-tuned for the financial field. The generation AI server 40B is equipped with Tsuzumi 41B, which is fine-tuned for the medical field.

[0087] The control unit 13B of the server device 10B has a selection unit 138B that selects Tsuzumi 41A, 41B, or ChatGPT 51 based on conference information about the conference in which the user is participating, the content of the conference, or predetermined rules. The selection unit 138B selects Tsuzumi 41A if the conference field is the financial field, and selects Tsuzumi 41B if the conference field is the medical field.

[0088] In this way, the server device 10B can create minutes that are more suitable for the conference by selecting a generation AI that matches the field of the conference.

[0089] As described above, in the first embodiment, an example has been described in which minutes of a meeting are automatically created and stored to construct the minutes DB 125. Here, it is also possible to support the progress of a meeting by using the minutes stored in the minutes DB 125 as source data to output annotation information on a Web conference screen or output search results for terms, etc. on the Web conference screen. In the following second embodiment, processing for automatic output of annotations will be described. In the third embodiment, processing for automatic search will be described.

[0090] [Embodiment 2] In the second embodiment, a case where annotation information is automatically output on a Web conference screen will be described. The annotation information is information for explaining words spoken during a conference. In the second embodiment, the minutes DB 125 constructed in the first embodiment is used as the source data to be searched.

[0091] In the second embodiment, when the server device detects a predetermined word (a time-related word such as "last week," a meeting topic, a person's name, etc.), it searches the minutes DB 125 based on the predetermined word and places the searched minutes or a link to the minutes in the annotations section of the web conference screen.

[0092] Fig. 12 is a diagram showing an example of the configuration of a processing system according to embodiment 2. As shown in Fig. 12, the processing system 201 according to embodiment 2 has a server device 210 having a control unit 213 instead of the server device 10. Fig. 13 is a diagram explaining the flow of processing of the processing system 201 according to embodiment 2. Fig. 14 is a diagram showing an example of a screen of the user terminal 20a.

[0093] The control unit 213 has an annotation unit 2131. The annotation unit 2131 searches the minutes DB 125 for annotation information for explaining the words uttered by users Ua to Ud during the conference, and automatically outputs the annotation information to an annotation column on the Web conference screen. The annotation unit 2131 has a determination unit 2132, an annotation information search unit 2133, an annotation display control unit 2134, and a correction receiving unit 2135.

[0094] In the server device 210, when the conference starts, the recognition unit 134 acquires, for example, voice data of the user Ua during the conference through communication with the user terminals 20a to 20d (step S81 in FIG. 13). The recognition unit 134 performs voice recognition on the acquired voice data of the user Ua and converts the voice data into text data (step S82 in FIG. 13). The recognition unit 134 outputs the text data obtained by converting the voice data of the user Ua to the annotation unit 2131.

[0095] The determination unit 2132 determines whether or not annotation information for explaining the words spoken during the conference needs to be displayed based on the text data converted from the voice data of each user Ua to Ud who participated in the conference (step S82 in FIG. 13).

[0096] The determination unit 2132 determines whether or not annotation information needs to be displayed and the words to be used when searching for annotation information based on the appearance of words related to the speaker or time, words related to people, and words characteristic of the meeting.

[0097] For example, the speakers are users Ua to Ud who are participants in the conference. Words related to time include "yesterday," "last week," "morning / afternoon," "next month," "last time," and "first." Words related to people include people's names, job titles, and names of affiliations. If the conference is about "progress in the development of a new service," words containing phrases characteristic of the conference include the name of the service, the concept of the service, topics, the technology that characterizes the service, the customer base of the service, and words related to services related to this service.

[0098] An example will be described in which user Ua makes a statement such as "Last week, the president of Company N mentioned Service Q." In this case, the determination unit 2132 determines that the minutes of a meeting held last week, which include a statement by the "president" of Company N about Service Q, should be displayed as annotation information. The determination unit 2132 then determines that the search target is the minutes of the meeting that user Ua attended last week, and that "Company N," "President," and "Service Q" are the words to be used when searching for annotation information.

[0099] Next, an example will be described in which the user Ua makes a statement that includes the phrase "The president said that last week." In this case, the determination unit 2132 determines that the minutes of the meeting held "last week" containing the statement of a person with the title "president" should be displayed as annotation information. The determination unit 2132 then determines that the search target is the minutes of the meeting that the user Ua attended "last week," and that the phrase is used by "president" when searching for annotation information. In this way, the determination unit 2132 determines that annotation information should be displayed even when a demonstrative pronoun appears.

[0100] The annotation information search unit 2133 searches for annotation information from the minutes DB 125, which stores the minutes of meetings that have been held so far (step S83 in FIG. 13). The annotation information search unit 2133 uses the wording to be used when searching for annotation information, determined by the determination unit 2132, to search the minutes DB 125 for minutes related to words spoken during the meeting as annotation information.

[0101] For example, a case will be described in which the determination unit 2132 determines that the minutes of the conference that the user Ua attended "last week" are the search target and that "President" is the phrase to be used when searching for annotation information. In this case, the annotation information search unit 2133 searches the minutes DB 125 for minutes of the conference that the user Ua attended "last week" and finds minutes D1 that include "President".

[0102] The annotation display control unit 2134 displays, as an annotation, the relevant portion of the minutes (step S84 in FIG. 13) found by the annotation information search unit 2133 on each of the user terminals 20a to 20d used by each of the users Ua to Ud (steps S85-1 to S85-4 in FIG. 13). The annotation display control unit 2134 may display, as an annotation, the minutes D1 found by the annotation information search unit 2133 or a link to the minutes D1 in the annotation column of the Web conference screen.

[0103] For example, the annotation display control unit 2134 causes the user terminal 20a to display annotation information C1 in the annotation column of the Web conference screen, stating, "Last week's meeting minutes (the part where President M spoke) are here," as shown on screen M1 in Fig. 14. When the user Ua selects the "here" part B1 in the annotation information C1, the user can transition to a viewing screen for the part where President M spoke in the corresponding minutes D1.

[0104] The correction receiving unit 2135 receives an instruction to correct the annotation from the user terminal operated by the user. The determining unit 2132 corrects the search conditions for the minutes DB 125 and causes the annotation information searching unit 2133 to search for the minutes again.

[0105] 14, when the user Ua selects the correction button C2 of the annotation information C1, the user Ua can input a correction instruction in the annotation field. The user Ua selects the correction button C2 and inputs a comment C3 in the annotation field, pointing out the mistake in the president's name, saying, "It's a mistake, it's President L."

[0106] In this case, the correction receiving unit 2135 receives an instruction to correct "President M" to "President L." Then, the determination unit 2132 determines that the minutes of the conference that user Ua attended "last week" are the search target, and that "President L" is the phrase to use when searching for annotation information. Then, the annotation information search unit 2133 searches for minutes D2 containing "President L" from the minutes of the conference that user Ua attended "last week" in the minutes DB 125. It is also possible to feed back the corrections made by user Ua to the determination algorithm of the determination unit 2132.

[0107] As shown in screen M1 of Fig. 14, the annotation display control unit 2134 causes the user terminal 20a to display annotation information C4 stating "Last week's meeting minutes (the part where President L spoke) are here" in the annotation column of the Web conference screen. When the user Ua selects the "here" part B2 in the annotation information C4, the user can transition to a viewing screen for the part where President L spoke in the corresponding minutes D2. When the user Ua selects the edit button C5 in the annotation information C4, the user can input edit instructions in the annotation column.

[0108] [Processing method] Next, a description will be given of a processing method executed by the processing system 201. Fig. 15 is a sequence diagram showing the processing steps of the processing method according to the second embodiment.

[0109] Steps S91-1 to S103 shown in Fig. 15 are the same processes as steps S21-1 to S33 shown in Fig. 7. After step S103 is completed, server device 210 receives the minutes created by generation AI 31, as in the first embodiment, and executes the processes of steps S35 to S39 in Fig. 7 to store the minutes in minutes DB 125.

[0110] In step S101, the server device 210 determines whether or not to display annotation information based on the text data converted from the voice data of the users Ua to Ud (step S104). If annotation information is not to be displayed (step S104: No), the server device 210 receives the voices of the user terminals 20a to 20d (steps S100-1 to S100-4), performs voice recognition (step S101), and performs the determination in step S104 again.

[0111] When the annotation information is to be displayed (step S104: Yes), the server device 210 determines the wording to be used when searching for the annotation information, and uses this wording to search for the corresponding minutes from the minutes DB 125 (step S105).

[0112] The server device 210 displays the minutes retrieved by the annotation information retrieval unit 2133 as annotations on the user terminals 20a to 20d used by the users Ua to Ud (steps S106 to S108-4).

[0113] For example, when the server device 210 receives an instruction to correct the annotation from the user terminal 20a (step S109), it corrects the search conditions for the annotation information (step S110) and searches for the minutes and displays the annotation information again.

[0114] [Effects of the second embodiment] The server device 210 detects predetermined words in the participants' comments, such as time-related words like "last week," words containing phrases characteristic of the meeting like meeting topics, and people's names.The server device 210 then searches the minutes DB 125 based on the detected predetermined words, and pastes the minutes or a link to the minutes in the annotations field on the Web conference screen.

[0115] For example, when the server device 210 recognizes the phrase "The president said that last week" during a meeting, it searches for minutes of meetings that the speaker attended "last week" that correspond to "the president," and displays the relevant part of the minutes in the comment section, or adds a link to the relevant minutes.

[0116] This allows all participants in the meeting to clearly recognize what the "president" said "last week" without having to go through the tedious process of checking what the "president" said "last week." Therefore, the second embodiment can help meetings proceed smoothly.

[0117] In the second embodiment, a link to news related to the meeting, which is stored in the minutes DB 125 in association with the relevant minutes, may be pasted in the comment field as annotation information. Therefore, according to the second embodiment, the user does not need to perform cumbersome processes such as searching for related news during the meeting.

[0118] [Modification 1 of Embodiment 2] In the second embodiment, the server device 210 determines whether or not an annotation needs to be displayed, but the generation AI 31 may determine whether or not an annotation needs to be displayed. The determination unit 2132 uses the generation AI 31 to determine whether or not an annotation needs to be displayed.

[0119] At this time, the server device 210 sets a prompt to the generation AI 31 to instruct it to determine whether or not annotation information needs to be displayed and the wording to be used when searching for annotation information, based on the context of the input text data.

[0120] Fig. 16 is a diagram illustrating a processing flow in Modification 1 of Embodiment 2. Fig. 17 is a sequence diagram illustrating a processing procedure of a processing method according to Modification 1 of Embodiment 2.

[0121] Steps S131-1 to S142 in FIG. 17 are the same processes as steps S91-1 to S103 in FIG.

[0122] Specifically, in the server device 210, when a conference starts, the recognition unit 134 acquires, for example, voice data of the user Ua during the conference through communication with the user terminals 20a to 20d (step S121 in FIG. 16). The recognition unit 134 performs voice recognition on the acquired voice data of the user Ua and converts the voice data into text data (step S122 in FIG. 16). The recognition unit 134 outputs the text data obtained by converting the voice data of the user Ua to the generation AI 31 (step S123 in FIG. 16). The generation AI 31 creates minutes of the conference (step S143 in FIG. 17).

[0123] Here, before the start of the conference, a prompt is set in the generation AI 31 by the server device 210 (steps S136 to S138 in FIG. 17). The prompt is a prompt that instructs the generation model to generate minutes of the conference based on the input text data, and information about the conference is added to it. The prompt also instructs the generation model to determine, based on the context of the input text data, whether or not annotation information needs to be displayed and the wording to be used when searching for annotation information.

[0124] Therefore, the generation AI31 determines whether or not annotation information needs to be displayed, and also determines the wording to be used when searching for annotation information, based on the context of the input text data (step S124 in FIG. 16, step S144 in FIG. 17). In this way, the determination unit 2132 uses the generation AI31 to determine whether or not annotation information needs to be displayed, and the wording to be used when searching for annotation information.

[0125] If annotation information is not to be displayed (step S144: No in Figure 17), the generation AI31 creates minutes based on the text data received from the server device 210 (step S143 in Figure 17) and again performs the judgment in step S144 in Figure 17.

[0126] When the annotation information is to be displayed (step S144 in FIG. 17: Yes), the generation AI 31 instructs the server device 210 to search for the annotation information (step S125 in FIG. 16, step S145 in FIG. 17). The instruction to search for the annotation information includes the wording to be used when searching for the annotation information.

[0127] The server device 210 searches for the relevant minutes from the minutes DB 125 in accordance with the annotation information search instruction (steps S126 and S127 in FIG. 16, and step S146 in FIG. 17). The server device 210 displays the minutes searched for by the annotation information search unit 2133 as annotations on the user terminals 20a to 20d used by the users Ua to Ud (steps S128-1 to S128-4 in FIG. 16, and steps S147 to S149-4 in FIG. 17).

[0128] For example, when the server device 210 receives an instruction to correct an annotation from the user terminal 20a (step S150 in FIG. 17), it corrects the search conditions for annotation information and searches for the minutes and displays the annotation information again. Furthermore, the server device 210 may have the generation AI 31 learn the content of the annotation correction, thereby feeding back the content of the correction to the generation AI 31 and improving the judgment accuracy regarding annotation search.

[0129] As in the first modification of the second embodiment, by having the generation AI 31 determine whether or not an annotation needs to be displayed, it is possible to omit setting a determination algorithm for the server device 210.

[0130] [Modification 2 of Embodiment 2] Next, a second modification of the second embodiment will be described. In the second modification of the second embodiment, one of a plurality of generation AIs is selected as the generation AI that determines whether or not an annotation needs to be displayed, thereby enabling a more appropriate determination to be made.

[0131] Fig. 18 is a diagram showing an example of the configuration of a processing system according to Modification 2 of Embodiment 2. As shown in Fig. 18, in a processing system 201A according to Modification 2 of Embodiment 2, a generation AI server 40 is installed within an in-house network 100 in which user terminals 20a to 20d and a server device 210A are installed. In addition, the server device 210A is capable of communicating with a generation AI server 50, which is an external server.

[0132] Tsuzumi41 of generation AI server 40 and ChatGPT51 of generation AI server 50 perform natural language processing on the input text data according to the set prompts, and based on the input text data, determine whether or not annotation information needs to be displayed and the wording to use when searching for annotation information.

[0133] The server device 210A selects Tsuzumi41 or ChatGPT51 based on the conference information about the conference the user is participating in, the content of the conference, or predetermined rules. Then, the server device 10A uses the selected generation AI to determine whether or not annotation information needs to be displayed and the wording to use when searching for annotation information.

[0134] This allows the server device 210A to determine whether or not it is necessary to display annotation information according to the situation, based on the selected generation AI.

[0135] [Server device] Server device 210A will be described. As shown in Fig. 18, server device 210A has a control unit 213A instead of control unit 213 of server device 210 shown in Fig. 12. Control unit 213A has a selection unit 2138A. Fig. 19 is a diagram illustrating the flow of processing in Variation 1 of Embodiment 2.

[0136] The selection unit 2138A selects one of a plurality of generation AIs, which are natural language processing models, as a generation AI to determine whether or not an annotation needs to be displayed, based on information about the meeting, the content of the meeting, or a predetermined rule. The information about the meeting and the content of the meeting are included, for example, in the meeting schedule or the meeting information input from the user terminal 20a.

[0137] For example, the selection unit 2138A selects Tsuzumi41 or ChatGPT51 as the generation AI that determines whether or not an annotation needs to be displayed before the meeting begins, based on the schedule and meeting information (steps S161 and S162 in FIG. 19). For example, in order to improve processing speed, the selection unit 2138A selects the same generation AI (for example, Tsuzumi41) as the generation AI that creates the minutes and the generation AI that determines whether or not an annotation needs to be displayed.

[0138] Alternatively, after the conference starts, the selection unit 2138A determines the content of the conference from text data converted from conference voice data by voice recognition and the minutes of the conference created by the generation AI, and selects Tsuzumi41 or ChatGPT51 as the generation AI that determines whether or not annotations need to be displayed. After the conference starts, the selection unit 2138A may switch the generation AI that determines the content of the conference and determines whether or not annotations need to be displayed.

[0139] The selection unit 2138A may also select different generation AIs as the generation AI that creates the minutes and the generation AI that determines whether or not annotations need to be displayed. For example, the selection unit 2138A may input information about the meeting and the contents of the meeting into the generation AI and cause the generation AI to determine the generation AI that creates the minutes and the generation AI that determines whether or not annotations need to be displayed. The generation AI that performs this determination may be generation AI31, Tsuzumi41, ChatGPT51, or another generation AI.

[0140] 8, the selection unit 2138A determines, based on information about the conference and the content of the conference, for example, the accuracy of the conference, the degree of response speed, the status of the conference, the industry related to the conference, whether the conference is in a specific field or emphasizes accuracy, or the level of confidentiality of the conference.The selection unit 2138A may then select Tsuzumi41 or ChatGPT51 based on the determined content.

[0141] Then, the prompt creation unit 133 sets a prompt that instructs the generation AI (e.g., Tsuzumi41) selected by the selection unit 2138A to determine whether or not annotation information needs to be displayed and the wording to use when searching for annotation information based on the context of the input text data.

[0142] Server device 210A converts the conference voice data into text data using voice recognition, and inputs the text data to Tsuzumi 41 selected by selection unit 2138A (steps S161 to S163 in FIG. 19).

[0143] Based on the text data converted from the conference voice data by speech recognition, Tsuzumi 41 determines whether or not annotation information needs to be displayed, as well as the wording to be used when searching for annotation information (step S164 in FIG. 19). If annotation information is to be displayed, Tsuzumi 41 instructs the server device 210 to search for annotation information (step S165 in FIG. 19). The instruction to search for annotation information includes the wording to be used when searching for annotation information.

[0144] In the server device 210, the annotation information search unit 2133 searches for the relevant minutes from the minutes DB 125 in accordance with the annotation information search instruction from Tsuzumi 41 (steps S166 and S167 in FIG. 19). In the server device 210, the annotation display control unit 2134 displays the minutes searched for by the annotation information search unit 2133 as annotations on the user terminals 20a to 20d used by the users Ua to Ud (steps S168-1 to S168-4 in FIG. 19).

[0145] Fig. 20 is a sequence diagram showing the processing procedure of the processing method according to Modification 2 of Embodiment 2. Fig. 20 illustrates an example in which the generation AI that creates the minutes and the generation AI that determines whether or not an annotation needs to be displayed are the same generation AI (Tsuzumi41).

[0146] Steps S171-1 to S175 in Fig. 20 are the same processes as steps S131-1 to S135 in Fig. 17. The server device 210A selects Tsuzumi41 or ChatGPT51 as the generating AI that determines whether or not an annotation needs to be displayed based on the schedule and meeting information (step S176).

[0147] A prompt is set by the server device 210 for the generation AI (e.g., Tsuzumi41) selected in step S178 (steps S177 to S179). The prompt is a prompt that instructs the generation model to generate minutes of the meeting based on the input text data, and information about the meeting is added. The prompt also instructs the generation model to determine, based on the context of the input text data, whether or not annotation information needs to be displayed and the wording to use when searching for annotation information.

[0148] Steps S180-1 to S180-4 and S181-1 to S182 in Fig. 20 are the same processes as steps S139-1 to S141 in Fig. 17. Server device 210A associates the time information and user information indicating the speaker with the converted text data, and inputs them to Tsuzumi41 selected in step S176 (step S183).

[0149] Tsuzumi41 performs steps S184 to S186 in Fig. 20, which are the same as steps S143 to S145 in Fig. 17. Steps S187 to S192 in Fig. 20 are the same as steps S146 to S151 in Fig. 17.

[0150] In this way, by selecting one of the multiple generation AIs as the generation AI that determines whether or not to display an annotation based on information about the meeting, the content of the meeting, or a predetermined rule, a more appropriate determination can be made, thereby further optimizing the annotation display on the user terminals 20a to 20d.

[0151] In this embodiment, Tsuzumi41 and ChatGPT51 have been used as examples of the generation AIs to be used, but other generation AIs may also be used, and the number of generation AIs is not limited to two, but may be any one of three or more generation AIs.

[0152] Figure 21 is a diagram showing another example configuration of a processing system according to Variation 2 of Embodiment 2. As shown in Figure 21, in processing system 201B, multiple generation AI servers 40A, 40B are installed within an in-house network 100. Generation AI server 40A is equipped with Tsuzumi 41A, which is fine-tuned for the financial field. Generation AI server 40B is equipped with Tsuzumi 41B, which is fine-tuned for the medical field.

[0153] Control unit 213B of server device 210B has selection unit 2138B that selects Tsuzumi 41A, 41B, or ChatGPT 51 based on conference information about the conference in which the user is participating, the content of the conference, or predetermined rules. Selection unit 2138B selects Tsuzumi 41A if the conference field is the financial field, and selects Tsuzumi 41B if the conference field is the medical field.

[0154] In this way, server device 210B can make a more appropriate determination by selecting a generation AI that matches the field of the conference.

[0155] [Embodiment 3] In the third embodiment, a case where search results are automatically output on a web conference screen will be described. The search results are search results for words spoken during a conference. In the third embodiment, the minutes DB 125 constructed in the first embodiment is used as one of the source data to be searched.

[0156] In the third embodiment, when the server device recognizes a question, a predetermined period of silence, or the repetition of the same words, it searches for the words asked in the question, the words immediately before the predetermined period of silence, or the repeated words, and displays the search results in the chat box on the web conference screen.

[0157] Fig. 22 is a diagram showing an example of the configuration of a processing system according to the third embodiment. As shown in Fig. 22, the processing system 301 according to the third embodiment has a server device 310 having a control unit 313 instead of the server device 10. The server device 310 is capable of communicating with an external server 60 that manages a search engine. Figs. 23 and 24 are diagrams explaining the flow of processing of the processing system 201 according to the third embodiment. Fig. 25 is a diagram showing an example of a screen of the user terminal 20a.

[0158] The control unit 313 has a search unit 3131. The search unit 3131 searches the minutes DB 125 or a search engine for words spoken during a conference, and automatically outputs the search results to a chat box on the Web conference screen. The search unit 3131 has a determination unit 3132, a search execution unit 3133, and a search result display control unit 3134 (display control unit).

[0159] In the server device 310, when a conference starts, the recognition unit 134 acquires, for example, voice data of the user Ua during the conference through communication with the user terminals 20a to 20d (step S201 in FIGS. 23 and 24). The recognition unit 134 performs voice recognition on the acquired voice data of the user Ua and converts the voice data into text data (step S202 in FIGS. 23 and 24). The recognition unit 134 outputs the text data converted from the voice data of the user Ua to the generation AI 31 (step S203 in FIGS. 23 and 24).

[0160] The determination unit 3132 determines whether or not it is necessary to search for words spoken during a conference, based on the text data converted from the conference voice data. If a search is necessary, the search execution unit 3133 searches for the words to be searched for.

[0161] Here, a prompt is set in the generation AI 31 to instruct it to determine whether a search is necessary and the search target words based on the context of the text data converted from the conference audio data. The determination unit 3132 determines whether a search is necessary and the search target words using the generation AI 31. The determination unit 3132 instructs the search execution unit 3133 to search for the search target words determined by the generation AI 31.

[0162] For example, the prompt is a prompt that instructs the generative model to generate minutes of a meeting based on input text data, with information about the meeting added.

[0163] The prompt then instructs the system to determine that a search is necessary if it recognizes a question, a predetermined period of silence, or repeated words based on the context of the text data converted from the conference audio data.The prompt then instructs the system to determine that the words being asked in the question, the words immediately before the predetermined period of silence, or the repeated words are the words to be searched.Furthermore, if a search is necessary, the prompt instructs the system to determine, based on the context of the text data, that a search should be performed using a search engine converted from the conference audio data or a search of the minutes DB 125.

[0164] As a result, when the generation AI 31 recognizes a question, a predetermined period of silence, or repetition of the same word based on the context of the text data converted from the conference audio data, it determines that a search is necessary (step S204 in Figures 23 and 24).

[0165] The generation AI 31 then determines that the interrogative phrase, the phrase immediately preceding the predetermined period of silence, or the repeated phrase is the phrase to be searched for. Furthermore, the generation AI 31 determines whether to perform a search using a search engine or a search of the minutes DB 125.

[0166] If the generation AI 31 determines that a search is necessary, it sends a search instruction to the determination unit 3123 to search the target phrase using a search engine or to search the minutes DB 125 (step S205-1 in FIG. 23, S205-2 in FIG. 24). If the generation AI 31 can determine the general meaning of the target phrase, it may return the semantic content to the determination unit 3132 as the search result.

[0167] The determination unit 3132 uses the generation AI 31 to instruct the search execution unit 3133 to search for the search target phrase, as well as to perform a search using a search engine or a search of the minutes DB 125 .

[0168] For example, referring to Fig. 23, a case will be described where an instruction to search the minutes DB 125 is given from the generation AI 31. In this case, the search execution unit 3133 searches the minutes DB 125 based on the search target wording (step S206-1 in Fig. 23), and obtains minutes that include the content of the search target wording as the search result (step S207-1 in Fig. 23). At this time, the search execution unit 3133 performs the search using a search formula created in response to the search of the minutes DB 125, thereby shortening the search time.

[0169] 24, an example will be described in which a search using a search engine is instructed by the generation AI 31. In this case, the search execution unit 3133 searches the search engine for the search target phrase via the external server 60 (step S206-2 in FIG. 24), and obtains the content of the search target phrase as the search result (step S207-2 in FIG. 24). At this time, the search execution unit 3133 performs the search using a search formula created in response to the search using the search engine, thereby shortening the search time.

[0170] The search result display control unit 3134 displays the search results obtained by the search execution unit 3133 in the chat boxes of the user terminals 20a to 20d used by the users Ua to Ud (steps S208-1 to S208-4 in FIGS. 23 and 24).

[0171] For example, the search result display control unit 3134 displays content C21 corresponding to the phrase "Company D's Code of Conduct," which is the phrase that occurred immediately before the predetermined period of silence, in the chat section of the Web conference screen, as shown on screen M2 in Fig. 25. "Company D's Code of Conduct" is information obtained by searching the minutes DB 125, for example.

[0172] Furthermore, the search result display control unit 3134 displays the definition C22 of "portfolio" repeated multiple times in the chat field of the web conference screen, as shown on screen M2 in Fig. 25. The definition of "portfolio" is, for example, information obtained by the search results of a search engine.

[0173] [Processing method] Next, a description will be given of a processing method executed by the processing system 301. Fig. 26 is a sequence diagram showing the processing steps of the processing method according to the third embodiment.

[0174] Steps S211-1 to S223 shown in Fig. 26 are the same processes as steps S21-1 to S33 shown in Fig. 7. After step S223 is completed, server device 310 receives the minutes created by generation AI 31, as in the first embodiment, and executes the processes of steps S35 to S39 in Fig. 7 to store the minutes in minutes DB 125.

[0175] In step S221, the generation AI 31 determines whether or not it is necessary to search for words spoken during the conference based on the text data converted from the voice data of the users Ua to Ud (step S224). If no search is necessary (step S224: No), the generation AI 31 receives the text data converted from the voice data of the users Ua to Ud from the server device 310, and performs the process of step S223 and the determination of step S224.

[0176] If a search is necessary (step S224: Yes), the generation AI 31 determines the word to be searched and instructs the server device 310 to search for this word (step S225). The generation AI 31 also instructs the server device 310 to search using a search engine or the minutes DB 125.

[0177] The server device 310 determines whether the instruction of the generated AI 31 is to search using a search engine or to search the minutes DB 125 (step S226).

[0178] If the search is in the minutes DB 125 (step S226: minutes DB), the server device 310 searches the minutes DB 125 based on the search target wording (step S227), and obtains minutes containing the content of the search target wording as the search result.

[0179] If the search is performed using a search engine (step S226: Search Engine), the server device 310 communicates with the external server 60 to perform a search using the search engine (step S228), and obtains the content of the search target phrase as the search result (step S229).

[0180] Then, the server device 310 displays the acquired search results on the user terminals 20a to 20d used by the users Ua to Ud (steps S230 to S232-4).

[0181] [Effects of the Third Embodiment] The server device 310 determines whether a search is necessary based on the text data converted from the conference voice data, and if a search is necessary, automatically executes the search and displays the search results on the screen of the user terminal 20.

[0182] Specifically, when the server device 310 recognizes a question, a predetermined period of silence, or repetition of the same word, it automatically detects the word being asked in the question, the word immediately before the predetermined period of silence, or the repeated word as a search target word.The server device 310 then automatically searches for the detected search target word and displays the search results in the chat box on the web conference screen.

[0183] Therefore, according to the third embodiment, the user himself / herself does not have to perform cumbersome processing such as searching for unknown words during the conference, and the conference can proceed smoothly.

[0184] [Modification 1 of the Third Embodiment] Next, a first modification of the third embodiment will be described. In the first modification of the third embodiment, one of a plurality of generation AIs is selected as the generation AI that determines whether or not a search is necessary, thereby enabling a more appropriate determination to be made.

[0185] Figure 27 is a diagram showing an example of the configuration of a processing system according to Modification 1 of Embodiment 3. As shown in Figure 27, in a processing system 301A according to Modification 1 of Embodiment 3, a generation AI server 40 is installed within an in-house network 100 in which user terminals 20a to 20d and a server device 310A are installed. In addition, the server device 310A is capable of communicating with a generation AI server 50, which is an external server.

[0186] Tsuzumi41 of generation AI server 40 and ChatGPT51 of generation AI server 50 perform natural language processing on the input text data according to the set prompts, and based on the input text data, determine whether a search is necessary and the words to be searched.

[0187] The server device 310A selects Tsuzumi41 or ChatGPT51 based on the conference information about the conference the user is participating in, the content of the conference, or predetermined rules. Then, the server device 10A uses the selected generation AI to determine whether a search is necessary and the words to be searched.

[0188] This allows server device 310A to determine whether search display is necessary depending on the situation, based on the selected generation AI.

[0189] [Server device] Server device 310A will be described. As shown in Fig. 27, server device 310A has a control unit 313A instead of control unit 313 of server device 310 shown in Fig. 22. Control unit 313A has a selection unit 3138A. Fig. 28 is a diagram illustrating the flow of processing in Variation 1 of Embodiment 3.

[0190] The selection unit 3138A selects one of a plurality of generation AIs, which are natural language processing models, as a generation AI to determine whether a search is necessary, based on information about the meeting, the content of the meeting, or a predetermined rule. The information about the meeting and the content of the meeting are included, for example, in the meeting schedule or the meeting information input from the user terminal 20a.

[0191] For example, the selection unit 3138A selects Tsuzumi41 or ChatGPT51 as the generation AI to determine whether a search is necessary before the meeting begins based on the schedule and meeting information (steps S241 and S242 in FIG. 28). For example, in order to improve processing speed, the selection unit 3138A selects the same generation AI (for example, Tsuzumi41) as the generation AI to create the minutes and the generation AI to determine whether a search is necessary.

[0192] Alternatively, after the conference starts, the selection unit 3138A determines the content of the conference from text data converted from the conference voice data by voice recognition and the minutes of the conference created by the generation AI, and selects Tsuzumi41 or ChatGPT51 as the generation AI that determines whether a search is necessary. After the conference starts, the selection unit 3138A may switch the generation AI that determines the content of the conference and determines whether a search is necessary.

[0193] In addition, the selection unit 3138A may select different generation AIs as the generation AI that creates the minutes and the generation AI that determines whether a search is necessary. For example, the selection unit 3138A may input information about the meeting and the content of the meeting into the generation AI and have the generation AI determine the generation AI that creates the minutes and the generation AI that determines whether a search is necessary. The generation AI that performs this determination may be generation AI31, Tsuzumi41, ChatGPT51, or another generation AI.

[0194] 8, the selection unit 3138A determines, based on information about the conference and the content of the conference, for example, the accuracy of the conference, the degree of response speed, the status of the conference, the industry related to the conference, whether the conference is in a specific field or emphasizes accuracy, or the level of confidentiality of the conference.The selection unit 3138A may then select Tsuzumi41 and ChatGPT51 based on the determined content.

[0195] Then, the prompt creation unit 133 sets a prompt to the generation AI (for example, Tsuzumi41) selected by the selection unit 3138A to determine whether a search is necessary and the words to be searched, based on the context of the input text data. The prompt is the same as the prompt set in the generation AI 31 in the third embodiment.

[0196] Server device 310A converts the conference voice data into text data using voice recognition, and inputs the text data to Tsuzumi 41 selected by selection unit 3138A (steps S241 to S243 in FIG. 28).

[0197] Tsuzumi 41 determines whether a search is necessary and the words to be searched based on the text data converted from the conference voice data by voice recognition (step S244 in FIG. 28). If Tsuzumi 41 determines that a search is necessary, the generation AI 31 instructs the server device 310 to perform a search (step S245 in FIG. 28). The search instruction includes the words to be searched. The search instruction may include content instructing a search using a search engine or a search of the minutes DB 125.

[0198] In the server device 310, the search execution unit 3133 executes a search in accordance with the search instruction from Tsuzumi41.

[0199] For example, referring to Fig. 28, a case will be described where Tsuzumi 41 instructs to search the minutes DB 125. In this case, the search execution unit 3133 searches the minutes DB 125 based on the search target wording (step S246 in Fig. 28), and obtains minutes that include the content of the search target wording as the search result (step S247 in Fig. 28). Also, when Tsuzumi 41 instructs to search using a search engine, the search execution unit 3133 searches the search engine for the search target wording, and obtains the content of the search target as the search result.

[0200] The search result display control unit 3134 displays the search results obtained by the search execution unit 3133 in the chat boxes of the user terminals 20a to 20d used by the users Ua to Ud (steps S248-1 to S248-4 in FIG. 28).

[0201] Fig. 29 is a sequence diagram showing the processing procedure of the processing method according to Modification 1 of Embodiment 3. Fig. 29 illustrates an example in which the generation AI that creates the minutes and the generation AI that determines whether a search is necessary are the same generation AI (Tsuzumi41).

[0202] Steps S251-1 to S255 in Fig. 29 are the same processes as steps S211-1 to S215 in Fig. 26. The server device 310A selects Tsuzumi41 or ChatGPT51 as the generating AI that determines whether or not an annotation needs to be displayed, based on the schedule and meeting information (step S256).

[0203] A prompt is set by the server device 310A for the generated AI (for example, Tsuzumi41) selected in step S256 (steps S257 to S259). The prompt is the same as the prompt that the server device 310 sets for the generated AI 31.

[0204] Steps S260-1 to S262 in Fig. 29 are the same processes as steps S219-1 to S211 in Fig. 26. Server device 310A associates the time information and user information indicating the speaker with the converted text data, and inputs them to Tsuzumi 41 selected in step S256 (step S263).

[0205] Tsuzumi41 performs the same processes as steps S243 to S245 in Fig. 26 as the processes of steps S264 to S266 in Fig. 29. Steps S267 to S273-4 in Fig. 28 are the same processes as steps S226 to S232-4 in Fig. 26.

[0206] In this way, by selecting one of the multiple generation AIs as the generation AI that determines whether a search is necessary based on information about the meeting, the content of the meeting, or a predetermined rule, a more appropriate determination can be made, which further optimizes the search display on the user terminals 20a to 20d.

[0207] In this embodiment, Tsuzumi41 and ChatGPT51 have been used as examples of the generation AIs to be used, but other generation AIs may also be used, and the number of generation AIs is not limited to two, but may be any one of three or more generation AIs.

[0208] Furthermore, the generation AI is just an example, and a server equipped with a plurality of other generation AIs may be further provided.

[0209] Figure 30 is a diagram showing another example configuration of a processing system according to Variation 2 of Embodiment 3. As shown in Figure 30, in processing system 301B, multiple generation AI servers 40A and 40B are installed within an in-house network 100. Generation AI server 40A is equipped with Tsuzumi 41A, which is fine-tuned for the financial field. Generation AI server 40B is equipped with Tsuzumi 41B, which is fine-tuned for the medical field.

[0210] Control unit 313B of server device 310B has selection unit 3138B that selects Tsuzumi 41A, 41B, or ChatGPT 51 based on conference information about the conference the user is participating in, the content of the conference, or predetermined rules. Selection unit 3138B selects Tsuzumi 41A if the conference field is the financial field, and selects Tsuzumi 41B if the conference field is the medical field.

[0211] In this way, server device 310B can make a more appropriate determination by selecting a generation AI that matches the field of the conference.

[0212] [Embodiment 4] Furthermore, the server device may be a combination of the functions of the server devices according to Embodiments 1, 2, and 3. Fig. 31 is a diagram showing an example of the configuration of a processing system according to Embodiment 4.

[0213] 31, in a processing system 401 according to the third embodiment, a control unit 413 of a server device 410 has the functions of, for example, an annotation unit 2131 and a search unit 3131 in addition to the functions of the control unit 113 of the server device 10. This enables the server 401 to automatically output both annotation information and search results to the Web conference screen. Note that in the fourth embodiment, minutes creation, annotation determination, or search determination may be performed using a generation AI or by selecting a generation AI from multiple generation AIs, similar to the server device 10A, server device 10B, server device 210A, server device 20B, server device 310A, and server device 310B in FIG. 8, without being limited to the server device 410.

[0214] [System configuration of the embodiment] The server devices 10, 10A, 10B, 210, 210A, 210B, 310, 310A, 310B, and 410 are conceptual functional representations and are not necessarily physically configured as shown in the drawings. In other words, the specific forms of distribution and integration of the functions of the server devices 10, 10A, 10B, 210, 210A, 210B, 310, 310A, 310B, and 410 are not limited to those shown in the drawings, and all or part of the server devices can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, and the like.

[0215] Furthermore, all or any part of the processes performed by the server devices 10, 10A, 10B, 210, 210A, 210B, 310, 310A, 310B, and 410 may be realized by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and a program analyzed and executed by the CPU and the GPU (Graphics Processing Unit). Furthermore, each process performed by the server devices 10, 10A, 10B, 210, 210A, 210B, 310, 310A, 310B, and 410 may be realized as hardware using wired logic.

[0216] Furthermore, among the processes described in the embodiments, all or part of the processes described as being performed automatically can be performed manually. Alternatively, all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters described above and illustrated can be changed as appropriate unless otherwise specified.

[0217] [program] 32 is a diagram showing an example of a computer in which server devices 10, 10A, 10B, 210, 210A, 210B, 310, 310A, 310B, and 410 are realized by executing a program. Computer 1000 includes, for example, memory 1010 and CPU 1020. Computer 1000 also includes a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0218] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM (Random Access Memory) 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.

[0219] The hard disk drive 1090 stores, for example, an operating system (OS) 1091, an application program 1092, a program module 1093, and program data 1094. That is, a program that defines each process of the server devices 10, 10A, 10B, 210, 210A, 210B, 310, 310A, 310B, and 410 is implemented as a program module 1093 in which code executable by the computer 1000 is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing processes similar to those of the functional configurations of the server devices 10, 10A, 10B, 210, 210A, 210B, 310, 310A, 310B, and 410 is stored in the hard disk drive 1090. The hard disk drive 1090 may be replaced with a solid state drive (SSD).

[0220] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in memory 1010 or hard disk drive 1090. Then, CPU 1020 reads program module 1093 and program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as necessary and executes them.

[0221] The program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090, but may also be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or a wide area network (WAN)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.

[0222] Although the present invention has been described above as an embodiment, the present invention is not limited to the descriptions and drawings that form part of the disclosure of the present invention. In other words, other embodiments, examples, and operational techniques that can be made by those skilled in the art based on the present invention are all included in the scope of the present invention. [Explanation of symbols]

[0223] 1,1A,1B,201,201A,201B,301,301A Processing System 10, 10A, 10B, 210, 210A, 210B, 310, 310A, 310B, 410 Server equipment 11 Communications Department 12 Storage section 13, 13A, 13B, 213, 213A, 213B, 313, 313A Control unit 20, 20a, 20b, 20c, 20d User terminal 26,126 recorded audio database 30, 40, 40A, 40B, 50 Generation AI Server 31 Generation AI 60 External Servers 100,200 Internal network 121 User Information 122 Schedule Information 123 Meeting Information 124 Text Data 125 Minutes DB 131 Schedule Acquisition Department 132 Meeting Information Acquisition Unit 133 Prompt Creation Department 134 Recognition part 135 Pretreatment section 136 Storage area 137 Mapping section 138, 138A, 138B, 2138A, 2138B, 3138A, 3138B selection section 2131 Commentary 2132,3132 Judgment section 2133 Annotation Information Search Unit 2134 Annotation display control section 2135 Correction Reception Department 3131 Search Department 3133 Search execution unit 3134 Search result display control unit

Claims

1. an acquisition unit that acquires information about a conference; a recognition unit that converts voice data of each user who has participated in the conference into text data and associates time information with each text; a determination unit that determines whether or not a search for phrases spoken during the conference is necessary based on the text data; a search execution unit that searches for a search target phrase when the search is necessary; a display control unit that displays the search results obtained by the search execution unit on each user terminal used by each user; A processing device comprising:

2. the recognition unit inputs the converted text data into a generation model that is a natural language processing model; a prompt is set in the generative model to instruct it to determine whether the search is necessary and the search target phrase based on the context of the text data; The processing device according to claim 1 , wherein the determining unit determines whether the search is necessary and the search target words using the generative model.

3. The generative model determines whether to search using a search engine or a minutes database that stores minutes of meetings that have been held, based on the context of the text data; 3. The processing device according to claim 2, wherein the determination unit instructs the search execution unit to execute a search using the search engine or a search of the minutes database using the generative model.

4. 4. The processing device according to claim 3, wherein the search execution unit performs a search using a search expression created for each of a search using the search engine and a search using the minutes database.

5. When the generative model recognizes a question, a predetermined period of silence, or repetition of the same word based on the context of the text data, it determines that the word being asked in the question, the word immediately before the predetermined period of silence, or the repeated word is the word to be searched for; The processing device according to claim 2 , wherein the determination unit instructs the search execution unit to search for the search target phrase determined by the generative model.

6. 3. The processing device according to claim 2, further comprising a selection unit that selects one of a plurality of generative models as the generative model based on information about the conference, the content of the conference, or a predetermined rule.

7. The processing device according to claim 6, characterized in that the selection unit determines the accuracy of the meeting, the degree of response speed, the status of the meeting, the industry related to the meeting, whether the meeting is in a specific field or a meeting that emphasizes accuracy, and the level of confidentiality of the meeting based on information about the meeting and / or the content of the meeting, and selects, based on the determined content, a first generation model that is a natural language processing model fine-tuned to a specific field, or a second generation model that is a large-scale natural language processing model, as the generation model.

8. A processing method executed by a processing device, obtaining pre-conference information regarding a conference in which the user will be participating; converting voice data of each user participating in the conference into text data and associating time information with each text; a step of determining whether or not it is necessary to search for words spoken during the conference based on the text data; If the search is necessary, a search is performed on the search target phrase; a step of displaying the search results obtained in the searching step on each user terminal used by each user; A processing method comprising:

9. obtaining pre-conference information regarding a conference in which the user will participate; converting voice data of each user participating in the conference into text data and associating time information with each text; a step of determining whether or not it is necessary to search for words spoken during the conference based on the text data; If the search is necessary, a search is performed on the search target phrase; a step of displaying the search results obtained in the searching step on each user terminal used by each user; A processing program that causes a computer to execute the above.

Citation Information

Patent Citations

  • Web meeting system, delegation method of organizer authority in web meeting system and web meeting system program

    JP2015230532A