Online meeting system, and program of online meeting system
The online meeting system addresses the challenge of reviewing video conference content by converting voice data to text and displaying it chronologically within the browser, allowing users to efficiently record and review discussions without needing to watch entire recordings.
Patent Information
- Application Number
- JP2023205335
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-19
AI Technical Summary
Conventional video conference systems make it difficult for users to easily confirm the content of online meetings after the fact, as they require viewing entire recording data, which is time-consuming and inefficient.
An online meeting system that includes a user terminal and a server, where the user terminal communicates with other user terminals via the server to conduct online meetings. The system features a browser-based interface with a 'minutes and transcription' button, allowing users to easily record and review discussions by converting voice data to text and displaying it chronologically, along with user identification information.
Enables users to efficiently record and review online meeting discussions by displaying text data of speech content in chronological order with user identification, allowing for easy identification of who said what and facilitating quick access to meeting summaries without needing to watch entire recordings.
Smart Images

Figure 2025091431000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an online meeting system in which a user terminal and a server are provided, and the user terminal communicates with other user terminals via the server to conduct an online meeting in the browser of the user terminal, and a program thereof.
Background Art
[0002] Conventionally, in a video conference system that can use additional functions with usage restrictions, when a first user who has no usage restrictions on the additional functions and a second user who has usage restrictions on the additional functions participate in a video conference, if the number of participants of the second user in the video conference is less than or equal to a predetermined allowable number, the use of the additional function is permitted, and if the number of participants of the second user in the video conference exceeds the allowable number, a video conference system including a restriction unit that restricts the use of the additional function is known (for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the above-described conventional video conference system, since the recording function of the video conference, which is an additional function, can be used by charging, in order for the user to confirm the content of the video conference, which is an online meeting, after the meeting, it is necessary to view the entire recording data over time after the meeting, and there is a problem that it is difficult to confirm easily.
[0005] Therefore, the present invention solves the problems of the prior art as described above. That is, the object of the present invention is to enable a user to easily record who said what and what kind of exchanges took place regarding the content of a video conference, which is an online meeting, without spending time watching the entire recorded data after the meeting. The present invention provides an online meeting system and its program for achieving this purpose.
Means for Solving the Problems
[0006] The invention according to claim 1 is an online meeting system including a user terminal and a server, where the user terminal communicates with other user terminals via the server and conducts an online meeting on the browser of the user terminal. The browser of the user terminal or the online meeting screen displayed on the browser has a minutes and transcription button. When the minutes and transcription button is operated, the browser displays a minutes and transcription window that can switchably display the minutes content and the transcription content. When the voice of a user is input via the microphone of any user terminal, the browser converts the input voice data into text data and displays the converted text data in the minutes and transcription window in chronological order together with the user identification information of the user terminal where the voice input occurred, thereby solving the above-described problems.
[0007] The invention according to claim 2, in addition to the configuration of the online meeting system described in claim 1, when there is a minutes creation operation in the browser or the minutes and transcription window, or at predetermined intervals, the browser uses the pre-trained artificial intelligence means of the server to generate minutes data including participant information, online meeting date information, start time and end time information, summary information of opinions by participant, summary information of current problems, and future issue information based on the text data of the transcription, and displays the generated minutes data in the minutes and transcription window, thereby further solving the above-described problems.
[0008] The invention according to claim 3, in addition to the configuration of the online meeting system described in claim 2, when there is a download operation in the browser or the minutes / transcription window, the browser can freely select at least one of the minutes data displayed in the minutes / transcription window and the text data of the transcription, and can freely select the file format to download to the user terminal where the download operation was performed, thereby further solving the above-described problems.
[0009] The invention according to claim 4, in addition to the configuration of the online meeting system described in claim 3, when there is a screen capture operation in the browser or the minutes / transcription window, the browser inserts the screen capture data into the time series of the text data of the transcription in the minutes / transcription window and displays it, thereby further solving the above-described problems.
[0010] The invention according to claim 5, in addition to the configuration of the online meeting system described in any one of claims 1 to 4, in the browser or the minutes / transcription window, the text data of the transcription can be modified, thereby further solving the above-described problems.
[0011] The invention according to claim 6, in addition to the configuration of the online meeting system described in any one of claims 2 to 4, in the browser or the minutes / transcription window, the minutes data can be modified, thereby further solving the above-described problems.
[0012] The invention according to claim 7, in addition to the configuration of the online meeting system described in any one of claims 1 to 4, the browser has a dictionary registration function for changing the text before change to the text after change, and when there is a same location as the text before change registered in the text data when converting from voice data to text data, the corresponding location is changed to the text after change and displayed in the minutes and transcription window, thereby further solving the above-described problems.
[0013] The invention according to claim 8, in addition to the configuration of the online meeting system described in any one of claims 1 to 4, the browser has an inappropriate content check function for checking whether the text data contains inappropriate words, and when there is a same location as the predetermined word registered in advance in the text data when converting from voice data to text data, the corresponding location is displayed in the minutes and transcription window with highlighting, and when the speaker is the user himself / herself of the user terminal, the browser displays that there is an inappropriate speech in the user's speech content, thereby further solving the above-described problems.
[0014] The invention according to claim 9, in addition to the configuration of the online meeting system described in any one of claims 1 to 4, when there is a video capture operation in the browser or the minutes and transcription window, the browser inserts the video capture data into the time series of the transcribed text data in the minutes and transcription window for display, and based on the video shooting start time information and the video shooting end time information, the location of the transcribed text data corresponding to the video capture is displayed as the corresponding location of the video capture, thereby further solving the above-described problems.
[0015] In addition to the configuration of the online meeting system described in any one of claims 1 to 4, in the administrator screen of the administrator terminal capable of communicating with the server according to the invention of claim 10, masking of word attributes included in text data from the browser of the user terminal to the pre-trained artificial intelligence means can be set freely. When the voice data input in the browser of the user terminal is converted into text data including the characters of the word attributes of the preset masking and transmitted to the server, the received server determines that the received text data includes the characters of the word attributes of the masking, and among the received data, after converting the characters of the word attributes of the masking into any other arbitrary characters of the same word attribute, the minutes data generation of the content based on the converted text data is executed using the pre-trained artificial intelligence means. Among the minutes data obtained by the execution, after changing the part of the other arbitrary characters converted at the time of transmission to the pre-trained artificial intelligence means back to the characters before conversion, the changed minutes data is transmitted to the user terminal, and the user terminal displays the received changed minutes data in the minutes and speech-to-text window, thereby further solving the above-described problems.
[0016] In addition to the configuration of the online meeting system described in any one of claims 1 to 4, in the administrator screen of the administrator terminal capable of communicating with the server according to the invention of claim 11, the dictionary registration function for changing the text before change to the text after change can be set freely, thereby further solving the above-described problems.
[0017] The invention according to claim 12, in addition to the configuration of the online meeting system described in any one of claims 1 to 4, when the user terminal stores at least one of the text data of the speech-to-text conversion and the minutes data in the server, the user can freely set a browsing restriction for at least one of the text data of the speech-to-text conversion and the minutes data to be stored. Along with at least one of the text data of the speech-to-text conversion and the minutes data to be stored, the user browsing restriction information is transmitted to the server. The server stores at least one of the received text data of the speech-to-text conversion and the minutes data, and in the browser of the user terminal, at least one of the text data of the speech-to-text conversion and the minutes data can be browsed within the range of the user browsing restriction based on the user browsing restriction information stored in the server, thereby further solving the above-mentioned problems.
[0018] The invention according to claim 13 is a program for an online meeting system in which a user terminal communicates with another user terminal via a server and conducts an online meeting in the browser of the user terminal. It includes a button operation determination step for determining whether or not a minutes / speech-to-text button provided on the browser of the user terminal or on the online meeting screen displayed on the browser has been operated, and a window display step in which when the minutes / speech-to-text button is operated, the browser displays a minutes / speech-to-text window that can freely switch between displaying the minutes content and the speech-to-text content. When the voice of a user is input via the microphone of any user terminal, the browser converts the input voice data into text data and displays the converted text data in chronological order in the minutes / speech-to-text window together with the user identification information of the user terminal where the voice input occurred, thereby solving the above-mentioned problems.
Advantages of the Invention
[0019] The online meeting system of the present invention includes a user terminal and a server. Thus, not only can the user terminal communicate with other user terminals via the server and conduct an online meeting in the browser of the user terminal, but also the following specific effects can be achieved.
[0020] According to the online meeting system of the invention according to claim 1, when one or more users speak in an online meeting, the text data of the speech content is displayed in chronological order in the minutes and transcription window together with the user identification information of the speaking user. Therefore, it is possible to easily record who said what and what kind of interaction took place. In particular, even when someone speaks simultaneously, the text data converted from the voice data is displayed in chronological order in the minutes and transcription window together with the user identification information of the user terminal where the voice input occurred. Therefore, the user can easily confirm who said what.
[0021] According to the online meeting system of the invention according to claim 2, in addition to the effects achieved by the invention according to claim 1, when a minutes creation operation is performed, or automatically at predetermined time intervals, minutes data including participant information, online meeting date information, start time and end time information, summary information of opinions by participant, summary information of current problems, and future issue information is generated based on the transcribed text data and displayed in the minutes and transcription window. Therefore, the user does not need to summarize the minutes himself / herself and can easily obtain the minutes data.
[0022] According to the online meeting system of the invention according to claim 3, in addition to the effects achieved by the invention according to claim 2, when a download operation is performed, at least one of the minutes data and the transcribed text data is downloaded as data. Therefore, the user can easily save at least one of the minutes and the transcription as a record.
[0023] According to the online meeting system of the invention according to claim 4, in addition to the effects achieved by the invention according to claim 3, when a screen capture operation occurs, the screen capture data is immediately inserted into the time series of the text data of the speech-to-text conversion. Therefore, when the user looks back later, they can easily understand who specifically spoke about what. Furthermore, only by the screen capture operation, the screen capture data is immediately inserted into the time series of the text data of the speech-to-text conversion without an insertion operation. Therefore, it is possible to avoid a time lag between capture and insertion and prevent a deviation between the speech and the screen capture.
[0024] According to the online meeting system of the invention according to claim 5, in addition to the effects achieved by the invention according to any one of claims 1 to 4, for example, it becomes possible to correct when there is a mistake in the conversion of Chinese characters from voice data to text data, and the accuracy of the content of the minutes is improved by generating the minutes data again after the correction. Therefore, the user can obtain accurate minutes data.
[0025] According to the online meeting system of the invention according to claim 6, in addition to the effects achieved by the invention according to any one of claims 2 to 4, for example, it becomes possible to correct when there is a mistake in the Chinese characters of the minutes data, and the accuracy of the content of the minutes is improved by downloading the minutes data after the correction. Therefore, the user can obtain accurate minutes data.
[0026] According to the online meeting system of the invention according to claim 7, in addition to the effects achieved by the invention according to any one of claims 1 to 4, for example, when it is desired to display in alphabetic characters what is more easily displayed in katakana, by registering the katakana display as the text before change and the alphabetic display as the text after change, although it is input by voice and is about to be displayed in katakana, the katakana part is changed to alphabetic characters and displayed in the minutes / transcription window, so that the user can obtain the transcription text data of the intended display characters, and further, the user can obtain the minutes data of the intended display.
[0027] According to the online meeting system of the invention according to claim 8, in addition to the effects achieved by the invention according to any one of claims 1 to 4, for example, when the text data contains discriminatory terms, that part is highlighted, so that the user can pay attention to the communication. Furthermore, when the user himself / herself makes an inappropriate statement, since it is displayed on the browser of the user's own user terminal that there is an inappropriate statement in the user's statement content, the user can immediately retract the inappropriate statement or take measures such as apologizing for the inappropriate statement.
[0028] According to the online meeting system of the invention according to claim 9, in addition to the effects achieved by the invention according to any one of claims 1 to 4, when there is a video capture start operation and an end operation as video capture operations, immediately the video capture data is inserted into the time series of the transcription text data, and the part of the transcription text data corresponding to the video capture is displayed as the corresponding part of the video capture, so that the user can easily understand which part of the communication from where to where is the video capture when looking back later. Furthermore, since the video capture data is inserted into the time series of the text data of the speech-to-text at just after the end operation without an insertion operation, only by the start / end operations of the video capture, it is possible to avoid a time lag between the capture and the insertion and prevent a deviation between the speech and the video from occurring.
[0029] According to the online meeting system of the invention according to claim 10, in addition to the effects achieved by the invention according to any one of claims 1 to 4, for example, by presetting masking for word attributes that can identify an individual, such as name, address, phone number, email address, etc., for text data from the user terminal to the server, when the text data converted from the input voice data by the browser contains characters of those word attributes, the server converts the characters of those word attributes into any other arbitrary characters of the same word attribute and then transmits the converted data to the pre-trained artificial intelligence means, and since the execution of the minutes data generation is performed based on the converted text data, it is possible to avoid ChatGPT, which is an example of pre-trained artificial intelligence means, from remembering information on characters of word attributes that can identify an individual. That is, since the content transmitted to ChatGPT, which is an example of pre-trained artificial intelligence means, is converted into characters that cannot identify an individual, it is possible to avoid leakage of information that can identify an individual to other terminals via ChatGPT, which is an example of pre-trained artificial intelligence means. Furthermore, when the server receives the minutes data from the pre-trained artificial intelligence means, among the received minutes data, a change is made to return the part of any other arbitrary characters converted at the time of transmission to the pre-trained artificial intelligence means to the characters before conversion, and the changed minutes data is transmitted to the user terminal, so that it is possible to obtain the changed minutes data that basically follows the intention of the exchange of the text data converted by the browser. That is, since the characters converted when the server sends them to the pre-trained artificial intelligence means are restored to the original characters when the server receives them from the pre-trained artificial intelligence means, it is possible to avoid involuntary information leakage to the pre-trained artificial intelligence means and involuntary information leakage to other terminals via the pre-trained artificial intelligence means, and to ensure the consistency between the content of the text data of the speech-to-text conversion from the user terminal and the content of the minutes data after the change to the user terminal.
[0030] According to the online meeting system of the invention according to claim 11, in addition to the effects achieved by the invention according to any one of claims 1 to 4, since there is no need for each user to set the dictionary registration function, and the content of the dictionary registration function set by the administrator is reflected in the browser of each user terminal, each user can obtain the speech-to-text text data of the display intended by the administrator, and further, can obtain the minutes data of the display intended by the user, and the variation in the accuracy of the speech-to-text text data and the minutes data can be made substantially uniform among each user.
[0031] According to the online meeting system of the invention according to claim 12, in addition to the effects achieved by the invention according to any one of claims 1 to 4, since user viewing restrictions are appropriately set for the speech-to-text text data and the minutes data of the online meeting and are made public within the set range, users can share information about the content of the online meeting within the set range.
[0032] According to the online meeting system of the invention according to claim 13, similar to the effect achieved by the invention according to claim 1, when one or more users speak in an online meeting, the text data of the speech content is displayed in chronological order in the minutes and speech-to-text window together with the user identification information of the speaking user, so that it is possible to easily record who said what and what kind of interaction took place. In particular, even when someone speaks simultaneously, the text data converted from the voice data is displayed in chronological order in the minutes and transcription window together with the user identification information of the user terminal where the voice input was made, so that the user can easily check who said what.
Brief Description of the Drawings
[0033]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Embodiments for Carrying Out the Invention
[0034] The online meeting system of the present invention includes a user terminal and a server. The browser of the user terminal or the online meeting screen displayed on the browser has a minutes and transcription button. When the minutes and transcription button is operated, the browser displays a minutes and transcription window that can freely switch between the minutes content and the transcription content. When the voice of a user is input via the microphone of any user terminal, the browser converts the input voice data into text data and displays the converted text data in the minutes and transcription window in chronological order together with the user identification information of the user terminal where the voice input occurred. As long as it is possible to easily record who said what and what kind of exchanges took place in the online meeting, the specific implementation manner can be any kind. Also, the program of the online meeting system of the present invention includes a button operation determination step for determining whether or not the minutes and transcription button provided on the browser of the user terminal or the online meeting screen displayed on the browser has been operated, a window display step for displaying a minutes and transcription window that can freely switch between the minutes content and the transcription content when the minutes and transcription button is operated, and a transcription text display step for converting the input voice data into text data and displaying the converted text data in the minutes and transcription window in chronological order together with the user identification information of the user terminal where the voice input occurred when the voice of a user is input via the microphone of any user terminal. As long as it is possible to easily record who said what and what kind of exchanges took place in the online meeting, the specific implementation manner can be any kind.
[0035] For example, the user terminal can send and receive information such as a desktop personal computer terminal, a notebook personal computer terminal, a smartphone terminal, or a tablet terminal, and can be connected to a server through a communication network including a wide area network such as the so-called Internet, a local network, a telephone line, or the like. Any device can be used as long as it can be connected to the server. Also, the server may be a single server or multiple servers on the cloud. The pre-trained artificial intelligence means is constituted by, for example, an interactive large language model also called ChatGPT (Generative Pre-trained Transformer) (hereinafter referred to as ChatGPT), etc., and may be provided in a configuration provided on a single server or multiple servers on the cloud.
Example
[0036] Hereinafter, the online meeting system 100 which is an embodiment of the present invention will be described with reference to FIGS. 1 to 18. Here, FIG. 1 is a diagram showing the concept of the online meeting system 100 which is an embodiment of the present invention, FIG. 2 is a chart diagram showing an operation example of the online meeting system 100 which is an embodiment of the present invention, FIG. 3(A) is a diagram showing an example of the online meeting screen 112 of the browser 111 of the user terminal 110 of the online meeting system 100 which is an embodiment of the present invention, FIG. 3(B) is a diagram showing an example of the first minutes and transcription window 113, FIG. 4 is a diagram showing an example of the transcribed text data TX of the first minutes and transcription window 113 of the online meeting system 100 which is an embodiment of the present invention, FIG. 5 is a diagram showing an example of the minutes data MN of the first minutes and transcription window 113 of the online meeting system 100 which is an embodiment of the present invention, FIG. 6 is a diagram showing an example of the download confirmation window 114 of the online meeting system 100 which is an embodiment of the present invention, FIG. 7(A) is a diagram showing the minutes data MN of the preview display screen PV of the browser 111 of the online meeting system 100 which is an embodiment of the present invention, FIG. 7(B) is a diagram showing the transcribed text data TX of the preview display screen PV, FIG. 8 is a diagram showing an example of the menu item window 115 as the screen capture operation of the online meeting system 100 which is an embodiment of the present invention, FIG. 9 is a diagram showing a state where the screen capture data CP1 is inserted into the transcribed text data TX of the first minutes and transcription window 113 of the online meeting system 100 which is an embodiment of the present invention, FIG. 10 is a diagram showing a state of the screen capture operation when the browser 111 of the online meeting system 100 which is an embodiment of the present invention is displaying an arbitrary screen other than the online meeting screen 112, FIG. 11(A) is a diagram showing an example of the dictionary registration window 117A of the online meeting system 100 which is an embodiment of the present invention, FIG. 11(B) is a diagram showing a state where the word DC1 to be converted in the transcribed text data TX is converted to the conversion candidate DC2 and displayed in the first minutes and transcription window 113, FIG. 12(A) isFIG. 0 is a diagram showing an example of an inappropriate speech registration window 117B of the online meeting system 100 according to an embodiment of the present invention. FIG. 12(B) is a diagram showing a state in which a word UL1 determined to be inappropriate in the first minutes / transcription window 113 is highlighted UL3. FIG. 12(C) is a diagram showing an example of a warning display window UL4. FIG. 13 is a diagram showing a state in which video capture data MV is inserted into the transcribed text data TX of the first minutes / transcription window 113 of the online meeting system 100 according to an embodiment of the present invention. FIG. 14 is a diagram showing an example of a masking setting 151 on the administrator screen 150 of the online meeting system 100 according to an embodiment of the present invention. FIG. 15 is a diagram for explaining an example of the operation of masking of the online meeting system 100 according to an embodiment of the present invention. FIG. 16 is a diagram showing an example of a dictionary registration setting and an inappropriate speech registration setting on the administrator screen 150 of the online meeting system 100 according to an embodiment of the present invention. FIG. 17 is a diagram showing an example of a server storage confirmation window 118 of the online meeting system 100 according to an embodiment of the present invention. FIG. 18 is a diagram showing an example of a minutes list screen 119 of the browser 111 of the user terminal 110 of the online meeting system 100 according to an embodiment of the present invention.
[0037] As shown in FIG. 1, the online meeting system 100 according to an embodiment of the present invention includes a user terminal 110, a first server 120, a second server 130 which are examples of servers, and an online meeting server (not shown), and another user terminal 140. Among these, the first server 120 has a database 121 and a bot which is also a display program. The second server 130 has ChatGPT which is an example of pre-trained artificial intelligence means 131. The online meeting server provides the service of the online meeting itself to the user terminal 110. Note that, as a technical concept, the first server 120, the second server 130, and the server for online meetings may have a physically common configuration with each other or may have physically different configurations. The user terminal 110 is provided to communicate with other user terminals 140 via the server for online meetings and conduct an online meeting on the browser 111 of the user terminal 110.
[0038] Specifically, the browser 111 of the user terminal 110 or the online meeting screen 112, which is an example of the online meeting screen 112 displayed on the browser 111, has a first minutes and transcription button 112a that is a minutes and transcription button. When the first minutes and transcription button 112a is operated, the browser 111 displays a first minutes and transcription window 113 that can freely switch and display the minutes content and the transcription content. Note that instead of operating the first minutes and transcription button 112a, an operation of displaying a menu by right-clicking and selecting an item may be used, or a shortcut key operation by keyboard input may also be used.
[0039] Furthermore, assume that the voice of another user is input via the microphone of another user terminal 140 as an example of any user terminal 110. Then, the other user terminal 140 with voice input transmits voice data and the user identification information US of the other user terminal 140 with the voice input to the user terminal 110 that is the destination via the server for online meetings. The browser 111 of the received user terminal 110 converts the input voice data into text data TX. Furthermore, the browser 111 is configured to display the converted text data TX in the first minutes and transcription window 113 in chronological order together with the user identification information US of the other user terminal 140 with the voice input.
[0040] As a result, when one or more users speak during an online meeting, the text data TX of the speech content is displayed in chronological order in the first minutes and transcription window 113 together with the user identification information US of the user who made the speech. As a result, it is possible to easily record who said what and what kind of conversation took place. In particular, even when multiple users speak simultaneously, the text data TX converted from the audio data is displayed in chronological order in the first minutes and transcription window 113 together with the user identification information US of the user terminal 110 where the audio input was made, so that the user can easily confirm who said what.
[0041] Next, the operation of the online meeting system 100 will be described in more detail. As shown in FIG. 2, in step S1, as an online meeting screen display determination step, the browser 111 of the user terminal 110 determines whether or not the online meeting screen 112 is being displayed. More specifically, as shown in FIG. 3(A), the browser 111 of the user terminal 110 determines whether or not the online meeting screen 112 is being displayed.
[0042] For example, when the user clicks on the URL for the online meeting described on the schedule screen (not shown), the browser 111 accesses the online meeting server and displays the online meeting screen 112. Here, the online meeting screen 112 is provided with a first minutes and transcription button 112a. If it is determined that the browser 111 is displaying the online meeting screen 112, the process proceeds to step S2. On the other hand, if it is determined that the screen is not yet displayed, step S1 is repeated.
[0043] In step S2, as a button operation determination step, the browser 111 of the user terminal 110 or the browser 111 determines whether the first minutes transcription button 112a, which is a minutes transcription button provided on the online meeting screen 112, an example of the online meeting screen 112 displayed on the browser 111, has been operated. More specifically, as shown in FIG. 3(A), the browser 111 determines whether the first minutes transcription button 112a on the online meeting screen 112 has been clicked or tapped with a mouse. If it is determined that the button has been operated, the process proceeds to step S3. On the other hand, if it is determined that the button has not been operated yet, step S2 is repeated. Note that the browser 111 may also determine whether the second minutes transcription button 111a, which is a minutes transcription button provided on the browser 111, has been operated.
[0044] In step S3, as a window display step, as shown in FIG. 3(B), the browser 111 displays a first minutes transcription window 113 that can freely switch between displaying minutes content and transcription content. Here, as shown in FIG. 4, the first minutes transcription window 113 is provided with, as an example, a meeting name input field 113a, a "Minutes" tab 113b, a "Transcription" tab 113c, a transcription start button 113d, a transcription pause button 113e, a transcription end button 113f, a reset button 113g, a minutes creation button 113h, a minutes file creation button 113k, a download button 113m, a server save button 113n, and other buttons 113p.
[0045] Among these, for the meeting name input field 113a, the user can freely input an arbitrary meeting name. Also, by selecting the "Minutes" tab 113b and the "Transcription" tab 113c, the display can be switched between the text data TX of the transcription and the minutes data MN. Furthermore, when the speech-to-text start button 113d is operated, it is provided such that the conversion from the voice data to the text data TX of the speech-to-text is started when there is voice input. Also, when the speech-to-text pause button 113e is operated, it is provided such that the conversion from the voice data by voice input to the text data TX of the speech-to-text is paused (interrupted). Furthermore, when the speech-to-text end button 113f is operated, it is provided such that the conversion from the voice data by voice input to the text data TX of the speech-to-text is completed, and a guidance for creating the minutes data MN is displayed.
[0046] Also, when the reset button 113g is operated, it is provided such that the text data TX of the speech-to-text displayed in the first minutes·speech-to-text window 113 is erased. Furthermore, when the minutes creation button 113h is operated, it is provided such that the minutes data MN is generated using ChatGPT, which is an example of the pre-trained artificial intelligence means 131, based on the content of the text data TX of the speech-to-text and the user identification information US. Also, when the minutes file creation button 113k is operated, it is provided such that the generated minutes data MN is created in a file such as PDF, simple text, cell display format (CSV, spreadsheet file).
[0047] Furthermore, when the download button 113m is operated, after selecting the target selection 114a, file format selection 114b, and minutes type selection 114c of the text data TX of the speech-to-text and the minutes data MN, it is provided such that the file of the selected content is downloaded to the user terminal 110. Also, when the server save button 113n is operated, after selecting the target selection 118a, minutes type selection 118b, and sharing range setting 118c of the text data TX of the speech-to-text and the minutes data MN, it is provided such that the data of the selected content is saved to the first server 120. Furthermore, when the other button 113p is operated, the menu item window 115 is displayed and provided so that operations can be freely performed on the items. In addition to operating these buttons, an operation of displaying a menu by right-clicking and selecting an item may be used, or a shortcut key operation by keyboard input may also be used.
[0048] And, as an example, the user may operate the speech recognition start button 113d of the first minutes and speech recognition window 113, or may be in a voice input waiting state in a speech recognition start state with the display of the first minutes and speech recognition window 113 by operating the first minutes and speech recognition button 112a.
[0049] In step S4, as a voice input presence / absence determination step, the browser 111 determines whether the user's voice has been input via the microphone of any of the user terminals 110. If it is determined that there is voice input, the process proceeds to step S5. On the other hand, if it is determined that there is still no voice input, step S4 is repeated. For example, assume that Emi Suzuki, who is a user, inputs voice via the microphone of the user terminal 110 saying, "Good morning, Mr. Tanaka." In this case, the browser 111 of the user terminal 110 recognizes the user identification information US "Emi Suzuki" and the voice data. Subsequently, assume that Ichiro Tanaka, who is another user, inputs voice via the microphone of another user terminal 140 saying, "Good morning, Ms. Suzuki. Please take care of me today." Then, the user identification information US "Ichiro Tanaka" and the voice data are transmitted from the other user terminal 140 to the user terminal 110 via the online meeting server.
[0050] In step S5, as a voice data text conversion step (speech recognition text display step), the browser 111 of the user terminal 110 of Emi Suzuki converts the input voice data into text data TX. More specifically, it is converted from the voice data into text data TX of the character transcription to the effect of "Good morning, Mr. Tanaka." Subsequently, it is converted from the voice data received from another user terminal 140 of Ichiro Tanaka into text data TX of the character transcription to the effect of "Good morning, Mr. Suzuki. Please take care of me today."
[0051] In step S6, as the character transcription text display step, the converted text data TX is displayed in chronological order in the first minutes - character transcription window 113 together with the user identification information US of the user terminal 110 where the voice input was made. More specifically, as shown in FIG. 4, the text data TX of the character transcription to the effect of "Good morning, Mr. Tanaka." is displayed in the first minutes - character transcription window 113 together with the user identification information US "Emi Suzuki". Subsequently, the text data TX of the character transcription to the effect of "Good morning, Mr. Suzuki. Please take care of me today." is displayed in the first minutes - character transcription window 113 together with the user identification information US "Ichiro Tanaka". At this time, the time information may also be displayed together with the character - transcribed text data TX and the user identification information US.
[0052] As a result, as described above, when one or more users speak in an online meeting, the text data TX of the speech content is displayed in chronological order in the first minutes - character transcription window 113 together with the user identification information US of the user who made the speech. As a result, it is possible to easily record who said what and what kind of interaction took place. In particular, even when someone speaks simultaneously, since the text data TX converted from the voice data is displayed in chronological order in the first minutes - character transcription window 113 together with the user identification information US of the user terminal 110 where the voice input was made, the user can easily confirm who said what.
[0053] Furthermore, in this embodiment, as shown in FIG. 4, as an example of the minutes creation operation in the browser 111 or the minutes / transcription window, when the minutes creation button 113h of the first minutes / transcription window 113 is operated, or, as an example every predetermined time, every 3 minutes, the browser 111 is provided to create minutes data MN. More specifically, the browser 111 of the user terminal 110 transmits the text data TX of the transcription selected and displayed in the "Transcription" tab 113c of the first minutes / transcription window 113 and the user identification information US to ChatGPT, which is an example of the pre-trained artificial intelligence means 131 of the second server 130, via, for example, the first server 120. Then, using ChatGPT, which is an example of the pre-trained artificial intelligence means 131 of the second server 130, minutes data MN including participant information PT, online meeting date information DT, start time / end time information SE, summary information OP of opinions by each participant, summary information PP of current problems, and future issue information AS is generated based on the text data TX of the transcription and the user identification information US. Subsequently, as shown in FIG. 5, the browser 111 is configured to display the generated minutes data MN in the first minutes / transcription window 113.
[0054] Thereby, when there is a minutes creation operation or automatically every predetermined time, minutes data MN including participant information PT, online meeting date information DT, start time / end time information SE, summary information OP of opinions by each participant, summary information PP of current problems, and future issue information AS is generated based on the text data TX of the transcription and is displayed in the first minutes / transcription window 113. As a result, the user does not need to summarize the minutes himself / herself and can easily obtain the minutes data MN.
[0055] Instead of operating the minutes creation button 113h in the first minutes transcription window 113, which is an example of the minutes creation operation, an operation of displaying a menu by right-clicking and selecting an item may be performed, or a shortcut key operation by keyboard input may be performed. Regarding the content of the minutes data MN, in addition to the participant information PT and the online meeting date information DT, it is sufficient if at least one of the following three types of information is included: the summary information OP of the opinions of each participant, the summary information PP of the current problems, and the future issue information AS. Furthermore, when the "Minutes" tab 113b is selected and displayed in the first minutes transcription window 113, a clipboard button 113q is provided so as to be displayed. When the clipboard button 113q is operated, the content of the minutes data MN, which is the content of the selection and display of the "Minutes" tab 113b in the first minutes transcription window 113, is configured to be temporarily saved in the clipboard. The user can paste the content of the minutes data MN into an email or the like by performing a paste operation in an email or the like and transmit it to a specific other user or the like. Also, when the browser 111 generates the minutes data MN based on the transcribed text data TX and the user identification information US using ChatGPT, which is an example of the pre-trained artificial intelligence means 131 of the second server 130, it may or may not pass through the first server 120.
[0056] Also, in this embodiment, when there is an operation of the download button 113m as an example of the download operation in the browser 111 or the minutes transcription window, the browser 111 is configured to download at least one of the minutes data MN displayed in the first minutes transcription window 113 and the transcribed text data TX to the user terminal 110 that has performed the download operation, freely selectable, and freely selectable in file format. More specifically, as shown in FIG. 5, assume that as an example of a download operation in the browser 111, an operation is performed on the download button 113m of the first minutes and transcription window 113.
[0057] Then, as shown in FIG. 6, the browser 111 displays a download confirmation window 114. In the download confirmation window 114, a target selection 114a for download, a file format selection 114b for download, and a minutes type selection 114c are displayed so that they can be freely selected. Further, a "Preview / Edit" button 114d and a "Download" button 114e are provided.
[0058] Thus, when there is a download operation, if the user makes a desired target selection 114a, file format selection 114b, and minutes type selection 114c and operates the "Download" button 114e, a file of the selected content is downloaded. That is, at least one of the minutes data MN and the transcribed text data TX is downloaded as data. As a result, the user can easily save at least one of the minutes and the transcription as a record. Note that instead of the operation on the download button 113m which is an example of the download operation, an operation of displaying a menu by right-clicking and selecting an item may be used, or a shortcut key operation by keyboard input may be used.
[0059] Furthermore, in this embodiment, in the browser 111 or the first minutes and transcription window 113, the transcribed text data TX is configured to be modifiable. More specifically, as shown in FIG. 4, when the "Transcription" tab 113c of the first minutes and transcription window 113 is selected and the transcribed text data TX is displayed in the first minutes and transcription window 113, when the user clicks on a desired location and selects it with the cursor, the selected location can be changed. Also, when the cursor is moved to a desired location and keyboard input is made, the content entered from the keyboard is configured to be inserted.
[0060] Also, as shown in FIG. 6, it is assumed that the "Preview / Edit" button 114d of the download confirmation window 114 is operated. Then, as shown in FIGS. 7(A) and 7(B), the browser 111 displays a preview of the target on the preview display screen PV based on the download target selection 114a in the download confirmation window 114. On this preview display screen PV, when a desired location is clicked and selected with the cursor, the selected location can be freely changed. Also, when the cursor is moved to a desired location and keyboard input is made, the content entered from the keyboard is configured to be inserted. That is, as shown in FIG. 7(B), the browser 111 displays the text data TX of the speech-to-text conversion on the preview display screen PV in a freely modifiable manner.
[0061] As a result, for example, when there is an error in the kanji conversion from audio data to text data TX, it can be corrected, and by generating the minutes data MN again after the correction, the accuracy of the minutes content is improved. As a result, the user can obtain accurate minutes data MN.
[0062] Also, in this embodiment, in the browser 111 or the first minutes / speech-to-text window 113, the minutes data MN is configured to be freely modifiable. More specifically, as shown in FIG. 5, when the "Minutes" tab 113b of the first minutes / speech-to-text window 113 is selected and the minutes data MN is displayed in the first minutes / speech-to-text window 113, when the user clicks on a desired location and selects it with the cursor, the selected location can be freely changed. Also, when the cursor is moved to a desired location and keyboard input is made, the content entered from the keyboard is configured to be inserted. Also, as shown in FIG. 6, it is assumed that the "Preview / Edit" button 114d of the download confirmation window 114 is operated.
[0063] Then, as shown in FIGS. 7(A) and 7(B), the browser 111 displays a preview of the target on the preview display screen PV based on the download target selection 114a in the download confirmation window 114. On this preview display screen PV, when a desired location is clicked and selected with the cursor, the selected location can be freely changed. Also, when the cursor is moved to a desired location and keyboard input is made, the content input by the keyboard is configured to be inserted. That is, as shown in FIG. 7(A), the browser 111 displays the location of the minutes data MN on the preview display screen PV so that it can be freely modified.
[0064] As a result, for example, when there is a mistake in the Chinese characters of the minutes data MN, it can be corrected, and by downloading the minutes data MN after the correction, the accuracy of the content of the minutes is improved. As a result, the user can obtain accurate minutes data MN. Note that a download button PV1 is provided on the preview display screen PV. When the download button PV1 is operated, the data of the display content at that time, or the data of the corrected content if it has been corrected, is configured to be downloaded.
[0065] Furthermore, in this embodiment, when there is a screen capture operation that is a still image capture in the browser 111 or the minutes / voice recognition window, the browser 111 is configured to insert and display the screen capture data CP1 in the time series of the recognized text data TX of the first minutes / voice recognition window 113. More specifically, as shown in FIG. 8, as an example of the screen capture operation, when the other button 113p of the first minutes transcription window 113 is operated, the browser 111 displays the menu item window 115.
[0066] The menu item window 115 is provided with, as an example, a "screen capture" item 115a, a "video capture" item 115b, a "dictionary" item 115c, and an "inappropriate speech" item 115d. Among these, when a selection operation of the "screen capture" item 115a is performed, the browser 111 executes a screen capture for the online meeting screen 112. Then, as shown in FIG. 9, the browser 111 inserts and displays the screen capture data CP1 obtained by execution into the time series of the transcribed text data TX of the first minutes transcription window 113.
[0067] Also, the user may be a presenter who is a speaker rather than a listener in an online meeting. In this case, as shown in FIG. 10, there may be a time when the browser 111 displays another arbitrary screen ST other than the online meeting screen 112 by switching tabs of the browser 111. At this time, as an example of the screen capture operation, when the second minutes transcription button 111a provided on the browser 111 is operated, the browser 111 displays a second minutes transcription window 116 similar to the first minutes transcription window 113. The basic display content of the second minutes transcription window 116 is the same as the display content of the first minutes transcription window 113, and the transcribed text data TX and the minutes data MN to be displayed are synchronized. The difference lies in the arrangement of the buttons.
[0068] In the second minutes transcription window 116, as an example, there are provided a minutes file creation button 116a, a download button 116b, a dictionary button 116c, a screen capture button 116d, and a video capture button 116e. Among these, when there is a selection operation of the screen capture button 116d, the browser 111 executes screen capture for any other arbitrary screen ST selected by the tab. As shown in FIG. 10, the browser 111 inserts and displays the screen capture data CP2 of any other arbitrary screen ST obtained by execution into the time series of the transcribed text data TX of the second minutes transcription window 116. Note that the screen capture data CP2 of any other arbitrary screen ST inserted into the second minutes transcription window 116 is also inserted into the first minutes transcription window 113 having a synchronization relationship with the second minutes transcription window 116.
[0069] Thereby, when there is a screen capture operation, the screen capture data CP1 (CP2) is immediately inserted into the time series of the transcribed text data TX. As a result, when the user looks back later, the user can easily understand who specifically spoke about what. Furthermore, only by the screen capture operation, the screen capture data CP1 (CP2) is immediately inserted into the time series of the transcribed text data TX without an insertion operation. As a result, it is possible to avoid a time lag between capture and insertion and a deviation between speech and screen capture. Note that instead of the selection operation of the “screen capture” item 115a, which is an example of the screen capture operation for a still image, or the selection operation of the screen capture button 116d, an operation of displaying a menu by a right click and selecting an item may be used, or a shortcut key operation by keyboard input may be used.
[0070] Also, in this embodiment, the browser 111 has a dictionary registration function for changing the text before change to the text after change. And, when converting the audio data to the text data TX, if there is a location in the text data TX that is the same as the text before change registered in the text data TX, the corresponding location is changed to the text after change and displayed in the first minutes transcription window 113 or the second minutes transcription window 116. More specifically, as shown in FIG. 8, when the "Dictionary" item 115c of the menu item window 115 is selected or, as shown in FIG. 10, when the dictionary button 116c of the second minutes transcription window 116 is operated, the browser 111 displays the dictionary registration window 117A as shown in FIG. 11(A).
[0071] The dictionary registration window 117A is provided with a word DC1 to be converted, which is the text before change, and a conversion candidate DC2, which is the text after change, so that they can be registered freely. For example, assume that "ChatGPT" is registered as an example of the word DC1 to be converted and "ChatGPT" is registered as an example of the conversion candidate DC2. And when the browser 111 converts the input audio data to the text data TX of the transcription, if "ChatGPT" exists in the text data TX of the transcription, this location is changed to "ChatGPT". And, as shown in FIG. 11(B), the browser 111 is configured to display the content after change in the first minutes transcription window 113.
[0072] Thus, for example, when it is desired to display in alphabetic characters something that is likely to be displayed in katakana, by registering the katakana display as the text before change and the alphabetic display as the text after change, although it is input by voice and is likely to be displayed in katakana, the katakana portion is changed to alphabetic characters and displayed in the first minutes transcription window 113 or the second minutes transcription window 116. As a result, the character recognition text data TX of the display intended by the user can be obtained, and furthermore, the minutes data MN of the display intended by the user can be obtained.
[0073] Furthermore, in this embodiment, the browser 111 has an inappropriate content check function for checking whether the text data TX is inappropriate and contains a predetermined word. When there is a location in the text data TX that is the same as a predetermined word registered in advance in the text data TX when the browser 111 converts the voice data into the text data TX, the corresponding location is highlighted and displayed in the first minutes / character recognition window 113 or the second minutes / character recognition window 116 by the highlight display UL3. Furthermore, when the speaker is the user himself / herself of the user terminal 110, the browser 111 is configured to display that there is an inappropriate speech in the user's speech content. More specifically, as shown in FIG. 8, when the "inappropriate speech" item 115d of the menu item window 115 is selected, as shown in FIG. 12(A), the browser 111 displays the inappropriate speech registration window 117B. In the inappropriate speech registration window 117B, a word UL1 to be considered inappropriate and an attribute UL2 to be considered inappropriate are provided so that they can be registered freely.
[0074] For example, assume that "debu" is registered as an example of the word UL1 to be considered inappropriate. And when the browser 111 converts the input voice data into the character recognition text data TX, assume that "debu" exists in the character recognition text data TX. In this case, as shown in FIG. 12(B), the browser 111 highlights and displays the location of "debu", which is an example of the word UL1 to be considered inappropriate in the character recognition text data TX, in the first minutes / character recognition window 113. The same applies to other registered words UL1 to be considered inappropriate and locations of the attributes UL2 to be considered inappropriate. Furthermore, as shown in FIG. 12(C), when the speaker regarding the inappropriate word UL1 or the inappropriate attribute UL2 is the user himself / herself of the user terminal 110, the browser 111 is configured to display a warning display window UL4 as an example of indicating that there is an inappropriate utterance in the user's utterance content.
[0075] As a result, for example, when the text data TX includes discriminatory terms, that part is highlighted as UL3. As a result, the user can pay attention to the communication. Furthermore, when the user himself / herself makes an inappropriate utterance, the browser 111 of the user terminal 110 of the user himself / herself displays that there is an inappropriate utterance in the user's utterance content. As a result, the user can immediately retract the inappropriate utterance or take measures such as apologizing for the inappropriate utterance.
[0076] Also, in this embodiment, when there is a video capture operation in the browser 111 or the minutes / transcription window, the browser 111 inserts and displays the video capture data MV in the time series of the transcribed text data TX of the first minutes / transcription window 113 and the second minutes / transcription window 116. At the same time, the browser 111 is configured to display the part of the transcribed text data TX corresponding to the video capture as the video capture corresponding part based on the video shooting start time information and the video shooting end time information.
[0077] More specifically, as shown in FIG. 8, as an example of the video capture operation, when there is an operation of the other button 113p of the first minutes / transcription window 113, the browser 111 displays the menu item window 115. Among these, when there is a selection operation of the "Video Capture" item 115b, the browser 111 executes video capture for the online meeting screen 112. Note that this operation is a video recording start operation. Performing this operation again will result in a video recording end operation. Then, when the video recording end operation is performed, as shown in FIG. 13, the browser 111 inserts and displays the video capture data MV obtained by execution into the time series of the text data TX of the speech-to-text in the first minutes-of-meeting speech-to-text window 113.
[0078] Also, similar to the screen capture described above, the user may be a presenter who is a speaker rather than a listener in an online meeting. In this case, as shown in FIG. 10, when a selection operation of the video capture button 116e is performed, the browser 111 performs video capture on any other screen ST. Note that this operation is a video recording start operation. Performing this operation again will result in a video recording end operation.
[0079] Then, when the video recording end operation is performed, the browser 111 inserts and displays the video capture data MV of any other screen ST obtained by execution into the time series of the text data TX of the speech-to-text in the second minutes-of-meeting speech-to-text window 116. Note that the video capture data MV of any other screen ST inserted into the second minutes-of-meeting speech-to-text window 116 is also inserted into the first minutes-of-meeting speech-to-text window 113 that is in a synchronization relationship with the second minutes-of-meeting speech-to-text window 116. Furthermore, as shown in FIG. 13, the browser 111 is configured to display the location of the text data TX of the speech-to-text corresponding to the video capture as a video recording range display RG that is the corresponding location of the video capture based on the video recording start time information and the video recording end time information.
[0080] As a result, when there is a video capture start operation and an end operation as video capture operations, the video capture data MV is immediately inserted into the time series of the text data TX of the speech-to-text conversion, and the location of the text data TX of the speech-to-text conversion corresponding to the video capture is displayed as the corresponding location of the video capture. As a result, when the user looks back later, the user can easily understand which part of the conversation from where to where is the video capture. Furthermore, only with the start / end operations of the video capture, the video capture data MV is inserted into the time series of the text data TX of the speech-to-text conversion immediately after the end operation without an insertion operation. As a result, there is no time lag between capture and insertion, and it is possible to avoid a deviation between the speech and the video. Note that instead of the selection operation of the "Video Capture" item 115b or the selection operation of the video capture button 116e, which is an example of the video capture operation, an operation of displaying a menu by a right click and selecting an item may be used, or a shortcut key operation by keyboard input may be used.
[0081] Furthermore, in this embodiment, as shown in FIG. 14, on the administrator screen 150 of the administrator terminal capable of communicating with the first server 120, a masking setting 151 for the word attributes included in the text data TX to ChatGPT, which is an example of the pre-trained artificial intelligence means 131, can be freely set from the browser 111 of the user terminal 110. Then, as shown in FIG. 15, the voice data input in the browser 111 of the user terminal 110 is converted into text data TX including the character WD1 of the pre-set masked word attribute and transmitted to the first server 120. Then, the received first server 120 determines that the received text data TX includes the character WD1 of the masked word attribute.
[0082] Then, among the received data, the first server 120 converts the character WD1 of the masking word attribute in the received data into another arbitrary character WD2 of the same word attribute, and then uses ChatGPT, which is an example of the pre-trained artificial intelligence means 131 of the second server 130, to generate minutes data MN of the content based on the converted text data TX. Furthermore, the first server 120 makes a change to restore the part of another arbitrary character WD2 that was converted when transmitting to ChatGPT, which is an example of the pre-trained artificial intelligence means 131 of the second server 130, in the minutes data MN obtained by execution, back to the character WD1 before conversion. Then, the first server 120 transmits the modified minutes data MN to the user terminal 110.
[0083] Then, the user terminal 110 is configured to display the received modified minutes data MN in the first minutes and transcription window 113. For example, as shown in FIG. 15, assume that the text data TX of the transcription includes name information, which is an example of the character WD1 of the masking word attribute as the employee information participating in the online meeting. Among the text data TX of the transcription and the user identification information US received by the first server 120 from the user terminal 110, the name information "Emi Suzuki, Taro Sato, Ichiro Tanaka..." which is an example of the character WD1 of the masking word attribute, is converted into another arbitrary character WD2 of the same word attribute, for example, "Name 1, Name 2, Name 3...".
[0084] Then, the first server 120 transmits the converted text data TX of the transcription and the converted user identification information US to the second server 130, and uses ChatGPT, which is an example of the pre-trained artificial intelligence means 131, to execute the generation of minutes data MN based on the converted text data TX of the transcription and the converted user identification information US. Subsequently, the first server 120 makes a change to revert the part "Name 1, Name 2, Name 3..." of another arbitrary character WD2 converted during the transmission to ChatGPT, which is an example of the pre-trained AI means 131 of the second server 130, back to the character WD1 "Emi Suzuki, Taro Sato, Ichiro Tanaka..." before conversion in the minutes data MN obtained by execution. Then, the first server 120 transmits the modified minutes data MN to the user terminal 110.
[0085] Thus, for example, by presetting masking for word attributes that can identify individuals such as names, addresses, phone numbers, email addresses, etc. for the text data TX from the user terminal 110 to the first server 120, if the text data TX converted from the voice data input by the browser 111 contains the character WD1 of those word attributes, the first server 120 converts the character WD1 of those word attributes to another arbitrary character WD2 of the same word attribute and then transmits the converted data to ChatGPT, which is an example of the pre-trained AI means 131, and the generation and execution of the minutes data MN are performed based on the converted text data TX. As a result, it is possible to avoid ChatGPT, which is an example of the pre-trained AI means 131, from memorizing the information of the character WD1 of the word attribute that can identify individuals. That is, the content transmitted to ChatGPT, which is an example of the pre-trained AI means 131, is converted into characters that cannot identify individuals. As a result, it is possible to avoid leakage of information that can identify individuals to other terminals via ChatGPT, which is an example of the pre-trained AI means 131.
[0086] Furthermore, when the first server 120 receives the minutes data MN from ChatGPT, which is an example of the pre-trained artificial intelligence means 131 of the second server 130, among the received minutes data MN, the part of another arbitrary character WD2 that was converted when transmitting to ChatGPT, which is an example of the pre-trained artificial intelligence means 131 of the second server 130, is changed back to the character WD1 before conversion, and the changed minutes data MN is transmitted to the user terminal 110. As a result, the browser 111 can obtain the changed minutes data MN that basically follows the intention of the exchange of the text data TX that it has converted. That is, the character WD2 that was converted when the first server 120 transmits to ChatGPT, which is an example of the pre-trained artificial intelligence means 131 of the second server 130, is reverted to the original character WD1 when the first server 120 receives from ChatGPT, which is an example of the pre-trained artificial intelligence means 131 of the second server 130. As a result, it is possible to avoid the involuntary leakage of information to ChatGPT, which is an example of the pre-trained artificial intelligence means 131, and the involuntary leakage of information to other terminals via ChatGPT, which is an example of the pre-trained artificial intelligence means 131, and to ensure the consistency between the content of the text data TX transcribed from the user terminal 110 and the content of the changed minutes data MN transmitted to the user terminal 110.
[0087] Also, in this embodiment, as shown in FIG. 16, in the administrator screen 150 of the administrator terminal that can communicate with the first server 120, the dictionary registration function for changing the text before the change to the text after the change can be freely set. More specifically, on the administrator screen 150, words to be converted DC1 and conversion candidates DC2 can be freely registered, similar to the dictionary registration window 117A of FIG. 11(A) described above.
[0088] As a result, it is not necessary for each user to set the dictionary registration function, and the content of the dictionary registration function set by the administrator is reflected in the browser 111 of each user terminal 110. As a result, each user can obtain the speech-to-text data TX of the display as intended by the administrator, and can further obtain the minutes data MN of the display as intended by the user. Among the users, the variation in the accuracy of the speech-to-text data TX and the minutes data MN can be made substantially uniform.
[0089] Furthermore, in this embodiment, as shown in FIG. 16, on the administrator screen 150 of the administrator terminal that can communicate with the first server 120, the inappropriate speech pointing function for pointing out inappropriate speech can be set freely. More specifically, on the administrator screen 150, words UL1 that are considered inappropriate and attributes UL2 that are considered inappropriate are provided so that they can be registered freely, similar to the inappropriate speech registration window 117B in FIG. 12(A) described above.
[0090] As a result, there is no need for each user to set the inappropriate speech pointing function, and the content of the inappropriate speech pointing function set by the administrator is reflected in the browser 111 of each user terminal 110. As a result, each user can obtain the minutes data MN with good manners, and among the users, the variation in the accuracy of the inappropriate speech in the speech-to-text data TX and the minutes data MN can be made substantially uniform.
[0091] Also, in this embodiment, when the user terminal 110 transmits and stores at least one of the speech-to-text data TX and the minutes data MN to the first server 120, the user browsing restriction for at least one of the speech-to-text data TX and the minutes data MN to be stored can be set freely. Then, the user terminal 110 transmits the user browsing restriction information to the first server 120 together with at least one of the speech-to-text data TX and the minutes data MN to be stored. Then, the first server 120 stores at least one of the received speech-to-text data TX and the minutes data MN.
[0092] Furthermore, in the browser 111 of the user terminal 110, at least one of the text data TX with speech recognition and the minutes data MN is viewable within the range of user viewing restrictions based on the user viewing restriction information stored in the first server 120. More specifically, as shown in FIG. 17, when the server save button 113n of the first minutes and speech recognition window 113 is operated, the browser 111 displays a server save confirmation window 118.
[0093] In the server save confirmation window 118, as an example, a target selection 118a, a minutes type selection 118b, and a sharing range setting 118c as user viewing restriction information are displayed so as to be selectable. Further, an "address book reference" button 118d, a sharer input field 118e, a "preview / editing" button 118f, and a "save to server" button 118g are provided. Among these, since the target selection 118a and the minutes type selection 118b are the same as those in the download confirmation window 114, their descriptions are omitted.
[0094] On the other hand, for the sharing range setting 118c, any one of all employees, within the department, within the section, within the group, and member designation is provided so as to be selectable. Among these, when member designation is selected and the "address book reference" button 118d is operated, the browser 111 displays the member information of the pre-registered organizational address book based on the organizational address book in the database 121 of the first server 120. When the user selects a member to share from among them, the name information of the selected member is provided to be input into the sharer input field 118e. Then, when the "preview / editing" button 118f is operated, similar to when the "preview / editing" button 114d of the download confirmation window 114 described above is operated, the browser 111 displays a preview display screen PV in which the text data TX with speech recognition and the minutes data MN can be freely modified. At this time, on the preview display screen PV, instead of the download button PV1, a save button (upload button) (not shown) to the server is displayed.
[0095] Also, when the "Save to Server" button 118g is operated, the user terminal 110 transmits (uploads) the text data TX with speech recognition and the minutes data MN, along with the user viewing restriction information, to the first server 120. Then, the first server 120 saves the received text data TX with speech recognition and the minutes data MN to the database 121. Furthermore, as shown in FIG. 18, when the user terminal 110 accesses the first server 120, the browser 111 displays the minutes list screen 119. On the minutes list screen 119, a search condition specification field 119a and a "Search" button 119b are provided. And on the minutes list screen 119 of the browser 111 of the user terminal 110, at least one of the text data TX with speech recognition and the minutes data MN is viewable within the range of the user viewing restriction based on the user viewing restriction information saved in the first server 120.
[0096] Thereby, user viewing restrictions are appropriately set for the text data TX with speech recognition and the minutes data MN of the online meeting and are publicly disclosed within the set range. As a result, within the set range, the user can share information about the content of the online meeting. Note that instead of the selection operation of the server save button 113n, which is an example of the save operation to the server, or the selection operation of the "Save to Server" button 118g, an operation of displaying a menu by right-clicking and selecting an item may be used, or a shortcut key operation by keyboard input may also be used.
[0097] Also, regarding the storage destinations of the transcribed text data TX and the minutes data MN displayed in the first minutes transcription window 113 and the second minutes transcription window 116 of the user terminal 110, as described above, it may be set to be stored based on a storage operation by the user to the server, or may be freely set by the administrator on the administrator screen 150. For example, for all online meetings without any room for user selection, the transcribed text data TX and the minutes data MN may be stored in the database 121 of the first server 120. In this case, for the transcribed text data TX and the minutes data MN of each online meeting, the administrator sets user viewing restriction information. For example, the user viewing restriction may be set in advance on the administrator screen 150, such as sharing within the section to which the users participating in the online meeting belong. Also, when the transcribed text data TX and the minutes data MN of the online meeting are forcibly stored in the database 121 of the first server 120 for the online meeting, the transcribed text data TX and the minutes data MN at the end of the online meeting are stored.
[0098] The online meeting system 100, which is an embodiment of the present invention obtained in this way, includes a user terminal 110, a first server 120 and a second server 130 as servers. The browser 111 of the user terminal 110, or the online meeting screen 112 displayed on the browser 111, has a first minutes transcription button 112a and a second minutes transcription button 111a, which are minutes transcription buttons. When the first minutes transcription button 112a and the second minutes transcription button 111a, which are minutes transcription buttons, are operated, the browser 111 displays a first minutes transcription window 113 and a second minutes transcription window 116 that can be switched freely between the minutes content and the transcribed content. When the user's voice is input via the microphone of any user terminal 110, the browser 111 converts the input voice data into text data TX, and displays the converted text data TX in the first minutes transcription window 113 and the second minutes transcription window 116 in chronological order together with the user identification information US of the user terminal 110 where the voice input occurred. In this way, it is possible to easily record who said what and what kind of interaction took place, and even when someone speaks simultaneously, the user can easily confirm who said what.
[0099] Furthermore, when there is a minutes creation operation in the browser 111 or the minutes transcription window (113, 116), or every predetermined time, the browser 111 uses ChatGPT, which is an example of the pre-trained artificial intelligence means 131 of the second server 130, to generate minutes data MN including participant information PT, online meeting date information DT, start time / end time information SE, summary information OP of opinions by participant, summary information PP of current problems, and future issue information AS based on the transcribed text data TX, and displays it in the first minutes transcription window 113 and the second minutes transcription window 116. In this way, the user does not need to summarize the minutes by themselves and can easily obtain the minutes data MN.
[0100] Also, when a download operation is performed in the browser 111 or the minutes / voice recognition window (113, 116), the browser 111 is configured to download at least one of the minutes data MN and the voice recognition text data TX displayed in the first minutes / voice recognition window 113 and the second minutes / voice recognition window 116 to the user terminal 110 that has performed the download operation in a freely selectable manner and with a freely selectable file format. As a result, the user can easily save at least one of the minutes and the voice recognition as a record.
[0101] Furthermore, when a screen capture operation is performed in the browser 111 or the minutes / voice recognition window (113, 116), the browser 111 is configured to insert and display the screen capture data CP1 (CP2) in the time series of the voice recognition text data TX in the first minutes / voice recognition window 113 and the second minutes / voice recognition window 116. As a result, the user can easily understand who specifically spoke about what when looking back later, and it is possible to avoid a time lag between the capture and the insertion and a deviation between the speech and the screen capture.
[0102] Also, since the voice recognition text data TX can be modified in the browser 111 or the first minutes / voice recognition window 113 and the second minutes / voice recognition window 116, the user can obtain accurate minutes data MN. Furthermore, since the minutes data MN can be modified in the browser 111 or the first minutes / voice recognition window 113 and the second minutes / voice recognition window 116, the user can obtain accurate minutes data MN.
[0103] In addition, the browser 111 has a dictionary registration function for changing the pre-change text to the post-change text. When there is a location in the text data TX that is the same as the pre-change text registered in the text data TX when converting from audio data to text data TX, the corresponding location is changed to the post-change text and displayed in the first minutes transcription window 113 and the second minutes transcription window 116. With this configuration, the user can obtain the transcribed text data TX of the intended display, and further, the user can obtain the minutes data MN of the intended display.
[0104] Furthermore, the browser 111 has an inappropriate content check function for determining whether the text data TX contains inappropriate words. When there is a location in the text data TX that is the same as a predetermined word registered in advance in the text data TX when converting from audio data to text data TX, the corresponding location is highlighted and displayed in the first minutes transcription window 113 and the second minutes transcription window 116 with the highlight display UL3. When the speaker is the user himself / herself of the user terminal 110, the browser 111 is configured to display that there is an inappropriate statement in the user's speech content. For example, when the text data TX contains discriminatory terms, the user can be cautious about the communication. Further, when the user himself / herself makes an inappropriate statement, the user can immediately retract the inappropriate statement or take measures such as apologizing for the inappropriate statement.
[0105] Also, when a video capture operation is performed in the browser 111 or the transcript / speech-to-text window (113, 116), the browser 111 inserts and displays the video capture data MV in the time series of the speech-to-text data TX of the first transcript / speech-to-text window 113 and the second transcript / speech-to-text window 116, and based on the video shooting start time information and the video shooting end time information, displays the location of the speech-to-text data TX corresponding to the video capture as the video capture corresponding location. With this configuration, the user can easily understand which part of the conversation from where to where is the video capture when looking back later. Furthermore, there is no time lag between capture and insertion, and it is possible to avoid the deviation between the speech and the video.
[0106] Furthermore, on the administrator screen 150 of the administrator terminal that can communicate with the first server 120, masking for the word attributes included in the text data TX to ChatGPT, which is an example of the pre-trained artificial intelligence means 131, can be set from the browser 111 of the user terminal 110. When the voice data input in the browser 111 of the user terminal 110 is converted into text data TX including the character WD1 of the pre-set masking word attribute and transmitted to the first server 120, the received first server 120 determines that the received text data TX includes the character WD1 of the masking word attribute, and among the received data, after converting the character WD1 of the masking word attribute into another arbitrary character WD2 of the same word attribute, it uses ChatGPT, which is an example of the pre-trained artificial intelligence means 131, to execute the generation of the minutes data MN based on the converted text data TX. Among the obtained minutes data MN, after changing the part of the other arbitrary character WD2 converted during the transmission to ChatGPT, which is an example of the pre-trained artificial intelligence means 131, back to the character WD1 before conversion, the changed minutes data MN is transmitted to the user terminal 110. The user terminal 110 is configured to display the received changed minutes data MN in the first minutes and transcription window 113 and the second minutes and transcription window 116. In this way, it is possible to avoid ChatGPT, which is an example of the pre-trained artificial intelligence means 131, from remembering the information of the character WD1 of the word attribute that can identify an individual, and it is possible to avoid the leakage of personally identifiable information to other terminals via ChatGPT, which is an example of the pre-trained artificial intelligence means 131. It is possible to obtain the changed minutes data MN that basically follows the intention of the text data TX exchange converted by the browser 111, and avoid the involuntary information leakage to ChatGPT, which is an example of the pre-trained artificial intelligence means 131, and the involuntary information leakage to other terminals via ChatGPT, which is an example of the pre-trained artificial intelligence means 131. At the same time, it is possible to ensure the consistency between the content of the text data TX of the transcription from the user terminal 110 and the content of the changed minutes data MN transmitted to the user terminal 110.
[0107] Also, in the administrator screen 150 of the administrator terminal that can communicate with the first server 120, since the dictionary registration function for changing the text before change to the text after change can be set freely, each user can obtain the character recognition text data TX of the display intended by the administrator, and further, can obtain the minutes data MN of the display intended by the user. Among each user, the variation in the accuracy of the character recognition text data TX and the minutes data MN can be made substantially uniform.
[0108] Furthermore, when the user terminal 110 saves at least one of the character recognition text data TX and the minutes data MN to the first server 120, the user viewing restriction for at least one of the character recognition text data TX and the minutes data MN to be saved can be set freely, and the user viewing restriction information is transmitted to the first server 120 together with at least one of the character recognition text data TX and the minutes data MN to be saved. The first server 120 saves at least one of the received character recognition text data TX and the minutes data MN, and in the browser 111 of the user terminal 110, at least one of the character recognition text data TX and the minutes data MN is viewable within the range of the user viewing restriction based on the user viewing restriction information saved in the first server 120. Thus, within the set range, the user can share information about the content of the online meeting.
[0109] Also, the program of the online meeting system 100 which is an embodiment of the present invention includes button operation determination steps S1 and S2 for determining whether the first minutes transcription button 112a and the second minutes transcription button 111a, which are the minutes transcription buttons provided on the browser 111 of the user terminal 110 or on the online meeting screen 112 displayed on the browser 111, are operated; a window display step S3 in which when the first minutes transcription button 112a and the second minutes transcription button 111a, which are the minutes transcription buttons, are operated, the browser 111 displays the first minutes transcription window 113 and the second minutes transcription window 116 that can freely switch between the minutes content and the transcription content; and a transcription text display step S4 to S6 in which when the voice of a user is input via the microphone of any user terminal 110, the browser 111 converts the input voice data into text data TX and displays the converted text data TX in chronological order in the first minutes transcription window 113 and the second minutes transcription window 116 together with the user identification information US of the user terminal 110 where the voice input occurred. By having these steps, it is possible to easily record who said what and what kind of exchanges took place. Even when someone speaks simultaneously, the user can easily confirm who said what, and the effect is very significant.
Explanation of Signs
[0110] 100 ··· Online meeting system 110 ··· User terminal 111 ··· Browser 111a ··· Second minutes transcription button 112 ··· Online meeting screen 112a ··· First minutes transcription button 113 ··· First minutes transcription window 113a ··· Meeting name input field 113b ··· "Minutes" tab 113c ··· "Transcription" tab 113d ··· Transcription start button 113e ··· Transcription pause button 113f ··· Transcription end button 113g ··· Reset button 113h ··· Minutes creation button 113k ··· Minutes file creation button (for the first minutes · transcription window) 113m ··· Download button (for the first minutes · transcription window) 113n ··· Server save button 113p ··· Other button 113q ··· Clipboard button 114 ··· Download confirmation window 114a ··· Target selection (for the download confirmation window) 114b ··· File format selection 114c ··· Minutes type selection (for the download confirmation window) 114d ··· "Preview · Edit" button (for the download confirmation window) 114e ··· "Download" button 115 ··· Menu item window 115a ··· "Screen capture" item 115b ··· "Video capture" item 115c ··· "Dictionary" item 115d ··· "Inappropriate speech" item 116 ··· Second minutes · transcription window 116a ··· Minutes file creation button (for the second minutes · transcription window) 116b ··· Download button (for the second minutes · transcription window) 116c ··· Dictionary button (for the second minutes · transcription window) 116d ··· Screen capture button (for the second minutes · transcription window) 116e ··· Video capture button (of the second meeting minutes transcription window) 117A ··· Dictionary registration window 117B ··· Inappropriate speech registration window 118 ··· Server save confirmation window 118a ··· Target selection (of the server save confirmation window) 118b ··· Minutes type selection (of the server save confirmation window) 118c ··· Sharing range setting (user viewing restriction information) 118d ··· "Address book reference" button 118e ··· Shared user input field 118f ··· "Preview / Edit" button (of the server save confirmation window) 118g ··· "Save to server" button 119 ··· Minutes list screen 119a ··· Search condition specification field 119b ··· "Search" button 120 ··· First server 121 ··· Database 130 ··· Second server 131 ··· Pre - learned artificial intelligence means (ChatGPT) 140 ··· Other user terminals 150 ··· Administrator screen (of the administrator terminal) 151 ··· Masking setting PV ··· Preview display screen PV1 ··· Download button (of the preview display screen) DC1 ··· Word (to be converted) DC2 ··· Conversion candidates UL1 ··· Word (to be considered inappropriate) UL2 ··· Attribute (to be considered inappropriate) UL3 ··· Highlight display UL4 ··· Warning display window TX ··· Text data (of the transcription) US ··· User identification information (of the user terminal with voice input) MN ··· Minutes data PT ··· Participant information DT ··· Online meeting date information SE ··· Start time · End time information OP ··· Summary information of opinions by participant PP ··· Summary information of current problems AS ··· Future issue information CP1 ··· Screen capture data (of the online meeting screen) ST ··· Any other arbitrary screen (other than the online meeting screen) CP2 ··· Screen capture data (of the other arbitrary screen) MV ··· Video capture data RG ··· Video recording range display (corresponding part of the video capture) WD1 ··· Word attribute character of masking (character before conversion) WD2 ··· Another arbitrary character of the same word attribute (another arbitrary character after conversion)
Claims
1. An online meeting system comprising a user terminal and a server, wherein the user terminal communicates with other user terminals via the server and conducts an online meeting on the browser of the user terminal, the browser of the user terminal, or the online meeting screen displayed on the browser, has a minutes / transcription button, when the minutes / transcription button is operated, the browser displays a minutes / transcription window that can freely switch between displaying minutes content and transcription content, When user voice is input via the microphone of any user terminal, the browser converts the input voice data into text data and displays the converted text data in chronological order in the minutes / transcription window together with the user identification information of the user terminal where the voice input occurred. An online meeting system characterized by this configuration.
2. When there is a minutes creation operation in the browser or the minutes / transcription window, or every predetermined time, the browser uses the pre-trained artificial intelligence means of the server to generate minutes data including participant information, online meeting date information, start time / end time information, summary information of opinions by participant, summary information of current problems, and future issue information based on the text data of the transcription, and displays it in the minutes / transcription window. The online meeting system according to claim 1, characterized by this configuration.
3. When there is a download operation in the browser or the minutes / transcription window, the browser freely selects at least one of the minutes data displayed in the minutes / transcription window and the text data of the transcription, and freely selects the file format to download to the user terminal where the download operation occurred. The online meeting system according to claim 2, characterized by this configuration.
4. When there is a screen capture operation in the browser or the minutes / voice recognition window, the browser inserts the screen capture data into the time series of the text data of the voice recognition in the minutes / voice recognition window and displays it. The online meeting system according to claim 3, characterized in that it has such a configuration.
5. The online meeting system according to any one of claims 1 to 4, characterized in that the text data of the voice recognition is configured to be modifiable in the browser or the minutes / voice recognition window.
6. The online meeting system according to any one of claims 2 to 4, characterized in that the minutes data is configured to be modifiable in the browser or the minutes / voice recognition window.
7. The browser has a dictionary registration function for changing the pre-change text to the post-change text. When there is the same location as the pre-change text registered in the text data when converting from voice data to text data, the corresponding location is changed to the post-change text and displayed in the minutes / voice recognition window. The online meeting system according to any one of claims 1 to 4, characterized in that it has such a configuration.
8. The browser has an inappropriate content check function for determining whether the text data contains a predetermined word and is inappropriate. When there is the same location as the predetermined word registered in advance in the text data when converting from voice data to text data, the corresponding location is highlighted and displayed in the minutes / voice recognition window. When the speaker is the user himself / herself of the user terminal, the browser is configured to display that there is an inappropriate statement in the user's speech content. The online meeting system according to any one of claims 1 to 4, characterized in that it has such a configuration.
9. When there is a video capture operation in the browser or the minutes / voice transcription window, the browser inserts the video capture data into the time series of the text data of the voice transcription in the minutes / voice transcription window and displays it, and based on the video shooting start time information and the video shooting end time information, the location of the text data of the voice transcription corresponding to the video capture is displayed as the corresponding location of the video capture. The online meeting system according to any one of claims 1 to 4, characterized in that it has such a configuration.
10. On the administrator screen of the administrator terminal capable of communicating with the server, masking of the word attributes included in the text data from the browser of the user terminal to the pre-trained artificial intelligence means can be set freely. When the voice data input in the browser of the user terminal is converted into text data including the characters of the masking word attributes set in advance and transmitted to the server, the receiving server determines that the received text data contains the characters of the masking word attributes, and among the received data, after converting the characters of the masking word attributes into any other arbitrary characters of the same word attribute, it uses the pre-trained artificial intelligence means to execute the generation of minutes data based on the converted text data, and among the obtained minutes data, it makes a change to return the part of the other arbitrary characters converted during the transmission to the pre-trained artificial intelligence means to the characters before conversion, and then transmits the changed minutes data to the user terminal. The online meeting system according to any one of claims 1 to 4, characterized in that the user terminal is configured to display the received changed minutes data in the minutes / voice transcription window.
11. On the administrator screen of the administrator terminal capable of communicating with the server, the dictionary registration function for changing the text before change to the text after change can be set freely. The online meeting system according to any one of claims 1 to 4, characterized in that it has such a configuration.
12. When the user terminal stores at least one of the text data of speech recognition and the minutes data in the server, user viewing restrictions regarding at least one of the text data of speech recognition and the minutes data to be stored can be set freely, and user viewing restriction information is transmitted to the server together with at least one of the text data of speech recognition and the minutes data to be stored. The server stores at least one of the received text data of speech recognition and the minutes data. The online meeting system according to any one of claims 1 to 4, characterized in that at least one of the text data of speech recognition and the minutes data can be viewed within the range of user viewing restrictions based on the user viewing restriction information stored in the server in the browser of the user terminal.
13. A program for an online meeting system in which a user terminal communicates with other user terminals via a server and holds an online meeting in the browser of the user terminal, A button operation determination step for determining whether or not a minutes / speech recognition button provided on the browser of the user terminal or on the online meeting screen displayed on the browser has been operated; A window display step in which when the minutes / speech recognition button is operated, the browser displays a minutes / speech recognition window that can freely switch between the minutes content and the speech recognition content; When the voice of a user is input via the microphone of any user terminal, the browser converts the input voice data into text data, and displays the converted text data in chronological order in the minutes / speech recognition window together with the user identification information of the user terminal where the voice input occurred. A program for an online meeting system, characterized by having a speech recognition text display step.
Citation Information
Patent Citations
Expansion absorbing type decorative joint board made of synthetic rubber
JP1988093955A