Web conference server and web conference system

The web conference server optimizes server selection for speech recognition and translation by considering language attributes, communication speed, and processing speed, improving the speed and accuracy of multi-user conferences.

JP7772359B2Active Publication Date: 2025-11-18ASIASTAR CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2021158558
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-29
Publication Date
2025-11-18
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

Existing web conferencing systems face challenges in improving the speed and accuracy of speech recognition and translation during multi-user conferences, particularly in determining the optimal speech and translation servers based on communication and processing speeds.

Method used

A web conference server dynamically selects and combines specific speech recognition and translation servers for each user based on language attributes, communication speed, and information processing speed, ensuring optimal performance and stability throughout the conference.

Benefits of technology

This approach enhances the speed and accuracy of web conferences by distributing processing and ensuring real-time, stable communication and information handling, optimizing server selection for each user.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007772359000001
    Figure 0007772359000001
  • Figure 0007772359000002
    Figure 0007772359000002
  • Figure 0007772359000003
    Figure 0007772359000003
Patent Text Reader

Abstract

To improve speed and accuracy of a web conference.SOLUTION: A web conference server includes: a speech recognition server speed determination unit that determines communication speed between the web conference server and a plurality of speech recognition servers and information processing speed of the plurality of speech recognition servers; a speech recognition server determination unit that determines a specific speech recognition server to be used for a specific user from among the plurality of speech recognition servers based on a language attribute, the communication speed with the plurality of speech recognition servers, and the information processing speed of the plurality of speech recognition servers; a translation server speed determination unit that determines communication speed between the specific speech recognition server to be used for the specific user and a plurality of translation servers and information processing speed of the plurality of translation servers; and a translation server determination unit that determines a specific translation server to be used for the specific user from among the plurality of translation servers based on the communication speed with the plurality of translation servers and the information processing speed of the plurality of translation servers.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a web conference server that provides a web conference service, a web conference method, and a web conference system. [Background technology]

[0002] Web conferencing systems are known in which multiple users in remote locations hold online conferences using individual user terminals connected via a network. In recent years, some web conferencing systems have been developed that not only allow users to talk over video feeds while viewing the video, but also simultaneously display text obtained by speech recognition of speech data on the screen. As web conferencing between remote locations becomes increasingly common due to the COVID-19 pandemic, there is a demand for technology that can translate speech recognition data into the languages ​​of the users in the conference and simultaneously display the translated text on the screen. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-153099 [Patent Document 2] Japanese Patent Application Publication No. 2019-061594 [Patent Document 3] Japanese Patent Application Laid-Open No. 2017-215931 [Patent Document 4] Patent No. 6795668 Summary of the Invention [Problem to be solved by the invention]

[0004] In Patent Document 1, when a video conferencing system is used, the device (location) that will perform the text conversion process is determined based on the communication speed and processing capacity of the device. In Patent Document 2, the user can arbitrarily specify the translation language, or the language is automatically determined based on the attributes of the attendees of the target conference or the users who have permission to view the minutes. In Patent Document 3, the input language and translation language are set based on user operation. In Patent Document 4, when conference participants select multiple translation dictionaries, they can set priorities for each translation dictionary, and the translation process is performed in the order of the set priorities.

[0005] In view of the above circumstances, an object of the present invention is to improve the speed and accuracy of web conferences. [Means for solving the problem]

[0006] A web conference server according to one aspect of the present invention includes: a voice acquisition unit that acquires voice data of a specific user among a plurality of users participating in a web conference from a user terminal of the specific user; a speech recognition unit that recognizes the speech data and determines a language attribute of the specific user; a speech recognition server speed determination unit that determines a communication speed between the web conference server and a plurality of speech recognition servers and an information processing speed of the plurality of speech recognition servers; a speech recognition server determination unit that determines a specific speech recognition server to be used for the specific user from the plurality of speech recognition servers based on the language attribute, a communication speed with the plurality of speech recognition servers, and an information processing speed of the plurality of speech recognition servers; a translation server speed determination unit that determines the communication speed between the specific speech recognition server used for the specific user and a plurality of translation servers, and the information processing speed of the plurality of translation servers; a translation server determination unit that determines a specific translation server to be used for the specific user from the plurality of translation servers based on a communication speed with the plurality of translation servers and an information processing speed of the plurality of translation servers; a voice data processing request unit that supplies the voice data of the specific user and identification information that identifies the specific translation server to the specific voice recognition server, thereby causing the specific voice recognition server to recognize the voice data of the specific user and generate voice recognition data, which is text data, and causing the specific translation server to translate the voice recognition data of the specific user and generate translation data, which is text data; Equipped with The combination of the specific speech recognition server and the specific translation server varies for each of the multiple users participating in the web conference. [Effects of the Invention]

[0007] According to the present invention, the speed and accuracy of web conferences can be improved. [Brief explanation of the drawings]

[0008] [Figure 1] 1 illustrates a web conferencing system according to one embodiment of the present invention. [Figure 2] 1 shows the functional configuration of a web conferencing system. [Figure 3] 10 shows a first operation flow of the web conference server. [Figure 4] 10 shows a second operation flow of the web conference server. [Figure 5] 10 shows a third operation flow of the web conference server. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0010] 1. Web conferencing system overview

[0011] FIG. 1 shows a web conferencing system according to an embodiment of the present invention.

[0012] The web conference system 1 includes multiple web conference servers 10, multiple speech recognition servers 20, and multiple translation servers 30. The multiple web conference servers 10, multiple speech recognition servers 20, and multiple translation servers 30 are connected to each other via a network N such as the Internet.

[0013] Multiple web conference servers 10 are installed in multiple different countries and regions. The multiple web conference servers 10 are computers that communicate with multiple user terminals 40 (personal computers, smartphones, tablet computers, wearable devices, etc.) used by multiple users participating in a web conference via a network N and provide web conference services to the multiple users. The web conference servers 10 accessed by the multiple user terminals 40 are different for each user terminal 40, even if the multiple users participate in the same web conference. Each web conference server 10 determines a specific speech recognition server 20 and a specific translation server 30 to use for each user. Therefore, the combination of a specific speech recognition server 20 and a specific translation server 30 is different for each user participating in the web conference, even if the multiple users participate in the same web conference. Each web conference server 10 acquires voice data of a specific user participating in the web conference from the user terminal 40 used by that user and provides the voice data to one of the voice recognition servers 20.

[0014] The multiple speech recognition servers 20 are installed in multiple different countries or regions. The multiple speech recognition servers 20 are typically provided by different providers and are computers running different speech recognition software. Each speech recognition server 20 acquires voice data of a specific user participating in the web conference from one of the web conference servers 10, performs speech recognition on the voice data to generate speech recognition data, which is text data, and supplies the speech recognition data to the web conference server 10 and a specific translation server 30 that are the source of the voice data.

[0015] The multiple translation servers 30 are installed in multiple different countries or regions. The multiple translation servers 30 are typically provided by different providers and are computers that run different translation software. Each translation server 30 acquires speech recognition data of a specific user participating in the web conference from one of the speech recognition servers 20, translates the speech recognition data to generate translation data, which is text data, and provides the translation data to the web conference server 10, which is the source of the speech data.

[0016] The web conference server 10 outputs speech recognition data and translation data obtained from the speech data of multiple users participating in the web conference in real time during the web conference to multiple user terminals 40 of the multiple users participating in the web conference. After the web conference ends, the web conference server 10 also creates minutes data of the speech recognition data and translation data and outputs the minutes data to multiple user terminals 40 of the multiple users participating in the web conference.

[0017] 2. Functional configuration of the web conferencing system

[0018] FIG. 2 shows the functional configuration of the web conference system.

[0019] The web conference server 10 operates as a voice acquisition unit 101, a voice recognition unit 102, a voice recognition server speed determination unit 103, a voice recognition server determination unit 104, a translation server speed determination unit 105, a content determination unit 106, an emotion determination unit 107, a translation server determination unit 108, a voice data processing request unit 109, a processing data acquisition unit 110, a context check unit 111, a real-time output unit 112, and a minutes creation unit 113 by having the CPU load an information processing program recorded in the ROM into the RAM and execute it. The web conference server 10 has a non-volatile or volatile storage device 120.

[0020] The user terminal 40 operates as a voice input unit 401, a real-time input unit 402, and a minutes acquisition unit 403 by the CPU loading an information processing program recorded in the ROM into the RAM and executing it. The user terminal 40 has an external or built-in microphone 411, an external or built-in display 412, and a non-volatile or volatile storage device 413.

[0021] 3. Web conferencing system operation flow

[0022] FIG. 3 shows a first operational flow of the web conference server (from the start to the end of the web conference).

[0023] First, at the start of a web conference, each of the user terminals 40 of the multiple users participating in the web conference selects a specific (one) web conference server 10 with the fastest communication speed from the multiple web conference servers 10 based on the IP address of the user terminal 40, accesses the selected web conference server 10, and requests to sign in to the web conference. Typically, each user terminal 40 selects a web conference server 10 with an IP address that identifies the country or region closest to the country or region identified by the IP address of the user terminal 40. Therefore, the specific web conference server 10 accessed by each user terminal 40 is different for each of the multiple users participating in the web conference. This allows each of the multiple user terminals 40 of the multiple users to use the web conference server 10 that is optimal in terms of speed, and also distributes processing throughout the web conference, ensuring stable communication and information processing.

[0024] The web conference server 10 accepts access and a sign-in request to the web conference from a user terminal 40 used by a specific user among multiple users participating in the web conference, and allows the user to sign in to the web conference. Unless otherwise specified, the following description will refer to one web conference server 10 and one user terminal 40 with which communication has been established, and the user using this user terminal 40. Furthermore, the uniquely determined web conference server 10, speech recognition server 20, translation server 30, and user terminal 40 may be referred to as a specific web conference server 10, specific speech recognition server 20, specific translation server 30, and specific user terminal 40.

[0025] When the web conference starts, the voice input unit 401 of the user terminal 40 starts to supply the user's voice data input by the user via the microphone 411 to one of the web conference servers 10 with which communication has been established.

[0026] The voice acquisition unit 101 of the web conference server 10 acquires the voice data of the user from the user terminal 40 (step S101).

[0027] The speech recognition unit 102 of the web conference server 10 recognizes the speech data acquired from the user terminal 40 and determines the language attributes of the specific user (step S102). The language attributes include, for example, language (e.g., English), dialect (e.g., Australian English), accent (e.g., Dutch accent), etc.

[0028] The speech recognition server speed determination unit 103 of the web conference server 10 determines the communication speed between the web conference server 10 and the plurality of speech recognition servers 20 and the information processing speed of the plurality of speech recognition servers 20 (step S103). The communication speed and the information processing speed each include not only a reference value but also a real-time speed.

[0029] The speech recognition server determination unit 104 of the web conference server 10 determines a specific speech recognition server 20 to be used for a specific user from the multiple speech recognition servers 20 based on the language attributes determined by the speech recognition unit 102, the communication speed between the web conference server 10 and the multiple speech recognition servers 20 determined by the speech recognition server speed determination unit 103, and the information processing speed of the multiple speech recognition servers 20 (step S104).

[0030] As one example, the speech recognition server determination unit 104 predefines one speech recognition server 20 that has high language recognition accuracy for a specific language attribute and is close to the web conference server 10 (resulting in a fast communication speed). If the information processing speed of this speech recognition server 20 is equal to or higher than a predetermined threshold, the speech recognition server determination unit 104 selects this speech recognition server 20. As another example, the speech recognition server determination unit 104 predefines multiple candidate speech recognition servers 20 that have high language recognition accuracy for a specific language attribute. From these candidates, the speech recognition server determination unit 104 selects one speech recognition server 20 that has the fastest communication speed with the multiple speech recognition servers 20, and if the information processing speed of the selected speech recognition server 20 is equal to or higher than a predetermined threshold, the speech recognition server determination unit 104 selects this speech recognition server 20.

[0031] The translation server speed determination unit 105 of the web conference server 10 determines the communication speed between the specific speech recognition server 20 to be used for the specific user, determined by the speech recognition server determination unit 104, and the multiple translation servers 30, and the information processing speed of the multiple translation servers 30 (step S105). The communication speed and information processing speed each include not only a reference value but also a real-time speed.

[0032] The content determination unit 106 of the web conference server 10 determines the content of the web conference based on the results of speech recognition by the speech recognition unit 102 (step S106). The content determination unit 106 may determine the content of the web conference, for example, by inputting the results of speech recognition by the speech recognition unit 102 into a pre-created AI model. The content of the web conference may be the overall theme of the web conference, such as sports or chemistry.

[0033] The emotion determination unit 107 of the web conference server 10 determines the user's emotion (irritated, calm, etc.) based on the result of the voice recognition by the voice recognition unit 102 (step S107). The emotion determination unit 107 may determine the user's emotion, for example, by inputting the result of the voice recognition by the voice recognition unit 102 into a pre-created AI model.

[0034] The translation server determination unit 108 of the web conference server 10 determines a specific translation server 30 to be used for a specific user from the multiple translation servers 30 based on the communication speed between the specific speech recognition server 20 and the multiple translation servers 30 determined by the translation server speed determination unit 105 and the information processing speed of the multiple translation servers 30. The translation server determination unit 108 may determine a specific translation server 30 to be used for a specific user based further on the content of the web conference determined by the content determination unit 106 and / or the emotion of the specific user determined by the emotion determination unit 107 (step S108).

[0035] As one example, the translation server determination unit 108 predefines one translation server 30 that has high translation accuracy for specific content and / or specific emotions and is located close to the web conference server 10 (resulting in a fast communication speed). The translation server determination unit 108 selects this translation server 30 if the information processing speed of this translation server 30 is equal to or higher than a predetermined threshold. As another example, the translation server determination unit 108 predefines multiple translation server 30 candidates that have high translation accuracy for specific content and / or specific emotions. The translation server determination unit 108 selects one translation server 30 from these candidates that has the fastest communication speed with the multiple translation servers 30, and selects this translation server 30 if the information processing speed of the selected translation server 30 is equal to or higher than a predetermined threshold.

[0036] The voice data processing request unit 109 of the web conference server 10 supplies the user's voice data acquired by the voice acquisition unit 101 and identification information identifying the translation server 30 determined by the translation server determination unit 108 to the voice recognition server 20 determined by the voice recognition server determination unit 104, and requests processing (step S109). The combination of a specific voice recognition server 20 and a specific translation server 30 is different for each of the multiple users participating in the web conference. This allows each of the multiple user terminals 40 of the multiple users to use the voice recognition server 20 and translation server 30 that are optimal in terms of both speed and accuracy, and also enables stable communication and information processing throughout the web conference because processing is distributed.

[0037] The speech recognition server 20 acquires the speech data of a specific user and identification information (such as an IP address) that identifies the translation server 30 determined by the translation server determination unit 108 from the speech data processing request unit 109 of the web conference server 10. The speech recognition server 20 recognizes the speech and generates speech recognition data, which is text data, and supplies the speech recognition data to the web conference server 10, which is the source of the speech data, and to the translation server 30 identified by the identification information.

[0038] The translation server 30 acquires the speech recognition data and identification information (such as an IP address) that identifies the web conference server 10 that is the source of the speech data from the speech recognition server 20. The translation server 30 translates the speech recognition data to generate translation data, which is text data, and supplies the translation data to the web conference server 10 identified by the identification information.

[0039] The processing data acquisition unit 110 of the web conference server 10 acquires the speech recognition data of the specific user from the speech recognition server 20, and acquires the translation data of the specific user from the specific translation server 30 (step S110).

[0040] The context check unit 111 of the web conference server 10 synchronizes the corresponding speech recognition data and translation data, checks the context of the corresponding speech recognition data and translation data, and modifies the corresponding speech recognition data and / or translation data according to the check result (step S111). The context check unit 111 may check the context of the speech recognition data and translation data, for example, by inputting the speech recognition data and translation data into a pre-created AI model. The context check unit 111 stores the corresponding speech recognition data and translation data after modification according to the check result in the storage device 120.

[0041] The storage device 120 may be either a non-volatile or volatile storage device, as long as it stores the speech recognition data and translation data at least during the web conference and for a predetermined period after the conference has ended. Note that the data corrected based on the check results includes data that is left uncorrected because the check results show that correction is not necessary.

[0042] The real-time output unit 112 of the web conference server 10 outputs the corresponding speech recognition data and translation data, corrected according to the check results, in real time during the web conference to the multiple user terminals 40 of the multiple users participating in the web conference (step S112). That is, the real-time output unit 112 outputs the speech recognition data and translation data not only to the user terminal 40 that input the speech data that is the basis for the data, but also to the multiple user terminals 40 of all users participating in the web conference.

[0043] The real-time input units 402 of the multiple user terminals 40 of all users participating in the web conference obtain the corresponding speech recognition data and translation data, corrected according to the check results, in real time during the web conference from the real-time output unit 112 of the web conference server 10. The real-time input units 402 display the speech recognition data and translation data, which are text data, on the display 412 in real time during the web conference.

[0044] After the web conference ends, the minutes-taking unit 113 of the web conference server 10 reads the corresponding speech recognition data and translation data corrected according to the check results from the storage device 120, and creates minutes data for the web conference based on the read speech recognition data and translation data. Specifically, the minutes-taking unit 113 of one of the web conference servers 10 acquires the speech recognition data and translation data of all users from the storage devices 120 of multiple web conference servers 10 accessed by multiple user terminals 40 of multiple users participating in the web conference, and creates minutes data by chronologically arranging the acquired speech recognition data and translation data. The minutes-taking unit 113 outputs the created minutes data to the multiple user terminals 40 of the multiple users participating in the web conference.

[0045] The minutes acquisition unit 403 of each user terminal 40 acquires minutes data from the minutes creation unit 113 of the web conference server 10 and stores it in the storage device 413. The storage device 413 only needs to store the speech recognition data and translation data for at least a predetermined period after the end of the web conference, and may be either a non-volatile or volatile storage device.

[0046] FIG. 4 shows a second operation flow of the web conferencing server (during a web conference).

[0047] During a web conference, the speech recognition server speed determination unit 103 of the web conference server 10 periodically (in a loop) determines the communication speed between the web conference server 10 and a specific (communicating) speech recognition server 20 and the information processing speed of the specific (communicating) speech recognition server 20 (step S201). The communication speed and the information processing speed are both real-time speeds. The speech recognition server speed determination unit 103 determines whether the communication speed between the web conference server 10 and the specific speech recognition server 20 and / or the information processing speed of the specific speech recognition server 20 has changed below a threshold value during the web conference (step S202). The threshold value is, for example, a value that is not allowable in terms of speed for conducting a smooth web conference.

[0048] If the communication speed between the web conference server 10 and a specific speech recognition server 20 and / or the information processing speed of the specific speech recognition server 20 changes to less than the threshold during the web conference (step S202, YES), the speech recognition server speed determination unit 103 determines the communication speed between the web conference server 10 and multiple speech recognition servers 20 (multiple speech recognition servers 20 other than the speech recognition server 20 currently communicating) and the information processing speed of the multiple speech recognition servers 20 (step S203). The communication speed and the information processing speed are both real-time speeds.

[0049] The speech recognition server determination unit 104 of the web conference server 10 newly determines a specific speech recognition server 20 to be used for a specific user from the multiple speech recognition servers 20 based on the language attributes determined by the speech recognition unit 102 (step S102), the communication speed between the web conference server 10 and the multiple speech recognition servers 20 determined by the speech recognition server speed determination unit 103, and the information processing speed of the multiple speech recognition servers 20 (step S203) (step S204). The speech recognition server determination unit 104 may newly determine a specific speech recognition server 20 in the same manner as the example described in step S104. Note that the newly determined speech recognition server 20 may not be changed from the speech recognition server 20 currently in communication.

[0050] When the specific speech recognition server 20 used for a specific user is changed during the web conference (step S205, YES), the translation server speed determination unit 105 of the web conference server 10 determines the communication speed between the newly determined specific speech recognition server 20 and the multiple translation servers 30, and the information processing speed of the multiple translation servers 30 (step S206). The communication speed and the information processing speed are both real-time speeds.

[0051] The reason for determining the communication speed between the newly determined specific speech recognition server 20 and the multiple translation servers 30 and the information processing speed of the multiple translation servers 30 (step S206) is that the translation server 30 most suitable for the new speech recognition server 20 (step S205, YES) may be a translation server 30 other than the currently communicating translation server 30. By selecting the translation server 30 most suitable for cooperating with the new speech recognition server 20, it is possible to select in real time the combination of the speech recognition server 20 and the translation server 30 that is optimal overall in terms of both speed and accuracy.

[0052] The translation server determination unit 108 of the web conference server 10 newly determines a specific translation server 30 to be used for a specific user from the multiple translation servers 30 based on the communication speed between the newly determined specific speech recognition server 20 and the multiple translation servers 30 and the information processing speed of the multiple translation servers 30 (step S207). The translation server determination unit 108 may newly determine a specific translation server 30 using a method similar to the example described in step S108. Note that the newly determined translation server 30 may not be changed from the translation server 30 currently in communication.

[0053] If the specific translation server 30 used for a specific user is not changed during the web conference (step S208, NO), the voice data processing request unit 109 supplies the newly determined specific voice recognition server 20 with the voice data of the specific user and identification information for identifying the specific (communicating) translation server 30 (step S209). This allows the specific voice recognition server 20 used for a specific user to be changed during the web conference, making it possible to use the voice recognition server 20 that is optimal in terms of both speed and accuracy in real time, and also enabling stable communication and information processing throughout the web conference as processing is distributed.

[0054] On the other hand, if the specific translation server 30 used for a specific user is changed during the web conference (step S208, YES), the voice data processing request unit 109 of the web conference server 10 supplies the voice data of the specific user and identification information for identifying the newly determined specific translation server 30 to the newly determined specific speech recognition server 20 (step S210). This further changes the specific translation server 30 used for a specific user during the web conference, allowing the use of the optimal translation server 30 in terms of both speed and accuracy in real time, and also enabling stable communication and information processing throughout the web conference as processing is distributed.

[0055] FIG. 5 shows a third operation flow of the web conferencing server (during a web conference).

[0056] During a web conference, the translation server speed determination unit 105 of the web conference server 10 periodically (in a loop) determines the communication speed between a specific (currently communicating) speech recognition server 20 and a specific (currently communicating) translation server 30, and the information processing speed of the specific (currently communicating) translation server 30 (step S301). The communication speed and information processing speed are both real-time speeds. The translation server speed determination unit 105 determines whether the communication speed between the specific speech recognition server 20 and a specific translation server 30 and / or the information processing speed of the specific translation server 30 has fallen below a threshold value during the web conference (step S302). The threshold value is, for example, a value that is not acceptable in terms of speed for conducting a smooth web conference.

[0057] If the communication speed between a specific speech recognition server 20 and a specific translation server 30 and / or the information processing speed of the specific translation server 30 changes to below the threshold during the web conference (step S302, YES), the translation server speed determination unit 105 determines the communication speed between the specific (communicating) speech recognition server 20 and multiple translation servers 30 (multiple translation servers 30 other than the communicating translation server 30) and the information processing speed of the multiple translation servers 30 (step S303). The communication speed and information processing speed are both real-time speeds.

[0058] The translation server determination unit 108 of the web conference server 10 newly determines a specific translation server 30 to be used for a specific user from the multiple translation servers 30 based on the communication speed between the specific speech recognition server 20 and the multiple translation servers 30 determined by the translation server speed determination unit 105 and the information processing speed of the multiple translation servers 30 (step S303) (step S304). The translation server determination unit 108 may newly determine a specific translation server 30 using a method similar to the example described in step S108. Note that the newly determined translation server 30 may not be changed from the translation server 30 currently in communication.

[0059] When the specific translation server 30 used for a specific user is changed during the web conference (step S305, YES), the speech recognition server speed determination unit 103 of the web conference server 10 determines the communication speed between the web conference server 10 and the multiple speech recognition servers 20 and the information processing speed of the multiple speech recognition servers 20 (step S306). The communication speed and the information processing speed are both real-time speeds.

[0060] The reason for determining the speeds of multiple speech recognition servers 20 (step S306) is that if the communication speed between a specific speech recognition server 20 and a specific translation server 30 and / or the information processing speed of a specific translation server 30 falls below a threshold during the web conference (step S302, YES), there may be a problem with the specific speech recognition server 20, and in that case, it may be better to change the speech recognition server 20 being used. In this way, by selecting the optimal speech recognition server 20 to cooperate with the new translation server 30, it is possible to select in real time a combination of a speech recognition server 20 and a translation server 30 that is optimal overall in terms of both speed and accuracy.

[0061] The speech recognition server determination unit 104 of the web conference server 10 determines a new specific speech recognition server 20 to be used for a specific user from the multiple speech recognition servers 20 based on the language attributes determined by the speech recognition unit 102 (step S102) and the communication speed between the web conference server 10 and the multiple speech recognition servers 20 and the information processing speed of the multiple speech recognition servers 20 determined by the speech recognition server speed determination unit 103 (step S306) (step S307). The speech recognition server determination unit 104 may determine a new specific speech recognition server 20 in the same manner as the example described in step S104. Note that the newly determined speech recognition server 20 may not be changed from the currently communicating speech recognition server 20.

[0062] If the specific speech recognition server 20 used for a specific user is not changed during the web conference (step S308, NO), the voice data processing request unit 109 supplies the specific (communicating) speech recognition server 20 with the voice data of the specific user and identification information for identifying the newly determined translation server 30 (step S309). This allows the specific translation server 30 used for a specific user to be changed during the web conference, making it possible to use the optimal translation server 30 in real time in terms of both speed and accuracy, and also enabling stable communication and information processing throughout the web conference as processing is distributed.

[0063] On the other hand, if the specific speech recognition server 20 used for a specific user is changed during the web conference (step S308, YES), the voice data processing request unit 109 of the web conference server 10 supplies the newly determined specific speech recognition server 20 with the voice data of the specific user and identification information for identifying the newly determined specific translation server 30 (step S310). This further changes the specific speech recognition server 20 used for a specific user during the web conference, allowing the use of the optimal speech recognition server 20 in real time in terms of both speed and accuracy, and also enabling stable communication and information processing throughout the web conference as a whole by distributing processing.

[0064] Although the embodiments and modified examples of the present technology have been described above, the present technology is not limited to the above-described embodiments, and it goes without saying that various modifications can be made within the scope of the gist of the present technology. [Explanation of symbols]

[0065] Web conferencing system 1 Web Conferencing Server 10 Speech Recognition Server 20 Translation Server 30 User terminal 40

Claims

1. 1. A web conferencing server, comprising: a voice acquisition unit that acquires voice data of a specific user among a plurality of users participating in a web conference from a user terminal of the specific user; a speech recognition unit that recognizes the speech data and determines a language attribute of the specific user; a speech recognition server speed determination unit that determines a communication speed between the web conference server and a plurality of speech recognition servers and an information processing speed of the plurality of speech recognition servers; a speech recognition server determination unit that determines a specific speech recognition server to be used for the specific user from the plurality of speech recognition servers based on the language attribute, a communication speed with the plurality of speech recognition servers, and an information processing speed of the plurality of speech recognition servers; a translation server speed determination unit that determines the communication speed between the specific speech recognition server used for the specific user and a plurality of translation servers, and the information processing speed of the plurality of translation servers; a translation server determination unit that determines a specific translation server to be used for the specific user from the plurality of translation servers based on a communication speed with the plurality of translation servers and an information processing speed of the plurality of translation servers; a voice data processing request unit that supplies the voice data of the specific user and identification information that identifies the specific translation server to the specific voice recognition server, thereby causing the specific voice recognition server to recognize the voice data of the specific user and generate voice recognition data, which is text data, and causing the specific translation server to translate the voice recognition data of the specific user and generate translation data, which is text data; Equipped with The combination of the specific speech recognition server and the specific translation server is different for each of the multiple users participating in the web conference. Web conferencing server.

2. 10. The web conferencing server of claim 1, the speech recognition server speed determination unit periodically determines, during the web conference, a communication speed between the web conference server and the specific speech recognition server and an information processing speed of the specific speech recognition server; If the communication speed between the web conference server and the specific speech recognition server and / or the information processing speed of the specific speech recognition server changes below a threshold during the web conference, the speech recognition server speed determination unit determines a communication speed between the web conference server and the plurality of speech recognition servers and an information processing speed of the plurality of speech recognition servers; the speech recognition server determination unit newly determines a specific speech recognition server to be used for the specific user from the plurality of speech recognition servers; The voice data processing request unit supplies the newly determined specific voice recognition server with the voice data of the specific user and identification information identifying the specific translation server, thereby changing the specific voice recognition server used for the specific user during the web conference. Web conferencing server.

3. 3. The web conferencing server of claim 2, If the specific speech recognition server used for the specific user is changed during the web conference, the translation server speed determination unit determines the communication speed between the newly determined specific speech recognition server and a plurality of translation servers, and the information processing speeds of the plurality of translation servers; the translation server determination unit determines a new specific translation server to be used for the specific user from the plurality of translation servers based on the communication speed between the newly determined specific speech recognition server and the plurality of translation servers and the information processing speed of the plurality of translation servers; The voice data processing request unit supplies the newly determined specific voice recognition server with the voice data of the specific user and identification information for identifying the newly determined specific translation server, thereby changing the specific translation server used for the specific user during the web conference. Web conferencing server.

4. 4. A web conferencing server according to any one of claims 1 to 3, comprising: the translation server speed determination unit periodically determines the communication speed between the specific speech recognition server and the specific translation server and the information processing speed of the specific translation server during the web conference; If the communication speed between the specific speech recognition server and the specific translation server and / or the information processing speed of the specific translation server changes below a threshold during the web conference, the translation server speed determination unit determines the communication speed between the specific speech recognition server and the plurality of translation servers and the information processing speed of the plurality of translation servers; the translation server determination unit newly determines a specific translation server to be used for the specific user from the plurality of translation servers; The voice data processing request unit supplies the newly determined specific voice recognition server with the voice data of the specific user and identification information for identifying the newly determined specific translation server, thereby changing the specific translation server used for the specific user during the web conference. Web conferencing server.

5. 4. The web conferencing server of claim 3, If the specific translation server used for the specific user is changed during the web conference, the speech recognition server speed determination unit determines the communication speed between the newly determined specific translation server and a plurality of speech recognition servers, and the information processing speeds of the plurality of speech recognition servers; the speech recognition server determination unit determines a new specific speech recognition server to be used for the specific user from the plurality of speech recognition servers based on the communication speed between the newly determined specific translation server and the plurality of speech recognition servers and the information processing speed of the plurality of speech recognition servers; The voice data processing request unit changes the specific voice recognition server used for the specific user during the web conference by supplying the newly determined specific voice recognition server with the voice data of the specific user and identification information that identifies the newly determined specific translation server. Web conferencing server.

6. 6. A web conferencing server according to any one of claims 1 to 5, comprising: a processing data acquisition unit that acquires the speech recognition data of the specific user from the specific speech recognition server and acquires the translation data of the specific user from the specific translation server; a context check unit that checks the context of the corresponding speech recognition data and the corresponding translation data and corrects the corresponding speech recognition data and / or the corresponding translation data in accordance with the check result; a real-time output unit that outputs the corresponding speech recognition data and the translation data after correction according to the check result to the plurality of user terminals of the plurality of users participating in the web conference in real time during the web conference; a web conferencing server further comprising:

7. 7. The web conferencing server of claim 6, a minutes creation unit that creates minutes data of the web conference based on the corresponding speech recognition data and the translation data after correction according to a check result, and outputs the minutes data to a plurality of user terminals of a plurality of users participating in the web conference; a web conferencing server further comprising:

8. 8. A web conferencing server according to any one of claims 1 to 7, comprising: a content determination unit that determines the content of the web conference based on the results of the speech recognition by the speech recognition unit; an emotion determination unit that determines the emotion of the specific user based on the result of the voice recognition by the voice recognition unit; Further comprising: The translation server determination unit determines a specific translation server to be used for the specific user based on the content of the web conference and / or the emotion of the specific user. Web conferencing server.

9. 9. A web conferencing server according to any one of claims 1 to 8, comprising: The user terminal with which the web conference server communicates selects a specific web conference server with the fastest communication speed from among the plurality of web conference servers based on the IP address of the user terminal, and accesses the selected specific web conference server; The specific web conference server accessed by the user terminal is different for each of the user terminals of the multiple users participating in the web conference. Web conferencing server.

10. a web conferencing server connected to each other via a network; A plurality of speech recognition servers; Multiple translation servers, Equipped with The web conferencing server: a voice acquisition unit that acquires voice data of a specific user among a plurality of users participating in a web conference from a user terminal of the specific user; a speech recognition unit that recognizes the speech data and determines a language attribute of the specific user; a speech recognition server speed determination unit that determines a communication speed between the web conference server and the plurality of speech recognition servers and an information processing speed of the plurality of speech recognition servers; a speech recognition server determination unit that determines a specific speech recognition server to be used for the specific user from the plurality of speech recognition servers based on the language attribute, a communication speed with the plurality of speech recognition servers, and an information processing speed of the plurality of speech recognition servers; a translation server speed determination unit that determines a communication speed between the specific speech recognition server used for the specific user and the plurality of translation servers, and an information processing speed of the plurality of translation servers; a translation server determination unit that determines a specific translation server to be used for the specific user from the plurality of translation servers based on a communication speed with the plurality of translation servers and an information processing speed of the plurality of translation servers; a voice data processing request unit that supplies the voice data of the specific user and identification information that identifies the specific translation server to the specific voice recognition server, thereby causing the specific voice recognition server to recognize the voice data of the specific user and generate voice recognition data, which is text data, and causing the specific translation server to translate the voice recognition data of the specific user and generate translation data, which is text data; and The combination of the specific speech recognition server and the specific translation server is different for each of the multiple users participating in the web conference. Web conferencing system.

Citation Information

Patent Citations

  • Method, device and voice translation equipment for achieving speech-to-speech translation

    CN108319591A

  • Hybrid Offline / Online Voice Translation System

    JP2016527587A

  • Information processor and information processing program

    JP2017049685A

  • Conference support system, conference support device, conference support method, and program

    JP2017215931A

  • Conference support system and conference support program

    JP2019061594A