Processing device, processing method, and processing program
The processing device addresses the challenge of honorific language usage in virtual reality conferences by converting utterances to correct honorific language, enhancing interaction quality.
Patent Information
- Application Number
- JP2022043151
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-17
- Publication Date
- 2025-10-06
- Estimated Expiration
- 2042-03-17
AI Technical Summary
Young individuals often struggle with using appropriate honorific language during virtual reality conferences, leading to discomfort and uncertainty in interactions with older individuals.
A processing device and method that constructs a virtual reality conference space, performs voice recognition, determines the use of honorific language, and converts utterances to correct honorific language if necessary, ensuring proper language usage is maintained.
Ensures that remarks are conveyed using correct honorific language, alleviating the burden on users and facilitating smooth interactions in virtual reality conferences.
Smart Images

Figure 0007749499000001 
Figure 0007749499000002 
Figure 0007749499000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a processing device, a processing method, and a processing program. [Background technology]
[0002] In recent years, web conferencing services that connect via applications or browsers have become widespread. In these services, the equipment that provides the conferencing service is installed on the network, and users participate in the conference using applications or browsers running on their devices. Furthermore, technology has been proposed that creates a virtual reality (VR) space that represents a conference room and enables conferences to be held in the VR space. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-230532 Summary of the Invention [Problem to be solved by the invention]
[0004] Older people and others in older age groups tend to be strict about the use of honorific language. In such cases, young people may find it difficult to use honorific language properly and may feel unsure whether they are using it correctly. This can lead to young people feeling burdened when talking with people of a different age during meetings.
[0005] The present invention has been made in consideration of the above, and aims to provide a processing device, processing method, and processing program that enable a user to convey their remarks to others using correct honorific language during a meeting in a virtual reality space. [Means for solving the problem]
[0006] In order to solve the above-mentioned problems and achieve the object, the processing device of the present invention is characterized by having: a construction unit that constructs a virtual reality space in which a conference can be held and provides the space to a first user terminal used by a first user and a second user terminal used by a second user; a voice recognition unit that performs voice recognition on the input voice data when voice data is input from the first user terminal in response to a first utterance by the first user; a judgment unit that determines whether correct honorific language is used in the first utterance based on the voice recognition result by the voice recognition unit; a conversion unit that converts the first utterance into a second utterance using correct honorific language if the judgment unit determines that correct honorific language is not used in the first utterance; and a transmission unit that transmits data corresponding to the second utterance to the second user terminal. [Effects of the Invention]
[0007] According to the present invention, during a meeting in a virtual reality space, a user's remarks can be conveyed to the other party using correct honorific language. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram illustrating an example of a configuration of a communication system according to an embodiment. [Figure 2] FIG. 2 is a block diagram illustrating an example of the configuration of the server device illustrated in FIG. [Figure 3] FIG. 3 is a diagram illustrating an example of the data configuration of the user information. [Figure 4] FIG. 4 is a diagram illustrating an example of the data configuration of the user information. [Figure 5] FIG. 5 is a diagram illustrating a conference VR space that the server device provides to the user terminal. [Figure 6] FIG. 6 is a sequence diagram illustrating an example of a processing procedure of communication processing according to the first embodiment. [Figure 7] FIG. 7 is a sequence diagram illustrating an example of a processing procedure of a communication process according to a modification of the first embodiment. [Figure 8] FIG. 8 is a block diagram showing an example of the configuration of the server device according to the second embodiment. [Figure 9] FIG. 9 is a diagram illustrating an example of the data configuration of the scoring rule. [Figure 10] FIG. 10 is a diagram illustrating an example of the data configuration of the classification information. [Figure 11] FIG. 11 is a diagram illustrating a conference VR space that the server device shown in FIG. 8 provides to a user terminal. [Figure 12] FIG. 12 is a sequence diagram illustrating an example of a processing procedure of communication processing according to the second embodiment. [Figure 13] FIG. 13 is a sequence diagram illustrating an example of a processing procedure of a communication process according to a modification of the second embodiment. [Figure 14] FIG. 14 is a diagram illustrating a computer that executes a program. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of a processing device, a processing method, and a processing program according to the present application will be described in detail with reference to the accompanying drawings. However, the processing device and the processing method according to the present application are not limited to these embodiments.
[0010] [Embodiment 1] First, a description will be given of embodiment 1. In embodiment 1, a description will be given of a communication system that provides users with a VR space where a conference can be held (conference VR space).
[0011] In the communication system according to the first embodiment, the server device determines whether or not correct honorific language is used in a statement made by a user (e.g., user A), and if the correct honorific language is not used in the statement, converts the statement into one using correct honorific language and transmits it to another user (e.g., user B) participating in the conference. Therefore, even if user A is not good at using honorific language during a conference in a VR space, a statement using correct honorific language is transmitted to user B. For example, if user A uses honorific language to himself / herself, the server device converts the part using this honorific language into a part using humble language and transmits it to user B.
[0012] [Communication system configuration] The configuration of a communication system according to an embodiment will be described below: Fig. 1 is a block diagram showing an example of the configuration of a communication system according to an embodiment.
[0013] As shown in FIG. 1, the communication system according to the embodiment includes a user terminal 20A (first user terminal) used by user A (first user), a user terminal 20B (second user terminal) used by user B (second user), and a server device 10 (processing device) of a VR provider. Note that the configuration shown in FIG. 1 is merely an example, and the specific configuration and the number of devices are not particularly limited. Also, the user terminals 20A and 20B may be collectively referred to as user terminal 20. In the embodiment, an example will be described in which users A and B participate in a conference, but the number of users is not limited to two and may be three or more.
[0014] The server device 10 is a server device of a VR provider that provides users with a conference VR space (first virtual reality space) where a conference can be held. For example, the server device 10 constructs a conference VR space that represents a conference room in a virtual space, and places avatars of each user in the conference VR space. The server device 10 transmits voices uttered by other users to a user terminal 20 used by the user, causing the user terminal 20 to output the voices. The server device 10 also transmits texts created by other users participating in the conference to the user terminal 20, causing the user terminal 20 to output the texts. In this way, the server device 10 provides each user with a conference in a VR space via each user terminal 20.
[0015] The server device 10 determines whether or not correct honorific language is used in a statement made by a user (e.g., user A), and if the correct honorific language is not used in the statement, converts the statement into a statement using correct honorific language and transmits it to a user terminal 20 used by another user (e.g., user B).
[0016] The user terminal 20 is an information processing device such as a notebook PC (Personal Computer) or a desktop PC, or a smart device such as a tablet or a smartphone. The user terminal 20 connects to the server device 10 via the network N and is provided with a VR space created by the server device 10. For example, a user can experience the conference VR space by wearing VR goggles and operating the user terminal 20.
[0017] [Server device] Next, the server device 10 will be described. Fig. 2 is a block diagram showing an example of the configuration of the server device 10 shown in Fig. 1. As shown in Fig. 2, the server device 10 has a communication unit 11 (transmission unit) that controls communication related to various information, a memory unit 12 that stores data and programs required for various processes by a control unit 13, and the control unit 13 that executes various processes.
[0018] The communication unit 11 is a communication interface for transmitting and receiving various information to and from other devices connected via a network, etc. The communication unit 11 is realized by a NIC (Network Interface Card) or the like, and performs communication between other devices and a control unit 13 (described later) via electric communication lines such as a LAN (Local Area Network) or the Internet.
[0019] For example, the communication unit 11 receives a request to enter the conference VR space from the user terminal 20 via the network N. In addition, the communication unit 11 receives various data such as voice and text transmitted from the user terminal 20 via the network N and transmits it to other user terminals 20 so that voice and data can be communicated between the user terminals 20.
[0020] Furthermore, the communication unit 11 receives voice data corresponding to the user's utterance from the user terminal 20. The communication unit 11 transmits the voice data corresponding to the user's utterance (first utterance (described later)) or data corresponding to the utterance converted by the conversion unit 136 (second utterance (described later)) to another user terminal 20. The data corresponding to the utterance converted by the conversion unit 136 is text data.
[0021] The storage unit 12 is a storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). The storage unit 12 may be a data-rewritable semiconductor memory such as a RAM (Random Access Memory), a flash memory, or an NVSRAM (Non-Volatile Static Random Access Memory). The storage unit 12 stores an OS (Operating System) and various programs executed by the server device 10. The storage unit 12 also stores various information used in the execution of the programs. The storage unit 12 stores construction information 121, user information 122, and an honorific dictionary 123.
[0022] The construction information 121 is information required to construct a VR space in a virtual space. The construction information 121 includes image processing conditions and image processing programs for constructing a conference VR space in a virtual space. The construction information 121 also includes photographs taken from a fixed point or 360° of the interior of the conference room, desks, chairs, whiteboard, etc. that are to be expressed in the VR space, or images of the VR space constructed by image processing these photographs. For example, when the VR space is constructed using CG (Computer Graphics), the construction information 121 is an image that reproduces the interior of the conference room that is to be expressed in the VR space.
[0023] The user information 122 is information about each user registered as a participant in the conference VR space. The user information 122 includes, for example, the user's ID, the user's affiliation, the user's job title, the entry history associated with each user ID, the IDs and job titles of other participants, the relationships between users participating in the conference, the importance of the conference, the number of conferences, etc.
[0024] 3 and 4 are diagrams showing an example of the data configuration of user information 122. Table 122-1 shown in Fig. 3 has items such as user ID, affiliation, and job level. For example, a user with user ID "A" belongs to "Company J, Department K" and has a job level of "general." A user with user ID "E" belongs to "Company L, Department M" and has a job level of "chief."
[0025] In addition, table 122-2 shown in Figure 4 has the following items: the ID, affiliation, and job level of the speaker and the user he / she is interacting with, the relationship between the speaker and the other user, and the importance and number of meetings held between the speaker and the other user.
[0026] For example, consider a conference in which user A (user ID "A") is the speaker and user B (user ID "B") is the other party. In this case, the relationship between the speaker and the other party is "subordinate and superior," the importance of the conference between the speaker and the other party is "standard," and the number of conferences is "15."
[0027] Also, for example, a conference will be described in which user A is the speaker and user E (user ID "E") is the other party. In this case, the relationship between the speaker and the other party is "person in charge and customer," the importance of the conference held between the speaker and the other party is "high," and the number of conferences is "1." The importance of the conference may be registered by the user himself or may be determined by the server device 10 based on user information, the history of past conferences, etc.
[0028] The honorific dictionary 123 associates correct honorific expressions with verbs and nouns, for example. The honorific dictionary 123 associates honorifics, humble expressions, and polite expressions that are commonly used in daily life with, for example, a person who performs an action or a target object. The honorific dictionary 123 also associates honorifics based on the honorific language guidelines of the Agency for Cultural Affairs ([online], [searched January 11, 2022], Internet<URL:https: / / www.bunka.go.jp / seisaku / bunkashingikai / kokugo / hokoku / pdf / keigo_tosin.pdf> ) may be associated with five categories of honorific language: honorific language, humble language I, humble language II, polite language, and beautified language.
[0029] The control unit 13 controls the entire server device 10. The control unit 13 is, for example, an electronic circuit such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit), or an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The control unit 13 also has an internal memory for storing programs that define various processing procedures and control data, and executes each process using the internal memory. The control unit 13 also functions as various processing units by running various programs.
[0030] The control unit 13 has a receiving unit 131, a construction unit 132, a voice recognition unit 133, a determination unit 135, and a conversion unit 136. As an explanation of the first embodiment, the function of the control unit 13 will be explained using an example in which a user A (first user) makes a statement to a user B (second user).
[0031] The reception unit 131 receives a request for a user to enter the conference VR space from the user terminal 20. The construction unit 132 constructs a conference VR space that represents a conference room in a virtual space using images or CG.
[0032] 5 is a diagram illustrating the conference VR space provided to the user terminal 20 by the server device 10 shown in FIG. 2. When the reception unit 131 receives a user's entry into the conference VR space, the construction unit 132 acquires the construction information 121 and constructs the conference VR space V1 ((1) in FIG. 5) for which the entry has been received in the virtual space. The construction unit 132 then provides the constructed VR space to the user terminals 20A and 20B. The construction unit 132 places avatars Ua and Ub of users A and B who have entered the conference VR space V1.
[0033] When voice data is input from user terminal 20A in response to the first utterance of user A, voice recognition unit 133 performs voice recognition on the input voice data. Voice recognition unit 133 converts the voice data into text and outputs the voice recognition result.
[0034] As shown in (1) of Fig. 5, an example will be described in which user A and user B are having a conference in VR space V1. Here, when voice data is input from user terminal 20A in response to a speech (first speech) from user A, voice recognition unit 133 performs voice recognition on the input voice data.
[0035] The determination unit 135 determines whether correct honorific language is used in the first utterance based on the speech recognition result by the speech recognition unit 133. For example, if the first utterance uses honorific language to describe the actions of user A himself, the determination unit 135 determines that the first utterance does not use correct honorific language. Also, if the first utterance uses humble language to describe the actions of user A himself, the determination unit 135 determines that the first utterance uses correct honorific language.
[0036] When the determination unit 135 determines that the first utterance does not use correct honorific language, the conversion unit 136 refers to the honorific language dictionary 123 and converts the first utterance into a second utterance using correct honorific language.
[0037] For example, if the first utterance uses honorific language to address user A himself / herself, the conversion unit 136 converts the part using this honorific language into a second utterance using humble language. If the determination unit 135 determines that the first utterance does not use correct honorific language, the conversion unit 136 converts the first utterance into a second utterance and generates text data indicating the content of the second utterance. Furthermore, the conversion unit 136 may convert the second utterance, such as by dividing it into short sentences, so that it is easier for user B to understand.
[0038] When the determination unit 135 determines that the first utterance does not use correct honorific language, the communication unit 11 transmits data corresponding to the second utterance to the user terminal 20B. The communication unit 11 transmits text data indicating the second utterance converted by the conversion unit 136 to the user terminal 20B.
[0039] As a result, in the VR space V1 used by user B, the text data T1 indicating the second utterance is displayed as a utterance by avatar Ua ((2) in FIG. 5). In this way, if user A does not use the correct honorific language, the utterance converted into the correct honorific language will be conveyed to user B.
[0040] A case will now be described in which the determination unit 135 determines that the first utterance uses correct honorific language. In this case, since the first utterance of user A uses correct honorific language, the communication unit 11 transmits the voice data of the first utterance input from user terminal 20A to user terminal 20B as is.
[0041] [Communication processing procedure] FIG. 6 is a sequence diagram illustrating an example of a processing procedure of communication processing according to the first embodiment.
[0042] As shown in FIG. 6, when the server device 10 receives an application to enter the conference VR space from the user terminals 20A and 20B (steps S1 and S2), the server device 10 constructs the conference VR space for which entry has been accepted in a virtual space (step S3) and provides it to the user terminals 20A and 20B (steps S4 and S5).
[0043] Then, when voice data is input from user terminal 20A in response to a utterance (first utterance) from user A (step S6), server device 10 performs voice recognition on the input voice data (step S7).
[0044] Based on the speech recognition result of the speech recognition process (step S7), server device 10 determines whether or not the correct honorific language is used in the first utterance (step S8). If the correct honorific language is used in the first utterance (step S8: Yes), server device 10 transmits the speech data of the first utterance input from user terminal 20A to user terminal 20B as is (step S9), and causes user terminal 20B to output the speech (step S10).
[0045] On the other hand, if the first utterance does not use correct honorific language (step S8: No), the server device 10 does not transmit the audio data of the first utterance to the user terminal 20B, and converts the first utterance into a second utterance using correct honorific language (step S11). The server device 10 transcribes the second utterance (step S12) and generates text data indicating the content of the second utterance. The server device 10 transmits the generated text data to the user terminal 20B (step S13) and causes the user terminal 20B to display the text (step S14).
[0046] [Effects of the First Embodiment] In this way, the server device 10 according to the first embodiment determines whether or not correct honorific language is used in the first utterance of the user A. If the correct honorific language is not used in the first utterance, the server device 10 converts the first utterance into a utterance using correct honorific language and transmits it to the user B who is participating in the conference. Therefore, even if the user A is not good at using honorific language during the conference in the VR space, the utterance using correct honorific language is transmitted to the user B.
[0047] Therefore, according to the first embodiment, if user A does not use the correct honorific language in his / her speech during a conference in the VR space, this speech is converted and conveyed to user B using the correct honorific language. Therefore, even if user B participating in the conference is an elderly person or the like with whom the correct honorific language should be used, user A does not need to carefully examine the use of honorific language, and can smoothly proceed with the conference without feeling burdened by the conversation with user B.
[0048] [Modification of the first embodiment] 7 is a sequence diagram showing an example of a processing procedure of a communication process according to a modification of Embodiment 1. Steps S21 to S31 shown in FIG. 7 are the same processes as steps S1 to S11 shown in FIG.
[0049] The conversion unit 136 generates voice data indicating the content of the second utterance based on the pre-recorded voice of the first user (step S32). In this case, the voice of the first user is recorded in advance. The conversion unit 136 then extracts sounds corresponding to each character indicating the second utterance from the recorded voice, and generates synthetic voice data by synthesizing the extracted sounds in an order corresponding to the content of the second utterance. The communication unit 11 transmits the synthetic voice data generated by the conversion unit 136 to the user terminal 20B (step S33) and causes it to be output (step S34).
[0050] In a variation of the first embodiment, when it is determined that the first utterance does not use the correct honorific language, the server device 10 does not transmit the voice data of the first utterance to the user terminal 20B, but generates synthetic voice data indicating the content of the second utterance based on the pre-recorded voice of the first user. The server device 10 then outputs this synthetic voice data to the user terminal 20B. In the variation of the first embodiment, the second utterance synthesized with the voice of the user A is output from the user terminal 20B, thereby achieving the same effect as the first embodiment.
[0051] The server device 10 may transmit both the text data and the synthetic voice data corresponding to the second utterance to the user terminal 20B. The server device 10 may also select either the text data or the synthetic voice data as the data corresponding to the second utterance and transmit it to the user terminal 20B.
[0052] Furthermore, the conversion unit 136 may generate voice data indicating the content of the second utterance using a model that learns human voice and reads a specified sentence. The model learns the pre-recorded voice of the first user and reproduces the content of the second utterance in the voice of this user.
[0053] [Embodiment 2] Next, a second embodiment will be described. In a communication system according to the second embodiment, a server device scores a speech made by a first user who is a speaker and a second user who is a partner in a conference in a VR space. Then, the server device converts the speech made by the first user into a type of honorific expression according to the scored score.
[0054] Fig. 8 is a block diagram showing an example of the configuration of a server device according to Embodiment 2. As shown in Fig. 8, in server device 210 according to Embodiment 2, storage unit 12 stores scoring rules 124 and classification information 125.
[0055] The scoring rules 124 are data that associate the counting target of the conference score with the score to be added. The scoring rules 124 are set in advance by an administrator of the VR space provider, etc., and updated as appropriate. The server device 210 may update the contents of the scoring rules 124 based on user information, the history of past conferences, etc.
[0056] Fig. 9 is a diagram showing an example of the data configuration of the scoring rules 124. As shown in table 124-1 in Fig. 9, the scoring rules 124 have items of counting targets and scores to be added. The counting targets include the relationship between users who hold a conference, the importance of the conference, the number of conferences, etc.
[0057] In table 124-1, for "relationship," the relationship between the user who is the speaker and the other user is set as, for example, "boss and subordinate," "subordinate and boss," or "person in charge and customer," and a score corresponding to each setting is associated. In table 124-1, the importance of the meeting is set as, for example, "low," "standard," or "high," and a score corresponding to each setting is associated. In table 124-1, a score corresponding to the number of meetings is associated.
[0058] The classification information 125 is data that associates the score of the meeting scored by the scoring unit 2134 (described later) with the type of honorific language corresponding to the score. The type of honorific language corresponding to the score is the type of honorific language that the speaker should use when speaking with the other party in the meeting. The classification information 125 is set in advance by an administrator of the VR space provider, etc., and updated as appropriate.
[0059] Fig. 10 is a diagram showing an example of the data configuration of the classification information 125. Table 125-1 shown in Fig. 10 has items such as score and type of honorific language. For example, table 125-1 indicates that honorific language is unnecessary if the score of the meeting scored by the scoring unit 2134 is between 0 and 15. Table 125-1 indicates that the type of honorific language to be used is "polite language" if the score of the meeting is between 16 and 49. Table 125-1 indicates that the type of honorific language to be used is "polite language, honorific language, and humble language" if the score of the meeting is 50 or higher.
[0060] 10, honorific language, humble language, and polite language that are commonly used are shown as examples of types of honorific language, but the types of honorific language indicated by the classification information 125 may be the five types of honorific language indicated in the honorific language guidelines of the Agency for Cultural Affairs: honorific language, humble language I, humble language II, polite language, and beautified language.
[0061] Server device 210 has a control unit 213 instead of control unit 13 shown in Fig. 2. Control unit 213 further has a scoring unit 2134. Control unit 213 also has a determination unit 2135 and a conversion unit 2136 instead of determination unit 135 and conversion unit 136 shown in Fig. 2.
[0062] The scoring unit 2134 assigns a score when the first user is the speaker and the other party in the conference is the second user based on at least one of the importance of the conference, the attributes of the speaker (first user), the attributes of the other party (second user), the relationship between the speaker and the other party, and the number of conferences between the speaker and the other party. The scoring unit 2134 refers to the user information 122 and calculates the score according to the scoring rules 124.
[0063] For example, a case will be described in which the speaker is user A (user ID "A") and the other party is user B (user ID "B"). The scoring unit 2134 refers to tables 122-1 and 122-2 (FIGS. 3 and 4). Then, the scoring unit 2134 determines that the relationship between user A and user B is "subordinate and superior," the importance of the meeting is "normal," and the number of meetings is "15."
[0064] Then, the scoring unit 2134 counts the score when the speaker is user A and the other party is user B, for example, according to table T124-1. Specifically, the scoring unit 2134 calculates the score in this case to be "20", which is the sum of the score "10" corresponding to "subordinate and superior", the score "10" corresponding to the importance of the meeting "standard", and the score "0" corresponding to the number of meetings "15".
[0065] Also, an example will be described in which the speaker is user A and the other party is user E (user ID "E"). The scoring unit 2134 refers to tables 122-1 and 122-2. The scoring unit 2134 determines that the relationship between user A and user E is "person in charge and customer," the importance of the meeting is "high," and the number of meetings is "0."
[0066] Then, the scoring unit 2134 counts the score in accordance with table T124-1 when the speaker is user A and the other party is user E. Specifically, the scoring unit 2134 calculates the score in this case to be "60", which is the sum of the score "20" corresponding to "person in charge and customer", the score "20" corresponding to the importance of the meeting "important", and the score "20" corresponding to the number of meetings "0".
[0067] The determination unit 2135 determines whether the correct type of honorific language is used in the user's speech based on the speech recognition result by the speech recognition unit 133. The determination unit 2135 has a first determination unit 21351 and a second determination unit 21352.
[0068] The first determination unit 21351 determines, based on the score scored by the scoring unit 2134, the type of honorific language that the user who is the speaker (first user) should use when speaking with the other user in the conference (second user).
[0069] For example, a case will be described in which the speaker is user A and the other party is user B. In this case, the first determination unit 21351 refers to the classification information 125 and determines that the "polite language" corresponding to the score "20" is the type of honorific language that user A should use with user B.
[0070] Next, a case will be described in which the speaker is user A and the other party is user E. In this case, the first determination unit 21351 refers to the classification information 125 and determines that the "polite language, honorific language, humble language" corresponding to the score "60" is the type of honorific language that user A should use with user E.
[0071] The second judgment unit 21352 judges, based on the speech recognition result by the speech recognition unit 133, whether or not the speech (first speech) of the user who is the speaker (first user) uses the type of honorific language that this user should use when speaking with the other user (second user).
[0072] For example, when the speaker is user A and the other party is user B, it is determined whether polite language is used in the speech of user A. When the speaker is user A and the other party is user E, it is determined whether polite language, honorific language, or humble language is used in the speech of user A. The type of honorific language that a user should use is determined by the first determination unit 21351.
[0073] If the second determination unit 21352 determines that the first utterance does not use the type of honorific language that a first user should use when speaking to a second user, the conversion unit 2136 refers to the honorific language dictionary 123 and converts the first utterance into a second utterance that uses the type of honorific language that the first user should use.
[0074] For example, if user A makes a statement without using polite language during a conference with user B, this statement is converted into a statement using polite language. Also, if user A makes a statement without using humble language during a conference with user E, user A's statement is converted into a statement using humble language. The conversion unit 2136 generates text data indicating the content of the converted statement. The communication unit 11 transmits the text data generated by the conversion unit 2136 to the user terminal 20B and displays it. Note that the conversion unit 2136 may further convert the converted statement, such as dividing it into short sentences, so that it is easier for the other party to understand.
[0075] Fig. 11 is a diagram illustrating a conference VR space provided to the user terminal 20 by the server device 210 shown in Fig. 8. When converting the user's speech into correct honorific language, the server device 210 selects the type of honorific language according to the other party, the number of meetings, etc., and converts the user's speech.
[0076] For example, the types of honorifics to be converted are classified as shown in Fig. 11. For example, when speaking with a boss (user B) with whom the user has spoken many times, the server device 210 converts the speech into a slightly informal honorific that uses only polite language ((2-1) in Fig. 11). When speaking with a person in charge of a business partner (user E) with whom the user is speaking for the first time, the server device 210 converts the speech into a proper honorific that uses polite language, honorific language, and humble language ((2-2) in Fig. 11).
[0077] [Communication processing procedure] 12 is a sequence diagram showing an example of a processing procedure of a communication process according to Embodiment 2. In FIG. 12, a case where a user A makes a statement to a user B will be described as an example.
[0078] Steps S41 to S47 shown in Fig. 12 are the same processes as steps S1 to S7 shown in Fig. 6. Server device 210 refers to user information 122 and scores the score when user A is the speaker and the other party is user B, according to scoring rule 124 (step S48).
[0079] Based on the score calculated in step S48, the server device 210 refers to the classification information 125 and determines the type of honorific language that the user A should use with the user B (step S49).
[0080] Based on the speech recognition result in the speech recognition process (step S47), the server device 210 determines whether the first utterance of the speaker, user A, uses the type of honorific language that user A should use when speaking to user B (step S50).
[0081] If the first utterance uses the type of honorific language that user A should use when speaking to user B (step S50: Yes), the audio data of the first utterance input from user terminal 20A is sent as is to user terminal 20B (step S51), and user terminal 20B is caused to output the audio (step S52).
[0082] On the other hand, if the first utterance does not use the type of honorific language that user A should use when speaking to user B (step S50: No), server device 210 does not transmit the audio data of the first utterance to user terminal 20B. Then, server device 210 refers to honorific language dictionary 123 and converts the first utterance into a second utterance that uses the type of honorific language that the first user should use (step S53). Then, server device 210 transcribes the second utterance (step S54) and generates text data indicating the content of the second utterance. Server device w10 transmits the generated text data to user terminal 20B (step S55) and causes user terminal 20B to display the text (step S56).
[0083] [Effects of the second embodiment] In this way, the server device 210 according to the second embodiment scores the score when a first user is the speaker and the other party is a second user in a conference in a VR space, and converts the speech into a type of honorific language according to the scored score. As a result, even if user A is not good at using honorific language during a conference in a VR space, a speech using the correct type of honorific language appropriate for the situation is transmitted to user B. Therefore, user A can smoothly proceed with the conference without feeling burdened by the conversation with user B.
[0084] [Modification of the second embodiment] 13 is a sequence diagram showing an example of a processing procedure of communication processing according to a modification of Embodiment 2. Steps S61 to S73 shown in FIG. 13 are the same processes as steps S41 to S53 shown in FIG.
[0085] The conversion unit 2136 generates voice data indicating the content of the second utterance using the pre-recorded voice of the first user or a model that has been trained on the voice of the first user in advance (step S74). The communication unit 11 transmits the synthesized voice data generated by the conversion unit 2136 to the user terminal 20B (step S75) and causes it to be output (step S76).
[0086] Server device 210 generates synthetic speech data indicating the second utterance in which the correct type of honorific language is used, and outputs this synthetic speech data to user terminal 20B, thereby achieving the same effect as in the second embodiment.
[0087] [System configuration, etc.] Furthermore, the components of each device shown in the figure are conceptual functional units and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of each device can be functionally or physically distributed and integrated in any unit depending on various loads and usage conditions. Furthermore, all or any part of the processing functions performed by each device can be realized by a CPU or GPU and a program analyzed and executed by the CPU or GPU, or can be realized as hardware using wired logic.
[0088] Furthermore, among the processes described in this embodiment, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method.In addition, the information including the processing procedures, control procedures, specific names, various data and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified.
[0089] [program] It is also possible to create a program in which the processes executed by the server devices 10 and 210 described in the above embodiments are written in a computer-executable language. For example, it is also possible to create a program in which the processes executed by the server devices 10 and 210 in the above embodiments are written in a computer-executable language. In this case, the same effects as those of the above embodiments can be obtained by having a computer execute the program. Furthermore, such a program may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be read and executed by a computer to realize processes similar to those of the above embodiments.
[0090] 14 is a diagram showing a computer that executes a program. As shown in the example of FIG. 14, a computer 1000 includes, for example, a memory 1010, a CPU 1020, a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070, and these components are connected by a bus 1080.
[0091] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM 1012, as exemplified in FIG. 14. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090, as exemplified in FIG. 14. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is connected to, for example, a display 1130.
[0092] 14, the hard disk drive 1090 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. That is, the above programs are stored on the hard disk drive 1090, for example, as program modules in which instructions to be executed by the computer 1000 are written.
[0093] The various data described in the above embodiment are stored as program data, for example, in the memory 1010 or the hard disk drive 1090. The CPU 1020 then reads the program module 1093 and the program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as needed, and executes various processing procedures.
[0094] Note that the program module 1093 and program data 1094 related to the program are not limited to being stored in the hard disk drive 1090, and may be stored in, for example, a removable storage medium and read by the CPU 1020 via a disk drive or the like. Alternatively, the program module 1093 and program data 1094 related to the program may be stored in another computer connected via a network (such as a LAN (Local Area Network) or WAN (Wide Area Network)) and read by the CPU 1020 via the network interface 1070.
[0095] The above-described embodiments and their modifications are included in the technology disclosed in this application, as well as in the scope of the invention described in the claims and their equivalents. [Explanation of symbols]
[0096] 10,210 server devices 11 Communications Department 12 Storage section 13,213 Control Unit 20A, 20B User terminal 121 Building Information 122 User Information 123 Honorific Dictionary 124 Scoring Rules 125 Classification information 131 Reception 132 Construction Department 133 Voice Recognition Unit 135,2135 Judgment section 136 Conversion Unit 2134 Scoring Department 21351 First Judgment Section 21352 Second Judgment Section
Claims
1. a construction unit that constructs a virtual reality space in which a conference can be held and provides the virtual reality space to a first user terminal used by a first user and a second user terminal used by a second user; a voice recognition unit that, when voice data is input from the first user terminal in response to a first utterance of the first user, performs voice recognition on the input voice data; a determination unit that determines whether correct honorific language is used in the first utterance based on a speech recognition result by the speech recognition unit; a conversion unit that converts the first utterance into a second utterance using correct honorific language when the determination unit determines that correct honorific language is not used in the first utterance; a transmitting unit that transmits data corresponding to the second utterance to the second user terminal; and the transmitting unit transmits the voice data to the second user terminal when the determining unit determines that correct honorific language is used in the first utterance; when the determination unit determines that correct honorific language is not used in the first utterance, the conversion unit converts the first utterance into the second utterance and generates voice data indicating the content of the second utterance based on a pre-recorded voice of the first user; The processing device is characterized in that, when the judgment unit determines that correct honorific language is not used in the first utterance, the transmission unit does not transmit the audio data of the first utterance to the second user terminal, and transmits the audio data generated by the conversion unit to the second user terminal.
2. a scoring unit configured to score a case in which the first user is a speaker and the other party is the second user based on at least one of the importance of the conference, the attributes of the first user, the attributes of the second user, the relationship between the first user and the second user, and the number of conferences between the first user and the second user; The determination unit a first determination unit that determines a type of honorific language that the first user should use when speaking to the second user, based on the score calculated by the scoring unit; and a second determination unit that determines whether or not the first utterance uses an honorific language that should be used by the first user when speaking to the second user, based on a speech recognition result by the speech recognition unit; and and The processing device according to claim 1, characterized in that, when the second judgment unit determines that the first utterance does not use the type of honorific language that the first user should use when speaking to the second user, the conversion unit converts the first utterance into the second utterance using the type of honorific language that the first user should use when speaking to the second user.
3. A processing method executed by a processing device, a construction step of constructing a virtual reality space in which a conference can be held and providing the space to a first user terminal used by a first user and a second user terminal used by a second user; a voice recognition step of, when voice data is input from the first user terminal in response to a first utterance of the first user, performing voice recognition on the input voice data; a determination step of determining whether or not correct honorific language is used in the first utterance based on a speech recognition result in the speech recognition step; a conversion step of converting the first utterance into a second utterance using correct honorific language when it is determined in the determination step that correct honorific language is not used in the first utterance; a transmitting step of transmitting data corresponding to the second comment to the second user terminal; Including, the transmitting step transmits the voice data to the second user terminal when it is determined in the determining step that correct honorific language is used in the first utterance; The conversion step includes converting the first utterance into the second utterance when it is determined in the determination step that correct honorific language is not used in the first utterance, and generating voice data indicating the content of the second utterance based on a pre-recorded voice of the first user; The processing method is characterized in that, when it is determined in the determination step that correct honorific language is not used in the first utterance, the transmission step does not transmit the audio data of the first utterance to the second user terminal, and transmits the audio data generated in the conversion step to the second user terminal.
4. a construction step of constructing a virtual reality space in which a conference can be held and providing the space to a first user terminal used by a first user and a second user terminal used by a second user; a speech recognition step of, when speech data is input from the first user terminal in response to a first utterance of the first user, performing speech recognition on the input speech data; a determination step of determining whether or not correct honorific language is used in the first utterance based on a speech recognition result in the speech recognition step; a conversion step of converting the first utterance into a second utterance using correct honorific language when it is determined in the determination step that correct honorific language is not used in the first utterance; a transmitting step of transmitting data corresponding to the second comment to the second user terminal; on the computer, the transmitting step, when it is determined in the determining step that correct honorific language is used in the first utterance, transmitting the voice data to the second user terminal; the converting step converts the first utterance into the second utterance when it is determined in the determining step that correct honorific language is not used in the first utterance, and generates voice data indicating the content of the second utterance based on a pre-recorded voice of the first user; The sending step is a processing program that, if it is determined in the determination step that correct honorific language is not used in the first utterance, does not send the audio data of the first utterance to the second user terminal, and sends the audio data generated in the conversion step to the second user terminal.
Citation Information
Patent Citations
Method, system and program for modifying image
JP2009077380A
Honorific expression correction device and reception answer support system using the same
JP2010092257A
Expression conversion device, method and program
JP2014071769A
Web meeting system, delegation method of organizer authority in web meeting system and web meeting system program
JP2015230532A
Telephone answering device
JP2021196462A