Audio transmission method for social virtual reality environment and related device
By converting voice speech into whisper audio in a social virtual reality environment and reducing the ambient voice volume of conversation users, the private conversation problem caused by broadcast transmission is solved, and accurate voice transmission and private conversation experience is achieved.
Patent Information
- Application Number
- CN202510252105.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
AI Technical Summary
In a social virtual reality environment, broadcast voice transmission cannot achieve private conversations, resulting in sound leakage and affecting the experience of other users.
By obtaining the conversation user designated by the speaker user, collecting the voice speech of the speaker user and converting it into whisper audio. Then, the ambient sound volume of the head-mounted display device of the conversation user is lowered and the whisper audio is played through the device.
It realizes accurate voice transmission in a social virtual reality environment, ensuring that private conversations do not affect the experience of other users, avoid unnecessary sound leakage, protect user privacy, and improve the clarity and concentration of conversations.
Smart Images

Figure CN120179205A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technologies, and particularly to an audio transmission method and related device for a social virtual reality environment. Background Art
[0002] In a social virtual reality environment, the voice communication method usually adopts broadcast transmission, so all nearby users can hear the conversation content. Although this method is suitable for some public interaction scenarios, it is not suitable enough in occasions where privacy or local communication is required. For example, in a virtual classroom, students may hope to privately discuss course content or seek help from classmates without disturbing others or affecting the classroom atmosphere. Suppose a student has a question about the content explained by the teacher. He may hope to quietly ask the classmate next to him without other classmates hearing the question or causing unnecessary interference. If the broadcast transmission is still used, it may lead to sound leakage, affecting the learning experience of others and even disrupting the classroom order.
[0003] Therefore, how to achieve precise voice transmission in a social virtual reality environment and ensure that private conversations do not affect the experience of other users is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0004] Based on the above problems, this application provides an audio transmission method and related device for a social virtual reality environment, which can achieve precise voice transmission in a social virtual reality environment and ensure that private conversations do not affect the experience of other users.
[0005] The embodiments of this application disclose the following technical solutions:
[0006] An audio transmission method for a social virtual reality environment, which is applied to a social virtual reality environment with multiple users participating; each user in the social virtual reality environment wears a head-mounted display device, and the method includes:
[0007] Obtain the conversation users specified by the speaking user, and collect the voice speech of the speaking user;
[0008] Convert the voice speech into whispering audio;
[0009] Reduce the ambient sound volume of the head-mounted display device of the conversation user to a preset volume, and play the whispering audio through the head-mounted display device of the conversation user.
[0010] In a possible implementation, each user participating in the social virtual reality environment corresponds to a unique virtual image in the social virtual reality environment;
[0011] The method further includes:
[0012] Set the user who initiates a specified conversation using the corresponding virtual avatar in the social virtual reality environment as the speaking user; any user controls the virtual avatar corresponding to him / her in the social virtual reality environment to emit a ray pointing to other virtual avatars through the character selection button on the controller of the social virtual reality environment to initiate the specified conversation;
[0013] Set the user corresponding to the virtual avatar specified by the specified conversation as the conversation user.
[0014] In a possible implementation manner, the collecting of the voice speech of the speaking user includes:
[0015] When the speaking user presses the speaking button on the controller of the social virtual reality environment, start the voice collection of the speaking user;
[0016] When the speaking button on the controller of the social virtual reality environment is released, end the voice collection of the speaking user;
[0017] Use the audio collected during the voice collection process as the voice speech of the speaking user.
[0018] In a possible implementation manner, the method further includes:
[0019] When the speaking user presses the speaking button on the controller of the social virtual reality environment, generate a first-person perspective image set; the first-person perspective image set includes the first initial position image of the speaking user in the first-person perspective for the conversation user, the speaking user's first moving closer image, and the first speaking user whispering image;
[0020] Among them, the first-person perspective image set is shown to the conversation user; the distance between the virtual avatar corresponding to the speaking user and the virtual avatar corresponding to the conversation user in the second speaking user whispering image is equal to the preset communication distance.
[0021] In a possible implementation manner, the environmental light intensity in the first initial position image of the speaking user is greater than the environmental light intensity in the first moving closer image of the speaking user;
[0022] The environmental light intensity in the first initial position image of the speaking user is greater than the environmental light intensity in the first whispering image of the speaking user;
[0023] Among them, highlight the virtual avatar of the speaking user in the first moving closer image and the first whispering image of the speaking user.
[0024] In a possible implementation, the method further includes:
[0025] When the head-mounted display device of the dialogue user plays the whispered audio, record the language speech of other users and convert it into text form to obtain a first recorded text, and convert the whispered audio into text form to obtain a second recorded text;
[0026] After the whispered audio playback ends, display the first recorded text and the second recorded text to the dialogue user.
[0027] In a possible implementation, at least two silent fans are provided on the head-mounted display device; the silent fans are provided on both sides of the head-mounted display device;
[0028] The method further includes:
[0029] When the head-mounted display device of the dialogue user plays the whispered audio, use the silent fans provided on the head-mounted display device of the dialogue user to simulate the breathing sensation during whispering; the breathing sensation is the slight tactile stimulation generated by the airflow generated by the silent fans on the user's ears and the surrounding skin, so as to simulate the breathing effect during whispering.
[0030] A dialogue transmission device for a social virtual reality environment, the device includes:
[0031] An acquisition unit, configured to acquire a dialogue user designated by a speaking user;
[0032] A collection unit, configured to collect the voice speech of the speaking user;
[0033] A voice speech conversion unit, configured to convert the voice speech into a whispered audio;
[0034] A volume reduction unit, configured to reduce the volume of the ambient sound of the head-mounted display device of the dialogue user to a preset volume;
[0035] A whispered audio playback unit, configured to play the whispered audio through the head-mounted display device of the dialogue user.
[0036] An audio transmission device for a social virtual reality environment, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the audio transmission method for the social virtual reality environment as described above is implemented.
[0037] A computer-readable storage medium, in which instructions are stored. When the instructions run on a terminal device, the terminal device is caused to execute the audio transmission method for the social virtual reality environment as described above.
[0038] Compared with the prior art, the present application has the following beneficial effects:
[0039] The present application provides an audio transmission method and related device for a social virtual reality environment. Specifically, when executing the audio transmission method for a social virtual reality environment applied to multiple users wearing head-mounted display devices, first, the conversation users specified by the speaking user can be obtained, and the voice speech of the speaking user can be collected. Then, the voice speech is converted into whispered audio. Next, the volume of the ambient sound of the head-mounted display device of the conversation user is reduced to a preset volume. Then, the whispered audio is played through the head-mounted display device of the conversation user. The transmission of the whispered audio in the present application is only directed to the specified conversation users, ensuring that only specific users can hear the speech content, thereby avoiding unnecessary sound leakage and protecting the privacy of users. At the same time, by reducing the volume of the ambient sound of the conversation user, the influence of background noise in the virtual environment on the whispered audio is avoided, ensuring that the speech content is clear and easy to understand. Even if there is noise around, the conversation can be made more focused and efficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] To more clearly illustrate the technical solutions in the embodiments or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 It is a flowchart of a method for an audio transmission method in a social virtual reality environment provided by an embodiment of the present application;
[0042] Figure 2 It is a schematic diagram of a first-person perspective image set provided by an embodiment of the present application;
[0043] Figure 3 It is a schematic diagram of an overhead image set provided by an embodiment of the present application;
[0044] Figure 4 It is a flowchart of a method for a voice speech collection method provided by an embodiment of the present application;
[0045] Figure 5 It is a schematic diagram of the structure of an audio transmission device for a social virtual reality environment provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] To facilitate understanding of the technical solutions provided by the embodiments of the present application, the following will first explain the background technology related to the embodiments of the present application.
[0047] A Social Virtual Reality (Social VR) environment is a digital interactive space that allows multiple users to enter together through virtual reality technology. In such an environment, users usually engage in social interactions, communications, entertainment, or cooperation in the virtual world through virtual avatars or characters (e.g., by embodying virtual personas). In a social virtual reality environment, users can interact with others in real time, whether through text, voice, or gestures, and can even exchange or cooperate through virtual items, simulating social behaviors in the real world. Users can obtain an immersive experience and feel the effect of being in a virtual space with others by wearing devices such as head-mounted displays (HMDs, Head-Mounted Displays) and headphones.
[0048] A head-mounted display is a wearable device that typically includes a display screen, sensors, and audio devices. After wearing it, users can experience a virtual environment visually and auditorily. Through the head-mounted display, users can see images in the virtual world, feel the depth of space, and track the movement of their heads through the sensors in the device, ensuring that in a social virtual reality environment, the users' experience is synchronized with their movements and gazes.
[0049] In a social virtual reality environment, voice communication usually adopts a broadcast transmission method, that is, each user can hear the conversation content of others nearby. This method is very suitable for some public interaction scenarios, such as social gatherings, group discussions, etc., because it can ensure the timely transmission of information and promote interaction among groups. However, in some occasions that require privacy or local communication, this broadcast transmission is not appropriate enough.
[0050] Taking a virtual meeting as an example, during a team presentation or discussion, participants may need to have private communication with specific participants without disturbing the main topic discussion. For example, during a team presentation, a technician may need to remind the speaker of certain details or provide instant feedback. However, if the broadcast transmission continues, other participants can also hear this content, which will not only lead to information leakage but may also affect the order of the meeting and the smoothness of the presentation.
[0051] To solve this problem, an audio transmission method and related device for a social virtual reality environment are provided in an embodiment of the present application. First, the dialogue user designated by the speaking user is obtained, and the voice speech of the speaking user is collected. Then, the voice speech is converted into whispered audio data. Next, the ambient sound volume of the head-mounted display device of the dialogue user is reduced to a preset volume, and the whispered audio is played through the head-mounted display device of the dialogue user. The whispered audio transmission of the present application focuses on the designated dialogue user, ensuring that only specific users can hear the speech content, thus effectively avoiding unnecessary sound leakage and enhancing the privacy protection between users. At the same time, by reducing the ambient sound volume of the dialogue user, the interference of background noise in the virtual environment on the whispered audio is reduced, further ensuring the clarity and comprehensibility of the speech content. Even if there are background noises around, this method helps to make the conversation more focused and efficient, improving the quality and effect of communication in the virtual environment.
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0053] See Figure 1 , which is a flowchart of an audio transmission method for a social virtual reality environment provided by an embodiment of the present application. This method is applied to a social virtual reality environment with multiple users participating, and each user in the social virtual reality environment wears a head-mounted display device.
[0054] As Figure 1 shown, the audio transmission method for this social virtual reality environment may include steps S101 - S103:
[0055] S101: Obtain the dialogue user designated by the speaking user, and collect the voice speech of the speaking user.
[0056] "Obtain the dialogue user designated by the speaking user, and collect the voice speech of the speaking user" means that the system first identifies and determines the dialogue object selected or designated by the speaking user. Then, the system records the voice content of the speaking user through a microphone or other voice collection devices for subsequent processing. The purpose of this process is to ensure that the system can accurately capture the communication content between the speaking user and the specific dialogue user without being interfered by other irrelevant sounds.
[0057] In a possible implementation manner, each user participating in the social virtual reality environment corresponds to a unique virtual image in the social virtual reality environment.
[0058] Specifically, the virtual avatar has uniqueness: each user has a unique virtual avatar in the virtual social environment, which represents the user's presence in the virtual space.
[0059] In one possible implementation, the method further includes:
[0060] Determination of the speaking user: When a user initiates a conversation in the environment through the virtual avatar, the system determines the user corresponding to the virtual avatar that initiates the conversation as the "speaking user". This means that whoever's virtual avatar initiates the conversation is the speaker.
[0061] Way of specifying a conversation: Any user can select a virtual avatar through a button on the controller in the social virtual reality environment and emit a ray (such as a virtual finger or light) with it, pointing at another virtual avatar to initiate a conversation. This process is equivalent to attracting the attention of another virtual avatar and starting a conversation by "emitting" a certain signal.
[0062] Determination of the conversation user: The user corresponding to the virtual avatar pointed at by the emitted ray will be set as the "conversation user" by the system, that is, this user becomes the recipient of the conversation.
[0063] Optionally, the ray can be controlled to point at other virtual characters through the trigger button or touchpad / button of the VR (Virtual Reality) controller in the social virtual reality environment:
[0064] Trigger button: The trigger button is usually located at the front of the VR controller, similar to a trigger button. After pressing this button, the controller emits a ray, pointing in the direction of the user's line of sight or the pointing direction of the controller.
[0065] Touchpad / Button:
[0066] In some VR devices, the controller may be equipped with a touchpad or other customizable buttons. Users can emit a ray by triggering the touchpad or touch button, control the direction of the ray, and then point at virtual characters or interact with them.
[0067] In one possible implementation, when the speaking user presses the speaking button on the controller of the social virtual reality environment, the system generates a set of first-person perspective image sets.
[0068] The first-person perspective image set includes but is not limited to three images:
[0069] The first initial position image of the speaking user from the first-person perspective (see Figure 2The first figure from the left in [Figure ID]: shows the position of the speaking user at the start of the conversation from the perspective of the conversing user;
[0070] The first moving-closer image of the speaking user (see Figure 2 the second figure from the left in [Figure ID]): shows the dynamic process of the speaking user gradually approaching the conversing user;
[0071] The whispering image of the speaking user (see Figure 2 the first figure from the right in [Figure ID], where Taylor is the avatar): shows the scenario of the speaking user during a close-range conversation.
[0072] In one possible implementation, in the first moving-closer image of the speaking user and the first whispering image of the speaking user, the avatar of the speaking user will be prominently displayed. Specifically:
[0073] The ambient light intensity in the first initial-position image of the speaking user is greater than that in the first moving-closer image of the speaking user, thus creating an effect of gradually approaching.
[0074] At the same time, the ambient light intensity in the first moving-closer image of the speaking user is greater than that in the first whispering image of the speaking user, further emphasizing the atmosphere of gradually approaching.
[0075] In particular, the surrounding environment of the avatar of the speaking user in the first moving-closer image of the speaking user and the second whispering image of the speaking user will be highlighted, similar to a spotlight highlighting the selected speaker, ensuring that the avatar of the speaker is more prominent in the scene and enhancing the realism and interactive experience of the conversation.
[0076] It should be noted that Figure 2 the change in ambient light intensity in [Figure ID] is not obvious. Therefore, this application also provides an overhead image set for easy understanding. Figure 3 as shown in Figure 3 The overhead image set shown currently is not presented to any user (it can also be presented to users according to actual needs). The overhead image set is proposed to facilitate understanding of the change in ambient light intensity in the first-person perspective image set. The overhead image set includes but is not limited to three images:
[0077] The second initial-position image of the speaking user from an overhead perspective (see Figure 3 the first figure from the left in [Figure ID], and the ambient light intensity of this figure is the same as that of the first initial-position image of the speaking user in the first-person perspective image set): shows the position of the speaking user at the start of the conversation from an overhead perspective;
[0078] The second moving-closer image of the speaking user (see Figure 3The second picture from the left in [reference], the ambient light intensity of this picture is equal to the ambient light intensity of the first moving approaching image of the speaking user in the first-person perspective image set): It shows the dynamic process of the speaking user gradually approaching the conversation user from a top-down perspective;
[0079] The second whispering image of the speaking user (see Figure 3 The first picture from the right in [reference], the ambient light intensity of this picture is equal to the ambient light intensity of the whispering image of the first speaking user in the first-person perspective image set): It shows the scene of the speaking user during a close conversation from a top-down perspective.
[0080] See Figure 4 , Figure 4 is the flowchart of a voice speech acquisition method provided by an embodiment of the present application. Correspondingly, the acquisition of the voice speech of the speaking user described in step S101 can be specifically implemented through steps A1 - A3:
[0081] A1: When the speaking user presses the speaking button on the controller of the social virtual reality environment, start the voice acquisition of the speaking user.
[0082] When the user presses the speaking button on the controller, the voice acquisition starts. This means that at the moment the user presses the button, the system starts recording the user's voice.
[0083] In a possible implementation, the user's speech can be activated through the microphone button or the menu / select button to activate the voice function.
[0084] Microphone Button: On some VR controllers, there are dedicated buttons or touch areas for activating the microphone, allowing the user to start speaking or issuing voice commands by pressing the button. Some devices may control whether to enable the voice function by long pressing or short pressing.
[0085] Menu / Select Button: In some social VR platforms, the menu button or select button is sometimes used to enable the voice chat function. When the user presses this button, the voice communication in the virtual environment can be activated, allowing voice interaction with other users.
[0086] A2: When the speaking button on the controller of the social virtual reality environment is released, end the voice acquisition of the speaking user.
[0087] When the user releases the speaking button, the voice acquisition process stops. This is to ensure that the voice acquisition is limited to the speaking time period of the user and avoid misacquisition of background noise or irrelevant sounds.
[0088] A3: Use the audio collected during the voice collection process as the voice speech of the speaking user.
[0089] The collected audio data will be processed and used as the voice speech of the speaking user during the voice speech process. That is, after the audio data is collected, the system will convert the audio into actual voice information for display in a virtual reality environment or sharing with other users.
[0090] Steps A1 - A3 implement a voice input mechanism based on physical button control, ensuring that the user starts voice input when pressing the button and ends voice input when releasing the button, and finally using the audio data for voice communication.
[0091] S102: Convert the voice speech into whispering audio.
[0092] The key to converting into whispering audio lies in audio processing. Whispering usually has some characteristics, such as lower volume, softer tone, and higher audio frequency. To achieve this effect, the voice signal can be processed as follows:
[0093] Lower the volume: The volume of whispering is usually lower than that of normal conversation, so the volume of the collected voice needs to be lowered to the preset "whispering" level;
[0094] Increase the high - frequency components: Whispering contains more high - frequency components, so the high - frequency part of the voice can be enhanced through a high - pass filter and the low - frequency part can be removed to simulate the high - frequency sound effect during whispering;
[0095] Adjust the pitch and speed of the voice: The voice of whispering usually sounds softer than normal speaking, and the pitch of the voice can be appropriately adjusted to make it sound like whispering.
[0096] In addition, if the whispering is played through a virtual reality device, to increase the immersion, the stereo effect can also be adjusted. The whispering audio may be played only through the sound of one ear, or simulate the sound coming from a certain direction. Using a stereo processing algorithm, the whispering can be simulated as coming from a specific position and distance.
[0097] Generally speaking, converting the voice into whispering audio requires audio filtering, volume adjustment, and frequency adjustment, all of which are to make the final effect sound like a whispering communication in a virtual environment.
[0098] In a possible implementation, voice recognition and voice cloning technologies can be used to convert the voice speech into whispering audio.
[0099] S103: Lower the ambient sound volume of the head-mounted display device of the conversation user to a preset volume, and play the whispered audio through the head-mounted display device of the conversation user.
[0100] In a virtual reality or augmented reality environment, to provide an immersive experience and ensure that users can clearly hear whispered audio during private conversations while reducing interference from external noise, the system needs to finely control the user's audio experience. Specifically, the system will automatically adjust the volume settings of the head-mounted display device according to the conversation scenario and the user's needs.
[0101] First, when the system detects that whispered audio is generated and a whispered or low-volume private conversation is about to occur, it will automatically identify the volume status of the current ambient sound. Ambient sound usually includes conversations of other users in the virtual world, background noise, and various sound effects generated when interacting with the virtual environment. To avoid interference from these noises with the clarity of the whispered conversation, the system will lower the ambient volume to a preset value, which is usually set to but not limited to 50% of the ambient volume. In this way, the system can effectively reduce interference from the environment and enable the whispered audio to be more clearly conveyed to the user.
[0102] Next, the system will play the whispered audio through the head-mounted display device of the conversation user. The volume of the whispered audio itself is low and is usually used for more private and intimate communication in the virtual environment. Therefore, to ensure that the transmission of the whisper is not affected by external noise, the system needs to ensure that the whispered audio is played at an appropriate volume through fine audio control. At this time, the whispered audio should not only be clear and natural but also coordinated with the adjustment of the ambient volume to avoid discomfort or auditory burden on the user due to sudden volume changes.
[0103] Through this control method, users can enjoy a realistic and private conversation experience in a highly immersive virtual environment. The adjustment of the ambient volume and the playback process of the whispered audio will change dynamically according to the actual situation to ensure that users can have sufficient clarity and comfort when whispering without being disturbed by other sounds. This intelligent audio control technology can greatly improve the quality of the user experience when interacting with others in virtual reality.
[0104] In a possible implementation manner, the method further includes the following key steps:
[0105] First, while the head-mounted display device of the conversation user plays the whispered audio, the system will automatically record the language speeches of other users. Through speech recognition technology, these language speeches will be converted into text form and stored as recorded text. This process ensures that all public conversation content can be recorded in a timely manner, whether it is a conversation that the user actively participates in or an exchange that is not directly heard.
[0106] Secondly, when the whispered audio playback ends, the system will display the recorded text. At this time, the system can display speech bubbles above the virtual avatars of the corresponding speakers to show the recorded text, where the common conversation content is displayed in standard text and the private messages are displayed in blue text. These speech bubbles are displayed for 10 seconds, enabling users to quickly understand the missed common conversation content and avoid missing important exchanges.
[0107] This method combines whispered audio and text recording functions, helping users maintain awareness of common conversations in a virtual environment while also providing effective management of private communications.
[0108] In a possible implementation, at least two silent fans are set in the head-mounted display device worn by the conversation users to simulate the breathing sensation during whispering. In the design of the head-mounted display device, the silent fans are installed on both the left and right sides of the device, usually near the ears on both sides of the device or in positions close to the ears.
[0109] When the whispered audio is played through this device, the silent fans will start to simulate the natural breathing feeling of the user during whispering. This simulation of the breathing sensation can enhance the immersion of whispered communication in the virtual environment, enabling users to experience a more realistic interaction.
[0110] The airflow generated by the silent fans gently blows on the user's ears and the surrounding skin, producing a tactile stimulus. This stimulus simulates the slight breathing sensation accompanied by whispering. Specifically, it describes the effect produced when the airflow of the silent fans acts on the ears and the surrounding skin. This airflow is as gentle and slight as the breathing during whispering. In this way, users can perceive a natural tactile effect, enhancing the immersion as if someone is whispering in their ear.
[0111] It should be noted that the fan noise of the silent fans is about 20 decibels, equivalent to the volume of calm breathing, and will not interfere with the audio experience.
[0112] In addition, to ensure the adaptability and effectiveness of the fans, the size of the silent fans is usually designed to be 3 cm in length, 3 cm in width, and 1 cm in height (it can also be other appropriate sizes). This size design not only enables the fans to be compactly installed in the head-mounted display device but also provides sufficient airflow simulation effect while maintaining silence, further improving the fidelity of the whispered audio.
[0113] Based on the content of S101 - S103, first obtain the conversation user specified by the speaking user. Then, collect the voice speech of the speaking user. Next, convert the voice speech into whispering audio. Finally, reduce the ambient sound volume of the head - mounted display device of the conversation user to a preset volume and play the whispering audio through the head - mounted display device of the conversation user. The whispering audio transmission method proposed in this application is only for the specified conversation user, ensuring that only specific users can hear the speech content, thus effectively avoiding unnecessary sound leakage and protecting the privacy of users. This method can provide a more private communication experience in a social virtual reality environment, avoiding the interference and privacy risks that may be brought by broadcast transmission. At the same time, by reducing the ambient sound volume of the conversation user, the interference of background noise in the virtual environment to the whispering audio is effectively reduced, ensuring that the voice content is clearer and more understandable. Even if there is background noise around, users can focus on the conversation with specific objects, further improving the efficiency and quality of communication. Such optimization not only enhances the privacy of voice communication but also improves the fluency and concentration of the conversation.
[0114] See Figure 5 , Figure 5 is a schematic structural diagram of an audio transmission device for a social virtual reality environment provided by an embodiment of this application. As Figure 5 shown, the audio transmission device for the social virtual reality environment includes:
[0115] An acquisition unit 501, configured to obtain the conversation user specified by the speaking user;
[0116] A collection unit 502, configured to collect the voice speech of the speaking user;
[0117] A voice speech conversion unit 503, configured to convert the voice speech into whispering audio;
[0118] A volume reduction unit 504, configured to reduce the ambient sound volume of the head - mounted display device of the conversation user to a preset volume;
[0119] A whispering audio playback unit 505, configured to play the whispering audio through the head - mounted display device of the conversation user.
[0120] In a possible implementation manner, each user participating in the social virtual reality environment corresponds to a unique virtual avatar in the social virtual reality environment.
[0121] In a possible implementation manner, the device further includes:
[0122] A speaking user setting unit for setting the user who initiates a specified conversation using the corresponding virtual avatar in the social virtual reality environment as the speaking user; any user controls the virtual avatar corresponding to him / her in the social virtual reality environment to emit a ray pointing to other virtual avatars through the character selection button on the controller of the social virtual reality environment to initiate the specified conversation;
[0123] A conversation user setting unit for setting the user corresponding to the virtual avatar specified in the specified conversation as the conversation user.
[0124] In a possible implementation manner, the acquisition unit 502 specifically includes:
[0125] A voice acquisition start unit for starting the voice acquisition of the speaking user when the speaking user presses the speaking button on the controller of the social virtual reality environment;
[0126] A voice acquisition end unit for ending the voice acquisition of the speaking user when the speaking button on the controller of the social virtual reality environment is released;
[0127] A setting unit for using the audio collected during the voice acquisition process as the voice speech of the speaking user.
[0128] In a possible implementation manner, the device further includes:
[0129] A conversation image set generation unit for generating a first-person perspective image set when the speaking user presses the speaking button on the controller of the social virtual reality environment; the first-person perspective image set includes the first initial position image of the speaking user in the first-person perspective for the conversation user, the speaking user's first moving closer image, and the first speaking user whispering image;
[0130] Among them, the first-person perspective image set is shown to the conversation user; the distance between the virtual avatar corresponding to the speaking user and the virtual avatar corresponding to the conversation user in the first speaking user whispering image is equal to the preset communication distance.
[0131] In a possible implementation manner, the environmental light intensity in the first initial position image of the speaking user is greater than the environmental light intensity in the first moving closer image of the speaking user;
[0132] The environmental light intensity in the first initial position image of the speaking user is greater than the environmental light intensity in the first whispering image of the speaking user;
[0133] Among them, the virtual avatar of the speaking user in the first moving closer image and the first whispering image of the speaking user is highlighted.
[0134] In a possible implementation, the device further includes:
[0135] An integration unit, when the head-mounted display device of the conversation user plays the ear audio, is configured to record the language speech of other users and convert it into a text form to obtain a first recorded text, and convert the ear audio into a text form to obtain a second recorded text;
[0136] A text display unit, when the playback of the ear audio ends, is configured to display the first recorded text and the second recorded text to the conversation user.
[0137] In a possible implementation, the head-mounted display device is provided with at least two silent fans; the silent fans are arranged on both sides of the head-mounted display device.
[0138] In a possible implementation, the device further includes:
[0139] A whispering simulation unit, when the ear audio is played through the head-mounted display device of the conversation user, is configured to simulate the breathing sensation during whispering by using the silent fans arranged on the head-mounted display device of the conversation user; the breathing sensation is a slight tactile stimulation generated by the airflow generated by the silent fans on the user's ears and the surrounding skin, so as to simulate the breathing effect during whispering.
[0140] In addition, an embodiment of the present application further provides an attack behavior detection device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the audio transmission method of the social virtual reality environment as described above is implemented.
[0141] In addition, an embodiment of the present application further provides a computer-readable storage medium, in which instructions are stored. When the instructions run on a terminal device, the terminal device is enabled to execute the audio transmission method of the social virtual reality environment as described above.
[0142] An embodiment of the present application provides an audio transmission device for a social virtual reality environment. First, the acquisition unit 501 is used to acquire the conversation user specified by the speaking user, and the collection unit 502 is used to collect the voice speech of the speaking user. The voice speech conversion unit 503 converts the voice speech into whispered audio. Then, the volume reduction unit 504 is used to reduce the volume of the ambient sound of the head-mounted display device of the conversation user to a preset volume, so that the whispered audio playback unit 505 can play the whispered audio through the head-mounted display device of the conversation user. The whispered audio transmitted by the present application can ensure that only the specified conversation user can hear the speech content, effectively avoiding unnecessary sound leakage, thereby protecting the user privacy. At the same time, by reducing the surrounding environment volume, the interference of background noise is reduced, ensuring that the voice content is clearer, improving the concentration and efficiency of the conversation, and maintaining good communication quality even in a noisy environment.
[0143] The above has introduced in detail an audio transmission method and related device provided by the present application. The various embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part. It should be noted that for those of ordinary skill in the art in the technical field of the present application, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
[0144] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0145] It should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
Claims
1. An audio transmission method for a social virtual reality environment, characterized in that: Applied to a social virtual reality environment in which multiple users participate; each user in the social virtual reality environment wears a head mounted display device, the method comprising: Acquire the dialogue user specified by the speaking user, and collect the voice speech of the speaking user; converting the voice utterance into whispered audio; The ambient sound volume of the head-mounted display device of the conversation user is reduced to a preset volume, and the whisper audio is played through the head-mounted display device of the conversation user.
2. The method according to claim 1, characterized in that Each user participating in the social virtual reality environment corresponds to a unique virtual image in the social virtual reality environment; The method further comprises: The user who initiates the designated conversation by using the corresponding virtual image in the social virtual reality environment is set as the speaking user; any user controls the virtual image corresponding to the user in the social virtual reality environment to emit a ray pointing to other virtual images through the character selection button on the controller of the social virtual reality environment to initiate the designated conversation; The user corresponding to the virtual image specified by the specified dialogue is set as the dialogue user.
3. The method according to claim 1, characterized in that The collecting of the voice speech of the speaking user includes: When the speaking user presses a speaking button on the controller of the social virtual reality environment, voice collection of the speaking user is started; When the speaking button on the controller of the social virtual reality environment is released, the voice collection of the speaking user is terminated; The audio collected during the voice collection process is used as the voice speech of the speaking user.
4. The method according to claim 2, characterized in that: The method further comprises: When the speaking user presses a speaking button on the controller of the social virtual reality environment, a first-person perspective image set is generated; the first-person perspective image set includes a first initial position image of the speaking user, a first moving approach image of the speaking user, and a first speaking user whispering image under the first-person perspective of the conversation user; Among them, the first-person perspective image set is displayed to the conversation user; the distance between the virtual image corresponding to the speaking user in the second whisper image of the speaking user and the virtual image corresponding to the conversation user is equal to the preset communication distance.
5. The method according to claim 4, characterized in that The ambient light intensity in the first initial position image of the speaking user is greater than the ambient light intensity in the first moving approach image of the speaking user; The ambient light intensity in the first initial position image of the speaking user is greater than the ambient light intensity in the first whispering image of the speaking user; The virtual image of the speaking user in the first moving approaching image of the speaking user and the first whispering image of the speaking user is highlighted.
6. The method according to claim 1, characterized in that The method further comprises: When the head mounted display device of the conversation user plays the whisper audio, the speech of the other user is recorded and converted into text form to obtain a first recorded text, and the whisper audio is converted into text form to obtain a second recorded text; After the whisper audio is played, the first record text and the second record text are displayed to the conversation user.
7. The method according to claim 1, characterized in that The head mounted display device is provided with at least two silent fans; the silent fans are arranged on both sides of the head mounted display device; The method further comprises: When the whisper audio is played through the head-mounted display device of the conversation user, the silent fan provided on the head-mounted display device of the conversation user is used to simulate the breathing feeling during the whisper; the breathing feeling is that the airflow generated by the silent fan produces slight tactile stimulation on the user's ears and surrounding skin, thereby simulating the breathing effect during the whisper.
8. A conversation transmission device for a social virtual reality environment, characterized in that: The device comprises: An acquisition unit, used for acquiring a dialogue user specified by a speaking user; A collection unit, used for collecting the voice speech of the speaking user; a voice utterance conversion unit, configured to convert the voice utterance into whisper audio; A volume reduction unit, configured to reduce the volume of the ambient sound of the head mounted display device of the conversation user to a preset volume; A whisper audio playback unit is used to play the whisper audio through a head-mounted display device of the conversation user.
9. An audio transmission device for a social virtual reality environment, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the audio transmission method for a social virtual reality environment as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device executes the audio transmission method for a social virtual reality environment as described in any one of claims 1 to 7.