Virtual space-based audio processing method, device and computer equipment
By distinguishing between interested and uninterested sound sources in virtual space and adjusting their audio accordingly, the problem of users having difficulty hearing interesting sounds in noisy environments is solved, achieving better sound effect differentiation and an immersive experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2022-01-13
- Publication Date
- 2026-07-24
AI Technical Summary
In virtual spaces, users often struggle to clearly hear sounds of interest in noisy environments. Existing technologies, through distance sensing processing and virtual stereo reconstruction, are insufficient to effectively distinguish sound sources.
By identifying the sound sources of interest and those of non-interest in the virtual space, the audio of the sound sources of interest is enhanced, while the audio of the sound sources of non-interest is attenuated and mixed to improve the sound effect differentiation of the sound sources of interest.
In noisy environments, users can more easily hear the sounds they are interested in, improving sound effect differentiation and enhancing the user's immersive auditory experience.
Smart Images

Figure CN114049871B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to an audio processing method, apparatus and computer device based on virtual space. Background Technology
[0002] Spatial sound effects utilize audio technology to enhance the sense of depth and spatial layering of sound. Played through headphones or a combination of two or more speakers, it recreates the auditory experience of a real-world environment, allowing listeners to clearly discern the location, distance, and movement of different acoustic objects. It also provides a feeling of being completely enveloped by sound, creating an immersive auditory experience as if the listener were actually in the environment. Therefore, proper audio processing is crucial to achieving the best possible sound experience for users.
[0003] In related technologies, audio processing is primarily achieved through virtual stereo reconstruction. Specifically, the volume of each sound source is adjusted based on its distance from the user, a process known as distance perception processing. Sound sources that are farther away have a lower volume setting, while those closer have a higher volume setting. This process emphasizes creating a sense of distance between different sound sources for the user, making it difficult for them to clearly hear sounds of interest, especially in noisy environments. Summary of the Invention
[0004] Therefore, it is necessary to provide a virtual space-based audio processing method, apparatus, and computer device that can make it easier for users to hear the sounds they are interested in, addressing the aforementioned technical problems.
[0005] A virtual space-based audio processing method, the method comprising:
[0006] In response to interactive operations in the virtual space, identify the sound sources of interest in the virtual space and treat all sound sources in the virtual space other than the sound sources of interest as non-interested sound sources.
[0007] Acquire the target audio corresponding to the sound source of interest, and the target non-interested audio corresponding to the sound source of non-interest.
[0008] Audio adjustment processing is performed on at least one of the target attention audio and the target non-attention audio to obtain a first intermediate audio of the attention source and a second intermediate audio of the non-attention source, wherein the sound effect discrimination of the first intermediate audio is greater than that of the second intermediate audio.
[0009] The first intermediate audio and the second intermediate audio are mixed to obtain the mixing result.
[0010] A virtual space-based audio processing device, the device comprising:
[0011] The determination module is used to respond to interactive operations in the virtual space, determine the sound sources of interest in the virtual space, and treat all sound sources in the virtual space other than the sound sources of interest as non-sound sources of interest.
[0012] The acquisition module is used to acquire the target audio of interest corresponding to the sound source of interest, and the target non-interest audio of non-interest sound sources;
[0013] An audio adjustment processing module is used to perform audio adjustment processing on at least one of the target attention audio and the target non-attention audio to obtain a first intermediate audio of the attention source and a second intermediate audio of the non-attention source, wherein the sound effect differentiation of the first intermediate audio is greater than that of the second intermediate audio.
[0014] The mixing module is used to mix the first intermediate audio and the second intermediate audio to obtain the mixing result.
[0015] In one embodiment, the determining module is configured to, in response to an interactive operation on the virtual space, determine a region of interest in the virtual space to which the interactive operation is directed; and designate a sound source in the region of interest as a sound source of interest.
[0016] In one embodiment, the acquisition module is used to perform mixing processing on the audio of each of the multiple sources of interest to obtain the target audio of interest; and to perform mixing processing on the audio of each of the multiple sources of non-interest to obtain the target audio of non-interest.
[0017] In one embodiment, the audio adjustment processing module is used to perform signal enhancement processing on the target audio of interest to obtain a first intermediate audio corresponding to the sound source of interest; and to perform signal attenuation processing on the target non-interest audio to obtain a second intermediate audio corresponding to the non-interest sound source.
[0018] In one embodiment, the audio adjustment processing module includes:
[0019] The first determining unit is used to determine the audio to be adjusted from the target attention audio and the target non-attention audio;
[0020] The second determining unit is used to determine the adjustment parameters corresponding to each sampling time of the audio to be adjusted;
[0021] The third determining unit is used to perform audio adjustment processing on the audio to be adjusted based on the adjustment parameters corresponding to each sampling time, and to determine the first intermediate audio of the source of interest and the second intermediate audio of the source of non-interest based on the audio adjustment results.
[0022] In one embodiment, the second determining unit is used to determine the target time period in which each sampling moment of the audio to be adjusted falls, the target time period being determined by the sound source switching process, which refers to the switching between the sound source of interest and the non-sound source of interest; and to obtain the calculation result of the adjustment function corresponding to the target time period in which each sampling moment falls, as the adjustment parameter corresponding to each sampling moment, wherein the value of an independent variable in the adjustment function is determined based on the sampling moment.
[0023] In one embodiment, the adjustment function includes an adjustment threshold, which is determined by the distance between the corresponding sound source and the virtual operation object in the virtual space. The virtual operation object is the mapping object in the virtual space of the target object that triggers the interactive operation.
[0024] In one embodiment, if the audio to be adjusted is a target attention audio, the target time period is one of the following: a focus-in period, a continuous attention period, and an exit attention period; wherein, the start time of the focus-in period is determined based on the time when the sound source is identified as a attention sound source, the start time of the continuous attention period is determined based on the end time of the focus-in period, and the start time of the exit attention period is determined based on the time when the attention sound source is switched to a non-attention sound source, or based on an unfocus command.
[0025] In one embodiment, the adjustment function corresponding to the period of interest is a monotonically increasing function, the adjustment function corresponding to the period of interest is a constant function, and the adjustment function corresponding to the period of interest is a monotonically decreasing function.
[0026] In one embodiment, the gradients of both the monotonically increasing and monotonically decreasing functions gradually decrease as the independent variable increases.
[0027] In one embodiment, if the audio to be adjusted is a target non-interested audio, the target time period is one of entering a non-interested time period, continuing a non-interested time period, and exiting a non-interested time period; wherein, the start time of entering a non-interested time period is determined based on the time when the sound source is identified as a non-interested sound source, the start time of continuing a non-interested time period is determined based on the end time of entering a non-interested time period, and the start time of exiting a non-interested time period is determined based on the time when the non-interested sound source is switched to an interested sound source, or based on an interested command.
[0028] In one embodiment, the adjustment function corresponding to entering a non-focused time period is a monotonically decreasing function, the adjustment function corresponding to continuing a non-focused time period is a constant function, and the adjustment function corresponding to exiting a non-focused time period is a monotonically increasing function.
[0029] In one embodiment, the third determining unit is configured to, when the target audio of interest is the audio to be adjusted, use the audio adjustment result corresponding to the target audio of interest as the first intermediate audio of the sound source of interest; otherwise, directly use the target audio of interest as the first intermediate audio; if the target non-interest audio is the audio to be adjusted, use the audio adjustment result corresponding to the target non-interest audio as the second intermediate audio of the non-interest sound source; otherwise, directly use the target non-interest audio as the second intermediate audio.
[0030] In one embodiment, each sound source in the virtual space corresponds to multiple channels; the device also includes a stereo reconstruction module; the stereo reconstruction module is used to obtain the mixing processing results corresponding to each of the multiple channels, perform stereo reconstruction based on the mixing processing results of the multiple channels, and output the reconstructed stereo in the virtual space.
[0031] In one embodiment, the interaction is captured by an interaction device, and the interaction type includes at least one of sensory pointing type, consciousness pointing type, and limb pointing type.
[0032] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0033] In response to interactive operations in the virtual space, identify the sound sources of interest in the virtual space and treat all sound sources in the virtual space other than the sound sources of interest as non-interested sound sources.
[0034] Acquire the target audio corresponding to the sound source of interest, and the target non-interested audio corresponding to the sound source of non-interest.
[0035] Audio adjustment processing is performed on at least one of the target attention audio and the target non-attention audio to obtain a first intermediate audio of the attention source and a second intermediate audio of the non-attention source, wherein the sound effect discrimination of the first intermediate audio is greater than that of the second intermediate audio.
[0036] The first intermediate audio and the second intermediate audio are mixed to obtain the mixing result.
[0037] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, performs the following steps:
[0038] In response to interactive operations in the virtual space, identify the sound sources of interest in the virtual space and treat all sound sources in the virtual space other than the sound sources of interest as non-interested sound sources.
[0039] Acquire the target audio corresponding to the sound source of interest, and the target non-interested audio corresponding to the sound source of non-interest.
[0040] Audio adjustment processing is performed on at least one of the target attention audio and the target non-attention audio to obtain a first intermediate audio of the attention source and a second intermediate audio of the non-attention source, wherein the sound effect discrimination of the first intermediate audio is greater than that of the second intermediate audio.
[0041] The first intermediate audio and the second intermediate audio are mixed to obtain the mixing result.
[0042] A computer program product or computer program includes computer instructions stored in a computer-readable storage medium, wherein a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the following steps:
[0043] In response to interactive operations in the virtual space, identify the sound sources of interest in the virtual space and treat all sound sources in the virtual space other than the sound sources of interest as non-interested sound sources.
[0044] Acquire the target audio corresponding to the sound source of interest, and the target non-interested audio corresponding to the sound source of non-interest.
[0045] Audio adjustment processing is performed on at least one of the target attention audio and the target non-attention audio to obtain a first intermediate audio of the attention source and a second intermediate audio of the non-attention source, wherein the sound effect discrimination of the first intermediate audio is greater than that of the second intermediate audio.
[0046] The first intermediate audio and the second intermediate audio are mixed to obtain the mixing result.
[0047] The aforementioned audio processing method, apparatus, computer device, storage medium, and computer program based on virtual space can perform audio adjustment processing on at least one of the audio sources corresponding to both interested and uninterested sound sources in the virtual space. This results in a greater sound effect differentiation between the audio sources corresponding to the interested sound sources and the uninterested sound sources, where the interested sound sources are precisely the sounds that the user is interested in, thus making it easier for the user to hear the sounds they are interested in clearly. Furthermore, even in noisy environments, the user can effectively distinguish the sounds that interest them. Attached Figure Description
[0048] Figure 1 This is an application environment diagram of the flowchart of stereo reconstruction methods in related technologies.
[0049] Figure 2 This is an application environment diagram of an audio processing method based on virtual space in one embodiment;
[0050] Figure 3This is a flowchart illustrating an audio processing method based on virtual space in one embodiment;
[0051] Figure 4 This is a schematic diagram of the simulated field of view area covered in the virtual space in one embodiment;
[0052] Figure 5 This is a schematic diagram of a sound source located within the region of interest in one embodiment;
[0053] Figure 6 This is a schematic diagram of a finger pointing action pointing to a virtual object in virtual space in one embodiment;
[0054] Figure 7 This is a schematic diagram illustrating a virtual object passively acting as a source of attention in one embodiment;
[0055] Figure 8 This is a schematic diagram illustrating the determination of sound sources of interest and non-interest in a virtual space in one embodiment;
[0056] Figure 9 This is a schematic diagram of the adjustment function corresponding to the target audio in one embodiment;
[0057] Figure 10 This is a schematic diagram of the adjustment function corresponding to the target non-interested audio in one embodiment;
[0058] Figure 11 This is a flowchart illustrating an audio processing method based on virtual space in another embodiment;
[0059] Figure 12 This is a structural block diagram of an audio processing device based on virtual space in one embodiment;
[0060] Figure 13 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0061] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0062] First, a brief explanation of the terms used in the embodiments of this application will be provided:
[0063] Metaverse: A virtual world created and linked through technological means, mapping and interacting with the real world, and a digital living space with a new social system. It integrates various new technologies to create a new type of internet application and social form that blends the virtual and the real. It provides immersive experiences based on extended reality technology, generates a mirror image of the real world based on digital twin technology, and builds an economic system based on blockchain technology. It closely integrates the virtual and real worlds in terms of economic, social, and identity systems, and allows each user to produce content and edit the world.
[0064] Spatial sound effects: These are audio effects that use certain audio technology to make users hear more three-dimensional and spatially layered sounds. Through headphones or a combination of two or more speakers, they reproduce the auditory scene of the actual scene, allowing listeners to clearly identify the location, distance, and movement trajectory of different acoustic objects. They also allow listeners to feel as if they are surrounded by sound from all sides, creating an immersive auditory experience as if they are in the actual environment.
[0065] The realization of the metaverse requires the integration of artificial intelligence virtual perception technologies, encompassing audio, video, and sensory processing, to construct a computer-generated virtual space that closely approximates real-world perception. Users can then experience sensory sensations indistinguishable from the real world using hardware devices (headphones, glasses, motion-sensing equipment). Virtual spatial sound effects are a crucial component, recreating binaural sound signals from a real environment. By wearing headphones, users can perceive a 3D sound experience, such as the sounds of different people talking, laughing, and walking from various directions, the engine sounds of approaching cars, pedestrian crossing warnings, and the sounds of wind and rain.
[0066] However, generating virtual stereo is extremely computationally intensive. To recreate the real-world experience, the virtual space needs to generate and mix different sound sources from different directions using HRTF (Head Related Transfer Function) virtual stereo reconstruction technology, and then transmit the audio to the listener's ears. Since a large number of sound sources need to undergo HRTF stereo reconstruction simultaneously, it results in enormous computational overhead, posing a significant challenge to real-time audio experiences.
[0067] In related technologies, such as Figure 1 As shown, the different sound sources in the virtual space mainly undergo distance perception processing, virtual stereo reconstruction, and stereo mixing to generate the final sound signal that enters the user's ears. Since the volume of each sound source is adjusted primarily based on its physical distance from the current user, that is... Figure 1 The "distance perception processing" shown is effective. However, when the ambient sound is noisy, it is difficult for users to hear the sounds of objects they are interested in.
[0068] To address the problems existing in related technologies, this application provides an audio processing method based on virtual space, which can be applied to, for example... Figure 2 In the application environment shown, the interactive device 202 can transmit data with the computer device 204. Specifically, the computer device 204 is used to construct a virtual space, which can be presented through the interactive device 202. Users can interact with the virtual space through the interactive device. Responding to the interaction, the computer device 204 can determine the sound sources of interest in the virtual space based on the acquired interaction data, and treat the remaining sound sources as non-interesting sound sources. The computer device 204 adjusts the audio of the sound sources to provide users with different sensory experiences. This process is the spatial sound effect adjustment process, and the virtual space mentioned above can be the metaverse mentioned in the above description.
[0069] It is understood that the methods provided in this application embodiment can be applied to virtual reality applications, 3D map applications, event simulation applications, and game applications without limitation. The aforementioned interactive device 202 can be a desktop computer, laptop computer, mobile phone, tablet computer, e-book reader, MP3 (Moving Picture Experts Group Audio Layer III) player, MP4 (Moving Picture Experts Group Audio Layer IV) player, VR device such as VR glasses, etc. This interactive device 202 can install and run applications that support virtual space, such as applications that support 3D virtual environments.
[0070] It should be noted that the interactive devices mentioned above are merely illustrative. In actual implementation, users are not limited to controlling only one interactive device, and this application does not specifically limit the type or number of interactive devices. For example, users can achieve visual interaction through VR glasses and physical interaction through motion-sensing devices, that is, interactive operation can be achieved through two interactive devices.
[0071] It should also be noted that the aforementioned computer device 204 can be a terminal or a server. The terminal can be a mobile terminal, an in-vehicle terminal, or other intelligent terminal. The server can be a physical server or a cloud server. Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. Cloud technology is a general term encompassing network technology, information technology, integration technology, management platform technology, and application technology applied to the cloud computing business model. It can form resource pools, providing flexibility and convenience on demand. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to a backend system for logical processing. Data at different levels will be processed separately, and various industry data will require robust system support, which can be achieved through cloud computing.
[0072] In some embodiments, the computer device 204 described above can also be implemented as a node in a blockchain system. Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and cryptographic algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer.
[0073] It is worth noting that the user exists as an interactive trigger in the virtual space, projecting themselves onto the virtual space through virtual manipulation objects. These virtual manipulation objects refer to movable objects within the virtual space. These movable objects can be virtual characters, virtual animals, and other movable virtual objects. For example, virtual manipulation objects can be characters, animals, plants, oil drums, walls, and stones displayed in a 3D virtual environment. Optionally, virtual manipulation objects are 3D models created based on animation skeletal technology. Each virtual manipulation object has its own shape and volume in the 3D virtual environment and can occupy a portion of the space within the 3D virtual environment.
[0074] In some embodiments, the above-described terminology and implementation environment are combined, such as Figure 3 As shown, a virtual space-based audio processing method is provided, which can be applied to... Figure 2 The following steps are used as an example of computer device 204 in a virtual reality application:
[0075] Step 302: In response to the interactive operation for the virtual space, identify the sound sources of interest in the virtual space, and designate all sound sources in the virtual space other than the sound sources of interest as non-interested sound sources.
[0076] Virtual space primarily refers to two-dimensional, three-dimensional, or higher-dimensional spaces constructed by computer equipment. It can present visual elements such as scenery, light and shadow, and water, as well as auditory elements such as human voices and ambient sounds, and even tactile elements. In this step, the interactive operations are mainly used to adjust the auditory elements within the virtual space. Background sound can exist in the virtual space, as can multiple virtual objects that can emit sound, such as multiple characters and scenery, serving as different sound sources.
[0077] Understandably, for sound sources generated in virtual reality applications within virtual space, some sound sources are of interest to the user, such as the voices of certain characters when they appear, while others are of no interest, such as ambient noise. Therefore, in this step, the computer device can first identify the sound sources of interest in the virtual space. These sound sources of interest are those that the user is interested in. And the sound sources of interest to the user are determined through the interactive operations mentioned in this step.
[0078] Interactive operations can be triggered by the user through an interactive device, specifically by limb movements. It is understood that the triggering method and type of interactive operations can be related to the application scenario and the type of interactive device, or to other factors; this application does not specifically limit this. For example, if the application scenario is a user playing a motion-sensing game using a motion-sensing device, the type of interactive operation can be the user's limb movements. As another example, if the application scenario is a user playing an immersive game using VR glasses, the type of interactive operation can be the user's head movements, such as turning their head. Alternatively, if the VR glasses do not capture the user's head movements but instead capture the user's eye movements through eye position acquisition, the type of interactive operation can be the user's eye movements.
[0079] It is also understandable that after the interactive device detects the user's interactive operation in the virtual space, it can convert the interactive operation into interactive action data that can be recognized by the computer device, such as the acoustic object data of the virtual operation object in the virtual reality application. Thus, the computer device can respond based on the interactive action data, that is, execute the process in this step.
[0080] It should be noted that since the sound sources in virtual space are usually captured in real time or pre-defined by virtual reality applications, the computer device can acquire data on all sound sources in the virtual space, meaning it can know all the sound sources present in the virtual space. Therefore, after identifying the sound sources of interest, the computer device can determine which sound sources remain and designate the remaining sound sources as non-interesting sound sources. Non-interesting sound sources are those that the user is not interested in. Of course, in actual implementation, the non-interesting sound sources in the virtual space may not necessarily be determined using the above elimination method; for example, background noise can be pre-defined as a non-interesting sound source. This application embodiment does not specifically limit this. It should also be noted that after identifying the sound sources of interest and non-interesting sound sources in the virtual space, the computer device can set labels for these two types of sound sources to facilitate subsequent processing.
[0081] Step 304: Obtain the target audio corresponding to the source of interest and the target non-interest audio corresponding to the source of non-interest.
[0082] It is understandable that the sound sources mentioned in the above steps are primarily from the perspective of human hearing. The sound source actually corresponds to virtual objects in the virtual space, and these virtual objects are typically represented by acoustic object data in virtual reality applications. The previous steps distinguish which virtual objects are objects of interest and which are not, while this step mainly involves acquiring the audio emitted by the virtual objects as sound sources. As explained in the above steps, sound sources in the virtual space are usually captured in real-time or pre-set by the virtual reality application; therefore, the audio acquired in this step can be real-time captured audio data or pre-recorded audio data.
[0083] Furthermore, the audio acquired in this step can be time-related. If the sound source in the virtual space is typically captured in real-time at the current moment, then the audio acquired in this step is the real-time audio data captured at that moment. If the sound source in the virtual space is pre-set by the virtual reality application, then the audio acquired in this step is the audio data output when the pre-set audio is played sequentially up to the current moment.
[0084] It should be noted that the terms "target attention audio" and "target non-attention audio" mentioned in this step are not limited to the audio itself, but mainly serve as a distinction between two different sources of sound. In actual implementation, each attention sound source can correspond to one target attention audio, and each non-attention sound source can correspond to one target non-attention audio.
[0085] Step 306: Perform audio adjustment processing on at least one of the target attention audio and the target non-attention audio to obtain a first intermediate audio of the attention source and a second intermediate audio of the non-attention source, wherein the sound effect differentiation of the first intermediate audio is greater than that of the second intermediate audio.
[0086] In this step, the computer device can perform audio adjustment processing only on the target audio of interest, only on the target non-interest audio, or simultaneously on both. The audio adjustment processing methods can include volume adjustment, code rate adjustment, and signal filtering, etc., which are not specifically limited in this embodiment. It is understood that there will be multiple sound sources, i.e., multiple audio sources, in the virtual space of this embodiment. The sound effect distinguishability of a particular audio mainly refers to the degree of difficulty in distinguishing that audio from all other audio sources within the range of human hearing.
[0087] The greater the sound effect differentiation, the easier it is for the human ear to distinguish the audio from all other audio sources. Conversely, a lower differentiation indicates that the audio is mixed in with all other audio sources and is more difficult for the human ear to distinguish. Therefore, "the sound effect differentiation of the first intermediate audio is greater than that of the second intermediate audio" can be specifically expressed as the sound effect differentiation of the source of interest caused by the combined processing results of the first and second signals among all other audio sources being greater than that caused by the combined processing results of the target source of interest signal and the target non-source of interest signal.
[0088] Step 308: Mix the first intermediate audio and the second intermediate audio to obtain the mixing result.
[0089] Understandably, for sounds from multiple sources in a virtual space, it is usually necessary to integrate them into a stereo track or mono track to output sound to the user. Therefore, in this step, the computer device can individually adjust the frequency, dynamics, timbre, localization, reverberation, and sound field of the first and second intermediate audio audios, outputting the mixing result to optimize the audio from each source.
[0090] In the above embodiments, audio adjustment processing can be applied to at least one of the audio sources corresponding to the focus sound source and the non-focus sound source in the virtual space, so that the sound effect differentiation of the audio corresponding to the focus sound source is greater than that of the audio corresponding to the non-focus sound source. Since the focus sound source is precisely the sound that the user is interested in, it makes it easier for the user to hear the sound that interests them. In addition, even in noisy environments, the user can effectively distinguish the sound that interests them.
[0091] Taking a virtual space as a three-dimensional space as an example, it can be understood that, apart from background noise, sounds in a virtual space usually originate from a certain direction, and this direction typically corresponds to a region. In other words, a sound source in a virtual space usually corresponds to a region. For example, if a person is speaking in a virtual space, that person, due to their shape and volume, will occupy a portion of the virtual space, and this portion of space can be considered the region corresponding to that person as a sound source. Since a sound source in a virtual space usually corresponds to a region, users can perform interactive operations based on the region corresponding to the sound source to select that sound source. Based on the above description, in some embodiments, in response to an interactive operation in the virtual space, determining a sound source of interest in the virtual space includes: in response to an interactive operation in the virtual space, determining a region of interest in the virtual space to which the interactive operation is directed; and selecting a sound source located within the region of interest as the sound source of interest.
[0092] The triggering methods and types of interactive operations can be referred to the processes in the above embodiments, and will not be repeated here. As can be seen from the above embodiments, the triggering methods and types of interactive operations can be associated with the application scenario and the type of interactive device. It is also understood that the determination of the area of interest can be associated with the triggering methods and types of interactive operations. For ease of understanding, the process of determining the area of interest and the source of interest will be explained below with reference to examples. It is understood that in actual implementation, with the enrichment of application scenarios, the expansion of interactive device types, and the upgrading of interactive operation triggering methods, there can be multiple derivative embodiments, all of which should be optional embodiments under the conceptual premise proposed in the embodiments of this application.
[0093] Taking an application scenario where a user plays an immersive game using VR glasses, the type of interactive device is VR glasses, and the triggering method for the interaction is a physical action, specifically a head-turning action, as an example, the area of focus can be determined by using the simulated field of vision covered by the user's VR glasses in the virtual space as the area of focus. Specifically, it can be as follows: Figure 4 As shown. In this example, the sound source corresponding to the virtual object in the simulated field of view area can be taken as the sound source of interest. Furthermore, since the location range information of the simulated field of view area is known, and the location information of the virtual objects in the virtual space is also known, it is possible to determine which virtual objects are located within the simulated field of view area by comparing their locations. The location information can be described using a latitude and longitude coordinate system, using longitude and latitude ranges, such as 30-45 degrees east longitude and 0-60 degrees north latitude; this embodiment of the application does not specifically limit this. Figure 5As shown, the sound sources located within the area of interest, such as cats, dogs, and dog walkers, are the sources of attention. It should be noted that the type of interactive device used in this example can be VR glasses or a VR headset; this embodiment does not specifically limit the type of device used.
[0094] Taking an application scenario where a user plays an immersive game using a control device, where the type of interactive device is a control device, and the triggering method for the interactive operation is a control action, specifically a finger pointing action, as an example, the method for determining the area of interest can be that the computer device determines the virtual object selected by the user in the virtual space through the finger pointing action, and the portion of the space occupied by that virtual object in the virtual space is taken as the area of interest. Specifically, since the position information of the portion of the space occupied by each virtual object in the virtual space is known, and the position information of the user's hand pointing can also be known, the method of position comparison can determine which virtual object the user selected in the virtual space through the finger pointing action, and the portion of the space occupied by that virtual object in the virtual space is taken as the area of interest. Specifically, it can be done as follows: Figure 6 As shown. The user points to... Figure 6 In a virtual space, the space occupied by a male virtual object can be considered as the region of interest. Since the male virtual object, as a sound source, is also located within this region of interest, the sound source corresponding to the male virtual object can be considered the sound source of interest. It should be noted that the pointing action in this example can be a real finger click, a virtual finger click similar to a mouse pointer, or a pointing action on the limbs; this embodiment does not specifically limit this. Therefore, the type of interactive device used in this example, in addition to being a control device, can further be a touch device, a game controller, or a visual motion capture device; this embodiment does not specifically limit this.
[0095] It should be noted that the process of identifying the source of attention in the examples mentioned above was all actively triggered by the user. Understandably, in actual implementation, there are also instances where the identification of the source of attention is passively triggered. Taking the example of the head-turning action, a user can identify a simulated field of vision area after each head-turning action. If the user stops moving after a certain head-turning action, the simulated field of vision area, i.e., the area of attention, is fixed. However, it is understandable that virtual objects in virtual space are usually not fixed but move. If a virtual object actively moves into the area of attention, then the sound source corresponding to that virtual object can also be considered a source of attention. For example... Figure 7 As shown in the figure, the male virtual object on the left was not originally in the attention area, but it actively moved into the attention area. Therefore, the sound source corresponding to the male virtual object can also be regarded as the attention sound source.
[0096] In the above embodiments, since users only need to perform interactive operations on the virtual space, they can determine the area of interest that they may be interested in, and automatically select the sound source in the area of interest as the sound source of interest, thereby making it convenient for users to identify the sound they are interested in.
[0097] like Figure 8 As shown, if all sound sources within the area of interest are considered as sources of interest, there will be three sources of interest. If all sound sources other than those of interest are considered as sources of non-interest, there will be two sources of non-interest. That is, in actual implementation, there may be multiple sources of both interest and non-interest. Figure 3 In the corresponding embodiments, each source of interest can correspond to a target audio source of interest, and each source of non-interested sound source can correspond to a target non-interested audio source. It is understood that, under the above premise, there can be multiple sources of interest requiring audio adjustment processing, multiple sources of non-interested sound sources requiring audio adjustment processing, or multiple sources of both interest and non-interested sound sources requiring audio adjustment processing.
[0098] Based on this, this application also provides an embodiment of performing audio mixing processing before audio adjustment processing. In some embodiments, obtaining target audio of interest corresponding to the sound source of interest and target non-interest audio corresponding to the sound source of non-interest includes: performing audio mixing processing on the audio of interest emitted by each of the multiple sound sources of interest to obtain target audio of interest; and performing audio mixing processing on the non-interest audio emitted by each of the multiple non-interest sound sources to obtain target non-interest audio.
[0099] It is understood that virtual spaces typically contain multiple sources of interest and multiple sources of non-interest. The term "multiple sources of interest" mentioned in this application's embodiments can refer to all sources of interest, or it can refer to some of the sources of interest among all sources of interest; this application's embodiments do not specifically limit this. The term "multiple sources of non-interest" mentioned in this application's embodiments can also be understood with reference to the above explanation. Furthermore, the term "mixing processing" mentioned in this application's embodiments can refer to integrating the audio from multiple sources into a single stereo track or mono track.
[0100] The target audio source obtained after mixing multiple audio sources can be denoted as: The target non-interest audio obtained after mixing multiple non-interested sound sources can be denoted as: Regardless of the sound source, the mixing process can employ one of the following algorithms: direct summation, averaging, clamping, normalization, adaptive weighted mixing, or automatic alignment. Taking the direct summation method for mixing multiple sound sources of interest as an example, the above mixing process can be referenced by the following formula (1):
[0101] (1)
[0102] in, This represents the mixing result of K sound sources of interest at the i-th sampling time. j represents the j-th sound source of interest. This represents the audio sample value of the j-th source of interest at the i-th sampling time.
[0103] In the above embodiments, since a target audio of interest that integrates multiple sources of interest and a target audio of non-interest that integrates multiple sources of non-interest can be formed, audio adjustment processing can be performed on at least one of the integrated target audio of interest and target audio of non-interest, without having to perform audio adjustment processing on the audio of each source individually, thereby reducing the workload of audio adjustment processing.
[0104] As can be seen from the content of the above embodiments, Figure 3 In the corresponding embodiment, the purpose of the audio adjustment processing is primarily to ensure that the sound effect differentiation of the source of interest among all sound sources, jointly caused by the intermediate processing results of the first and second signals, is greater than that jointly caused by the target source of interest signal and the target non-source of interest signal. To achieve this goal, audio adjustment processing can be performed only on the target audio corresponding to the source of interest, only on the target non-source of interest audio corresponding to the non-source of interest, or both can be performed simultaneously.
[0105] In some embodiments, for methods that simultaneously perform audio adjustment processing on both the target audio source and the target non-target audio source, audio adjustment processing is performed on at least one of the audio sources to obtain a first intermediate audio source of interest and a second intermediate audio source of non-target audio source. This includes: performing signal enhancement processing on the target audio source of interest to obtain a first intermediate audio source corresponding to the target audio source of interest; and performing signal attenuation processing on the target non-target audio source to obtain a second intermediate audio source corresponding to the non-target audio source of interest.
[0106] The signal enhancement and signal attenuation processing methods can include at least one of Wiener filtering, subspace enhancement algorithm, spectral subtraction, adaptive cancellation, hidden Markov model, short-time amplitude spectrum estimation, or wavelet transform. This application does not specifically limit the methods used. When a computer device enhances the target audio signal, the resulting increase in volume or clarity of the target sound source can be observed. Conversely, when a computer device attenuates the target non-target audio signal, the resulting decrease in volume or clarity of the non-target sound source can be observed. This application does not specifically limit the manifestation of the sound changes after signal enhancement and signal attenuation processing.
[0107] In the above embodiments, since signal enhancement processing can be performed on the target audio of interest and signal attenuation processing can be performed on the target non-interest audio at the same time, adjustments can be made in two directions to make the sound emitted by the sound source of interest achieve higher sound effect differentiation among all sound sources, making it easier for users to hear the sound they are interested in.
[0108] It is understandable that in actual implementation, audio is usually adjusted based on adjustment parameters. Based on this description, in some embodiments, audio adjustment processing is performed on at least one of the target audio of interest and the target non-interest audio to obtain a first intermediate audio of the sound source of interest and a second intermediate audio of the non-interest sound source, including: determining the audio to be adjusted from the target audio of interest and the target non-interest audio; determining the adjustment parameters corresponding to each sampling time of the audio to be adjusted; performing audio adjustment processing on the audio to be adjusted based on the adjustment parameters corresponding to each sampling time; and determining the first intermediate audio of the sound source of interest and the second intermediate audio of the non-interest sound source based on the audio adjustment result.
[0109] Specifically, only the target audio of interest can be used as the audio to be adjusted, or only the target non-interested audio can be used as the audio to be adjusted, or both target audio of interest and target non-interested audio can be used as the audio to be adjusted simultaneously. This application does not impose specific limitations on this method. Sound is actually an energy wave, and waves are infinitely smooth. A string can be considered as being composed of countless points. Since storage space is relatively limited, points on the string are usually sampled during digital encoding. Therefore, it can be understood that the audio to be adjusted corresponds to the time axis, and each sampling moment corresponds to the adjustment parameters of the sound signal.
[0110] Therefore, in this embodiment, the computer device mainly performs audio adjustment processing on the adjustment parameters corresponding to each sampling moment of the audio to be adjusted. The adjustment parameters may include at least one of the following: the amplitude of the audio waveform, the fundamental frequency of the audio, or the number of harmonics in the audio waveform; this embodiment does not specifically limit the specific parameters. The amplitude of the audio waveform corresponds to the volume, the fundamental frequency of the audio corresponds to the pitch, and the number of harmonics in the audio waveform corresponds to the timbre. The computer device can adjust the adjustment parameters corresponding to each sampling moment of the audio to be adjusted to obtain the audio adjustment result, and then determine the first intermediate audio of the sound source of interest and the second intermediate audio of the non-sound source of interest based on the audio adjustment result.
[0111] In the above embodiments, since audio adjustment processing can be performed on at least one of the target attention audio and the target non-attention audio based on the adjustment parameters corresponding to each sampling time, precise adjustment can be achieved.
[0112] In some embodiments, determining the adjustment parameters corresponding to each sampling time of the audio to be adjusted includes: determining the target time period into which each sampling time of the audio to be adjusted falls, the target time period being determined by the sound source switching process, the sound source switching process referring to the switching between the sound source of interest and the non-sound source of interest; obtaining the calculation result of the adjustment function corresponding to the target time period into which each sampling time falls, as the adjustment parameter corresponding to each sampling time, wherein the value of one independent variable in the adjustment function is determined based on the sampling time.
[0113] As can be seen from the above embodiments, the sound source in the virtual space is usually captured in real time or preset by the virtual reality application. The start time and duration of the target time period can be predetermined. For example, if the sound source in the virtual space is preset by the virtual reality application, both the start time and duration of the target time period can be preset. Furthermore, once the start time and duration are determined, the end time of the target time period can also be determined.
[0114] For clarity, let's take the example of a virtual space sound source being captured in real time. If the audio to be adjusted is a target audio source of interest, then for the current sampling moment, the computer device can first determine the target time period that the current sampling moment falls into. It's understandable that a sound source in the virtual space isn't always a source of interest or a source of non-interest; its state changes. For example, a sound source might be a source of interest at one moment and switch to a source of non-interest at the next. Therefore, since the target audio source of interest is the audio corresponding to the source of interest, and the target time period is determined by the sound source switching process, the computer device can determine the start time of the target time period, which can be related to the moment when the source of interest last switched to "source of interest." The end time of the target time period can be related to the moment after the current sampling moment when the source of interest last switched out of "source of non-interest."
[0115] For example, taking sound source A as an example, the current sampling time is 11:30 AM on December 21, 2021. At the current sampling time, the most recent time A became a sound source of interest was at 11:25 AM on December 21, 2021; before that, it was a non-interested sound source. Therefore, the start time of the target time period can be related to 11:25 AM on December 21, 2021, and can be directly determined as that time. If, at 11:35 AM on December 21, 2021, after the current sampling time, A becomes a non-interested sound source, then the end time of the target time period can also be related to 11:35 AM on December 21, 2021, and can also be directly determined as that time.
[0116] Furthermore, if the audio to be adjusted is a target non-interested audio, the computer device can first determine the target time period that the current sampling time falls into. In this case, the target non-interested audio is the audio corresponding to a non-interested sound source. The determination of the start and end times of the target time period can be found in the above embodiments. For example, the start time of the target time period can be related to the moment when the non-interested sound source last switched to being a "non-interested sound source." The end time of the target time period can be related to the moment when the interested sound source last switched out of being a "interested sound source" after the current sampling time.
[0117] As can be seen from the above embodiments, if the audio to be adjusted is a target audio of interest, then the target time period can be understood as a complete lifecycle of the sound source as a target audio source of interest. If the audio to be adjusted is a target non-target audio of interest, then the target time period can be understood as a complete lifecycle of the sound source as a non-target audio source of interest. It should be noted that the audio to be adjusted may include both target audio of interest and target non-target audio of interest. Therefore, for the current sampling time, audio adjustment processing needs to be performed on both types of audio at the same time. In this case, although the time periods may overlap, the target audio of interest and the target non-target audio of interest can each correspond to a target time period, and both perform their own audio adjustment processing based on their respective adjustment functions within their respective target time periods.
[0118] It should also be noted that, in the above content, the target time period corresponding to the audio source of interest is defined based on a complete lifecycle of the sound source as a source of interest; that is, a complete lifecycle of the sound source as a source of interest corresponds to only one target time period. Similarly, a complete lifecycle of the sound source as a source of non-interest also corresponds to only one target time period. However, in actual implementation, a complete lifecycle of the sound source as a source of interest can be divided into multiple time periods. The time period in which the current sampling moment falls can be considered the target time period. Likewise, a complete lifecycle of the sound source as a source of non-interest can also be divided into multiple time periods, and the time period in which the current sampling moment falls can be considered the target time period.
[0119] In this embodiment, a target time period corresponds to an adjustment function. It is understood that different sampling times can correspond to different calculation results of the adjustment function. Therefore, at least one independent variable in the adjustment function can be associated with the value of the sampling time. The computer device substitutes the current sampling time into the adjustment function corresponding to the target time period into which the current sampling time falls, and the calculated result can be used as the adjustment parameter corresponding to the current sampling time.
[0120] In the above embodiments, since the adjustment parameters corresponding to the sampling time can be determined based on the adjustment function corresponding to the target time period in which the sampling time of the audio to be adjusted falls, different adjustment parameters can be applied to different times, thereby achieving precise adjustment.
[0121] It is understandable that the distance between the listener and the sound source affects the perceived sound quality, such as clarity. Therefore, in practical implementation, the distance between the listener and the sound source can be used as a reference when determining the adjustment parameters. In this embodiment, in the real space, the "listener" can be the target object triggering the interactive operation, i.e., the user using the interactive device. In the virtual space, the user using the interactive device is typically mapped to a mapping object, which is the virtual operation object. Based on this description, in some embodiments, the adjustment function includes an adjustment threshold, which is determined by the distance between the corresponding sound source and the virtual operation object in the virtual space. The virtual operation object is the mapping object of the target object triggering the interactive operation in the virtual space.
[0122] As can be seen from the above embodiments, the adjustment function can contain an independent variable corresponding to the sampling time. In this embodiment, in addition to this independent variable, the adjustment function can also include another independent variable determined based on the distance between the virtual operation object and the sound source; this independent variable is the adjustment threshold. As can be seen from the above embodiments, "target attention audio and target non-attention audio can each correspond to a target time period, and both perform their own audio adjustment processing based on their respective adjustment functions within their respective target time periods," and each target attention audio and target non-attention audio corresponds to a sound source. Therefore, each sound source can correspond to an adjustment function. Based on this, for any sound source, the adjustment threshold included in the adjustment function corresponding to that sound source can be determined by the distance between the sound source and the virtual operation object in the virtual space.
[0123] The adjustment threshold can be determined using a lookup table mapping method, which is not specifically limited in this embodiment. For example, different distance value ranges can be preset. Each distance value range corresponds to an adjustment threshold for a source of interest, and each distance value range also corresponds to an adjustment threshold for a source of non-interest. Subsequently, based on the distance value range that the actual distance falls into, and the type of sound source corresponding to the actual distance (whether it is a source of interest or a source of non-interest), the adjustment threshold corresponding to the actual distance can be determined.
[0124] In the above embodiments, the distance between the listener and the sound source can affect the value of the adjustment parameters, thus affecting the sound effect produced by the sound source. Therefore, the audio emitted by the sound source can be adjusted based on the distance, making it easier for the user to hear the sounds they are interested in.
[0125] As can be seen from the above embodiments, a complete lifecycle of a sound source as a source of interest can correspond to multiple time periods. A complete lifecycle of a sound source of interest typically refers to the period from when it is identified as a source of interest to when it is no longer identified as a source of interest. In practice, this period can be divided into three time periods: the period of initial interest identification, the period of continuous interest identification, and the period of withdrawal from interest identification. The initial interest identification period refers to the initial stage when the sound source is first identified as a source of interest, the period of continuous interest identification refers to the stable stage after the sound source is identified as a source of interest, and the period of withdrawal from interest identification refers to the initial stage when the sound source is first identified as a source of interest, and also the final stage when the sound source is no longer identified as a source of interest.
[0126] Based on the above description, in some embodiments, if the audio to be adjusted is the target audio to be focused on, the target time period is one of the following: a focus-in period, a continuous focus period, and a focus-out period. The start time of the focus-in period is determined based on the time when the sound source is identified as a focus source, the start time of the continuous focus period is determined based on the end time of the focus-in period, and the start time of the focus-out period is determined based on the time when the focus source is switched to a non-focus source, or based on a cancel focus command.
[0127] The start time of the period of focus can be the moment when the sound source is identified as a source of interest, or it can be a moment before or after that moment, such as a moment a few seconds later. Similarly, the start time of the period of continuous focus can be the end time of the period of focus, or it can be a moment before or after that moment. The start time of the period of exiting focus can be the moment when the sound source of interest is switched to a non-interested sound source, or it can be a moment before or after that moment. In addition, the end time of the period of continuous focus can be the moment when the sound source of interest is switched to a non-interested sound source, or it can be a moment before or after that moment; this embodiment of the application does not limit this.
[0128] In the above description, no specific basis is given for the end time of entering and exiting the attention period. In actual implementation, the end time of each period can be determined by setting its own duration; this application embodiment does not specifically limit this. It is understood that "the attention source is switched to a non-attention source" emphasizes the passive switching process of the sound source. For example, the attention source is usually movable; when the attention source leaves the attention area, it will be switched to a non-attention source. In actual implementation, the attention source can also be actively switched to a non-attention source through command triggering. Therefore, in this application embodiment, the start time of exiting the attention period can also be determined based on the unfollow command.
[0129] In the above embodiments, since a complete life cycle of a sound source as a source of interest can correspond to multiple time periods, the corresponding adjustment parameters can be determined based on the different time periods in which the sampling time falls. Subsequently, audio adjustment processing is performed based on the adjustment parameters, so that audio adjustment can be carried out in stages for a complete life cycle of the sound source of interest, thereby making the adjustment process more refined and producing better sound effects.
[0130] In some embodiments, the adjustment function corresponding to the period of initial attention is a monotonically increasing function, the adjustment function corresponding to the period of continuous attention is a constant function, and the adjustment function corresponding to the period of exiting attention is a monotonically decreasing function.
[0131] Among them, a monotonically increasing function refers to a function whose calculated result, or adjustment parameter, gradually increases over time. A constant function refers to a function whose calculated result, or adjustment parameter, remains unchanged at a certain value over time. A monotonically decreasing function refers to a function whose calculated result, or adjustment parameter, gradually decreases over time.
[0132] Understandably, the adjustment function for the focus entry period is set as a monotonically increasing function because the focus entry period represents the initial stage when the sound source is first identified as the focus source. As time progresses, the value of the adjustment parameter increases, which aligns with the expectation of a gradual transition as the focus source is integrated into the auditory world. Similarly, the adjustment function for the sustained focus period is set as a constant function because the sustained focus period refers to the stable stage after the sound source is identified as the focus source. As time progresses, the value of the adjustment parameter remains unchanged, which also aligns with the expectation of a stable transition as the focus source is integrated into the auditory world. Finally, the adjustment function for the exit focus period is set as a monotonically decreasing function because the exit focus period refers to the final stage when the sound source is no longer the focus source. As time progresses, the value of the adjustment parameter decreases, which also aligns with the expectation of a gradual transition as the focus source fades from the auditory world.
[0133] In the above embodiments, since the adjustment parameters can gradually increase, remain unchanged, or gradually decrease over time, the gradual change process can conform to the switching process between the focus sound source and the non-focus sound source, thereby making the adjustment process more precise and producing better sound effects.
[0134] In some embodiments, the gradients of both the monotonically increasing and monotonically decreasing functions gradually decrease as the independent variable increases.
[0135] In this application, the independent variables mentioned in the embodiments refer to those determined based on the sampling time. It can be understood that the initial attention period is the initial stage when the sound source is first identified as the source of attention. As time progresses, the value of the adjustment parameter increases, which aligns with the expectation of gradual processing as the source of attention is gradually integrated into the auditory world. Since the period following the initial attention period is a stable, continuous attention period, a continuous increase in the adjustment parameter at the initial rate as time progresses would not align with the expectation of subsequent stabilization. A trend of gradually increasing value but with a gradually slowing rate of increase is more consistent with the expectation of subsequent stabilization. Therefore, the monotonically increasing function can be set such that its gradient gradually decreases as time progresses.
[0136] Similarly, the exit-attention period is the initial stage when a sound source is first identified as a non-attention source, and also refers to the final stage when the sound source is no longer considered a focus source. As time progresses, the value of the adjustment parameter decreases, which aligns with the expectation of a gradual fading of the focus source from the auditory world. Since the exit-attention period can be followed by a stable, continuous non-attention period, a continuous decrease in the adjustment parameter value at the initial rate does not necessarily align with the expectation of a subsequent stabilizing gradual process. A gradual decrease in value at a slower rate is more consistent with the expectation of a subsequent stabilizing gradual process. Therefore, the monotonically decreasing function can also be set to have a gradient that gradually decreases over time.
[0137] In the above embodiments, since the speed increase or decrease of the adjustment parameter can change over time to adapt to the expected gradual processing, the adjustment process is more refined and the resulting sound effect is better.
[0138] The above embodiments primarily explain, from a principle perspective, how the adjustment parameters are determined when the audio to be adjusted is the target audio of interest. The determination process of the adjustment parameters in the above embodiments will now be explained with specific examples; please refer to [link / reference]. Figure 9 .by Figure 9 Taking the time unit in the example of seconds, Figure 9The intervals from 0 to 0.5 seconds correspond to the period of initial focus, 0.5 to T corresponds to the period of continuous focus, and T to T+0.5 corresponds to the period of exiting focus. Note that the 0th and 0.5th seconds do not refer to actual times, but are used for ease of measurement. Based on the above embodiments, it can be understood that the 0th second can be the moment when the sound source is identified as a source of focus, and the Tth second can be the moment when the source of focus is switched to a source of no longer focus.
[0139] Specifically, the adjustment function corresponding to the time period of focus can be referenced as follows (2):
[0140] (2)
[0141] in, This represents the adjustment function corresponding to the time period of focus, where t represents the independent variable determined based on the current sampling time. The value is determined based on the duration of the period of focus. Corresponding to the adjustment threshold mentioned in the above embodiments, its value is determined by the distance between the corresponding sound source and the virtual operation object in the virtual space; , and Both represent constants. Furthermore, The value can be greater than 1. The value of can be 1. The value of can be 1. The value can be 0.5. In actual implementation, the value of the above constant can be set according to requirements, and this application embodiment does not impose specific limitations on it. For example Figure 9 As shown, the adjustment function corresponding to the time period of focus is an increasing function and the gradient gradually decreases as time progresses.
[0142] The adjustment function corresponding to the continuous monitoring period can be referenced from the following formula (3):
[0143] (3)
[0144] in, This indicates the adjustment function corresponding to the period of continuous monitoring. As an adjustment threshold, its value can be set according to requirements. In the embodiments of this application, The value can be related to Similarly, the embodiments in this application do not specifically limit this. For example... Figure 9 As shown, the adjustment function corresponding to the continuous monitoring period is a constant function.
[0145] The adjustment function for the time period to be exited can be found in the following formula (4):
[0146] (4)
[0147] in, This represents the adjustment function corresponding to the time period of exiting the focus, where t represents the independent variable determined based on the current sampling time; The value is determined based on the duration of the period after the user leaves the area. Corresponding to the adjustment threshold mentioned in the above embodiments, its value is determined by the distance between the corresponding sound source and the virtual operation object in the virtual space; , and Both represent constants. Furthermore, The value can be greater than 1. The value of can be 1. The value of can be 1. The value can be 0.5. In actual implementation, the value of the above constant can be set according to requirements, and this application embodiment does not impose specific limitations on it. For example Figure 9 As shown, the adjustment function corresponding to the exit focus period is a decreasing function and the gradient gradually decreases as time progresses.
[0148] The above content mainly describes the embodiments corresponding to the audio to be adjusted being a target audio source. Based on the above embodiments, it can be understood that the audio to be adjusted can also be a target non-target audio source. Therefore, in some embodiments, if the audio to be adjusted is a target non-target audio source, the target time period is one of entering a non-target time period, continuing a non-target time period, and exiting a non-target time period. The start time of entering a non-target time period is determined based on the time when the sound source is identified as a non-target sound source; the start time of continuing a non-target time period is determined based on the end time of entering a non-target time period; and the start time of exiting a non-target time period is determined based on the time when the non-target sound source is switched to a target sound source, or based on a target command.
[0149] The start time of entering a non-interested period can be the moment when the sound source is identified as a non-interested source, or it can be a moment before or after that moment, such as a moment a few seconds later. Similarly, the start time of a continuous non-interested period can be the end time of entering a non-interested period, or it can be a moment before or after that moment. The start time of exiting a non-interested period can be the moment when a non-interested sound source is switched to an interested sound source, or it can be a moment before or after that moment. Furthermore, the end time of a continuous non-interested period can be the moment when a non-interested sound source is switched to an interested sound source, or it can be a moment before or after that moment; this embodiment of the application does not limit this.
[0150] In the above description, no specific basis is given for the end time of entering and exiting the non-focused time period. In actual implementation, the end time can be determined by setting the duration of entering and exiting the non-focused time period, and this application embodiment does not specifically limit this. Similar to the above embodiment, "a non-focused sound source is switched to a focus sound source" focuses on the passive switching process of the sound source. For example, a non-focused sound source is usually movable, and when a non-focused sound source enters the focus area, it will be switched to a focus sound source. In actual implementation, a non-focused sound source can also be actively switched to a focus sound source by triggering a command. Therefore, in this application embodiment, the start time of exiting the non-focused time period can also be determined based on the focus command.
[0151] In the above embodiments, since a complete life cycle of a sound source as a non-interested sound source can correspond to multiple time periods, the corresponding adjustment parameters can be determined based on the different time periods in which the sampling time falls. Subsequently, audio adjustment processing is performed based on the adjustment parameters. Thus, for a complete life cycle of a non-interested sound source, audio adjustment can be performed in stages, thereby making the adjustment process more refined and producing better sound effects.
[0152] In some embodiments, the adjustment function corresponding to entering a non-focused time period is a monotonically decreasing function, the adjustment function corresponding to continuing a non-focused time period is a constant function, and the adjustment function corresponding to exiting a non-focused time period is a monotonically increasing function.
[0153] The definitions of monotonically increasing functions, constant functions, and monotonically increasing functions can be found in the explanations in the corresponding embodiments for the audio to be adjusted as the target audio, and will not be repeated here.
[0154] Similarly, it's understandable that the adjustment function corresponding to the transition into the non-focused period is set as a monotonically decreasing function. This is because the transition into the non-focused period refers to the final stage when the sound source is no longer considered a focus source. As time progresses, the value of the adjustment parameter decreases, which aligns with the expectation of a gradual fading of the focus source from the auditory world. The adjustment function corresponding to the continuous non-focused period is set as a constant function because this period refers to the stable stage after the sound source is identified as a non-focused source. As time progresses, the value of the adjustment parameter remains unchanged, which also aligns with the expectation of a stable process where the non-focused source has been incorporated into the auditory world. The adjustment function corresponding to the exit from the non-focused period is set as a monotonically increasing function because exiting the non-focused period is the initial stage when the sound source is first identified as a focus source. As time progresses, the value of the adjustment parameter increases, which also aligns with the expectation of a gradual fading of the focus source into the auditory world.
[0155] Furthermore, the gradients of both the monotonically increasing and monotonically decreasing functions mentioned in the embodiments of this application can gradually decrease as the independent variable increases. The independent variable mentioned here is also determined based on the sampling time. It can be understood that entering the non-interested time period is the initial stage when the sound source is first identified as a non-interested sound source. As time progresses, the value of the adjustment parameter becomes smaller and smaller, which can conform to the expectation of a gradual fading of the non-interested sound source from the auditory world. Since the non-interested time period is followed by a stable, continuous non-interested time period, and the value of the adjustment parameter continues to decrease at the initial decreasing rate as time progresses, this obviously does not conform to the expectation of a subsequent gradual stabilization. A trend of gradually decreasing value but with a gradually slowing rate of decrease is more in line with the expectation of a subsequent gradual stabilization. Therefore, the monotonically decreasing function can be set so that its gradient gradually decreases as time progresses.
[0156] Similarly, exiting the non-focused period marks the initial stage when a sound source is identified as a focus source, and also the final stage when the sound source is no longer considered a non-focused source. As time progresses, the value of the adjustment parameter increases, which aligns with the expectation of gradual processing as the focus source is gradually integrated into the auditory world. Since the period following the exit from the non-focused period can be a stable, continuous focus period, a continuous increase in the adjustment parameter value at the initial rate over time would not align with the expectation of subsequent stabilization. A trend of gradually increasing value but at a gradually slowing rate is more consistent with the expectation of subsequent stabilization. Therefore, the monotonically increasing function can also be set to have a gradient that gradually decreases over time.
[0157] In the above embodiments, since the adjustment parameter can gradually decrease, remain constant, or gradually increase over time, the gradual processing can conform to the switching process between non-interested sound sources and interested sound sources, thus making the adjustment process more precise and producing better sound effects. Furthermore, since the increase or decrease in the adjustment parameter can change over time to adapt to the expected gradual processing, the adjustment process can also be made more precise, resulting in better sound effects.
[0158] The above embodiments primarily explain, from a principle perspective, how the adjustment parameters are determined when the audio to be adjusted is a target non-interested audio. The determination process of the adjustment parameters in the above embodiments will now be explained with specific examples; please refer to [link / reference]. Figure 10 .by Figure 10 Taking the time unit in the example of seconds, Figure 10The intervals from 0 to 0.5 seconds correspond to the period of transitioning to a non-interested period, from 0.5 to T seconds correspond to the period of continuous non-interested period, and from T to T+0.5 seconds correspond to the period of exiting a non-interested period. Note that the 0th and 0.5th seconds do not refer to actual times, but are used for ease of measurement. Based on the above embodiments, it can be understood that the 0th second can be the moment when the sound source is identified as a non-interested sound source, and the Tth second can be the moment when the non-interested sound source is switched to an interested sound source.
[0159] Specifically, the adjustment function corresponding to the non-focused time period can be referenced as follows (5):
[0160] (5)
[0161] in, This represents the adjustment function corresponding to the non-interested time period, where t represents the independent variable determined based on the current sampling time; The value is determined based on the duration of the non-interested time period; Corresponding to the adjustment threshold mentioned in the above embodiments, its value is determined by the distance between the corresponding sound source and the virtual operation object in the virtual space; , and Both represent constants. Furthermore, The value of can be less than 1. The value can be related to same, The value of can be 1. The value can be 0.5. In actual implementation, the value of the above constant can be set according to requirements, and this application embodiment does not impose specific limitations on it. For example Figure 10 As shown, the adjustment function corresponding to the non-interested time period is a decreasing function and the gradient gradually decreases as time progresses.
[0162] The adjustment function for the continuous non-focus period can be referenced by the following formula (6):
[0163] (6)
[0164] in, This represents the adjustment function corresponding to the continuous non-focused period. As an adjustment threshold, its value can be set according to requirements. In the embodiments of this application, The value can be related to Similarly, the embodiments in this application do not specifically limit this. For example... Figure 10 As shown, the adjustment function corresponding to the continuous non-focus period is a constant function.
[0165] The adjustment function for exiting the non-focused time period can be referenced in the following formula (7):
[0166] (7)
[0167] in, This indicates the adjustment function corresponding to the time period to be excluded from the non-focused period, where t represents the independent variable determined based on the current sampling time; The value is determined based on the duration of the period of exiting non-focused time; Corresponding to the adjustment threshold mentioned in the above embodiments, its value is determined by the distance between the corresponding sound source and the virtual operation object in the virtual space; , and Both represent constants. Furthermore, The value of can be less than 1. The value of and They can be the same. The value of can be 1. The value can be 0.5. In actual implementation, the value of the above constant can be set according to requirements, and this application embodiment does not impose specific limitations on it. For example Figure 10 As shown, the adjustment function corresponding to the exit from the non-focused time period is an increasing function and the gradient gradually decreases as time progresses.
[0168] It should also be noted that the audio to be adjusted can simultaneously include target-interest audio and target-non-interest audio. In this case, audio adjustment processing can be performed on the target-interest audio and the target-non-interest audio separately according to the processing methods mentioned in the embodiments. Additionally, by... Figure 9 and Figure 10 It can be seen that the end time of exiting the non-focused period and the start time of entering the focused period can be connected. For example, the start time of exiting the non-focused period can be the moment when the non-focused sound source is switched to a focus sound source, and the start time of entering the focused period can be the end time of exiting the non-focused period.
[0169] In some embodiments, determining a first intermediate audio source of interest and a second intermediate audio source of non-interest based on the audio adjustment result includes: if the target audio source of interest is the audio to be adjusted, then the audio adjustment result corresponding to the target audio source of interest is used as the first intermediate audio source of interest; otherwise, the target audio source of interest is directly used as the first intermediate audio source of interest. If the target non-interest audio source of interest is the audio to be adjusted, then the audio adjustment result corresponding to the target non-interest audio source of interest is used as the second intermediate audio source of non-interest; otherwise, the target non-interest audio source of interest is directly used as the second intermediate audio source of interest.
[0170] Specifically, if the audio to be adjusted is the target audio of interest, then the target audio of interest needs to undergo audio adjustment processing, and the computer device can use the audio adjustment result corresponding to the target audio of interest as the first intermediate audio of the source of interest. Otherwise, it means that the target audio of interest does not need to undergo audio adjustment processing, and the computer device can directly use the target audio of interest as the first intermediate audio. The same applies to the target non-interest audio. In addition, if both the target audio of interest and the target non-interest audio are audio to be adjusted, then the process of obtaining the first intermediate audio and the second intermediate audio after audio adjustment processing can be referred to the following formulas (8) and (9):
[0171] (8)
[0172] (9)
[0173] in, Let represent the target audio at the i-th sampling time, and t represent the independent variable determined based on the i-th sampling time. This represents the first intermediate audio at the i-th sampling time. This represents the adjustment parameter at the i-th sampling time, assuming the audio to be adjusted is the target audio. To determine how to determine this, refer to the above formulas (2), (3) and (4). This represents the target non-interest audio at the i-th sampling time. This represents the second intermediate audio at the i-th sampling time. This represents the adjustment parameters at the i-th sampling time, assuming the audio to be adjusted is the target non-interested audio. To determine how to determine this, refer to the above formulas (5), (6) and (7).
[0174] In the above embodiments, since audio adjustment processing can be performed on at least one of the target attention audio and the target non-attention audio based on the adjustment parameters corresponding to each sampling time, precise adjustment can be achieved.
[0175] It is understandable that audio signals played in virtual space typically have multiple channels, thus requiring multi-channel sound reconstruction in actual implementation. Therefore, in some embodiments, such as... Figure 11 As shown, a virtual space-based audio processing method is provided, which can be applied to... Figure 2 The following steps are used as an example of computer device 204 in a virtual reality application:
[0176] Step 1102: In response to the interactive operation for the virtual space, identify the sound sources of interest in the virtual space, and designate all sound sources in the virtual space other than the sound sources of interest as non-interested sound sources.
[0177] Step 1104: Obtain the target audio corresponding to the source of interest and the target non-interest audio corresponding to the source of non-interest.
[0178] Step 1106: Perform audio adjustment processing on at least one of the target attention audio and the target non-attention audio to obtain a first intermediate audio of the attention source and a second intermediate audio of the non-attention source, wherein the sound effect differentiation of the first intermediate audio is greater than that of the second intermediate audio.
[0179] Step 1108: Mix the first intermediate audio and the second intermediate audio to obtain the mixing result.
[0180] For a detailed explanation of the above steps, please refer to the above. Figure 3 The content of the corresponding embodiment.
[0181] Step 1110: Obtain the mixing results corresponding to each of the multiple channels, perform stereo reconstruction based on the mixing results of the multiple channels, and output the reconstructed stereo in virtual space.
[0182] Each channel can undergo audio adjustment processing according to the process mentioned in the above embodiments to obtain the mixing processing result for each channel. Furthermore, stereo reconstruction can be implemented using HRTF virtual stereo reconstruction technology; this application does not specifically limit this implementation.
[0183] In the above embodiments, since audio adjustment processing can be performed on at least one audio source corresponding to the sound source of interest and the audio source of non-interest in each channel, the sound effect differentiation of the audio source corresponding to the sound source of interest is greater than that of the audio source of non-interest in interest. Since the sound source of interest is the sound that the user is interested in, it makes it easier for the user to hear the sound that interests them. In addition, in noisy environments, the user can also effectively distinguish the sound that interests them.
[0184] In addition to the triggering methods and types of interactive operations mentioned in the above embodiments, interactive operations may also have other triggering methods and types. Therefore, in some embodiments, interactive operations are captured by an interactive device, and the operation type of the interactive operation includes at least one of sensory pointing type, consciousness pointing type, and limb pointing type.
[0185] Among them, sensory-directed interactive operations refer to interactive operations triggered by the senses, such as those triggered by vision, as in the above embodiments. Consciousness-directed interactive operations refer to interactive operations triggered by consciousness, such as those triggered by brainwaves. Limb-directed interactive operations refer to interactive operations triggered by limb movements, such as those triggered by hand manipulation.
[0186] In the above embodiments, the interaction operations of various types are available, which enriches the ways to trigger the sound source.
[0187] This application also provides an application scenario in which the above-described virtual space-based audio processing method is applied. Specifically, the virtual space-based audio processing method is applied in this scenario as follows:
[0188] When a user wears VR glasses and connects to a host device to play an immersive game, the user's head-turning motion is mapped in the virtual reality game program on the host device as a change in the simulated field of view of a virtual object in virtual space. The host device treats sound sources located within the changed simulated field of view (i.e., the area of interest) in virtual space as focused sound sources, while treating the remaining sound sources in virtual space as unfocused sound sources.
[0189] The host device performs mixing processing on all sound sources of interest to obtain the target audio for each sound source. It then performs mixing processing on all non-sound sources of interest to obtain the target non-sound sources of interest. The host device determines the time interval within which the audio sampling time falls, determines the corresponding adjustment function based on this interval, determines the adjustment parameters based on the adjustment function, and performs signal enhancement processing on the target audio of interest based on these parameters to obtain the first intermediate audio. Similarly, the host device performs signal attenuation processing on the target non-sound sources of interest to obtain the second intermediate audio. The host device then mixes the first and second intermediate audio to obtain the mixing result. Finally, the host device performs stereo reconstruction on the mixing result for each channel and outputs stereo sound to the VR glasses. The user can then hear the stereo sound through the headphones included with the VR glasses.
[0190] This application also provides another application scenario in which the above-described virtual space-based audio processing method is applied. Specifically, the virtual space-based audio processing method is applied in this scenario as follows:
[0191] When a user connects a motion capture device to a host device to play motion-sensing games, the user's hand movements are captured by the motion capture device and mapped into a virtual object in the motion-sensing game program. For example, the user might point the virtual object at a non-player character in the game. The host device can then designate the space occupied by this virtual object in the virtual space as the area of interest, and the sound source corresponding to that virtual object within that area as the sound source of interest. In this way, the user can specify multiple sound sources of interest. The remaining sound sources in the virtual space are designated as non-interested sound sources.
[0192] The host device performs mixing processing on all sound sources of interest to obtain the target audio for each sound source. It then performs mixing processing on all non-sound sources of interest to obtain the target non-sound sources of interest. The host device determines the time period within which the audio sampling time falls, determines the corresponding adjustment function based on this time period, determines the adjustment parameters based on the adjustment function, and performs signal enhancement processing on the target audio of interest based on these parameters to obtain the first intermediate audio. Similarly, the host device performs signal attenuation processing on the target non-sound sources of interest to obtain the second intermediate audio. The host device mixes the first and second intermediate audio to obtain the mixing result. Finally, the host device performs stereo reconstruction on the mixing result for each channel and outputs stereo sound to speakers or headphones worn by the user, allowing the user to hear stereo sound.
[0193] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0194] Based on the same inventive concept, this application also provides a virtual space-based audio processing apparatus for implementing the aforementioned virtual space-based audio processing method. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more virtual space-based audio processing apparatus embodiments provided below can be found in the limitations of the virtual space-based audio processing method described above, and will not be repeated here.
[0195] In some embodiments, such as Figure 12 As shown, a virtual space-based audio processing device 1200 is provided. This device can be a software module, a hardware module, or a combination of both as part of a computer device. Specifically, the device includes: a determining module 1202, an acquiring module 1204, an audio adjustment processing module 1206, and a mixing processing module 1208, wherein:
[0196] The determination module 1202 is used to respond to interactive operations on the virtual space, determine the sound sources of interest in the virtual space, and treat all sound sources in the virtual space other than the sound sources of interest as non-sound sources of interest.
[0197] The acquisition module 1204 is used to acquire the target audio of interest corresponding to the sound source of interest, and the target non-interest audio corresponding to the non-interest sound source;
[0198] The audio adjustment processing module 1206 is used to perform audio adjustment processing on at least one of the target attention audio and the target non-attention audio to obtain a first intermediate audio of the attention sound source and a second intermediate audio of the non-attention sound source, wherein the sound effect differentiation of the first intermediate audio is greater than that of the second intermediate audio.
[0199] The mixing module 1208 is used to mix the first intermediate audio and the second intermediate audio to obtain the mixing result.
[0200] In some embodiments, the determining module 1202 is configured to, in response to an interactive operation on the virtual space, determine the region of interest in the virtual space to which the interactive operation is directed; and designate a sound source in the region of interest as a sound source of interest.
[0201] In some embodiments, the acquisition module 1204 is used to perform mixing processing on the audio of interest emitted by each of multiple sound sources of interest to obtain target audio of interest; and to perform mixing processing on the audio of non-interest emitted by each of multiple sound sources of non-interest to obtain target audio of non-interest.
[0202] In some embodiments, the audio adjustment processing module 1206 is used to perform signal enhancement processing on the target audio of interest to obtain a first intermediate audio corresponding to the sound source of interest; and to perform signal attenuation processing on the target non-interest audio to obtain a second intermediate audio corresponding to the non-interest sound source.
[0203] In some embodiments, the audio adjustment processing module 1206 includes:
[0204] The first determining unit is used to determine the audio to be adjusted from the target attention audio and the target non-attention audio;
[0205] The second determining unit is used to determine the adjustment parameters corresponding to each sampling time of the audio to be adjusted;
[0206] The third determining unit is used to perform audio adjustment processing on the audio to be adjusted based on the adjustment parameters corresponding to each sampling time, and to determine the first intermediate audio of the source of interest and the second intermediate audio of the source of non-interest based on the audio adjustment results.
[0207] In some embodiments, the second determining unit is used to determine the target time period in which each sampling time of the audio to be adjusted falls, the target time period being determined by the sound source switching process, the sound source switching process referring to the switching between the sound source of interest and the non-sound source of interest; and to obtain the calculation result of the adjustment function corresponding to the target time period in which each sampling time falls, as the adjustment parameter corresponding to each sampling time, wherein the value of an independent variable in the adjustment function is determined based on the sampling time.
[0208] In some embodiments, the adjustment function includes an adjustment threshold, which is determined by the distance between the corresponding sound source and the virtual operation object in the virtual space. The virtual operation object is the mapping object in the virtual space of the target object that triggers the interactive operation.
[0209] In some embodiments, if the audio to be adjusted is a target attention audio, the target time period is one of the following: a focus-in period, a continuous attention period, and an exit attention period; wherein, the start time of the focus-in period is determined based on the time when the sound source is identified as a attention sound source, the start time of the continuous attention period is determined based on the end time of the focus-in period, and the start time of the exit attention period is determined based on the time when the attention sound source is switched to a non-attention sound source, or based on an unfocus instruction.
[0210] In some embodiments, the adjustment function corresponding to the period of initial attention is a monotonically increasing function, the adjustment function corresponding to the period of continuous attention is a constant function, and the adjustment function corresponding to the period of exiting attention is a monotonically decreasing function.
[0211] In some embodiments, the gradients of both the monotonically increasing and monotonically decreasing functions gradually decrease as the independent variable increases.
[0212] In some embodiments, if the audio to be adjusted is a target non-interested audio, the target time period is one of entering a non-interested time period, continuing a non-interested time period, and exiting a non-interested time period; wherein, the start time of entering a non-interested time period is determined based on the time when the sound source is identified as a non-interested sound source, the start time of continuing a non-interested time period is determined based on the end time of entering a non-interested time period, and the start time of exiting a non-interested time period is determined based on the time when the non-interested sound source is switched to an interested sound source, or based on an interested command.
[0213] In some embodiments, the adjustment function corresponding to entering a non-focused time period is a monotonically decreasing function, the adjustment function corresponding to continuing a non-focused time period is a constant function, and the adjustment function corresponding to exiting a non-focused time period is a monotonically increasing function.
[0214] In some embodiments, the third determining unit is configured to, when the target audio of interest is the audio to be adjusted, use the audio adjustment result corresponding to the target audio of interest as the first intermediate audio of the sound source of interest, otherwise use the target audio of interest directly as the first intermediate audio; if the target non-interest audio is the audio to be adjusted, use the audio adjustment result corresponding to the target non-interest audio as the second intermediate audio of the non-interest sound source, otherwise use the target non-interest audio directly as the second intermediate audio.
[0215] In some embodiments, each sound source in the virtual space corresponds to multiple channels; the device also includes a stereo reconstruction module; the stereo reconstruction module is used to acquire the mixing processing results corresponding to each of the multiple channels, perform stereo reconstruction based on the mixing processing results of the multiple channels, and output the reconstructed stereo in the virtual space.
[0216] In some embodiments, the interactive operation is captured by an interactive device, and the operation type of the interactive operation includes at least one of sensory pointing type, consciousness pointing type, and limb pointing type.
[0217] For specific limitations on the virtual space-based audio processing device, please refer to the limitations on the virtual space-based audio processing method above, which will not be repeated here. Each module in the aforementioned virtual space-based audio processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0218] In one embodiment, a computer device is provided, which may be a terminal or a server. Its internal structure diagram may be as follows: Figure 13 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a virtual space-based audio processing method.
[0219] Those skilled in the art will understand that Figure 13The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0220] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0221] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0222] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the steps in the above method embodiments.
[0223] It should be noted that the user information (including but not limited to user device information, user operation information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0224] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these. The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of these technical features are not contradictory, they should be considered within the scope of this specification.
[0225] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. An audio processing method based on virtual space, characterized in that, The method includes: In response to an interactive operation targeting a region corresponding to a sound source in a virtual space, a region of interest in the virtual space to which the interactive operation is directed is determined, a sound source in the region of interest is designated as a sound source of interest, and all sound sources in the virtual space other than the sound source of interest are designated as sound sources of non-interest. The interactive operation is captured by an interactive device. The audio signals emitted by multiple sources of interest are mixed to obtain the target audio signal, and the audio signals emitted by multiple sources of non-interest are mixed to obtain the target audio signal. Based on the target time period falling into which the sampling time of the audio to be adjusted falls, the adjustment parameters corresponding to each sampling time are determined. The audio to be adjusted includes at least one of target attention audio and target non-attention audio. The target time period is determined by the switching process between the attention source and the non-attention source. If the audio to be adjusted is target attention audio, the target time period is defined based on a complete lifecycle of the sound source as a attention source. The target time period is one of a focus entry time period, a continuous attention time period, and a focus exit time period. The start time of the focus entry time period is determined based on the moment the sound source is identified as a attention source. The start time of the exit from the attention period is determined based on the moment when the attention source is switched to a non-attention source, or based on an unattention command; if the audio to be adjusted is a target non-attention audio, the target time period is defined based on a complete life cycle of the sound source as a non-attention source, and the target time period is one of entering a non-attention period, continuing a non-attention period, and exiting a non-attention period; wherein, the start time of entering a non-attention period is determined based on the moment when the sound source is determined to be a non-attention source, and the start time of exiting a non-attention period is determined based on the moment when the non-attention source is switched to a attention source, or based on an attention command; If the audio to be adjusted is the target audio of interest, then the target audio of interest is subjected to audio adjustment processing based on the adjustment parameters corresponding to each sampling time of the target audio of interest to obtain the first intermediate audio; wherein, the adjustment parameters corresponding to the entry of the attention period gradually increase, the adjustment parameters corresponding to the continuous attention period remain unchanged, and the adjustment parameters corresponding to the exit of the attention period gradually decrease. If the audio to be adjusted is the target non-interested audio, then based on the adjustment parameters corresponding to each sampling time of the target non-interested audio, audio adjustment processing is performed on the target non-interested audio to obtain the second intermediate audio; wherein, the adjustment parameters corresponding to the entry into the non-interested time period gradually decrease, the adjustment parameters corresponding to the continuous non-interested time period remain unchanged, and the adjustment parameters corresponding to the exit from the non-interested time period gradually increase. The first intermediate audio and the second intermediate audio are mixed to obtain the mixing result.
2. The method according to claim 1, characterized in that, The adjustment parameters corresponding to each sampling time are determined based on the target time period in which the sampling time of the audio to be adjusted falls, including: Determine the target time period that each sampling moment of the audio to be adjusted falls into; The calculation results of the adjustment function corresponding to the target time period that each sampling time falls into are obtained and used as the adjustment parameters corresponding to each sampling time. The value of one independent variable in the adjustment function is determined based on the sampling time.
3. The method according to claim 2, characterized in that, The adjustment function includes an adjustment threshold, which is determined by the distance between the corresponding sound source and the virtual operation object in the virtual space. The virtual operation object is the mapping object in the virtual space of the target object that triggers the interactive operation.
4. The method according to claim 2, characterized in that, The start time of the continuous attention period is determined based on the end time of the initiation attention period.
5. The method according to claim 4, characterized in that, The adjustment function corresponding to the period of initial attention is a monotonically increasing function, the adjustment function corresponding to the period of continuous attention is a constant function, and the adjustment function corresponding to the period of exiting attention is a monotonically decreasing function.
6. The method according to claim 5, characterized in that, The gradients of both the monotonically increasing and monotonically decreasing functions gradually decrease as the independent variable increases.
7. The method according to claim 2, characterized in that, The start time of the continuous non-interest period is determined based on the end time of the cut-in non-interest period.
8. The method according to claim 7, characterized in that, The adjustment function corresponding to the entry into the non-focused time period is a monotonically decreasing function, the adjustment function corresponding to the continuous non-focused time period is a constant function, and the adjustment function corresponding to the exit from the non-focused time period is a monotonically increasing function.
9. The method according to any one of claims 1 to 8, characterized in that, Each sound source in the virtual space corresponds to multiple sound channels, and the method further includes: The mixing results for each of the multiple channels are obtained, stereo reconstruction is performed based on the mixing results of the multiple channels, and the reconstructed stereo is output in the virtual space.
10. The method according to any one of claims 1 to 8, characterized in that, The operation type of the interactive operation includes at least one of sensory pointing type, consciousness pointing type and limb pointing type.
11. An audio processing device based on virtual space, characterized in that, The device includes: The determination module is used to respond to an interactive operation targeting a region corresponding to a sound source in a virtual space, determine the region of interest in the virtual space to which the interactive operation is directed, identify the sound source in the region of interest as the sound source of interest, and identify all sound sources in the virtual space other than the sound source of interest as the sound source of non-interest, wherein the interactive operation is captured by an interactive device. The acquisition module is used to mix the audio of each of the multiple audio sources of interest to obtain the target audio of interest, and to mix the audio of each of the multiple audio sources of non-interest to obtain the target audio of non-interest. An audio adjustment processing module is used to determine the adjustment parameters corresponding to each sampling moment based on the target time period into which the sampling moment of the audio to be adjusted falls. The audio to be adjusted includes at least one of target attention audio and target non-attention audio. The target time period is determined by the switching process between attention sources and non-attention sources. If the audio to be adjusted is target attention audio, the target time period is defined based on a complete lifecycle of the sound source as a attention source. The target time period is one of a focus entry time period, a continuous attention time period, and a focus exit time period. The start time of the focus entry time period is based on the moment when the sound source is determined to be a attention source. It is determined that the start time of the exit from the attention period is determined based on the moment when the attention source is switched to a non-attention source, or based on an unfollow command; if the audio to be adjusted is a target non-attention audio, then the target time period is defined based on a complete life cycle of the sound source as a non-attention source, and the target time period is one of entering a non-attention period, continuous non-attention period, and exiting a non-attention period; wherein, the start time of entering a non-attention period is determined based on the moment when the sound source is determined to be a non-attention source, and the start time of exiting a non-attention period is determined based on the moment when the non-attention source is switched to a attention source, or based on an attention command; The audio adjustment processing module is further configured to, when the audio to be adjusted is a target audio of interest, perform audio adjustment processing on the target audio of interest based on the adjustment parameters corresponding to each sampling time of the target audio of interest to obtain a first intermediate audio; wherein, the adjustment parameters corresponding to the entry into the attention period gradually increase, the adjustment parameters corresponding to the continuous attention period remain unchanged, and the adjustment parameters corresponding to the exit from the attention period gradually decrease; when the audio to be adjusted is a target non-attention audio, perform audio adjustment processing on the target non-attention audio based on the adjustment parameters corresponding to each sampling time of the target non-attention audio to obtain a second intermediate audio; wherein, the adjustment parameters corresponding to the entry into the non-attention period gradually decrease, the adjustment parameters corresponding to the continuous non-attention period remain unchanged, and the adjustment parameters corresponding to the exit from the non-attention period gradually increase; The audio mixing module is used to mix the first intermediate audio and the second intermediate audio to obtain the audio mixing result.
12. The apparatus according to claim 11, characterized in that, The audio adjustment processing module is further configured to determine the target time period into which each sampling time of the audio to be adjusted falls; obtain the calculation result of the adjustment function corresponding to the target time period into which each sampling time falls, and use it as the adjustment parameter corresponding to each sampling time, wherein the value of one independent variable in the adjustment function is determined based on the sampling time.
13. The apparatus according to claim 12, characterized in that, The adjustment function includes an adjustment threshold, which is determined by the distance between the corresponding sound source and the virtual operation object in the virtual space. The virtual operation object is the mapping object in the virtual space of the target object that triggers the interactive operation.
14. The apparatus according to claim 12, characterized in that, The start time of the continuous attention period is determined based on the end time of the initiation attention period.
15. The apparatus according to claim 14, characterized in that, The adjustment function corresponding to the period of initial attention is a monotonically increasing function, the adjustment function corresponding to the period of continuous attention is a constant function, and the adjustment function corresponding to the period of exiting attention is a monotonically decreasing function.
16. The apparatus according to claim 15, characterized in that, The gradients of both the monotonically increasing and monotonically decreasing functions gradually decrease as the independent variable increases.
17. The apparatus according to claim 12, characterized in that, The start time of the continuous non-interest period is determined based on the end time of the cut-in non-interest period.
18. The apparatus according to claim 17, characterized in that, The adjustment function corresponding to the entry into the non-focused time period is a monotonically decreasing function, the adjustment function corresponding to the continuous non-focused time period is a constant function, and the adjustment function corresponding to the exit from the non-focused time period is a monotonically increasing function.
19. The apparatus according to any one of claims 11 to 18, characterized in that, The audio adjustment and processing module is also used to obtain the mixing processing results corresponding to each of the multiple channels, perform stereo reconstruction based on the mixing processing results of the multiple channels, and output the reconstructed stereo in the virtual space.
20. The apparatus according to any one of claims 11 to 18, characterized in that, The operation type of the interactive operation includes at least one of sensory pointing type, consciousness pointing type and limb pointing type.
21. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 10.
22. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.