Audio adjusting method, apparatus, device and computer readable storage medium
Patent Information
- Application Number
- CN202211218017.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2042-09-30
AI Technical Summary
[0003]本申请的主要目的在于提供一种音频调节方法、电子设备及计算机可读存储介质,旨在解决现有技术中空间音频的音效难以满足用户的个性化需求的技术问题
[0057] This application provides an audio adjustment method, apparatus, device, and computer-readable storage medium. Specifically, when a user interacts with a virtual space through a first interactive operation, in response to the user's first interactive operation on the virtual space, at least one target sound source in the virtual space is identified. The first interactive operation includes at least one of eye-tracking interaction and gesture interaction. Then, when the user interacts with the target sound source through a second interactive operation, in response to the user's second interactive operation on the target sound source, corresponding audio adjustment processing is performed on the target sound source. Therefore, in this application, users can adjust the volume, sound effects, and other audio parameters of the target sound source in the virtual space according to their own needs, satisfying their personalized requirements.
Smart Images

Figure CN115494953B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of virtual reality technology, and in particular to an audio adjustment method, apparatus, device, and computer-readable storage medium. Background Technology
[0002] With the rapid development of science and technology and people's pursuit of better sound effects, spatial audio technology has emerged. Spatial audio technology refers to transmitting spatial information data of the audio source simultaneously with the audio data transmission process. By combining this spatial information data with the user's motion data, it creates a three-dimensional spatial effect, providing the user with an immersive auditory experience. However, to achieve the realism of the three-dimensional spatial effect, current spatial audio technologies typically adjust the sound effects holistically, which makes it difficult to meet the personalized needs of users. Summary of the Invention
[0003] The main purpose of this application is to provide an audio adjustment method, electronic device, and computer-readable storage medium, which aims to solve the technical problem that the sound effects of spatial audio in the prior art cannot meet the personalized needs of users.
[0004] To achieve the above objectives, this application provides an audio adjustment method, applied to an audio adjustment device, the audio adjustment method comprising:
[0005] In response to a first interactive operation in a virtual space, at least one target sound source in the virtual space is identified, wherein the first interactive operation includes at least one of an eye-tracking interactive operation and a gesture interactive operation.
[0006] In response to a second interactive operation on the target audio source, corresponding audio adjustment processing is performed on the target audio source.
[0007] Optionally, the first interactive operation includes an eye-tracking interactive operation, and the step of determining at least one target sound source in the virtual space in response to the first interactive operation for the virtual space includes:
[0008] By tracking the user's eyes, the user's first gaze direction is detected;
[0009] When the gaze duration in the first gaze direction reaches a preset duration threshold, the virtual sound source in the virtual space in the first gaze direction is determined as the target sound source.
[0010] Optionally, the step of determining the virtual sound source in the virtual space along the first gaze direction as the target sound source further includes, in the audio adjustment method:
[0011] When there is only one virtual sound source in the virtual space in the first gaze direction, the virtual sound source is taken as the target sound source.
[0012] When the number of virtual sound sources in the virtual space in the first gaze direction is at least two, the target virtual region corresponding to the first gaze direction is obtained, and the target virtual region is locally magnified to obtain the magnified target region;
[0013] By performing eye tracking on the user again, the user's second gaze direction in the magnified target area is detected;
[0014] Use the virtual sound source in the second gaze direction as the target sound source.
[0015] Optionally, the first interactive operation includes a gesture interactive operation, and the step of determining at least one target sound source in the virtual space in response to the first interactive operation for the virtual space further includes:
[0016] Acquire the user's first gesture information and identify the direction of the instruction corresponding to the first gesture information;
[0017] When the indicated direction points to a virtual sound source in the virtual space, the virtual sound source is determined to be the target sound source.
[0018] Optionally, the step of performing corresponding audio adjustment processing on the target audio source in response to the user's second interactive operation on the target audio source includes:
[0019] Acquire the user's second gesture information and identify the gesture movement direction and amplitude corresponding to the second gesture information;
[0020] Based on the direction of the gesture movement, determine the volume adjustment direction of the target sound source, and based on the amplitude of the gesture movement, determine the volume adjustment amplitude of the target sound source.
[0021] The target sound source is audio adjusted according to the volume adjustment direction and the volume adjustment amplitude.
[0022] Optionally, the step of determining the volume adjustment direction of the target sound source based on the direction of the gesture movement, and determining the volume adjustment amplitude of the target sound source based on the amplitude of the gesture movement, includes:
[0023] When the gesture moves upward, the volume adjustment direction of the target sound source is determined to be increasing, and the increase in the volume adjustment of the target sound source is calculated based on the gesture movement amplitude.
[0024] When the gesture moves downwards, the volume adjustment direction of the target sound source is determined to be decreasing, and the decrease in volume adjustment of the target sound source is calculated based on the gesture movement amplitude.
[0025] Optionally, the step of performing corresponding audio adjustment processing on the target audio source in response to the user's second interactive operation on the target audio source further includes:
[0026] In response to a sound effect adjustment gesture for the target sound source, a preset sound effect mode for the target sound source is displayed in the virtual space;
[0027] In response to a gesture selecting the preset sound effect mode, the specified sound effect mode of the target sound source is determined;
[0028] The target sound source is subjected to audio adjustment processing based on the adjustment parameters corresponding to the specified sound effect mode.
[0029] Furthermore, to achieve the above objectives, this application also provides an audio adjustment device, which is applied to an audio adjustment equipment. The audio adjustment device includes:
[0030] A sound source determination module is configured to determine at least one target sound source in the virtual space in response to a first interactive operation for the virtual space, wherein the first interactive operation includes at least one of an eye-tracking interactive operation and a gesture interactive operation.
[0031] An audio adjustment module is used to perform corresponding audio adjustment processing on the target audio source in response to a second interactive operation on the target audio source.
[0032] Optionally, the sound source determination module is also used for:
[0033] By tracking the user's eyes, the user's first gaze direction is detected;
[0034] When the gaze duration in the first gaze direction reaches a preset duration threshold, the virtual sound source in the virtual space in the first gaze direction is determined as the target sound source.
[0035] Optionally, the sound source determination module is also used for:
[0036] When there is only one virtual sound source in the virtual space in the first gaze direction, the virtual sound source is taken as the target sound source.
[0037] When the number of virtual sound sources in the virtual space in the first gaze direction is at least two, the target virtual region corresponding to the first gaze direction is obtained, and the target virtual region is locally magnified to obtain the magnified target region;
[0038] By performing eye tracking on the user again, the user's second gaze direction in the magnified target area is detected;
[0039] Use the virtual sound source in the second gaze direction as the target sound source.
[0040] Optionally, the sound source determination module is also used for:
[0041] Acquire the user's first gesture information and identify the direction of the instruction corresponding to the first gesture information;
[0042] When the indicated direction points to a virtual sound source in the virtual space, the virtual sound source is determined to be the target sound source.
[0043] Optionally, the audio adjustment module is also used for:
[0044] Acquire the user's second gesture information and identify the gesture movement direction and amplitude corresponding to the second gesture information;
[0045] Based on the direction of the gesture movement, determine the volume adjustment direction of the target sound source, and based on the amplitude of the gesture movement, determine the volume adjustment amplitude of the target sound source.
[0046] The target sound source is audio adjusted according to the volume adjustment direction and the volume adjustment amplitude.
[0047] Optionally, the audio adjustment module is also used for:
[0048] When the gesture moves upward, the volume adjustment direction of the target sound source is determined to be increasing, and the increase in the volume adjustment of the target sound source is calculated based on the gesture movement amplitude.
[0049] When the gesture moves downwards, the volume adjustment direction of the target sound source is determined to be decreasing, and the decrease in volume adjustment of the target sound source is calculated based on the gesture movement amplitude.
[0050] Optionally, the audio adjustment module is also used for:
[0051] In response to a sound effect adjustment gesture for the target sound source, a preset sound effect mode for the target sound source is displayed in the virtual space;
[0052] In response to a gesture selecting the preset sound effect mode, the specified sound effect mode of the target sound source is determined;
[0053] The target sound source is subjected to audio adjustment processing based on the adjustment parameters corresponding to the specified sound effect mode.
[0054] This application also provides an electronic device, which includes: a memory, a processor, and a program for an audio adjustment method stored in the memory and executable on the processor. When the program for the audio adjustment method is executed by the processor, it can implement the steps of the audio adjustment method as described above.
[0055] This application also provides a computer-readable storage medium storing a program that implements an audio adjustment method. When the program is executed by a processor, it implements the steps of the audio adjustment method as described above.
[0056] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the audio adjustment method described above.
[0057] This application provides an audio adjustment method, apparatus, device, and computer-readable storage medium. Specifically, when a user interacts with a virtual space through a first interactive operation, in response to the user's first interactive operation on the virtual space, at least one target sound source in the virtual space is identified. The first interactive operation includes at least one of eye-tracking interaction and gesture interaction. Then, when the user interacts with the target sound source through a second interactive operation, in response to the user's second interactive operation on the target sound source, corresponding audio adjustment processing is performed on the target sound source. Therefore, in this application, users can adjust the volume, sound effects, and other audio parameters of the target sound source in the virtual space according to their own needs, satisfying their personalized requirements. Attached Figure Description
[0058] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0059] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0060] Figure 1 This is a flowchart illustrating the first embodiment of the audio adjustment method of this application;
[0061] Figure 2 This is a flowchart illustrating the second embodiment of the audio adjustment method of this application;
[0062] Figure 3 This is a schematic diagram of the device structure of the hardware operating environment involved in the audio adjustment method in the embodiments of this application.
[0063] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0064] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0065] Example 1
[0066] This application provides an audio adjustment method. In the first embodiment of the audio adjustment method of this application, refer to... Figure 1 Audio adjustment methods include:
[0067] Step S10: In response to the first interactive operation for the virtual space, at least one target sound source in the virtual space is determined;
[0068] Step S20: In response to the second interactive operation for the target sound source, perform corresponding audio adjustment processing on the target sound source.
[0069] In this embodiment, it should be noted that the audio adjustment method can be applied to any device with audio adjustment needs, such as VR (Virtual Reality) devices (e.g., VR glasses), smartphones, computers, and servers.
[0070] Additionally, it should be noted that virtual space is primarily a two-dimensional, three-dimensional, or higher-dimensional space constructed by computer equipment. Virtual space can present visual elements, such as people, objects, scenery, light and shadow, and water, as well as auditory elements, such as human voices (speaking, singing, etc.), object sounds (music, vehicle sounds, knocking sounds, etc.), and ambient sounds (wind sounds, rain sounds, birdsong, etc.).
[0071] In this embodiment, the first interactive operation targeting the virtual space is mainly used to determine the target sound source that the user wishes to adjust within the virtual space. The second interactive operation targeting the virtual space is mainly used to perform audio adjustment processing on the target sound source within the virtual space (such as adjusting volume, adding or removing decorative sound effects, etc.). It is understood that different sound sources have corresponding virtual representations within the virtual space; for example, the virtual representation of flowing water is a river, the virtual representation of birdsong is a bird, the virtual representation of conversation and singing is a human, and the virtual representation of musical instrument sounds is a musical instrument, etc. The first interactive operation includes at least one of gesture interaction or eye-tracking interaction. The second interactive operation can be performed using gesture interaction, voice interaction, etc.
[0072] For example, when the first interaction is an eye-tracking interaction, eye movement can be tracked by measuring the position of the user's gaze point or the movement of the eyeballs relative to the head. This monitors the user's eye movement and gaze direction when looking at a specific target, and uses the user's gaze direction to identify the virtual sound source (i.e., the virtual image of the sound source) in the virtual space as the target sound source. In another embodiment, for example, when the first interaction is a gesture interaction, the user's gesture can be detected. When the user's gesture is a directional gesture, the direction corresponding to the directional gesture can be determined, and the virtual sound source (i.e., the virtual image of the sound source) in the virtual space in that direction can be identified as the target sound source. Furthermore, for example, the user's head movement can be detected to determine the user's facing direction, and then a region of interest in the virtual space can be determined based on the facing direction. The virtual sound source within that region of interest can then be identified as the target sound source.
[0073] Understandably, since users may need to adjust audio from multiple virtual sound sources, if there are two or more target sound sources, and the distance between them is relatively short, and if there is no virtual sound source in the virtual space at the target direction (such as the gaze direction, pointing direction, etc.) corresponding to the first interactive operation, and the gaze duration in the target direction reaches a preset duration threshold, then the spatial intersection of the target direction and the virtual space can be determined, and all virtual sound sources whose relative distance to the spatial intersection is within a preset distance threshold can be used as target sound sources. For example, if the virtual space is a concert scene, and the user wants to adjust the audio of a guitar and drum set that are close together in the virtual space, the user can use eye-tracking interaction to gaze at the position between the guitar and drum set, or use gesture interaction to point to the position between the guitar and drum set, so that both the guitar and drum set can be used as target sound sources, without having to frequently point to or gaze at multiple sound sources to determine two or more target sound sources, thus improving the convenience and efficiency of the user in determining target sound sources.
[0074] For example, when the interaction method of the second interactive operation is gesture interaction, after determining the target sound source in the virtual space, the user's adjustment intention can be determined by detecting the user's adjustment gesture. It is understood that the adjustment gesture includes volume adjustment gestures and sound effect adjustment gestures. For example, if the adjustment gesture is a volume adjustment gesture, an upward gesture will increase the volume of the target sound source; a downward gesture will decrease the volume of the target sound source. If the adjustment gesture is a sound effect adjustment gesture (such as clicking the target sound source, pointing at the target sound source for a duration exceeding a preset switching duration, etc.), then the icon corresponding to the preset sound effect mode will be displayed, and the user's selection gesture for the preset sound effect mode (such as clicking the icon corresponding to a preset sound effect mode) will be received, thereby determining the specified sound effect mode of the target sound source. According to the adjustment parameters corresponding to the specified sound effect mode, the target sound source will be subjected to audio adjustment processing.
[0075] It is also understood that after detecting the user's first and second interactive operations in the virtual space, the first and second interactive operations can be converted into interactive action data that can be recognized by the computer device, so that the computer device can respond based on the interactive action data, that is, execute the process in this embodiment.
[0076] In step S10, the first interaction operation includes an eye-tracking interaction operation. The step of determining at least one target sound source in the virtual space in response to the first interaction operation in the virtual space includes:
[0077] Step A10: Detect the user's first gaze direction by performing eye tracking on the user;
[0078] Step A20: When the gaze duration in the first gaze direction reaches a preset duration threshold, the virtual sound source in the virtual space in the first gaze direction is determined as the target sound source.
[0079] Steps A10 to A20 include: when the user is gazing at the desired virtual sound source in the virtual space, the user's first gaze direction can be detected by eye tracking. When the gaze duration in the first gaze direction reaches a preset duration threshold (e.g., 2 seconds, 3 seconds, etc.), it indicates that the virtual sound source in the virtual space in the first gaze direction is the virtual sound source the user desires to adjust, and thus the virtual sound source can be determined as the target sound source.
[0080] In step S10, the first interactive operation includes a gesture interactive operation. The step of determining at least one target sound source in the virtual space in response to the first interactive operation on the virtual space further includes:
[0081] Step B10: Obtain the user's first gesture information and identify the direction of indication corresponding to the first gesture information;
[0082] Step B20: When the indicated direction points to the virtual sound source in the virtual space, the virtual sound source is determined to be the target sound source.
[0083] Steps B10 to B20 include: when the user points to a desired virtual sound source within the virtual space, the VR device's sensors can capture the user's hand contour to obtain first gesture information. This first gesture information includes a directional gesture for identifying the target sound source the user desires to adjust within the virtual space. Then, by recognizing the direction corresponding to the first gesture information, if the direction points to the virtual sound source in the virtual space, it indicates that the virtual sound source is the one the user desires to adjust, thus confirming the virtual sound source as the target sound source.
[0084] In step S20, the step of performing corresponding audio adjustment processing on the target audio source in response to the user's second interactive operation on the target audio source includes:
[0085] Step C10: Obtain the user's second gesture information and identify the gesture movement direction and gesture movement amplitude corresponding to the second gesture information;
[0086] Step C20: Determine the volume adjustment direction of the target sound source based on the direction of the gesture movement, and determine the volume adjustment range of the target sound source based on the amplitude of the gesture movement;
[0087] Step C30: Adjust the audio of the target sound source according to the volume adjustment direction and the volume adjustment amplitude.
[0088] Steps C10 to C30 include: when a user wishes to adjust the volume of a target sound source, the VR device's sensors can capture the user's hand contour to obtain second gesture information. This second gesture information includes adjustable gestures for audio adjustment processing of the target sound source in virtual space. Then, the gesture movement direction and amplitude corresponding to the second gesture information are identified. Based on the gesture movement direction, the volume adjustment direction of the target sound source is determined, and based on the gesture movement amplitude, the volume adjustment amplitude of the target sound source is determined. It is understood that the gesture movement direction and volume adjustment direction have a preset correspondence. For example, if the gesture movement direction is upward, the volume adjustment direction is to increase; if the gesture movement direction is downward, the volume adjustment direction is to decrease; similarly, if the gesture movement direction is to the right, the volume adjustment direction is to increase; and if the gesture movement direction is to the left, the volume adjustment direction is to decrease. The correspondence between the gesture movement direction and the volume adjustment direction can be set according to requirements, and this embodiment does not impose any limitations on this. The gesture movement can be a distance (i.e., the entire hand being raised or lowered) or an angle (i.e., the palm swinging upwards or downwards with the wrist as a fulcrum). Then, based on the volume adjustment direction and the volume adjustment amplitude, the target sound source is audio adjusted. Furthermore, during the audio adjustment of the target sound source, the current volume of the target sound source can be displayed near the target sound source to provide the user with a more intuitive understanding of the audio adjustment status.
[0089] Step C20, which involves determining the volume adjustment direction of the target sound source based on the direction of the gesture movement and determining the volume adjustment amplitude of the target sound source based on the amplitude of the gesture movement, includes:
[0090] Step C21: When the gesture moves upward, determine that the volume adjustment direction of the target sound source is to increase, and calculate the increase in the volume adjustment of the target sound source based on the gesture movement amplitude.
[0091] Step C22: When the gesture movement direction is downward, determine that the volume adjustment direction of the target sound source is decreasing, and calculate the decrease in volume adjustment of the target sound source based on the gesture movement amplitude.
[0092] Steps C21 and C22 include: when the gesture movement direction is upward, determining that the volume adjustment direction of the target sound source is to increase, and calculating the increase magnitude of the target sound source volume adjustment based on the gesture movement amplitude. When the gesture movement direction is downward, determining that the volume adjustment direction of the target sound source is to decrease, and calculating the decrease magnitude of the target sound source volume adjustment based on the gesture movement amplitude. For example, the volume adjustment amplitude corresponding to the gesture movement amplitude can be obtained by querying a preset mapping table of gesture movement amplitude and volume adjustment amplitude, or by using a preset formula relating gesture movement amplitude and volume adjustment amplitude. In this embodiment, the volume of the target sound source is adjusted by the user's upward or downward gesture, allowing the user to adjust the volume of the target sound source in the virtual space according to their own needs, thus satisfying the user's personalized requirements.
[0093] In step S20, the step of performing corresponding audio adjustment processing on the target audio source in response to the user's second interactive operation on the target audio source further includes:
[0094] Step D10: In response to a sound effect adjustment gesture for the target sound source, display the preset sound effect mode of the target sound source in the virtual space;
[0095] Step D20: In response to the gesture of selecting the preset sound effect mode, determine the specified sound effect mode of the target sound source;
[0096] Step D30: Perform audio adjustment processing on the target sound source according to the adjustment parameters corresponding to the specified sound effect mode.
[0097] Steps D10 to D30 include: when a user wishes to adjust the preset sound effect mode of a target sound source, they can interact with the target sound source in the virtual space using sound effect adjustment gestures. The preset sound effect mode is a sound effect mode processed through audio signal correction, compensation, spatial surround, sound field widening, and dynamic enhancement. The user can interact with the target sound source in the virtual space using sound effect adjustment gestures (such as clicking the target sound source or pointing at the target sound source for a duration exceeding a preset switching time). The VR device, in response to the sound effect adjustment gesture for the target sound source, can display an icon corresponding to the preset sound effect mode in the virtual space. It then receives the user's selection gesture for the preset sound effect mode (such as clicking or long-pressing the icon corresponding to the preset sound effect mode) and, in response to the selection gesture, determines the specified sound effect mode for the target sound source. It is understood that different preset sound effect modes have corresponding pre-set adjustment parameters. Therefore, after determining the specified sound effect mode, the target sound source can be processed using the adjustment parameters corresponding to the specified sound effect mode. This allows users to adjust the preset sound effect modes of the target sound source in the virtual space according to their own needs, thus satisfying their personalized requirements.
[0098] Furthermore, in another embodiment, after step S10, which determines at least one target sound source in the virtual space in response to a first interactive operation for the virtual space, the audio adjustment method further includes:
[0099] Step S30: Acquire the user's voice command signal and identify the adjustment command corresponding to the voice command signal;
[0100] Step S31: Perform corresponding audio adjustment processing on the target sound source according to the adjustment instruction.
[0101] After identifying at least one target audio source in the virtual space, the user's voice command signal can be acquired, and the adjustment command (i.e., the user's intention to adjust the target audio) can be identified from the voice command signal using voice recognition technology, such as "adjust the volume to 50%" or "switch to classical mode." Then, according to the adjustment command, the target audio source is subjected to corresponding audio adjustment processing. In this embodiment, after identifying at least one target audio source in the virtual space, the user's voice command signal is acquired, and the adjustment command corresponding to the voice command signal is identified. Then, according to the adjustment command, the target audio source is subjected to corresponding audio adjustment processing. Users can adjust the audio of the target audio source using voice commands. Compared to gesture control, users can issue more and more complex adjustment commands through voice commands, making the audio adjustment processing of the target audio source more tailored to the user's needs and satisfying their personalized requirements.
[0102] The first embodiment of this application provides an audio adjustment method. Specifically, when a user interacts with a virtual space through a first interactive operation, in response to the user's first interactive operation on the virtual space, at least one target sound source in the virtual space is identified. The first interactive operation includes at least one of eye-tracking interaction and gesture interaction. Then, when the user interacts with the target sound source through a second interactive operation, in response to the user's second interactive operation on the target sound source, corresponding audio adjustment processing is performed on the target sound source. Therefore, in this application, users can adjust the volume, sound effects, and other audio parameters of the target sound source in the virtual space according to their own needs, satisfying their personalized requirements.
[0103] Example 2
[0104] Furthermore, referring to Figure 2 Based on the first embodiment of this application, in another embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description and will not be repeated hereafter. In step A20, the step of determining the virtual sound source in the virtual space along the first gaze direction as the target sound source, the audio adjustment method further includes:
[0105] Step A21: When the number of virtual sound sources in the virtual space in the first gaze direction is one, the virtual sound source is taken as the target sound source;
[0106] Step A22: When the number of virtual sound sources in the virtual space in the first gaze direction is at least two, the target virtual region corresponding to the first gaze direction is obtained, and the target virtual region is locally magnified to obtain the magnified target region.
[0107] Step A23: By performing eye tracking on the user again, detect the user's second gaze direction in the magnified target area;
[0108] Step A24: Use the virtual sound source in the second gaze direction as the target sound source.
[0109] In this embodiment, it should be noted that when the location of the virtual sound source is far from the user's location in the virtual space, or when the virtual sound sources are close to each other, there may be more than one virtual sound source in the first gaze direction. In this case, if the virtual sound source in the first gaze direction is directly used as the target sound source, it is very likely to misidentify the target sound source.
[0110] As an example, steps A21 to A24 include: when there is only one virtual sound source in the virtual space along the first gaze direction, it indicates that the virtual sound source is the one the user wishes to adjust, and thus the virtual sound source can be used as the target sound source; when there are at least two virtual sound sources in the virtual space along the first gaze direction, a target virtual region corresponding to the first gaze direction is obtained, and the target virtual region is locally magnified to obtain a magnified target region; by performing eye tracking on the user again, a second gaze direction of the user in the magnified target region is detected; and the virtual sound source in the second gaze direction is used as the target sound source. Since the second gaze direction is determined by eye tracking based on the magnified target region, it points more accurately to the virtual sound source the user actually wishes to adjust compared to the first gaze direction.
[0111] A second embodiment of this application provides an audio adjustment method. Specifically, when there is one virtual sound source in the virtual space along the first gaze direction, this virtual sound source is used as the target sound source. When there are at least two virtual sound sources in the virtual space along the first gaze direction, a target virtual region corresponding to the first gaze direction is obtained, and the target virtual region is locally magnified to obtain a magnified target region. The user's eye is tracked again to detect the user's second gaze direction within the magnified target region. The virtual sound source along the second gaze direction is used as the target sound source. Since the second gaze direction is determined by eye tracking based on the magnified target region, it points more accurately to the virtual sound source the user actually wants to adjust compared to the first gaze direction, thereby improving the accuracy of target sound source determination.
[0112] Example 3
[0113] This application embodiment also provides an audio adjustment device, which is applied to an audio adjustment equipment. The audio adjustment device includes:
[0114] A sound source determination module is configured to determine at least one target sound source in the virtual space in response to a first interactive operation for the virtual space, wherein the first interactive operation includes at least one of an eye-tracking interactive operation and a gesture interactive operation.
[0115] An audio adjustment module is used to perform corresponding audio adjustment processing on the target audio source in response to a second interactive operation on the target audio source.
[0116] Optionally, the sound source determination module is also used for:
[0117] By tracking the user's eyes, the user's first gaze direction is detected;
[0118] When the gaze duration in the first gaze direction reaches a preset duration threshold, the virtual sound source in the virtual space in the first gaze direction is determined as the target sound source.
[0119] Optionally, the sound source determination module is also used for:
[0120] When there is only one virtual sound source in the virtual space in the first gaze direction, the virtual sound source is taken as the target sound source.
[0121] When the number of virtual sound sources in the virtual space in the first gaze direction is at least two, the target virtual region corresponding to the first gaze direction is obtained, and the target virtual region is locally magnified to obtain the magnified target region;
[0122] By performing eye tracking on the user again, the user's second gaze direction in the magnified target area is detected;
[0123] Use the virtual sound source in the second gaze direction as the target sound source.
[0124] Optionally, the sound source determination module is also used for:
[0125] Acquire the user's first gesture information and identify the direction of the instruction corresponding to the first gesture information;
[0126] When the indicated direction points to a virtual sound source in the virtual space, the virtual sound source is determined to be the target sound source.
[0127] Optionally, the audio adjustment module is also used for:
[0128] Acquire the user's second gesture information and identify the gesture movement direction and amplitude corresponding to the second gesture information;
[0129] Based on the direction of the gesture movement, determine the volume adjustment direction of the target sound source, and based on the amplitude of the gesture movement, determine the volume adjustment amplitude of the target sound source.
[0130] The target sound source is audio adjusted according to the volume adjustment direction and the volume adjustment amplitude.
[0131] Optionally, the audio adjustment module is also used for:
[0132] When the gesture moves upward, the volume adjustment direction of the target sound source is determined to be increasing, and the increase in the volume adjustment of the target sound source is calculated based on the gesture movement amplitude.
[0133] When the gesture moves downwards, the volume adjustment direction of the target sound source is determined to be decreasing, and the decrease in volume adjustment of the target sound source is calculated based on the gesture movement amplitude.
[0134] Optionally, the audio adjustment module is also used for:
[0135] In response to a sound effect adjustment gesture for the target sound source, a preset sound effect mode for the target sound source is displayed in the virtual space;
[0136] In response to a gesture selecting the preset sound effect mode, the specified sound effect mode of the target sound source is determined;
[0137] The target sound source is subjected to audio adjustment processing based on the adjustment parameters corresponding to the specified sound effect mode.
[0138] The audio adjustment device provided in this application, employing the audio adjustment method described in Embodiment 1 or Embodiment 2 above, solves the technical problem that the sound effects of spatial audio in the prior art are difficult to meet the personalized needs of users. Compared with the prior art, the beneficial effects of the audio adjustment device provided in this application are the same as those of the audio adjustment method provided in the above embodiments, and other technical features of this audio adjustment device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0139] Example 4
[0140] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the audio adjustment method in the first embodiment described above.
[0141] The following is for reference. Figure 3 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0142] like Figure 3As shown, an electronic device may include a processing unit (such as a central processing unit, graphics processing unit, etc.) that can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) or a program loaded from a storage device into random access memory (RAM). The RAM also stores various programs and data required for the operation of the electronic device. The processing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0143] Typically, the following systems can be connected to the I / O interface: input devices including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices including, for example, magnetic tapes, hard disks, etc.; and communication devices. Communication devices allow electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although electronic devices with various systems are shown in the figures, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems may be implemented alternatively.
[0144] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, it performs the functions defined above in the methods of embodiments of this disclosure.
[0145] The electronic device provided in this application, employing the audio adjustment method described in Embodiment 1 or Embodiment 2 above, solves the technical problem that the sound effects of spatial audio in the prior art are difficult to meet the personalized needs of users. Compared with the prior art, the beneficial effects of the electronic device provided in this application are the same as those of the audio adjustment method provided in Embodiment 1 above, and other technical features of the electronic device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0146] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.
[0147] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0148] Example 5
[0149] This embodiment provides a computer-readable storage medium having computer-readable program instructions stored thereon, which are used to perform the sound effect adjustment method in the first embodiment described above.
[0150] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0151] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.
[0152] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to: determine at least one target sound source in the virtual space in response to a first interactive operation for the virtual space, wherein the first interactive operation includes at least one of an eye-tracking interactive operation and a gesture interactive operation; and perform corresponding audio adjustment processing on the target sound source in response to a second interactive operation for the target sound source.
[0153] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0154] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0155] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0156] The computer-readable storage medium provided in this application stores computer-readable program instructions for performing the above-described audio adjustment method, solving the technical problem that the sound effects of spatial audio in the prior art are difficult to meet the personalized needs of users. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the audio adjustment method provided in Embodiment 1 or Embodiment 2 above, and will not be repeated here.
[0157] Example 6
[0158] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the audio adjustment method described above.
[0159] The computer program product provided in this application solves the technical problem that the sound effects of spatial audio in the prior art are difficult to meet the personalized needs of users. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiments of this application are the same as the beneficial effects of the audio adjustment methods provided in Embodiment 1 or Embodiment 2 above, and will not be repeated here.
[0160] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.
Claims
1. An audio conditioning method, characterized by, The audio adjustment method includes: In response to a first interactive operation in a virtual space, at least one target sound source in the virtual space is identified, wherein the first interactive operation includes at least one of an eye-tracking interactive operation and a gesture interactive operation. The step of determining at least one target sound source in the virtual space in response to a first interactive operation in the virtual space includes: By tracking the user's eyes, the user's first gaze direction is detected; If the gaze duration in the first gaze direction reaches a preset duration threshold, and the number of virtual sound sources in the virtual space in the first gaze direction is one, then the virtual sound source is used as the target sound source. If the number of virtual sound sources in the virtual space in the first gaze direction is at least two, then the target virtual region corresponding to the first gaze direction is obtained, and the target virtual region is locally magnified to obtain the magnified target region; By performing eye tracking on the user again, the user's second gaze direction in the magnified target area is detected; Use the virtual sound source in the second gaze direction as the target sound source; Furthermore, when multiple virtual sound sources are close to each other, there is no virtual sound source in the target direction corresponding to the first interaction operation, and the gaze duration in the first gaze direction reaches a preset duration threshold, the spatial intersection of the first interaction operation and the virtual space is determined, and all virtual sound sources whose distance to the spatial intersection is within the preset threshold are taken as the target sound source. After identifying at least one target sound source in the virtual space, the user's voice command signal is acquired, and the adjustment command corresponding to the voice command signal is identified; according to the adjustment command, the target sound source is subjected to corresponding audio adjustment processing. In response to a second interactive operation on the target sound source, the target sound source is subjected to corresponding audio adjustment processing; The step of performing corresponding audio adjustment processing on the target audio source in response to the user's second interactive operation on the target audio source further includes: In response to a sound effect adjustment gesture for the target sound source, a preset sound effect mode for the target sound source is displayed in the virtual space; in response to a selection gesture for the preset sound effect mode, a specified sound effect mode for the target sound source is determined; and audio adjustment processing is performed on the target sound source according to the adjustment parameters corresponding to the specified sound effect mode.
2. The audio adjustment method as described in claim 1, characterized in that, The first interactive operation includes a gesture interactive operation, and the step of determining at least one target sound source in the virtual space in response to the first interactive operation for the virtual space further includes: Acquire the user's first gesture information and identify the direction of the instruction corresponding to the first gesture information; When the indicated direction points to a virtual sound source in the virtual space, the virtual sound source is determined to be the target sound source.
3. The audio adjustment method as described in claim 1, characterized in that, The step of performing corresponding audio adjustment processing on the target audio source in response to the user's second interactive operation on the target audio source includes: Acquire the user's second gesture information and identify the gesture movement direction and amplitude corresponding to the second gesture information; Based on the direction of the gesture movement, determine the volume adjustment direction of the target sound source, and based on the amplitude of the gesture movement, determine the volume adjustment amplitude of the target sound source. The target sound source is audio adjusted according to the volume adjustment direction and the volume adjustment amplitude.
4. The audio adjustment method as described in claim 3, characterized in that, The steps of determining the volume adjustment direction of the target sound source based on the direction of the gesture movement, and determining the volume adjustment amplitude of the target sound source based on the amplitude of the gesture movement, include: When the gesture moves upward, the volume adjustment direction of the target sound source is determined to be increasing, and the increase in the volume adjustment of the target sound source is calculated based on the gesture movement amplitude. When the gesture moves downwards, the volume adjustment direction of the target sound source is determined to be decreasing, and the decrease in volume adjustment of the target sound source is calculated based on the gesture movement amplitude.
5. An audio adjustment device, characterized in that, The audio adjustment device is used in an audio adjustment equipment, and the audio adjustment device includes: A sound source determination module is configured to, in response to a first interactive operation in a virtual space, determine at least one target sound source in the virtual space, wherein the first interactive operation includes at least one of an eye-tracking interactive operation and a gesture interactive operation. Specifically, the sound source determination module is configured to detect the user's first gaze direction by performing eye tracking on the user; if the gaze duration in the first gaze direction reaches a preset duration threshold, and if the number of virtual sound sources in the virtual space in the first gaze direction is one, then that virtual sound source is designated as the target sound source; if the number of virtual sound sources in the virtual space in the first gaze direction is at least two, then the first gaze direction is determined. The target virtual region corresponding to the viewing direction is magnified to obtain the magnified target region; the user's eye is tracked again to detect the user's second gaze direction in the magnified target region; the virtual sound source in the second gaze direction is taken as the target sound source; and when multiple virtual sound sources are close to each other, there is no virtual sound source in the target direction corresponding to the first interaction operation, and the gaze duration in the first gaze direction reaches a preset duration threshold, the spatial intersection of the first interaction operation and the virtual space is determined, and all virtual sound sources whose distance to the spatial intersection is within the preset threshold are taken as the target sound source; An audio adjustment module is used to perform corresponding audio adjustment processing on the target audio source in response to a second interactive operation on the target audio source. Specifically, the audio adjustment module is used to display a preset sound effect mode of the target audio source in the virtual space in response to a sound effect adjustment gesture on the target audio source; to determine a specified sound effect mode of the target audio source in response to a selection gesture on the preset sound effect mode; and to perform audio adjustment processing on the target audio source according to the adjustment parameters corresponding to the specified sound effect mode. The audio adjustment device further includes: after determining at least one target sound source in the virtual space, acquiring the user's voice command signal and identifying the adjustment command corresponding to the voice command signal; and performing corresponding audio adjustment processing on the target sound source according to the adjustment command.
6. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the steps of the audio adjustment method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program for implementing the audio adjustment method, which is executed by a processor to implement the steps of the audio adjustment method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Virtua reality equipment parameter adjustment method and device
CN107402634A
Systems and methods configured to provide gaze-based audio in interactive experiences
US20200097246A1