Audio processing apparatus and method, and extended reality apparatus, device and storage medium
By integrating audio processing devices in VR devices, using positioning devices and processors to process user location and sound source information in real time, and generating personalized audio filtering results, the problem that existing VR devices cannot truly restore spatial audio, and a more immersive 3D audio experience is achieved.
Patent Information
- Application Number
- PCT/CN2023/121888
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-09-27
- Publication Date
- 2025-05-22
AI Technical Summary
Existing VR products cannot truly restore spatial audio based on changes in human and sound source positions in virtual scenes, resulting in a reduced user immersive 3D experience.
An audio processing device is designed, including a positioning device and a processor, which can track the user's location in real time, determine the relationship between the user and the virtual sound source, obtain the user's human parameters, generate personalized audio filtering results based on this information, and reconstruct the audio to be broadcast close to natural sound.
It realizes that in virtual reality scenarios, the audio effects are updated in real time according to the changes in user location and sound source location, enhance the user's immersion and 3D surround sense, and improve the spatial audio experience.
Smart Images

Figure CN2023121888_22052025_PF_FP_ABST
Abstract
Description
Audio processing device, method, extended reality device, equipment and storage medium Technical Field
[0001] The present disclosure belongs to the field of virtual reality display and sensing technology, and specifically relates to an audio processing device, method, extended reality device, equipment and storage medium. Background Art
[0002] When people use virtual reality (VR) or augmented reality (AR) devices to watch movies or play games, in addition to hoping that the system can provide three-dimensional (3D) display effects, they also hope to provide 3D spatial audio, so as to experience a more comprehensive and all-round immersive spatial effect with integrated sound and picture.
[0003] However, the sound-producing devices on existing VR products are generally universal headphones or external speakers, which can only provide two-channel horizontal surround sound or spatial sound field effects with poor adaptability. They cannot truly restore spatial audio according to the changes in the positions of people and sound sources in VR scenes, reducing the user's immersive 3D experience.
[0004] Summary of the Invention
[0005] The present disclosure aims to solve at least one of the technical problems existing in the prior art and provides an audio processing device, method, extended reality device, equipment and storage medium.
[0006] In a first aspect, the technical solution adopted to solve the technical problem of the present disclosure is an audio processing device, comprising:
[0007] a positioning device configured to track and locate a user in a real space in real time, determine the user's positioning information, and send the positioning information to the processor;
[0008] The processor is configured to determine the user's target positioning information based on pre-established spatial information and the user's real-time positioning information; determine the association relationship information between the user and a virtual sound source in the virtual space based on the pre-established spatial information and the user's real-time target positioning information; obtain the user's human body parameters, and determine an audio filtering result that matches the user based on the association relationship information, the human body parameters, and the audio frequency of the virtual sound source; and reconstruct the audio to be played based on the audio pre-configured for the virtual sound source and the audio filtering result.
[0009] In some embodiments, the positioning device includes an inertial measurement unit and a positioning sensor; the positioning sensor includes one of an optical camera, a wireless positioning device, and a microphone array; the positioning information includes first positioning information and second positioning information;
[0010] The inertial measurement unit is configured to track and locate a user in real space in real time, determine first positioning information of the user, and send the first positioning information to the processor;
[0011] The positioning sensor is configured to regularly obtain changes in the real scene according to a preset period, determine second positioning information of the target object in the real scene, and send the second positioning information to the processor;
[0012] The processor determines the target positioning information of the user, specifically including receiving the first positioning information and the second positioning information, and determining the target positioning information based on the user's real-time first positioning information, the second positioning information and pre-built spatial information.
[0013] In some embodiments, the processor determines the target positioning information of the user, specifically including receiving the first positioning information and the second positioning information, and determining calibration information based on the first positioning information and the second positioning information; calibrating the first positioning information using the calibration information to determine the precise positioning information; and determining the target positioning information based on the precise positioning information and pre-established spatial information.
[0014] In some embodiments, the human body parameters include parameters of the user's auricle; the audio processing device further includes a depth camera;
[0015] The depth camera is configured to collect an auricle depth map of the user's auricle and send the map to the processor;
[0016] The processor is further configured to receive the auricle depth map and determine the user's auricle parameters based on the depth information of the auricle in the auricle depth map.
[0017] In some embodiments, the body parameters include auricle parameters and head parameters of the user;
[0018] The audio processing device further includes a depth camera;
[0019] The depth camera is configured to collect a depth map of the user's head and send it to the processor; the head depth map includes at least left ear features and right ear features;
[0020] The processor is further configured to receive the head depth map, and determine the head parameters and auricle parameters of the user based on the depth information of the head and the depth information of the auricle in the head depth map.
[0021] In some embodiments, the audio processor device further comprises a receiving port; the receiving port is electrically connected to an external processing device;
[0022] The receiving port is configured to obtain the user's body parameters uploaded by the external processing device and send them to the processor.
[0023] In some embodiments, the body parameters include auricle parameters, head parameters, neck parameters, and shoulder parameters.
[0024] In some embodiments, the association relationship information includes a relative distance between the user and the virtual sound source, a relative pitch angle between the user and the virtual sound source, and a relative horizontal angle between the user and the virtual sound source;
[0025] The processor determines the association relationship information, specifically including determining a relative distance between the user and the virtual sound source based on a first position coordinate of the user indicated in the real-time target positioning information of the user and a second position coordinate of the virtual sound source indicated in the spatial information; determining a relative pitch angle between the user and the virtual sound source based on a first pitch angle of the user indicated in the real-time target positioning information of the user and a second pitch angle of the virtual sound source indicated in the spatial information; and determining a relative horizontal angle between the user and the virtual sound source based on a first horizontal angle of the user indicated in the real-time target positioning information of the user and a second horizontal angle of the virtual sound source indicated in the spatial information.
[0026] In some embodiments, the processor determines the audio filtering results that match the user, specifically including screening out the audio filtering results that match the user from a pre-set audio filtering database based on the association relationship information, the human body parameters and the audio frequency of the virtual sound source.
[0027] In some embodiments, the processor determines an audio filtering result that matches the user, specifically including inputting the association relationship information, the human body parameters, and the audio frequency of the virtual sound source into a pre-trained target network model, and outputting the audio filtering result that matches the user;
[0028] The pre-trained target network model is obtained by training based on various reference audio filter data and their corresponding human body parameters, the association relationship information and audio frequency in a pre-set audio filter database.
[0029] In some embodiments, the audio filtering result includes a left ear filtering result and a right ear filtering result; the to-be-played audio includes the to-be-played audio corresponding to the left ear and the to-be-played audio corresponding to the right ear;
[0030] The processor reconstructs the audio to be played, specifically including performing a Fourier transform on the audio pre-configured for the virtual sound source to obtain frequency response information of the audio; determining the audio to be played for the left ear based on the frequency response information, the left ear filtering result, and the relative distance between the user and the virtual sound source; and determining the audio to be played for the right ear based on the frequency response information, the right ear filtering result, and the relative distance between the user and the virtual sound source.
[0031] In some embodiments, the processor determines the audio to be played for the left ear, specifically including multiplying the frequency response information and the gain of the corresponding frequency in the left ear filtering result to obtain a first intermediate result; performing an inverse Fourier transform on the first intermediate result to obtain first audio time domain information; processing the first audio time domain information according to the relative distance between the user and the virtual sound source to obtain a left ear audio signal; and processing the left ear audio signal to obtain the audio to be played for the left ear.
[0032] In some embodiments, the processor determines the audio to be played for the right ear, specifically including multiplying the frequency response information and the gain of the corresponding frequency in the right ear filtering result to obtain a second intermediate result; performing an inverse Fourier transform on the second intermediate result to obtain second audio time domain information; processing the second audio time domain information according to the relative distance between the user and the virtual sound source to obtain a right ear audio signal; and processing the right ear audio signal to obtain the audio to be played for the right ear.
[0033] In a second aspect, an embodiment of the present disclosure further provides an extended reality device, comprising the audio processing device and a speaker as described in any one of the first aspects; the audio processing device comprises a positioning device and a processor;
[0034] The positioning device is configured to track and locate a user in real space in real time, determine the user's positioning information, and send it to the processor;
[0035] The processor is configured to determine target location information of the user based on pre-established spatial information and the user's real-time location information; determine association information between the user and a virtual sound source in the virtual space based on the pre-established spatial information and the user's real-time location information; obtain human body parameters of the user, and determine an audio filtering result that matches the user based on the association information, the human body parameters, and the audio frequency of the virtual sound source; and reconstruct the audio to be played based on the audio pre-configured for the virtual sound source and the audio filtering result.
[0036] The speaker is configured to receive and play the audio to be played.
[0037] In some embodiments, the extended reality device further includes a positioning sensor; the positioning sensor includes one of an optical camera, a wireless positioning device, and a microphone array; the positioning sensor element is fixed on the outside of the outer shell.
[0038] In some embodiments, the audio processing device further includes a depth camera; the depth camera is fixed on the outside of the outer shell.
[0039] In some embodiments, the audio processing device further includes a depth camera and a handle communicatively connected to the processor; the depth camera is fixed to the handle.
[0040] In some embodiments, the audio processing device further includes a receiving port; the receiving port is electrically connected to the processor.
[0041] In a third aspect, the present disclosure further provides an audio processing method, including:
[0042] Tracking and locating a user in real space in real time, determining the user's location information, and sending it to a processor;
[0043] Determine the target location information of the user based on the pre-built spatial information and the user's real-time location information;
[0044] Determining, based on pre-established spatial information and the user's real-time target positioning information, an association relationship between the user and a virtual sound source in the virtual space;
[0045] Acquiring a user's body parameters, and determining an audio filtering result that matches the user based on the association relationship information, the body parameters, and the audio frequency of the virtual sound source;
[0046] The audio to be played is reconstructed based on the audio pre-configured for the virtual sound source and the audio filtering result.
[0047] In a fourth aspect, an embodiment of the present disclosure further provides a computer device comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the steps of the audio processing method described in the third aspect are performed.
[0048] In a fifth aspect, an embodiment of the present disclosure further provides a computer non-volatile readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the audio processing method described in the third aspect are executed. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] FIG1 is a schematic diagram of an audio processing device provided by an embodiment of the present disclosure;
[0050] FIG2a is a schematic diagram of some parameters of the head, neck, and shoulders among the human body parameters provided by an embodiment of the present disclosure;
[0051] FIG2 b is a schematic diagram of auricle parameters among the human body parameters provided by an embodiment of the present disclosure;
[0052] FIG3 is a schematic diagram of an exemplary audio processing device provided by an embodiment of the present disclosure;
[0053] FIG4 is a schematic diagram of adjusting the position of a depth camera on a VR device according to an embodiment of the present disclosure;
[0054] FIG5 is a schematic diagram of another exemplary audio processing device provided by an embodiment of the present disclosure;
[0055] FIG6 is a schematic diagram of another exemplary audio processing device provided by an embodiment of the present disclosure;
[0056] FIG7 is a schematic diagram of a positioning device provided by an embodiment of the present disclosure;
[0057] FIG8 is a system architecture diagram of an audio processing device for determining audio filtering results according to an embodiment of the present disclosure;
[0058] FIG9 is a model framework diagram of a target network model provided by an embodiment of the present disclosure;
[0059] FIG10 is a system structure diagram of another audio processing device for determining audio filtering results provided by an embodiment of the present disclosure;
[0060] FIG11 is a diagram illustrating a system architecture for generating personalized left-ear audio capture and right-ear audio playback according to an embodiment of the present disclosure;
[0061] FIG12 is a schematic diagram of an extended reality device provided by an embodiment of the present disclosure;
[0062] FIG13 is a flowchart of an audio processing method provided by an embodiment of the present disclosure;
[0063] FIG14 is a schematic structural diagram of a computer device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. The components of the embodiments of the present disclosure generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the disclosure for which protection is sought, but merely represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present disclosure.
[0065] Unless otherwise defined, the technical or scientific terms used in this disclosure should have the usual meanings understood by people with ordinary skills in the field to which this disclosure belongs. The words "first", "second" and similar words used in this disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "one", "an" or "the" do not indicate a quantity limitation, but rather indicate the existence of at least one. Words such as "include" or "comprise" mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0066] In this disclosure, "multiple or several" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0067] In related technologies, when users use headphones, the way they receive sound is different from natural sound reception. The sound from in-ear headphones is directly "poured" into the ear canal without passing through the outer ear at all, so the human brain perceives the sound reception as unnatural. In order to restore the true natural sound of audio, it is often necessary to combine human body parameters related to the head. For example, traditional audio reconstruction technology uses Head Related Transfer Functions (HRTFs) or other time-domain forms of Head Related Impulse Response (HRIR). HRTFs are the core technology of virtual auditory display. They describe the transfer function between the sound source and the two ears in a free field and contain key information for virtual sound source playback.
[0068] However, current technology using HRTFs to reconstruct spatial audio is primarily used on mobile phones. This often involves capturing personalized body parameters through mobile phone photography, followed by audio processing via mobile software. Mobile phones are unable to detect real-time changes in the position between the user and the sound source, resulting in a relatively flat spatial audio reproduction. The audio effect fails to truly replicate the natural sound reception of the human ear, hindering the user's ability to experience a truly immersive experience.
[0069] In view of this, embodiments of the present disclosure provide an audio processing device, specifically comprising a positioning device and a processor. The positioning device is capable of tracking and locating a user in real space in real time, thereby determining the user's location information. The processor is capable of determining the user's target location information based on pre-established spatial information and the user's real-time location information. Based on this pre-established spatial information and the user's real-time location information, the processor can quickly determine the association relationship between the user and a virtual sound source, which can be understood as information about the relative position change between the user and the virtual sound source. Based on this, the processor obtains the user's body parameters and, based on the user's real-time association relationship information, body parameters, and the audio frequency of the virtual sound source, updates the user's frequency response in real time, thereby generating personalized audio filtering results for each user in real time. The processor then reconstructs the audio pre-configured for the virtual sound source based on the user's real-time audio filtering results, ensuring that the reconstructed audio to be played is closer to the natural sound perceived by the user's ears. This optimizes the user's spatial audio experience and, combined with virtual reality scenarios, enhances the user's sense of immersion and 3D surround sound, improving the user experience.
[0070] The audio processing device provided by the present disclosure is mainly used in virtual reality scenes or augmented reality scenes. The audio processing device can be a component integrated in an extended reality device. An extended reality device is, for example, a VR / RA device. Taking a VR device as an example, the main hardware structure of a common VR device includes, for example, a head display and a handle. The head display is used for data processing and display, and the handle is used for virtual operation and control. Compared with the conventional structure, the VR device provided by the present disclosure adds an audio processing device. Some functions of the audio processing device are implemented by the original VR device, some functions are integrated in the processor of the head display, and some functions are implemented by additional hardware.
[0071] The audio processing device provided by the embodiment of the present disclosure is described in detail below.
[0072] FIG1 is a schematic diagram of an audio processing apparatus provided by an embodiment of the present disclosure. As shown in FIG1 , the apparatus specifically includes a positioning device 10 and a processor 20, wherein:
[0073] The positioning device 10 is configured to track and locate a user in real space in real time and determine the user's positioning information. The user in real space can be understood as the wearer of the VR device in a real scene.
[0074] For example, the positioning device 10 may be an inertial measurement unit (IMU) integrated into a head-mounted display. The IMU may be, for example, a nine-axis IMU, specifically including a gyroscope, an accelerometer, and a geomagnetic compass, which can respectively obtain the user's angular velocity, acceleration, and spatial coordinates in three directions.
[0075] When the positioning device 10 is an IMU, the IMU can detect the position change of the user in real time, thereby determining the positioning information of the user in the real space. The positioning information specifically includes the position information and posture information of the user in the real space.
[0076] For example, the positioning device 10 may be a combination of an IMU and other sensors. The other sensors may be, for example, an optical camera, a wireless positioning device, or an array microphone. The wireless positioning device may be, for example, a Global Positioning System (GPS), a Global Navigation Satellite System (GNSS), a Beidou Navigation Satellite System (BDS), or the like.
[0077] When the positioning device 10 is an IMU + other sensors, the IMU can detect the user's position changes in real time, thereby determining the positioning information of the IMU; other sensors can obtain changes in the real scene, thereby assisting in verifying the positioning information of the IMU.
[0078] The user's positioning information includes, for example, the position and posture of the center of the left auricle and / or the position and posture of the center of the right auricle.
[0079] The processor 20 is configured to determine the user's target location information based on pre-established spatial information and the user's real-time location information. The virtual space is a virtual space pre-established by the processor 20 in a virtual reality scenario. The spatial information corresponding to the virtual space includes the size and location information of all virtual objects in the virtual space. The user's real-time location information is also the location information tracked in real time by the positioning device 10.
[0080] The positioning information determined by the positioning device 10 is the actual position and posture of the user in the real space scene. Based on the spatial information, the positioning information is mapped to the virtual space to determine the target positioning information of the user in the virtual space. The target positioning information includes position information and posture information.
[0081] The processor 20 is further configured to determine association relationship information between the user and the virtual sound source in the virtual space based on pre-built space information and the user's real-time target positioning information.
[0082] The pre-built space information includes the size and positioning information of all virtual objects in the pre-built virtual space, such as the sound source positioning information (including spatial coordinates and posture) of a virtual sound source. The user's real-time target positioning information is the most recently determined target positioning information under real-time tracking by the positioning device 10.
[0083] The processor 20 can determine the relative position change between the user and the virtual sound source based on the target positioning information and the sound source positioning information. The association relationship information is used to represent the relative position change between the user and the virtual sound source, including but not limited to relative position information and relative posture information.
[0084] Exemplarily, the association relationship information includes, for example, first association relationship information between the left ear and the virtual sound source, and / or second association relationship information between the right ear and the virtual sound source. The present disclosure takes one of them as an example for illustration, and the repeated parts are not repeated here.
[0085] Afterwards, the processor 20 is further configured to obtain the user's body parameters, and determine an audio filtering result that matches the user based on the association relationship information, the body parameters, and the audio frequency of the virtual sound source.
[0086] For example, FIG2a is a schematic diagram of some parameters of the head, neck, and shoulders among the human body parameters provided in the embodiment of the present disclosure, and FIG2b is a schematic diagram of the auricle parameters among the human body parameters provided in the embodiment of the present disclosure. As shown in FIG2a and FIG2b, the human body parameters include, but are not limited to, auricle parameters, head parameters, neck parameters, and shoulder parameters. For example, the head parameters include x1 to x5 in FIG2a, the head circumference x16 (not shown in the figure), and the shoulder circumference x17 (not shown in the figure); the neck parameters include, for example, x6 to x8 in FIG2a; the shoulder parameters include, for example, x9 to x13 in FIG2a; the auricle parameters include, for example, d1 to d8 and α1 to α2 in FIG2b. Alternatively, the human body parameters also include the user's height x14 in an upright state (not shown in the figure) and the user's height x15 in a sitting state (not shown in the figure).
[0087] The processor 20 may obtain the human body parameters by, for example, an external processing device providing the human body parameters and transmitting the human body parameters to the processor 20. For another example, the processor 20 may obtain human body related information and calculate the human body parameters.
[0088] The audio filtering results in the embodiments of the present disclosure include, for example, a left-ear filtering result that matches the user's left ear and / or a right-ear filtering result that matches the user's right ear.
[0089] Specifically, the processor 20 may determine an audio filtering result using a first preset algorithm model based on the association information, human body parameters, and the audio frequency of the virtual sound source. If the association information is the first association information, the audio filtering result is a left-ear filtering result corresponding to the left ear; if the association information is the second association information, the audio filtering result is a right-ear filtering result corresponding to the right ear.
[0090] It should be noted that, when the virtual sound source is pre-constructed, the audio frequency of the virtual sound source is known and stored.
[0091] The first preset algorithm model can be, for example, a multivariate linear regression model or a neural network model. The audio filtering result can be, for example, a frequency response curve, which plots the gain of the virtual audio as it changes with frequency. The frequency response curve data includes data on each frequency and its gain across the full frequency range of the virtual audio.
[0092] The processor 20 is further configured to reconstruct the audio to be played based on the audio pre-configured for the virtual sound source and the audio filtering result.
[0093] Specifically, the processor 20 can reconstruct the audio to be played based on the audio and audio filtering results pre-configured for the virtual sound source using a second preset algorithm model. The second preset algorithm model, for example, integrates at least part of a Fourier transform algorithm, a weighting algorithm, an inverse Fourier transform algorithm, an amplifier, and an equalization algorithm.
[0094] When the audio filtering result is a left-ear filtering result, the audio to be played corresponding to the left ear is reconstructed. When the audio filtering result is a right-ear filtering result, the audio to be played corresponding to the right ear is reconstructed.
[0095] Based on the pre-established spatial information and the user's real-time target positioning information, the processor 20 can quickly determine the association relationship information between the user and the virtual sound source; on this basis, based on the user's real-time association relationship information, human body parameters and the audio frequency of the virtual sound source, the frequency response to be received by the user can be updated in real time, and personalized left-ear filtering results and / or right-ear filtering results can be generated in real time for different users. The audio pre-configured for the virtual sound source is reconstructed based on the user's real-time left-ear filtering result, so that the audio to be played is closer to the natural sound received by the user's left ear; similarly, the audio pre-configured for the virtual sound source is reconstructed based on the user's real-time right-ear filtering result, so that the audio to be played is closer to the natural sound received by the user's right ear, thereby optimizing the user's spatial audio experience, combining with the virtual reality scene, enhancing the user's sense of immersion and 3D surround feeling, and improving the user experience.
[0096] The following is a detailed introduction to each functional module in the embodiment of the present disclosure.
[0097] In some embodiments, the human body parameters include the user's auricle parameters, as shown in FIG2 b .
[0098] FIG3 is a schematic diagram of an exemplary audio processing device provided by an embodiment of the present disclosure. As shown in FIG3 , the audio processing device further includes a depth camera 30 , wherein:
[0099] The depth camera 30 is configured to collect an auricle depth map of the user's auricle and send the map to the processor 20 .
[0100] Exemplarily, the depth camera 30 is set on the VR headset shell or temples. The depth camera 30 includes, for example, a binocular camera, a TOF or a structured light depth camera. No specific limitation is made here. As long as the device can obtain a point cloud map or a depth map, it is within the scope of protection of this disclosure.
[0101] Figure 4 is a schematic diagram of adjusting the position of a depth camera on a VR device according to an embodiment of the present disclosure. As shown in Figure 4, the position of depth camera 30 on the VR device is adjusted to ensure that depth camera 30 can capture all auricle parameters of the user's auricle. The relationship between the horizontal and vertical extension distances of depth camera 30 relative to the VR headset and the auricle parameters is shown in Table 1.
[0102] Table 1
[0103] The horizontal and vertical directions are based on the position of the temple directly above the ear as the origin, with the left direction being the horizontal positive direction and the downward direction being the vertical positive direction. A "√" indicates that the depth camera 30 is located at a position where the corresponding auricle parameters can be obtained; an "×" indicates that the depth camera 30 is located at a position where the corresponding auricle parameters cannot be obtained.
[0104] As can be seen from Table 1, the depth camera 30 is installed within a range of 10 cm to 15 cm horizontally or 3 cm to 7 cm vertically from the VR headset, and all the auricle parameters of the user can be obtained.
[0105] For example, the depth camera 30 is installed at a distance of 15 cm from the VR head display in the horizontal direction. For example, the depth camera 30 is installed at a distance of 5 cm from the VR head display in the vertical direction.
[0106] The processor 20 is further configured to receive the auricle depth map and determine the auricle parameters of the user based on the depth information of the auricle in the auricle depth map.
[0107] The auricle depth map is accompanied by the depth information of the auricle features, and the user's auricle parameters are determined according to the ratio of depth to standard size.
[0108] Exemplarily, the processor 20 is, for example, a VR head-mounted display processor.
[0109] In this embodiment, the auricle feature map collected by the depth camera 30 can directly extract the corresponding auricle parameters. This method is direct and efficient and can quickly obtain the auricle parameters.
[0110] In some embodiments, the human body parameters include auricle parameters and head parameters of the user, as shown in Figures 2a and 2b.
[0111] FIG5 is a schematic diagram of another exemplary audio processing device provided by an embodiment of the present disclosure. As shown in FIG5 , the audio processing device further includes a depth camera 40, wherein:
[0112] The depth camera 40 is configured to collect a depth map of the user's head and send it to the processor 20; the head depth map includes at least left ear features and right ear features.
[0113] Exemplarily, the depth camera 40 is set on the VR handle. The depth camera 40 includes, for example, a binocular camera, a TOF or a structured light depth camera. No specific limitation is made here. As long as the device can obtain a point cloud map or a depth map, it is within the scope of protection of this disclosure.
[0114] The processor 20 is further configured to determine head parameters and auricle parameters of the user based on the depth information of the head and the depth information of the auricle in the head depth map.
[0115] The head depth map includes depth information of head features, and the user's head parameters are determined by the ratio of depth to standard size. The ear pinna depth map includes depth information of ear pinna features, and the user's ear pinna parameters are determined by the ratio of depth to standard size.
[0116] Exemplarily, the processor 20 is, for example, a VR handle processor or a VR head display processor.
[0117] In this embodiment, since the VR handle and the VR headset are designed to be separated, the VR handle has a larger movable space, which allows multi-angle shooting, thereby obtaining more comprehensive human body parameters, such as auricle parameters, head parameters, neck parameters and shoulder parameters, thereby improving the accuracy of subsequent audio filtering results.
[0118] In some embodiments, FIG6 is a schematic diagram of another exemplary audio processing device provided by embodiments of the present disclosure. As shown in FIG6 , the audio processing device further includes a receiving port 50 ; the receiving port 50 is electrically connected to an external processing device. The receiving port 50 is configured to obtain the user's body parameters uploaded by the external processing device and transmit them to the processor 20 .
[0119] Exemplarily, the external processing device is a device with certain computing capabilities. The external processing device captures an image of the user and determines the dimensions of the head, neck, shoulders, and auricle based on a reference object of known size, features of the reference object in the image, and features of the head, neck, shoulders, and auricle in the image. For example, if the width of the reference object is A, the width of the reference object in the image is B pixels, and the width of a certain parameter of the auricle in the image is C pixels, the actual dimension of the auricle parameter can be determined to be A×C÷B.
[0120] For example, the external processing device is a mobile phone, which has both camera and processing functions. Alternatively, the external processing device is a monocular camera integrated into the handle, and the processor on the handle is used to calculate the body parameters. Alternatively, the external processing device is a monocular camera integrated into the handle, and the processor on the head-mounted display is used to calculate the body parameters.
[0121] This embodiment uses relatively few cameras, and compared to depth camera solutions, monocular cameras are more cost-effective. Furthermore, this embodiment fully utilizes external high-performance processing resources, conserving the computing power of the headset processor, enabling rapid acquisition of human parameters and improving the overall efficiency of audio processing.
[0122] In some embodiments, FIG7 is a schematic diagram of a positioning device provided by an embodiment of the present disclosure. As shown in FIG7 , the positioning device 10 specifically includes an inertial measurement unit IMU and a positioning sensor 101, wherein:
[0123] The inertial measurement unit (IMU) is configured to track and locate the user in real space in real time, and determine first positioning information of the user. The first positioning information includes, for example, coordinates and posture information in real space.
[0124] Exemplarily, the first positioning information includes, for example, information on the spatial coordinates and posture of the left ear, and / or information on the spatial coordinates and posture of the right ear.
[0125] However, IMU positioning estimation has cumulative errors. The longer the working time, the more obvious the zero bias and drift phenomena are. Therefore, in order to obtain more accurate target positioning information, IMU generally needs to be integrated with other sensors (i.e., positioning sensor 101) for positioning to compensate and correct the IMU errors.
[0126] The positioning sensor 101 is configured to periodically detect changes in the real scene at a preset interval, determine second positioning information of a target object in the real scene, and transmit the information to the processor 20. The target object may be, for example, a reference object such as a specific marker or a fixed sound source in the real scene. The second positioning information may include, for example, spatial coordinates and posture information.
[0127] The processor 20 determines the target positioning information of the user, specifically including receiving the first positioning information and the second positioning information, and determining the target positioning information based on the first positioning information, the second positioning information and the pre-established spatial information.
[0128] Specifically, the information to be calibrated can be determined based on the first positioning information and the second positioning information; the first positioning information can be calibrated using the information to be calibrated to determine the precise positioning information; based on the spatial information, the precise positioning information can be mapped to the virtual space to determine the target positioning information of the user in the virtual space.
[0129] When the first positioning information is the positioning information of the left ear, the precise positioning information obtained by calibration is the positioning information of the left ear, and the corresponding target positioning information is the target positioning information of the left ear. When the first positioning information is the positioning information of the right ear, the precise positioning information obtained by calibration is the positioning information of the right ear, and the corresponding target positioning information is the target positioning information of the right ear.
[0130] Exemplarily, the audio processing device pre-stores the actual positioning information of the target object marked in the real scene. The calibration information is, for example, the difference between the actual positioning information of the target object and the second positioning information. The difference between the actual positioning information of the target object and the second positioning information is used to assist in correcting the first positioning information measured by the IMU, thereby determining more accurate and precise positioning information. For example, the first positioning information measured by the IMU is corrected by using the spatial displacement difference; or, the first positioning information measured by the IMU is corrected by using a certain posture difference; or, the first positioning information measured by the IMU is corrected by using a weighted value of the spatial displacement difference and the posture difference.
[0131] Exemplarily, the positioning sensor 101 includes, but is not limited to, an optical camera fixed to the VR headset housing, a wireless positioning device integrated within the VR headset, and a microphone array fixed to the earphones, headset housing, or temples. The optical camera may be, for example, a monocular camera, a binocular camera, a multi-lens camera, a visible light camera, an infrared camera, or a depth camera such as a time-of-flight (TOF) or structured light camera. No specific limitations are imposed herein; any optical sensor capable of detecting changes in a target object falls within the scope of protection of this disclosure. The wireless positioning device may be, for example, a GPS, GNSS, or BDS.
[0132] This embodiment uses the positioning sensor 101 to correct and compensate for the drift of the IMU, thereby improving the accuracy of the precise positioning information.
[0133] In some embodiments, the association relationship information includes a relative distance between the user and the virtual sound source, a relative pitch angle between the user and the virtual sound source, and a relative horizontal angle between the user and the virtual sound source.
[0134] The processor 20 determines the relative distance r in the association relationship information, specifically including: determining the relative distance r between the user and the virtual sound source based on the first position coordinates of the user indicated in the user's real-time target positioning information and the second position coordinates of the virtual sound source indicated in the spatial information. Exemplarily, the first position coordinates are, for example, the spatial coordinates r1 (x1, y1, z1) under the virtual space coordinate system O-XYZ. The second position coordinates include, for example, the spatial coordinates r2 (x2, y2, z2) under the virtual space coordinate system O-XYZ. Based on this, the relative distance between the user and the virtual sound source is determined. It should be noted that the user's first position coordinates here can be the first position coordinates of the left ear or the first position coordinates of the right ear. If r1(x1, y1, z1) represents the first position coordinates of the left ear contour center, r represents the relative distance between the left ear contour center and the virtual sound source. If r1(x1, y1, z1) represents the first position coordinates of the right ear contour center, r represents the relative distance between the right ear contour center and the virtual sound source.
[0135] The processor 20 determines the relative pitch angle in the association relationship information Specifically, the method includes: determining the relative pitch angle between the user and the virtual sound source according to the first pitch angle of the user indicated in the user's real-time target positioning information and the second pitch angle of the virtual sound source indicated in the spatial information; For example, the first pitch angle is, for example, the pitch angle corresponding to the two-dimensional coordinate system XOZ and YOZ in the virtual space coordinate system O-XYZ. The second pitch angle is, for example, the pitch angle corresponding to the two-dimensional coordinate systems XOZ and YOZ in the virtual space coordinate system O-XYZ. Based on this, the relative pitch angle between the user and the virtual sound source is determined Here, the user's first pitch angle can be the first pitch angle of the left ear or the first pitch angle of the right ear. represents the first pitch angle of the left ear, then represents the relative pitch angle between the left ear and the virtual sound source. represents the first pitch angle of the right ear, then Indicates the relative pitch angle between the right ear and the virtual sound source.
[0136] The processor 20 determines the relative horizontal angle θ in the association relationship information, specifically by determining the relative horizontal angle between the user and the virtual sound source based on the user's first horizontal angle indicated in the user's real-time target positioning information and the second horizontal angle of the virtual sound source indicated in the spatial information. Exemplarily, the first horizontal angle is the horizontal angle θ1 corresponding to the two-dimensional coordinate system XOY in the virtual spatial coordinate system O-XYZ. The second horizontal angle is the horizontal angle θ2 corresponding to the two-dimensional coordinate system XOY in the virtual spatial coordinate system O-XYZ. Based on this, the relative horizontal angle θ between the user and the virtual sound source is determined to be θ=θ1-θ2. Here, the user's first horizontal angle can be the first horizontal angle of the left ear or the first horizontal angle of the right ear. If θ1 represents the first horizontal angle of the left ear, then θ represents the relative horizontal angle between the left ear and the virtual sound source. If θ1 represents the first horizontal angle of the right ear, then θ represents the relative horizontal angle between the right ear and the virtual sound source.
[0137] In some embodiments, the processor 20 determines the audio filtering results that match the user, specifically including screening out the audio filtering results that match the user from a pre-set audio filtering database based on the association relationship information, human body parameters and the audio frequency of the virtual sound source.
[0138] Exemplarily, the human body parameters include at least the user's auricle parameters; the auricle parameters include left ear auricle parameters and right ear auricle parameters. When the first preset algorithm model is a multiple linear regression model, the processor 20 is specifically configured to filter out a left ear filter result that matches the user's left ear from a preset audio filter database based on the first association relationship information, other parameters of the human body parameters except the right ear auricle parameters, and the audio frequency of the virtual sound source; and filter out a right ear filter result that matches the user's right ear from a preset audio filter database based on the second association relationship information, other parameters of the human body parameters except the left ear auricle parameters, and the audio frequency of the virtual sound source.
[0139] The audio filter database (also known as the HRTF database) includes a variety of pre-generated reference audio filter data. The reference audio filter data, also known as HRTF data, is determined based on the HRTF algorithm. Each HRTF data in the audio filter database has a corresponding set of condition information (r, θ, f, a).
[0140] Based on the HRTF algorithm, the process of determining HRTF data is shown in the following formula 1:
[0141] Among them, H L Indicates the HRTF data of the left ear; H R represents the HRTF data of the right ear; P L Indicates the sound pressure intensity of the signal actually received by the left ear; P R represents the sound pressure intensity of the actual signal received by the right ear; P0 represents the sound pressure intensity of the actual speaker output signal; r represents the relative distance between the user and the virtual sound source; θ represents the relative horizontal angle between the user and the virtual sound source; represents the relative pitch angle between the user and the virtual sound source; f represents the audio frequency of the virtual sound source; a represents the human body parameter.
[0142] For example, the human body parameter a in this embodiment is, for example, an auricle parameter, which improves the processing efficiency of the deep neural network compared to when all parameters are involved in the calculation.
[0143] For example, the human body parameter a in this embodiment includes some key head parameters and auricle parameters. Compared with all parameters participating in the calculation, it can not only ensure the accuracy of the output left ear filtering results and the right ear filtering results, but also improve the processing efficiency of the deep neural network.
[0144] In one case, in order to ensure that the left ear filtering results and the right ear filtering results are more accurate, the human body parameters include auricle parameters, head parameters, neck parameters and shoulder parameters. For example, the relative distance r, relative horizontal angle θ, relative pitch angle The human body parameters a (auricle parameters, head parameters, neck parameters and shoulder parameters) and the audio frequency f of the virtual sound source are input into the multivariate linear regression model, and combined with the HRTF data in the HRTF database and its corresponding condition information (r, θ, f. a), perform statistical analysis and filter out the audio filtering results that match the user's left or right ear from the HRTF database.
[0145] It should be noted that the r, θ, and f are related to r, θ, The same as f, the human body parameters a are at least partially the same. For example, the human body parameters a corresponding to the HRTF data only need to match the ear parameters or some key parameters such as the head, and at the same time satisfy r, θ, The same as f, which can be used as the audio filtering result matched with the user.
[0146] In another embodiment, each HRTF data in the HRTF database is pre-set as dimensionality-reduced HRTF data Hn×m. Specifically, the HRTF data after dimensionality reduction is reconstructed based on principal component analysis (PCA) or independent component analysis (ICA), and related HRTF dimensionality reduction and extraction algorithms such as weight coefficients.
[0147] For example, the independent component analysis algorithm ICA and related HRTF dimensionality reduction and extraction algorithms such as weight coefficients are used to reconstruct the reduced dimensionality HRTF data. The calculation process of the reduced dimensionality HRTF data is shown in the following formula 2:
[0148] in, A represents HRTF data, which is also a reference audio filter data in the audio filter database. n×d Represents the weight coefficient matrix, including multiple dimensionality reduction weights. IC d×m Represents the human body parameter matrix, including some key physiological parameters in the human body parameters. Formula 2 can reflect that the HRTF data after dimension reduction is related to the human body parameter matrix IC and the weight coefficient matrix A. A with unique corresponding n×d .
[0149] Specifically, the relative distance r, the relative horizontal angle θ, and the relative pitch angle Some key physiological parameters a1 in human body parameters and the audio frequency f of virtual sound source are input into the multivariate linear regression model, combined with each HRTF data in the HRTF database and its corresponding condition information (r, θ, f, a1), and statistical analysis can be performed to filter out audio filtering results that match the user from the dimensionality-reduced HRTF database. For example, first, the multivariate linear regression model can preliminarily filter out the audio filtering results that match the input condition information (r, θ, f) Matching part Secondly, from this part Based on input a1 and weight coefficient matrix A n×d , use formula 2 to determine a candidate Calculate the candidate and A n×d Corresponding The similarity of Filter out matching
[0150] The "dimensionality reduction" disclosed in the present invention can be understood as reducing the number of physiological parameters in the human body parameters involved in the calculation. Therefore, the HRTF database with reduced dimension is used to improve the regression efficiency of the multivariate linear regression model while ensuring the accuracy of the audio filtering results.
[0151] Figure 8 is a system architecture diagram of the audio processing device provided in an embodiment of the present disclosure for determining the audio filtering results. As shown in Figure 8, it includes a pre-built virtual space, a depth camera 40, an inertial measurement unit IMU, a positioning sensor 101, a processor 20, and a HRTF database after dimensionality reduction.
[0152] Exemplarily, the pre-built virtual space is a virtual space pre-built by the VR processor and stored in the memory. The HRTF database after dimensionality reduction is a resource stored in the memory. The processor 20 is integrated with a multivariate linear regression model.
[0153] For example, as shown in FIG8 , the pre-built virtual space can provide the sound source positioning information of the virtual sound source (for example, the second position coordinate r2 (x2, y2, z2), the horizontal angle θ2 and the pitch angle ) and audio frequency f. The depth camera 40 can obtain a head depth map of the user's auricle. The processor 20 can determine the user's body parameter a. The inertial measurement unit IMU can track and determine the user's first positioning information in real time. The positioning sensor 101 can determine the second positioning information of the target object in the real scene. The processor 20 uses the second positioning information to assist in verifying the first positioning information, obtains precise positioning information, and maps the precise positioning information to the virtual space to obtain target positioning information. The processor 20 can also calculate the target positioning information based on the target positioning information (such as the first position coordinate r1 (x1, y1, z1), the horizontal angle θ1 and the pitch angle ) and sound source localization information (such as the second position coordinate r2 (x2, y2, z2), horizontal angle θ2 and pitch angle ), determine in real time the relationship information between the user and the virtual sound source in the virtual space (such as relative distance r, relative horizontal angle θ and relative pitch angle ). And, the processor 20 receives r, θ, f, a, using the multiple linear regression model and combining the HRTF database and A n×d , filter out the marriage filter result with the highest matching degree from the HRTF database.
[0154] In some embodiments, the processor 20 determines the audio filtering result that matches the user, specifically including: inputting the association relationship information, human body parameters and the audio frequency of the virtual sound source into a pre-trained target network model, and outputting the audio filtering result that matches the user.
[0155] Exemplarily, the human body parameters include at least the user's auricle parameters; the auricle parameters include at least left ear auricle parameters and right ear auricle parameters. When the first preset algorithm model is a neural network model, the processor 20 inputs the first association relationship information, the human body parameters, and the audio frequency of the virtual sound source into a pre-trained target network model, and outputs a left ear filtering result that matches the user's left ear; and inputs the second association relationship information, the human body parameters, and the audio frequency of the virtual sound source into the pre-trained target network model, and outputs a right ear filtering result that matches the user's right ear.
[0156] Among them, the pre-trained target network model is trained based on various reference audio filter data (ie HRTF data) in a pre-set audio filter database and its corresponding human body parameters, association relationship information and audio frequency.
[0157] Exemplarily, the target network model is, for example, a deep neural network (DNN).
[0158] For example, in this embodiment, the human body parameter a input into the target network model is, for example, an auricle parameter, which improves the processing efficiency of the deep neural network compared to when all parameters participate in the calculation.
[0159] For example, in this embodiment, the human body parameters a input into the target network model include some key head parameters and auricle parameters. Compared with all parameters participating in the calculation, it can not only ensure the accuracy of the output left ear filtering results and the right ear filtering results, but also improve the processing efficiency of the deep neural network.
[0160] Exemplarily, the output audio filtering result is personalized HRTF data determined by combining human body parameters using the HRTF algorithm, which is expressed as a frequency response curve for the virtual sound source, including gain data of the virtual sound source at all frequencies.
[0161] FIG9 is a model framework diagram of the target network model provided by an embodiment of the present disclosure. As shown in FIG9 , the target network model includes a first sub-network 21, a second sub-network 22, and a third sub-network 23. The first sub-network 21 includes a normalization layer, an input layer, and three hidden layers connected in sequence. The second sub-network 22 includes an input layer and a tensor flatten operation connected in sequence. The third sub-network 23 includes three hidden layers connected in sequence.
[0162] As shown in Figure 9, the human body parameters are input into the first sub-network 21; the relationship association information is input into the second sub-network 22; the output results of the first sub-network 21 and the output results of the second sub-network 22 are input into the third sub-network 23; and the audio filtering results personalized with the user are output from the third sub-network 23.
[0163] FIG10 is a system structure diagram of another audio processing device for determining audio filtering results provided by an embodiment of the present disclosure. As shown in FIG10 , the device includes a pre-built virtual space, a depth camera 30, an inertial measurement unit (IMU), a positioning sensor 101, a processor 20, and a reduced HRTF database. The target network model is integrated into the processor 20.
[0164] For example, as shown in FIG10 , the pre-built virtual space can provide the sound source positioning information of the virtual sound source (for example, the second position coordinate r2 (x2, y2, z2), the horizontal angle θ2 and the pitch angle ) and audio frequency f. The depth camera 30 can obtain a depth map of the user's auricle. The processor 20 can determine the user's body parameter a. The inertial measurement unit IMU can track and determine the user's first positioning information in real time. The positioning sensor 101 can determine the second positioning information of the target object in the real scene. The processor 20 uses the second positioning information to assist in verifying the first positioning information to obtain more accurate target positioning information. The processor 20 can calculate the target positioning information based on the target positioning information (such as the first position coordinate r1 (x1, y1, z1), the horizontal angle θ1 and the pitch angle ) and sound source localization information (such as the second position coordinate r2 (x2, y2, z2), horizontal angle θ2 and pitch angle ), determine in real time the relationship information between the user and the virtual sound source in the virtual space (such as relative distance r, relative horizontal angle θ and relative pitch angle ). The processor 20 can receive r, θ, f. a. Utilize the target network model pre-trained based on the HRTF database to output audio filtering results that are personalized to the user.
[0165] In some embodiments, the audio filtering result includes a left-ear filtering result and a right-ear filtering result; and the audio to be played includes the audio to be played corresponding to the left ear and the audio to be played corresponding to the right ear.
[0166] Processor 20 reconstructs the audio to be played by performing a Fourier transform on the audio pre-configured for the virtual sound source to obtain frequency response information for the audio. This frequency response information includes a frequency response curve, i.e., gain data for the audio emitted by the virtual sound source across all frequencies. The audio to be played for the left ear is determined based on the frequency response information, the left-ear filtering results, and the relative distance between the user and the virtual sound source. The audio to be played for the right ear is determined based on the frequency response information, the right-ear filtering results, and the relative distance between the user and the virtual sound source.
[0167] Determining the audio to be played for the left ear specifically includes multiplying the frequency response information with the gain at the corresponding frequency in the left-ear filtering result to obtain a first intermediate result. As can be seen from the above description, the left-ear filtering result is personalized HRTF data determined using a HRTF algorithm in conjunction with human body parameters. It is represented by a frequency response curve for the virtual sound source, including gain data for all frequencies of the virtual sound source. Therefore, using the product algorithm integrated in the first filtering subunit 1421, the gain at each frequency indicated by the frequency response information is multiplied by the gain at each frequency indicated by the left-ear filtering result to obtain a product result corresponding to each frequency, a new frequency response curve, denoted as the first intermediate result. The first intermediate result is then subjected to an inverse Fourier transform (IFFT) to obtain first audio time domain information. Here, the frequency domain signal represented by the first intermediate result is converted to a time domain signal through an inverse Fourier transform (IFFT), denoted as the first audio time domain information. The first audio time domain information is then processed based on the relative distance between the user and the virtual sound source to obtain the left-ear audio signal. Exemplarily, based on the relative distance r=|r2-r1| between the center of the left ear auricle and the virtual sound source, and the relationship formula between the sound pressure level difference and the distance difference LP1-LP2=20lg(r1 / r2), the sound pressure gain corresponding to the distance difference is scaled (gain increased) by the first amplifier 1423 to truly simulate the sound intensity at different distances and obtain the left ear audio signal. Here, r1 represents the distance between the center of the user's left ear contour and the origin O in the virtual space coordinate system. Afterwards, the left ear audio signal is processed to obtain the audio to be played for the left ear. For example, the left ear audio signal is equalized using the equalization algorithm integrated in the processor 20 so that the sound pressure intensity of the audio to be played is uniform across the entire frequency band. Exemplarily, the audio to be played for the left ear can be played through the left ear speaker. The left ear speaker can be a left ear speaker in an external headset, or a left ear speaker integrated in an audio processing device.
[0168] Determining the audio to be played for the right ear specifically includes multiplying the frequency response information with the gain at the corresponding frequency in the right-ear filtering result to obtain a second intermediate result. As can be seen from the above description, the right-ear filtering result is personalized HRTF data determined using a HRTF algorithm in conjunction with human body parameters. It is expressed as a frequency response curve for the virtual sound source, including gain data for all frequencies of the virtual sound source. Therefore, using the multiplication algorithm integrated into the second filtering subunit 1431, the gain at each frequency indicated by the frequency response information is multiplied by the gain at each frequency indicated by the right-ear filtering result to obtain a product result corresponding to each frequency, a new frequency response curve, denoted as the second intermediate result. The second intermediate result is then subjected to an inverse Fourier transform (IFFT) to obtain second audio time domain information. Here, the frequency domain signal represented by the first intermediate result is converted to a time domain signal through an inverse Fourier transform (IFFT), denoted as the first audio time domain information. The second audio time domain information is then processed based on the relative distance between the user and the virtual sound source to obtain the right-ear audio signal. Exemplarily, based on the relative distance r=|r2-r1| between the center of the right ear auricle and the virtual sound source, and the relationship formula between the sound pressure level difference and the distance difference LP1-LP2=20lg(r1 / r2), the sound pressure gain corresponding to the distance difference is scaled (gain increased) by the second amplifier 1433 to truly simulate the sound intensity at different distances and obtain the right ear audio signal. Here, r1 represents the distance between the center of the user's right ear contour and the origin O in the virtual space coordinate system. Afterwards, the right ear audio signal is processed to obtain the audio to be played for the right ear. For example, the right ear audio signal is equalized using the equalization algorithm integrated in the processor 20 so that the sound pressure intensity of the audio to be played is uniform across the entire frequency band. Exemplarily, the audio to be played for the right ear can be played through the right ear speaker. The right ear speaker can be the right ear speaker in an external headset, or it can be the right ear speaker integrated in the audio processing device.
[0169] Exemplarily, FIG11 is a system architecture diagram for personalized generation of left-ear arrest audio and right-ear to-be-played audio provided by an embodiment of the present disclosure. As shown in FIG11 , it includes a depth camera 30, an inertial measurement unit IMU, a positioning sensor 101, a processor 20, a left-ear speaker 61, and a right-ear speaker 61.
[0170] Among them, the depth camera 40 is used to collect the head depth map; the processor 20 is used to extract the human body parameter a. The inertial measurement unit IMU is used to collect the user's first positioning information, and the positioning sensor 101 is used to collect the target object's second positioning information. The processor 20 is also used to use the second positioning information to assist in verifying the first positioning information, obtain precise positioning information, and map the precise positioning information to the virtual space to obtain target positioning information; determine the association relationship information between the user and the virtual sound source; determine the personalized left ear filtering result (left ear HRTF data) and the right ear filtering result (right ear HRTF data). For the processing process of determining one of the left ear audio to be played and the right ear audio to be played, specifically, the processor 20 includes combining the audio filtering result and the Fourier transform (FFT) result of the audio configured for the virtual sound source for multiplication, and sequentially performing inverse Fourier transform (IFFT), scaling processing and equalization processing to output the audio to be played. Among them, for the left ear audio to be played, the left ear speaker 61 is used for playback; for the right ear audio to be played, the right ear speaker 62 is used for playback.
[0171] In addition, an embodiment of the present disclosure further provides an extended reality device, which includes an audio processing device as described in any one of the above embodiments.
[0172] Extended reality devices, such as VR / RA devices, typically include a headset and controllers. The headset is used for data processing and display, while the controllers are used for virtual operation and control. Compared to conventional structures, the VR device provided by this disclosure incorporates an audio processing device. Some of the functions of this audio processing device are implemented in the original VR device, some are integrated into the headset's processor, and some are implemented in additional hardware.
[0173] Figure 12 is a schematic diagram of a VR device provided by an embodiment of the present disclosure. As shown in Figure 12, the VR device includes an audio processing device 100 and a speaker 200; wherein, the audio processing device 100 includes a positioning device 10 and a processor 20.
[0174] The positioning device 10 is configured to track and locate the user in the real space in real time, determine the user's positioning information, and send it to the processor 20; the processor 20 is configured to determine the user's target positioning information based on the pre-established spatial information and the user's real-time positioning information; determine the association relationship information between the user and the virtual sound source in the virtual space based on the pre-established spatial information and the user's real-time target positioning information; obtain the user's human body parameters, and determine the audio filtering results that match the user based on the association relationship information, the human body parameters and the audio frequency of the virtual sound source; reconstruct the audio to be played based on the audio pre-configured for the virtual sound source and the audio filtering results; the speaker 200 is configured to receive the audio to be played and play it.
[0175] Exemplarily, the speaker 200 includes a left ear speaker 61 and a right ear speaker 62. The audio to be played for the left ear is played through the left ear speaker 61. The audio to be played for the right ear is played through the right ear speaker.
[0176] In some embodiments, as shown in Figure 12, the positioning device 10 includes an inertial measurement unit IMU and a positioning sensor 101; the positioning sensor 101 includes one of an optical camera, a wireless positioning device and a microphone array; the inertial measurement unit IMU is located in the accommodation space defined by the outer shell of the extended reality device (not shown in the figure); the positioning sensor 101 is fixed to the outside of the outer shell.
[0177] Optical cameras include, for example, monocular cameras, binocular cameras, multi-lens cameras, visible light cameras, infrared cameras, and other cameras. They may also be depth cameras such as time-of-flight (TOF) and structured light cameras. These are not specifically limited here; any optical sensor capable of detecting changes in a target object falls within the scope of protection of this disclosure. Wireless positioning devices include, for example, GPS, GNSS, or BDS.
[0178] In some embodiments, as shown in FIG12 , the audio processing device further includes a depth camera 30 ; the depth camera 30 is fixed to the outside of the outer shell of the augmented reality device.
[0179] The depth camera 30 is fixed to the head-mounted display or the temples of the glasses. The depth camera 30 includes, for example, a binocular camera, a time-of-flight (TOF) camera, or a structured light camera. No specific limitation is imposed herein. Any device capable of obtaining a point cloud image or a depth image is within the scope of protection of this disclosure.
[0180] In some embodiments, the audio processing device further includes a depth camera 40 and a handle 70 communicatively connected to the processor 20; the depth camera 40 is fixed to the handle 70. The depth camera 40 may include, for example, a binocular camera, a time-of-flight (TOF) camera, or a structured light camera, and is not specifically limited here. Any device that can obtain a point cloud image or a depth map is within the scope of protection of this disclosure.
[0181] The processor 20 here is, for example, a head display processor located in the head display shell, or a handle processor located in the handle.
[0182] When the processor is a handle processor, the handle processor is communicatively connected to the head display processor, and the human body parameters determined by the feature extraction unit 112 can be uploaded to the head display processor.
[0183] In addition, the present disclosure further provides an audio processing method. FIG13 is a flow chart of an audio processing method provided by the present disclosure. As shown in FIG13 , the audio processing method specifically includes steps S11 to S15, wherein:
[0184] S11. Track and locate the user in the real space in real time, determine the user's location information, and send it to the processor.
[0185] S12. Determine the user's target positioning information based on the pre-built spatial information and the user's real-time positioning information.
[0186] S13. Determine the association relationship information between the user and the virtual sound source in the virtual space based on the pre-established space information and the user's real-time target positioning information.
[0187] S14: Obtain the user's body parameters, and determine an audio filtering result that matches the user based on the association relationship information, the body parameters, and the audio frequency of the virtual sound source.
[0188] S15: reconstruct the audio to be played based on the audio pre-configured for the virtual sound source and the audio filtering result.
[0189] For the above steps S11 to S15 , please refer to the specific implementation process of the positioning device 10 and the processor 20 in the above audio processing device, and the repeated parts will not be repeated.
[0190] The audio processing method provided by the embodiment of the present disclosure is mainly applied to virtual reality scenes, and can quickly obtain the user's human body parameters; it can track and locate the user in the virtual space in real time, thereby determining the user's target positioning information. Based on the pre-established spatial information and the user's real-time target positioning information, the association relationship information between the user and the virtual sound source can be quickly determined. The association relationship information can be understood as information about the relative position change between the user and the virtual sound source; on this basis, the frequency response to be received by the user can be updated in real time based on the user's real-time association relationship information, human body parameters and the audio frequency of the virtual sound source, and personalized audio filtering results can be generated in real time for different users. In addition, based on the user's real-time audio filtering results, the audio pre-configured for the virtual sound source is reconstructed, so that the reconstructed audio to be played is closer to the natural sound received by the user's human ear, thereby optimizing the user's spatial audio experience, combining with the virtual reality scene, enhancing the user's immersion and 3D surround feeling, and improving the user experience.
[0191] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0192] Figure 14 is a schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. As shown in Figure 14, the computer device provided in an embodiment of the present disclosure includes: one or more processors 201, memory 202, and one or more I / O interfaces 203. Memory 202 stores one or more programs. When executed by the one or more processors, the one or more processors implement the audio processing method described in any of the above embodiments. One or more I / O interfaces 203 are connected between the processors and memory and are configured to facilitate information exchange between the processors and memory.
[0193] Among them, the processor 201 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory 202 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read-write interface) 203 is connected between the processor 201 and the memory 202, and can realize information interaction between the processor 201 and the memory 202, including but not limited to a data bus (Bus), etc.
[0194] In some embodiments, the processor 201 , the memory 202 , and the I / O interface 203 are connected to each other via a bus 204 , and further connected to other components of the computing device.
[0195] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium is further provided, wherein the non-transitory computer-readable storage medium stores a computer program, wherein when the program is executed by a processor, the steps of the audio processing method in any of the above embodiments are implemented.
[0196] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), the above-mentioned functions defined in the system of the present disclosure are executed.
[0197] It should be noted that the computer non-transitory readable medium shown in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any non-transitory computer-readable storage medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the non-transitory computer-readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination thereof.
[0198] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architectures, functions and operations of the devices, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the aforementioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two connected boxes can actually represent execution in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0199] It is understood that the above embodiments are merely exemplary embodiments for illustrating the principles of the present disclosure, and the present disclosure is not limited thereto. Those skilled in the art may make various modifications and improvements without departing from the spirit and substance of the present disclosure, and such modifications and improvements are also considered to be within the scope of protection of the present disclosure.
Claims
1. An audio processing device, in, include: A positioning device, configured to track and locate a user in a real space in real time, determine the positioning information of the user, and send it to a processor; The processor is configured to determine the target location information of the user based on the pre-built space information and the real-time location information of the user; determine the association relationship information between the user and the virtual sound source in the virtual space based on the pre-built space information and the real-time location information of the user; Acquire the human body parameters of the user, and determine an audio filtering result matching the user based on the association relationship information, the human body parameters and the audio frequency of the virtual sound source; The audio to be played is reconstructed based on the audio pre-configured for the virtual sound source and the audio filtering result.
2. The audio processing device according to claim 1, in, The positioning device includes an inertial measurement unit and a positioning sensor; the positioning sensor includes one of an optical camera, a wireless positioning device and a microphone array; the positioning information includes first positioning information and second positioning information; The inertial measurement unit is configured to track and locate a user in real space in real time, determine first positioning information of the user, and send it to the processor; The positioning sensor is configured to obtain the change of the real scene at a regular interval according to a preset period, determine the second positioning information of the target object in the real scene, and send it to the processor; The processor determines the target positioning information of the user, specifically including receiving the first positioning information and the second positioning information, and determining the target positioning information based on the user's real-time first positioning information, the second positioning information and pre-built spatial information.
3. The audio processing device according to claim 2, in, The processor determines the target positioning information of the user, specifically including receiving the real-time first positioning information and the second positioning information of the user, and determining calibration information based on the real-time first positioning information and the second positioning information of the user; Using the calibration information to calibrate the first positioning information to determine the precise positioning information; Based on the pre-built space information, the precise positioning information is mapped to the virtual space to determine the target positioning information.
4. The audio processing device according to claim 1, in, The human body parameters include auricle parameters of the user; the audio processing device also includes a depth camera; The depth camera is configured to collect an auricle depth map of the user's auricle and send it to the processor; The processor is further configured to receive the auricle depth map, and determine the user's auricle parameters based on the depth information of the auricle in the auricle depth map.
5. The audio processing device according to claim 1, in, The human body parameters include auricle parameters and head parameters of the user; The audio processing device also includes a depth camera; The depth camera is configured to collect a head depth map of the user and send it to the processor; the head depth map includes at least left ear features and right ear features; The processor is further configured to receive the head depth map, and determine the head parameters and auricle parameters of the user based on the depth information of the head and the depth information of the auricle in the head depth map.
6. The audio processing device according to claim 1, in, The audio processor device also includes a receiving port; the receiving port is electrically connected to an external processing device; The receiving port is configured to obtain the human body parameters of the user uploaded by the external processing device and send them to the processor.
7. The audio processing device according to any one of claims 1 to 6, in, The human body parameters include auricle parameters, head parameters, neck parameters and shoulder parameters.
8. The audio processing device according to claim 1, in, The association relationship information includes a relative distance between the user and the virtual sound source, a relative pitch angle between the user and the virtual sound source, and a relative horizontal angle between the user and the virtual sound source; The processor determines the association relationship information, specifically including determining a relative distance between the user and the virtual sound source according to a first position coordinate of the user indicated in the real-time target positioning information of the user and a second position coordinate of the virtual sound source indicated in the spatial information; determining a relative pitch angle between the user and the virtual sound source according to a first pitch angle of the user indicated in the real-time target positioning information of the user and a second pitch angle of the virtual sound source indicated in the spatial information; determining a relative pitch angle between the user and the virtual sound source according to a first horizontal angle of the user indicated in the real-time target positioning information of the user and a second horizontal ... The second horizontal angle of the virtual sound source indicated in the information is used to determine the relative horizontal angle between the user and the virtual sound source.
9. The audio processing device according to claim 1, in, The processor determines the audio filtering result matching the user, specifically including screening out the audio filtering result matching the user from a preset audio filtering database based on the association relationship information, the human body parameters and the audio frequency of the virtual sound source.
10. The audio processing device according to claim 1, in, The processor determines an audio filtering result that matches the user, specifically including inputting the association relationship information, the human body parameters, and the audio frequency of the virtual sound source into a pre-trained target network model, and outputting an audio filtering result that matches the user; The pre-trained target network model is obtained by training based on various reference audio filter data and their corresponding human body parameters, the association relationship information and audio frequency in a pre-set audio filter database.
11. The audio processing device according to claim 1, in, The audio filtering result includes a left ear filtering result and a right ear filtering result; the to-be-played audio includes the to-be-played audio corresponding to the left ear and the to-be-played audio corresponding to the right ear; The processor reconstructs the audio to be played, specifically comprising performing Fourier transformation on the audio pre-configured for the virtual sound source to obtain frequency response information of the audio; and determining the audio to be played for the left ear based on the frequency response information, the left ear filtering result, and the relative distance between the user and the virtual sound source; The audio to be played for the right ear is determined based on the frequency response information, the right ear filtering result, and the relative distance between the user and the virtual sound source.
12. The audio processing device according to claim 11, in, The processor determines the audio to be played for the left ear, specifically including multiplying the frequency response information and the gain of the corresponding frequency in the left ear filtering result to obtain a first intermediate result; performing an inverse Fourier transform on the first intermediate result to obtain first audio time domain information; processing the first audio time domain information according to the relative distance between the user and the virtual sound source to obtain a left ear audio signal; and processing the left ear audio signal to obtain the audio to be played for the left ear.
13. The audio processing device according to claim 11, in, The processor determines the audio to be played for the right ear, specifically including multiplying the frequency response information and the gain of the corresponding frequency in the right ear filtering result to obtain a second intermediate result; performing an inverse Fourier transform on the second intermediate result to obtain second audio time domain information; processing the second audio time domain information according to the relative distance between the user and the virtual sound source to obtain a right ear audio signal; and processing the right ear audio signal to obtain the audio to be played for the right ear.
14. An extended reality device, comprising an audio processing device and a speaker as claimed in any one of claims 1 to 13; the audio processing device comprises a positioning device and a processor; The positioning device is configured to track and locate a user in real space in real time, determine the positioning information of the user, and send it to the processor; The processor is configured to determine the target location information of the user based on the pre-built space information and the real-time location information of the user; determine the association relationship information between the user and the virtual sound source in the virtual space based on the pre-built space information and the real-time location information of the user; Acquire the human body parameters of the user, and determine an audio filtering result matching the user based on the association relationship information, the human body parameters and the audio frequency of the virtual sound source; Reconstructing the audio to be played based on the audio pre-configured for the virtual sound source and the audio filtering result; The speaker is configured to receive and play the audio to be played.
15. The extended reality device according to claim 14, in, The positioning device includes an inertial measurement unit and a positioning sensor; the positioning sensor includes one of an optical camera, a wireless positioning device and a microphone array; the inertial measurement unit is located in a housing space defined by an outer shell of the extended reality device; and the positioning sensor is fixed on the outside of the outer shell.
16. The extended reality device according to claim 14, in, The audio processing device also includes a depth camera; the depth camera is fixed on the outside of the outer shell of the extended reality device.
17. The extended reality device according to claim 14, in, The audio processing device also includes a depth camera and a handle that is communicatively connected to the processor; the depth camera is fixed on the handle.
18. The extended reality device according to claim 14, in, The audio processing device further includes a receiving port; the receiving port is electrically connected to the processor.
19. An audio processing method, in, include: Tracking and locating a user in real space in real time, determining the location information of the user, and sending it to a processor; Determine the target location information of the user based on the pre-built spatial information and the real-time location information of the user; Determine the association relationship information between the user and the virtual sound source in the virtual space based on the pre-built space information and the real-time target positioning information of the user; Acquire the human body parameters of the user, and determine an audio filtering result matching the user based on the association relationship information, the human body parameters and the audio frequency of the virtual sound source; The audio to be played is reconstructed based on the audio pre-configured for the virtual sound source and the audio filtering result.
20. A computer device, in, include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the steps of the audio processing method as described in claim 19 are performed.
21. A computer non-transitory readable storage medium, in, The computer non-transitory readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the audio processing method as claimed in claim 19 are executed.