Sound-effect processing device, sound-effect processing method and extended reality apparatus
By using a neckband-style audio processing device and a photometric stereo algorithm, the head width is accurately measured and a three-dimensional model of the human ear is reconstructed, solving the measurement challenge of personalized HRTF and improving the accuracy and experience of audio processing.
Patent Information
- Application Number
- PCT/CN2025/093433
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-25
- Filing Date
- 2025-05-08
- Publication Date
- 2026-01-02
AI Technical Summary
Existing technologies struggle to implement personalized head-related transfer functions (HRTFs), requiring measurements to be performed on each individual, which leads to inconvenience and insufficient accuracy.
Using a neckband-style sound processing device, the device acquires the user's head width and three-dimensional ear model through an image acquisition structure and light source combination. It then uses photometric stereo algorithms and triangle geometry theorems to accurately measure head width and reconstruct the three-dimensional ear model, generating a personalized HRTF function.
It enables precise measurement and sound processing of personalized HRTF, improving the accuracy and effect of the sound experience.
Smart Images

Figure CN2025093433_02012026_PF_FP_ABST
Abstract
Description
Sound effect processing device, sound effect processing method and extended reality device
[0001] Cross-reference to related applications
[0002] The present application claims priority from Chinese Patent Application No. 202410831209.6 filed on June 25, 2024 in China, the contents of which are incorporated herein by reference in its entirety. TECHNICAL FIELD
[0003] The present disclosure relates to the technical field of spatial audio, and in particular to a sound effect processing device, a sound effect processing method and an extended reality device. BACKGROUND
[0004] Spatial audio is implemented based on HRTF, and the head-related transfer function (HRTF) describes the transmission process of sound waves from a sound source to both ears. It is the result of the comprehensive filtering of sound waves by the physiological structure of a person, such as the head, pinna, and torso. The size and shape of the head, pinna, and torso of each person are different, and to achieve personalized HRTF, the relevant parameters of each person need to be measured. SUMMARY
[0005] To solve the above technical problems, the present disclosure provides a sound effect processing device, a sound effect processing method and an extended reality device.
[0006] To achieve the above purpose, the technical solution adopted by the embodiments of the present disclosure is as follows: a sound effect processing device, comprising:
[0007] a neck-mounted body;
[0008] an image acquisition structure, comprising a first image acquisition structure and a second image acquisition structure arranged at two ends of the neck-mounted body, the first image acquisition structure and the second image acquisition structure are respectively configured to acquire corresponding ear images;
[0009] a light source, comprising a first light source arranged around the four sides of the first image acquisition structure and a second light source arranged around the four sides of the second image acquisition structure;
[0010] controlling the first light source to be turned on to make the first image acquisition structure acquire a left ear image, and controlling the second light source to be turned on to make the second image acquisition structure acquire a right ear image;
[0011] a first processing module, configured to acquire human body parameters of a user according to images acquired by the image acquisition structure, to determine an audio input signal corresponding to the user, the parameters including head width and a three-dimensional model of the ear;
[0012] The second processing module is configured to control a loudspeaker to output corresponding sound effects according to the audio input signal.
[0013] Optionally, the first light source comprises a plurality of first sub-light sources capable of independently emitting light, and the plurality of first sub-light sources are configured to be illuminated at different times so that the first image acquisition structure obtains image information of the left ear; the second light source comprises a plurality of second sub-light sources capable of independently emitting light, and the plurality of second sub-light sources are configured to be illuminated at different times so that the second image acquisition structure obtains image information of the right ear.
[0014] Optionally, the first image acquisition structure and the second image acquisition structure have the same structure, and the first image acquisition structure comprises a monocular camera or a binocular camera.
[0015] Optionally, the first light source has a ring structure, and the plurality of sub-light sources are equally divided from the ring structure.
[0016] Optionally, a coordinate system is established, a position of the first image acquisition structure is set as point A, a position of the second image acquisition structure is set as point B, a position of a rotation shaft of the left and right head swinging of the user is set as point O, the points A, B and O form an isosceles triangle, ∠OAB and ∠OBA are the same, and a value of ∠OAB is set as θ.
[0017] A position of the left ear of the user is set as point C, a position of the right ear of the user is set as point D, a distance OC from the point O to the left ear of the user and a distance OD from the point O to the right ear of the user are the same, and a length of the OC is set as r.
[0018] The first image acquisition structure comprises a monocular camera, and the first processing module comprises a first head width measurement unit, the first head width measurement unit is configured to obtain a value of ∠COD by using a theorem of triangles, and obtain head width information of the user according to the value of ∠COD and the value of r, wherein the head width information comprises a distance between the position C of the left ear of the user and the position D of the right ear of the user.
[0019] Optionally, the first head width measurement unit comprises:
[0020] A first sub-processing unit, the first processing unit is configured to obtain the values of r and θ according to the following formula:
[0021] Wherein, the value of ∠CAO is θ1, the value of ∠DBO is θ2, and the ratio n1 of the distance p1 between the first image acquisition structure and the left ear of the user and the distance p2 between the second image acquisition structure and the right ear of the user; the value of ∠CAO is θ3, the value of ∠DBO is θ4, and the ratio n2 of the distance p3 between the second image acquisition structure and the left ear of the user and the distance p4 between the second image acquisition structure and the right ear of the user;
[0022] The second sub-processing unit is configured to obtain the value of ∠COD according to the r value and the θ value obtained by the first processing unit.
[0023] Optionally, the first processing module comprises a human ear model reconstruction unit, and the human ear model reconstruction unit comprises left ear model reconstruction units and right ear model reconstruction units which are identical in structure.
[0024] The first processing unit is configured to process the image collected by the first image acquisition structure to obtain normal vector information of the left ear.
[0025] The second processing unit is configured to take the p1 value obtained by the first head width measurement unit as a starting point, integrate the left ear model in combination with the normal vector information, and reconstruct a three-dimensional model of the left ear.
[0026] Optionally, the first image acquisition structure is a binocular camera, and the processing module comprises a human ear model reconstruction unit, and the human ear model reconstruction unit comprises left ear model reconstruction units and right ear model reconstruction units which are identical in structure.
[0027] The third processing unit is configured to process the image collected by the first image acquisition structure and reconstruct a three-dimensional model of the left ear.
[0028] Optionally, the left ear model reconstruction unit further comprises a fourth processing unit, and the fourth processing unit is configured to integrate the depth map of the three-dimensional reconstruction of the left ear obtained by the third processing unit to supplement the depth information of the occluded and textureless areas.
[0029] Optionally, the processing module further comprises a second head width measurement unit, and the second head width measurement unit comprises:
[0030] The fifth processing unit is configured to convert the depth information of the left ear into point cloud data based on the depth map of the three-dimensional reconstruction of the left ear obtained by the third processing unit.
[0031] The sixth processing unit is configured to obtain the head width information of the user according to the point cloud data.
[0032] The embodiment of the present disclosure further provides an audio effect processing method applied to the audio effect processing device, comprising the following steps.
[0033] Collecting image information of multi-angle images of the left ear and the right ear of the user under multi-angle light sources;
[0034] Obtaining human body parameters of the user according to the image information to determine an audio input signal corresponding to the user, wherein the parameters include head width and a three-dimensional model of human ears;
[0035] Controlling a loudspeaker to output corresponding audio effects according to the audio input signal.
[0036] Optionally, the first image collection structure and the second image collection structure are of the same structure, and the first image collection structure comprises a monocular camera.
[0037] The position of the first image collection structure is set as point A, the position of the second image collection structure is set as point B, and the position of the rotation shaft of the left and right head swinging of the user is set as point O, the points A, B and O form an isosceles triangle, ∠OAB and ∠OBA are the same, and the value of ∠OAB is set as θ.
[0038] The position of the left ear of the user is set as point C, the position of the right ear of the user is set as point D, the distance OC from point O to the left ear of the user and the distance OD from point O to the right ear of the user are the same, and the length of OC is set as r.
[0039] Obtaining human body parameters of the user according to the image information, wherein the parameters include head width and a three-dimensional model of human ears, and specifically include:
[0040] Obtaining the value of ∠COD by using the triangle geometric theorem, and obtaining the head width information of the user according to the value of ∠COD and the value of r, wherein the head width information includes the distance between the position C of the left ear of the user and the position D of the right ear of the user.
[0041] Optionally, the obtaining of the value of ∠COD by using the triangle geometric theorem and the obtaining of the head width information of the user according to the value of ∠COD and the value of r specifically include:
[0042] The values of r and θ are obtained according to the following formula:
[0043] wherein the value of ∠CAO is θ1, the value of ∠DBO is θ2, the ratio n1 of the distance p1 between the first image collection structure and the left ear of the user to the distance p2 between the second image collection structure and the right ear of the user; the value of ∠CAO is θ3, the value of ∠DBO is θ4, the ratio n2 of the distance p3 between the second image collection structure and the left ear of the user to the distance p4 between the second image collection structure and the right ear of the user.
[0044] The value of ∠COD is obtained according to the r value and the θ value obtained by the first processing unit.
[0045] Optionally, a human body parameter of the user is obtained according to the image information, and the parameter includes a head width and a three-dimensional model of the human ear, and specifically includes:
[0046] The image of the human ear is processed, and normal vector information of the left ear and the right ear is obtained.
[0047] The p1 value obtained is taken as a starting point, and the left ear model and the right ear model are integrated respectively in combination with the corresponding normal vector information, so as to reconstruct the three-dimensional models of the left ear and the right ear.
[0048] Optionally, a human body parameter of the user is obtained according to the image information, and the parameter includes a head width and a three-dimensional model of the human ear, and specifically includes:
[0049] The image of the human ear is processed, and the left ear and the right ear are reconstructed in three dimensions respectively.
[0050] Optionally, the method further includes: integrating the depth map of the three-dimensional reconstruction of the left ear and the right ear respectively, so as to supplement the depth information of the occluded and non-texture areas.
[0051] Optionally, the method further includes:
[0052] The depth information of the left ear and the right ear is converted into point cloud data respectively based on the depth map of the three-dimensional reconstruction of the left ear and the right ear.
[0053] The head width information of the user is obtained according to the point cloud data.
[0054] The embodiment of the disclosure further provides an extended reality device, which includes the sound effect processing device.
[0055] The embodiment of the disclosure further provides a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the sound effect processing method when executing the computer program.
[0056] The embodiment of the disclosure further provides a computer readable storage medium, which stores a computer program, and the computer program implements the steps of the sound effect processing method when executed by a processor.
[0057] The beneficial effects of the present disclosure are that the head width and the three-dimensional model of the human ear play an important role in the calculation of personalized HRTF. The embodiments of the present disclosure provide an audio processing device and an audio processing method, which realize the collection of information required for personalized HRTF, have a simple structure, improve the accuracy of HRTF measurement, and control the corresponding audio effect according to the obtained personalized HTTF function Control the speaker output to improve the audio effect. BRIEF DESCRIPTION OF DRAWINGS
[0058] Figure 1 shows a schematic diagram of the use state of the audio processing device in the embodiments of the present disclosure;
[0059] Figure 2 shows a schematic diagram of the use state of the audio processing device in the embodiments of the present disclosure;
[0060] Figure 3 shows a schematic diagram of head width measurement in the posture of the human head turning right in the embodiments of the present disclosure;
[0061] Figure 4 shows a schematic diagram of head width measurement in the posture of the human head turning left in the embodiments of the present disclosure. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical scheme and advantages of the embodiments of the present disclosure clearer, the technical scheme of the embodiments of the present disclosure will be described clearly and completely below in conjunction with the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present disclosure.
[0063] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure should be understood as the usual meaning understood by those skilled in the art to which the present disclosure belongs. The terms "first", "second" and the like used in the present disclosure do not represent any order, number or importance, but are only used to distinguish different components. Similarly, "one", "an" or "the" and the like do not represent a quantity limitation, but represent the existence of at least one. "Include" or "contain" and the like mean that the elements or objects before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connected" or "connected" and the like are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to represent relative positional relationships, and when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0064] The head width and the three-dimensional model of the human ear play an important role in the calculation of personalized HRTF. The head width, three-dimensional model of the human ear and other parameters corresponding to the user are obtained, and the HRTF function corresponding to the user is obtained, and then the corresponding sound effect is output by the loudspeaker according to the HRTF function, so as to improve the personal sound experience. Referring to FIGS. 1 and 2, the embodiment provides a sound effect processing device, which comprises:
[0065] a neck-mounted body 10;
[0066] an image acquisition structure, comprising a first image acquisition structure 3 and a second image acquisition structure 4 arranged at two ends of the neck-mounted body 10, and the first image acquisition structure 3 and the second image acquisition structure 4 are respectively configured to acquire corresponding human ear images;
[0067] a light source, comprising a first light source 5 arranged around the four sides of the first image acquisition structure 3 and a second light source 6 arranged around the four sides of the second image acquisition structure 4;
[0068] the first light source 5 is controlled to be turned on to make the first image acquisition structure 3 acquire a left ear image, and the second light source 6 is controlled to be turned on to make the second image acquisition structure 4 acquire a right ear image;
[0069] a first processing module, configured to acquire human body parameters of a user according to images acquired by the image acquisition structure, so as to determine an audio input signal corresponding to the user, wherein the parameters comprise a head width and a three-dimensional model of the human ear;
[0070] a second processing module, configured to control a loudspeaker to output a corresponding sound effect according to the audio input signal.
[0071] The neck-mounted body 10 is adopted, so that the sound effect processing device can be placed on the neck, and the image acquisition structure and the light source are arranged on the neck-mounted body 10, so that the structure is simple and the measurement is convenient. The first image acquisition structure 3 and the second image acquisition structure 4 are arranged at two ends of the neck-mounted body 10 respectively, so as to facilitate the acquisition of images of the left ear 1 and the right ear 2 of the user respectively.
[0072] In the embodiment, the first light source 5 comprises a plurality of first sub-light sources capable of emitting light independently, and the plurality of first sub-light sources are configured to be turned on at different times, so that the first image acquisition structure 3 obtains image information of the left ear; the second light source 6 comprises a plurality of second sub-light sources capable of emitting light independently, and the plurality of second sub-light sources are configured to be turned on at different times, so that the second image acquisition structure 4 obtains image information of the right ear.
[0073] In this embodiment, the first image acquisition structure 3 is surrounded by the first light source 5, and the second image acquisition structure 4 is surrounded by the second light source 6. The first light source 5 is inclined to emit light towards the first image acquisition structure 3, and the second light source 6 is inclined to emit light towards the second image acquisition structure 4. The first light source 5 and the second light source 6 each include a plurality of independently emitting sub-sources. The plurality of first sub-sources of the first light source 5 are time-divisionally lit to provide multi-angle light sources, so that the first image acquisition structure 3 can obtain multi-view images of the left ear 1. The plurality of second sub-sources of the second light source 6 are time-divisionally lit to provide multi-angle light sources, so that the second image acquisition structure 4 can obtain multi-view images of the left ear 1, and then obtain a detailed ear gradient map, so as to facilitate the reconstruction of a three-dimensional ear model.
[0074] The first processing module processes the images obtained by the image acquisition structure to obtain head width information and reconstruct a human ear model, and realizes the collection of information required for personalized HRTF. The second processing module controls the loudspeaker to output corresponding sound effects according to the personalized HRTF function, so as to improve the personal sound effect experience.
[0075] It should be noted that the audio input signal includes an HRTF function.
[0076] It should be noted that, in this embodiment, the first light source 5 corresponding to the first image acquisition structure 3 includes a plurality of first sub-sources, and the plurality of first sub-sources are time-divisionally lit, i.e., different first sub-sources are controlled to be lit in a predetermined order, so that left ear images of different angles and different views under different angle light sources can be obtained. Similarly, the second light source 6 corresponding to the second image acquisition structure 4 includes a plurality of second sub-sources, and the plurality of second sub-sources are time-divisionally lit, i.e., different second sub-sources are controlled to be lit in a predetermined order, so that right ear images of different angles and different views under different angle light sources can be obtained.
[0077] The first processing module adopts photometric stereo algorithm when processing the images of the image acquisition structure. In the photometric stereo algorithm, the images captured under a plurality of light sources need to be preprocessed first, including image calibration, alignment and denoising operations. Then, by comparing the brightness and color information of the images under different light sources, the depth information of different objects in the scene can be inferred, and more accurate depth estimation than traditional stereo vision algorithm can be provided.
[0078] In an exemplary embodiment, the first image acquisition structure 3 and the second image acquisition structure 4 have the same structure, and the first image acquisition structure 3 includes a monocular camera or a binocular camera.
[0079] The monocular camera and the plurality of sub-light sources around the monocular camera form a photometric stereo vision system. By time-sharing lighting and driving the camera to take pictures, a detailed human ear gradient map can be obtained. The two binocular cameras on one side and the plurality of light sources form a space-time stereo matching system. The binocular cameras are synchronized, and after each light source is turned on, the corresponding camera takes pictures. Then, one frame of binocular image is obtained by the camera multiple times and the light source synchronization exposure. The plurality of images form a description of the sub-stereo matching, which is more robust and has better details.
[0080] In an exemplary embodiment, the first light source 5 and the second light source 6 have the same structure, the first light source 5 has a ring structure, and the plurality of sub-light sources are equally divided into the ring structure.
[0081] The following describes a method for measuring head width by using two postures of the human head 100 through the sound effect processing device.
[0082] FIG. 1 is a schematic diagram of head width measurement in the posture of the user turning his head to the right, and FIG. 2 is a schematic diagram of head width measurement in the posture of the user turning his head to the left.
[0083] In an exemplary embodiment, a coordinate system is established, the position of the first image acquisition structure 3 is set as point A, the position of the second image acquisition structure 4 is set as point B, and the position of the rotation axis of the user turning his head left and right is set as point O (it should be noted that FIG. 1 and FIG. 2 are front views, and at this time, the rotation axis of turning the head left and right is represented as a point due to the view relationship), the points A, B and O form an isosceles triangle, O1 is the midpoint of the line connecting the optical centers of the first image acquisition structure 3 and the second image acquisition structure 4, and the AOB triangle does not change when the human head 100 rotates. ∠OAB and ∠OBA are the same, and the value of ∠OAB is set as θ;
[0084] The position of the left ear 1 of the user is set as point C, and the position of the right ear 2 of the user is set as point D. When the user turns his head left and right, the distance from the rotation axis to the left ear 1 and the right ear 2 is fixed, and no matter how the human head 100 turns, the left ear 1 and the right ear 2 rotate on a circle with the rotation axis as the center and the radius as r. That is, the distance OC from the O point to the left ear 1 of the user and the distance OD from the O point to the right ear 2 of the user are the same, and the length of OC is set as r.
[0085] The first image acquisition structure 3 includes a monocular camera, and the first processing module includes a first head width measurement unit.
[0086] The value of ∠COD is obtained by using the theorem of triangles, and the head width information of the user is obtained according to the value of ∠COD and the value of r, wherein the head width information includes the distance between the position C of the left ear of the user and the position D of the right ear of the user.
[0087] The first head width measurement unit includes:
[0088] a first processing unit configured to obtain the r value and the θ value according to the following formula:
[0089] wherein the value of ∠CAO is θ1, the value of ∠DBO is θ2, and the ratio n1 of the distance p1 between the first image acquisition structure and the left ear of the user and the distance p2 between the second image acquisition structure and the right ear of the user; the value of ∠CAO is θ3, the value of ∠DBO is θ4, and the ratio n2 of the distance p3 between the second image acquisition structure and the left ear of the user and the distance p4 between the second image acquisition structure and the right ear of the user;
[0090] a second processing unit configured to obtain the value of ∠COD according to the r value and the θ value obtained by the first processing unit.
[0091] Further specifically, the first head width measurement unit comprises:
[0092] a first processing unit configured to obtain the first sub-parameters of the first posture when turning the head to the right, the first sub-parameters comprising the r value, the value θ1 of ∠CAO, the value θ2 of ∠DBO, and the ratio n1 of the distance p1 between the first image acquisition structure 3 and the left ear 1 of the user and the distance p2 between the second image acquisition structure 4 and the right ear 2 of the user (see FIG. 1);
[0093] a second processing unit configured to obtain the second sub-parameters of the second posture when turning the head to the left, the second sub-parameters comprising the value θ3 of ∠CAO, the value θ4 of ∠DBO, and the ratio n2 of the distance p3 between the second image acquisition structure 4 and the left ear 1 of the user and the distance p4 between the second image acquisition structure 4 and the right ear 2 of the user (see FIG. 2);
[0094] a third processing unit configured to obtain the r value and the θ value according to the following formula:
[0095] a fourth processing unit configured to obtain the distance p1 between the first image acquisition structure 3 and the left ear 1 of the user and the distance p2 between the second image acquisition structure 4 and the right ear 2 of the user according to the r value and the θ value obtained by the third processing unit;
[0096] a fifth processing unit configured to obtain the value of ∠COD according to the r value and the θ value obtained by the third processing unit, and the p1 and p2 obtained by the fourth processing unit;
[0097] The sixth sub-processing unit is configured to obtain head width information of the user according to the value of ∠COD and the value of r, wherein the head width information includes the distance between the position C of the left ear 1 of the user and the position D of the right ear 2 of the user.
[0098] Referring to FIG. 1, according to the cosine theorem of geometry, the above formulas (1) and (2) can be obtained, and the above formulas (1) and (2) can be simplified as r=f(θ1,θ2,θ,n1,t) (5).
[0099] Referring to FIG. 2, according to the cosine theorem of geometry, the above formulas (3) and (4) can be obtained, and the above formulas (3) and (4) can be simplified as r=f(θ3,θ4,θ,n2,t) (6)
[0100] Wherein θ1, θ2, θ3, θ4, n1, n2 are known quantities; by simultaneously solving the above formulas (5) and (6), r and θ can be obtained.
[0101] After r and θ are obtained, the distance p1 between the first image acquisition structure 3 and the left ear 1 of the user and the distance p2 between the second image acquisition structure 4 and the right ear 2 of the user can be calculated, and according to the geometric formula, the values of ∠AOC, ∠DOB and ∠AOB can be calculated, and then the value of ∠COD is obtained, and the head width CD can be obtained according to the cosine theorem.
[0102] It should be noted that when measuring the head width by using the above scheme, two postures of the human head 100 can be used to achieve the measurement, or multiple postures can be used to establish multiple equations, and the least square method can be used to calculate the head width.
[0103] In an exemplary embodiment, the processing module includes a human ear model reconstruction unit, and the human ear model reconstruction unit includes a left ear 1 model reconstruction unit and a right ear 2 model reconstruction unit which are the same in structure.
[0104] The first processing unit is configured to process different angle images acquired by the first image acquisition structure 3 under illumination of a multi-angle light source, and obtain normal vector information of the left ear 1 according to photometric stereo algorithm.
[0105] The second processing unit is configured to take the p1 value obtained by the head width measurement structure as a starting point, integrate the left ear model in combination with the normal vector information, and perform three-dimensional model reconstruction on the left ear 1.
[0106] The left ear model reconstruction unit and the right ear model reconstruction unit are the same in structure, and the method for three-dimensional model reconstruction of the left ear and the method for three-dimensional model reconstruction of the right ear are also the same. The three-dimensional model reconstruction of the human ear is described below by taking the left ear 1 as an example.
[0107] In this embodiment, the camera and the surrounding multiple sub-light source groups form a photometric stereo vision system. The principle of photometric stereo vision is to use multiple light sources to illuminate the target object at different angles and time points, and to obtain depth and surface normal vector information by analyzing the brightness changes of the target object under different lighting conditions. By driving the camera to capture images under these lighting conditions, rich gradient information can be obtained, which helps to more accurately represent the details of the human ear surface.
[0108] Based on the analysis of the images obtained by the photometric stereo vision system, the head width accuracy is improved. The p1 value obtained is used as the starting point, and the corresponding normal vector information is combined to generate a three-dimensional model of the human ear by the method of integral reconstruction of the human ear model. The accuracy is higher. Integral reconstruction is a commonly used three-dimensional reconstruction technique that uses depth and normal vector information to integrate and reconstruct the surface of an object.
[0109] It should be noted that after obtaining the preliminary three-dimensional model, optimization and refinement can be performed. This includes removing noise, filling in missing parts, smoothing the surface, and other operations to obtain a more accurate and realistic three-dimensional model of the human ear.
[0110] In an exemplary embodiment, the first image acquisition structure 3 is a binocular camera, and the first processing module includes a human ear model reconstruction unit, which includes a left ear model reconstruction unit and a right ear model reconstruction unit with the same structure. The left ear model reconstruction unit includes:
[0111] The third processing unit is configured to process the images captured by the first image acquisition structure 3 and perform three-dimensional reconstruction of the left ear 1 through a stereo matching algorithm.
[0112] The stereo matching algorithm is used to match the images captured by the binocular camera, find the corresponding pixel points, and calculate the disparity between them. Common stereo matching algorithms include region-based methods, feature-based methods, and deep learning methods.
[0113] According to the calculated disparity information, the depth map of the human ear surface can be derived. The depth map reflects the distance from the pixel points at different positions to the camera, and can be used to represent the three-dimensional shape of the human ear.
[0114] Based on the depth map, three-dimensional reconstruction operations can be performed to map the shape of the human ear from two-dimensional image space to three-dimensional space. This usually involves point cloud reconstruction, surface reconstruction, and other techniques to ultimately generate a three-dimensional model of the human ear.
[0115] In this embodiment, two binocular cameras on one side and multiple light sources are combined to achieve a stronger stereo matching effect through synchronous operation. This system utilizes the synchronization between the binocular cameras and the synergy of multiple light sources. Since each frame of image contains information under different lighting conditions, richer image information can be obtained, thereby improving the robustness of matching, especially when dealing with complex scenes or objects with rich texture details, better matching results can be obtained.
[0116] After each corresponding sub-light source is lit, the corresponding binocular camera is used to capture images, thereby obtaining a frame of binocular images. Through multiple synchronous exposures of sub-light sources and cameras, multiple frames of images can be obtained, which have information under different lighting conditions, which helps to improve the quality and accuracy of stereo matching.
[0117] In an exemplary embodiment, the left ear 1 model reconstruction unit further includes a fourth processing unit configured to perform integral processing on the depth map of the left ear 1 three-dimensional reconstruction obtained by the third processing unit to supplement the depth information of occluded and textureless areas.
[0118] The multi-angle light source provides information under different lighting conditions, and the camera captures this information. Through photometric stereo technology, a detailed gradient map can be obtained. By performing integral processing on the depth map, the depth information of occluded and textureless areas can be effectively supplemented, and the overall three-dimensional reconstruction quality can be improved.
[0119] In an exemplary embodiment, the first processing module further includes a second head width measurement unit, which includes:
[0120] The fifth processing unit is configured to convert the depth information of the left ear 1 into point cloud data based on the depth map of the left ear 1 three-dimensional reconstruction obtained by the third processing unit;
[0121] The sixth processing unit is configured to obtain the head width information of the user according to the point cloud data.
[0122] The embodiments of the present disclosure also provide an audio effect processing method applied to the above-mentioned audio effect processing device, which includes the following steps:
[0123] Collecting image information of multi-angle light source irradiation under multi-view images of the left ear 1 and the right ear 2 of the user;
[0124] Obtaining human body parameters of the user according to the image information to determine the audio input signal corresponding to the user, the parameters including head width and three-dimensional model of human ear;
[0125] Controlling the loudspeaker to output the corresponding audio effect according to the audio input signal. For example, the audio input signal includes an HRTF function.
[0126] In an exemplary embodiment, the first image acquisition structure 3 and the second image acquisition structure 4 are of the same structure, and the first image acquisition structure 3 comprises a monocular camera.
[0127] A coordinate system is established, and the position of the first image acquisition structure 3 is set as point A, the position of the second image acquisition structure 4 is set as point B, and the position of the rotation axis of the user's left and right head swinging is set as point O. The points A, B and O form an isosceles triangle, and ∠OAB and ∠OBA are the same, and the value of ∠OAB is set as θ.
[0128] The position of the user's left ear 1 is set as point C, and the position of the user's right ear 2 is set as point D. The distance OC from point O to the user's left ear 1 and the distance OD from point O to the user's right ear 2 are the same, and the length of OC is set as r.
[0129] According to the image information, the human body parameters of the user are obtained, which include head width and a human ear three-dimensional model, and specifically include:
[0130] The first sub-parameters of the first posture when swinging the head to the right are obtained, which include the value r, the value θ1 of ∠CAO, the value θ2 of ∠DBO, and the ratio n1 of the distance p1 between the first image acquisition structure 3 and the user's left ear 1 to the distance p2 between the second image acquisition structure 4 and the user's right ear 2.
[0131] The second sub-parameters of the second posture when swinging the head to the left are obtained, which include the value θ3 of ∠CAO, the value θ4 of ∠DBO, and the ratio n2 of the distance p3 between the second image acquisition structure 4 and the user's left ear 1 to the distance p4 between the second image acquisition structure 4 and the user's right ear 2.
[0132] The values of r and θ are obtained according to the following formula:
[0133] The distance p1 between the first image acquisition structure 3 and the user's left ear 1 and the distance p2 between the second image acquisition structure 4 and the user's right ear 2 are obtained according to the values of r and θ obtained by the third processing unit.
[0134] The value of ∠COD is obtained according to the values of r and θ obtained by the third processing unit and the values of p1 and p2 obtained by the fourth processing unit.
[0135] The head width information of the user is obtained according to the value of ∠COD and the value of r, wherein the head width information includes the distance between the position C of the user's left ear 1 and the position D of the user's right ear 2.
[0136] In an exemplary embodiment, according to the image information, the human body parameters of the user are obtained, which include head width and a human ear three-dimensional model, and specifically include:
[0137] The image of the human ear is processed, and normal vector information of the left ear 1 and the right ear 2 is obtained respectively according to a photometric stereo algorithm;
[0138] Taking the p1 value obtained by the first head width measurement unit as a starting point, the left ear 1 and the right ear 2 models are integrated respectively according to the corresponding normal vector information to reconstruct the three-dimensional models of the left ear 1 and the right ear 2.
[0139] In an exemplary embodiment, human body parameters of the user are obtained according to the image information, and the parameters include head width and a three-dimensional model of the human ear, and specifically include:
[0140] The image of the human ear is processed, and the left ear 1 and the right ear 2 are three-dimensionally reconstructed respectively through a stereo matching algorithm.
[0141] In an exemplary embodiment, the sound effect processing method further includes: performing integral processing on the depth map of the three-dimensional reconstruction of the left ear 1 and the right ear 2 respectively to supplement the depth information of the occluded and textureless areas.
[0142] It should be noted that in the stereo matching algorithm, left-right consistency check is a commonly used technique for verifying the accuracy of the disparity map obtained through the matching algorithm. In stereo vision, left-right consistency refers to the fact that the same object in the left and right images should have the same disparity value because they are the same object in the real world. The basic idea of left-right consistency check is that for each pixel point, its corresponding matching point is found in the left image, and then the reverse matching point of the matching point is found in the right image. If the disparity value between the two reverse matching points is close to the disparity value of the initial matching point, then the matching is considered reliable. If the matching on the left and right sides is inconsistent, it may be due to mis-matching or other errors, and further processing or screening is needed. Through left-right consistency check, the accuracy of the stereo matching algorithm can be improved, and the situation of mis-matching can be reduced, so that a more reliable disparity map can be obtained. Left-right consistency check is usually an important step in stereo matching algorithm, which helps to improve the performance and stability of the stereo vision system.
[0143] The depth in the embodiment is obtained based on a photometric stereo vision system, that is, the first image acquisition structure and the second image acquisition structure are both binocular cameras, the first light source corresponding to the first image acquisition structure includes a plurality of first sub-light sources arranged around the first image acquisition structure, the second light source corresponding to the second image acquisition structure includes a plurality of second sub-light sources arranged around the second image acquisition structure, the parallax images obtained by the first image acquisition structure and the second image acquisition structure are obtained under the corresponding plurality of sub-light sources, the light-emitting angles of the plurality of first sub-light sources are different, and the light-emitting angles of the plurality of second sub-light sources are also different. In this way, in the step of performing left-right consistency checking, the false matching is reduced, the matching accuracy is improved, and the accuracy of the depth map obtained is also higher. On this basis, the depth loss in the occluded or textureless area of the depth map is supplemented by integration, so that a more complete depth map is obtained.
[0144] In an example embodiment, the sound effect processing method further comprises:
[0145] Converting the depth information of the left ear 1 and the right ear 2 into point cloud data based on the depth map of the three-dimensional reconstruction of the left ear 1 and the right ear 2;
[0146] Obtaining the head width information of the user according to the point cloud data.
[0147] The embodiment of the present disclosure also provides an extended reality device, comprising the sound effect processing device.
[0148] The embodiment of the present disclosure also provides a computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the sound effect processing method.
[0149] The embodiment of the present disclosure also provides a computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the steps of the sound effect processing method.
[0150] The processor can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), a PLA (Programmable Logic Array). The processor can also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also referred to as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor can be integrated with a GPU (Graphics Processing Unit) that is responsible for rendering and drawing the content required to be displayed on the display screen. In some embodiments, the processor can further include an I (Artificial Intelligence) processor for processing computing operations related to machine learning. The memory can include one or more computer-readable storage media, which can be non-transitory. The memory can also include a high-speed random access memory and a non-volatile memory such as one or more disk storage devices, flash memory devices. In this embodiment, the memory is used at least to store the following computer programs, wherein the computer programs are loaded and executed by the processor, and the computer programs are executed by the processor to implement the steps of the portable sound effect processing method according to any one of the above embodiments. In addition, the resources stored in the memory can also include an operating system and data, and the storage mode can be temporary storage or permanent storage. The operating system can include Windows, Unix, Linux, and the like. The data can include, but is not limited to, data corresponding to test results, and the like.
[0151] The following points need to be explained:
[0152] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can be referred to the general design.
[0153] (2) For the sake of clarity, the thickness of the layers or regions is exaggerated or reduced in the drawings used to describe the embodiments of the present disclosure, that is, the drawings are not drawn according to the actual proportion. It can be understood that when an element such as a layer, a film, a region or a substrate is referred to as being located "on" or "under" another element, the element can be "directly" located on or under another element or there can be an intermediate element.
[0154] (3) In the case of no conflict, the embodiments of the disclosure and the features in the embodiments can be combined to obtain new embodiments.
[0155] It can be understood that the above implementation is only an exemplary implementation adopted for illustrating the principles of the disclosure, however the disclosure is not limited thereto. Various modifications and improvements can be made by those of ordinary skill in the art without departing from the spirit and essence of the disclosure, and these modifications and improvements are also considered as the protection scope of the disclosure.
Claims
1. An audio processing apparatus, characterized by comprising: The application relates to a neck-hanging type body, image acquisition structures, light sources, a first processing module and a second processing module. The image acquisition structures include first and second image acquisition structures arranged at two ends of the neck-hanging type body, and the first and second image acquisition structures are configured to acquire corresponding ear images. The light sources include first and second light sources arranged around the first and second image acquisition structures. The first light source is controlled to be lighted to make the first image acquisition structure acquire a left ear image, and the second light source is controlled to be lighted to make the second image acquisition structure acquire a right ear image. The first processing module is used for acquiring human body parameters of a user according to images acquired by the image acquisition structures, determining an audio input signal of the user, and the parameters include head width and a three-dimensional model of human ears. The second processing module is used for controlling a loudspeaker to output corresponding sound effects according to the audio input signal. The first light source includes a plurality of first sub-light sources capable of independently emitting light, and the plurality of first sub-light sources are configured to be lighted at different times to make the first image acquisition structure acquire image information of a left ear.
2. The sound effect processing apparatus according to claim 1, wherein The first and second image acquisition structures have the same structure, and the first image acquisition structure includes a monocular camera or a binocular camera.
3. The sound effect processing apparatus according to claim 1, wherein The first light source is in a ring structure, and a plurality of sub-light sources are formed by equisection of the ring structure.
4. The sound effect processing apparatus according to claim 1, wherein A coordinate system is established, and positions of the first and second image acquisition structures are set as points A and B, respectively.
5. The sound effect processing apparatus according to claim 1, wherein The first processing module includes a first head width measurement unit, and the first head width measurement unit is used for acquiring a value of the angle COD by using a triangle geometric theorem and acquiring head width information of the user according to the value of the angle COD and the value of r, wherein the head width information includes a distance between a left ear position C of the user and a right ear position D of the user. The first head width measurement unit includes: The first head width measurement unit includes:
6. The sound effect processing apparatus according to claim 4, wherein A second sub-processing unit is used for acquiring the value of the angle COD according to the value of r and the value of theta acquired by the first processing unit. a first processing unit, the first processing unit configured to obtain r and θ values according to the following equations: 7. The sound effect processing apparatus according to claim 6, wherein The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises:
8. The sound effect processing apparatus according to claim 1, wherein The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises:
9. The sound effect processing apparatus according to claim 8, wherein The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises:
10. The sound effect processing apparatus according to claim 8, wherein The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises:
11. A sound effect processing method applied to the sound effect processing device of any one of claims 1-10, characterized in that, The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises:
12. The sound effect processing method of claim 11, wherein, The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises:
13. The sound effect processing method of claim 12, wherein, The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a right ear model reconstruction unit which are structurally identical, and the left ear model reconstruction unit comprises: The first processing module comprises a human ear model reconstruction unit, the human ear model reconstruction unit comprises a left ear model reconstruction unit and a The r and theta values are obtained according to the following equations: Wherein, the value of ∠CAO is θ1, the value of ∠DBO is θ2, the ratio n1 of the distance p1 between the first image acquisition structure and the left ear of the user and the distance p2 between the second image acquisition structure and the right ear of the user; the value of ∠CAO is θ3, the value of ∠DBO is θ4, the ratio n2 of the distance p3 between the second image acquisition structure and the left ear of the user and the distance p4 between the second image acquisition structure and the right ear of the user; The value of ∠COD is obtained according to the r value and the θ value obtained by the first processing unit.
14. The sound effect processing method of claim 12, wherein, The human body parameters of the user are obtained according to the image information, and the parameters include head width and three-dimensional model of human ears, specifically including: The image of the human ear is processed, and the normal vector information of the left ear and the right ear is obtained; The p1 value obtained is taken as the starting point, and the left ear and the right ear models are integrated respectively in combination with the corresponding normal vector information to reconstruct the three-dimensional models of the left ear and the right ear.
15. The sound effect processing method of claim 11, wherein, The human body parameters of the user are obtained according to the image information, and the parameters include head width and three-dimensional model of human ears, specifically including: The image of the human ear is processed, and the left ear and the right ear are three-dimensionally reconstructed respectively.
16. The sound effect processing method of claim 15, wherein, Further comprising: The depth information of the occluded and non-texture areas is supplemented by integrating the depth maps of the three-dimensional reconstruction of the left ear and the right ear respectively.
17. The sound effect processing method of claim 15, wherein, Further comprising: The depth information of the left ear and the right ear is converted into point cloud data respectively based on the depth maps of the three-dimensional reconstruction of the left ear and the right ear. The head width information of the user is obtained according to the point cloud data.
18. An extended reality device, comprising: The sound effect processing device of any one of claims 1-10.
19. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to realize the steps of the sound effect processing method of any one of claims 11-17.
20. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the steps of the sound effect processing method of any one of claims 11-17.
Citation Information
Patent Citations
Method of improving localization of surround sound
CN112005559A
Personalized hrtfs via optical capture
CN112470497A
Audio output device and audio output system using same
CN114175142A
Personalized equalization of audio output using 3D reconstruction of user's ears
CN114270879A
System and method for generating head-related transfer function
JP2020201479A