Sound effect adjustment method and operation device for sound effect adjustment
By using a processor to determine the direction of the sound source and changes in head posture, and by using a direction recognition model and an equalizer to adjust the sound effects, the problem of unsuitable sound effects when the head is not facing the center of the screen is solved, thus improving the auditory experience.
Patent Information
- Application Number
- CN202410556316.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-07
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technology no longer applies spatial audio adjustment when the head is not facing the center of the screen, resulting in a poor auditory experience.
The processor determines the direction of the sound source and changes in head posture. It then uses a direction recognition model and an equalizer to adjust the frequency response and phase of the sound signal to adapt to changes in orientation caused by head rotation.
It provides appropriate spatial sound effects variations, enhances the user's auditory experience, and ensures that the sound effects adapt to changes in head posture.
Smart Images

Figure CN120916103A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to a processing technique of sound signals, and more particularly, to a sound effect adjustment method and a computing device for sound effect adjustment. BACKGROUND
[0002] Spatial sound effect is to transfer sound signals to a surround sound field composed of multiple virtual speakers, to adjust the response and delay of virtual sound signals from different directions, and to transfer sound signals to a stereo sound field accordingly. It is worth noting that the setting of the above-mentioned spatial sound effect is usually assumed to be the situation that the user wears headphones and the head is oriented to the center of the screen of the computer. However, when the head is not oriented to the center of the screen, the above-mentioned adjustment for spatial sound effect will no longer be applicable. SUMMARY
[0003] The present invention is directed to a sound effect adjustment method and a computing device for sound effect adjustment, which can be applied to sound field adjustment for changes in the posture of the head.
[0004] According to an embodiment of the present invention, a sound effect adjustment method is implemented by a processor. The sound effect adjustment method comprises the following steps: determining a sound source direction corresponding to a sound feature of a sound signal, wherein the sound feature is related to at least one of an amplitude and a phase of the sound signal, and the sound signal is recorded from a sound source located at the sound source direction; determining a change in the posture of a head, wherein the change in the posture comprises a rotation angle of the head rotated from a first orientation to a second orientation, and the head is used to wear a sound playing device; and adjusting the sound feature of the sound signal according to a directional difference between the sound source direction and the second orientation, wherein the directional difference is an included angle between the sound source direction and a corrected second orientation, the corrected second orientation is an orientation after the sound source direction is changed by the change in the posture, and the adjusted sound signal is used to be played by the sound playing device.
[0005] According to an embodiment of the present invention, a computing device for sound effect adjustment comprises a memory and a processor. The memory is used to store program codes. The processor is coupled to the memory. The processor is configured to: determine a sound source direction corresponding to a sound feature of a sound signal, wherein the sound feature is related to at least one of an amplitude and a phase of the sound signal, and the sound signal is recorded from a sound source located at the sound source direction; determine a change in the posture of a head, wherein the change in the posture comprises a rotation angle of the head rotated from a first orientation to a second orientation, and the head is used to wear a sound playing device; and adjust the sound feature of the sound signal according to a directional difference between the sound source direction and the second orientation, wherein the directional difference is an included angle between the sound source direction and a corrected second orientation, the corrected second orientation is an orientation after the sound source direction is changed by the change in the posture, and the adjusted sound signal is used to be played by the sound playing device.
[0006] Based on the above, the sound effect adjustment method and the operation device for sound effect adjustment can take the sound source direction as a reference direction, determine a corrected orientation after head rotation according to the reference direction, and provide sound effect adjustment suitable for the corrected orientation. Thus, appropriate spatial sound effect changes can be provided, and a user can be given an immersive auditory experience. BRIEF DESCRIPTION OF DRAWINGS
[0007] The accompanying drawings are included to provide a further understanding of the present application, and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments of the present application and, together with the description, serve to explain the principles of the present application.
[0008] FIG. 1A is a component block diagram of a system according to an embodiment of the present application;
[0009] FIG. 1B is a schematic diagram illustrating an application context according to an embodiment of the present application;
[0010] FIG. 2 is a flowchart of a sound effect adjustment method according to an embodiment of the present application;
[0011] FIGS. 3A-3C is a schematic diagram illustrating a sound propagation path according to an embodiment of the present application;
[0012] FIG. 4A is a schematic diagram illustrating an environment for sample collection according to an embodiment of the present application;
[0013] FIG. 4B is a schematic diagram illustrating model training and inference according to an embodiment of the present application;
[0014] FIG. 5 is a schematic diagram illustrating a pose according to an embodiment of the present application;
[0015] FIG. 6A is a schematic diagram illustrating a first orientation and a sound source direction according to an embodiment of the present application;
[0016] FIG. 6B is a schematic diagram illustrating a first orientation, a second orientation, and a sound source direction according to an embodiment of the present application;
[0017] FIGS. 7A-7G is a schematic diagram illustrating parameters of an equalizer for multiple orientations according to an embodiment of the present application;
[0018] FIG. 8A and FIG. 8B is a schematic diagram illustrating parameters of an equalizer for two channels according to an embodiment of the present application;
[0019] FIG. 9A is a frequency response plot of two channels at an orientation of zero degrees according to an embodiment of the present application;
[0020] FIG. 9B is a frequency response graph of a different monaural sound according to an embodiment of the present application, in which the direction is zero degrees;
[0021] FIGS. 10A-10F is a frequency response graph of a left sound according to an embodiment of the present application, in which the direction is different angles;
[0022] FIGS. 11A-11F is a frequency response graph of a right sound according to an embodiment of the present application, in which the direction is different angles.
[0023] BRIEF DESCRIPTION OF DRAWINGS
[0024] 10: sound playing device;
[0025] 30: image capturing device;
[0026] 50: arithmetic device;
[0027] 51: memory;
[0028] 52: processor;
[0029] H: head;
[0030] S210-S230: step;
[0031] E: ear;
[0032] P1-P9: propagation path;
[0033] S1, S3: sound source;
[0034] LE: left ear;
[0035] RE: right ear;
[0036] S2: reference sound source;
[0037] RL: dummy;
[0038] SS1: reference sound signal;
[0039] LFR, RFR, 910, 920, 1010, 1020, 1110, 1120: frequency response;
[0040] CR: time delay;
[0041] DIM: direction identification model;
[0042] SD1, SD2: sound source direction;
[0043] α H : yaw angle;
[0044] β H: pitch angle;
[0045] γ H : roll angle;
[0046] D1: first orientation;
[0047] D2: second orientation;
[0048] θ H : rotation angle;
[0049] ED2: corrected second orientation;
[0050] θ D : direction difference. DETAILED DESCRIPTION
[0051] Reference will now be made in detail embodiments of the application, examples of which are illustrated in the accompanying drawings. Wherever possible, the same reference numbers are used in the drawings and the description to refer to the same or like parts.
[0052] FIG. 1A is a component block diagram of a system according to an embodiment of the application. Please refer to FIG. 1A , the system comprises a sound playing device 10, an image capturing device 30 and a computing device 50.
[0053] The sound playing device 10 can be a headphone or a wearable playing device. FIG. 1B is a schematic diagram illustrating an application context according to an embodiment of the application. Please refer to FIG. 1B , the sound playing device 10 is wearable for a head H of a user. A (in-ear or in-canal) speaker unit of the sound playing device 10 can be directed towards both ears on the head H. In an embodiment, the sound playing device 10 is configured to play a sound signal.
[0054] The image capturing device 30 can be a camera, a video camera or a circuit with image acquisition function. Please refer to FIG. 1B , the image capturing device 30 is built-in or externally connected to the image capturing device 30. A lens of the image capturing device 30 can be directed towards the head H. In an embodiment, the image capturing device 30 is configured to capture an image. For example, the image capturing device 30 captures the head and generates a head image (i.e. acquires an image of the head H) accordingly. FIG. 1B
[0055] The computing device 50 can be a smartphone, a tablet computer, a desktop computer, a notebook computer, a smart assistant device, a wearable device, a smart television, or other electronic device. The computing device 50 is communicatively connected to the sound playing device 10 and the image capturing device 30. For example, the computing device 50 is loaded with a USB, UART, or other wired transmission interface (not shown in the figure), or is loaded with a Wi-Fi, Bluetooth, or other wireless communication transceiver (not shown in the figure), and is able to transmit or receive signals. For example, the image capturing device 30 transmits a signal carrying an image to the computing device 50, or the computing device 50 transmits a sound signal to the sound playing device 10.
[0056] The computing device 50 includes, but is not limited to, a memory 51 and a processor 52.
[0057] The memory 51 can be any type of fixed or removable Radom Access Memory (RAM), Read Only Memory (ROM), flash memory, Hard Disk Drive (HDD), Solid-State Drive (SSD), or similar component. In embodiments, the memory 51 is used to store program codes, software modules, configurations, data (e.g., sound signals, head images, or algorithm parameters), or files, and embodiments thereof will be described in detail later.
[0058] The processor 52 is coupled to the memory 51. The processor 52 can be a Central Processing Unit (CPU), a Graphic Processing unit (GPU), or other programmable general-purpose or special-purpose microprocessors (Microprocessor), Digital Signal Processors (DSP), programmable controllers, Field Programmable Gate Arrays (FPGA), Application-Specific Integrated Circuits (ASIC), neural network accelerators, or other similar components or combinations thereof. In embodiments, the processor 52 is used to perform all or part of the operations of the computing device 100, and can load and execute each program code, software module, file, and data stored in the memory 51. In embodiments, the processor 52 can control the image capturing device 30 to take images. In other embodiments, the processor 52 can control the playing functions (e.g., playing, pausing, switching tracks, fast-forwarding, or rewinding) of the sound playing device 10. In some embodiments, the functions of the processor 52 can be implemented by software or chips.
[0059] For application scenarios, FIG. 1B For example, the computing device 50 is a laptop computer, and the head H is facing the laptop computer's display. However, the user's position and / or orientation may vary.
[0060] The methods described in this embodiment of the invention will be explained in conjunction with the components and modules in the sound playback device 10, image capture device 30, and computing device 50. The various processes of this method may be adjusted according to the implementation situation, and are not limited thereto.
[0061] FIG. 2 This is a flowchart of a sound effect adjustment method according to an embodiment of the present invention. Please refer to... FIG. 2 The processor 52 determines the direction of the sound source corresponding to the sound characteristics of the sound signal (step S210). Specifically, the sound signal is the signal that the processing device 50 intends to transmit to the sound playback device 10 and play through the sound playback device 10. The content of the sound signal can be music, speech, lecture, or broadcast, and is not limited thereto.
[0062] Sound features relate to at least one of the amplitude and phase of a sound signal. In an embodiment, sound features include a frequency response. The frequency response is the response of the sound signal in the frequency domain, or it may be the amplitude of the sound signal corresponding to multiple frequencies. The processor 52 may measure the frequency response of the sound signal. For example, its response in the frequency domain may be measured by inputting an impulse response, but it is not limited thereto.
[0063] In an embodiment, the sound features also include signal delay. Signal delay is the time difference of a sound signal between two channels (e.g., left and right channels). For example, the cross-correlation between the sound signals of the two channels is calculated, and the amount of delay (as signal delay) is determined based on the peak value of the cross-correlation function.
[0064] It is worth noting that sound waves may propagate along different paths due to obstruction or interference from objects. For example, FIGS. 3A-3C This is a schematic diagram illustrating the sound propagation paths P1 to P6 according to an embodiment of the present invention. Please refer to... FIG. 3AThe pinna surface of the ear E includes multiple curved surfaces with different curvatures. The propagation paths P1, P2 originate from a far distance in the horizontal direction. However, the propagation path P1 can be reflected by the pinna to the ear canal. Alternatively, the propagation path P3 can directly enter the ear canal. The propagation paths P1, P2 originate from a far distance in the vertical direction. However, the propagation path P3 can directly enter the ear canal. Alternatively, the propagation path P4 can be reflected by the pinna to the ear canal. Sound waves from different directions also have different distribution characteristics in frequency. The frequency response can reflect the above distribution characteristics. That is, sound waves from different directions can correspond to different frequency responses, in which the amplitudes / intensities of responses at some frequencies can be different.
[0065] In addition, please refer to FIG. 3B For the same sound source S1 (for example, a loudspeaker), the sound signals directly reaching the left ear LE and the right ear RE via the propagation paths P5, P6 are different, and the propagation times of the two propagation paths P5, P6 can also be different. That is, the time for the sound signal originating from the sound source S1 to directly reach the left ear LE and the right ear RE can be different. The time difference in propagation / arrival time (i.e., signal delay) can affect the phase of the sound signal.
[0066] On the other hand, please refer to FIG. 3C For the same sound source S1 (for example, a loudspeaker), the sound signals directly or via reflection reaching the left ear LE and the right ear RE via the propagation paths P7, P8, P9 are different, and the propagation times of the three propagation paths P7-P9 can also be different. That is, the time for the sound signal originating from the sound source S1 to directly or via reflection reach the left ear LE and the right ear RE can be different. The time difference in propagation / arrival time (i.e., signal delay) can affect the phase of the sound signal. Sound waves from different directions can also correspond to different signal delays on the two channels.
[0067] In an embodiment, the sound signal is recorded for a sound source located at a position in a sound source direction. That is, the microphone is located at a reference center, and the sound source direction is the direction of the sound source relative to the reference center. The sound source direction can include a horizontal direction and / or a vertical direction. The sound source can be a person, an instrument, an animal, a loudspeaker, a device, wind, or water, and the like. For example, a person sings in front of a microphone, and the microphone records the person's voice to generate a sound signal. The distance between the sound source and the reference center can be 20 cm, 50 cm, or 100 cm, and the like.
[0068] In an embodiment, the processor 52 can analyze the sound characteristics of the sound signal. For example, the frequency response and / or the signal delay of the two channels. Sound signals from different directions have different frequency responses and / or different signal delays. The processor 52 can identify or estimate the sound source direction of the sound signal according to the sound characteristics of the sound signal.
[0069] In an embodiment, the processor 52 can train the direction identification model by a machine learning algorithm, and learn the association between the positions of the reference sound sources located in the plurality of reference directions and the corresponding sound features. The machine learning algorithm is, for example, a multiple layer perception (MLP), a convolutional neural network (CNN), a recurrent neural network (RNN), or a temporal convolutional network (TCN) (e.g., Conv-TasNet), but is not limited thereto. The machine learning algorithm can train the direction identification model to understand the labeled samples (e.g., sound features of the determined reference directions) to establish the association between the sound signals / sound features (i.e., input of the model) and the reference directions (i.e., output of the model). For example, based on the CNN-based model training, a feature map of the labeled samples can be obtained. The direction identification model is a model constructed after learning, and can be used to infer the to-be-evaluated data (e.g., to-be-evaluated sound signals / sound features) to determine the corresponding direction (as the sound source direction) of the to-be-evaluated signals. For example, a linear classifier is used to determine the correct classification (as the sound source direction).
[0070] For example, FIG. 4A is a schematic diagram illustrating an environment for sample collection according to an embodiment of the present application. Please refer to FIG. 4A , it is assumed that a plurality of loudspeakers (as reference sound sources S2) are arranged in the experimental space, and the relative directions of these reference sound sources S2 with respect to the head of the dummy RL (as reference directions) are known. The reference directions can be defined in advance.
[0071] FIG. 4B is a schematic diagram illustrating model training and inference according to an embodiment of the present application. Please refer to FIG. 4A and FIG. 4BThe dummy RL has microphones set at its ears to receive reference sound signals SSI. The reference sound signals SSI are played by reference sound sources S2 respectively. The reference sound signals SSI can be human voice, music, or synthesized sound, without limitation. Sound features are extracted from the reference sound signals SSI received by the two microphones. For example, the two microphones correspond to left and right channels respectively, and the sound features are frequency responses LFR, RFR of the reference sound signals SSI received by the two microphones respectively and / or signal delay CR between the two reference sound signals. The sound features of the reference sound signals SSI and the reference direction corresponding to the reference sound sources S2 are training samples for the direction identification model DIM. The sound features of the reference sound sources S2 corresponding to other reference directions and the reference directions thereof are also other training samples. The training samples are used to train the direction identification model DIM.
[0072] The processor 52 can train the direction identification model DIM or obtain a trained direction identification model DIM from other devices. Then, the processor 52 can input sound features of a sound signal to the direction identification model DIM to determine a sound source direction SD1 corresponding to the sound features by the direction identification model DIM. The output of the direction identification model DIM can be a specific direction (e.g., 30, 45, or 90 degrees, without limitation) (directly as the sound source direction SD1) or probabilities corresponding to multiple reference directions (the probability with the highest value or the arithmetic mean of the probabilities with the highest values can be taken as the sound source direction SD1).
[0073] In another embodiment, the association between the positions of the multiple reference directions and the corresponding sound features can be recorded as a lookup table or converted into an equation. The processor 52 can determine the sound source direction corresponding to the sound features of a sound signal by looking up the lookup table or plugging into the equation.
[0074] Please refer to FIG. 2 The processor 52 determines a change in the pose of the head (step S220). Specifically, the head is used to wear the sound playing device 10. As shown in FIG. 1B , the head H wears ear muff-type headphones (i.e., an example of the sound playing device 10). The rotation of the head will cause a change in the pose. The change in the pose includes a rotation angle of the head from a first orientation to a second orientation. For example, the head is at a first orientation at time point t, and the head is at a second orientation at time point t+1.
[0075] FIG. 5 is a schematic diagram illustrating the pose according to an embodiment of the present application. Please refer to FIG. 5 The rotation angle of the head H includes a yaw angle (Yaw) a H , a pitch angle (Pitch) b H , and a roll angle (Roll) g H .
[0076] FIG. 6A is a schematic diagram illustrating the first orientation D1 and the sound source direction SD2 according to an embodiment of the application. Please refer to FIG. 6A , the head front of the user wearing the sound playing device 10 is facing the first orientation D1. The sound source direction SD2 is the direction of the sound source S3 relative to the recording position (e.g. the center as mentioned above). For example, the left channel corresponds to 30 degrees; the right channel is 180 degrees apart from the left channel, and the right channel corresponds to -30 degrees.
[0077] FIG. 6B is a schematic diagram illustrating the first orientation D1, the second orientation D2 and the sound source direction SD2 according to an embodiment of the application. Please refer to FIG. 6B , the rotation angle θ H is the yaw angle a H is 20 degrees (e.g. the left channel corresponds to 20 degrees; the right channel is 180 degrees apart from the left channel, and the right channel corresponds to -20 degrees). At this time, the head front of the user is facing the second orientation D2.
[0078] In an embodiment, the processor 52 can identify the change of the posture according to the head image. The processor 52 can take the head by the image capturing device 30, and acquire the head image therefrom. As shown in FIG. 1B , the head H is in front of the image capturing device 30. And the lens field of view of the image capturing device 30 covers the head H. The image features of the head image can be used to identify the change of the posture. The image features are, for example, Histogram of Oriented Gradient (HOG), Scale-Invariant Feature Transform (SIFT), Harr, or Speeded Up Robust Features (SURF). The image features can also be feature maps acquired by a machine learning model.
[0079] The head image is the image acquired by rotating the head from the first orientation to the second orientation. As shown in FIG. 6A , the posture of the head H facing the first orientation D1 to FIG. 6B , the posture of the head H facing the second orientation D2, the image capturing device 30 can continuously acquire the head image. The frequency of image acquisition can be 24, 30 or 60 per second, and is not limited thereto. The image capturing device 30 can also trigger the image acquisition function based on predetermined conditions (e.g. user operation or sound).
[0080] The processor 52 can identify a face in the head image. The identification can be based on object detection techniques. For example, the processor 52 can implement object detection based on a neural network based algorithm (e.g., YOLO (You only look once), Region Based Convolutional Neural Networks (R-CNN), or Fast R-CNN) or a feature matching based algorithm (e.g., Histogram of Oriented Gradients (HOG), Scale-Invariant Feature Transform (SIFT), Harr, or Speeded Up Robust Features (SURF) feature matching).
[0081] The processor 52 can also identify facial organs (e.g., eyes, mouth, or nose) in the head image. With a fixed lens of the image capture device 30, the head can not be able to capture all facial organs in some poses.
[0082] The processor 52 can define feature points for the head image. For example, the feature points are at the corners of the mouth, the tip of the nose, the upper rim of the ear, or the eyes, without limitation. The processor 52 can track the positions of one or more feature points in a plurality of consecutive head images. Changes in the pose of the head will be reflected in the changes in the positions of these feature points. For example,
[0083]
[0084] βH= RP' nose-x -RP nose-x …(2)
[0085] γH= RP nose-y -RP' nose-y …(3)
[0086] RP L-eye-y is the position of the left eye feature point on the vertical axis of the head image, RP R-eye-y is the position of the right eye feature point on the vertical axis of the head image, RP L-eye-x is the position of the left eye feature point on the horizontal axis of the head image, RP R-eye-x is the position of the right eye feature point on the horizontal axis of the head image, RP nose-x is the position of the nose feature point on the horizontal axis of the head image when the head is in the second orientation, RP nose-x is the position of the nose feature point on the horizontal axis of the head image when the head is in the first orientation, RP nose-x is the position of the nose feature point on the vertical axis of the head image when the head is in the second orientation, RP nose-y is the position of the nose feature point on the vertical axis of the head image when the head is in the first orientation.
[0087] In other embodiments, processor 52 may also apply neural network-based algorithms (e.g., YOLO, region-based convolutional neural networks (R-CNN), or Fast R-CNN) or feature matching-based algorithms (e.g., feature matching using Histogram of Oriented Gradients (HOG), Scale-Invariant Feature Transform (SIFT), Haar, or Speed-Up Robust Feature (SURF)). For example, the neural network is trained to learn the associations between multiple reference poses / rotation angles and image features. Alternatively, a lookup table records the associations between multiple reference poses / rotation angles and image features. Yet another example is a transformation function that records the associations between multiple reference poses / rotation angles and image features.
[0088] In another embodiment, the sound playback device 10 is provided with a motion sensor (e.g., a gyroscope, accelerometer, or inertial detection unit). The sensing data from the motion sensor can be used to analyze attitude changes.
[0089] Please refer to FIG. 2 The processor 52 adjusts the sound characteristics of the sound signal based on the directional difference between the sound source direction and the second orientation (step S230). Specifically, the directional difference is the angle between the sound source direction and the corrected second orientation (i.e., the rotation angle corresponding to the posture change, or the angle between the first orientation and the second orientation), and the corrected second orientation is the orientation after the posture change from the sound source direction. It is worth noting that, compared to traditional spatial audio settings that use the direction of the head towards the center of the computer screen as the sound source direction, the actual location of the sound source of the sound signal is not necessarily directly in front of the reference center. Therefore, the initial orientation of the posture change should be corrected to the sound source direction.
[0090] by FIG. 6B For example, the rotation angle from the first orientation D1 to the second orientation is θ. H (For example, including yaw angle α) H Pitch angle β H and roll angle γ H The initial orientation is corrected to the sound source direction SD2. The direction is then rotated by an angle θ from the sound source direction SD2. H This is the corrected second orientation ED2. Assume a rotation angle θ. H The angle is 20 degrees (corresponding to the left channel and -20 degrees to the right channel), and the sound source direction is 30 degrees (corresponding to the left channel and -30 degrees to the right channel). Therefore, the corrected second orientation ED2 is 10 degrees (i.e., 30 degrees - 20 degrees). Furthermore, the angle between the corrected second orientation ED2 and the sound source direction SD2 (i.e., the directional difference θ) is... D Same as the rotation angle θ H .
[0091] In an embodiment, the processor 52 can apply corresponding spatial effects to the plurality of orientation configurations of the head. In an embodiment, the processor 52 can apply the spatial effects or other effects through equalizers. The parameters of the equalizers can be gains / powers (for increasing or decreasing the responses of corresponding frequencies / bands) at the plurality of frequencies / bands. Different orientations can configure different parameters and be used to provide the spatial effects or other effects. Taking the spatial effects as an example, the processor 52 can transfer the two-channel sound signals to a surround sound field with a plurality of virtual speakers, adjust the frequency responses and / or phases from different directions based on the theory of Head Related Transfer Functions (HRTF), and then transfer the adjusted sound signals back to the two-channel stereo sound field signals.
[0092] For example, FIGS. 7A-7G is a diagram illustrating the parameters of the equalizers for the plurality of orientations according to an embodiment of the present application. Please refer to FIGS. 7A-7G , which are the parameters of the equalizers for the orientations of the head of 15 degrees, 30 degrees, 45 degrees, 60 degrees, 75 degrees, 90 degrees, and -15 degrees, respectively. Taking FIG. 7A and FIG. 7F as examples, compared with the parameters for the orientation of 15 degrees, the parameters for the orientation of 90 degrees have higher gains / powers (i.e., larger amplitudes) at high frequency bands (e.g., frequencies of 1K to 20K Hertz (Hz)).
[0093] FIG. 8A and FIG. 8B are diagrams illustrating the parameters of the equalizers for the two channels according to an embodiment of the present application. Please refer to FIG. 8A and FIG. 8B , FIG. 8A are the parameters of the left channel, and FIG. 8B are the parameters of the right channel. In an embodiment, in response to an increase in the parameter (e.g., gain / power) of the left channel at a certain frequency / band, the processor 52 can decrease the parameter of the right channel at the same frequency / band. Alternatively, in response to a decrease in the parameter (e.g., gain / power) of the left channel at a certain frequency / band, the processor 52 can increase the parameter of the right channel at the same frequency / band. In another embodiment, in response to an increase in the parameter (e.g., gain / power) of the right channel at a certain frequency / band, the processor 52 can decrease the parameter of the left channel at the same frequency / band. Alternatively, in response to a decrease in the parameter (e.g., gain / power) of the right channel at a certain frequency / band, the processor 52 can increase the parameter of the left channel at the same frequency / band. The equalizers of the two channels compensate for each other to maintain the overall power and keep the balance of the sound field. For example, when the head turns to the left, the power of the left channel increases and the power of the right channel decreases; when the head turns to the right, the power of the right channel increases and the power of the left channel decreases.
[0094] It should be noted that,FIGS. 7A-7F , FIG. 8A and FIG. 8B The parameters shown are for illustrative purposes only, and their values can be adjusted according to actual needs.
[0095] In this embodiment, the processor 52 can adjust the frequency response of the sound signal using a first parameter of the equalizer. The direction of the sound source corresponds to a second parameter of the equalizer, and the corrected second direction corresponds to a third parameter of the equalizer. The first, second, and third parameters have corresponding gain / power at one or more frequencies / bands. FIGS. 7A-7F , FIG. 8A and FIG. 8B As shown, different orientations have different parameter configurations.
[0096] The first parameter represents the gain / power difference between the second and third parameters at multiple frequencies / bands. For example, the mathematical expression is as follows:
[0097]
[0098] These are the first parameters of the left and right channels at frequency f, respectively. These are for the direction of the sound source (the left channel corresponds to...). And the right channel corresponds to The second parameter of the left and right channels at frequency f, and The second orientation for the correction (rotation angle of the left channel corresponding to) And the right channel corresponds to The third parameter for the left and right channels at frequency f.
[0099] by FIG. 7A and FIG. 7C For example, let's assume the rotation angle corresponding to the left channel. It is 30 degrees (rotated from 15 degrees to 45 degrees). FIG. 7A The parameter for 15 degrees is the second parameter, and FIG. 7C The parameter for 45 degrees is the third parameter. Therefore, the first parameter is... FIG. 7A The second parameter and FIG. 7C The third parameter is the gain / power difference at one or more frequencies / bands.
[0100] Another FIGS. 7A-7D and FIG. 7G For example, suppose the head rotates by an angle θ H It is 15 degrees, and the direction of the sound source is θ. S It can be 30 degrees or 60 degrees. Regarding existing technology, without considering the direction of the sound source θ... S In this case, the power adjustment parameters of the equalizer will be adopted. FIG. 7GThe parameters shown are for an orientation of -15 degrees. However, in this embodiment of the invention, if the sound source direction θ S If the temperature is 30 degrees, the power adjustment parameters of the equalizer will be based on... FIG. 7B The parameters shown are for an orientation of 30 degrees and FIG. 7A The gain / power difference between the parameters is shown for a 15-degree orientation (i.e., the corrected second orientation); if the sound source direction θ S If the angle is 60 degrees, the power adjustment parameters of the equalizer will be based on... FIG. 7D The parameters shown are for an orientation of 60 degrees and FIG. 7C The gain / power difference between parameters is shown for a 45-degree orientation (i.e., the corrected second orientation). The equalizer parameters used in this embodiment of the invention will differ from those used in the prior art.
[0101] In one embodiment, the processor 52 can adjust the signal delay of the two channels of the audio signal to a corrected delay. This corrected delay is the difference between a first delay and a second delay. The direction of the sound source corresponds to the first delay, and the corrected second direction corresponds to the second delay. For example, the mathematical expression is as follows:
[0102] Δτ=τ(θ S -θ H )-τ(θ S (6)
[0103] Δτ is the correction delay, τ(θ) S ) corresponds to the direction θ of the sound source. S The first delay, and τ(θ) S -θ H ) corresponds to the corrected second orientation (from the sound source direction θ) S After rotation angle θ H The processor 52 can delay at least one of the two-channel audio signals so that the signal delay of the two-channel audio signals is the same as the corrected delay. For example, the delay of the audio signals can be achieved by means of a buffer or delay circuit.
[0104] The adjusted sound signal (with spatial or other sound effects corresponding to the corrected second orientation) is used for playback by the sound playback device 10. For example, the processing unit 50 transmits the adjusted sound signal to the sound playback device 10. The sound playback device 10 can then play the adjusted sound signal.
[0105] FIG. 9A This is an illustration of the frequency response diagram of two channels at a zero-degree orientation, according to an embodiment of the present invention. Please refer to... FIG. 9A Experiments have shown that corresponding sound effect settings are provided for the two audio channels respectively. Although there are differences between the frequency response 910 of the left channel and the frequency response 920 of the right channel, they are both within an acceptable range.
[0106] FIG. 9B This is an illustration of the frequency response diagrams of different monomers at an orientation of zero degrees, according to an embodiment of the present invention. Please refer to... FIG. 9B Experiments have shown that even when using different individual units or different wearing methods (corresponding to different solid or dashed line segments in the figure), these differences in frequency response remain within an acceptable range.
[0107] FIGS. 10A-10F This is an illustration of the frequency response diagram of the left channel at different orientation angles according to an embodiment of the present invention. Please refer to... FIGS. 11A-11F These are the frequency responses of the head with an orientation of 0 degrees, 1010 (same as above). FIG. 9A The frequency response of the sound signal adjusted for the left channel was measured when the head orientation was 15 degrees, 30 degrees, 45 degrees, 60 degrees, 75 degrees and 90 degrees, respectively.
[0108] FIGS. 11A-11F This is an illustration of the frequency response diagram of the right channel at different orientation angles according to an embodiment of the present invention. Please refer to... FIGS. 9A These are the frequency responses 1110 (same) when the head is facing 0 degrees. FIGS. 10A-10F The frequency response of the sound signal adjusted for the right channel (920) and the frequency response of the head orientation (1120) were measured at 15 degrees, 30 degrees, 45 degrees, 60 degrees, 75 degrees and 90 degrees, respectively.
[0109] Depend on FIGS. 11A-11F and As can be seen, the actual measurement results are consistent with the theory. Furthermore, the greater the change in head rotation angle / posture, the more significant the impact of sound field changes on the high-frequency band (e.g., 2K to 10K Hz) of the sound signal (as shown in the figure, the greater the difference in sound pressure).
[0110] In summary, in the sound effect adjustment method and the computing device for sound effect adjustment in the embodiments of the present invention, the sound source direction of the sound signal is detected, a corrected orientation corresponding to head rotation is determined based on the sound source direction, and the sound characteristics of the sound signal are adjusted (e.g., to impart spatial or other sound effects) based on the sound source direction and the corrected orientation. Thus, suitable sound effects can be provided, and the auditory experience can be enhanced.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An audio effect adjustment method, characterized by, Suitable for processor implementation, the sound effect adjustment method comprises: determining a sound source direction corresponding to a sound feature of a sound signal, wherein the sound feature is related to at least one of an amplitude and a phase of the sound signal, and the sound signal is recorded from a sound source located at the sound source direction; determining a pose change of a head, wherein the pose change comprises a rotation angle of the head rotated from a first orientation to a second orientation, and the head is used to wear a sound playing device; and adjusting the sound feature of the sound signal according to a direction difference between the sound source direction and the second orientation, wherein the direction difference is an included angle between the sound source direction and a corrected second orientation, the corrected second orientation is an orientation after the sound source direction is changed by the pose change, and the adjusted sound signal is used to be played by the sound playing device.
2. The sound effect adjustment method of claim 1, wherein the sound feature comprises a frequency response, the frequency response is an amplitude of the sound signal corresponding to a plurality of frequencies, and the step of adjusting the sound feature of the sound signal according to the direction difference between the sound source direction and the second orientation comprises: adjusting the frequency response of the sound signal by a first parameter of an equalizer, wherein the sound source direction corresponds to a second parameter of the equalizer, the corrected second orientation corresponds to a third parameter of the equalizer, the first parameter, the second parameter and the third parameter have corresponding gains at the frequencies, and the first parameter is a gain difference of the second parameter and the third parameter at the frequencies respectively.
3. The sound effect adjustment method of claim 1, wherein the sound feature comprises a signal delay, the signal delay is a time difference between two channels of the sound signal, and the step of adjusting the sound feature of the sound signal according to the direction difference between the sound source direction and the second orientation comprises: adjusting the signal delay of the two channels of the sound signal to a corrected delay, wherein the corrected delay is a difference between a first delay and a second delay, the sound source direction corresponds to the first delay, and the corrected second orientation corresponds to the second delay.
4. The sound effect adjustment method of claim 1, wherein the step of determining the sound source direction corresponding to the sound feature of the sound signal comprises: determining the sound source direction by a direction identification model through inputting the sound feature of the sound signal to the direction identification model, wherein the direction identification model is trained by a machine learning algorithm to learn an association between a plurality of reference directions where a reference sound source is located and corresponding sound features.
5. The sound effect adjustment method of claim 1, wherein the step of determining the pose change of the head comprises: recognizing the pose change according to a plurality of head images, wherein the head images are images acquired when the head is rotated from the first orientation to the second orientation.
6. An arithmetic device for sound effect adjustment, characterized by comprising: comprises: a memory to store program code; and a processor coupled to the memory and configured to: determining a sound source direction corresponding to a sound feature of a sound signal, wherein the sound feature is related to at least one of an amplitude and a phase of the sound signal, and the sound signal is recorded from a sound source located at the sound source direction; determining a head pose change, wherein the head pose change comprises a rotation angle of the head rotated from a first orientation to a second orientation, and the head is used to wear a sound playing device; and adjusting the sound feature of the sound signal according to a direction difference between the sound source direction and the second orientation, wherein the direction difference is an included angle between the sound source direction and a modified second orientation, the modified second orientation is an orientation of the second orientation after the head pose change, and the adjusted sound signal is used to be played by the sound playing device.
7. The operation device for sound effect adjustment of claim 6, wherein the sound feature comprises a frequency response, the frequency response is an amplitude of the sound signal corresponding to a plurality of frequencies, and the processor is further configured to: adjust the frequency response of the sound signal by a first parameter of an equalizer, wherein the sound source direction corresponds to a second parameter of the equalizer, the modified second orientation corresponds to a third parameter of the equalizer, the first parameter, the second parameter and the third parameter have corresponding gains at the frequencies, and the first parameter is a gain difference of the second parameter and the third parameter at the frequencies respectively.
8. The operation device for sound effect adjustment of claim 6, wherein the sound feature comprises a signal delay, the signal delay is a time difference between two channels of the sound signal, and the processor is further configured to: adjust the signal delay of the two channels of the sound signal to a modified delay, wherein the modified delay is a difference between a first delay and a second delay, the sound source direction corresponds to the first delay, and the modified second orientation corresponds to the second delay.
9. The operation device for sound effect adjustment of claim 6, wherein the processor is further configured to: determine the sound source direction by inputting the sound feature of the sound signal to a direction identification model, wherein the direction identification model is trained by a machine learning algorithm to learn an association between a plurality of reference directions where a reference sound source is located and corresponding sound features.
10. The operation device for sound effect adjustment of claim 6, wherein the processor is further configured to: identify the head pose change according to a plurality of head images, wherein the head images are images taken from the head rotated from the first orientation to the second orientation.