An audio object-based sound reproduction method, apparatus, device, and storage medium
By decomposing and compensating the audio signal and transmitting it to different speakers to solve the problem of incorrect representation of the spatial information of the sound object, a more delicate and accurate audio playback effect is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUOGUANG ELECTRIC COMPANY LIMITED
- Filing Date
- 2022-10-31
- Publication Date
- 2026-04-28
AI Technical Summary
Within the framework of audio path, when the listening position deviates from the central axis of the left and right symmetrical speakers, existing technology cannot effectively adjust the relative output gain and delay between the speakers, resulting in incorrect representation of the spatial information of the sound object.
The left, middle, and right channel signals are decomposed into audio object and location-related metadata, and compensated according to the speaker arrangement in front of the listening position. The signals are then transmitted to the corresponding speakers for playback. At the same time, the surround sound signal is compensated and transmitted to the rear speakers for playback, and the subwoofer signal is transmitted to the third speaker for playback.
It enables individual adjustment of each sound object within the audio object framework, improving the subtlety and accuracy of sound spatial information, especially providing better playback details when the listening position is off-center from the vertical direction between the left and right symmetrical speakers.
Smart Images

Figure CN115696176B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing technology, and in particular to a sound reproduction method, apparatus, device, and storage medium based on audio objects. Background Technology
[0002] Within the audio path-based framework, each path in a multi-path signal contains multiple sound objects, and the spatial information of these sounds is expressed through the relative relationships between the path signals.
[0003] When the listening position deviates from the central axis of the left and right symmetrical speakers, it is necessary to use a certain sound object as a reference to adjust the relative output gain and delay between each speaker so that its spatial information meets the requirements, and then output the path signal from the speaker.
[0004] Adjusting the relative output gain and delay between speakers with a certain sound object as a reference will result in the incorrect representation of the spatial information of other sound objects. Summary of the Invention
[0005] This invention provides a sound reproduction method, apparatus, device, and storage medium based on audio objects, in order to solve the problem of expressing the spatial information of sound objects in a delicate and accurate manner.
[0006] According to one aspect of the present invention, a sound reproduction method based on an audio object is provided, the method comprising:
[0007] Receives audio signals from multiple channels, including left channel signal, center channel signal, right channel signal, surround sound signal, and subwoofer signal;
[0008] The left channel signal, the middle channel signal, and the right channel signal are decomposed into audio objects and positioning-related metadata;
[0009] The audio object is compensated based on the arrangement of the first speaker located in front of the listening position and the metadata, and then transmitted to the first speaker for playback;
[0010] The surround sound signal is compensated and transmitted to a second speaker located behind the listening position for playback;
[0011] The subwoofer signal is transmitted to the third speaker for playback.
[0012] According to another aspect of the present invention, an audio object-based sound reproduction apparatus is provided, the apparatus comprising:
[0013] An audio signal receiving module is used to receive audio signals from multiple channels, including left channel signal, middle channel signal, right channel signal, surround sound signal, and subwoofer signal.
[0014] The signal decomposition module is used to decompose the left channel signal, the middle channel signal and the right channel signal into audio objects and positioning-related metadata.
[0015] An audio object compensation module is used to compensate the audio object according to the arrangement of the first speaker located in front of the listening position and the metadata, and transmit it to the first speaker for playback;
[0016] A surround sound signal compensation module is used to compensate the surround sound signal and transmit it to a second speaker located behind the listening position for playback;
[0017] The subwoofer transmission and playback module is used to transmit the subwoofer signal to the third speaker for playback.
[0018] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0019] At least one processor; and
[0020] A memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method described in any embodiment of the present invention.
[0022] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program configured to cause a processor to execute the method described in any embodiment of the present invention.
[0023] In this embodiment, multiple channels of audio signals are received, including left channel, center channel, right channel, surround sound, and subwoofer signals. The left, center, and right channel signals are decomposed into audio objects and location-related metadata. The audio objects are compensated according to the arrangement of the first speaker located in front of the listening position and the metadata, and then transmitted to the first speaker for playback. The surround sound signal is compensated and transmitted to the second speaker located behind the listening position for playback. The subwoofer signal is transmitted to the third speaker for playback. This audio object-based sound reproduction method allows for individual adjustment of each sound object, resulting in a more nuanced and accurate representation of the spatial information of the sound object. By decomposing the channel signals into audio objects and location-related metadata, compensating the audio objects and surround sound signals according to their location, and then transmitting them to the speakers for playback, better sound reproduction details can be provided to the listener when the listening position deviates from the vertical position between the left and right symmetrical speakers within the audio object framework.
[0024] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a flowchart of an audio object-based sound reproduction method provided in Embodiment 1 of the present invention;
[0027] Figure 2 This is a speaker position relationship diagram provided according to Embodiment 1 of the present invention;
[0028] Figure 3 This is a schematic diagram of the structure of an audio object-based sound reproduction device according to Embodiment 2 of the present invention;
[0029] Figure 4 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. Detailed Implementation
[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0032] Example 1
[0033] Figure 1 This is a flowchart of a sound reproduction method based on audio objects provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where each sound object is adjusted individually to make the spatial information of the sound object more detailed. This method can be executed by a device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0034] Step 101: Receive audio signals from multiple channels, including left channel signal, center channel signal, right channel signal, surround sound signal, and subwoofer signal.
[0035] Audio signals represent mechanical waves and carry information about the wavelength and intensity changes of these waves. Based on their characteristics, mechanical waves can be categorized into regular and irregular signals. Regular signals can be further divided into categories such as music. A regular signal is a continuously changing analog signal, which can be represented by a continuous curve. A sine wave has three important parameters: frequency, amplitude, and phase.
[0036] The purpose of audio signals is to represent mechanical waves. Their strength is reflected in the intensity of the mechanical wave, while the perceived pitch is reflected in its wavelength. When representing mechanical waves, the signal is a continuous analog signal in both time and amplitude. Mechanical waves possess wave characteristics such as reflection, refraction, and diffraction. Analysis of mechanical waves reveals that they are composed of many components with different wavelengths, making them a composite signal. An important parameter of audio signals is bandwidth, which describes the range of wavelengths that make up the composite signal.
[0037] Wavelength refers to the distance between two adjacent peaks or troughs of a wave. Human perception of wavelength manifests as pitch, or tone in music. Pitch is determined by wavelength.
[0038] Within the framework based on audio frequency channels, each of the multiple channel signals contains various sound objects. The spatial information of various sounds is expressed through the relative relationship between the channel signals. When the listening position deviates from the central axis of the left and right symmetrical speakers, in order to obtain better sound reproduction details or results, multiple channels of audio signals are first received. The audio signals include the left channel signal, the middle channel signal, the right channel signal, the surround sound signal, and the subwoofer signal.
[0039] Step 102: Decompose the left channel signal, middle channel signal, and right channel signal into audio objects and location-related metadata.
[0040] After receiving the left, middle, and right channel signals, each signal is decomposed into distinct audio objects and location-related metadata. An audio object can be understood as a different sound source; decomposing the left, middle, and right channel signals into these audio objects represents the number of different sound sources. The metadata uses angles to indicate the location of the corresponding audio object.
[0041] In one embodiment of the present invention, step 102 may include the following steps:
[0042] Step 1021: Perform short-time Fourier transform on the left channel signal, the middle channel signal, and the right channel signal respectively to obtain the left channel spectrum signal, the middle channel spectrum signal, and the right channel spectrum signal.
[0043] L(m,k) = FFT(L*w,NFFT), C(m,k) = FFT(C*w,NFFT), R(m,k) = FFT(R*w,NFFT). Where L(m,k) is the NFFT-point short-time Fourier transform of the left channel signal, C(m,k) is the NFFT-point short-time Fourier transform of the middle channel signal, and R(m,k) is the NFFT-point short-time Fourier transform of the right channel signal.
[0044] The Short-Time Fourier Transform (SFT) is a mathematical transformation related to the Fourier Transform, used to determine the frequency and phase of a local sinusoidal wave in a time-varying signal. It is also known as the Windowed Fourier Transform. Before performing the Fourier Transform on the left, middle, and right channel signals, a time-finite window function is applied. It is assumed that the non-stationary left, middle, and right channel signals are stationary within the time interval of the analysis window. By shifting the window function along the time axis, segment-by-segment analysis is performed on the left, middle, and right channel signals to obtain a set of local "spectrums." After obtaining the spectral information of the left, middle, and right channel signals, filtering is applied to remove secondary frequency information considered noise. An inverse transform is then performed to obtain the denoised left, middle, and right channel signals.
[0045] Step 1022: Take the first logarithm of the ratio between the absolute value of the left acoustic spectrum signal and the absolute value of the middle acoustic spectrum signal, and amplify the first logarithm to obtain the first reference value.
[0046] Furthermore, the left channel signal is subjected to a short-time Fourier transform to obtain the left channel spectrum signal, and the middle channel signal is subjected to a short-time Fourier transform to obtain the middle channel spectrum signal. The first logarithm is taken between the absolute values of the left and middle channel spectrum signals. This first logarithm is then amplified and used as the first reference value. The first reference value represents the ratio of the frequency domain data of the left and middle channel signals. This ratio will differ depending on the location of the audio object. However, the first reference value will fluctuate within a certain range. By grouping these fluctuating first reference values into a single category—that is, a category representing a single audio object—the process is completed.
[0047] Step 1023: Take the second logarithm of the ratio between the absolute value of the right acoustic spectrum signal and the absolute value of the middle acoustic spectrum signal, and amplify the second logarithm to obtain the second reference value.
[0048] Furthermore, the right channel signal is subjected to a short-time Fourier transform to obtain the right channel spectrum signal, and the middle channel signal is subjected to a short-time Fourier transform to obtain the middle channel spectrum signal. The ratio between the absolute values of the right channel spectrum signal and the absolute values of the middle channel spectrum signal is taken as the second logarithm. This second logarithm is then amplified and used as the second reference value. The second reference value represents the ratio of the frequency domain data of the right channel signal and the middle channel signal. If the location of the audio object is different, the ratio of the second reference value will be different. However, the second reference value will fluctuate within a certain range. By grouping the second reference values fluctuating within a certain range into one category, that is, a category is an audio object.
[0049] For example, after performing a short-time Fourier transform on the left channel signal, middle channel signal, and right channel signal to obtain the left channel spectrum signal L(m,k), the middle channel spectrum signal C(m,k), and the right channel spectrum signal R(m,k), the ratio between the absolute value of the left channel spectrum signal and the absolute value of the middle channel spectrum signal is taken as the first logarithm, and then amplified to obtain... Let P1(m,k) record the ratio of the absolute value of the left acoustic spectrum signal to the amplified absolute value of the middle acoustic spectrum signal. Take the second logarithm of the ratio between the absolute values of the right acoustic spectrum signal and the middle acoustic spectrum signal, and amplify this second logarithm to obtain... Use P2(m,k) to record the ratio of the absolute value of the right acoustic spectrum signal to the absolute value of the middle acoustic spectrum signal after amplification.
[0050] Step 1024: Determine the first reference range and the second reference range respectively.
[0051] When short-time Fourier transforms are performed on the left, middle, and right channel signals to obtain the left, middle, and right audio spectrum signals, a first reference range and a second reference range are determined. The first reference value fluctuates within the first reference range, and values fluctuating within a certain range are grouped into one category. The second reference value fluctuates within the second reference range, and values fluctuating within a certain range are grouped into one category.
[0052] In one embodiment of the present invention, step 1024 may include the following steps:
[0053] Step 10241: Generate multiple first histograms of the first group for the left channel signal and the middle channel signal.
[0054] If the location of the audio object is different, the obtained first reference value will be different, and the first reference value fluctuates within a certain range. The fluctuating first reference values within this range are grouped into one category, and each category is represented by a group from the different groups in the first histogram. Furthermore, multiple first histograms for the left channel and middle channel signals are generated for each first group. Data with similar characteristics and fluctuating first reference values within a certain range are grouped into one of the first groups in the first histogram.
[0055] The first midpoint of the first x-axis of the first histogram in each group is calculated using the following formula:
[0056]
[0057] Among them, Z n L1 is the first midpoint of the first x-axis of the first histogram of the nth group, L1 is the number of the first group, and α is the first adjustment coefficient.
[0058] The first ordinate of the first histogram in each group is calculated using the following formula:
[0059]
[0060]
[0061] Where H1(n) is the first ordinate of the first histogram of the first group in the nth iteration, m is time, k is frequency, and P1(m,k) is the first reference value. Ω represents the value when ||P1(m,k)-Z n | Satisfies less than or equal to The case is the set of (m, k). H1(n) is represented as the ordinate of the histogram, and its value is equal to the sum of the number of elements in (m, k) in the set Ω.
[0062] Step 10242: Construct the first objective function using the following formula. When the value of the first objective function is maximized, divide the multiple sets of first histograms into multiple first categories in sequence:
[0063]
[0064]
[0065]
[0066]
[0067]
[0068] The sequences in the histogram are divided into M classes, and the boundary points of the sequences are denoted as t1, t2, ..., t3. M-1 The first class C1 contains a sequence of histogram groups [1, t1], ..., the Mth class C M The sequence of groups containing histograms is [t] M-1 +1, L1]. Let be the first objective function. In the M classes, calculate the first objective function in each different class, compare the first objective function values of different classes, and obtain the largest first objective function value.
[0069] Step 10245: Generate the first range using the following formula:
[0070]
[0071] Among them, Z n Let α be the first midpoint of the first x-axis of the first histogram of the nth group, α be the first adjustment coefficient, m be time, k be frequency, and t be the first midpoint of the first x-axis. iLet M be the first histogram of the boundary point of the first class, where M is the number of the first class.
[0072] Step 10246: Generate multiple second histograms for the middle channel signal and the right channel signal.
[0073] If the location of the audio object is different, the obtained second reference value will be different, and the second reference value fluctuates within a certain range. These fluctuating second reference values are grouped into one category and represented by a group from different sets of the second histogram. Furthermore, multiple second histograms are generated for the middle channel and right channel signals. Second reference values fluctuating within a certain range, i.e., data with similar characteristics, are grouped into one of the second groups within the second histogram.
[0074] The second midpoint of the second x-axis in each of the second groups of the second histogram is calculated using the following formula:
[0075]
[0076] Among them, Z n ' is the second midpoint of the second horizontal axis of the second histogram of the nth group, L1 is the number of the second group, and α is the second adjustment coefficient.
[0077] The first ordinate of the first histogram in each second group is calculated using the following formula:
[0078]
[0079]
[0080] Where H1'(n) is the second ordinate of the nth second histogram of the second group, m is time, k is frequency, and P2(m',k') is the second reference value. Ω is represented as ||P2(m',k')-Z n | Satisfies less than or equal to The case is the set of (m', k'). H1'(n) is represented as the ordinate of the histogram, and its value is equal to the sum of the number of elements in the set Ω' of (m', k').
[0081] The second objective function is constructed using the following formula. When the value of the second objective function is maximized, the multiple sets of second histograms are sequentially divided into multiple second classes:
[0082]
[0083]
[0084]
[0085]
[0086]
[0087] The sequences in the histogram are divided into M' classes, and the boundary points of the sequences are denoted as t1, t2, ..., tt. M’-1 The first class C i The sequence of groups containing histograms is [1, t1], ..., the M'th class C M The sequence of groups containing histograms is [t] M;-1 +1, L2]. Let M' be the second objective function. In each of the M' classes, the second objective function is calculated. The values of the second objective function in different classes are compared, and the maximum value of the second objective function is obtained.
[0088] Step 102410: Generate the second range using the following formula:
[0089]
[0090] in, Let L2 be the second midpoint of the second x-axis of the second histogram of the n'th group, L2 be the number of the second group, α' be the second adjustment coefficient, and t be the second midpoint of the second x-axis of the second histogram. i ' is the second histogram at the boundary point of the second class at the i'th point, and M' is the number of the second class.
[0091] Step 1025: If the first reference value is within the first reference range, then decompose the left acoustic spectrum signal and the middle acoustic spectrum signal into audio objects and location-related metadata.
[0092] A first reference value is obtained, which is the amplified value obtained by taking the first logarithm of the ratio between the absolute values of the left and middle acoustic spectrum signals. When the first reference value is within the first reference range, the left and middle acoustic spectrum signals are decomposed into audio objects and metadata. The metadata represents the horizontal azimuth angle of the audio object, i.e., the position of the audio object.
[0093] The left and middle acoustic spectrum signals are decomposed into frequency domain signals of the audio object using the following formula:
[0094]
[0095] The left and middle acoustic spectrum signals are decomposed to generate location-related metadata using the following formula:
[0096]
[0097]
[0098] Where L(m,k) is the left acoustic spectrum signal, C(m,k) is the middle acoustic spectrum signal, and S... i (m,k) represents the audio object, g i For gain, θ i Here, j is the metadata, ∠L(m,k) is the imaginary number, and ∠L(m,k) is the phase of the left acoustic spectrum signal;
[0099] Step 1026: If the second reference value is within the second reference range, then the right acoustic spectrum signal and the middle acoustic spectrum signal are decomposed into audio objects and positioning-related metadata in the first manner.
[0100] A second reference value is obtained, which is the second logarithm of the ratio between the absolute values of the right-hand and middle-hand acoustic spectrum signals, amplified from the second logarithm. When the second reference value is within the second reference range, the right-hand and middle-hand acoustic spectrum signals are decomposed into audio objects and metadata according to the first method. The metadata represents the horizontal azimuth angle of the audio object, i.e., the position of the audio object.
[0101] In one embodiment of the present invention, step 1026 may include the following steps:
[0102] Step 10261: Decompose the middle acoustic spectrum signal and the right acoustic spectrum signal into the frequency domain signal of the audio object using the following formula:
[0103] The following formula is the first method to decompose the right acoustic spectrum signal and the middle acoustic spectrum signal into audio objects and metadata. The first method contains two formula steps, namely step 10261 and step 10262.
[0104]
[0105] Step 10262: Decompose the center acoustic spectrum signal and the right acoustic spectrum signal to generate location-related metadata using the following formula:
[0106]
[0107]
[0108] Where R(m,k) is the right acoustic spectrum signal, C(m,k) is the middle acoustic spectrum signal, and S i (m,k) represents the audio object, g i For gain, θ i Here, j is the metadata, ∠R(m,k) is the imaginary number, and ∠R(m,k) is the phase of the right acoustic spectrum signal;
[0109] Step 1027: If the first reference value is less than the first reference range or the second reference value is less than the second reference range, then the right acoustic spectrum signal and the middle acoustic spectrum signal are decomposed into audio objects and positioning-related metadata in the second manner.
[0110] Obtain the first reference value and the second reference value, and use the calculated first reference range and the second reference range as the basis for judging the range of the first reference value and the second reference value. When the first reference value is less than the first reference range or the second reference value is less than the second reference range, decompose the right acoustic spectrum signal and the middle acoustic spectrum signal into audio objects and metadata according to the second method. The metadata represents the horizontal azimuth angle of the audio object, that is, the position of the audio object.
[0111] In one embodiment of the present invention, step 1027 may include the following steps:
[0112] Step 10271: Decompose the middle acoustic spectrum signal and the right acoustic spectrum signal into the frequency domain signal of the audio object using the following formula:
[0113] The following formula decomposes the right acoustic spectrum signal and the middle acoustic spectrum signal into audio objects and metadata in a second way, which contains two steps, namely step 10271 and step 10272.
[0114]
[0115] Where R(m,k) is the right acoustic spectrum signal, C(m,k) is the middle acoustic spectrum signal, and S i (m,k) is the audio object, j is the imaginary number, and ∠R(m,k) is the phase of the right-hand audio spectrum signal;
[0116] Step 10272: Set the location-related metadata to 0°.
[0117] When the metadata is set to 0°, that is, when the angle between the listening position and the audio object position is 0, S i (m,k) represents the frequency domain signal of the audio object directly in front. Obtain metadata at 0°. When the left and right audio signals are decomposed into audio objects, extract metadata at 0° and merge the audio object signals with metadata at 0°. Using the area directly in front of the listening position (i.e., with the listening position as the origin of the coordinate axis) as the origin, the area directly in front (with metadata at 0°) is defined as the positive half-axis of the Y-axis.
[0118] Step 103: Compensate the audio object according to the arrangement of the first speaker located in front of the listening position and the metadata, and transmit it to the first speaker for playback.
[0119] like Figure 2The diagram shows the listening position and the arrangement of the first speakers, which, from left to right, include a left speaker L, a center speaker C, and a right speaker R. The audio is compensated based on the arrangement of the left speaker L, center speaker C, and right speaker R in front of the listening position and the horizontal azimuth of the audio object, and then transmitted to the corresponding left speaker L, center speaker C, and right speaker R for playback.
[0120] In one embodiment of the present invention, step 103 may include the following steps:
[0121] Step 1031: Measure the first direction vector from the listening position to the left speaker, the third direction vector to the middle speaker, and the second direction vector to the right speaker.
[0122] like Figure 2 As shown, the first direction vector from the listening position Q to the left speaker L is measured, the third direction vector from the listening position Q to the middle speaker C is measured, and the second direction vector from the listening position Q to the right speaker R is measured.
[0123] Step 1032: Measure the first angle between the front of the listening position and the left speaker, the third angle between the front and the middle speaker, and the second angle between the front and the right speaker.
[0124] like Figure 2 As shown, the first angle γ1 between the listening position Q and the left speaker is measured, the third angle γ3 between the listening position Q and the center speaker is measured, and the first and second angles γ2 between the listening position Q and the center speaker are measured. The listening position is taken as the origin, and the area directly in front of the listening position is taken as the baseline (taking the listening position as the origin of the coordinate axis as an example, the angle starting from the positive direction of the Y-axis and rotating clockwise is the negative angle, with a range of (0°, -180°), and the angle starting from the positive direction of the Y-axis and rotating counterclockwise is the positive angle, with a range of (0°, 180°). The first angle γ1, the second angle γ2, and the third angle γ3 are measured.
[0125] Step 1033: If the metadata is greater than or equal to the first angle, then transmit all audio objects to the left speaker for playback.
[0126] Obtain the audio object and metadata (horizontal azimuth angle of the audio object) from the decomposed left channel signal. Extract the first angle between the listening position and the left speaker, and compare the metadata (horizontal azimuth angle of the audio object) with the first angle. If the metadata is greater than or equal to the first angle, then transmit the entire audio object to the left speaker for playback.
[0127] Step 1034: If the metadata is less than or equal to the second angle, then transmit all audio objects to the right speaker for playback.
[0128] Obtain the audio object and metadata (horizontal azimuth angle of the audio object) from the right channel signal. Extract the second angle between the listening position and the right speaker, and compare the metadata (horizontal azimuth angle of the audio object) with the second angle. If the metadata is less than or equal to the second angle, then transmit the entire audio object to the right speaker for playback.
[0129] Step 1035: If the metadata is greater than 0 and less than the first angle, or if the metadata is less than 0 and greater than the third angle, then the audio object is compensated according to the first direction vector and the third direction vector and transmitted to the left speaker and the middle speaker for playback.
[0130] Obtain the decomposed audio object and its metadata (horizontal azimuth angle of the audio object). Extract the first angle from the listening position to the left speaker and the third angle from the listening position to the center speaker. Compare the metadata (horizontal azimuth angle of the audio object) with 0°. If the metadata (horizontal azimuth angle of the audio object) is greater than 0 and less than the first angle, or less than 0 and greater than the third angle, then compensate the audio object according to the first direction vector and the third direction vector and transmit it to the left and center speakers for playback.
[0131] The first coefficient function is constructed using the following formula:
[0132] When the frequency of the audio object is less than or equal to a preset first threshold, a first coefficient function is constructed.
[0133]
[0134] A1 2 +A2 2 =1
[0135] The second coefficient function is constructed using the following formula:
[0136] When the frequency of the audio object is greater than a preset first threshold, a second coefficient function is constructed.
[0137]
[0138] A3 2 +A4 2 =1
[0139] in, The first direction vector, The first distance between the listening position and the left speaker For a third-direction vector, The third distance between the listening position and the right speaker, θ iThe data consists of A1 (first coefficient), A2 (second coefficient), A3 (third coefficient), and A4 (fourth coefficient).
[0140] By solving the first coefficient function, we obtain the first and second coefficients. By solving the second coefficient function, we obtain the third and fourth coefficients.
[0141] In one embodiment of the present invention, step 1035 may include the following steps:
[0142] Step 10351: When the frequency of the audio object is less than or equal to a preset first threshold, calculate the signal transmitted to the left speaker and the signal transmitted to the center speaker using the following formulas:
[0143] When the frequency of the audio object is less than or equal to the preset first threshold, the first coefficient and the second coefficient are obtained by solving the first coefficient function. The first coefficient, the second coefficient, the first direction vector, the second direction vector, the third direction vector, the preset hyperparameters, and the extracted audio object are substituted into the following formula to calculate the signal transmitted to the left speaker and the middle speaker.
[0144]
[0145]
[0146] Where L is the signal transmitted from the audio object to the left speaker, C is the signal transmitted from the audio object to the center speaker, and S is the signal transmitted from the audio object to the center speaker. i For audio objects, A1 is the first coefficient. The first direction vector, The first distance between the listening position and the left speaker The second distance between the listening position and the right speaker. For a third-direction vector, A1 represents the third distance between the listening position and the center loudspeaker, A2 represents the second coefficient, q represents the hyperparameter, j represents the imaginary number, and f represents the frequency.
[0147] Step 10352: When the frequency of the audio object is greater than a preset first threshold, calculate the signal transmitted to the left speaker and the signal transmitted to the center speaker using the following formulas:
[0148] When the frequency of the audio object is greater than the preset first threshold, the third and fourth coefficients are obtained by solving the second coefficient function. The third and fourth coefficients, the first direction vector, the second direction vector, the third direction vector, the preset hyperparameters, and the extracted audio object are substituted into the following formula to calculate the signal transmitted to the left speaker and the middle speaker.
[0149]
[0150]
[0151] Where L is the mid-to-high frequency signal transmitted from the audio object to the left speaker, C is the mid-to-high frequency signal transmitted from the audio object to the middle speaker, and S... i For audio objects, A3 is the third coefficient. The first direction vector, The first distance between the listening position and the left speaker The second distance between the listening position and the right speaker. For a third-direction vector, A4 is the third distance between the listening position and the center speaker, P is the hyperparameter, j is the imaginary number, and f is the frequency.
[0152] Step 1036: If the metadata is greater than the second included angle and less than the third included angle, then the audio object is compensated according to the second direction vector and the third direction vector and transmitted to the middle speaker and the right speaker for playback.
[0153] Extract the decomposed audio object and its metadata (horizontal azimuth angle of the audio object). Compare the decomposed metadata (horizontal azimuth angle of the audio object) with 0°. If the decomposed metadata (horizontal azimuth angle of the audio object) is greater than the second included angle and less than the third included angle, then the audio object is compensated according to the first direction vector and the third direction vector and transmitted to the left and center speakers for playback.
[0154] The third coefficient function is constructed using the following formula:
[0155] When the frequency of the audio object is less than or equal to a preset first threshold, a third coefficient function is constructed.
[0156] The third coefficient function is constructed using the following formula:
[0157]
[0158] A'3 2 +A'4 2 =1
[0159] The fourth coefficient function is constructed using the following formula:
[0160] When the frequency of the audio object is greater than the preset first threshold, a fourth coefficient function is constructed.
[0161]
[0162] A'3 2 +A'4 2 =1
[0163] in, For the second direction vector, The first distance between the listening position and the left speaker For a third-direction vector, A1' is the third distance between the listening position and the right speaker, A2' is the fifth coefficient, and A'3 is the sixth coefficient. 2 The seventh coefficient, A'4 2 The eighth coefficient, θ i Metadata;
[0164] By solving the third coefficient function, we obtain the fifth and sixth coefficients. By solving the fourth coefficient function, we obtain the seventh and eighth coefficients.
[0165] In one embodiment of the present invention, step 1036 may include the following steps:
[0166] Step 10361: When the frequency of the audio object is less than or equal to a preset threshold, calculate the signal transmitted to the middle speaker and the signal transmitted to the right speaker respectively using the following formula:
[0167] When the frequency of the audio object is less than or equal to the preset first threshold, the fifth and sixth coefficients are obtained by solving the third coefficient function. The fifth and sixth coefficients, the first direction vector, the second direction vector, the third direction vector, the preset hyperparameters, and the extracted audio object are substituted into the following formula to calculate the signal transmitted to the middle speaker and the right speaker.
[0168]
[0169]
[0170] Where C is the signal transmitted from the audio object to the center speaker, R is the signal transmitted from the audio object to the right speaker, and S... i For audio objects, A1' is the fifth coefficient. The first direction vector, The first distance between the listening position and the left speaker The second distance between the listening position and the right speaker. For a third-direction vector, A2' is the third distance between the listening position and the center loudspeaker, J is the hyperparameter, j is the imaginary number, and f is the frequency.
[0171] When the frequency of the audio object is greater than a preset threshold, the mid-high frequency signals transmitted to the center speaker and the mid-high frequency signals transmitted to the right speaker are calculated using the following formulas:
[0172] When the frequency of the audio object is greater than the preset first threshold, the seventh and eighth coefficients are obtained by solving the fourth coefficient function. The seventh and eighth coefficients, the first direction vector, the second direction vector, the third direction vector, the preset hyperparameters, and the extracted audio object are substituted into the following formula to calculate the signal transmitted to the middle speaker and the right speaker.
[0173]
[0174]
[0175] Where R is the signal transmitted from the audio object to the right speaker, C is the signal transmitted from the audio object to the middle speaker, and S... i For audio objects, A'3 2 The seventh coefficient The first direction vector, The first distance between the listening position and the left speaker The second distance between the listening position and the right speaker. For a third-direction vector, The third distance between the listening position and the center speaker, A'4 2 is the eighth coefficient, F is the hyperparameter, j is the imaginary number, and f is the frequency.
[0176] Step 104: Compensate the surround sound signal and transmit it to the second speaker located behind the listening position for playback.
[0177] Surround sound refers to the pre-encoding of the left, right, center, and rear surround channels into a two-channel signal. During playback, the sound is enhanced by a decoder and power amplifier, with the center speaker and rear surround speakers boosting the stereo effect, resulting in a more realistic sound than a two-channel live performance.
[0178] like Figure 2 As shown, the second loudspeaker, from left to right, includes a left surround speaker LS and a right surround speaker RS. The surround signal includes a left surround signal and a right surround signal. Surround signal compensation is calculated based on the positions of the second loudspeakers. First, the fourth distance between the listening position and the left surround speaker LS, and the fifth distance between the listening position and the right surround speaker RS are measured. The left surround signal is compensated based on the fourth and fifth distances and then transmitted to the left surround speaker LS for playback.
[0179] In an embodiment of the present invention, step 104 may include the following steps:
[0180] Step 1041: Compensate the surround sound signal using the following formula and then output it to the left surround speaker for playback:
[0181]
[0182] Among them, S' LS For the compensated surround sound signal, S LS For the left surround signal, r4 is the fourth distance, r5 is the fifth distance, j is the imaginary part, f is the frequency, and G is the hyperparameter;
[0183] The right surround sound signal is compensated based on the fourth and fifth distances and then transmitted to the right surround speaker RS for playback.
[0184] Step 1042: Compensate the surround sound signal using the following formula and then output it to the right surround speaker for playback:
[0185]
[0186] Among them, S' RS For the compensated surround sound signal, S RS For the right surround signal, r4 is the fourth distance, r5 is the fifth distance, j is the imaginary part, f is the frequency, and Q is the hyperparameter.
[0187] Step 105: Transmit the subwoofer signal to the third speaker for playback.
[0188] Subwoofers are audible in the lowest frequency range of 20 Hz to 120 Hz. Their function is to enhance the low-frequency range, that is, to boost the long-wave portion of the audio signal. Therefore, a larger third speaker is needed to produce a powerful low-frequency effect.
[0189] In this embodiment, multiple channels of audio signals are received, including left channel, center channel, right channel, surround sound, and subwoofer signals. The left, center, and right channel signals are decomposed into audio objects and location-related metadata. The audio objects are compensated according to the arrangement of the first speaker located in front of the listening position and the metadata, and then transmitted to the first speaker for playback. The surround sound signal is compensated and transmitted to the second speaker located behind the listening position for playback. The subwoofer signal is transmitted to the third speaker for playback. This audio object-based sound reproduction method allows for individual adjustment of each sound object, resulting in a more nuanced and accurate representation of the spatial information of the sound object. By decomposing the channel signals into audio objects and location-related metadata, compensating the audio objects and surround sound signals according to their location, and then transmitting them to the speakers for playback, better sound reproduction details can be provided to the listener when the listening position deviates from the vertical position between the left and right symmetrical speakers within the audio object framework.
[0190] Example 2
[0191] Figure 3 This is a schematic diagram of the structure of a sound reproduction method apparatus based on an audio object provided in Embodiment 2 of the present invention. Figure 3 As shown, the device includes:
[0192] The audio signal receiving module 301 is used to receive audio signals from multiple channels, including left channel signal, middle channel signal, right channel signal, surround sound signal and subwoofer signal;
[0193] The signal decomposition module 302 is used to decompose the left channel signal, the middle channel signal and the right channel signal into audio objects and positioning-related metadata.
[0194] The audio object compensation module 303 is used to compensate the audio object according to the arrangement of the first speaker located in front of the listening position and the metadata, and transmit it to the first speaker for playback;
[0195] The surround sound signal compensation module 304 is used to compensate the surround sound signal and transmit it to the second speaker located behind the listening position for playback;
[0196] The subwoofer transmission and playback module 305 is used to transmit the subwoofer signal to the third speaker for playback.
[0197] In one embodiment of the present invention, the signal decomposition module 302 includes:
[0198] The signal transformation module is used to perform short-time Fourier transform on the left channel signal, the middle channel signal and the right channel signal respectively to obtain the left channel spectrum signal, the middle channel spectrum signal and the right channel spectrum signal;
[0199] The first logarithmic amplification module is used to take the first logarithm of the ratio between the absolute value of the left acoustic spectrum signal and the absolute value of the middle acoustic spectrum signal, and amplify the first logarithm to obtain a first reference value.
[0200] The second logarithmic amplification module is used to take the second logarithm of the ratio between the absolute value of the right acoustic spectrum signal and the absolute value of the middle acoustic spectrum signal, and amplify the second logarithm to obtain a second reference value;
[0201] The range determination module is used to determine the first reference range and the second reference range, respectively.
[0202] The signal decomposition module is used to decompose the left acoustic spectrum signal and the middle acoustic spectrum signal into audio objects and positioning-related metadata if the first reference value is within the first reference range.
[0203] The first spectrum signal decomposition module is used to decompose the right acoustic spectrum signal and the middle acoustic spectrum signal into audio objects and positioning-related metadata in a first manner if the second reference value is within the second reference range.
[0204] The second spectrum signal decomposition module is used to decompose the right acoustic spectrum signal and the middle acoustic spectrum signal into audio objects and positioning-related metadata in a second manner if the first reference value is less than the first reference range or the second reference value is less than the second reference range.
[0205] In one embodiment of the present invention, the range determination module includes:
[0206] The first histogram generation module is used to generate multiple first histograms of the first group for the left channel signal and the middle channel signal;
[0207] The first horizontal coordinate calculation module is used to calculate the first midpoint of the first horizontal coordinate for each of the first histograms in the first group using the following formula:
[0208]
[0209] The first ordinate calculation module is used to calculate the first ordinate of the first histogram in each first group using the following formula:
[0210]
[0211]
[0212] The first histogram partitioning module is used to construct a first objective function using the following formula. When the value of the first objective function is maximized, multiple sets of the first histograms are sequentially divided into multiple first categories:
[0213]
[0214]
[0215]
[0216]
[0217]
[0218] The first range generation module is used to generate the first range using the following formula:
[0219]
[0220] Among them, Z nLet L1 be the first midpoint of the first abscissa of the first histogram of the nth group, α be the number of the first group, α be the first adjustment coefficient, H1(n) be the first ordinate of the first histogram of the nth first group, m be time, k be frequency, P1(m,k) be the first reference value, and C be the first reference value. i For the i-th class of the first type, δ1 is the first objective function, and t i The first histogram is located at the i-th boundary point of the first class, and M is the number of the first class;
[0221] The second histogram generation module is used to generate multiple second groups of second histograms for the middle channel signal and the right channel signal;
[0222] The second horizontal coordinate calculation module is used to calculate the second midpoint of the second horizontal coordinate for each of the second histograms in the second group using the following formula:
[0223]
[0224] The first ordinate calculation module is used to calculate the first ordinate of the first histogram in each second group using the following formula:
[0225]
[0226]
[0227] The second histogram partitioning module is used to construct a second objective function using the following formula. When the value of the second objective function is maximized, multiple sets of second histograms are sequentially divided into multiple second categories:
[0228]
[0229]
[0230]
[0231]
[0232]
[0233] The second range generation module is used to generate the second range using the following formula:
[0234]
[0235] in, Let L2 be the second midpoint of the second abscissa of the second histogram of the n'th group, α' be the number of the second group, α' be the second adjustment coefficient, H1'(n) be the second ordinate of the second histogram of the n'th second group, m be time, k be frequency, P2 be the first reference value, and C be the second midpoint of the second abscissa of the second histogram of the n'th group. i ' represents the i-th class of the first type, δ1' represents the second objective function, and t i ' is the second histogram at the boundary point of the second class at the i'th point, and M' is the number of the second class;
[0236] The signal decomposition module includes:
[0237] The third spectrum signal decomposition module is used to decompose the left acoustic spectrum signal and the middle acoustic spectrum signal into frequency domain signals of the audio object using the following formula:
[0238]
[0239] The first metadata generation module is used to decompose the left acoustic spectrum signal and the middle acoustic spectrum signal to generate positioning-related metadata using the following formula:
[0240]
[0241]
[0242] Where L(m,k) is the left acoustic spectrum signal, C(m,k) is the middle acoustic spectrum signal, and S i (m,k) represents the audio object, g i For gain, θ i Here, j is the metadata, ∠L(m,k) is the imaginary number, and ∠L(m,k) is the phase of the left acoustic spectrum signal;
[0243] The first spectrum signal decomposition module includes:
[0244] The fourth spectrum decomposition module is used to decompose the middle acoustic spectrum signal and the right acoustic spectrum signal into frequency domain signals of the audio object using the following formula:
[0245]
[0246] The second metadata generation module is used to decompose the middle acoustic spectrum signal and the right acoustic spectrum signal to generate positioning-related metadata using the following formula:
[0247]
[0248]
[0249] Where R(m,k) is the right acoustic spectrum signal, C(m,k) is the middle acoustic spectrum signal, and S i (m,k) represents the audio object, g i For gain, θ i Here, j is the metadata, ∠R(m,k) is the imaginary number, and ∠R(m,k) is the phase of the right acoustic spectrum signal;
[0250] The second spectrum signal decomposition module includes:
[0251] The fifth spectrum decomposition module is used to decompose the middle acoustic spectrum signal and the right acoustic spectrum signal into frequency domain signals of the audio object using the following formula:
[0252]
[0253] Where R(m,k) is the right acoustic spectrum signal, C(m,k) is the middle acoustic spectrum signal, and S i (m,k) is the audio object, j is the imaginary number, and ∠R(m,k) is the phase of the right-hand audio spectrum signal;
[0254] Set the location-related metadata to 0°.
[0255] In one embodiment of the present invention, the first loudspeaker includes, from left to right, a left loudspeaker, a center loudspeaker, and a right loudspeaker, and the audio object compensation module 303 includes:
[0256] The vector measurement module is used to measure the first direction vector from the listening position to the left speaker, the third direction vector to the middle speaker, and the second direction vector to the right speaker.
[0257] Angle measurement module is used to measure the first angle between the front of the listening position and the left speaker, the third angle between the listening position and the middle speaker, and the second angle between the listening position and the right speaker.
[0258] The first audio object playback module is used to transmit all the audio objects to the left speaker for playback if the metadata is greater than or equal to the first angle.
[0259] The second audio object playback module is used to transmit all the audio objects to the right speaker for playback if the metadata is less than or equal to the second included angle.
[0260] The third audio object playback module is used to compensate the audio object according to the first direction vector and the third direction vector and transmit it to the left speaker and the middle speaker for playback if the metadata is greater than 0 and less than the first angle, or if the metadata is less than 0 and greater than the third angle.
[0261] The fourth audio object playback module is used to compensate the audio object according to the second direction vector and the third direction vector and transmit it to the middle speaker and the right speaker for playback if the metadata is greater than or equal to the third included angle and less than or equal to the second included angle.
[0262] In one embodiment of the present invention, the third audio object playback module includes:
[0263] The first coefficient function construction module is used to construct the first coefficient function using the following formula:
[0264]
[0265] A1 2 +A2 2 =1
[0266] The second coefficient function construction module is used to construct the second coefficient function using the following formula:
[0267]
[0268] A3 2 +A4 2 =1
[0269] in, For the first direction vector, The first distance between the listening position and the left speaker. For the third-direction vector, The third distance θ between the listening position and the right speaker i For the aforementioned metadata, A1 is the first coefficient, A2 is the second coefficient, A3 is the third coefficient, and A4 is the fourth coefficient;
[0270] The first coefficient function solving module is used to solve the first coefficient function to obtain the first coefficient and the second coefficient;
[0271] The second coefficient function solving module is used to solve the second coefficient function to obtain the third coefficient and the fourth coefficient;
[0272] The first signal calculation module is used to calculate the signal transmitted to the left speaker and the signal transmitted to the middle speaker respectively using the following formula when the frequency of the audio object is less than or equal to a preset first threshold:
[0273]
[0274]
[0275] Wherein, L is the signal transmitted from the audio object to the left speaker, C is the signal transmitted from the audio object to the center speaker, and S... i For the audio object, A1 is the first coefficient. For the first direction vector, The first distance between the listening position and the left speaker. The second distance between the listening position and the right speaker. For the third-direction vector, A1 is the third distance between the listening position and the loudspeaker, A2 is the second coefficient, q is the hyperparameter, j is the imaginary number, and f is the frequency.
[0276] The fifth audio object playback module is used to calculate the signal transmitted to the left speaker and the signal transmitted to the middle speaker respectively using the following formula when the frequency of the audio object is greater than a preset first threshold:
[0277]
[0278]
[0279] Wherein, L is the mid-to-high frequency signal transmitted from the audio object to the left speaker, C is the mid-to-high frequency signal transmitted from the audio object to the middle speaker, and S... i For audio objects, A3 is the third coefficient. For the first direction vector, The first distance between the listening position and the left speaker. The second distance between the listening position and the right speaker. For the third-direction vector, A4 is the third distance between the listening position and the loudspeaker, P is the hyperparameter, j is the imaginary number, and f is the frequency.
[0280] The first signal calculation module includes:
[0281] The third coefficient function construction module is used to construct a third coefficient function using the following formula:
[0282]
[0283] A'3 2 +A'4 2 =1
[0284] The fourth coefficient function construction module is used to construct the fourth coefficient function using the following formula:
[0285]
[0286] A'3 2 +A'4 2 =1
[0287] in, For the second direction vector, The first distance between the listening position and the left speaker. For the third-direction vector, A1' is the third distance between the listening position and the right speaker, A2' is the fifth coefficient, and A'3 is the sixth coefficient. 2 The seventh coefficient, A'4 2 The eighth coefficient, θ i For the metadata;
[0288] The third coefficient function solving module is used to solve the third coefficient function to obtain the fifth coefficient and the sixth coefficient;
[0289] The fourth coefficient function solving module is used to solve the fourth coefficient function to obtain the seventh coefficient and the eighth coefficient;
[0290] The second signal calculation module is used to calculate the signal transmitted to the middle speaker and the signal transmitted to the right speaker respectively using the following formula when the frequency of the audio object is less than or equal to a preset threshold:
[0291]
[0292]
[0293] Wherein, C is the signal transmitted from the audio object to the middle speaker, R is the signal transmitted from the audio object to the right speaker, and S... i For audio objects, A1' is the fifth coefficient. For the first direction vector, The first distance between the listening position and the left speaker. The second distance between the listening position and the right speaker. For the third-direction vector, A2' is the third distance between the listening position and the loudspeaker, J is the hyperparameter, j is the imaginary number, and f is the frequency.
[0294] The third signal calculation module is used to calculate the mid- and high-frequency signals transmitted to the middle speaker and the mid- and high-frequency signals transmitted to the right speaker respectively, using the following formulas, when the frequency of the audio object is greater than a preset threshold:
[0295]
[0296]
[0297] Wherein, R is the signal transmitted from the audio object to the right speaker, C is the signal transmitted from the audio object to the middle speaker, and S... i For audio objects, A'3 2 The seventh coefficient For the first direction vector, The first distance between the listening position and the left speaker. The second distance between the listening position and the right speaker. For the third-direction vector, The third distance A'4 between the listening position and the loudspeaker 2 is the eighth coefficient, F is the hyperparameter, j is the imaginary number, and f is the frequency.
[0298] In one embodiment of the present invention, the second loudspeaker includes, from left to right, a left surround loudspeaker and a right surround loudspeaker, and the surround signal includes a left surround signal and a right surround signal; the surround sound signal compensation module 304 includes...
[0299] The distance measurement module is used to measure the fourth distance between the listening position and the left surround speaker, and the fifth distance between the listening position and the right surround speaker.
[0300] The first signal compensation module is used to compensate the left surround sound signal according to the fourth distance and the fifth distance, and transmit it to the left surround speaker for playback;
[0301] The second signal compensation module is used to compensate the right surround sound signal based on the fourth distance and the fifth distance, and transmit it to the right surround speaker for playback.
[0302] In one embodiment of the present invention,
[0303] The first signal compensation module includes:
[0304] The first compensation signal transmission module is used to compensate the surround sound signal using the following formula and then output it to the left surround speaker for playback:
[0305]
[0306] Among them, S' LS For the compensated surround sound signal, S LS Let r4 be the left surround signal, r5 be the fourth distance, r5 be the fifth distance, j be the imaginary part, f be the frequency, and G be the hyperparameter.
[0307] The second signal compensation module includes:
[0308] The second compensation signal transmission module is used to compensate the surround sound signal using the following formula and then output it to the right surround speaker for playback:
[0309]
[0310] Among them, S' RS For the compensated surround sound signal, S RS Let r4 be the right surround signal, r5 be the fourth distance, r5 be the fifth distance, j be the imaginary part, f be the frequency, and Q be the hyperparameter.
[0311] The apparatus provided in the embodiments of the present invention can execute the methods provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the methods.
[0312] Example 3
[0313] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0314] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0315] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0316] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as methods.
[0317] In some embodiments, the method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the method by any other suitable means (e.g., by means of firmware).
[0318] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0319] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0320] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0321] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0322] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0323] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0324] Example 4
[0325] This invention also provides a computer program product comprising a computer program that, when executed by a processor, implements the method provided in any embodiment of this invention.
[0326] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0327] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and no limitation is imposed herein.
[0328] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A sound reproduction method based on audio objects, characterized in that, include: Receives audio signals from multiple channels, including left channel signal, center channel signal, right channel signal, surround sound signal, and subwoofer signal; The left channel signal, the middle channel signal, and the right channel signal are decomposed into audio objects and positioning-related metadata; The audio object is compensated based on the arrangement of the first speaker located in front of the listening position and the metadata, and then transmitted to the first speaker for playback; The surround sound signal is compensated and transmitted to a second speaker located behind the listening position for playback; The subwoofer signal is transmitted to the third speaker for playback. The step of decomposing the left channel signal, the middle channel signal, and the right channel signal into audio objects and positioning-related metadata includes: Short-time Fourier transforms are performed on the left channel signal, the middle channel signal, and the right channel signal respectively to obtain the left channel spectrum signal, the middle channel spectrum signal, and the right channel spectrum signal; Take the first logarithm of the ratio between the absolute value of the left acoustic spectrum signal and the absolute value of the middle acoustic spectrum signal, and amplify the first logarithm to obtain a first reference value; Take the second logarithm of the ratio between the absolute value of the right acoustic spectrum signal and the absolute value of the middle acoustic spectrum signal, and amplify the second logarithm to obtain a second reference value; Determine the first reference range and the second reference range respectively; If the first reference value is within the first reference range, then the left acoustic spectrum signal and the middle acoustic spectrum signal are decomposed into audio objects and positioning-related metadata. If the second reference value is within the second reference range, then the right acoustic spectrum signal and the middle acoustic spectrum signal are decomposed into audio objects and positioning-related metadata in a first manner; If the first reference value is less than the first reference range or the second reference value is less than the second reference range, then the right acoustic spectrum signal and the middle acoustic spectrum signal are decomposed into audio objects and positioning-related metadata in a second manner.
2. The method according to claim 1, characterized in that, The determination of the first reference range and the second reference range includes: Generate multiple first histograms of the first group for the left channel signal and the middle channel signal; The first midpoint of the first x-axis of the first histogram in each first group is calculated using the formula. The first ordinate is calculated for the first histogram in each first group using the formula; A first objective function is constructed using a formula. When the value of the first objective function is maximized, multiple sets of the first histograms are sequentially divided into multiple first categories. The first range is generated using a formula; Generate multiple second groups of second histograms for the middle channel signal and the right channel signal; The second midpoint of the second x-axis is calculated for the second histogram in each second group using the formula. The first ordinate is calculated for the first histogram in each second group using the formula; A second objective function is constructed using a formula. When the value of the second objective function is maximized, multiple sets of the second histograms are sequentially divided into multiple second categories. The second range is generated using a formula; The process of decomposing the left acoustic spectrum signal and the middle acoustic spectrum signal into audio objects and location-related metadata includes: The left acoustic spectrum signal and the middle acoustic spectrum signal are decomposed into frequency domain signals of the audio object using the following formula: The left acoustic spectrum signal and the middle acoustic spectrum signal are decomposed to generate positioning-related metadata using the following formula: in, The signal is a left-side acoustic spectrum signal. The mid-sound spectrum signal. For audio objects, For gain, For metadata, j is an imaginary number. The phase of the left acoustic spectrum signal; The step of decomposing the right acoustic spectrum signal and the middle acoustic spectrum signal into audio objects and location-related metadata according to a first method includes: The middle acoustic spectrum signal and the right acoustic spectrum signal are decomposed into frequency domain signals of the audio object using the following formula: The positioning-related metadata is generated by decomposing the middle acoustic spectrum signal and the right acoustic spectrum signal using the following formula: in, This is a right-side acoustic spectrum signal. The mid-sound spectrum signal. For audio objects, For gain, For metadata, j is an imaginary number. The phase of the right acoustic spectrum signal; The step of decomposing the right acoustic spectrum signal and the middle acoustic spectrum signal into audio objects and location-related metadata according to the second method includes: The middle acoustic spectrum signal and the right acoustic spectrum signal are decomposed into frequency domain signals of the audio object using the following formula: in, This is a right-side acoustic spectrum signal. The mid-sound spectrum signal. For audio objects, j is an imaginary number. The phase of the right acoustic spectrum signal; Set the location-related metadata to 0°.
3. The method according to claim 1, characterized in that, The first speaker, from left to right, includes a left speaker, a center speaker, and a right speaker. The step of compensating the audio object based on the arrangement of the first speakers located in front of the listening position and the metadata, and then transmitting it to the first speaker for playback, includes: Measure the first direction vector from the listening position to the left speaker, the third direction vector to the middle speaker, and the second direction vector to the right speaker; Measure the first angle between the front of the listening position and the left speaker, the third angle between the listening position and the middle speaker, and the second angle between the listening position and the right speaker; If the metadata is greater than or equal to the first angle, then all the audio objects are transmitted to the left speaker for playback; If the metadata is less than or equal to the second angle, then all audio objects are transmitted to the right speaker for playback; If the metadata is greater than 0 and less than the first angle, or if the metadata is less than 0 and greater than the third angle, then the audio object is compensated according to the first direction vector and the third direction vector and transmitted to the left speaker and the middle speaker for playback; If the metadata is greater than the second included angle and less than the third included angle, then the audio object is compensated according to the second direction vector and the third direction vector and transmitted to the middle speaker and the right speaker for playback.
4. The method according to claim 3, characterized in that, The step of compensating the audio object based on the first direction vector and the third direction vector and transmitting it to the left speaker and the center speaker for playback includes: The first coefficient function is constructed using the following formula: The second coefficient function is constructed using the following formula: in, For the first direction vector, The first distance between the listening position and the left speaker. For the third-direction vector, The third distance between the listening position and the loudspeaker. For the metadata, For the first coefficient, As the second coefficient, The third coefficient, It is the fourth coefficient; Solving the first coefficient function yields the first coefficient and the second coefficient; Solving the second coefficient function yields the third and fourth coefficients; When the frequency of the audio object is less than or equal to a preset first threshold, the signal transmitted to the left speaker and the signal transmitted to the center speaker are calculated respectively using the following formulas: Wherein, L is the signal transmitted from the audio object to the left speaker, and C is the signal transmitted from the audio object to the center speaker. For the audio object, For the first coefficient, For the first direction vector, The first distance between the listening position and the left speaker. The second distance between the listening position and the right speaker. For the third-direction vector, The third distance between the listening position and the loudspeaker. is the second coefficient, q is the hyperparameter, j is the imaginary number, and f is the frequency; When the frequency of the audio object is greater than a preset first threshold, the signal transmitted to the left speaker and the signal transmitted to the center speaker are calculated using the following formulas: Wherein, L represents the mid-to-high frequency signal transmitted from the audio object to the left speaker, and C represents the mid-to-high frequency signal transmitted from the audio object to the middle speaker. For audio objects, For the third coefficient, For the first direction vector, The first distance between the listening position and the left speaker. The second distance between the listening position and the right speaker. For the third-direction vector, The third distance between the listening position and the loudspeaker. The fourth coefficient is P, the hyperparameter is j, the imaginary number is f, and the frequency is f. The step of compensating the audio object based on the second direction vector and the third direction vector and transmitting it to the center speaker and the right speaker for playback includes: The third coefficient function is constructed using the following formula: The fourth coefficient function is constructed using the following formula: in, For the second direction vector, The second distance between the listening position and the right speaker. For the third-direction vector, The third distance between the listening position and the loudspeaker. It is the fifth coefficient. It is the sixth coefficient. It is the seventh coefficient. It is the eighth coefficient. For the metadata; Solving the third coefficient function yields the fifth and sixth coefficients; Solving the fourth coefficient function yields the seventh and eighth coefficients; When the frequency of the audio object is less than or equal to a preset threshold, the signals transmitted to the middle speaker and the right speaker are calculated respectively using the following formulas: Wherein, C is the signal transmitted from the audio object to the middle speaker, and R is the signal transmitted from the audio object to the right speaker. For audio objects, For the fifth coefficient, For the first direction vector, The first distance between the listening position and the left speaker. The second distance between the listening position and the right speaker. For the third-direction vector, The third distance between the listening position and the loudspeaker. is the sixth coefficient, J is the hyperparameter, j is the imaginary number, and f is the frequency; When the frequency of the audio object is greater than a preset threshold, the mid-high frequency signals transmitted to the middle speaker and the mid-high frequency signals transmitted to the right speaker are calculated using the following formulas: Wherein, R is the signal transmitted from the audio object to the right speaker, and C is the signal transmitted from the audio object to the middle speaker. For audio objects, The seventh coefficient For the first direction vector, The first distance between the listening position and the left speaker. The second distance between the listening position and the right speaker. For the third-direction vector, The third distance between the listening position and the loudspeaker. is the eighth coefficient, F is the hyperparameter, j is the imaginary number, and f is the frequency.
5. The method according to any one of claims 1-4, characterized in that, The second speaker, from left to right, includes a left surround speaker and a right surround speaker, and the surround sound signal includes a left surround sound signal and a right surround sound signal; the step of compensating the surround sound signal and transmitting it to the second speaker located behind the listening position for playback includes: Measure the fourth distance between the listening position and the left surround speaker, and the fifth distance between the listening position and the right surround speaker; The left surround sound signal is compensated based on the fourth distance and the fifth distance, and then transmitted to the left surround speaker for playback; The right surround sound signal is compensated based on the fourth distance and the fifth distance, and then transmitted to the right surround speaker for playback.
6. The method according to claim 5, characterized in that, The step of compensating the left surround sound signal based on the fourth distance and the fifth distance, and transmitting it to the left surround speaker for playback, includes: The surround sound signal is compensated using the following formula and then output to the left surround speaker for playback: in, The compensated surround sound signal, The left surround sound signal, For the fourth distance, Let j be the fifth distance, f be the imaginary part, and G be the frequency; The step of compensating the right surround sound signal based on the fourth distance and the fifth distance, and transmitting it to the right surround speaker for playback, includes: The surround sound signal is compensated using the following formula and then output to the right surround speaker for playback: in, The compensated surround sound signal, The right surround sound signal, For the fourth distance, Let j be the fifth distance, j be the imaginary part, f be the frequency, and Q be the hyperparameter.
7. A sound reproduction device based on an audio object, characterized in that, include: An audio signal receiving module is used to receive audio signals from multiple channels, including left channel signal, middle channel signal, right channel signal, surround sound signal, and subwoofer signal. The signal decomposition module is used to decompose the left channel signal, the middle channel signal and the right channel signal into audio objects and positioning-related metadata. An audio object compensation module is used to compensate the audio object according to the arrangement of the first speaker located in front of the listening position and the metadata, and transmit it to the first speaker for playback; A surround sound signal compensation module is used to compensate the surround sound signal and transmit it to a second speaker located behind the listening position for playback; The subwoofer transmission and playback module is used to transmit the subwoofer signal to the third speaker for playback; The signal decomposition module includes a signal transformation module, which is specifically used to perform short-time Fourier transform on the left channel signal, the middle channel signal and the right channel signal respectively to obtain the left channel spectrum signal, the middle channel spectrum signal and the right channel spectrum signal; The first logarithmic amplification module is used to take the first logarithm of the ratio between the absolute value of the left acoustic spectrum signal and the absolute value of the middle acoustic spectrum signal, and amplify the first logarithm to obtain a first reference value. The second logarithmic amplification module is used to take the second logarithm of the ratio between the absolute value of the right acoustic spectrum signal and the absolute value of the middle acoustic spectrum signal, and amplify the second logarithm to obtain a second reference value; The range determination module is used to determine the first reference range and the second reference range, respectively. The signal decomposition module is used to decompose the left acoustic spectrum signal and the middle acoustic spectrum signal into audio objects and positioning-related metadata if the first reference value is within the first reference range.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the audio object-based sound reproduction method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the audio object-based sound reproduction method according to any one of claims 1-6.
Citation Information
Patent Citations
Localization of sound in a speaker system
CN112005558A