An audio processing method, apparatus, device, storage medium, and product
By acquiring the positioning data of the audio acquisition device, determining its positioning jitter data, and performing compensation processing, the problem of jitter noise was solved, and the accuracy and quality of sound source positioning of the audio data were improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2024-12-13
- Publication Date
- 2026-06-16
AI Technical Summary
Audio acquisition equipment is susceptible to jitter and noise during movement, which can lead to unstable sound source localization and reduced audio quality.
By acquiring the positioning data of the audio acquisition device, its positioning jitter data is determined, and the audio data is compensated based on this data to obtain jitter-compensated audio data.
It improves the accuracy of sound source localization in audio data and ensures the quality of audio data.
Smart Images

Figure CN122227150A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio processing technology, and in particular to an audio processing method, apparatus, device, storage medium and product. Background Technology
[0002] Because of their portability, audio acquisition devices are susceptible to jitter noise due to unstable factors during movement. The presence of jitter noise makes the sound source localization of the audio data unstable, resulting in poor audio quality. Summary of the Invention
[0003] This invention provides an audio processing method, apparatus, device, storage medium, and product to solve the problem of audio data being affected by jitter noise, improve the accuracy of sound source localization of audio data, and thus ensure the audio quality of audio data.
[0004] In a first aspect, embodiments of the present invention provide an audio processing method, the method comprising:
[0005] Acquire first audio data collected by the audio acquisition device and first positioning data of the audio acquisition device; wherein, the first positioning data represents a set of location information of the audio acquisition device during the process of collecting the first audio data, and the first audio data and the first positioning data have the same acquisition timestamp;
[0006] Based on the first positioning data, the positioning jitter data of the audio acquisition device is determined; wherein, the positioning jitter data represents a set of position fluctuation information of the audio acquisition device during the process of acquiring the first audio data;
[0007] Based on the positioning jitter data, the first audio data is compensated to obtain jitter-compensated second audio data.
[0008] Secondly, embodiments of the present invention also provide an audio processing apparatus, the apparatus comprising:
[0009] The first positioning data acquisition module is used to acquire first audio data collected by the audio acquisition device and first positioning data of the audio acquisition device; wherein, the first positioning data represents a set of location information of the audio acquisition device during the process of acquiring the first audio data, and the acquisition timestamp of the first audio data and the first positioning data are the same.
[0010] The positioning jitter data determination module is used to determine the positioning jitter data of the audio acquisition device based on the first positioning data; wherein, the positioning jitter data represents a set of position fluctuation information of the audio acquisition device during the process of acquiring the first audio data;
[0011] The second audio data determination module is used to perform compensation processing on the first audio data based on the positioning jitter data to obtain jitter-compensated second audio data.
[0012] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising:
[0013] At least one processor; and
[0014] A memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the audio processing method according to any embodiment of the present invention.
[0016] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing computer instructions that are used to cause a processor to execute the audio processing method described in any embodiment of the present invention.
[0017] Fifthly, embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the audio processing method described in any embodiment of the present invention.
[0018] The technical solution of this invention determines the positioning jitter data of the audio acquisition device based on the first positioning data of the audio acquisition device, and performs compensation processing on the first audio data acquired by the audio acquisition device based on the positioning jitter data to obtain jitter-compensated second audio data. The first audio data and the first positioning data have the same acquisition timestamp, which solves the problem of audio data being affected by jitter noise interference, improves the accuracy of sound source positioning of audio data, and thus ensures the audio quality of audio data. Attached Figure Description
[0019] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0020] Figure 1 A flowchart illustrating an audio processing method provided in one embodiment of the present invention;
[0021] Figure 2 A flowchart illustrating a specific example of an audio processing method provided in an embodiment of the present invention;
[0022] Figure 3 A flowchart illustrating another audio processing method provided in one embodiment of the present invention;
[0023] Figure 4 This is a schematic diagram of positioning adjustment data for a surround sound signal provided in one embodiment of the present invention;
[0024] Figure 5 A flowchart illustrating a specific example of another audio processing method provided in an embodiment of the present invention;
[0025] Figure 6 This is a schematic diagram of the structure of an audio processing device provided in one embodiment of the present invention;
[0026] Figure 7 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present invention. Detailed Implementation
[0027] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the invention. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the invention.
[0028] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0029] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0030] It should be noted that the concepts of "first" and "second" mentioned in this invention are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0031] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0032] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0033] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0034] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application program, server, or storage medium executing the operation of this invention, based on the prompt message.
[0035] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0036] It is understood that the above notification and user authorization process is merely illustrative and does not constitute a limitation on the implementation of the present invention. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present invention.
[0037] Figure 1 This is a flowchart illustrating an audio processing method according to an embodiment of the present invention. This embodiment is applicable to processing audio data acquired by an audio acquisition device. The method can be executed by an audio processing device, which can be implemented in hardware and / or software. This audio processing device can be configured in an electronic device, typically a mobile terminal or tablet computer. Figure 1 As shown, the method includes:
[0038] S110. Obtain the first audio data collected by the audio acquisition device and the first positioning data of the audio acquisition device.
[0039] Specifically, an audio acquisition device refers to a tool used to acquire sound information. For example, an audio acquisition device can be a microphone or a voice recorder, but is not limited to the example scenario.
[0040] Specifically, the first audio data refers to signal data used for digitally recording sound information. For example, the audio signal types of the first audio data include, but are not limited to, WAV audio signals, MP3 audio signals, FLAC audio signals, AAC audio signals, mono signals, stereo signals, Ambisonic signals, or surround sound signals, etc.
[0041] In this embodiment, the first positioning data represents a set of location information of the audio acquisition device during the acquisition of the first audio data, and the acquisition timestamps of the first audio data and the first positioning data are the same. Specifically, the first audio data includes a first audio signal corresponding to at least one acquisition time, and the first positioning data includes a first positioning location corresponding to at least one acquisition time.
[0042] In one optional embodiment, the first positioning data is acquired by an inertial sensor mounted on the audio acquisition device. An inertial sensor is a device used to measure changes in the motion state of an object; exemplary devices include, but are not limited to, accelerometers, gyroscopes, or inertial measurement units (IMUs).
[0043] S120. Based on the first positioning data, determine the positioning jitter data of the audio acquisition device.
[0044] In this embodiment, the positioning jitter data represents a set of fluctuation information about the position of the audio acquisition device during the acquisition of the first audio data. Specifically, the positioning jitter data includes jitter data corresponding to at least one acquisition moment.
[0045] Specifically, the low-frequency data in the first positioning data represents the motion position data of the audio acquisition device, while the high-frequency data represents the positioning jitter data of the audio acquisition device. The motion position data represents directional and regular changes in the first positioning data, generated by the purposeful movement of the audio acquisition device. The positioning jitter data represents small-amplitude, high-frequency, and irregular changes in the first positioning data.
[0046] For example, positioning jitter data can be caused by at least one of the following: external environmental factors, installation factors, or user factors. External environmental factors could include wind, vibrations generated by the operation of mechanical equipment placed around the audio acquisition device, etc. Installation factors could include the structural stability of the fixing components used to secure the audio acquisition device and its mounting device, the structural stability of the mounting device, etc. The causes of positioning jitter data are illustrated here as examples and are not intended to limit the scope of the problem.
[0047] In an optional embodiment, determining the positioning jitter data of the audio acquisition device based on the first positioning data includes: performing wavelet transform processing on the first positioning data to obtain high-frequency coefficient data, and filtering the first positioning data based on the high-frequency coefficient threshold and the high-frequency coefficient data to obtain the positioning jitter data of the audio acquisition device.
[0048] Specifically, the high-frequency coefficient data includes high-frequency detail coefficients corresponding to each location data point in the first location data. For example, the high-frequency coefficient threshold can be the standard deviation of the first location data.
[0049] In another optional embodiment, determining the positioning jitter data of the audio acquisition device based on the first positioning data includes: performing low-pass filtering on the first positioning data to obtain filtered positioning data, and using the difference between the first positioning data and the filtered positioning data as the positioning jitter data of the audio acquisition device.
[0050] Specifically, the filtered positioning data represents a set of motion position information of the audio acquisition device during the acquisition of the first audio data. For example, the low-pass filtering process may use a Butterworth filter, a Chebyshev filter, or an elliptic filter, but is not limited to the example scenario.
[0051] In another optional embodiment, determining the positioning jitter data of the audio acquisition device based on the first positioning data includes: performing high-pass filtering on the first positioning data to obtain the positioning jitter data of the audio acquisition device.
[0052] For example, the filters used in high-pass filtering can be discrete-time high-pass filters, finite impulse response high-pass filters, Butterworth high-pass filters, weighted moving average high-pass filters, or first-order difference filters, etc., but are not limited to the example cases.
[0053] It is understood that the above is only an exemplary description of the method for determining positioning jitter data and is not intended to limit it. In other embodiments, any other implementation method that can extract positioning jitter data from the first positioning data may be adopted.
[0054] Specifically, positioning jitter data can be represented as jitter coordinates, for example, (Δx, Δy, Δz), where Δx, Δy, and Δz represent the jitter displacement of the audio acquisition device on the X, Y, and Z axes, respectively. Positioning jitter data can also be represented as jitter angles, for example, (Δpitch, Δroll, Δyaw), where Δpitch, Δroll, and Δyaw represent the jitter angles of the audio acquisition device on the X, Y, and Z axes, respectively. Positioning jitter data can also be represented as quaternions, for example, (w, x, y, z), where w represents a scalar defining the jitter angle, and (x, y, z) represents a vector defining the jitter direction.
[0055] S130. Based on the positioning jitter data, the first audio data is compensated to obtain jitter-compensated second audio data.
[0056] In an optional embodiment, when the compensation processing scenario is an audio playback scenario, the second audio data includes second audio signals corresponding to at least one playback moment. Accordingly, the step of compensating the first audio data based on the positioning jitter data to obtain jitter-compensated second audio data includes: for each playback moment, acquiring second positioning data of the audio playback device at that playback moment, and acquiring a first audio signal in the first audio data matching the playback moment and target jitter data in the positioning jitter data matching the playback moment; compensating the second positioning data based on the target jitter data to obtain third positioning data; decoding the first audio signal based on the third positioning data to obtain a second audio signal corresponding to the playback moment; and determining the jitter-compensated second audio data based on the second audio signals corresponding to at least one playback moment.
[0057] In this embodiment, the second positioning data represents the actual position information of the audio playback device at the playback time, and the third positioning data represents the position information of the audio playback device after jitter compensation at the playback time.
[0058] Specifically, an audio playback device refers to a tool used to convert audio data into sound that can be perceived by the human ear. For example, an audio playback device can be a speaker or headphones, but is not limited to the example scenario.
[0059] Specifically, in order to ensure the complete reproduction of the first audio data on the time scale, the playback time series of the audio playback device has the same time length as the acquisition time series of the first audio data, and the playback moment refers to any moment in the playback time series.
[0060] Specifically, when the time boundaries of the playback time series and the acquisition time series are the same, such as both being 0:00-00:30, the first audio signal represents the audio signal at the playback time in the first audio data, and the target jitter data represents the jitter data at the playback time in the positioning jitter data. For example, when the playback time is 0:29, the first audio signal is the audio signal at 0:29 in the first audio data, and the target jitter data is the jitter data at 0:29 in the positioning jitter data. When the time boundaries of the playback time series and the acquisition time series are different, such as the playback time series being 00:00-00:30 and the acquisition time series being 8:00-8:30, the first audio signal represents the audio signal at the acquisition time in the first audio data that is aligned with the playback time, and the target jitter data represents the jitter data at the acquisition time in the positioning jitter data that is aligned with the playback time. Taking the above example again, when the playback time is 0:29, the first audio signal is the audio signal at 8:29 in the first audio data, and the target jitter data is the jitter data at 8:29 in the positioning jitter data.
[0061] In an optional embodiment, the step of compensating the second positioning data based on the target jitter data to obtain the third positioning data includes: inverting the target jitter data to obtain jitter compensation data, and determining a compensation rotation parameter based on the jitter compensation data; and using the product of the second positioning data and the compensation rotation parameter as the third positioning data.
[0062] Specifically, the compensation rotation parameter is a quaternion or compensation rotation matrix used to represent jitter compensation. The quaternion representing jitter compensation can be expressed as (w, -x, -y, -z), and the compensation rotation matrix includes rotation matrices R corresponding to the X, Y, and Z coordinate axes, respectively. pitch R roll and R yaw .
[0063] For example, when the compensation rotation parameter is a quaternion used to represent jitter compensation, the third positioning data Q ′ It can be represented as:
[0064] Q ′ = (w, -x, -y, -z)Q
[0065] Where Q represents the second positioning data.
[0066] Specifically, the audio signal type of the second audio data can be the same as or different from that of the first audio data. For example, the decoding process may employ algorithms including, but not limited to, head-related impulse response (HIR) algorithms, head-related transfer function (HRT) algorithms, or wavelength synthesis algorithms, which can be customized according to the audio signal type of the second audio data.
[0067] Figure 2 This is a flowchart illustrating a specific example of an audio processing method provided in an embodiment of the present invention. Specifically, a microphone collects sound information and encodes the collected sound information to obtain first audio data. During the process of the microphone collecting sound information, a gyroscope mounted on the microphone synchronously collects the microphone's first positioning data. The first positioning data is then high-pass filtered to obtain the microphone's positioning jitter data. For each playback moment, a gyroscope mounted on the speaker collects the speaker's second positioning data at the playback moment. Based on the target jitter data in the positioning jitter data that matches the playback moment, jitter compensation is performed on the second positioning data to obtain third positioning data. The first audio signal in the first audio data that matches the playback moment is then decoded based on the third positioning data to obtain a second audio signal corresponding to the playback moment. Based on the second audio signals corresponding to at least one playback moment, jitter-compensated second audio data is determined. In this embodiment, the second audio data can be directly input into an audio playback device for sound playback.
[0068] The advantage of jitter compensation for the second positioning data is that the first audio data can be stored in its original form, eliminating the computational burden of jitter compensation for the first audio data during the audio recording stage, reducing the computational resources occupied by the audio processing flow in the audio playback scenario, and reducing the maintenance cost of the audio signal.
[0069] The technical solution of this embodiment determines the positioning jitter data of the audio acquisition device based on the first positioning data of the audio acquisition device, and performs compensation processing on the first audio data acquired by the audio acquisition device based on the positioning jitter data to obtain jitter-compensated second audio data. The first audio data and the first positioning data have the same acquisition timestamp, which solves the problem of audio data being affected by jitter noise interference, improves the accuracy of sound source positioning of audio data, and thus ensures the audio quality of audio data.
[0070] Figure 3This is a flowchart of another audio processing method provided by an embodiment of the present invention. This embodiment further refines the step of "compensating the first audio data according to the positioning jitter data to obtain jitter-compensated second audio data" in the above embodiment. In this embodiment, when the processing scenario of the compensation processing is an audio recording scenario, the second audio data includes second audio signals corresponding to at least one recording time. The step of compensating the first audio data according to the positioning jitter data to obtain jitter-compensated second audio data includes: for each recording time, acquiring the first audio signal of the recording time in the first audio data and the target jitter data of the recording time in the positioning jitter data; compensating the first audio signal according to the target jitter data to obtain the second audio signal corresponding to the recording time; and determining the jitter-compensated second audio data according to the second audio signals corresponding to at least one recording time. Figure 3 As shown, the method includes:
[0071] S210. Obtain the first audio data collected by the audio acquisition device and the first positioning data of the audio acquisition device.
[0072] S220. Based on the first positioning data, determine the positioning jitter data of the audio acquisition device.
[0073] S210-S220 in this embodiment are the same as those in the above embodiments. Figure 1 The S110-S120 shown are the same or similar, and will not be described again in this embodiment.
[0074] S230. For each recording moment, acquire the first audio signal of the recording moment in the first audio data and the target jitter data of the recording moment in the positioning jitter data.
[0075] Specifically, the recording time refers to any moment in the acquisition time sequence of the first audio data. For example, assuming the acquisition time sequence of the first audio data is 0:00-00:30 and the recording time is 0:29, then the first audio signal is the audio signal at 0:29 in the first audio data, and the target jitter data is the jitter data at 0:29 in the positioning jitter data.
[0076] S240. The first audio signal is compensated according to the target jitter data to obtain a second audio signal corresponding to the recording time.
[0077] In an optional embodiment, the step of compensating the first audio signal based on the target jitter data to obtain a second audio signal corresponding to the recording time includes: when the first audio signal is an audio signal based on a spherical harmonic function, inverting the target jitter data to obtain jitter compensation data; determining a compensation rotation matrix based on the jitter compensation data and the spherical harmonic order of the first audio signal; and using the product of the first audio signal and the compensation rotation matrix as the second audio signal corresponding to the recording time.
[0078] Specifically, spherical harmonics represent an orthogonal function system defined on a sphere, used to decompose the distribution of sound waves in three-dimensional space. Spherical harmonic coefficients express the components of sound at different frequencies and directions. The first audio signal is composed of spherical harmonic coefficients corresponding to different frequency components and directional components.
[0079] Specifically, the compensation rotation matrix includes rotation matrices R corresponding to the X, Y, and Z coordinate axes, respectively. pitch R roll and R yaw When the spherical harmonic order is n, the compensation rotation matrix consists of three M×M matrices, where M = (n+1) 2 For example, when the spherical harmonic order is first, the compensation rotation matrix includes three 4×4 rotation matrices; when the spherical harmonic order is second, the compensation rotation matrix includes three 9×9 rotation matrices.
[0080] For example, the second audio signal S out It can be represented as: S out =R pitch R roll R yaw S in , of which S in This represents the first audio signal. It is understood that the order of the products of the three rotation matrices in the compensation rotation matrix and the first audio signal is merely illustrative and not intended to be limiting.
[0081] In another optional embodiment, the step of compensating the first audio data based on the target jitter data to obtain a second audio signal corresponding to the recording time includes: when the first audio signal is a surround sound signal, acquiring speaker layout data of the audio acquisition device; determining the jitter gain weight of each speaker based on the speaker layout data and the target jitter data; and compensating the channel signal corresponding to each speaker in the first audio signal according to the jitter gain weight of each speaker to obtain a second audio signal corresponding to the recording time; wherein the jitter gain weight corresponds one-to-one with the channel signal.
[0082] Specifically, the speaker layout data represents information such as the number, location, and angle of the speakers corresponding to the surround sound signal. For example, when the surround sound signal is 5.1 surround sound, the speaker layout data includes the location information of the speakers corresponding to the front left channel, front center channel, front right channel, rear left surround channel, and rear right surround channel. When the surround sound signal is 7.1 surround sound, the speaker layout data includes the location information of the speakers corresponding to the front left channel, front center channel, front right channel, left surround channel, right surround channel, rear left surround channel, and rear right surround channel.
[0083] In an optional embodiment, determining the jitter gain weight of each speaker based on the speaker layout data and the target jitter data includes: determining positioning adjustment data corresponding to each speaker based on the speaker layout data and the target jitter data; for each positioning adjustment data, determining at least two target speakers matching the positioning adjustment data based on the speaker layout data and the positioning adjustment data, and determining component gain weights corresponding to the at least two target speakers based on the positioning adjustment data and the positioning data corresponding to the at least two target speakers respectively in the speaker layout data; wherein, the component gain weight represents the component coefficient of the positioning adjustment data on the target speaker; and for each speaker, the summation result of at least one component gain weight corresponding to the speaker is used as the jitter gain weight of the speaker.
[0084] Specifically, based on the speaker layout data and the target jitter data, determining the positioning adjustment data corresponding to each speaker includes: rotating the speaker layout data according to the target jitter data to obtain the positioning adjustment data corresponding to each speaker; or, inverting the target jitter data to obtain jitter compensation data, and rotating the speaker layout data according to the jitter compensation data to obtain the positioning adjustment data corresponding to each speaker. Specifically, the positioning adjustment data represents the speaker jitter offset data or jitter compensation data in the first audio signal.
[0085] Figure 4 This is a schematic diagram of surround sound signal positioning adjustment data provided in one embodiment of the present invention. Figure 4 Taking 5.1 surround sound as an example. Specifically, Figure 4 The left image shows the speaker positions for the front left channel (L), front center channel (C), front right channel (R), rear left surround channel (Ls), and rear right surround channel (Rs) in the 5.1 surround sound speaker layout data. Figure 4In the middle diagram, the dashed circles represent the jitter offset positions of the five speakers after rotation based on the target jitter data, and α represents the target jitter angle. Figure 4 The dashed circles in the right-hand diagram represent the jitter compensation positions of the five speakers after rotation based on jitter compensation data, and -α represents the jitter compensation angle.
[0086] In an alternative embodiment, determining at least two target loudspeakers that match the positioning adjustment data based on the loudspeaker layout data and the positioning adjustment data includes: for each offset direction of the surround sound plane, determining the target loudspeaker that is closest to the positioning adjustment data of the loudspeaker in the offset direction based on the loudspeaker layout data.
[0087] For example, when the surround sound signal is 5.1 surround sound, the offset direction of the surround sound plane includes a clockwise offset direction and a counterclockwise offset direction. Figure 4 For example, in the clockwise offset direction, the speaker closest to the jitter offset data corresponding to the front center channel is the speaker corresponding to the front center channel (C). In the counterclockwise offset direction, the speaker closest to the jitter offset data corresponding to the front center channel is the speaker corresponding to the front left channel (L). In the clockwise offset direction, the speaker closest to the jitter compensation data corresponding to the front center channel is the speaker corresponding to the front right channel (R). In the counterclockwise offset direction, the speaker closest to the jitter offset data corresponding to the front center channel is the speaker corresponding to the front center channel (C).
[0088] In this embodiment, the component gain weight represents the component coefficient of the positioning adjustment data on the target speaker. Specifically, based on the positioning data corresponding to the at least two target speakers in the speaker layout data, the positioning adjustment data is vector decomposed, and the component coefficients on the target speaker obtained from the vector decomposition are used as the component gain weight of the target speaker.
[0089] For example, suppose there are two target speakers, in order to Figure 4 Taking the jitter offset data corresponding to the center channel of the front speaker as an example, the jitter offset data (x c ,y c ) can be represented as: x c =g1x1+g2x2; y c = g1y1 + g2y2, where (x1, y1) represents the positioning data of the speaker corresponding to the front center channel (C), (x2, y2) represents the positioning data of the speaker corresponding to the front left channel (L), and g1 represents the jitter offset data (x1, y1) + g2y2. c ,y c The component gain weights on the speaker corresponding to the front center channel (C), g2 represents the jitter offset data (x).c ,y c The component gain weights on the speaker corresponding to the front left channel (L). By solving the above equation, we can obtain [g1, g2].
[0090] Specifically, the jitter gain weight of the speaker corresponding to the front left channel (L) is the sum of the component gain weight of the jitter offset data corresponding to the front center channel on the speaker corresponding to the front left channel (L) and the component gain weight of the jitter offset data corresponding to the front right channel (R) on the speaker corresponding to the front left channel (L).
[0091] In one optional embodiment, the channel signal corresponding to each speaker in the first audio signal is corrected according to the jitter gain weight of each speaker to obtain a second audio signal corresponding to the recording time, including: when the positioning adjustment data represents the jitter offset data of the speaker, for each speaker, the product of the jitter gain weight of the speaker and the channel signal of the speaker is used as the jitter gain signal, and the difference between the channel signal and the jitter gain signal is used as the compensated channel signal corresponding to the speaker; and the second audio signal is determined according to the corrected channel signal corresponding to each speaker.
[0092] In another optional embodiment, the channel signal corresponding to each speaker in the first audio signal is corrected according to the jitter gain weight of each speaker to obtain a second audio signal corresponding to the recording time, including: when the positioning adjustment data represents the jitter compensation data of the speaker, for each speaker, the product of the jitter gain weight of the speaker and the channel signal of the speaker is used as the compensation gain signal, and the sum of the channel signal and the compensation gain signal is used as the compensated channel signal corresponding to the speaker; and the second audio signal is determined according to the corrected channel signal corresponding to each speaker.
[0093] In an optional embodiment, the step of compensating the first audio signal based on the target jitter data to obtain a second audio signal corresponding to the recording time includes: when the first audio signal is an object audio signal, compensating the first position data of the sound object in the first audio signal based on the target jitter data to obtain a second audio signal corresponding to the recording time.
[0094] Specifically, the object audio signal consists of the location data and audio information of the sound object, wherein the number of sound objects can be one or more.
[0095] For example, the target jitter data is (Δx, Δy, Δz), and the first position data of the i-th sound object in the first audio signal is represented as (x...). i y i, z i If the jitter compensation of the i-th sound object in the second audio signal is given by (x), then the second position data is represented as (x). i ′ y i ′ , z i ′ ), where x i ′ =x i -Δx, y i ′ =y i -Δy, z i ′ =z i -Δz.
[0096] S250. Determine the second audio data for jitter compensation based on the second audio signal corresponding to at least one recording time.
[0097] Figure 5 This is a flowchart illustrating a specific example of another audio processing method provided in an embodiment of the present invention. Specifically, a microphone collects sound information and encodes the collected sound information to obtain first audio data. During the process of the microphone collecting sound information, a gyroscope mounted on the microphone synchronously collects the microphone's first positioning data. The first positioning data is then high-pass filtered to obtain positioning jitter data. Based on the positioning jitter data, the first audio data is compensated to obtain jitter-compensated second audio data.
[0098] The technical solution of this embodiment compensates the first audio signal of the recording time in the first audio data according to the target jitter data of the recording time in the positioning jitter data for each recording time, thereby obtaining a second audio signal corresponding to the recording time. Based on the second audio signals corresponding to at least one recording time, the jitter-compensated second audio data is determined. This reduces the computational load of the post-processing stage after audio recording, thereby ensuring the processing efficiency of the audio post-processing stage. Especially in multi-device scenarios of audio playback, it achieves the consistency of audio signals used by multiple devices, thereby helping to maintain the stability of audio quality and spatial effects.
[0099] The following are embodiments of the audio processing apparatus provided in this invention. This apparatus and the audio processing method described in the above embodiments belong to the same inventive concept. For details not described in detail in the embodiments of the audio processing apparatus, please refer to the content about the audio processing method in the above embodiments.
[0100] Figure 6 This is a schematic diagram of the structure of an audio processing device according to an embodiment of the present invention. Figure 6As shown, the device includes: a first positioning data acquisition module 310, a positioning jitter data determination module 320, and a second audio data determination module 330.
[0101] The first positioning data acquisition module 310 is used to acquire first audio data and first positioning data of the audio acquisition device; wherein the first positioning data represents a set of location information of the audio acquisition device during the acquisition of the first audio data, and the acquisition timestamp of the first audio data and the first positioning data are the same.
[0102] The positioning jitter data determination module 320 is used to determine the positioning jitter data of the audio acquisition device based on the first positioning data; wherein, the positioning jitter data represents a set of position fluctuation information of the audio acquisition device during the process of acquiring the first audio data;
[0103] The second audio data determination module 330 is used to perform compensation processing on the first audio data based on the positioning jitter data to obtain jitter-compensated second audio data.
[0104] The technical solution of this embodiment determines the positioning jitter data of the audio acquisition device based on the first positioning data of the audio acquisition device, and performs compensation processing on the first audio data acquired by the audio acquisition device based on the positioning jitter data to obtain jitter-compensated second audio data. The first audio data and the first positioning data have the same acquisition timestamp, which solves the problem of audio data being affected by jitter noise interference, improves the accuracy of sound source positioning of audio data, and thus ensures the audio quality of audio data.
[0105] In an optional embodiment, when the compensation processing scenario is an audio playback scenario, the second audio data includes second audio signals corresponding to at least one playback moment. Correspondingly, the second audio data determination module 330 includes:
[0106] The second positioning data acquisition unit is used to acquire, for each playback moment, the second positioning data of the audio playback device at the playback moment, and to acquire the first audio signal in the first audio data that matches the playback moment and the target jitter data in the positioning jitter data that matches the playback moment; wherein, the second positioning data represents the real position information of the audio playback device at the playback moment;
[0107] The third positioning data determination unit is used to perform compensation processing on the second positioning data based on the target jitter data to obtain third positioning data; wherein, the third positioning data represents the jitter-compensated position information of the audio playback device at the playback time;
[0108] The first audio signal decoding unit is used to decode the first audio signal according to the third positioning data to obtain a second audio signal corresponding to the playback time.
[0109] The second audio data determining unit is used to determine the second audio data for jitter compensation based on the second audio signals corresponding to at least one playback time.
[0110] In one optional embodiment, the third positioning data determining unit is specifically used for:
[0111] The target jitter data is inverted to obtain jitter compensation data, and the compensation rotation parameters are determined based on the jitter compensation data.
[0112] The product of the second positioning data and the compensation rotation parameter is used as the third positioning data.
[0113] In an optional embodiment, when the compensation processing scenario is an audio recording scenario, the second audio data includes second audio signals corresponding to at least one recording moment. Accordingly, the second audio data determination module 330 includes:
[0114] The target jitter data acquisition unit is used to acquire, for each recording moment, the first audio signal of the recording moment in the first audio data and the target jitter data of the recording moment in the positioning jitter data;
[0115] The second audio signal determination unit is used to compensate the first audio signal according to the target jitter data to obtain a second audio signal corresponding to the recording time.
[0116] The second audio data determining unit is used to determine the second audio data for jitter compensation based on the second audio signals corresponding to at least one recording time.
[0117] In one optional embodiment, the second audio signal determining unit includes:
[0118] The second audio signal determination subunit is used to invert the target jitter data to obtain jitter compensation data when the first audio signal is an audio signal based on a spherical harmonic function.
[0119] The compensation rotation matrix is determined based on the jitter compensation data and the spherical harmonic order of the first audio signal;
[0120] The product of the first audio signal and the compensation rotation matrix is used as the second audio signal corresponding to the recording time.
[0121] In one optional embodiment, the second audio signal determining unit includes:
[0122] The speaker layout data acquisition subunit is used to acquire speaker layout data of the audio acquisition device when the first audio signal is a surround sound signal.
[0123] The jitter gain weight determination subunit is used to determine the jitter gain weight of each speaker based on the speaker layout data and the target jitter data.
[0124] The second audio signal determination subunit is used to compensate the channel signal corresponding to each speaker in the first audio signal according to the jitter gain weight of each speaker, so as to obtain the second audio signal corresponding to the recording time; wherein the jitter gain weight corresponds one-to-one with the channel signal.
[0125] In one optional embodiment, the jitter gain weighting determination subunit is specifically used for:
[0126] Based on the speaker layout data and the target jitter data, determine the positioning adjustment data corresponding to each speaker;
[0127] For each positioning adjustment data, at least two target speakers matching the positioning adjustment data are determined based on the speaker layout data and the positioning adjustment data. Then, based on the positioning adjustment data and the positioning data corresponding to the at least two target speakers in the speaker layout data, component gain weights corresponding to the at least two target speakers are determined. The component gain weights represent the component coefficients of the positioning adjustment data on the target speakers.
[0128] For each speaker, the summation of at least one component gain weight corresponding to the speaker is used as the jitter gain weight of the speaker.
[0129] In one optional embodiment, the second audio signal determining unit includes:
[0130] The second audio signal determination subunit is used to compensate the first position data of the sound object in the first audio signal according to the target jitter data when the first audio signal is the target audio signal, so as to obtain the second audio signal corresponding to the recording time.
[0131] The audio processing apparatus provided in the embodiments of the present invention can execute the audio processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0132] The following is for reference. Figure 7The diagram illustrates a structural schematic of an electronic device (e.g., a terminal device or a server) 400 suitable for implementing embodiments of the present invention. The terminal device in the embodiments of the present invention may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0133] like Figure 7 As shown, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 406 into random access memory (RAM) 403. RAM 403 also stores various programs and data required for the operation of electronic device 400. Processing device 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.
[0134] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 406 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0135] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 409, or installed from a storage device 406, or installed from a ROM 402. When the computer program is executed by the processing device 401, it performs the functions defined in the methods of the embodiments of the present invention.
[0136] It should be noted that the computer-readable medium described above in this invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0137] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0138] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0139] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire at least two Internet Protocol (IP) addresses; send a node evaluation request including the at least two IP addresses to a node evaluation device, wherein the node evaluation device selects an IP address from the at least two IP addresses and returns it; and receive the IP address returned by the node evaluation device; wherein the acquired IP address indicates an edge node in a content delivery network.
[0140] Alternatively, the aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: receive a node evaluation request including at least two Internet Protocol (IP) addresses; select an IP address from the at least two IP addresses; and return the selected IP address; wherein the received IP address indicates an edge node in the content delivery network.
[0141] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0143] The units described in the embodiments of the present invention can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0144] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0145] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0146] According to one or more embodiments of the present invention, Example 1 provides an audio processing method, comprising:
[0147] Acquire first audio data collected by the audio acquisition device and first positioning data of the audio acquisition device; wherein, the first positioning data represents a set of location information of the audio acquisition device during the process of collecting the first audio data, and the first audio data and the first positioning data have the same acquisition timestamp;
[0148] Based on the first positioning data, the positioning jitter data of the audio acquisition device is determined; wherein, the positioning jitter data represents a set of position fluctuation information of the audio acquisition device during the process of acquiring the first audio data;
[0149] Based on the positioning jitter data, the first audio data is compensated to obtain jitter-compensated second audio data.
[0150] According to one or more embodiments of the present invention, Example 2, based on the method described in Example 1, when the processing scenario of the compensation processing is an audio playback scenario, the second audio data includes second audio signals corresponding to at least one playback moment. Correspondingly, the step of compensating the first audio data based on the positioning jitter data to obtain jitter-compensated second audio data includes:
[0151] For each playback moment, the second positioning data of the audio playback device at the playback moment is obtained, and the first audio signal in the first audio data that matches the playback moment and the target jitter data in the positioning jitter data that matches the playback moment are obtained; wherein, the second positioning data represents the real position information of the audio playback device at the playback moment;
[0152] The second positioning data is compensated based on the target jitter data to obtain the third positioning data; wherein, the third positioning data represents the jitter-compensated position information of the audio playback device at the playback time.
[0153] The first audio signal is decoded based on the third positioning data to obtain a second audio signal corresponding to the playback time.
[0154] The second audio data for jitter compensation is determined based on the second audio signal corresponding to at least one playback moment.
[0155] According to one or more embodiments of the present invention, Example 3, based on the method described in Example 2, wherein the step of compensating the second positioning data based on the target jitter data to obtain the third positioning data includes:
[0156] The target jitter data is inverted to obtain jitter compensation data, and the compensation rotation parameters are determined based on the jitter compensation data.
[0157] The product of the second positioning data and the compensation rotation parameter is used as the third positioning data.
[0158] According to one or more embodiments of the present invention, Example 4, based on the method described in Example 1, when the processing scenario of the compensation processing is an audio recording scenario, the second audio data includes second audio signals corresponding to at least one recording time. Correspondingly, the step of compensating the first audio data based on the positioning jitter data to obtain jitter-compensated second audio data includes:
[0159] For each recording moment, acquire the first audio signal of the recording moment in the first audio data and the target jitter data of the recording moment in the positioning jitter data;
[0160] The first audio signal is compensated based on the target jitter data to obtain a second audio signal corresponding to the recording time.
[0161] The second audio data for jitter compensation is determined based on the second audio signal corresponding to at least one recording time.
[0162] According to one or more embodiments of the present invention, Example 5 describes the method described in Example 4, wherein the step of compensating the first audio signal based on the target jitter data to obtain a second audio signal corresponding to the recording time includes:
[0163] When the first audio signal is an audio signal based on a spherical harmonic function, the target jitter data is inverted to obtain jitter compensation data;
[0164] The compensation rotation matrix is determined based on the jitter compensation data and the spherical harmonic order of the first audio signal;
[0165] The product of the first audio signal and the compensation rotation matrix is used as the second audio signal corresponding to the recording time.
[0166] According to one or more embodiments of the present invention, Example 6 describes the method described in Example 4, wherein the step of compensating the first audio signal based on the target jitter data to obtain a second audio signal corresponding to the recording time includes:
[0167] When the first audio signal is a surround sound signal, the speaker layout data of the audio acquisition device is acquired;
[0168] Based on the speaker layout data and the target jitter data, determine the jitter gain weight for each speaker;
[0169] The channel signals corresponding to each speaker in the first audio signal are compensated according to the jitter gain weight of each speaker to obtain the second audio signal corresponding to the recording time; wherein, the jitter gain weight corresponds one-to-one with the channel signal.
[0170] According to one or more embodiments of the present invention, Example 7 describes the method described in Example 6, wherein determining the jitter gain weight of each speaker based on the speaker layout data and the target jitter data includes:
[0171] Based on the speaker layout data and the target jitter data, determine the positioning adjustment data corresponding to each speaker;
[0172] For each positioning adjustment data, at least two target speakers matching the positioning adjustment data are determined based on the speaker layout data and the positioning adjustment data. Then, based on the positioning adjustment data and the positioning data corresponding to the at least two target speakers in the speaker layout data, component gain weights corresponding to the at least two target speakers are determined. The component gain weights represent the component coefficients of the positioning adjustment data on the target speakers.
[0173] For each speaker, the summation of at least one component gain weight corresponding to the speaker is used as the jitter gain weight of the speaker.
[0174] According to one or more embodiments of the present invention, Example 8 describes the method described in Example 4, wherein the step of compensating the first audio signal based on the target jitter data to obtain a second audio signal corresponding to the recording time includes:
[0175] When the first audio signal is an object audio signal, the first position data of the sound object in the first audio signal is compensated according to the target jitter data to obtain a second audio signal corresponding to the recording time.
[0176] According to one or more embodiments of the present invention, Example 9 provides an audio processing apparatus, comprising:
[0177] The first positioning data acquisition module is used to acquire first audio data collected by the audio acquisition device and first positioning data of the audio acquisition device; wherein, the first positioning data represents a set of location information of the audio acquisition device during the process of acquiring the first audio data, and the acquisition timestamp of the first audio data and the first positioning data are the same.
[0178] The positioning jitter data determination module is used to determine the positioning jitter data of the audio acquisition device based on the first positioning data; wherein, the positioning jitter data represents a set of position fluctuation information of the audio acquisition device during the acquisition of the first audio data;
[0179] The second audio data determination module is used to perform compensation processing on the first audio data based on the positioning jitter data to obtain jitter-compensated second audio data.
[0180] According to one or more embodiments of the present invention, Example 10 provides an electronic device, comprising:
[0181] At least one processor; and
[0182] A memory communicatively connected to the at least one processor; wherein,
[0183] The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the audio processing method described in any one of Examples 1-8.
[0184] According to one or more embodiments of the present invention, Example 11 provides a computer-readable storage medium storing computer instructions for causing a processor to execute the audio processing method described in any one of Examples 1-7.
[0185] According to one or more embodiments of the present invention, Example 12 provides a computer program product including a computer program that, when executed by a processor, implements the audio processing method according to any one of Examples 1-8.
[0186] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.
[0187] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the invention. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0188] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An audio processing method, characterized in that, include: Acquire first audio data collected by the audio acquisition device and first positioning data of the audio acquisition device; wherein, the first positioning data represents a set of location information of the audio acquisition device during the process of collecting the first audio data, and the first audio data and the first positioning data have the same acquisition timestamp; Based on the first positioning data, the positioning jitter data of the audio acquisition device is determined; wherein, the positioning jitter data represents a set of position fluctuation information of the audio acquisition device during the process of acquiring the first audio data; Based on the positioning jitter data, the first audio data is compensated to obtain jitter-compensated second audio data.
2. The method according to claim 1, characterized in that, When the compensation processing scenario is an audio playback scenario, the second audio data includes second audio signals corresponding to at least one playback moment. Accordingly, the step of compensating the first audio data based on the positioning jitter data to obtain jitter-compensated second audio data includes: For each playback moment, the second positioning data of the audio playback device at the playback moment is obtained, and the first audio signal in the first audio data that matches the playback moment and the target jitter data in the positioning jitter data that matches the playback moment are obtained; wherein, the second positioning data represents the real position information of the audio playback device at the playback moment; The second positioning data is compensated based on the target jitter data to obtain the third positioning data; wherein, the third positioning data represents the jitter-compensated position information of the audio playback device at the playback time. The first audio signal is decoded based on the third positioning data to obtain a second audio signal corresponding to the playback time. The second audio data for jitter compensation is determined based on the second audio signal corresponding to at least one playback moment.
3. The method according to claim 2, characterized in that, The step of compensating the second positioning data based on the target jitter data to obtain the third positioning data includes: The target jitter data is inverted to obtain jitter compensation data, and the compensation rotation parameters are determined based on the jitter compensation data. The product of the second positioning data and the compensation rotation parameter is used as the third positioning data.
4. The method according to claim 1, characterized in that, When the compensation processing scenario is an audio recording scenario, the second audio data includes second audio signals corresponding to at least one recording moment. Accordingly, the step of compensating the first audio data based on the positioning jitter data to obtain jitter-compensated second audio data includes: For each recording moment, acquire the first audio signal of the recording moment in the first audio data and the target jitter data of the recording moment in the positioning jitter data; The first audio signal is compensated based on the target jitter data to obtain a second audio signal corresponding to the recording time. The second audio data for jitter compensation is determined based on the second audio signal corresponding to at least one recording time.
5. The method according to claim 4, characterized in that, The step of compensating the first audio signal based on the target jitter data to obtain a second audio signal corresponding to the recording time includes: When the first audio signal is an audio signal based on a spherical harmonic function, the target jitter data is inverted to obtain jitter compensation data; The compensation rotation matrix is determined based on the jitter compensation data and the spherical harmonic order of the first audio signal; The product of the first audio signal and the compensation rotation matrix is used as the second audio signal corresponding to the recording time.
6. The method according to claim 4, characterized in that, The step of compensating the first audio signal based on the target jitter data to obtain a second audio signal corresponding to the recording time includes: When the first audio signal is a surround sound signal, the speaker layout data of the audio acquisition device is acquired; Based on the speaker layout data and the target jitter data, determine the jitter gain weight for each speaker; The channel signals corresponding to each speaker in the first audio signal are compensated according to the jitter gain weight of each speaker to obtain the second audio signal corresponding to the recording time; wherein, the jitter gain weight corresponds one-to-one with the channel signal.
7. The method according to claim 6, characterized in that, The step of determining the jitter gain weight of each speaker based on the speaker layout data and the target jitter data includes: Based on the speaker layout data and the target jitter data, determine the positioning adjustment data corresponding to each speaker; For each positioning adjustment data, at least two target speakers matching the positioning adjustment data are determined based on the speaker layout data and the positioning adjustment data. Then, based on the positioning adjustment data and the positioning data corresponding to the at least two target speakers in the speaker layout data, component gain weights corresponding to the at least two target speakers are determined. The component gain weights represent the component coefficients of the positioning adjustment data on the target speakers. For each speaker, the summation of at least one component gain weight corresponding to the speaker is used as the jitter gain weight of the speaker.
8. The method according to claim 4, characterized in that, The step of compensating the first audio signal based on the target jitter data to obtain a second audio signal corresponding to the recording time includes: When the first audio signal is an object audio signal, the first position data of the sound object in the first audio signal is compensated according to the target jitter data to obtain a second audio signal corresponding to the recording time.
9. An audio processing device, characterized in that, include: The first positioning data acquisition module is used to acquire first audio data collected by the audio acquisition device and first positioning data of the audio acquisition device; wherein, the first positioning data represents a set of location information of the audio acquisition device during the process of acquiring the first audio data, and the acquisition timestamp of the first audio data and the first positioning data are the same. The positioning jitter data determination module is used to determine the positioning jitter data of the audio acquisition device based on the first positioning data; wherein, the positioning jitter data represents a set of position fluctuation information of the audio acquisition device during the acquisition of the first audio data; The second audio data determination module is used to perform compensation processing on the first audio data based on the positioning jitter data to obtain jitter-compensated second audio data.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the audio processing method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the audio processing method of any one of claims 1-8.
12. A computer program product comprising a computer program that, when executed by a processor, implements the audio processing method according to any one of claims 1-8.