Metadata calibration method and device of audio data and storage medium

CN121890107APending Publication Date: 2026-04-17BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING XIAOMI MOBILE SOFTWARE CO LTD
Filing Date
2024-08-16
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

If the input audio signal is inaccurate, the accuracy of the metadata cannot be guaranteed, affecting the reconstruction effect of the audio signal.

Method used

By determining the difference parameter between the performance data of the first audio data and the expected performance data, the output metadata is calibrated based on the difference parameter to improve the accuracy of the metadata.

Benefits of technology

It enables metadata calibration when the input audio data is inaccurate, improving the accuracy of the metadata and thus enhancing the quality of audio signal reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121890107A_ABST
    Figure CN121890107A_ABST
Patent Text Reader

Abstract

The invention provides a metadata calibration method and device for audio data and a storage medium. According to the method and the device, the difference parameter between the performance data of the first audio data and the expected performance data is determined, so that the metadata of the first audio data is calibrated based on the difference parameter, and the metadata calibration of the first audio data is realized under the condition that the difference exists between the performance of the first audio data and the expected performance; the metadata accuracy of the first audio data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Metadata calibration method and device for audio data, and storage medium TECHNICAL FIELD

[0001] The present disclosure relates to the field of communication, and in particular, to a metadata calibration method and device for audio data, and a storage medium. BACKGROUND

[0002] Parametric audio, as a technology for describing audio content by using mathematical models and algorithms, can realize the description and reconstruction of audio signals by using less data. As one of the key parameters for reconstructing audio signals, the accuracy of metadata plays a crucial role in the reconstruction result of audio signals.

[0003] SUMMARY

[0004] In order to calibrate the metadata of the data audio signal in the case that the input audio signal is inaccurate, the present disclosure provides a metadata calibration method and device for audio data, and a storage medium.

[0005] According to a first aspect of the present disclosure, a metadata calibration method for audio data is provided, the method comprising:

[0006] determining a difference parameter between performance data of first audio data and expected performance data;

[0007] calibrating metadata output based on the first audio data based on the difference parameter.

[0008] According to a second aspect of the present disclosure, a communication device is provided, comprising:

[0009] a processing module configured to determine a difference parameter between performance data of first audio data and expected performance data;

[0010] the processing module is further configured to calibrate metadata output based on the first audio data based on the difference parameter.

[0011] According to a third aspect of the present disclosure, a communication device is provided, comprising:

[0012] one or more processors;

[0013] The communication device is configured to perform the metadata calibration method for audio data according to the first aspect.

[0014] According to a fourth aspect of the present disclosure, a storage medium is provided, the storage medium storing instructions, when the instructions are run on a communication device, causing the communication device to perform the metadata calibration method for audio data according to the first aspect.

[0015] The embodiments of the present disclosure determine a difference parameter between the performance data of the first audio data and the expected performance data, and calibrate the metadata output based on the first audio data based on the difference parameter, so as to realize calibration of the output metadata in the case that there is a gap between the performance of the first audio data and the expected performance, and improve the accuracy of the output metadata.

[0016] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, further serve to explain the principles of the present disclosure.

[0018] FIG. 1 is a schematic diagram of an application scenario of a metadata calibration method of audio data according to an embodiment of the present disclosure.

[0019] FIG. 2 is a schematic diagram of a flow of a metadata calibration method of audio data according to an embodiment of the present disclosure.

[0020] FIG. 3 is a schematic diagram of a structure of a communication device according to an embodiment of the present disclosure.

[0021] FIG. 4A is a schematic diagram of a structure of a communication device 4100 according to an embodiment of the present disclosure.

[0022] FIG. 4B is a schematic diagram of a structure of a chip 4200 according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] The exemplary embodiments will be described in detail herein with reference to the attached drawings. The following description is with reference to the drawings, in which like numerals refer to like elements throughout. The embodiments described below are not meant to represent all embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0024] The terminology used in the present disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used in the present disclosure and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0025] It should be understood that, although the terms first, second, third, etc. can be employed in this disclosure to describe various messages, these messages should not be limited to these terms. These terms are only used to distinguish one message from another. For example, a first message can also be termed a second message, and, similarly, a second message can also be termed a first message, without departing from the scope of this disclosure. Depending on the context, the word "if' as used herein can be interpreted as meaning "when" or "in response to determining" or "in response to ascertaining".

[0026] The embodiments of the present disclosure provide a metadata calibration method and device for audio data and a storage medium.

[0027] In a first aspect, the embodiments of the present disclosure provide a metadata calibration method for audio data, the method comprising:

[0028] determining a difference parameter between performance data of the first audio data and expected performance data;

[0029] calibrating metadata output based on the first audio data based on the difference parameter.

[0030] In the above embodiment, by determining the difference parameter between the performance data of the first audio data and the expected performance data, the metadata output based on the first audio data is calibrated based on the difference parameter, so as to realize the calibration of the output metadata in the case that there is a gap between the performance of the first audio data and the expected performance, thereby improving the accuracy of the output metadata.

[0031] In combination with some embodiments of the first aspect, in some embodiments, the difference parameter between the performance data of the first audio data and the expected performance data comprises at least one of:

[0032] a difference value between the performance data of the first audio data and the expected performance data;

[0033] a proportional relationship between the performance data of the first audio data and the expected performance data.

[0034] In the above embodiment, multiple optional parameter types of the difference parameter between the performance data of the first audio data and the expected performance data are provided, so that the determination of the difference parameter can be realized in multiple ways, thereby improving the flexibility of the difference parameter determination process.

[0035] In combination with some embodiments of the first aspect, in some embodiments, the first audio data comprises audio signals on multiple input channels;

[0036] The determination of the difference parameter between the performance data of the first audio data and the expected performance data comprises:

[0037] determine a difference parameter between the performance data and the expected performance data of the audio signal on each input channel, to obtain the difference parameter corresponding to each of the plurality of input channels.

[0038] In the above embodiment, the difference parameter corresponding to each of the plurality of input channels is determined by determining the difference parameter between the performance data and the expected performance data of the audio signal on each input channel, so that the difference determination can be realized from each input channel, thereby ensuring that the calibration of the output metadata is realized based on the difference parameters on the plurality of input channels, to improve the calibration effect of the output metadata.

[0039] In combination with some embodiments of the first aspect, in some embodiments, the method further includes:

[0040] obtaining performance data of the first audio data.

[0041] In the above embodiment, the performance data of the first audio data is obtained, so that the determination of the difference parameter based on the obtained performance data and the expected performance data is ensured, thereby ensuring the smooth progress of the metadata calibration process of the audio data.

[0042] In combination with some embodiments of the first aspect, in some embodiments, the obtaining of the performance data of the first audio data includes any one of the following:

[0043] in response to obtaining the first audio data, performing audio processing based on the first audio data to obtain the performance data of the first audio data;

[0044] obtaining the performance data of the first audio data based on pre-stored performance data corresponding to the first audio data;

[0045] determining the performance data of the first audio data based on a device parameter of a device used to collect the first audio data.

[0046] In the above embodiment, a plurality of optional implementation manners of obtaining the performance data of the first audio data are provided, so that the performance data of the first audio data can be obtained in a plurality of ways, and the flexibility of the performance data obtaining process of the audio data is improved.

[0047] In combination with some embodiments of the first aspect, in some embodiments, the device parameter includes at least one of the following:

[0048] directivity;

[0049] frequency response curve parameter;

[0050] sensitivity;

[0051] delay parameter;

[0052] distortion parameter;

[0053] noise.

[0054] In the above embodiment, a plurality of optional device parameters are provided, so that the determination of the performance data of the first audio data can be implemented based on the plurality of device parameters, to improve the flexibility of the process of implementing the determination of the performance data of the audio data based on the device parameters.

[0055] In combination with some embodiments of the first aspect, in some embodiments, the expected performance data is agreed upon by a protocol, or the expected performance data is customized.

[0056] In the above embodiment, by providing a plurality of optional implementation manners of configuring the expected performance data, the configuration of the expected performance data can be implemented in a plurality of ways, so that the flexibility of the process of configuring the expected performance data can be improved.

[0057] In combination with some embodiments of the first aspect, in some embodiments, different types of audio data correspond to different expected performance data; or,

[0058] Different audio data processing scenarios correspond to different expected performance data.

[0059] In the above embodiment, by providing a plurality of optional implementation manners of configuring the expected performance data, the configuration of the expected performance data can be implemented not only according to the type of the audio data, but also according to the audio data processing scenario, so that the flexibility and diversity of the configuration manner of the expected performance data are improved.

[0060] In combination with some embodiments of the first aspect, in some embodiments, the first audio data includes audio signals on a plurality of input channels, and each input channel corresponds to a difference parameter;

[0061] The calibration of the metadata output based on the first audio data based on the difference parameter includes:

[0062] Calibration of the metadata output based on the audio signals on the plurality of input channels based on the difference parameter corresponding to each input channel.

[0063] In the above embodiment, in the case where the first audio data includes audio signals on a plurality of input channels, the metadata output based on the audio signals on the plurality of input channels is calibrated based on the difference parameter corresponding to each input channel, to implement the calibration of the metadata output based on the first audio data, which can improve the calibration effect of the metadata, thereby improving the accuracy of the output metadata.

[0064] In some embodiments of the first aspect, in some embodiments, the calibrating the metadata output based on the first audio data based on the difference parameter comprises any one of the following:

[0065] calibrating the first audio data based on the difference parameter, and performing calculation based on the calibrated first audio data to obtain the calibrated metadata;

[0066] performing calculation based on the difference parameter and the first audio data to obtain the calibrated metadata;

[0067] correcting the metadata calculated based on the first audio data based on the difference parameter to obtain the calibrated metadata;

[0068] performing calculation based on the first audio data, and processing the metadata calculated based on the first audio data based on the first processing and the difference parameter.

[0069] In the above embodiments, multiple optional implementation manners of calibrating the output metadata based on the difference parameter are provided, so that the metadata calibration of the audio data can be implemented in multiple ways, and the flexibility of the metadata calibration process of the audio data is improved.

[0070] In some embodiments of the first aspect, in some embodiments, the metadata comprises at least one of the following:

[0071] direction of arrival (DOA);

[0072] direct sound to total energy ratio;

[0073] diffusion coefficient;

[0074] distance;

[0075] diffuse sound to total energy ratio;

[0076] surround coefficient;

[0077] residual sound to total energy ratio.

[0078] In the above embodiments, multiple optional metadata types are provided to implement calibration of multiple types of metadata, and the flexibility and comprehensiveness of the metadata calibration process are improved.

[0079] In a second aspect, the embodiments of the present disclosure provide a communication device, comprising:

[0080] a processing module configured to determine a difference parameter between performance data of first audio data and expected performance data;

[0081] The processing module is further configured to calibrate metadata output based on the first audio data based on the difference parameter.

[0082] In a third aspect, an embodiment of the present disclosure provides a communication device, comprising:

[0083] one or more processors;

[0084] The communication device is configured to perform the metadata calibration method of audio data as described in the first aspect and any one of the embodiments of the first aspect.

[0085] In a fourth aspect, an embodiment of the present disclosure provides a storage medium storing instructions that, when executed on a communication device, cause the communication device to perform the metadata calibration method of audio data as described in the first aspect and any one of the embodiments of the first aspect.

[0086] In a fifth aspect, an embodiment of the present disclosure provides a program product that, when executed on a communication device, causes the communication device to perform the metadata calibration method of audio data as described in the first aspect and any one of the embodiments of the first aspect.

[0087] In a sixth aspect, an embodiment of the present disclosure provides a computer program that, when executed on a computer, causes the computer to perform the metadata calibration method of audio data as described in the first aspect and any one of the embodiments of the first aspect.

[0088] In a seventh aspect, an embodiment of the present disclosure provides a chip or chip system. The chip or chip system comprises processing circuitry configured to perform the metadata calibration method of audio data as described in the first aspect and any one of the embodiments of the first aspect.

[0089] It can be understood that the above-mentioned communication device, storage medium, program product, computer program, chip or chip system are all used to perform the method proposed in the embodiments of the present disclosure. Therefore, the beneficial effects that can be achieved are referred to the beneficial effects in the corresponding method, which will not be described here.

[0090] The embodiments of the present disclosure provide a metadata calibration method and device of audio data, and a storage medium. In some embodiments, the metadata calibration method of audio data, the information processing method, the communication method, the data processing method, the audio data processing method, and the metadata calibration method of parametric audio can be replaced with each other, the metadata calibration device of audio data, the information processing device, the communication device, the data processing device, the audio data processing device, and the metadata calibration device of parametric audio can be replaced with each other, and the information processing system and the communication system can be replaced with each other.

[0091] The embodiments of the present disclosure are not exhaustive, but only illustrate some embodiments, and are not specific limitations on the protection scope of the present disclosure. In the case of no contradiction, each step in an embodiment can be implemented as an independent embodiment, and the steps can be combined arbitrarily, for example, the scheme after removing part of the steps in an embodiment can also be implemented as an independent embodiment, and the order of the steps in an embodiment can be exchanged arbitrarily, in addition, the optional implementation manners in an embodiment can be combined arbitrarily; in addition, the embodiments can be combined arbitrarily, for example, part or all steps of different embodiments can be combined arbitrarily, an embodiment can be combined with optional implementation manners of other embodiments arbitrarily.

[0092] In each embodiment of the present disclosure, the terms and / or descriptions between the embodiments are consistent if there is no special description and logical conflict, and can be referred to each other, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.

[0093] The terms used in the embodiments of the present disclosure are only for the purpose of describing the specific embodiments, and not as a limitation on the present disclosure.

[0094] In the embodiments of the present disclosure, unless otherwise specified, the elements expressed in singular form, such as "one", "a", "the", "above", "said", "preceding", "this" and the like, can represent "one and only one", and can also represent "one or more", "at least one" and the like. For example, in the case of using articles such as "a", "an", "the" and the like in English, the noun after the article can be understood as singular expression, and can also be understood as plural expression.

[0095] In the embodiments of the present disclosure, "a plurality of" means two or more.

[0096] In some embodiments, the terms "at least one of", "one or more", "a plurality of", "multiple" and the like can be replaced with each other.

[0097] In some embodiments, "at least one of A, B", "A and / or B", "in one case A, in another case B", "responsive to case A, responsive to case B" and the like, can be interpreted to include both cases, A and B, in some embodiments, A (A is performed regardless of B), in some embodiments, B (B is performed regardless of A), in some embodiments, selected from the group consisting of A and B (the selection between A and B is an option), in some embodiments, A and B (both A and B are performed).

[0098] In some embodiments, "A or B" and the like, can be interpreted to include both cases, A and B, in some embodiments, A (A is performed regardless of B), in some embodiments, B (B is performed regardless of A), in some embodiments, selected from the group consisting of A and B (the selection between A and B is an option).

[0099] The prefix words "first", "second" and the like in the embodiments of the present disclosure are merely intended to distinguish different description objects, and do not constitute limitation on the position, order, priority, quantity or content of the description objects. The description objects are described in the claims or embodiments in the context, and should not be construed as redundant limitation because of the use of the prefix words. For example, the description object is "field", and the ordinal words before "field" in "first field" and "second field" do not limit the position or order between "fields", and "first" and "second" do not limit whether the "fields" modified thereby are in the same message or not, nor limit the order of "first field" and "second field". For another example, the description object is "level", and the ordinal words before "level" in "first level" and "second level" do not limit the priority between "levels". For another example, the quantity of the description object is not limited by the ordinal words, and can be one or more. For example, "first device", wherein the quantity of "device" can be one or more. In addition, the objects modified by different prefix words can be the same or different, for example, the description object is "device", and "first device" and "second device" can be the same device or different devices, and the types thereof can be the same or different. For another example, the description object is "information", and "first information" and "second information" can be the same information or different information, and the contents thereof can be the same or different.

[0100] In some embodiments, "including A", "containing A", "for indicating A", "carrying A" can be interpreted as directly carrying A, or indirectly indicating A.

[0101] In some embodiments, the terms "in response to", "in response to determining", "in the case of", "when", "when", "if", "if" and the like can be replaced with each other.

[0102] In some embodiments, the terms "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not lower than", "above", and the like can be replaced with each other, and the terms "less than", "less than or equal to", "not greater than", "fewer than", "fewer than or equal to", "not more than", "lower than", "lower than or equal to", "not higher than", "below", and the like can be replaced with each other.

[0103] In some embodiments, the apparatuses and devices can be interpreted as physical or virtual, and their names are not limited to the names described in the embodiments, and in some cases can also be understood as "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "subject", and the like.

[0104] In some embodiments, "network" can be interpreted as an apparatus included in the network, such as an access network device, a core network device, and the like.

[0105] In some embodiments, an “access network device (AN device)” can also be referred to as a “radio access network device (RAN device),” a “base station (BS),” a “radio base station,” a “fixed station,” and in some embodiments can also be understood as a “node,” an “access point,” a “transmission point (TP),” a “reception point (RP),” a “transmission / reception point (TRP),” a “panel,” an “antenna panel,” an “antenna array,” a “cell,” a “macro cell,” a “small cell,” a “femto cell,” a “pico cell,” a “sector,” a “cell group,” a “serving cell,” a “carrier,” a “component carrier,” a “bandwidth part (BWP),” and the like.

[0106] In some embodiments, a "terminal" or "terminal device" can be referred to as a "user equipment" (UE), a "user terminal," a "mobile station" (MS), a "mobile terminal" (MT), a subscriber station, a mobile unit, a subscriber unit, a wireless unit, a remote unit, a mobile device, a wireless device, a wireless communication device, a remote device, a mobile subscriber station, an access terminal, a mobile terminal, a wireless terminal, a remote terminal, a handset, a user agent, a mobile client, a client, and / or the like.

[0107] In some embodiments, data, information, and / or the like can be obtained in compliance with laws and regulations of a country in which the data, information, and / or the like is obtained.

[0108] In some embodiments, data, information, and / or the like can be obtained after consent of a user is obtained.

[0109] Further, each element, each row, or each column in a table of embodiments of the present disclosure can be implemented as an independent embodiment, and a combination of any element, any row, or any column can be implemented as an independent embodiment.

[0110] FIG. 1 is a schematic diagram of an application scenario of a metadata calibration method of audio data according to embodiments of the present disclosure. As shown in FIG. 1, the metadata calibration method of audio data provided by embodiments of the present disclosure can be applied in a scenario including a terminal 101 and a server 102.

[0111] In some embodiments, the terminal 101 includes at least one of a mobile phone, a wearable device, an Internet of Things device, a communication-capable automobile, a smart automobile, a tablet (Pad), a wireless transceiver-equipped computer, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal device in industrial control, a wireless terminal device in self-driving, a wireless terminal device in remote medical surgery, a wireless terminal device in a smart grid, a wireless terminal device in transportation safety, a wireless terminal device in a smart city, a wireless terminal device in a smart home, and the like, but is not limited thereto.

[0112] In some embodiments, the server 102 includes at least one of a virtual server, an edge server, a proxy server, a database server, an application server, a storage server, a load balancer, a container platform, a server cluster, a cloud computing platform, and the like, but is not limited thereto.

[0113] It can be understood that the application scenarios described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the present disclosure, and do not constitute a limitation on the technical solutions proposed by the present disclosure. Those skilled in the art can know that, as the system architecture evolves and new business scenarios appear, the technical solutions proposed by the present disclosure are also applicable to similar technical problems.

[0114] The following embodiments of the present disclosure can be applied to all or part of the subjects shown in FIG. 1, but are not limited thereto. The subjects shown in FIG. 1 are exemplary, and the application scenarios can include all or part of the subjects in FIG. 1, or other subjects other than those in FIG. 1. The number and form of each subject is arbitrary, each subject can be real or virtual, the connection relationship between each subject is exemplary, each subject can be connected or not connected, and the connection can be in any manner, can be direct or indirect, can be wired or wireless.

[0115] Embodiments of the present disclosure can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G new radio (NR), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New radio access (NX), Future generation radio access (FX), Global System for Mobile communications (GSM (registered trademark)), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, Ultra-WideBand (UWB), Bluetooth (Bluetooth (registered trademark)), Public Land Mobile Network (PLMN) network, Device-to-Device (D2D) system, Machine to Machine (M2M) system, Internet of Things (IoT) system, Vehicle-to-Everything (V2X), system using other communication methods, next-generation system expanded based thereon, and the like. In addition, a plurality of systems can be combined (for example, combination of LTE or LTE-A and 5G, and the like).

[0116] In the related art, parametric audio, which is a technology of describing audio content by parameters, can describe and reconstruct an audio signal by using a mathematical model and an algorithm with less data.

[0117] In some embodiments, metadata-assisted spatial audio format (MASA) audio is one of the audio formats supported by the Immersive Voice and Audio Services (IVAS) codec as a kind of parametric audio.

[0118] In some embodiments, the metadata of the MASA audio is calculated by the audio data converted from the audio input signal, and thus the accuracy of the input audio signal directly affects the accuracy of the metadata of the audio data. When the input audio signal is inaccurate, the metadata of the audio data is inaccurate, thereby affecting the accuracy of the audio data reconstructed based on the metadata and the final audio effect.

[0119] Therefore, the embodiments of the present disclosure aim to provide a metadata calibration method of audio data to solve the problem of inaccurate metadata determined in the case of non-standard audio data signal.

[0120] FIG. 2 is a flowchart of a metadata calibration method of audio data according to an embodiment of the present disclosure. As shown in FIG. 2, the embodiments of the present disclosure relate to a metadata calibration method of audio data, and the method can be executed by a communication device, and the method comprises the following steps:

[0121] In step S2101, first audio data is obtained.

[0122] In some embodiments, the communication device obtains audio data input therein as the first audio data. That is, the first audio data can be the input audio of the communication device.

[0123] In some embodiments, the communication device can obtain the first audio data by collecting an audio signal.

[0124] Optionally, the first audio data can be an audio input in any format; or the first audio data can be an audio input in a standard format, for example, the first audio data can be Ambisonic audio, binaural audio, stereo audio, multi-channel audio, etc., but is not limited thereto.

[0125] In some embodiments, the communication device is a terminal, or the communication device is a server.

[0126] In some embodiments, the communication device is a terminal, and the first audio data can be pre-stored audio data in the terminal, or the first audio data can be audio data collected by a built-in or externally connected audio collection component of the terminal, but is not limited thereto.

[0127] In some embodiments, the communication device is a server, and the first audio data can be pre-stored audio data in the server, or the first audio data can be audio data uploaded to the server by a terminal, but the present application is not limited thereto.

[0128] In some embodiments, the name of the first audio data is not limited, and is, for example, “input audio data”, “first audio signal”, “input audio signal”, etc.

[0129] In some embodiments, the first audio data includes audio signals on multiple input channels.

[0130] For example, the first audio data is a first order ambisonic (FOA) signal, and the first audio data can be composed of audio signals on four input channels, which are, respectively, an omnidirectional (W) channel, a front-back (X) channel, a left-right (Y) channel, and an up-down (Z) channel, to represent the distribution of sound in the horizontal and vertical directions through the four channels.

[0131] In some embodiments, the names of signals and the like are not limited to the names described in the embodiments, and terms such as “signal”, “information”, “message”, “signaling”, “report”, “configuration”, “indication”, “instruction”, “command”, “channel”, “parameter”, “domain”, “field”, “symbol”, “symbol”, “codebook”, “codeword”, “codepoint”, “bit”, “data”, “program”, “chip”, etc. can be replaced with each other.

[0132] In step S2102, performance data of the first audio data is obtained.

[0133] In some embodiments, the performance data of the first audio data includes audio signal energy and / or audio signal intensity of the first audio data, but the present application is not limited thereto.

[0134] In some embodiments, the first audio data includes audio signals on multiple input channels, and the communication device can obtain performance data of the audio signals on each input channel, respectively.

[0135] For example, the first audio data is a FOA signal composed of audio signals on four input channels, and the communication device can obtain performance data of the audio signals on the W channel, the X channel, the Y channel and the Z channel respectively to obtain the performance data of the first audio data.

[0136] In some embodiments, the communication device can obtain the performance data of the first audio data by measuring the first audio data in response to obtaining the first audio data.

[0137] That is, the communication device can obtain the performance data of the first audio data by analyzing the first audio data when detecting the input first audio data through real-time detection of the input audio, so as to ensure that the performance data of the input audio can be adjusted in real time according to the real-time fluctuation of the system, thereby ensuring the accuracy of the determined audio performance data.

[0138] In some embodiments, the communication device can store pre-stored performance data corresponding to the first audio data, and the communication device can obtain the performance data of the first audio data based on the pre-stored performance data corresponding to the first audio data.

[0139] For example, the first audio data can be obtained by the communication device in the past, and the communication device has obtained the performance data of the first audio data by measuring the first audio data in advance, and has stored the pre-measured performance data as the pre-stored performance data corresponding to the first audio data. The communication device can obtain the performance data of the first audio data based on the pre-stored performance data corresponding to the first audio data.

[0140] In some embodiments, the communication device determines the performance data of the first audio data based on device parameters of a device used to collect the first audio data.

[0141] For example, the communication device can determine the device parameters of the device used to collect the first audio data based on the product design of the device used to collect the first audio data, so as to determine the performance data of the first audio data based on the device parameters of the device used to collect the first audio data, thereby achieving the purpose of estimating the performance data through the product design of the device used to collect the first audio data.

[0142] In some embodiments, the device parameters include at least one of directivity, frequency response curve parameters, sensitivity, delay parameters, distortion parameters, and noise, but are not limited thereto.

[0143] Step S2103: determining a difference parameter between the performance data of the first audio data and the expected performance data.

[0144] In some embodiments, the expected performance data is agreed upon by a protocol, or the expected performance data is self-defined.

[0145] In some embodiments, the name of the expected performance data is not limited, which is, for example, “ideal performance data”, “standard performance data”, and the like.

[0146] In some embodiments, different types of audio data correspond to different expected performance data, or different audio data processing scenarios correspond to different expected performance data.

[0147] In some embodiments, different types of audio data corresponding to the expected performance data can be agreed in the protocol, and the communication device can determine the expected performance data of the first audio data according to the type of the first audio data.

[0148] In some embodiments, different types of audio data corresponding to the expected performance data can be customized by the manufacturer of the communication device, and the communication device can determine the expected performance data of the first audio data according to the type of the first audio data.

[0149] In some embodiments, different audio data processing scenarios corresponding to the expected performance data can be agreed in the protocol, and the communication device can determine the expected performance data of the first audio data according to the current audio data processing scenario.

[0150] In some embodiments, different audio data processing scenarios corresponding to the expected performance data can be customized by the manufacturer of the communication device, and the communication device can determine the expected performance data of the first audio data according to the current audio data processing scenario.

[0151] It should be noted that the above is only a few exemplary ways of configuring the expected performance data, and does not constitute a determination of the embodiments of the present disclosure.

[0152] In some embodiments, the audio data includes audio signals on multiple input channels, and the expected performance data of the audio signals on each input channel can be agreed in the protocol, or the expected performance data of the audio signals on each input channel can be customized.

[0153] In some embodiments, the expected performance data of the audio signals on different input channels can be the same or different, and the embodiments of the present disclosure do not limit this.

[0154] In some embodiments, the difference parameter between the performance data of the first audio data and the expected performance data can be the difference between the performance data of the first audio data and the expected performance data; or the difference parameter between the performance data of the first audio data and the expected performance data can be the proportional relationship between the performance data of the first audio data and the expected performance data; or the difference parameter between the performance data of the first audio data and the expected performance data can include the proportional relationship and the difference between the performance data of the first audio data and the expected performance data.

[0155] In some embodiments, the communication device can determine a difference value of the performance data of the first audio data and the expected performance data as the difference parameter between the performance data of the first audio data and the expected performance data.

[0156] In some embodiments, the communication device can determine a ratio relationship of the performance data of the first audio and the expected performance data as the difference parameter between the performance data of the first audio data and the expected performance data.

[0157] In some embodiments, the communication device can determine a ratio relationship and a difference value of the performance data of the first audio and the expected performance data as the difference parameter between the performance data of the first audio data and the expected performance data.

[0158] In some embodiments, the first audio data includes audio signals on a plurality of input channels, and the communication device can determine the difference parameter between the performance data of the audio signals on each input channel and the expected performance data respectively, to obtain a plurality of difference parameters respectively corresponding to the plurality of input channels.

[0159] Step S2104, calibrating the output metadata based on the difference parameter.

[0160] In some embodiments, the communication device can calibrate the first audio data based on the difference parameter, to calculate calibrated metadata based on the calibrated first audio data.

[0161] That is, the difference parameter can be directly applied to the first audio data to obtain accurate first audio data, so as to determine accurate metadata based on the accurate first audio data.

[0162] In some embodiments, the communication device can calculate calibrated metadata based on the difference parameter and the first audio data.

[0163] That is, the difference parameter can be directly applied to the calculation of the metadata, and the calculation of the corresponding metadata can be changed to directly output correct metadata.

[0164] In some embodiments, the communication device can correct the metadata calculated based on the first audio data based on the difference parameter, to obtain calibrated metadata.

[0165] That is, the calculation process of the metadata can not be changed, and the metadata can be calibrated after being output to directly adjust the metadata through the difference parameter.

[0166] In some embodiments, the communication device can perform a calculation based on the first audio data to process metadata calculated based on the first audio data based on the first processing and the difference parameter. The first processing can be a subsequent processing process, for example, the first processing can be a rendering process, but is not limited thereto.

[0167] That is, the difference parameter can be passed into the subsequent processing process to participate in the calculation, so as to adjust the output audio based on the metadata and the difference parameter in the subsequent processing process. For example, the difference parameter can be passed into the subsequent rendering process to participate in the calculation, so as to adjust the output audio when rendering.

[0168] In some embodiments, the first audio data includes audio signals on multiple input channels, and each input channel corresponds to a difference parameter. The communication device can calibrate the output metadata based on the difference parameter corresponding to each input channel. That is, the communication device can calibrate the metadata determined based on the audio signals on different input channels based on the difference parameter corresponding to each input channel.

[0169] In some embodiments, the metadata includes a direction of arrival (DOA), a direct-to-total energy ratio, a spread coherence, a distance, a diffuse-to-total energy ratio, a surround coherence, and a remainder-to-total energy ratio.

[0170] For example, the first audio data is a FOA signal composed of audio signals on four channels, and the horizontal direction of arrival can be calculated by the following formula (1):

[0171] Wherein, θ represents the horizontal direction of arrival, i y represents the signal intensity of the first audio data in the Y direction, i x represents the signal intensity of the first audio data in the X direction.

[0172] For example, the first audio data is a FOA signal composed of audio signals on four channels, and the direct-to-total energy ratio can be calculated by the following formula (2):

[0173] Wherein, r dirrepresents a diffuse sound to total energy ratio, i represents a signal strength of the first audio data, e band represents a total energy of the first audio data.

[0174] For example, the first audio data is a FOA signal composed of audio signals on four channels, the diffuse sound to total energy ratio can be calculated by the following formula (3): r diff = 1 - r dir (3)

[0175] wherein r diff represents a diffuse sound to total energy ratio, r dir represents a direct sound to total energy ratio.

[0176] For example, the first audio data is a FOA signal composed of audio signals on four channels, the diffuse coefficient can be calculated by the following formula (4): C spr = r dir C gen (4)

[0177] wherein C spr represents a diffuse coefficient, r dir represents a direct sound to total energy ratio, C gen represents a general coherence parameter.

[0178] For example, the first audio data is a FOA signal composed of audio signals on four channels, the surround coefficient can be calculated by the following formula (5): C sur = r diff C gen (5)

[0179] wherein C sur represents a surround coefficient, r diff represents a diffuse sound to total energy ratio, C gen represents a general coherence parameter.

[0180] wherein C gen can be calculated by the following formula (6):

[0181] wherein C gen represents a general coherence parameter, b x represents an audio signal of the first audio data on an X channel, b y represents an audio signal of the first audio data on a Y channel, b z represents an audio signal of the first audio data on a Z channel, b w represents an audio signal of the first audio data on a W channel.

[0182] For example, the first audio data is a FOA signal composed of audio signals on four channels,

[0183] It should be noted that the residual sound to total energy ratio can represent non-scene energy (e.g., known microphone noise), and optionally, the residual sound to total energy ratio can be set to 0, but is not limited thereto,

[0184] In some embodiments, "obtain", "get", "receive", "transmit", "bidirectionally transmit", "send and / or receive" can be replaced with each other, and can be interpreted as receiving from other subjects, obtaining from protocols, obtaining from higher layers, processing by itself, autonomously implementing, and the like.

[0185] In some embodiments, the terms "send", "transmit", "report", "issue", "transmit", "bidirectionally transmit", "send and / or receive" can be replaced with each other.

[0186] In some embodiments, the terms "certain", "preset", "pre-set", "set", "indicated", "certain", "arbitrary", "first", and the like can be replaced with each other, and "certain A", "preset A", "pre-set A", "set A", "indicated A", "certain A", "arbitrary A", "first A" can be interpreted as A specified in advance in protocols and the like, or can be interpreted as A obtained by setting, configuring, or indicating, or can be interpreted as certain A, certain A, arbitrary A, or first A, but is not limited thereto.

[0187] In some embodiments, determination or judgment can be performed by a value represented by 1 bit (0 or 1), or by a true or false value (Boolean value) represented by true or false, or by comparison of numerical values (e.g., comparison with a predetermined value), but is not limited thereto.

[0188] In some embodiments, "not expecting to receive" can be interpreted as not receiving on time domain resources and / or frequency domain resources, or can be interpreted as not performing subsequent processing on the data and the like after receiving the data and the like; "not expecting to send" can be interpreted as not sending, or can be interpreted as sending but not expecting the receiving party to respond to the content of the sending.

[0189] The communication method related to the embodiments of the present disclosure can include at least one of steps S2101-S2104. For example, step S2103 can be implemented as an independent embodiment, step S2104 can be implemented as an independent embodiment, steps S2101+S2103 can be implemented as an independent embodiment, steps S2102+S2103 can be implemented as an independent embodiment, steps S2101+S2104 can be implemented as an independent embodiment, steps S2102+S2104 can be implemented as an independent embodiment, steps S2103+S2104 can be implemented as an independent embodiment, steps S2101+S2102+S2103 can be implemented as an independent embodiment, steps S2101+S2102+S2104 can be implemented as an independent embodiment, steps S2102+S2103+S2104 can be implemented as an independent embodiment, but not limited thereto.

[0190] In some embodiments, steps S2101, S2102, S2103 are optional, and one or more of these steps can be omitted or replaced in different embodiments.

[0191] In some embodiments, steps S2101, S2102, S2104 are optional, and one or more of these steps can be omitted or replaced in different embodiments.

[0192] In some embodiments, other optional implementations described before or after the description corresponding to FIG. 2 can be referred to.

[0193] According to the scheme provided by the embodiments of the present disclosure, the audio metadata can be calibrated (or compensated) according to the difference between the actual performance of the input signal and the standard performance, so as to solve the problem of inaccurate metadata in the case that the input audio signal of the UE and other civil equipment is not standard.

[0194] In some embodiments, the metadata is calculated based on the input audio data, so the error of the metadata can be obtained by determining the error of the input audio data, so as to compensate the metadata based on the error of the metadata.

[0195] In some embodiments, the error of the metadata can be calculated according to the difference between the actual performance of the input audio data and the standard performance requirement (i.e. expected performance data), since the audio system can generally be considered as a linear system, so the error of the metadata can be calculated according to the difference between the actual performance of the input audio data and the standard performance requirement (i.e. expected performance data), so as to calibrate (or compensate) the metadata to obtain accurate metadata.

[0196] In some embodiments, the actual performance of the input audio data is determined by pre-measuring the audio data to obtain pre-stored performance data; or the actual performance of the input audio data can be determined through product design; or the actual performance of the input audio data can be determined through real-time detection of the input audio, so as to adjust the actual performance of the input audio data in real time according to real-time fluctuations of the system.

[0197] In some embodiments, the difference parameter can be directly used in the calculation of the metadata to change the calculation of the corresponding metadata, and accurate metadata can be directly output; or the calculation of the metadata can not be changed, and the output metadata can be calibrated after the metadata is output, that is, the output metadata is adjusted through the difference parameter; or the difference parameter can be passed to the subsequent module for calculation, for example, the difference parameter can be passed to the subsequent rendering module to output the audio adjusted by the metadata when rendering.

[0198] For ease of understanding, the scheme provided by the embodiments of the present disclosure is further described below with reference to a specific example of quantized audio defined in the 3GPP standard.

[0199] Taking an example of input including four signals on channels W, X, Y and Z, if there is an error in the X signal, the input data will be inaccurate, thereby causing the metadata to be inaccurate.

[0200] For a general coherence parameter C involved in the calculation of the metadata gen , the general coherence parameter can be calculated by the formula , where b x , b y , b z , b w are actual input audios on the X, Y, Z and W channels respectively. Assuming that Δb x , Δb y , Δb z , Δb w are difference parameters of the actual input audios on the X, Y, Z and W channels respectively, b′ x , b′ y , b′ z , b′ w are accurate input audios on the X, Y, Z and W channels respectively, if b x = f (Δb x , b′ x ) = Δb x + b′ x , then b′ x = b x - Δb x , and further, the input signal on the X channel is the accurate input signal b′x The value of the general coherence parameter in the case of the difference parameter can be C gen (b′ x ) = C gen (b x - Δb x ).

[0201] Alternatively, the difference parameter can be used to adjust the metadata output based on the first audio data, so that after the metadata calculation is implemented based on inaccurate audio data, the inaccurate metadata is adjusted by the difference parameter, by the following formula:

[0202] Alternatively, the difference parameter can be directly applied to the calculation of the metadata, so that the calculation of the accurate metadata can be implemented based on the inaccurate audio data and the difference parameter, by the following formula:

[0203] For example, for the directional intensity i involved in the metadata calculation process, the directional intensity can be calculated by the formula , where b x , b y , b z , b w are the actual input audio on the X, Y, Z, W channels respectively. Assuming that Δb x , Δb y , Δb z , Δb w are the difference parameters of the actual input audio on the X, Y, Z, W channels respectively, b′ x , b′ y , b′ z , b′ w are the accurate input audio on the X, Y, Z, W channels respectively, and b x , b y , b z , b w can be calculated by the following formula respectively:

[0204] b′ x = b x * Δb x , b′ y = b y * Δb y , b′ z = b z * Δb z , b′ w = b w * Δbw Further, the value of the directional intensity i in the case that the input signal on the X, Y, Z, W channel is the accurate input signal can be

[0205] Optionally, the difference parameter can be directly applied to the adjustment of the input signal to adjust the accurate input signal b' x , b' y , b' z , b' w as the input for calculating the directional intensity i, so as to realize the direct adjustment of the first audio data through the following formula:

[0206] Optionally, the difference parameter can be directly applied to the calculation of the directional intensity i to achieve the purpose of adjusting the output metadata through the difference parameter through the following formula:

[0207] Further, the calculation on the three components of the directional intensity can be realized through the following formula to achieve the purpose of adjusting the output metadata through the difference parameter:

[0208] In the embodiments of the present disclosure, part or all of the steps, and optional implementation manners thereof, can be combined with part or all of the steps in other embodiments, or can be combined with optional implementation manners of other embodiments.

[0209] The embodiments of the present disclosure also propose a device for realizing any of the above methods, for example, a device is proposed, and the above device includes units or modules for realizing each step performed by the communication device (such as a terminal, a server, etc.) in any of the above methods.

[0210] It should be understood that the division of each unit or module in the above apparatus is only a logical function division, and all or part of them can be integrated into a physical entity or physically separated in actual implementation. In addition, the units or modules in the apparatus can be implemented in the form of processor calling software: for example, the apparatus includes a processor, the processor is connected with a memory, the memory stores instructions, and the processor calls the instructions stored in the memory to realize the functions of any of the above methods or the units or modules of the above apparatus, wherein the processor is a general processor such as a central processing unit (CPU) or a microprocessor, and the memory is a memory in the apparatus or a memory outside the apparatus. Alternatively, the units or modules in the apparatus can be implemented in the form of hardware circuit, and the functions of part or all of the units or modules can be realized by the design of the hardware circuit. The above hardware circuit can be understood as one or more processors; for example, in one implementation, the above hardware circuit is an application-specific integrated circuit (ASIC), and the functions of part or all of the units or modules are realized by the design of the logical relationship between the elements in the circuit; for another example, in another implementation, the above hardware circuit is a programmable logic device (PLD), and a field programmable gate array (FPGA) is taken as an example, which can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by a configuration file, so as to realize the functions of part or all of the units or modules. All units or modules of the above apparatus can be all implemented in the form of processor calling software, or all implemented in the form of hardware circuit, or part implemented in the form of processor calling software and the remaining part implemented in the form of hardware circuit.

[0211] In embodiments of the present disclosure, the processor is a circuit with signal processing capability. In one implementation, the processor can be a circuit with instruction reading and running capability, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), a digital signal processor (DSP), or the like. In another implementation, the processor can implement certain functions through a logical relationship of hardware circuits, and the logical relationship of the hardware circuits is fixed or reconfigurable. For example, the processor is a hardware circuit implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the processor loads a configuration document to implement the configuration of the hardware circuit. It can be understood that the processor loads instructions to implement the functions of some or all of the units or modules described above. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), or the like.

[0212] FIG. 3 is a structural schematic diagram of a communication device according to an embodiment of the present disclosure. As shown in FIG. 3, the communication device 3100 can include at least a processing module 3101. In some embodiments, the processing module 3101 is configured to determine a difference parameter between performance data of first audio data and expected performance data, and the processing module 3101 is further configured to calibrate metadata output based on the first audio data based on the difference parameter. Optionally, the processing module 3101 is configured to perform at least one of other steps (for example, steps S2101, S2102, S2103, and S2104, but not limited thereto) performed by the communication device in any of the methods described above. Details are not described herein again. In some embodiments, the communication device 3100 can further include a transceiver module. Optionally, the transceiver module is configured to perform at least one of the communication steps (such as sending and / or receiving) performed by the communication device in any of the methods described above. Details are not described herein again.

[0213] In some embodiments, the transceiving module can include a transmitting module and / or a receiving module, which can be separate or integrated together. Alternatively, the transceiving module can be mutually replaced with a transceiver.

[0214] In some embodiments, the processing module can be one module or include multiple sub-modules. Alternatively, the multiple sub-modules perform all or part of the steps required to be performed by the processing module respectively. Alternatively, the processing module can be mutually replaced with a processor.

[0215] FIG. 4A is a structural schematic diagram of a communication device 4100 according to an embodiment of the present disclosure. The communication device 4100 can be a server, a terminal (e.g., a user equipment, etc.), a chip, a chip system, or a processor supporting the server to implement any of the above methods, or a chip, a chip system, or a processor supporting the terminal to implement any of the above methods. The communication device 4100 can be used to implement the methods described in the above method embodiments, which can be referred to the descriptions in the above method embodiments.

[0216] As shown in FIG. 4A, the communication device 4100 includes one or more processors 4101. The processor 4101 can be a general-purpose processor or a special-purpose processor, for example, a baseband processor or a central processing unit. The baseband processor can be used to process communication protocols and communication data, and the central processing unit can be used to control the communication device (e.g., a base station, a baseband chip, a terminal device, a terminal device chip, a DU or a CU, etc.), execute programs, and process data of the programs. The communication device 4100 is used to implement any of the above methods.

[0217] In some embodiments, the communication device 4100 further includes one or more memories 4102 for storing instructions. Alternatively, all or part of the memory 4102 can also be outside the communication device 4100.

[0218] In some embodiments, the communication device 4100 further includes one or more transceivers 4103. When the communication device 4100 includes one or more transceivers 4103, the transceiver 4103 performs at least one of the communication steps (e.g., transmitting and / or receiving in the above methods) and the processor 4101 performs at least one of the other steps (e.g., steps S2101, S2102, S2103, S2104, but not limited to).

[0219] In some embodiments, the transceiver can include a receiver and / or a transmitter, which can be separate or integrated together. Optionally, the terms transceiver, transceiving unit, transceiver, transceiving circuit, etc. can be replaced by each other, the terms transmitter, transmitting unit, transmitter, transmitting circuit, etc. can be replaced by each other, and the terms receiver, receiving unit, receiver, receiving circuit, etc. can be replaced by each other.

[0220] In some embodiments, the communication device 4100 can include one or more interface circuits 4104. Optionally, the interface circuit 4104 is connected with the memory 4102, and the interface circuit 4104 can be used to receive signals from the memory 4102 or other devices, and can be used to send signals to the memory 4102 or other devices. For example, the interface circuit 4104 can read instructions stored in the memory 4102 and send the instructions to the processor 4101.

[0221] The communication device 4100 described in the above embodiments can be a network device or a terminal, but the scope of the communication device 4100 described in the present disclosure is not limited thereto, and the structure of the communication device 4100 can not be limited by FIG. 4A. The communication device can be a standalone device or can be part of a larger device. For example, the communication device can be: 1) a standalone integrated circuit (IC), or a chip, or a chip system or subsystem; (2) a set of one or more ICs, which can optionally also include storage components for storing data, programs; (3) an ASIC, such as a Modem; (4) a module that can be embedded in other devices; (5) a receiver, a terminal device, a smart terminal device, a cellular phone, a wireless device, a handset, a mobile unit, a vehicle-mounted device, a network device, a cloud device, an artificial intelligence device, etc.; (6) other devices, etc.

[0222] FIG. 4B is a structural schematic diagram of a chip 4200 according to an embodiment of the present disclosure. For the case where the communication device 4100 is a chip or a chip system, the structural schematic diagram of the chip 4200 shown in FIG. 4B can be referred to, but is not limited thereto.

[0223] The chip 4200 includes one or more processors 4201, and the chip 4200 is configured to execute any of the above methods.

[0224] In some embodiments, the chip 4200 further includes one or more interface circuits 4202. Optionally, the interface circuit 4202 is connected with the memory 4203, and the interface circuit 4202 can be used to receive signals from the memory 4203 or other devices, and can be used to send signals to the memory 4203 or other devices. For example, the interface circuit 4202 can read instructions stored in the memory 4203 and send the instructions to the processor 4201.

[0225] In some embodiments, the interface circuit 4202 performs at least one of the communication steps of transmission and / or reception in the above-described methods, and the processor 4201 performs at least one of the other steps (for example, steps S2101, S2102, S2103, S2104, but not limited thereto).

[0226] In some embodiments, the terms of interface circuit, interface, transceiver pin, transceiver, and the like can be replaced by each other.

[0227] In some embodiments, the chip 4200 further includes one or more memories 4203 for storing instructions. Optionally, all or part of the memories 4203 can be outside the chip 4200.

[0228] The disclosure also proposes a storage medium having instructions stored thereon, which, when executed on the communication device 4100, causes the communication device 4100 to perform any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but is not limited thereto, and it can also be a storage medium readable by other devices. Optionally, the storage medium can be a non-transitory storage medium, but is not limited thereto, and it can also be a transitory storage medium.

[0229] The disclosure also proposes a program product, which, when executed by the communication device 4100, causes the communication device 4100 to perform any of the above methods. Optionally, the program product is a computer program product.

[0230] The disclosure also proposes a computer program, which, when executed on a computer, causes the computer to perform any of the above methods.

[0231] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. The disclosure is intended to cover any variations, uses or adaptations of the disclosure following, in general, the principles of the disclosure and including such departures from the present disclosure as come within known or customary practice in the art to which the disclosure pertains. The specification and examples are to be regarded as illustrative only, and the true scope and spirit of the disclosure are indicated by the following claims.

[0232] It should be understood that the present disclosure is not limited to the precise structures herein described and illustrated in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is indicated by the appended claims, and their equivalents.

Claims

1. A metadata calibration method of audio data, characterized by, The method comprises: determining a difference parameter between performance data of first audio data and expected performance data; calibrating metadata output based on the first audio data based on the difference parameter.

2. The method of claim 1, wherein, The difference parameter between the performance data of the first audio data and the expected performance data comprises at least one of: a difference between the performance data of the first audio data and the expected performance data; a proportional relationship between the performance data of the first audio data and the expected performance data.

3. The method according to claim 1 or 2, characterized in that, The first audio data comprises audio signals on a plurality of input channels; The determining of the difference parameter between the performance data of the first audio data and the expected performance data comprises: determining a difference parameter between the performance data of the audio signals on each input channel and the expected performance data, respectively, to obtain the difference parameters corresponding to the plurality of input channels, respectively.

4. The method according to any one of claims 1 to 3, characterized in that, The method further comprises: obtaining the performance data of the first audio data.

5. The method of claim 4, wherein, The obtaining of the performance data of the first audio data comprises any one of: in response to obtaining the first audio data, performing audio processing based on the first audio data to obtain the performance data of the first audio data; obtaining the performance data of the first audio data based on pre-stored performance data corresponding to the first audio data; determining the performance data of the first audio data based on device parameters of a device used to collect the first audio data.

6. The method of claim 5, wherein, The device parameters comprise at least one of: directivity; frequency response curve parameters; sensitivity; delay parameters; distortion parameters; noise.

7. The method according to any one of claims 1 to 6, characterized in that, The expected performance data is agreed upon by a protocol, or the expected performance data is self-defined.

8. The method according to any one of claims 1 to 7, characterized in that, Different types of audio data correspond to different expected performance data; or Different audio data processing scenarios correspond to different expected performance data.

9. The method according to any one of claims 1 to 8, characterized in that, The first audio data comprises audio signals on a plurality of input channels, and each input channel corresponds to a difference parameter; The calibrating of the metadata output based on the first audio data based on the difference parameter comprises: calibrating metadata output based on the audio signals on the plurality of input channels based on the difference parameter corresponding to each input channel.

10. The method according to any one of claims 1 to 9, characterized in that, The calibrating of the metadata output based on the first audio data based on the difference parameter comprises any one of: calibrating the first audio data based on the difference parameter, to calculate based on the calibrated first audio data to obtain calibrated metadata; calculating based on the difference parameter and the first audio data to obtain calibrated metadata; correcting metadata calculated based on the first audio data based on the difference parameter to obtain calibrated metadata; calculating based on the first audio data, and processing metadata calculated based on the first audio data based on first processing and the difference parameter.

11. The method according to any one of claims 1 to 10, characterized in that, The metadata comprises at least one of: direction of arrival (DOA); direct sound to total energy ratio; diffusion coefficient; distance; diffuse sound to total energy ratio; surrounding coefficient; residual sound to total energy ratio.

12. A communication device, characterized by comprises: a processing module configured to determine a difference parameter between performance data of the first audio data and expected performance data; the processing module is further configured to calibrate metadata output based on the first audio data based on the difference parameter.

13. A communication device, characterized by comprise: one or more processors; wherein the communication device is configured to perform the metadata calibration method of any one of claims 1-11.

14. A storage medium, the storage medium storing instructions, wherein, the instructions, when executed on the communication device, cause the communication device to perform the metadata calibration method of any one of claims 1-11.