Processing method and apparatus, and storage medium

WO2025160827A1PCT designated stage Publication Date: 2025-08-07BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/075042
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-31
Publication Date
2025-08-07

Smart Images

  • Figure CN2024075042_07082025_PF_FP_ABST
    Figure CN2024075042_07082025_PF_FP_ABST
Patent Text Reader

Abstract

A processing method and apparatus, and a storage medium. The processing method is executed by an encoder, and comprises: downmixing sound channels in at least one sound channel group to generate sound channel information of each sound channel group, wherein the at least one sound channel group includes N first sound channel groups and M second sound channel groups, each first sound channel group comprises three sound channels, each second sound channel group comprises two sound channels, N is 1, and M is a non-negative integer. The method solves the problem caused by generating corresponding sound channel information after downmixing sound channels. The present invention provides a solution for generating sound channel information of each sound channel group, which ensures that each sound channel group has corresponding sound channel information and ensures the accuracy of generated sound channel groups, thereby ensuring the reliability of encoding and decoding.
Need to check novelty before this filing date? Find Prior Art

Description

Processing method, device and storage medium Technical Field

[0001] The present disclosure relates to the field of communication technologies, and in particular to a processing method, device, and storage medium. Background Art

[0002] With the rapid development of multimedia technology, audio data can be encoded and transmitted, and then the encoded data can be decoded to obtain the corresponding audio data, ensuring that the audio data can be transmitted efficiently.

[0003] Summary of the Invention

[0004] The present disclosure solves the problem of less information in audio data by generating first data for describing the audio data, thereby ensuring the accuracy of the description of the audio data and further ensuring the accuracy of subsequent machine analysis or machine training of the audio data.

[0005] The embodiments of the present disclosure provide a processing method, an apparatus, and a storage medium.

[0006] According to a first aspect of an embodiment of the present disclosure, a processing method is provided. The method is performed by a first device, and the method includes:

[0007] A plurality of first data are generated based on the audio data, where the first data are used to describe the audio data and are also used to perform machine analysis or machine training on the audio data.

[0008] According to a second aspect of an embodiment of the present disclosure, a processing method is proposed. The method is performed by an encoder, and the method includes:

[0009] At least one of the first data and the audio data is encoded to obtain encoded data, where the first data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data.

[0010] According to a third aspect of an embodiment of the present disclosure, a processing method is proposed, where the method is executed by a decoder, and the method includes:

[0011] The encoded data is decoded to obtain at least one of second data and audio data, where the second data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data.

[0012] According to a fourth aspect of the embodiments of the present disclosure, a processing method is proposed, the method comprising:

[0013] The first device generates a plurality of first data based on the audio data, where the first data is used to describe the audio data and is further used to perform machine analysis or machine training on the audio data;

[0014] The encoder encodes at least one of the first data and the audio data to obtain encoded data;

[0015] The decoder decodes the encoded data to obtain at least one of second data and audio data, where the second data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data.

[0016] According to a fifth aspect of an embodiment of the present disclosure, a processing device is provided, including:

[0017] The processing module is used to generate a plurality of first data based on the audio data, where the first data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data.

[0018] According to a sixth aspect of an embodiment of the present disclosure, a processing device is provided, including:

[0019] The processing module is used to encode at least one of the first data and the audio data to obtain encoded data, wherein the first data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data.

[0020] According to a seventh aspect of the embodiments of the present disclosure, a processing device is provided, including:

[0021] The processing module is used to decode the encoded data to obtain at least one of second data and audio data, where the second data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data.

[0022] According to an eighth aspect of the embodiments of the present disclosure, a processing device is provided, including:

[0023] one or more processors;

[0024] Wherein, the processing device is used to execute any method described in the first aspect.

[0025] According to a ninth aspect of the embodiments of the present disclosure, a processing device is provided, including:

[0026] one or more processors;

[0027] Wherein, the processing device is used to execute any method described in the second aspect.

[0028] According to a tenth aspect of an embodiment of the present disclosure, a processing device is provided, including:

[0029] one or more processors;

[0030] Wherein, the processing device is used to execute any method described in the third aspect.

[0031] According to an eleventh aspect of the present disclosure, a coding and decoding system is proposed, including:

[0032] A first device, an encoder and a decoder, wherein the first device is configured to implement the processing method described in the first aspect, the encoder is configured to implement the processing method described in the second aspect, and the decoder is configured to implement the processing method described in the third aspect.

[0033] According to the twelfth aspect of an embodiment of the present disclosure, a storage medium is proposed, which stores instructions. When the instructions are executed on a communication device, the communication device executes a method as described in the first aspect or the second aspect or any one of the second aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The drawings described herein are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the present disclosure. The illustrative embodiments of the embodiments of the present disclosure and their descriptions are used to explain the embodiments of the present disclosure and do not constitute an improper limitation on the embodiments of the present disclosure. In the drawings:

[0035] FIG1 is a schematic diagram of the architecture of a coding and decoding system according to an embodiment of the present disclosure;

[0036] FIG2A is an interactive schematic diagram illustrating a processing method according to an embodiment of the present disclosure;

[0037] FIG2B is an interactive schematic diagram illustrating a processing method according to an embodiment of the present disclosure;

[0038] FIG2C is a schematic diagram showing first data according to an embodiment of the present disclosure;

[0039] FIG2D is a schematic diagram showing first data according to an embodiment of the present disclosure;

[0040] FIG3A is a schematic flow chart of a processing method according to an embodiment of the present disclosure;

[0041] FIG3B is a flow chart of a processing method according to an embodiment of the present disclosure;

[0042] FIG4A is a schematic flow chart of a processing method according to an embodiment of the present disclosure;

[0043] FIG4B is a flow chart of a processing method according to an embodiment of the present disclosure;

[0044] FIG5 is a flow chart of a processing method according to an embodiment of the present disclosure;

[0045] FIG6 is a flow chart of a processing method according to an embodiment of the present disclosure;

[0046] FIG7 is a flow chart of a processing method according to an embodiment of the present disclosure;

[0047] FIG8A is a schematic structural diagram of a processing device proposed in an embodiment of the present disclosure;

[0048] FIG8B is a schematic structural diagram of a processing device proposed in an embodiment of the present disclosure;

[0049] FIG8C is a schematic diagram of the structure of a processing device proposed in an embodiment of the present disclosure;

[0050] FIG9A is a schematic structural diagram of a communication device proposed in an embodiment of the present disclosure;

[0051] FIG9B is a schematic diagram of the structure of the chip proposed in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0052] The present disclosure provides a processing method, an apparatus, and a storage medium.

[0053] According to a first aspect of an embodiment of the present disclosure, a processing method is provided. The method is performed by a first device, and the method includes:

[0054] A plurality of first data are generated based on the audio data, where the first data are used to describe the audio data and are also used to perform machine analysis or machine training on the audio data.

[0055] In the above embodiment, first data for describing audio data can be generated, which solves the problem of less information in audio data. By generating first data for describing audio data, the accuracy of the description of the audio data is guaranteed, thereby ensuring the accuracy of subsequent machine analysis or machine training of the audio data.

[0056] In combination with some embodiments of the first aspect, in some embodiments, at least one of the audio data or the first data is further used by an encoder for encoding and transmission.

[0057] In the above embodiment, at least one of the audio data or the first data needs to be encoded by an encoder, and the encoded data is transmitted, so as to save transmission resources and improve transmission efficiency.

[0058] In combination with some embodiments of the first aspect, in some embodiments, the first data is encoded using Huffman coding.

[0059] In the above embodiment, Huffman coding may be used to encode the first data, so as to facilitate transmission of the encoded first data, save transmission resources, and improve transmission efficiency.

[0060] In combination with some embodiments of the first aspect, in some embodiments, each first data corresponds to a channel, and the channel includes the audio data corresponding to the first data.

[0061] In combination with some embodiments of the first aspect, in some embodiments, the first data includes reference information, where the reference information is used to indicate reference audio data corresponding to the first data for matching with the audio data to be tested.

[0062] In the above embodiment, the first data indicates the role of the audio data corresponding to the first data through the included reference information, that is, the audio data corresponding to the first data is reference audio data, which is used to match the audio data to be tested to determine whether the audio data to be tested meets the requirements, thereby improving the accuracy of indicating the type of audio data.

[0063] In combination with some embodiments of the first aspect, in some embodiments, the first data belongs to a first group, the first group includes multiple first data, and the first data included in the first group all include the reference information.

[0064] In the above embodiment, the same first data may be grouped, and the first data in the same group all indicate that the corresponding audio data is reference audio data, thereby ensuring the accuracy of the grouping and further ensuring the accuracy of the type of the indicated audio data.

[0065] In combination with some embodiments of the first aspect, in some embodiments, the first data includes information to be tested, and the information to be tested is used to indicate that audio data to be tested corresponding to the first data is used to match with reference audio data.

[0066] In the above embodiment, the first data indicates the role of the audio data corresponding to the first data by including the information to be tested, that is, the audio data corresponding to the first data is the audio data to be tested, which is used to match with the reference audio data to determine whether the audio data to be tested meets the requirements, thereby improving the accuracy of indicating the type of audio data.

[0067] In combination with some embodiments of the first aspect, in some embodiments, the first data belongs to a second group, the second group includes multiple first data, and the first data included in the second group all include the information to be tested.

[0068] In the above embodiment, the same first data may be grouped, and the first data in the same group all indicate that the corresponding audio data is the audio data to be tested, thereby ensuring the accuracy of the grouping and further ensuring the accuracy of the type of the indicated audio data.

[0069] In combination with some embodiments of the first aspect, in some embodiments, the first data includes position information, and the position information is used to indicate a collection position of audio data corresponding to the first data.

[0070] In the above embodiment, the first data indicates the collection location of the corresponding audio data by including location information, which can accurately point out the location when the audio data was collected, ensuring that subsequent machine analysis of the audio data can be based on the location information analysis, ensuring the comprehensiveness of the description of the audio data, and thus ensuring the accuracy of the analyzed audio data.

[0071] In combination with some embodiments of the first aspect, in some embodiments, the first data includes environmental information, and the environmental information is used to indicate the environment in which the audio data corresponding to the first data is collected.

[0072] In the above embodiment, the first data indicates the collection environment of the corresponding audio data by including environmental information, which can accurately point out the environment in which the audio data was collected, ensuring that subsequent machine analysis of the audio data can be based on the environmental information analysis, ensuring the comprehensiveness of the description of the audio data, and thus ensuring the accuracy of the analyzed audio data.

[0073] In conjunction with some embodiments of the first aspect, in some embodiments, the environmental information includes at least one of the following:

[0074] Temperature information;

[0075] Humidity information;

[0076] Air pressure information.

[0077] In combination with some embodiments of the first aspect, in some embodiments, the first data includes any one of a character string, a numerical value, or a combination of the character string and the numerical value.

[0078] In combination with some embodiments of the first aspect, in some embodiments, the first data is carried in a MAE (Metadata Audio Elements) extension domain.

[0079] In a second aspect, an embodiment of the present disclosure provides a processing method, which is performed by an encoder and includes:

[0080] At least one of the first data and the audio data is encoded to obtain encoded data, where the first data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data.

[0081] In conjunction with some embodiments of the second aspect, in some embodiments, encoding at least one of the first data and the audio data to obtain the encoded data includes:

[0082] The first data is encoded using Huffman encoding to obtain the encoded data.

[0083] In conjunction with some embodiments of the second aspect, in some embodiments, encoding at least one of the first data and the audio data to obtain the encoded data includes:

[0084] The audio data is encoded based on the first data to obtain the encoded data.

[0085] In conjunction with some embodiments of the second aspect, in some embodiments, encoding at least one of the first data and the audio data to obtain the encoded data includes:

[0086] Encode the first data of the character type using a character encoding method to obtain the encoded data; or,

[0087] The first data of the character type is encoded in a one-to-one mapping manner from characters to numerical values ​​to obtain the encoded data.

[0088] In combination with some embodiments of the second aspect, in some embodiments, the first data is carried in the MAE extension domain.

[0089] In a third aspect, an embodiment of the present disclosure provides a processing method, which is executed by a decoder and includes:

[0090] The encoded data is decoded to obtain at least one of second data and audio data, where the second data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data.

[0091] In conjunction with some embodiments of the third aspect, in some embodiments, decoding the encoded data to obtain at least one of the second data and the audio data includes:

[0092] The coded data is encoded using Huffman coding to obtain the second data.

[0093] In conjunction with some embodiments of the third aspect, in some embodiments, decoding the encoded data to obtain at least one of the second data and the audio data includes:

[0094] Decoding the encoded data in a character encoding manner to obtain the second data; or

[0095] The encoded data is decoded using a one-to-one mapping method from characters to numerical values ​​to obtain the second data.

[0096] In a fourth aspect, an embodiment of the present disclosure provides a processing method, the method comprising:

[0097] The first device generates a plurality of first data based on the audio data, where the first data is used to describe the audio data and is further used to perform machine analysis or machine training on the audio data;

[0098] The encoder encodes at least one of the first data and the audio data to obtain encoded data;

[0099] The decoder decodes the encoded data to obtain at least one of second data and audio data, where the second data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data.

[0100] In a fifth aspect, an embodiment of the present disclosure provides a processing device, which includes at least one of a transceiver module and a processing module; wherein the first device is used to execute the optional implementation methods of the first and fourth aspects.

[0101] In a sixth aspect, an embodiment of the present disclosure provides a processing device, which includes at least one of a transceiver module and a processing module; wherein the encoder is used to execute the optional implementation methods of the second aspect and the fourth aspect.

[0102] In the seventh aspect, an embodiment of the present disclosure provides a processing device, which includes at least one of a transceiver module and a processing module; wherein the decoder is used to execute the optional implementation methods of the third aspect and the fourth aspect.

[0103] In an eighth aspect, an embodiment of the present disclosure provides a processing device, including:

[0104] one or more processors;

[0105] The processing device is used to execute the method described in any one of the first and fourth aspects.

[0106] In a ninth aspect, an embodiment of the present disclosure provides a processing device, including:

[0107] one or more processors;

[0108] The processing device is used to execute the method described in any one of the second aspect and the fourth aspect.

[0109] In a tenth aspect, an embodiment of the present disclosure provides a processing device, including:

[0110] one or more processors;

[0111] Wherein, the processing device is used to execute the method described in any one of the third aspect and the fourth aspect.

[0112] In the eleventh aspect, an embodiment of the present disclosure provides a storage medium storing first information. When the first information is run on a communication device, the communication device executes the method as described in any one of the first, second and third aspects.

[0113] In a twelfth aspect, an embodiment of the present disclosure proposes a program product. When the program product is executed by a communication device, the communication device executes any one of the methods described in the first aspect, the second aspect, and the third aspect.

[0114] In a thirteenth aspect, an embodiment of the present disclosure proposes a computer program, which, when executed on a communication device, enables the communication device to execute any one of the methods described in the first, second, and third aspects.

[0115] In a fourteenth aspect, an embodiment of the present disclosure provides a chip or a chip system, wherein the chip or chip system includes a processing circuit configured to execute any one of the methods described in the first, second, and third aspects.

[0116] It is understandable that the above-mentioned encoder, storage medium, program product, computer program, chip or chip system are all used to execute the method proposed in the embodiment of the present disclosure. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding method and will not be repeated here.

[0117] The present disclosure provides processing methods, devices, and storage media. In some embodiments, the terms "processing method" and "information processing method" are interchangeable, "processing device" and "information processing device" are interchangeable, and "indicating device" and "information processing system" are interchangeable.

[0118] The embodiments of the present disclosure are not exhaustive and are merely illustrative of some embodiments, and are not intended to be a specific limitation on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment can be arbitrarily exchanged. In addition, the optional implementation methods in a certain embodiment can be arbitrarily combined; in addition, the embodiments can be arbitrarily combined. For example, some or all steps of different embodiments can be arbitrarily combined, and a certain embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.

[0119] In each embodiment of the present disclosure, unless otherwise specified or provided for by logic, the terms and / or descriptions between the embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form a new embodiment based on their inherent logical relationships.

[0120] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure.

[0121] In the embodiments of the present disclosure, unless otherwise specified, elements expressed in the singular, such as "a", "an", "the", "above", "said", "the", "the", etc., may mean "one and only one", or "one or more", "at least one", etc. For example, when using articles such as "a", "an", "the" in English in translation, the noun following the article may be understood as a singular expression or a plural expression.

[0122] In the embodiments of the present disclosure, “plurality” refers to two or more.

[0123] In some embodiments, the terms "at least one," "one or more," "a plurality of," "multiple," etc. can be used interchangeably.

[0124] In some embodiments, descriptions such as "at least one of A and B," "A and / or B," "A in one case, B in another case," or "in response to one case A, in response to another case B" may include the following technical solutions depending on the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed); and in some embodiments, A and B (both A and B are executed). The above is also applicable when there are more branches such as A, B, and C.

[0125] In some embodiments, "A or B" and other descriptions may include the following technical solutions depending on the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed). The above is also applicable when there are more branches such as A, B, C, etc.

[0126] The prefixes such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different description objects and do not constitute any restriction on the position, order, priority, quantity or content of the description objects. For the statement of the description object, please refer to the description in the context of the claims or embodiments, and no unnecessary restriction should be constituted due to the use of prefixes. For example, if the description object is a "field", the ordinal number before the "field" in the "first field" and the "second field" does not limit the position or order between the "fields". "First" and "second" do not limit whether the "fields" they modify are in the same message, nor do they limit the order of the "first field" and the "second field". For another example, if the description object is a "level", the ordinal number before the "level" in the "first level" and the "second level" does not limit the priority between the "levels". For another example, the number of description objects is not limited by the ordinal number and can be one or more. Taking "first device" as an example, the number of "devices" can be one or more. In addition, the objects modified by different prefixes can be the same or different. For example, if the description object is "device", then the "first device" and the "second device" can be the same device or different devices, and their types can be the same or different; for another example, if the description object is "information", then the "first information" and the "second information" can be the same information or different information, and their contents can be the same or different.

[0127] In some embodiments, “including A,” “comprising A,” “used to indicate A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.

[0128] In some embodiments, terms such as "time / frequency" and "time / frequency domain" refer to the time domain and / or the frequency domain.

[0129] In some embodiments, terms such as "in response to...", "in response to determining...", "in the case of...", "at the time of...", "when...", "if...", "if...", etc. can be used interchangeably.

[0130] In some embodiments, terms such as "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not less than", and "above" can be replaced with each other, and terms such as "less than", "less than or equal to", "not greater than", "less than", "less than or equal to", "not more than", "lower than", "lower than or equal to", "not higher than", and "below" can be replaced with each other.

[0131] In some embodiments, devices and equipment can be interpreted as physical or virtual, and their names are not limited to the names recorded in the embodiments. In some cases, they can also be understood as "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "subject", etc.

[0132] In some embodiments, "network" can be interpreted as devices included in the network, such as access network equipment, core network equipment, etc.

[0133] In some embodiments, "access network device (AN device)" may also be referred to as "radio access network device (RAN device)", "base station (BS)", "radio base station", "fixed station", and in some embodiments may also be understood as "node", "access point", "transmission point (TP)", "reception point (RP)", "transmission and / or reception point (TRP)" "panel", "antenna panel", "antenna array", "cell", "macro cell", "small cell", "femto cell", "pico cell", "sector", "cell group", "serving cell", "carrier", "component carrier", "bandwidth part (BWP)", etc.

[0134] In some embodiments, "encoder (terminal)" or "encoder device (terminal device)" may be referred to as "user equipment (encoder)", "user terminal" "mobile station (MS)", "mobile terminal (MT)", subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, etc.

[0135] In some embodiments, obtaining data, information, etc. may comply with the laws and regulations of the country where the data is obtained.

[0136] In some embodiments, data, information, etc. may be obtained with the user's consent.

[0137] In addition, each element, each row, or each column in the table of the embodiment of the present disclosure can be implemented as an independent embodiment, and the combination of any elements, any rows, and any columns can also be implemented as an independent embodiment.

[0138] FIG1 is a schematic diagram of the architecture of a codec system according to an embodiment of the present disclosure. As shown in FIG1 , the method provided in the embodiment of the present disclosure can be applied to a codec system 100, which can include an encoder 101 and a decoder 102. It should be noted that the codec system 100 can also include other devices, and the present disclosure does not limit the devices included in the codec system 100.

[0139] In some embodiments, the encoder 101 and the decoder 102 are both provided in a terminal. In some embodiments, the terminal can be various devices. For example, the terminal can be a mobile phone, a wearable device, an Internet of Things device, a car with communication functions, a smart car, a tablet computer, a computer with wireless transceiver functions, a virtual reality (VR) encoder device, an augmented reality (AR) encoder device, a wireless encoder device in industrial control, a wireless encoder device in self-driving, a wireless encoder device in remote medical surgery, a wireless encoder device in a smart grid, a wireless encoder device in transportation safety, a wireless encoder device in a smart city, and a wireless encoder device in a smart home, but is not limited thereto.

[0140] It can be understood that the encoding and decoding system described in the embodiment of the present disclosure is for the purpose of more clearly illustrating the technical solution of the embodiment of the present disclosure, and does not constitute a limitation on the technical solution proposed in the embodiment of the present disclosure. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solution proposed in the embodiment of the present disclosure is also applicable to similar technical problems.

[0141] The following embodiments of the present disclosure may be applied to the codec system 100 shown in FIG1 , or a portion thereof, but are not limited thereto. The entities shown in FIG1 are illustrative only. The codec system may include all or part of the entities shown in FIG1 , or may include other entities outside of FIG1 . The number and form of the entities are arbitrary, and the entities may be physical or virtual. The connection relationship between the entities is illustrative only. The entities may be connected or disconnected, and the connection may be in any manner, including direct or indirect, wired or wireless.

[0142] The embodiments of the present disclosure can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G New Radio (NR), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New Radio Access (NX), Future Generation Radio Access (FX), Global System for Mobile Communications (GSM (registered trademark)), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, Ultra-WideBand (UWB), Bluetooth (registered trademark), Public Land Mobile Network (PLMN) networks, Device-to-Device (D2D) systems, Machine-to-Machine (M2M) systems, Internet of Things (IoT) systems, Vehicle-to-Everything (V2X), systems utilizing other processing methods, and next-generation systems based on and extending these systems. Furthermore, a combination of multiple systems (e.g., a combination of LTE or LTE-A with 5G) may also be employed.

[0143] In some embodiments, the present disclosure further includes a first device, or may also be referred to as an acquisition device. Optionally, the first device is used to acquire audio data and generate first data of the audio data, and the audio data is described by the first data. Optionally, the first device is an audio sensor device.

[0144] FIG2A is an interactive diagram of a processing method according to an embodiment of the present disclosure. As shown in FIG2A , the embodiment of the present disclosure relates to a processing method, which includes:

[0145] Step S2101: The first device collects audio data.

[0146] In some embodiments, the first device refers to a device with a collection function. For example, the first device may also be referred to as a collection device, a collection terminal, etc., which is not limited in the embodiments of the present disclosure. In some embodiments, the audio data refers to audio, or it can also be understood as sound collected by the first device. In some embodiments, the audio data is used to indicate the audio generated by the device during operation.

[0147] In some embodiments, the audio data is used for machine analysis or machine training. Alternatively, it can be understood that machine analysis of the audio data can determine whether the operating status of the device indicated by the audio data is normal. Alternatively, it can be understood that machine training of the audio data can be used to predict the data based on the machine training. Alternatively, it can be understood that the audio data indicates the operating status of the device, and whether the device is operating normally can be determined based on the audio data.

[0148] In some embodiments, the embodiments of the present disclosure are applied to automated production and testing scenarios, where the first device collects audio data of the produced products (components), and then performs machine analysis on the collected audio data to determine whether the generated products (components) are qualified.

[0149] It should be noted that the present disclosure uses audio data as an example for illustration. Furthermore, after the decoder decodes and obtains audio data in the following embodiments, the second device analyzes the audio data obtained by the decoder to determine whether the device is operating normally.

[0150] In some embodiments, the embodiments of the present disclosure are applied in product testing, see Figure 2B, where audio data of the product is collected by a first device, the audio data is encoded using an encoder, the encoded data is sent to a decoder, and the decoder decodes the encoded data to determine whether the product corresponding to each piece of audio data is qualified.

[0151] Step S2102: The first device generates a plurality of first data based on the audio data.

[0152] In some embodiments, the first data is used to describe the audio data. In some embodiments, the first data is used to describe the audio data means that the first data is used to interpret the audio data. Alternatively, it can also be understood that the first data includes information related to the audio data.

[0153] In some embodiments, the first data can also be understood as descriptive information of the audio data. In some embodiments, the name of the first data is not limited, and it can be, for example, metadata of the audio data, or other data. In some embodiments, different audio data correspond to different first data, or it can also be understood that the audio data and the first data have a one-to-one correspondence. In other words, the first data is used to describe the audio data corresponding to the first data.

[0154] In some embodiments, at least one of the audio data or the first data is also used by an encoder to encode and transmit. In an embodiment of the present disclosure, at least one of the audio data or the first data can be encoded by an encoder, and the encoder transmits the encoded data to another device, such as a decoder. In some embodiments, the first data is encoded using Huffman coding. In other words, the encoder can encode the first data using Huffman coding.

[0155] In some embodiments, each first data item corresponds to a channel, and the channel includes audio data corresponding to the first data item. In some embodiments, the channel refers to a channel for transmitting data in the communication system. Alternatively, the channel can be understood as a medium for transmitting audio data. For example, the communication system includes three channels, each channel corresponding to audio data and first data.

[0156] In some embodiments, the information included in the first data is also different, and different types of information have different functions. The information included in the first data is described below.

[0157] In some embodiments, the first data includes reference information, and the reference information is used to indicate that the reference audio data corresponding to the first data is used to match the audio data to be tested. In the embodiment of the present disclosure, the reference information means that the audio data corresponding to the first data is reference audio data. Optionally, the reference audio data can also be understood as standard audio data, and the device indicated by the reference audio data is normal, that is, the audio data generated when the working state of the device is normal. Optionally, by matching the audio data to be tested with the reference audio data, it can be determined whether the audio data to be tested matches the reference audio data, and then it can be determined whether the audio data to be tested is normal.

[0158] In some embodiments, if the first data corresponding to the channel includes reference information, it is indicated that the audio data included in the channel is reference audio data. Optionally, the channel has corresponding first data, and the channel includes reference audio data. In some embodiments, the reference audio data and the audio data to be measured are determined by similarity to determine whether the similarity between the reference audio data and the audio data to be measured exceeds a similarity threshold, thereby determining whether the reference audio data and the audio data to be measured are similar, and then determining whether the audio data to be measured meets the requirements. In some embodiments, the reference audio data and the audio data to be measured are determined by determining a difference to determine whether the reference audio data and the audio data to be measured are similar. If the difference is greater than the difference threshold, it is indicated that the reference audio data and the audio data to be measured are not similar. If the difference is not greater than the difference threshold, it is indicated that the reference audio data and the audio data to be measured are similar.

[0159] In some embodiments, the first data belongs to a first group, the first group includes multiple first data, and the first data included in the first group all include reference information. In an embodiment of the present disclosure, the first group includes multiple first data, and each first data corresponds to a channel, that is, the first group includes multiple channels, each channel corresponds to a first data, and each channel includes an audio data. In some embodiments, the first data included in the first group all indicate that the audio data is reference audio data, that is, the audio data in the channels included in the first group are all reference audio data.

[0160] In some embodiments, the first data includes information to be tested, and the information to be tested is used to indicate that the audio data to be tested corresponding to the first data is used to match the reference audio data. In the embodiment of the present disclosure, the information to be tested means that the audio data corresponding to the first data is the audio data to be tested. Optionally, the audio data to be tested can also be understood as audio data that needs to be analyzed and tested. Optionally, by matching the audio data to be tested with the reference audio data, it can be determined whether the audio data to be tested matches the reference audio data, and then it can be determined whether the audio data to be tested is normal.

[0161] In some embodiments, if the first data corresponding to the channel includes the information to be tested, it is indicated that the audio data included in the channel is the audio data to be tested. Optionally, the channel has the corresponding first data, and the channel includes the audio data to be tested. In some embodiments, the reference audio data and the audio data to be tested are determined by means of similarity to determine whether the similarity between the reference audio data and the audio data to be tested exceeds a similarity threshold, thereby determining whether the reference audio data and the audio data to be tested are similar, and then determining whether the audio data to be tested meets the requirements. In some embodiments, the reference audio data and the audio data to be tested are determined by means of determining a difference to determine whether the reference audio data and the audio data to be tested are similar. If the difference is greater than the difference threshold, it is indicated that the reference audio data and the audio data to be tested are not similar. If the difference is not greater than the difference threshold, it is indicated that the reference audio data and the audio data to be tested are similar.

[0162] In some embodiments, referring to FIG2C , the first data of the first audio data indicates that the first audio data is reference audio data, the first data of the second audio data indicates that the second audio data and the first data of the third audio data indicate that the third audio data are audio data to be tested, and then the second audio data and the third audio data are identified by an identifier to determine that the product corresponding to the second audio data is qualified and the product corresponding to the third audio data is unqualified.

[0163] In some embodiments, the first data belongs to the second group, the second group includes multiple first data, and the first data included in the second group all include information to be tested. In an embodiment of the present disclosure, the second group includes multiple first data, and each first data corresponds to a channel, that is, the second group includes multiple channels, each channel corresponds to a first data, and each channel includes an audio data. In some embodiments, the first data included in the second group all indicate that the audio data is audio data to be tested, that is, the audio data in the channels included in the second group are all audio data to be tested.

[0164] It should be noted that in the embodiments of the present disclosure involving a first group and a second group, the first group and the second group include the same amount of first data, and there is a one-to-one correspondence between the first data included in the first group and the first data included in the second group. In some embodiments, the first group includes first data 1, first data 2, and first data 3, and the second group includes first data 4, first data 5, and first data 6. Then, first data 1 corresponds to first data 4, first data 2 corresponds to first data 5, and first data 3 corresponds to first data 6.

[0165] In some embodiments, the first data includes any one of a string, a numerical value, or a combination of a string and a numerical value. Optionally, if the first data includes a string. Wherein, the information included in the first data is reference information, and the reference information is represented by groundtruth audio. If the information included in the first data is information to be tested, the information to be tested is represented by product audio to be tested. For example, channels include channel 1, channel 2...channel N, where N is a positive integer, and the corresponding first data is groundtruth audio, product1 audio to be tested...productN audio to be tested. Optionally, if the first data includes a numerical value. The reference information is represented by 0, that is, 0 corresponds to groundtruth audio. If the information included in the first data is information to be tested, the information to be tested is represented by 1, that is, 1 represents product audio to be tested. For example, channels include channel 1, channel 2...channel N, where N is a positive integer, and the corresponding first data is 0: groundtruth audio, 1: product audio to be tested...1: product audio to be tested. Optionally, if the first data includes a combination of strings and numerical values. The reference information is represented by groundtruth audio (real audio). If the information included in the first data is the information to be tested, the information to be tested is represented by product audio to be tested (test product audio). In addition, the reference information can also be represented by 0, that is, 0 corresponds to groundtruth audio (real audio). If the information included in the first data is the information to be tested, the information to be tested can also be represented by 1, that is, 1 represents product audio to be tested (test product audio). For example, the channels include channel 1, channel 2...channel N, N is a positive integer, and the corresponding first data are "groundtruth audio", "product1 audio" and "0:groundtruth audio", "1:product audio to be tested".

[0166] In some embodiments, the first data includes location information, and the location information is used to indicate the collection location of the audio data corresponding to the first data. In the embodiment of the present disclosure, when collecting audio data, the first device collects audio data at different locations. Therefore, the location information can be included in the first data to indicate the location when the audio data was collected. Optionally, the location information can be represented by a rectangular coordinate system (x, y, z). Optionally, the location information can be represented by a polar coordinate system (r, θ, φ).

[0167] In some embodiments, referring to Figure 2D, the first data of the audio data indicates that the first audio data 1, the second audio data 2, and the third audio data 3 are reference audio data, and their positions are (x1, y1, z1), (x2, y2, z2), and (x3, y3, z3), respectively. The first data of the audio data indicates that the fourth audio data 4, the fifth audio data 5, and the sixth audio data 6 are audio data to be tested, and their positions are (x4, y4, z4), (x5, y5, z5), and (x6, y6, z6), respectively. Among them, (x1, y1, z1) and (x4, y4, z4) are in the same position, (x2, y2, z2) and (x5, y5, z5) are in the same position, and (x3, y3, z3) and (x6, y6, z6) are in the same position, and then the audio data to be tested are identified as qualified by the identifier.

[0168] In some embodiments, combined with the above-mentioned first grouping and second grouping embodiments, it can be explained that the position information of the first data in the first group is the same as the corresponding first data in the second group, that is, the position information of the one-to-one corresponding first data in the first group and the second group is the same.

[0169] In some embodiments, the first data includes any one of a string, a numerical value, or a combination of a string and a numerical value. Optionally, if the first data includes a string, the first data can be represented as position1, x1, y1, z1. Optionally, if the first data includes a numerical value, the first data can be represented as x1=0, y1=0, z1=0.2m. Optionally, if the first data includes a combination of a string and a numerical value, the first data can be represented as groundtruth position; x1=0, y1=0, z1=0.2m. Alternatively, the first data can be represented as product1 position; x2=0, y2=0.2m, z2=0.2m.

[0170] In some embodiments, the first data includes environmental information, where the environmental information indicates the environment in which the audio data corresponding to the first data was collected. In embodiments of the present disclosure, when collecting audio data, the first device records the environment in which the audio data was collected. Therefore, the environmental information can be included in the first data to indicate the environment in which the audio data was collected.

[0171] In some embodiments, the environmental information includes at least one of the following:

[0172] (1) Temperature information;

[0173] (2) Humidity information;

[0174] (3) Air pressure information.

[0175] Optionally, the first data includes temperature 1, humidity 1 and air pressure 1. Optionally, the first data includes temperature 2, humidity 2 and air pressure 2. In some embodiments, the first data includes any one of a character string, a numerical value, and a combination of a character string and a numerical value. Optionally, if the first data is a character string, it can be expressed as environment1: 15°C, %30, 101.325kPa. Optionally, if the first data is a numerical value, it can be expressed as temperature1 (temperature 1) (°C) = 15, humidity (humidity) (%) = 30, air pressure (air pressure) (kPa) = 101.325", "temperature2 (°C) = 15, humidity (%) = 30, air pressure (kPa) = 101.325. Optionally, the first data is a combination of a character string and a numerical value, and can be expressed as "groundtruth environment", "product1 environment" and "temperature1 (℃) = 15, humidity (%) = 30, air pressure (kPa) = 101.325", "temperature2 (℃) = 15, humidity (%) = 30, air pressure (kPa) = 101.325".

[0176] Step S2103: The first device sends audio data and multiple first data to the encoder.

[0177] In some embodiments, after the first device collects audio data and generates a plurality of first data based on the audio data, the audio data and the plurality of first data may be sent to the encoder.

[0178] In some embodiments, the encoder receives audio data and a plurality of first data sent by the first device. In some embodiments, the first device sends audio data and a plurality of first data. In some embodiments, the encoder receives audio data and a plurality of first data.

[0179] Step S2104: The encoder encodes at least one of the first data and the audio data to obtain encoded data.

[0180] In the embodiment of the present disclosure, during transmission, data may be encoded and the resulting encoded data may be transmitted to reduce the amount of data required for transmission.

[0181] In some embodiments, the encoder may encode the audio data to obtain encoded data, which is then sent to the decoder in step S2105. In some embodiments, the encoder may encode the first data to obtain encoded data, which is then sent to the decoder in step S2105. In some embodiments, the encoder may encode the first data and audio data to obtain encoded data, which is then sent to the decoder in step S2105.

[0182] In some embodiments, the first data is encoded using Huffman coding to obtain encoded data.

[0183] In some embodiments, the audio data is encoded based on the first data to obtain the encoded data. In the embodiment of the present disclosure, the first data is used to describe the audio data, so the audio data can be encoded according to the information indicated by the first data to obtain the encoded data.

[0184] In some embodiments, the transmission between the encoder and the decoder may be performed using different encoding methods, and then the encoded data is transmitted. Different encoding methods are described below.

[0185] In some embodiments, the character-type first data is encoded using a character encoding method to obtain encoded data. Optionally, the character encoding method includes at least one of ASCII (American Standard Code for Information Interchange), GB2312 (Chinese Character Coded Character Set for Information Interchange), and UTF-8 (Universal Character Set / Unicode Transformation Format, 8-bit).

[0186] In some embodiments, the character-type first data is encoded using a one-to-one mapping between characters and numerical values ​​to obtain encoded data. Optionally, if the character string includes "groundtruth audio" and "audio to be testing," the corresponding numerical values ​​can be set to 0 and 1, such that 0 corresponds to "groundtruth audio" and 1 corresponds to "audio to be testing."

[0187] In some embodiments, the first data is carried in a MAE extension field. In some embodiments, the MAE extension field refers to an extension field of MPEG-H Audio (an audio standard). In some embodiments, MPEG-H Audio refers to the audio standard of ISO / IEC 23008-3. Optionally, MPEG-H data includes static data and dynamic data. Static data, also known as Data Audio Elements (MAE), optionally has four types: 1. Descriptive data: information about the presence of objects in the bitstream and high-level properties of the objects. 2. Restrictive data: information about how content creators can enable or enable interaction. 3. Positional data and the ability to render to specific speakers, as well as signal channel content as objects. 4. Structured data: grouping and combining objects. Dynamic data is optionally used for object-based signals.

[0188] In some embodiments, the first data may be of the type shown in Table 1, where ID_CONFIG_EXT_ACOM_INFO is used to indicate reference information, audio data, location information, and environment information of the first data.

[0189] Table 1

[0190] In some embodiments, the first data may be of the type shown in Table 2, where ID_CONFIG_EXT_ACOM_TYPE_INFO is used to indicate reference information of the first data, ID_CONFIG_EXT_ACOM_POSITION_INFO is used to indicate position information of the first data, and ID_CONFIG_EXT_ACOM_ENVIRONMENT_INFO is used to indicate environment information of the first data:

[0191] Table 2

[0192] Step S2105: The encoder carries the encoded data in a data stream and sends it to the decoder.

[0193] In some embodiments, if the encoder encodes the audio data to obtain encoded data, the first data and the encoding are carried in a data stream and sent to the decoder. In some embodiments, if the encoder encodes the audio data and the first data to obtain encoded data, the encoded data are carried in a data stream and sent to the decoder. In some embodiments, if the encoder encodes the first data to obtain encoded data, the encoded data and the audio data are carried in a data stream and sent to the decoder.

[0194] Step S2106: The decoder decodes the encoded data to obtain at least one of the second data and the audio data.

[0195] In some embodiments, the second data obtained by decoding the decoder can be understood as the first data in the above-mentioned embodiments. It should be noted that, because the first data may be lost during the process of encoding and decoding the first data to obtain the second data, the first data and the second data are not completely identical. In addition, the audio data obtained by decoding the decoder and the audio data encoded by the encoder can also be understood as not completely identical data.

[0196] In some embodiments, the second data is used to describe the audio data and is also used for machine analysis or machine training of the audio data.

[0197] In some embodiments, the encoded data is encoded using Huffman coding to obtain second data.

[0198] In some embodiments, the encoded data is decoded using a character encoding method to obtain the second data.

[0199] In some embodiments, the encoded data is decoded using a one-to-one mapping method from characters to numerical values ​​to obtain the second data.

[0200] In some embodiments, the names of information, etc. are not limited to the names described in the embodiments, and terms such as "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "symbol", "codeword", "codebook", "codeword", "codepoint", "bit", "data", "program", and "chip" can be used interchangeably.

[0201] In some embodiments, terms such as "uplink", "uplink", "physical uplink" can be interchangeable with each other, and terms such as "downlink", "downlink", "physical downlink" can be interchangeable with each other, and terms such as "side", "sidelink", "side communication", "sidelink communication", "direct connection", "direct link", "direct communication", "direct link communication" can be interchangeable with each other.

[0202] In some embodiments, "obtain", "get", "get", "receive", "transmit", "bidirectional transmission", "send and / or receive" can be interchangeable, and can be interpreted as receiving from other entities, obtaining from protocols, obtaining from higher layers, obtaining by self-processing, autonomous implementation, etc.

[0203] In some embodiments, terms such as "send", "transmit", "report", "download", "transmit", "bidirectional transmission", "send and / or receive" can be used interchangeably.

[0204] In some embodiments, terms such as "moment", "time point", "time", and "time position" can be replaced with each other, and terms such as "duration", "period", "time window", "window", and "time" can be replaced with each other.

[0205] In some embodiments, terms such as "certain", "preset", "preset", "setting", "indicated", "a certain", "any", and "first" can be interchangeable. "Specific A", "preset A", "preset A", "setting A", "indicated A", "a certain A", "any A", and "first A" can be interpreted as A pre-specified in a protocol, etc., or as A obtained through setting, configuration, or indication, etc., or as specific A, a certain A, any A, or first A, etc., but not limited to this.

[0206] The processing method involved in the embodiment of the present disclosure may include at least one of steps S2101 to S2106. For example, step S2101 can be implemented as an independent embodiment, step S2102 can be implemented as an independent embodiment, step S2103 can be implemented as an independent embodiment, step S2104 can be implemented as an independent embodiment, step S2105 can be implemented as an independent embodiment, step S2106 can be implemented as an independent embodiment, step S2101 and step S2102 can be implemented as independent embodiments, step S2103 and step S2104 can be implemented as independent embodiments, step S2105 and step S2106 can be implemented as independent embodiments, step S2101, step S2102, step S2103, and step S2104 can be implemented as independent embodiments, step S2101, step S2102, step S2105, and step S2106 can be implemented as independent embodiments, but is not limited to this.

[0207] In some embodiments, step S2101 is optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0208] In some embodiments, step S2102 is optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0209] In some embodiments, step S2103 is optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0210] In some embodiments, step S2104 is optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0211] In some embodiments, step S2105 is optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0212] In some embodiments, step S2106 is optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0213] In some embodiments, step S2101 and step S2102 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0214] In some embodiments, step S2103 and step S2104 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0215] In some embodiments, step S215 and step S2106 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0216] In some embodiments, reference may be made to other optional implementations described before or after the description corresponding to FIG. 2A .

[0217] FIG3A is a flow chart of a processing method according to an embodiment of the present disclosure, which is applied to a first device. As shown in FIG3A , the embodiment of the present disclosure relates to a processing method, which includes:

[0218] Step S3101: The first device collects audio data.

[0219] The optional implementation of step S3101 can refer to the optional implementation of step S2101 in Figure 2A and other related parts in the embodiment involved in Figure 2A, which will not be repeated here.

[0220] Step S3102: The first device generates multiple first data based on the audio data.

[0221] The optional implementation of step S3102 can refer to the optional implementation of step S2102 in Figure 2A and other related parts in the embodiment involved in Figure 2A, which will not be repeated here.

[0222] Step S3103: The first device sends audio data and multiple first data to the encoder.

[0223] The optional implementation of step S3103 can refer to the optional implementation of step S2103 in Figure 2A and other related parts in the embodiment involved in Figure 2A, which will not be repeated here.

[0224] The processing method involved in the embodiments of the present disclosure may include at least one of steps S3101 to S3103. For example, step S3101 may be implemented as an independent embodiment, step S3102 may be implemented as an independent embodiment, and step S3103 may be implemented as an independent embodiment, or at least two steps may be combined, but the present invention is not limited thereto.

[0225] In some embodiments, step S3101 is optional, step S3102 is optional, and step S3103 is optional, and one or more of these steps may be omitted or replaced in different embodiments, but the present invention is not limited thereto.

[0226] FIG3B is a flow chart of a processing method according to an embodiment of the present disclosure, which is applied to a first device. As shown in FIG3B , an embodiment of the present disclosure relates to a processing method, which includes:

[0227] Step S3201: The first device generates multiple first data based on audio data.

[0228] Optional implementations of step S3201 may refer to step S2102 in FIG. 2A , step S3102 in FIG. 3A , and other related parts in the embodiments involved in FIG. 2A and FIG. 3A , which will not be described in detail here.

[0229] FIG4A is a flow chart of a processing method according to an embodiment of the present disclosure, which is applied to an encoder. As shown in FIG4A , the embodiment of the present disclosure relates to a processing method, which includes:

[0230] Step S4101: The encoder encodes at least one of the first data and the audio data to obtain encoded data.

[0231] Optional implementations of step S4101 may refer to step S2104 in FIG. 2A and other related parts of the embodiment involved in FIG. 2A , which will not be described in detail here.

[0232] In step S4102, the encoder carries the encoded data in a data stream and sends it to the decoder.

[0233] The optional implementation of step S4102 can be found in step S2105 of FIG. 2A and other related parts of the embodiment involved in FIG. 2A , which will not be described in detail here.

[0234] The processing method involved in the embodiment of the present disclosure may include at least one of steps S4101 and S4102. For example, step S4101 may be implemented as an independent embodiment, step S4102 may be implemented as an independent embodiment, or at least two steps may be combined, but the present invention is not limited thereto.

[0235] In some embodiments, step S4101 is optional, step S4102 is optional, and one or more of these steps may be omitted or replaced in different embodiments, but the present invention is not limited thereto.

[0236] FIG4B is a flow chart of a processing method according to an embodiment of the present disclosure, which is applied to an encoder. As shown in FIG4B , the embodiment of the present disclosure relates to a processing method, which includes:

[0237] Step S4201: The encoder encodes at least one of the first data and the audio data to obtain encoded data.

[0238] Optional implementations of step S4201 may refer to step S2104 in FIG. 2A , step S4101 in FIG. 4A , and other related parts in the embodiments involved in FIG. 2A and FIG. 4A , which will not be described in detail here.

[0239] FIG5 is a flow chart of a processing method according to an embodiment of the present disclosure, which is applied to a decoder. As shown in FIG5 , the embodiment of the present disclosure relates to a processing method, which includes:

[0240] Step S5101: The decoder decodes the encoded data to obtain at least one of the second data and the audio data.

[0241] The optional implementation of step S5101 can be found in step S2106 of FIG. 2A and other related parts of the embodiment involved in FIG. 2A , which will not be described in detail here.

[0242] FIG6 is a flow chart of a processing method according to an embodiment of the present disclosure. As shown in FIG6 , the embodiment of the present disclosure relates to a processing method, which includes:

[0243] Step S6101: The first device generates multiple first data based on audio data.

[0244] In some embodiments, the first data is used to describe audio data and is also used for machine analysis or machine training of the audio data.

[0245] Step S6102: The encoder encodes at least one of the first data and the audio data to obtain encoded data.

[0246] Step S6103: The decoder decodes the encoded data to obtain at least one of the second data and the audio data.

[0247] In some embodiments, the second data is used to describe audio data and is also used for machine analysis or machine training of the audio data.

[0248] In some embodiments, the above method may include the methods of the above embodiments of the communication system side, the first device side, the encoder side, the decoder side, etc., which will not be repeated here.

[0249] FIG7 is a flow chart of a processing method according to an embodiment of the present disclosure. As shown in FIG7 , the embodiment of the present disclosure relates to a processing method, which includes:

[0250] Step S7101: Determine metadata required for automatic detection.

[0251] In some embodiments, decoded audio channel 1 has corresponding metadata 1. Metadata 1 contains channel information: groundtruth audio.

[0252] The decoded audio channel 2 has corresponding metadata 2. Metadata 2 contains the channel information to be tested: product1 audio.

[0253] The decoded audio channel 3 has corresponding metadata 3. Metadata 3 contains the channel information to be tested: product2 audio.

[0254] In some embodiments, the metadata of the channel information may be a string or a numerical value, or a combination of the two.

[0255] For example, the format of the string is: "groundtruth audio", "product1 audio to be testing"...

[0256] For example, the value format is: "0: groundtruth audio", "1: audio of the product to be tested"...

[0257] For example, the combination format of string and value is: "groundtruth audio", "product1 audio" and "0:groundtruth audio", "1:test product audio"...

[0258] In some embodiments, the detector uses a certain judgment method to compare the audio channel of the product to be tested with the audio true value channel.

[0259] The difference between the audio channel of the product under test and the audio true bottom channel does not exceed the threshold, and product 1 is judged to be qualified.

[0260] If the difference between the audio channel of the product under test and the audio true bottom channel exceeds the threshold, the product is judged to be unqualified.

[0261] In some embodiments, the metadata also records the different locations of the audio data.

[0262] The detector takes two sets of audio channels and corresponding metadata.

[0263] In some embodiments, group 1 contains decoded audio channels 1 to 3 and corresponding metadata, which are labeled as groundtruth audio.

[0264] In some embodiments, group 2 includes decoded audio channels 4 to 6 and corresponding metadata, marked as product 1 audio to be tested.

[0265] In some embodiments, decoded audio for channel 1 is captured at position 1: x1, y1, z1.

[0266] The decoded audio of channel 2 is captured at position 2: x2, y2, z2.

[0267] The decoded audio of channel 3 is captured at position 3: x3, y3, z3.

[0268] The decoded audio of channel 4 is captured at position 4: x4, y4, z4.

[0269] The decoded audio of channel 5 is captured at position 5: x5, y5, z5.

[0270] The decoded audio of channel 6 is captured at position 6: x6, y6, z6.

[0271] Among them, position 1 and position 4 are the same, position 2 and position 5 are the same, and position 3 and position 6 are the same.

[0272] The coordinate system can be a rectangular coordinate system or a polar coordinate system, and (x, y, z) can be replaced by (r, θ, φ).

[0273] In some embodiments, the metadata of the channel information may be a string or a numerical value, or a combination of the two.

[0274] For example, the format of the string is: "position1,x1,y1,z1", "position2,x2,y2,z2"...

[0275] For example, the format of the value is: "x1=0,y1=0,z1=0.2m", "x2=0,y2=0.2m,z2=0.2m"...

[0276] For example, the combination format of string and value is: "groundtruth position", "product1 position" and "x1=0,y1=0,z1=0.2m", "x2=0,y2=0.2m,z2=0.2m"

[0277] In some embodiments, the detector compares two groups of audio channels and audio true value channels of the product to be tested, and adopts a certain judgment method.

[0278] Optionally, if the difference between the two audio channel groups of the product to be tested and the audio true bottom channel does not exceed a threshold, the product 1 is judged to be qualified.

[0279] Optionally, the difference between the audio channel of the product under test and the audio true value channel at a certain location will affect the final result of the product

[0280] In some embodiments, the metadata also records the context of the audio data.

[0281] In some embodiments, decoded audio channel 1 has corresponding metadata 1. Metadata 1 includes channel information: groundtruth audio, sensor position information: position1 (x1, y1, z1), and environmental information: temperature 1, humidity 1, and air pressure 1.

[0282] In some embodiments, decoded audio channel 2 has corresponding metadata 2. Metadata 2 includes channel information: groundtruth audio, sensor position information: position1 (x2, y2, z2), and environmental information: temperature 2, humidity 2, and air pressure 2.

[0283] In some embodiments, the metadata of the channel information may be a string or a numerical value, or a combination of the two.

[0284] For example, the string format is "environment1:15℃,%30,101.325kPa", "environment:15℃,%30,101.325kPa"...

[0285] For example, the numerical format is: "temperature1(℃)=15,humidity(%)=30,air pressure(kPa)=101.325", "temperature2(℃)=15,humidity(%)=30,air pressure(kPa)=101.325"...

[0286] For example, the combination format of string and value is: "groundtruth environment", "product1 environment" and "temperature1(℃)=15,humidity(%)=30,air pressure(kPa)=101.325", "temperature2(℃)=15,humidity(%)=30,air pressure(kPa)=101.325"...

[0287] In the embodiments of the present disclosure, some or all of the steps and their optional implementations may be arbitrarily combined with some or all of the steps in other embodiments, or may be arbitrarily combined with the optional implementations of other embodiments.

[0288] The present disclosure also provides an apparatus for implementing any of the above methods. For example, a device is provided that includes units or modules for implementing each step performed by an encoder in any of the above methods. For another example, another device is provided that includes units or modules for implementing each step performed by a decoder in any of the above methods.

[0289] It should be understood that the division of the various units or modules in the above device is merely a division of logical functions. In actual implementation, they may be fully or partially integrated into a physical entity, or they may be physically separated. In addition, the units or modules in the device may be implemented in the form of a processor calling software: for example, the device includes a processor, the processor is connected to a memory, and the memory stores instructions. The processor calls the instructions stored in the memory to implement any of the above methods or implement the functions of the various units or modules of the above device, wherein the processor is, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory is a memory within the device or a memory outside the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits, and the functions of some or all of the units or modules can be realized by designing the hardware circuits. The above-mentioned hardware circuits can be understood as one or more processors; for example, in one implementation, the above-mentioned hardware circuit is an application-specific integrated circuit (ASIC), and the functions of some or all of the above units or modules are realized by designing the logical relationship of the components in the circuit; for example, in another implementation, the above-mentioned hardware circuit can be realized by a programmable logic device (PLD). Taking a field programmable gate array (FPGA) as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by a configuration file, thereby realizing the functions of some or all of the above units or modules. All units or modules of the above devices can be realized in the form of software called by the processor, or in the form of hardware circuits, or in part by software called by the processor, and the rest by hardware circuits.

[0290] In the embodiments of the present disclosure, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationship of the hardware circuit. The logical relationship of the above-mentioned hardware circuit is fixed or reconfigurable. For example, the processor is a hardware circuit implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and implementing the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), etc.

[0291] Figure 8A is a structural diagram of the processing device proposed in an embodiment of the present disclosure. As shown in Figure 8A, the processing device 8100 may include: at least one of a transceiver module 8101, a processing module 8102, etc. In some embodiments, the processing module 8102 is used to generate a plurality of first data based on the audio data, and the first data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data. Optionally, the above-mentioned transceiver module 8101 is used to execute at least one of the communication steps such as sending and / or receiving performed by the encoder in any of the above methods (such as step S2101 but not limited to this), which will not be repeated here. Optionally, the above-mentioned processing module is used to execute at least one of the other steps performed by the encoder in any of the above methods, which will not be repeated here.

[0292] Optionally, the processing module 8102 is used to execute at least one of the communication steps such as processing performed by the encoder in any of the above methods, which will not be repeated here.

[0293] FIG8B is a schematic diagram of the structure of the processing device proposed in an embodiment of the present disclosure. As shown in FIG8B , the processing device 8200 may include: at least one of a transceiver module 8201 and a processing module 8202. In some embodiments, the processing module 8202 is used to encode at least one of the first data and the audio data to obtain encoded data, wherein the first data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data. Optionally, the above-mentioned transceiver module is used to perform at least one of the communication steps such as sending and / or receiving performed by the decoder in any of the above methods, which will not be repeated here.

[0294] Optionally, the processing module 8202 is used to execute at least one of the communication steps such as processing performed by the decoder in any of the above methods, which will not be repeated here.

[0295] In some embodiments, the transceiver module may include a transmitting module and / or a receiving module, and the transmitting module and the receiving module may be separate or integrated. Optionally, the transceiver module may be interchangeable with the transceiver.

[0296] In some embodiments, the processing module can be a single module or can include multiple submodules. Optionally, the multiple submodules respectively execute all or part of the steps required to be executed by the processing module. Optionally, the processing module can be interchangeable with the processor.

[0297] FIG8C is a schematic diagram of the structure of the processing device proposed in an embodiment of the present disclosure. As shown in FIG8C , the processing device 8300 may include: at least one of a transceiver module 8301 and a processing module 8302. In some embodiments, the processing module 8302 is used to decode the encoded data to obtain at least one of the second data and the audio data, wherein the second data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data. Optionally, the above-mentioned transceiver module is used to perform at least one of the communication steps such as sending and / or receiving performed by the decoder in any of the above methods, which will not be repeated here.

[0298] Optionally, the processing module 8302 is used to execute at least one of the communication steps such as processing performed by the decoder in any of the above methods, which will not be repeated here.

[0299] In some embodiments, the transceiver module may include a transmitting module and / or a receiving module, and the transmitting module and the receiving module may be separate or integrated. Optionally, the transceiver module may be interchangeable with the transceiver.

[0300] In some embodiments, the processing module can be a single module or can include multiple submodules. Optionally, the multiple submodules respectively execute all or part of the steps required to be executed by the processing module. Optionally, the processing module can be interchangeable with the processor.

[0301] Figure 9A is a schematic diagram of the structure of a communication device 9100 proposed in an embodiment of the present disclosure. Communication device 9100 can be a first device, an encoder, a decoder, or a chip, a chip system, or a processor that supports the first device, encoder, or decoder in implementing any of the above methods. Communication device 9100 can be used to implement the methods described in the above method embodiments. For details, please refer to the description of the above method embodiments.

[0302] As shown in Figure 9A, communication device 9100 includes one or more processors 9101. Processor 9101 can be a general-purpose processor or a dedicated processor, for example, a baseband processor or a central processing unit. The baseband processor can be used to process communication protocols and communication data, while the central processing unit can be used to control the processing device, execute programs, and process program data. Communication device 9100 is configured to perform any of the above methods.

[0303] In some embodiments, the communication device 9100 further includes one or more memories 9102 for storing instructions. Optionally, all or part of the memories 9102 may be located outside the communication device 9100.

[0304] In some embodiments, the communication device 9100 further includes one or more transceivers 9103. When the communication device 9100 includes one or more transceivers 9103, the transceiver 9103 performs at least one of the communication steps of sending and / or receiving in the above method.

[0305] In some embodiments, a transceiver may include a receiver and / or a transmitter. The receiver and transmitter may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, and transceiver circuit may be used interchangeably; the terms transmitter, transmitting unit, transmitter, and transmitting circuit may be used interchangeably; and the terms receiver, receiving unit, receiver, and receiving circuit may be used interchangeably.

[0306] In some embodiments, the communication device 9100 may include one or more interface circuits 9104. Optionally, the interface circuit 9104 is connected to the memory 9102. The interface circuit 9104 may be configured to receive signals from the memory 9102 or other devices, and may be configured to send signals to the memory 9102 or other devices. For example, the interface circuit 9104 may read instructions stored in the memory 9102 and send the instructions to the processor 9101.

[0307] The communication device 9100 described in the above embodiments may be a first device, a decoder, or an encoder, but the scope of the communication device 9100 described in the present disclosure is not limited thereto, and the structure of the communication device 9100 may not be limited by FIG. 9A. The communication device may be an independent device or may be part of a larger device. For example, the communication device may be: 1) an independent integrated circuit IC, or a chip, or a chip system or subsystem; (2) a collection of one or more ICs, optionally, the above IC collection may also include a storage component for storing data or programs; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, an encoder device, an intelligent encoder device, a cellular phone, a wireless device, a handheld device, a mobile unit, an in-vehicle device, a decoder, a cloud device, an artificial intelligence device, etc.; (6) others, etc.

[0308] 9B is a schematic diagram of the structure of a chip 9200 according to an embodiment of the present disclosure. If the communication device 9100 can be a chip or a chip system, please refer to the schematic diagram of the structure of the chip 9200 shown in FIG9B , but the present disclosure is not limited thereto.

[0309] The chip 9200 includes one or more processors 9201 , and the chip 9200 is configured to execute any of the above methods.

[0310] In some embodiments, the chip 9200 further includes one or more interface circuits 9202. Optionally, the interface circuit 9202 is connected to the memory 9203. The interface circuit 9202 can be used to receive signals from the memory 9203 or other devices, and can be used to send signals to the memory 9203 or other devices. For example, the interface circuit 9202 can read instructions stored in the memory 9203 and send the instructions to the processor 9201.

[0311] In some embodiments, the interface circuit 9202 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processor 9201 performs at least one of the other steps.

[0312] In some embodiments, terms such as interface circuit, interface, transceiver pin, and transceiver may be used interchangeably.

[0313] In some embodiments, the chip 9200 further includes one or more memories 9203 for storing instructions. Alternatively, all or part of the memories 9203 may be located outside the chip 9200.

[0314] The present disclosure also proposes a storage medium having instructions stored thereon, which, when executed on the communication device 9100, causes the communication device 9100 to execute any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but is not limited thereto and may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but is not limited thereto and may also be a temporary storage medium.

[0315] The present disclosure also provides a program product, which, when executed by the communication device 9100, enables the communication device 9100 to perform any of the above methods. Optionally, the program product is a computer program product.

[0316] The present disclosure also proposes a computer program, which, when executed on a computer, causes the computer to perform any one of the above methods.

Claims

1. A processing method, characterized in that: The method is performed by a first device, and includes: A plurality of first data are generated based on the audio data, where the first data are used to describe the audio data and are also used to perform machine analysis or machine training on the audio data.

2. The method according to claim 1, characterized in that At least one of the audio data or the first data is further used by an encoder for encoding and transmission.

3. The method according to claim 2, characterized in that The first data is encoded using Huffman coding.

4. The method according to any one of claims 1 to 3, characterized in that: Each first data corresponds to a channel, and the channel includes the audio data corresponding to the first data.

5. The method according to any one of claims 1 to 4, characterized in that: The first data includes reference information, where the reference information is used to indicate reference audio data corresponding to the first data for matching with the audio data to be tested.

6. The method according to claim 5, characterized in that The first data belongs to a first group, the first group includes a plurality of first data, and the first data included in the first group all include the reference information.

7. The method according to any one of claims 1 to 4, characterized in that: The first data includes information to be tested, where the information to be tested is used to indicate that audio data to be tested corresponding to the first data is used to be matched with reference audio data.

8. The method according to claim 7, characterized in that The first data belongs to a second group, the second group includes a plurality of the first data, and the first data included in the second group all include the information to be tested.

9. The method according to any one of claims 1 to 8, characterized in that: The first data includes location information, and the location information is used to indicate a collection location of audio data corresponding to the first data.

10. The method according to any one of claims 1 to 9, characterized in that: The first data includes environmental information, and the environmental information is used to indicate the environment in which the audio data corresponding to the first data is collected.

11. The method according to claim 10, characterized in that The environmental information includes at least one of the following: Temperature information; Humidity information; Air pressure information.

12. The method according to any one of claims 1 to 11, characterized in that: The first data includes any one of a character string, a numerical value, or a combination of the character string and the numerical value.

13. The method according to any one of claims 1 to 12, characterized in that: The first data is carried in the metadata audio element MAE extension field.

14. A processing method, characterized in that: The method is performed by an encoder, and includes: At least one of the first data and the audio data is encoded to obtain encoded data, where the first data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data.

15. The method according to claim 14, characterized in that The step of encoding at least one of the first data and the audio data to obtain encoded data includes: The first data is encoded using Huffman encoding to obtain the encoded data.

16. The method according to claim 14, characterized in that The step of encoding at least one of the first data and the audio data to obtain encoded data includes: The audio data is encoded based on the first data to obtain the encoded data.

17. The method according to claim 14, characterized in that The step of encoding at least one of the first data and the audio data to obtain encoded data includes: Encode the first data of the character type using a character encoding method to obtain the encoded data; or, The first data of the character type is encoded in a one-to-one mapping manner from characters to numerical values to obtain the encoded data.

18. The method according to any one of claims 14 to 17, characterized in that The first data is carried in the MAE extension domain.

19. A processing method, characterized in that: The method is performed by a decoder, and comprises: The encoded data is decoded to obtain at least one of second data and audio data, where the second data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data.

20. The method according to claim 19, characterized in that The decoding of the encoded data to obtain at least one of the second data and the audio data includes: The coded data is encoded using Huffman coding to obtain the second data.

21. The method according to claim 19, wherein The decoding of the encoded data to obtain at least one of the second data and the audio data includes: Decoding the encoded data in a character encoding manner to obtain the second data; or The encoded data is decoded using a one-to-one mapping method from characters to numerical values to obtain the second data.

22. A processing method, characterized in that: The method comprises: The first device generates a plurality of first data based on the audio data, where the first data is used to describe the audio data and is further used to perform machine analysis or machine training on the audio data; The encoder encodes at least one of the first data and the audio data to obtain encoded data; The decoder decodes the encoded data to obtain at least one of second data and audio data, where the second data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data.

23. A processing device, characterized in that The processing device comprises: The processing module is used to generate a plurality of first data based on the audio data, where the first data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data.

24. A processing device, characterized in that The processing device comprises: The processing module is used to encode at least one of the first data and the audio data to obtain encoded data, wherein the first data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data.

25. A processing device, characterized in that The processing device comprises: The processing module is used to decode the encoded data to obtain at least one of second data and audio data, where the second data is used to describe the audio data and is also used to perform machine analysis or machine training on the audio data.

26. A processing device, characterized in that The processing device comprises: one or more processors; The processor is configured to execute the processing method according to any one of claims 1 to 13.

27. A processing device, characterized in that The processing device comprises: one or more processors; The processor is configured to execute the processing method according to any one of claims 14 to 18.

28. A processing device, characterized in that The processing device comprises: one or more processors; The processor is configured to execute the processing method according to any one of claims 19 to 21.

29. A coding and decoding system, characterized in that: The invention comprises a first device, an encoder and a decoder, wherein the first device is configured to implement the processing method described in any one of claims 1 to 13, the encoder is configured to implement the processing method described in any one of claims 14 to 18, and the decoder is configured to implement the processing method described in any one of claims 19 to 21.

30. A storage medium storing instructions, characterized in that: When the instruction is executed on a communication device, the communication device executes the processing method according to any one of claims 1 to 13, or the processing method according to any one of claims 14 to 18, or the processing method according to any one of claims 19 to 21.

Citation Information

Patent Citations

  • Intelligent audio abnormal sound detection method

    CN109300483A

  • Audio signal analysis system based on machine learning

    CN115862647A

  • Method and apparatus for prediciting a map object based upon audio data

    US20210372813A1

  • Apparatus and method for audio data analysis

    US20220111294A1

  • Encoder generation method, fingerprint extraction method, medium, and electronic device

    WO2023134549A1