Data processing method, data processing apparatus, communication device and storage medium

By decoding and separating audio files from encoded datasets and processing them according to AI task type identifiers, the problem of low data processing efficiency for different AI tasks is solved, achieving more efficient data processing and model transfer.

WO2026060566A1PCT designated stage Publication Date: 2026-03-26BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively identify and process the audio files required for different types of AI tasks, resulting in low data processing efficiency.

Method used

By decoding the encoded dataset, audio files and their corresponding metadata, including AI task type identifiers, are obtained, and the audio files required for each AI task are separated and distributed based on these identifiers.

Benefits of technology

It improves data processing efficiency, simplifies data processing workflows, reduces repetitive work, facilitates model migration and comparison, supports automated tools and processes, and simplifies dataset expansion and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024119470_26032026_PF_FP_ABST
    Figure CN2024119470_26032026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a data processing method, a data processing apparatus, a communication device, and a storage medium. The data processing method comprises: decoding an encoded data set to obtain audio files and metadata corresponding to the audio files, wherein the metadata comprises AI task type identifiers corresponding to the audio files; and determining, on the basis of the metadata, an audio file respectively corresponding to each AI task type. By means of the embodiments of the present disclosure, the data processing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing method, data processing apparatus, communication device and storage medium TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of communication, and particularly relates to a data processing method, a data processing apparatus, a communication device and a storage medium. BACKGROUND

[0002] With the rapid development of artificial intelligence (AI) technology, the application of AI is more and more extensive. AI tasks such as automatic speech recognition (ASR), automatic speaker verification (ASV), emotion recognition (ER), and audio event classification (AEC) can be performed by using an AI model.

[0003] SUMMARY

[0004] How to determine the audio files required by each AI task from a data set is a problem to be solved.

[0005] Embodiments of the present disclosure provide a data processing method, a data processing apparatus, a communication device and a storage medium.

[0006] According to a first aspect of embodiments of the present disclosure, a data processing method is provided, comprising: decoding an encoded data set to obtain an audio file and metadata corresponding to the audio file, the metadata comprising an AI task type identifier corresponding to the audio file; determining the audio file corresponding to each AI task type based on the metadata.

[0007] According to a second aspect of embodiments of the present disclosure, a data processing apparatus is provided, comprising: a processing module configured to decode an encoded data set to obtain an audio file and metadata corresponding to the audio file, the metadata comprising an AI task type identifier corresponding to the audio file; and determine the audio file corresponding to each AI task type based on the metadata.

[0008] According to a third aspect of embodiments of the present disclosure, a communication device is provided, comprising: one or more processors; wherein the communication device is configured to perform the data processing method of the first aspect.

[0009] According to a fourth aspect of embodiments of the present disclosure, a storage medium is provided, the storage medium storing instructions, when the instructions are executed on a communication device, causing the communication device to perform the method of the first aspect.

[0010] According to a fifth aspect of the embodiments of the present disclosure, a computer program is provided, which, when executed by a communication device, causes the communication device to perform the method of the first aspect.

[0011] According to the embodiments of the present disclosure, the encoded data set is decoded to obtain an audio file and metadata corresponding to the audio file, the metadata including an AI task type identifier corresponding to the audio file, and the audio file corresponding to each AI task type is obtained based on the task type identifier, so that each AI task can be trained, verified or tested using the audio file aligned therewith, and the data processing efficiency can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following describes the drawings required for the embodiments, and the following drawings are only some embodiments of the present disclosure, and do not specifically limit the protection scope of the present disclosure.

[0013] FIG. 1A is a schematic diagram of an architecture of a communication system according to an embodiment of the present disclosure.

[0014] FIG. 1B is a schematic diagram of a data processing process according to an embodiment of the present disclosure.

[0015] FIG. 1C is a schematic diagram of a data processing process according to an embodiment of the present disclosure.

[0016] FIG. 2A is a schematic diagram of a data processing method according to an embodiment of the present disclosure.

[0017] FIG. 2B is a schematic diagram of a data processing method according to an embodiment of the present disclosure.

[0018] FIG. 2C is a schematic diagram of a data processing method according to an embodiment of the present disclosure.

[0019] FIG. 2D is a schematic diagram of a data processing method according to an embodiment of the present disclosure.

[0020] FIG. 3 is a schematic diagram of a data processing method according to an embodiment of the present disclosure.

[0021] FIG. 4 is a schematic diagram of a data processing method according to an embodiment of the present disclosure.

[0022] FIG. 5 is a schematic diagram of a data processing apparatus according to an embodiment of the present disclosure.

[0023] FIG. 6A is a schematic diagram of a communication device according to an embodiment of the present disclosure.

[0024] FIG. 6B is a schematic diagram of a chip according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] The embodiments of the present disclosure propose a data processing method, a data processing apparatus, a communication device and a storage medium.

[0026] In a first aspect, the embodiments of the present disclosure propose a data processing method, comprising: decoding an encoded data set to obtain an audio file and metadata corresponding to the audio file, the metadata comprising an AI task type identifier corresponding to the audio file; determining, based on the metadata, audio files corresponding to each AI task type.

[0027] In the above embodiments, the encoded data set comprises data of multiple AI task types, and audio files corresponding to each AI task type are obtained based on the task type identifier, which facilitates subsequent training, verification or testing of each AI task using the audio file aligned therewith, and improves data processing efficiency.

[0028] In combination with some embodiments of the first aspect, in some embodiments, the metadata further comprises at least one of the following: an identifier of a data set type, the data set type comprising at least one of a training data set, a verification data set and a test data set; a label of an AI task; a unique identifier of an audio file.

[0029] In combination with some embodiments of the first aspect, in some embodiments, the encoded audio file and the metadata corresponding to the audio file are associated through the unique identifier of the audio file.

[0030] In combination with some embodiments of the first aspect, in some embodiments, the audio file and the metadata corresponding to the audio file are associated through the unique identifier of the audio file.

[0031] In combination with some embodiments of the first aspect, in some embodiments, the method further comprises: determining, based on the unique identifier of the audio file, the metadata corresponding to the audio file.

[0032] In combination with some embodiments of the first aspect, in some embodiments, the determining, based on the metadata, of audio files corresponding to each AI task type comprises: grouping the audio files based on the AI task type identifier to obtain audio files corresponding to each AI task type.

[0033] In combination with some embodiments of the first aspect, in some embodiments, the audio file corresponds to one or more groups of metadata, and each AI task type corresponds to one or more labels.

[0034] In combination with some embodiments of the first aspect, in some embodiments, the label is used for supervised training of an AI task.

[0035] In some embodiments combined with the first aspect, in some embodiments, the method further includes: storing the audio file of each AI task type and the corresponding label into a storage container corresponding to each AI task type.

[0036] In some embodiments combined with the first aspect, in some embodiments, the method further includes: distributing the audio file corresponding to each AI task type to an AI model corresponding to each AI task type, and processing the audio file based on the AI model.

[0037] In a second aspect, the embodiments of the present disclosure provide a data processing apparatus, including: a processing module configured to decode an encoded data set to obtain an audio file and metadata corresponding to the audio file, the metadata including an AI task type corresponding to the audio file; and determine an audio file corresponding to each AI task type based on the metadata.

[0038] In a third aspect, the embodiments of the present disclosure provide a communication device, including: one or more processors; and wherein the communication device is configured to perform the data processing method of the first aspect.

[0039] In a fourth aspect, the embodiments of the present disclosure provide a storage medium, which stores instructions, and when the instructions are executed on a communication device, the communication device performs any of the above data processing methods.

[0040] In a fifth aspect, the embodiments of the present disclosure provide a program product, and when the program product is executed on a communication device, the communication device performs any of the above data processing methods.

[0041] In a sixth aspect, the embodiments of the present disclosure provide a computer program, and when the computer program is executed on a communication device, the communication device performs any of the above data processing methods.

[0042] In a seventh aspect, the embodiments of the present disclosure provide a chip or chip system. The chip or chip system includes a processing circuit configured to perform any of the above data processing methods.

[0043] It can be understood that the above communication device, storage medium, program product, computer program, chip or chip system are all used to execute the method proposed by the embodiments of the present disclosure. Therefore, the beneficial effects that can be achieved are referred to the beneficial effects in the corresponding method, which will not be repeated here.

[0044] The embodiments of the present disclosure propose a data processing method, a data processing apparatus, a communication system and a storage medium. In some embodiments, the data processing method and the information sending method, the information receiving method and other terms can be replaced with each other.

[0045] The embodiments of the present disclosure are not exhaustive, but only illustrate some embodiments, and are not specific limitations on the protection scope of the present disclosure. In the case of no contradiction, each step in an embodiment can be implemented as an independent embodiment, and the steps can be combined arbitrarily, for example, the scheme after removing part of the steps in an embodiment can also be implemented as an independent embodiment, and the order of the steps in an embodiment can be exchanged arbitrarily, in addition, the optional implementation in an embodiment can be combined arbitrarily; in addition, the embodiments can be combined arbitrarily, for example, part or all steps of different embodiments can be combined arbitrarily, an embodiment can be combined with optional implementation of other embodiments arbitrarily.

[0046] In each embodiment of the present disclosure, the terms and / or descriptions between the embodiments are consistent if there is no special description and logical conflict, and can be referred to each other, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.

[0047] The terms used in the embodiments of the present disclosure are only for the purpose of describing the specific embodiments, and not as a limitation on the present disclosure.

[0048] In the embodiments of the present disclosure, unless otherwise specified, the elements expressed in singular form, such as "one", "a", "the", "above", "said", "preceding", "this" and the like, can represent "one and only one", and can also represent "one or more", "at least one" and the like. For example, in the case of using articles such as "a", "an", "the" and the like in English, the noun after the article can be understood as singular expression, and can also be understood as plural expression.

[0049] In the embodiments of the present disclosure, "plurality" means two or more.

[0050] In some embodiments, the terms "at least one of", "one or more", "a plurality of", "multiple" and the like can be replaced with each other.

[0051] In some embodiments, "at least one of A, B", "A and / or B", "in one case A, in another case B", "responsive to case A, responsive to case B" and the like, can be interpreted to include both cases, A and B, in some embodiments, A (A is performed regardless of B), in some embodiments, B (B is performed regardless of A), in some embodiments, selected from the group consisting of A and B (the selection between A and B is an option), in some embodiments, A and B (both A and B are performed).

[0052] In some embodiments, "A or B" and the like, can be interpreted to include both cases, A and B, in some embodiments, A (A is performed regardless of B), in some embodiments, B (B is performed regardless of A), in some embodiments, selected from the group consisting of A and B (the selection between A and B is an option).

[0053] In some embodiments, the prefix words "first", "second" and the like in the disclosure do not limit the position, order, priority, number or content of the described objects, and the description of the described objects should be understood in the context of the claims or embodiments, and should not be construed as redundant limitations. For example, the described object is "field", and the ordinal words before "field" in "first field" and "second field" do not limit the position or order between "fields", and "first" and "second" do not limit whether the "fields" modified by them are in the same message or not, nor do they limit the order of "first field" and "second field". For another example, the described object is "level", and the ordinal words before "level" in "first level" and "second level" do not limit the priority between "levels". For another example, the number of described objects is not limited by ordinal words, and can be one or more. For example, "first device", where the number of "devices" can be one or more. In addition, objects modified by different prefix words can be the same or different, for example, the described object is "device", and "first device" and "second device" can be the same device or different devices, and their types can be the same or different; for another example, the described object is "information", and "first information" and "second information" can be the same information or different information, and their contents can be the same or different.

[0054] In some embodiments, "including A", "containing A", "for indicating A", "carrying A" can be interpreted as directly carrying A, or indirectly indicating A.

[0055] In some embodiments, the terms "in response to", "in response to determining", "in the case of", "when", "when", "if", "if" and the like can be replaced with each other.

[0056] In some embodiments, the terms "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not lower than", "above", and the like can be replaced with each other, and the terms "less than", "less than or equal to", "not greater than", "fewer than", "fewer than or equal to", "not more than", "lower than", "lower than or equal to", "not higher than", "below", and the like can be replaced with each other.

[0057] In some embodiments, an apparatus and the like can be interpreted as an entity, and can also be interpreted as virtual, and the name thereof is not limited to the name recited in the embodiments, and the terms "apparatus", "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "subject", and the like can be replaced with each other.

[0058] In some embodiments, "network" can be interpreted as an apparatus (for example, an access network device, a core network device, and the like) included in the network.

[0059] In some embodiments, the terms “access network device (AN device),” “radio access network device (RAN device),” “base station (BS),” “radio base station,” “fixed station,” “node,” “access point,” “transmission point (TP),” “reception point (RP),” “transmission / reception point (TRP),” “panel,” “antenna panel,” “antenna array,” “cell,” “macro cell,” “small cell,” “femto cell,” “pico cell,” “sector,” “cell group,” “serving cell,” “carrier,” “component carrier,” “bandwidth part (BWP),” and the like can be used interchangeably.

[0060] In some embodiments, the terms "terminal," "terminal device," "user equipment (UE)," "user terminal," "mobile station (MS)," "mobile terminal (MT)," "subscriber station," "mobile unit," "subscriber unit," "wireless unit," "remote unit," "mobile device," "wireless device," "wireless communication device," "remote device," "mobile subscriber station," "access terminal," "mobile terminal," "wireless terminal," "remote terminal," "handset," "user agent," "mobile client," "client," and so on can be replaced with each other.

[0061] In some embodiments, an access network device, a core network device, or a network device can be replaced with a terminal. For example, the embodiments of the present disclosure can also be applied to a structure in which communication between an access network device, a core network device, or a network device and a terminal is replaced with communication between a plurality of terminals (e.g., device-to-device (D2D), vehicle-to-everything (V2X), etc.). In this case, the structure in which the terminal has all or part of the functions of the access network device can also be provided. In addition, the terms "uplink," "downlink," and the like can be replaced with terms corresponding to the inter-terminal communication (e.g., "side"). For example, an uplink channel, a downlink channel, and the like can be replaced with a side channel, and an uplink, a downlink, and the like can be replaced with a sidelink.

[0062] In some embodiments, a terminal can be replaced with an access network device, a core network device, or a network device. In this case, the structure in which the access network device, the core network device, or the network device has all or part of the functions of the terminal can also be provided.

[0063] In some embodiments, the data, information, etc. can be obtained in compliance with the laws and regulations of the country where the location is located.

[0064] In some embodiments, the data, information, etc. can be obtained after obtaining the consent of the user.

[0065] In addition, each element, each row, or each column in the table of the embodiments of the present disclosure can be implemented as an independent embodiment, and any combination of any element, any row, or any column can also be implemented as an independent embodiment.

[0066] FIG. 1A is a schematic diagram of an architecture of a communication system according to an embodiment of the present disclosure.

[0067] As shown in FIG. 1A, the communication system 100 includes an encoding end 101 and a decoding end 102.

[0068] In some embodiments, the encoding end 101 can be any electronic device with processing capability, and the encoding end 101 can be, for example, a server or a terminal. The decoding end 102 can be any electronic device with processing capability, and the decoding end 102 can be, for example, a server or a terminal.

[0069] In some embodiments, the encoding end 101 can perform encoding processing (which can also be referred to as compression processing) on original data. The type of the original data can be one or more of a training set, a validation set, and a test set, and the content of the original data can include audio data and label data.

[0070] In some embodiments, the encoding end 101 performs compression processing on the original data to meet the transmission or storage needs, and the output of the encoding end 101 can be a bitstream. For example, in a scenario where the transmission bandwidth of a network or a bus cannot meet the real-time transmission of the original data, the encoding end 101 performs compression processing on the original data. For another example, in a scenario where a storage device with limited storage space cannot meet the storage of the original data, the encoding end 101 performs compression processing on the original data.

[0071] In some embodiments, the decoding end 102 can receive the encoded data set sent by the encoding end 101, and the encoded data set can be transmitted in the form of a bitstream. The decoding end 102 can decode the data set to obtain an audio file and metadata corresponding to the audio file, the metadata including an AI task type identifier corresponding to the audio file; the decoding end 102 can determine the audio file corresponding to each AI task type based on the metadata, and distribute the audio file to each AI task for processing.

[0072] FIG. 1B is a schematic diagram of a data processing process according to an embodiment of the present disclosure.

[0073] As shown in FIG. 1B, a transmission process module compresses the original data (one or more of the training set, the validation set, and the test set) to meet the transmission or storage needs, and the output of the module is a bitstream. Application scenario one: a scenario in which a network or a bus with limited transmission bandwidth cannot meet the real-time transmission of the original data; application scenario two: a scenario in which a storage device with limited storage space cannot meet the storage of the original data.

[0074] The input of the local process module is the bitstream, and the local module decodes and converts the format of the bitstream. The format conversion is used to adapt to the data format of the back-end AI task.

[0075] The back-end AI task module is a user of the original data, and the module includes a plurality of AI task corresponding models. The AI model set uses data for training, verification, testing, etc.

[0076] FIG. 1C is a schematic diagram of a data processing process according to an embodiment of the present disclosure.

[0077] FIG. 1C shows a specific process of processing the bitstream, the local process module in FIG. 1B includes a decoder and a format conversion module in FIG. 1C, and the back-end AI task module in FIG. 1B includes an ER module, an ASR module, an ASV module, and an AEC module in FIG. 1C.

[0078] As shown in FIG. 1C, the output of the format conversion module includes audio data (also referred to as an audio file) and a label. The back-end AI task model set includes ASR, ASV, ER, AEC, etc.

[0079] When the decoder parses the bitstream, the bitstream can include an AI task type identifier, such as ASR, ASV, ER, AEC, etc. The form of the AI task type identifier in the bitstream includes a string and an integer, which can be directly expressed or indirectly mapped. The AI task type identifier is used in the local process module to store or distribute the bitstream including a plurality of data sets through the AI task type identifier.

[0080] When the decoder parses the bitstream, the bitstream can include labels of AI tasks, such as "happy", "sad", "positive" and the like in the ER task. The labels of AI tasks in the bitstream include: a string, an integer, which can be directly expressed or indirectly mapped. The labels can be used as the input of the AI model of the back-end supervised learning.

[0081] When the decoder parses the bitstream, the bitstream can include unique identifiers of audio files corresponding to the labels of AI tasks. The unique identifiers of the audio files are used for the unique correspondence between the labels and the audio files.

[0082] The format conversion module converts the transmission format of the data into a data storage format. The purpose of the format conversion module is to adapt to the training of the back-end task. Since the audio data and the labels are encoded and decoded separately, the format conversion module is used to organize the data.

[0083] In some embodiments, the terminal can be a user equipment (UE), such as at least one of a mobile phone, a wearable device, an Internet of Things device, a communication-capable car, a smart car, a tablet computer (Pad), a wireless transceiver-equipped computer, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal device in industrial control, a wireless terminal device in self-driving, a wireless terminal device in remote medical surgery, a wireless terminal device in a smart grid, a wireless terminal device in transportation safety, a wireless terminal device in a smart city, a wireless terminal device in a smart home, but is not limited thereto.

[0084] It can be understood that the communication system described in the embodiments of the present disclosure is for more clearly illustrating the technical solutions of the embodiments of the present disclosure, and does not constitute a limitation on the technical solutions proposed by the embodiments of the present disclosure. Those skilled in the art can know that, as the system architecture evolves and new business scenarios appear, the technical solutions proposed by the embodiments of the present disclosure are also applicable to similar technical problems.

[0085] The following embodiments of the present disclosure can be applied to the communication system 100 shown in FIG. 1A or part of the subjects, but are not limited thereto. The subjects shown in FIG. 1A are illustrative, and the communication system can include all or part of the subjects in FIG. 1A, or other subjects other than those in FIG. 1A. The number and form of each subject is arbitrary, each subject can be physical or virtual, the connection relationship between each subject is illustrative, each subject can not be connected or can be connected, and the connection can be in any manner, can be direct connection or indirect connection, can be wired connection or wireless connection.

[0086] Embodiments of the present disclosure can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G new radio (NR), 6th generation mobile communication system (6G), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New radio access (NX), Future generation radio access (FX), Global System for Mobile communications (GSM (registered trademark)), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, Ultra-WideBand (UWB), Bluetooth (registered trademark), Public Land Mobile Network (PLMN) network, Device-to-Device (D2D) system, Machine to Machine (M2M) system, Internet of Things (IoT) system, Vehicle-to-Everything (V2X), system using other data processing methods, next-generation system expanded based thereon, and the like. In addition, a plurality of systems can be combined (for example, combination of LTE or LTE-A and 5G, and the like).

[0087] According to different frequencies of sound waves, sound waves can be divided into several frequency bands. Sound waves with a frequency lower than 20 Hz are called infrasound waves; sound waves with a frequency between 20 Hz and 20 KHz are called audible sound waves; and sound waves with a frequency higher than 20 KHz are called ultrasonic waves.

[0088] Ultrasound frequencies typically used for medical diagnostics range from 1 to 5 MHz. Ultrasound has the characteristics of good directionality, strong penetration, easy to obtain relatively concentrated sound energy, and long propagation distance in water. It can be used for ranging, speed measurement, cleaning, welding, stone, etc. It can be used for ultrasonic welding, ultrasonic chemistry, ultrasonic cleaning, ultrasonic processing (drilling, carving, polishing, etc.), ultrasonic therapy, ultrasonic surgery, ultrasonic beauty, ultrasonic motor and ultrasonic suspension.

[0089] In areas where light and radio waves cannot reach, such as oceans and strata, infrasound waves can freely enter and exit. Due to its characteristics, it can be used to explore deeply buried ore deposits, measure the distribution of cold and hot air masses in the stratosphere, and check for hidden dangers in running machines. It can also be used to predict natural phenomena such as tsunamis, storms, volcanic eruptions, and magnetic storms. Therefore, infrasound can be used to detect weather, earthquakes, and predict typhoons and tsunamis.

[0090] The frequency range of sound waves received by the human ear is 20 Hz-20 kHz. Current lossy audio encoding algorithms are designed based on the frequency range of the human ear, and use the masking effect to make the human ear less sensitive to distortion. However, in an increasing number of automated fields, computers are used to "listen" to audio data, such as automatic detection in automated manufacturing plants. Sensors collect audio data during product production. Computers analyze audio data to determine whether products are defective and take action. For subsequent processing, traditional lossy audio encoding algorithms are not suitable because effective audio frequency components will be lost.

[0091] In addition, audio signals in the medical field, such as brain waves (brain waves refer to electrical oscillations produced by nerve cell activity in the human brain. Because this oscillation appears on scientific instruments, it looks like a wave, so it is called brain waves) have a different frequency range than the human ear. According to the frequency, brain waves can be divided into five categories: beta waves (consciousness 14-30 Hz), alpha waves (bridge consciousness 8-14 Hz), theta waves (subconsciousness 4-8 Hz) and delta waves (subconsciousness 4 Hz or lower) and gamma waves (focus on something (30 Hz or above) and so on. The combination of these consciousnesses forms a person's internal and external behavior, emotions, and learning performance. Compressing brain wave signals using traditional lossy audio encoding algorithms also loses audio components.

[0092] Therefore, in the field of using more and more automated control, a new lossless or near-lossless sound encoding algorithm is needed to effectively store sound data and transmit it to computers for analysis.

[0093] The Advanced Audio Coding (AAC) compression algorithm used in the Moving Pictures Experts Group (MPEG) standard uses a psychoacoustic model to achieve lossy compression. The encoding and decoding block diagram is shown in the following figure. The algorithm takes into account the importance of audio components based on the sensitive frequency band of the human ear and the masking effect, and compresses signal components that are not noticeable to the human ear. For a machine listener, this compression will lose effective audio components.

[0094] MPEG-4 SLS (Scalable to Lossless) or MPEG-4 Scalable to Lossless according to "ISO / IEC 14496-3:2005 / Amd 3:2006 (Scalable to Lossless Coding)" is an extension of the MPEG-4 Part 3 (MPEG-4 Audio) standard that allows lossless audio compression of the MPEG-4 General Audio Coding method (e.g. variants of AAC) that is scalable to lossy.

[0095] The MPEG-4 SLS codec utilizes a lossless coding method based on an integer modified discrete cosine transform (IntMDCT), where the IntMDCT spectral data is encoded with two complementary layers, namely a core MPEG-4 AAC layer and a Lossless Enhancement (LLE) layer, the core MPEG-4 AAC layer generates an AAC compliant bitstream at a predefined bit rate that is the smallest rate / quality unit of the lossless bitstream, and the Lossless Enhancement (LLE) layer utilizes a bitplane encoding method to produce a fine granularity of the lossless part of the scalable to lossless bitstream.

[0096] The core layer AAC encoder in MPEG-4 SLS follows the informative AAC encoding specification described in ISO / IEC 14496-3:2009. The encoded information in the core AAC bitstream is then removed from the IntMDCT spectral data by an error mapping process; and the resulting IntMDCT spectral residuals are encoded in the LLE encoder. The error mapping process also manages to preserve in the IntMDCT residuals a probability distribution skew of the original IntMDCT coefficients that approximates a Lapacian distribution, so that they can be very efficiently encoded by the entropy encoder used in the LLE layer.

[0097] In particular, for high sampling rate (96 kHz and above) inputs, the performance of this scalable system is further enhanced by so-called oversampling techniques. In this way, the LLE encoder can run with a preferred longer transform length, while the AAC core codec can run at a more appropriate, lower sampling rate. For example, for an oversampling factor (osf) of 2, the AAC core can run at 48 kHz, while the LLE encoder runs at 96 kHz, with a frame length twice that of the AAC core (i.e. 2048 samples). In this case, the lower 1024 IntMDCT spectrum values can be used as an approximation of the MDCT spectrum values required by the AAC encoder. The error mapping process is still valid. In this case, the quantized AAC spectrum is mapped to the lower portion of the oversampled IntMDCT spectrum.

[0098] The MPEG-4 SLS codec provides a non-core mode for applications that only require lossless quality. This is achieved by simply disabling the AAC core used in the MPEG-4 SLS codec. Tests have shown that the lossless compression performance of MPEG-4 SLS improves by 1-5% if the AAC core is not provided, and since the AAC encoder and decoder do not need to be implemented in the MPEG-4 SLS non-core mode codec, the computational complexity and implementation cost are also significantly reduced.

[0099] The 3GPP Enhanced Voice Services (EVS) compression algorithm uses a multi- core algorithm to achieve efficient compression of different types of audio signals. The algorithm uses a signal analysis module to analyze and classify audio frames. After this module, speech frames are compressed by a linear precoding (LP) based encoding core, music signals are compressed by a frequency domain encoding core, and silent frames are compressed by an Inactive Signal Coding / CNG core. The algorithm takes into account the importance of audio components based on the sensitive frequency band and masking effect of the human ear, and compresses signal components that are not noticeable to the human ear. For a machine listener, compression based on the sensitive frequency band and masking effect of the human ear can lose effective audio components.

[0100] BS.2127 is the Audio Definition Model (ADM) renderer for advanced audio systems. It specifies a reference renderer for use with advanced audio systems as specified in Recommendation ITU-R BS.2051-2 and audio-related metadata as specified in Recommendation ITU-R BS.2076-1 "Audio Definition Model (ADM)", including for programme exchange. The audio renderer converts a set of audio signals with associated metadata into a different configuration of audio signals and metadata based on the provided content metadata and local environment metadata.

[0101] The overall architecture consists of several core components and processing steps. There are three types of data: input metadata, target environment and audio channels. At the initialization of the target environment behavior, the user can select a speaker layout from the ones specified in Recommendation ITU-R BS.2051-2. The rendering itself is divided into sub-components (object renderer, HOA (Higher Order Ambisonics) renderer, Direct Speakers renderer) depending on the typeDefinition of the project.

[0102] The ADM structure contains different levels of description metadata for the content and properties: audioProgramme, audioContent, audioObject. The ADM structure also uses different formats to establish the relationship between the content and the transport / storage channels: audioPackFormat, audioChannelFormat, audioBlockFormat.

[0103] MPEG metadata is a subset of ADM. However, ADM only defines metadata for the rendering system and is not enough for metadata types for machine listening such as types and positions of sensors, temperature, humidity, air pressure, etc.

[0104] MPEG-H Multichannel Coding Tool (MCT) is an adaptive pairing and downmix method that relies on inter-channel correlation. It first finds the two channels with the highest correlation to pair, then searches for the two channels with the highest correlation among the remaining channels, and iterates until all remaining channels have not reached the correlation threshold for pairing. It can effectively eliminate the redundancy between channels.

[0105] AI datasets can be used on multiple AI tasks, while each AI dataset can employ multiple AI models. The current situation is that the dataset formats used by different AI models for the same purpose are not compatible. A unified AI audio dataset format brings many benefits in multiple AI audio tasks, as follows:

[0106] Simplify data processing: A unified format ensures that all audio data is stored under the same structure and standard, reducing the complexity of data preprocessing and data cleaning.

[0107] Improve efficiency: Using a unified format in multiple AI audio tasks can reduce repetitive work, such as format conversion and label alignment.

[0108] Promote model migration and comparison: Models trained under a unified format can be more easily migrated between different tasks, improving the effectiveness of transfer learning. Different researchers can conduct experiments under the same data format, ensuring that model performance comparisons are fair and consistent.

[0109] Dataset expansion and maintenance: When the dataset needs to be expanded with new data or labels, a unified format makes expansion easier.

[0110] Support automation: Standardized data formats can better support automated tools and processes, such as automatic labeling, data augmentation, and data analysis.

[0111] Overall, a unified dataset format helps to promote the progress of AI audio technology research and application.

[0112] Therefore, the present disclosure provides a data processing method, which acquires an encoded dataset, the dataset including data of multiple AI task types, and obtains audio files corresponding to each AI task type based on task type identifiers, facilitating subsequent training, verification, or testing of each AI task using audio files aligned with it, and improving data processing efficiency.

[0113] FIG. 2A is a flowchart of a data processing method according to an embodiment of the present disclosure.

[0114] In the present disclosure, the data processing method can be executed by any device with specific data processing capabilities, such as a server, a terminal, or a combination of a server and a terminal, which is not limited in the present disclosure.

[0115] As shown in FIG. 2A, the present disclosure relates to a data processing method, and the method includes the following steps:

[0116] Step S2101: Acquire an encoded dataset.

[0117] The encoded data set corresponds to a plurality of artificial intelligence (AI) task types, and the encoded data set includes an encoded audio file and encoded metadata.

[0118] In some embodiments, an encoding end can encode an original data set to obtain an encoded data set. The original data set can include an audio file and corresponding metadata, and the encoded data set can be transmitted in the form of a bit stream. A decoding end can receive the encoded data set transmitted by the encoding end, and the encoded data set includes an encoded audio file and encoded metadata. It can be understood that in the embodiments of the present disclosure, the audio file and the metadata are encoded and decoded separately.

[0119] In some embodiments, the encoding end can use a lossy encoding algorithm for encoding, or can use a lossless encoding algorithm for encoding, and the present disclosure does not limit the encoding algorithm.

[0120] In some embodiments, the original data set includes an audio file and corresponding metadata, which can be used for a plurality of AI task types. The AI task types can include at least one of ER, ASR, ASV, and AEC.

[0121] In some embodiments, the encoded metadata can include at least one of the following: an AI task type identifier; a data set type identifier; an AI task label; and a unique identifier of the audio file. The AI task type identifier is used to identify the AI task type corresponding to the audio file. The data set type identifier is used to identify the data set type in which the audio file is located, and the data set type includes at least one of a training data set, a validation data set, and a test data set. The AI task label is used as an input of an AI model of a back-end supervised learning, and the unique identifier of the audio file is used to identify the audio file corresponding to the label in the metadata.

[0122] In some embodiments, the encoded audio file and the encoded metadata are associated through the unique identifier of the audio file. After decoding the encoded data set, the metadata corresponding to the audio file can be determined based on the unique identifier of the audio file, or the audio file corresponding to the metadata can be determined based on the unique identifier of the audio file.

[0123] In some embodiments, when the obtained encoded data set corresponds to a plurality of AI task types, the encoded data set can be grouped to obtain a file corresponding to each AI task type, to facilitate subsequent execution of each AI task.

[0124] Step S2102: decoding the encoded data set to obtain an audio file and metadata corresponding to the audio file.

[0125] In some embodiments, the encoded data set can be decoded to obtain at least one audio file and metadata corresponding to each audio file.

[0126] In some embodiments, the metadata further comprises at least one of: the metadata comprises an AI task type identifier corresponding to the audio file; an identifier of a data set type, the data set type comprising at least one of a training data set, a validation data set, and a test data set; a label of the AI task; and a unique identifier of the audio file.

[0127] In some embodiments, the AI task type identifier is used to identify the AI task type corresponding to the audio file, for example, the AI task type identifier is ER.

[0128] In some embodiments, the identifier of the data set type is used to identify the data set type in which the audio file is located, for example, the data set type is a training data set.

[0129] In some embodiments, the label is used for supervised training of the AI task. The label of the AI task can be used as input of the AI model for supervised learning in the back end. For example, the label of the AI task is "happy"

[0130] In some embodiments, the unique identifier of the audio file is used to identify the audio file corresponding to the label in the metadata. For example, the unique identifier of the audio file is "001".

[0131] In some embodiments, the audio file and the metadata corresponding to the audio file are associated through the unique identifier of the audio file. For example, the unique identifier (audio_id) of the audio file included in the metadata is "001", then the metadata is associated with the audio file identified as "001", and the data included in the metadata and the audio file identified as "001" form a data set, which is used for subsequent training, validation or testing of the AI model.

[0132] In some embodiments, the method can further comprise determining the metadata corresponding to the audio file based on the unique identifier of the audio file.

[0133] In some embodiments, the encoded data set can be decoded to obtain at least one audio file and at least one metadata, based on the unique identifier of the audio file included in the metadata, the audio file corresponding to the metadata can be determined, or the metadata corresponding to the audio file can be determined. Thus, each audio file and its corresponding metadata can be obtained, and when grouping based on the AI task type identifier in the subsequent step, the audio file and its corresponding metadata can be grouped as a whole.

[0134] Step S2103, determining the audio file corresponding to each AI task type based on the metadata.

[0135] In some embodiments, the audio files can be grouped according to the metadata corresponding to the audio files to obtain audio files corresponding to each AI task type.

[0136] For example, the data set includes N audio files with an AI task type of ER and M audio files with an AI task type of ASR, where N and M are positive integers; based on the metadata, the N audio files corresponding to the AI task type of ER are grouped into one group, and the M audio files corresponding to the AI task type of ASR are grouped into one group, so that the N audio files corresponding to the AI task type of ER can be subsequently distributed to an AI model corresponding to ER for processing, and the M audio files corresponding to the AI task type of ASR can be distributed to an AI model corresponding to ASR for processing, which can improve processing efficiency.

[0137] In some embodiments, determining the audio files corresponding to each AI task type based on the metadata includes grouping the audio files based on the AI task type identifier to obtain the audio files corresponding to each AI task type.

[0138] In some embodiments, the metadata includes an AI task type identifier corresponding to the audio file, and the audio file can be grouped based on the AI task type identifier to obtain the audio files corresponding to each AI task type.

[0139] For example, if the AI task type identifier included in the metadata is ER, the audio file corresponding to the metadata can be divided into a group corresponding to ER.

[0140] In some embodiments, the audio file corresponds to one or more groups of metadata, and each AI task type corresponds to one or more labels.

[0141] In an example, one audio file can correspond to one group of metadata, i.e., one task type. For example, the audio file "001" corresponds to the AI task type of ER, the audio file "002" corresponds to the AI task type of ASR, the audio file "003" corresponds to the AI task type of ASV, and the audio file "004" corresponds to the AI task type of AEC. Each audio file corresponds to a different task type.

[0142] In another example, one audio file can correspond to multiple groups of metadata, i.e., multiple task types. For example, the audio file "001" corresponds to the AI task type of ER, the audio file "001" corresponds to the AI task type of ASR, the audio file "001" corresponds to the AI task type of ASV, and the audio file "001" corresponds to the AI task type of AEC. Each AI task type can use the audio file "001" for training.

[0143] In an example, one task type can correspond to one label. For example, the label of the AI task type of ER is "Happy", and the label of the AI task type of ASR is "Hello, world!".

[0144] In another example, one task type can correspond to multiple labels. For example, the labels of the AI task type of ER are "Happy" and "positive".

[0145] In step S2104, the audio file of each AI task type and the corresponding label are stored in the storage container corresponding to each AI task type.

[0146] In the embodiments of the present disclosure, the storage container can be a folder, for example, and each AI task type can correspond to one storage container.

[0147] For example, one or more audio files of the AI task type of ER and the corresponding labels are stored in the storage container corresponding to ER, and one or more audio files of the AI task type of ASR and the corresponding labels are stored in the storage container corresponding to ASR.

[0148] In some embodiments, the audio file of each AI task type and the corresponding metadata can be stored in the storage container corresponding to each AI task type.

[0149] In step S2105, the audio file corresponding to each AI task type is distributed to the AI model corresponding to each AI task type, and the audio file is processed based on the AI model.

[0150] For example, one or more audio files in the storage container corresponding to ER and the corresponding labels are distributed to the ER model, and the ER model can perform corresponding training, verification or testing on the audio file according to the data set type corresponding to the audio file; one or more audio files in the storage container corresponding to ASR and the corresponding labels are distributed to the ASR model, and the ASR model can perform corresponding training, verification or testing on the audio file according to the data set type corresponding to the audio file.

[0151] In the embodiments of the present disclosure, at the decoding end, the received data set includes data of multiple AI task types, the data set is grouped by the task type identifier, and the file corresponding to each AI task type is obtained, which facilitates subsequent training, verification or testing.

[0152] FIG. 2B is a schematic diagram of a data processing method according to an embodiment of the present disclosure.

[0153] As shown in FIG. 2B, the decoder receives a bitstream 1, and the data file included in the bitstream 1 corresponds to multiple task types. Meanwhile, there is one label under the task type corresponding to one data file. The decoder decodes the bitstream 1 to obtain an audio file and multiple task types, the unique identifier (audio id) of the audio file is "001", and the metadata of the audio file includes: the task type (task_type) is "ER", the label (Label) is "Happy", the task type (task_type) is "ASR", the label (Label) is "Hello, world!", the task type (task_type) is "ASV", and the label (Label) is "Speaker1", the task type (task_type) is "AEC", and the label (Label) is "Conversation".

[0154] The format conversion module processes the audio file and the metadata obtained by the decoder to obtain a first group of data: "audio_id": "001", "task_type": "ER", and "label": "happy"; a second group of data: "audio_id": "001", "task_type": "ASR", and "label": "Hello, world!"; a third group of data: "audio_id": "001", "task_type": "ASV", and "label": "speaker1"; and a fourth group of data: "audio_id": "001", "task_type": "AEC", and "label": "Conversation". The first group of data is input to the ER model for training (or verification or testing), the second group of data is input to the ASR model for training (or verification or testing), the third group of data is input to the SAV model for training (or verification or testing), and the fourth group of data is input to the AEC model for training (or verification or testing).

[0155] FIG. 2C is a schematic diagram of a data processing method according to an embodiment of the present disclosure.

[0156] As shown in FIG. 2C, in an example, the decoder receives bitstream 1-bitstream 4, each of which includes a data file corresponding to a task type. Meanwhile, there is one label under the task type corresponding to one data file. The decoder decodes bitstream 1 to obtain an audio file and a task type, the unique identifier (audio id) of the audio file being "001", the task type task_type being "ER", and the label Label being "Happy". The decoder decodes bitstream 2 to obtain an audio file and a task type, the unique identifier of the audio file being "002", the task type task_type being "ASR", and the label Label being "Hello, world!". The decoder decodes bitstream 3 to obtain an audio file and a task type, the unique identifier of the audio file being "003", the task type task_type being "ASV", and the label Label being "Speaker1". The decoder decodes bitstream 4 to obtain an audio file and a task type, the unique identifier of the audio file being "004", the task type task_type being "AEC", and the label Label being "Conversation".

[0157] The format conversion module processes the audio files and metadata decoded by the decoder to obtain first data "audio id": "001", "task type": "ER", "label": "Happy"; second data "audio id": "002", "task type": "ASR", "label": "Hello, world!"; third data "audio id": "003", "task type": "ASV", "label": "Speaker1"; and fourth data "audio id": "004", "task type": "AEC", "label": "Conversation". The first data is input to the ER model for training (or verification or testing), the second data is input to the ASR model for training (or verification or testing), the third data is input to the SAV model for training (or verification or testing), and the fourth data is input to the AEC model for training (or verification or testing).

[0158] As shown in FIG. 2C, in another example, the decoder receives bitstream 1-bitstream 4, each of which includes a data file corresponding to a task type. Meanwhile, there is one or more labels under the task type corresponding to one data file. The decoder decodes bitstream 1 to obtain an audio file and a task type, the unique identifier of the audio file being "001", the task type task_type being "ER", the label Label being "Happy", "Positive". The decoder decodes bitstream 2 to obtain an audio file and a task type, the unique identifier of the audio file being "002", the task type task_type being "ASR", the label Label being "Hello, world!". The decoder decodes bitstream 3 to obtain an audio file and a task type, the unique identifier of the audio file being "003", the task type task_type being "ASV", the label Label being "Speaker1". The decoder decodes bitstream 4 to obtain an audio file and a task type, the unique identifier of the audio file being "004", the task type task_type being "AEC", the label Label being "Conversation", "street".

[0159] The format conversion module processes the audio file and the metadata obtained by the decoder to obtain first data "audio_id": "001", "task_type": "ER", "label": "Happy", "Positive"; second data "audio_id": "002", "task_type": "ASR", "label": "Hello, world!"; third data "audio_id": "003", "task_type": "ASV", "label": "Speaker1"; and fourth data "audio_id": "004", "task_type": "AEC", "label": "Conversation", "street". The first data is input to the ER model for training (or verification or testing), the second data is input to the ASR model for training (or verification or testing), the third data is input to the SAV model for training (or verification or testing), and the fourth data is input to the AEC model for training (or verification or testing).

[0160] FIG. 2D is a schematic diagram of a data processing method according to an embodiment of the present disclosure.

[0161] As shown in FIG. 2D, the decoder receives 4*N bit streams, each of which includes a data file corresponding to a task type. Meanwhile, there is one or more labels under the task type corresponding to one data file. The decoder decodes each bit stream to obtain an audio file, a task type, and one or more labels.

[0162] In an example, the format conversion module processes the audio file, the task type, and the label obtained by the decoder to obtain first group data: "task_type": "ER", "audio_id": "1" + corresponding label,..., "audio_id": "N" + corresponding label; second group data: "task_type": "ASR", "audio_id": "1" + corresponding label,..., "audio_id": "N" + corresponding label; third group data: "task_type": "ASV", "audio_id": "1" + corresponding label,..., "audio_id": "N" + corresponding label; and fourth group data: "task_type": "AEC", "audio_id": "1" + corresponding label,..., "audio_id": "N" + corresponding label. The first group data is input to the ER model for training (or verification or testing), the second group data is input to the ASR model for training (or verification or testing), the third group data is input to the SAV model for training (or verification or testing), and the fourth group data is input to the AEC model for training (or verification or testing).

[0163] In another example, the format conversion module processes the audio file decoded by the decoder, the task type and the label to obtain a first group of data: "task_type": "ER", "task_id": "1", "data_type": "training set", "audio_id": "1" + corresponding label, …, "audio_id": "N" + corresponding label; a second group of data: "task_type": "ASR", "task_id": "2", "data_type": "training set", "audio_id": "1" + corresponding label, …, "audio_id": "N" + corresponding label; a third group of data: "task_type": "ASV", "task_id": "3", "data_type": "training set", "audio_id": "1" + corresponding label, …, "audio_id": "N" + corresponding label; and a fourth group of data: "task_type": "AEC", "task_id": "4", "data_type": "training set", "audio_id": "1" + corresponding label, …, "audio_id": "N" + corresponding label. The first group of data is input to the ER model for training (or verification or testing), the second group of data is input to the ASR model for training (or verification or testing), the third group of data is input to the SAV model for training (or verification or testing), and the fourth group of data is input to the AEC model for training (or verification or testing).

[0164] The data processing method according to the embodiments of the present disclosure can include at least one of steps S2101-S2105. For example, steps S2102-S2103 can be implemented as independent embodiments, but are not limited thereto.

[0165] In some embodiments, step S2101 is optional, and one or more of the steps can be omitted or replaced in different embodiments.

[0166] In some embodiments, step S2104 is optional, and one or more of the steps can be omitted or replaced in different embodiments.

[0167] In some embodiments, step S2105 is optional, and one or more of the steps can be omitted or replaced in different embodiments.

[0168] In some embodiments, other optional implementations described before or after the description of FIG. 2A can be referred to.

[0169] In some embodiments, the name of information and the like is not limited to the name described in the embodiments, and the terms of "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "symbol", "codebook", "codeword", "codepoint", "bit", "data", "program", "chip", and the like can be replaced with each other.

[0170] In some embodiments, the terms of "time", "time point", "time", "time position", and the like can be replaced with each other, and the terms of "time length", "time period", "time window", "window", "time", and the like can be replaced with each other.

[0171] In some embodiments, "acquire", "obtain", "get", "receive", "transmit", "bidirectional transmission", "send and / or receive", and the like can be replaced with each other, and can be interpreted as receiving from other subjects, acquiring from protocols, obtaining from high layers, obtaining by self-processing, autonomously implementing, and the like.

[0172] In some embodiments, the terms of "send", "transmit", "report", "issue", "transmit", "bidirectional transmission", "send and / or receive", and the like can be replaced with each other.

[0173] In some embodiments, the terms of "certain", "preset", "preset", "set", "indicated", "certain", "arbitrary", "first", and the like can be replaced with each other, and "certain A", "preset A", "preset A", "set A", "indicated A", "certain A", "arbitrary A", "first A" can be interpreted as A specified in advance in protocols and the like, can be interpreted as A obtained by setting, configuring, or indicating, and the like, and can be interpreted as certain A, certain A, arbitrary A, or first A, but is not limited thereto.

[0174] In some embodiments, the determining or judging can be performed by a value represented by 1 bit (0 or 1), a true or false value (Boolean value) represented by true or false, or a comparison of numerical values (for example, a comparison with a predetermined value), but is not limited thereto.

[0175] In some embodiments, "not expecting to receive" can be interpreted as not receiving on the time domain resource and / or the frequency domain resource, or as not performing subsequent processing on the data, etc. after receiving the data; "not expecting to send" can be interpreted as not sending, or as sending but not expecting the receiver to respond to the content of the sending.

[0176] FIG. 3 is a flow diagram of a data processing method according to an embodiment of the present disclosure.

[0177] As shown in FIG. 3, the present disclosure relates to a data processing method, and the method comprises:

[0178] In step S3101, a coded data set is obtained.

[0179] The optional implementation of step S3101 can refer to the optional implementation of step S2101 in FIG. 2A and other associated parts in the embodiments involved in FIG. 2A, which will not be repeated here.

[0180] In step S3102, the coded data set is decoded to obtain an audio file and metadata corresponding to the audio file.

[0181] The optional implementation of step S3102 can refer to the optional implementation of step S2102 in FIG. 2A and other associated parts in the embodiments involved in FIG. 2A, which will not be repeated here.

[0182] In step S3103, the audio file corresponding to each AI task type is determined based on the metadata.

[0183] The optional implementation of step S3103 can refer to the optional implementation of step S2103 in FIG. 2A and other associated parts in the embodiments involved in FIG. 2A, which will not be repeated here.

[0184] In step S3104, the audio file corresponding to each AI task type is distributed to the AI model corresponding to each AI task type, and the audio file is processed based on the AI model.

[0185] The optional implementation of step S3104 can refer to the optional implementation of step S2105 in FIG. 2A and other associated parts in the embodiments involved in FIG. 2A, which will not be repeated here.

[0186] The data processing method related to the embodiments of the present disclosure can include at least one of steps S3101-S3104. For example, steps S3102-S3103 can be implemented as independent embodiments, but are not limited thereto.

[0187] In some embodiments, step S3101 is optional, and one or more of the steps can be omitted or replaced in different embodiments.

[0188] In some embodiments, step S3104 is optional, and one or more of the steps can be omitted or replaced in different embodiments.

[0189] FIG. 4 is a flowchart of a data processing method according to an embodiment of the present disclosure.

[0190] As shown in FIG. 4, the embodiments of the present disclosure relate to a data processing method, and the method includes:

[0191] Step S4101 decodes the encoded data set to obtain an audio file and metadata corresponding to the audio file.

[0192] The optional implementation of step S4101 can refer to the optional implementation of step S2102 of FIG. 2A and other associated parts of the embodiments related to FIG. 2A, which will not be repeated here.

[0193] Step S4102 determines the audio file corresponding to each AI task type based on the metadata.

[0194] The optional implementation of step S4102 can refer to the optional implementation of step S2103 of FIG. 2A and other associated parts of the embodiments related to FIG. 2A, which will not be repeated here.

[0195] In some embodiments, the bitstream received by the decoder contains a data compression file corresponding to multiple task types. Meanwhile, there is one label under the task type corresponding to one data file.

[0196] The audio file with the unique file identifier audio_id of “001” carries the following metadata:

[0197] The task type task_type is “ER”, and the label Label is “Happy”

[0198] The task type task_type is “ASR”, and the label Label is “Hello, world!”

[0199] The task type task_type is “ASV”, and the label Label is “Speaker1”

[0200] Task type task_type is "AEC", label is "Conversation"

[0201] In some embodiments, the data compression files contained in the bitstream received by the decoder correspond to one task type. Meanwhile, there is one label under the task type corresponding to one data file.

[0202] The audio file with unique file identifier audio_id "001" carries the following metadata:

[0203] Task type task_type is "ER"

[0204] Multiple labels Labels are "Happy"

[0205] The audio file with unique file identifier audio_id "002" carries the following metadata:

[0206] Task type task_type is "ASR"

[0207] Multiple labels Labels are "Hello, world!"

[0208] The audio file with unique file identifier audio_id "003" carries the following metadata:

[0209] Task type task_type is "ASV"

[0210] Multiple labels Labels are "Speaker1"

[0211] The audio file with unique file identifier audio_id "004" carries the following metadata:

[0212] Task type task_type is "AEC"

[0213] Multiple labels Labels are "Conversation"

[0214] In some embodiments, the data compression files contained in the bitstream received by the decoder correspond to one task type. Meanwhile, there are multiple labels under the task type corresponding to one data file.

[0215] The case of one audio file corresponding to multiple labels exists in file 001 and file 004.

[0216] The audio file with unique file identifier audio_id "001" carries the following metadata:

[0217] task_type is "ER"

[0218] Labels are "Happy" and "Positive"

[0219] Audio files with unique file identifier audio_id "004" carry the following metadata:

[0220] task_type is "AEC"

[0221] Labels are "Conversation" and "Street"

[0222] In some embodiments, for the case of receiving multiple data file bitstreams, the decoder decoder decodes and the format conversion module puts data files of the same task_type and corresponding labels into the same storage container (such as a folder), forming a training data set for subsequent tasks.

[0223] Specifically, 4*N bitstreams are input, decoded by the decoder, and the output of the decoder is processed by the format conversion module.

[0224] The format conversion module puts the audio file set with task_type "ER" (file unique identifier "1"-"N") into storage container 1, a total of N files and corresponding labels for each file, for use in training the ER model on the back end.

[0225] The format conversion module puts the audio file set with task_type "ASR" (file unique identifier "N+1"-"2*N") into storage container 2, a total of N files and corresponding labels for each file, for use in training the ASR model on the back end.

[0226] The format conversion module puts the audio file set with task_type "ASV" (file unique identifier "2*N+1"-"3*N") into storage container 3, a total of N files and corresponding labels for each file, for use in training the ASV model on the back end.

[0227] The format conversion module puts the audio file set with task_type "AEC" (file unique identifier "3*N+1"-"4*N") into storage container 4, a total of N files and corresponding labels for each file, for use in training the AEC model on the back end.

[0228] The multi-task multi-label format obtained by the decoder from the bit stream is as follows, including: decoded task type, decoded task type serial number, decoded data set type, decoded data file unique identification number, and multiple labels of the decoded data file.

[0229] The format is as follows:

[0230] Task type

[0231] Task serial number

[0232] Data set type

[0233] Data file unique identification number

[0234] Multiple labels of data file

[0235] Specific examples are as follows:

[0236] Task type: ER

[0237] Task serial number: 1

[0238] Data set type: Training Set

[0239] Data file unique identification number: 001

[0240] Multiple labels of data file

[0241] The AI training data file for multiple tasks is encoded to form a bit stream, and the bit stream is decoded and format-converted by the decoder to obtain multiple training data sets for multiple AI tasks.

[0242] The bit stream includes an encoding format of the task type, and the decoder decodes the encoding format of the task type to obtain the decoded task type.

[0243] The bit stream includes an encoding format of the task type serial number, and the decoder decodes the encoding format of the task type serial number to obtain the decoded task type serial number.

[0244] The bit stream includes an encoding format of the data set type, and the data set type includes “Training Set”, “Validation Set”, and “Test Set”. The decoder decodes the encoding format of the data set type to obtain the decoded task type serial number.

[0245] The bit stream includes an encoding format of the data file unique identification number, and the decoder decodes the encoding format of the data file unique identification number to obtain the decoded data file unique identification number.

[0246] The bitstream includes the encoding format of the plurality of labels of the data file, and the decoder decodes the encoding format of the plurality of labels of the data file to obtain the plurality of decoded labels of the data file.

[0247] The bitstream includes the encoding format of one or more of the above, and the decoded format after decoding includes one or more of 2-6. The encoding format is lossless. The encoding format can also be slightly damaged, which does not affect the performance of the corresponding model of the backend task.

[0248] The embodiments of the present disclosure further provide a device for implementing any of the above methods, for example, a device including units or modules for implementing the steps performed by the terminal in any of the above methods. For another example, another device is further provided, including units or modules for implementing the steps performed by the network equipment (such as access network equipment, core network function node, core network equipment, etc.) in any of the above methods.

[0249] It should be understood that the division of each unit or module in the above apparatus is only a logical function division, and all or part of them can be integrated into a physical entity or physically separated in actual implementation. In addition, the units or modules in the apparatus can be implemented in the form of processor calling software: for example, the apparatus includes a processor, the processor is connected with a memory, the memory stores instructions, and the processor calls the instructions stored in the memory to realize any of the above methods or realize the functions of each unit or module of the above apparatus, wherein the processor is a general processor such as a central processing unit (CPU) or a microprocessor, and the memory is a memory in the apparatus or a memory outside the apparatus. Alternatively, the units or modules in the apparatus can be implemented in the form of hardware circuit, and the functions of part or all of the units or modules can be realized by the design of hardware circuit. The above hardware circuit can be understood as one or more processors; for example, in one implementation, the above hardware circuit is an application-specific integrated circuit (ASIC), and the functions of part or all of the units or modules are realized by the design of the logical relationship of elements in the circuit; for another example, in another implementation, the above hardware circuit is a programmable logic device (PLD), and a field programmable gate array (FPGA) is taken as an example, which can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by a configuration file, so as to realize the functions of part or all of the above units or modules. All units or modules of the above apparatus can be all implemented in the form of processor calling software, or all implemented in the form of hardware circuit, or part implemented in the form of processor calling software and the remaining part implemented in the form of hardware circuit.

[0250] In the embodiments of the present disclosure, the processor is a circuit with signal processing capability. In one implementation, the processor can be a circuit with instruction reading and running capability, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), a digital signal processor (DSP), or the like. In another implementation, the processor can implement certain functions through a logical relationship of hardware circuits, and the logical relationship of the hardware circuits is fixed or can be reconfigured. For example, the processor is a hardware circuit implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In the reconfigurable hardware circuit, the processor loads a configuration document to implement the configuration of the hardware circuit. It can be understood that the processor loads instructions to implement the functions of the above part or all units or modules. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), and the like.

[0251] FIG. 5 is a structural schematic diagram of a data processing apparatus according to an embodiment of the present disclosure. As shown in FIG. 5, the data processing apparatus 5100 can include a processing module 5101. In some embodiments, the processing module 5101 is configured to decode an encoded data set to obtain an audio file and metadata corresponding to the audio file, the metadata including an AI task type corresponding to the audio file; and determine, based on the metadata, audio files corresponding to each AI task type.

[0252] In some embodiments, the apparatus described above can further include an obtaining module configured to obtain an encoded data set.

[0253] In some embodiments, the metadata further includes at least one of the following: an identification of a data set type, the data set type including at least one of a training data set, a validation data set, and a test data set; a label of an AI task; and a unique identifier of an audio file.

[0254] In some embodiments, the encoded audio file and the metadata corresponding to the audio file are associated through the unique identifier of the audio file.

[0255] In some embodiments, the processing module is configured to determine metadata corresponding to the audio file based on the unique identifier of the audio file.

[0256] In some embodiments, the processing module is configured to group the audio files based on the AI task type identifier, to obtain audio files corresponding to each AI task type.

[0257] In some embodiments, the audio file corresponds to one or more groups of metadata, and each AI task type corresponds to one or more labels.

[0258] In some embodiments, the label is used for supervised training of the AI task.

[0259] In some embodiments, the processing module is configured to store the audio files of each AI task type and the corresponding labels in a storage container corresponding to each AI task type.

[0260] In some embodiments, the processing module is configured to distribute the audio files corresponding to each AI task type to an AI model corresponding to each AI task type, and process the audio files based on the AI model.

[0261] FIG. 6A is a structural schematic diagram of a communication device 6100 according to an embodiment of the present disclosure. The communication device 6100 can be a network device (such as an access network device, a core network device, etc.), a terminal (such as a user equipment, etc.), a chip, a chip system, or a processor supporting the implementation of the above-mentioned any method by the network device, a chip, a chip system, or a processor supporting the implementation of the above-mentioned any method by the terminal, etc. The communication device 6100 can be used to implement the methods described in the above method embodiments, and specific implementation can be referred to the descriptions in the above method embodiments.

[0262] As shown in FIG. 6A, the communication device 6100 includes one or more processors 6101. The processor 6101 can be a general-purpose processor or a special-purpose processor, etc., such as a baseband processor or a central processing unit. The baseband processor can be used to process communication protocols and communication data, and the central processing unit can be used to control the communication device (such as a base station, a baseband chip, a terminal device, a terminal device chip, a DU or a CU, etc.), execute programs, and process data of the programs. Optionally, the communication device 6100 is configured to execute any of the above methods. Optionally, the one or more processors 6101 are configured to invoke instructions to cause the communication device 6100 to execute any of the above methods.

[0263] In some embodiments, the communication device 6100 further includes one or more transceivers 6102. When the communication device 6100 includes one or more transceivers 6102, the transceiver 6102 performs at least one of the communication steps of transmitting and / or receiving in the above-described methods, and the processor 6101 performs at least one of the other steps. In alternative embodiments, the transceiver can include a receiver and / or a transmitter, which can be separate or integrated together. Alternatively, the terms transceiver, transceiving unit, transceiver, transceiving circuit, interface circuit, interface, etc. can be replaced by each other, the terms transmitter, transmitting unit, transmitter, transmitting circuit, etc. can be replaced by each other, and the terms receiver, receiving unit, receiver, receiving circuit, etc. can be replaced by each other.

[0264] In some embodiments, the communication device 6100 further includes one or more memories 6103 for storing data. Alternatively, all or part of the memory 6103 can also be outside the communication device 6100. In alternative embodiments, the communication device 6100 can include one or more interface circuits 6104. Alternatively, the interface circuit 6104 is connected with the memory 6103, and the interface circuit 6104 can be used to receive data from the memory 6103 or other devices, and can be used to send data to the memory 6103 or other devices. For example, the interface circuit 6104 can read the data stored in the memory 6103 and send the data to the processor 6101.

[0265] The communication device 6100 described in the above embodiments can be a network device or a terminal, but the scope of the communication device 6100 described in the present disclosure is not limited thereto, and the structure of the communication device 6100 can not be limited by Figure 6A. The communication device can be a standalone device or can be part of a larger device. For example, the communication device can be: 1) a standalone integrated circuit (IC), or a chip, or a chip system or subsystem; (2) a set of one or more ICs, which can optionally include storage components for storing data, programs; (3) an ASIC, such as a Modem; (4) a module that can be embedded in other devices; (5) a receiver, a terminal device, a smart terminal device, a cellular phone, a wireless device, a handset, a mobile unit, a vehicle-mounted device, a network device, a cloud device, an artificial intelligence device, etc.; (6) others, etc.

[0266] Figure 6B is a structural schematic diagram of a chip 6200 according to an embodiment of the present disclosure. For the case where the communication device 6100 is a chip or a chip system, the structural schematic diagram of the chip 6200 shown in Figure 6B can be referred to, but is not limited thereto.

[0267] The chip 6200 includes one or more processors 6201. The chip 6200 is configured to execute any of the above methods.

[0268] In some embodiments, the chip 6200 further includes one or more interface circuits 6202. Optionally, the terms interface circuit, interface, transceiver pin, etc. can replace each other. In some embodiments, the chip 6200 further includes one or more memories 6203 for storing data. Optionally, all or part of the memory 6203 can be outside the chip 6200. Optionally, the interface circuit 6202 is connected with the memory 6203, the interface circuit 6202 can be used to receive data from the memory 6203 or other devices, and the interface circuit 6202 can be used to send data to the memory 6203 or other devices. For example, the interface circuit 6202 can read the data stored in the memory 6203 and send the data to the processor 6201.

[0269] In some embodiments, the interface circuit 6202 performs at least one of the communication steps (for example, step S2101, but not limited to this) of the above-mentioned method of sending and / or receiving. The interface circuit 6202 performing the communication steps such as sending and / or receiving in the above-mentioned method means that the interface circuit 6202 performs data interaction between the processor 6201, the chip 6200, the memory 6203 or the transceiver device. In some embodiments, the processor 6201 performs at least one of the other steps.

[0270] The modules and / or devices described in each of the embodiments of the virtual device, the physical device, the chip, etc. can be combined or separated as appropriate. Optionally, part or all of the steps can also be performed by multiple modules and / or devices, which are not limited here.

[0271] The disclosure also proposes a storage medium, and the above-mentioned storage medium stores instructions, when the above-mentioned instructions run on the communication device 6100, the communication device 6100 executes any one of the above-mentioned methods. Optionally, the above-mentioned storage medium is an electronic storage medium. Optionally, the above-mentioned storage medium is a computer readable storage medium, but not limited to this, it can also be a storage medium readable by other devices. Optionally, the above-mentioned storage medium can be a non-transitory storage medium, but not limited to this, it can also be a transitory storage medium.

[0272] The disclosure also proposes a program product, and the above-mentioned program product is executed by the communication device 6100, so that the communication device 6100 executes any one of the above-mentioned methods. Optionally, the above-mentioned program product is a computer program product.

[0273] The disclosure also proposes a computer program, when it runs on a computer, so that the computer executes any one of the above-mentioned methods.

Claims

1. A data processing method, characterized by, The method comprises: decoding the encoded data set to obtain an audio file and metadata corresponding to the audio file, the metadata comprising an AI task type identifier corresponding to the audio file; determining, based on the metadata, audio files corresponding to each AI task type.

2. The method of claim 1, wherein, The metadata further comprises at least one of: an identifier of a data set type, the data set type comprising at least one of a training data set, a validation data set, and a test data set; a label of an AI task; a unique identifier of the audio file.

3. The method of claim 2, wherein, The audio file and the metadata corresponding to the audio file are associated through the unique identifier of the audio file.

4. The method of claim 3, wherein, The method further comprises: determining, based on the unique identifier of the audio file, the metadata corresponding to the audio file.

5. The method according to claim 1 or 2, characterized in that, The determining, based on the metadata, of audio files corresponding to each AI task type comprises: grouping the audio files based on the AI task type identifier to obtain audio files corresponding to each AI task type.

6. The method of claim 5, wherein, The audio file corresponds to one or more groups of metadata, and each AI task type corresponds to one or more labels.

7. The method according to claim 2 or 6, characterized in that, The label is used for supervised training of an AI task.

8. The method of claim 5, wherein, The method further comprises: storing the audio files of each AI task type and their corresponding labels in a storage container corresponding to each AI task type.

9. The method according to any one of claims 1 to 8, characterized in that, The method further comprises: distributing the audio files corresponding to each AI task type to an AI model corresponding to each AI task type, and processing the audio files based on the AI model.

10. A data processing apparatus, characterized by, Comprise: a processing module configured to decode an encoded data set to obtain an audio file and metadata corresponding to the audio file, the metadata comprising an AI task type identifier corresponding to the audio file; determine, based on the metadata, audio files corresponding to each AI task type.

11. The apparatus of claim 10, wherein, The processing module is further configured to determine, based on the unique identifier of the audio file, the metadata corresponding to the audio file.

12. The apparatus of claim 10, wherein, The processing module is configured to group the audio files based on the AI task type identifier to obtain audio files corresponding to each AI task type.

13. The apparatus of claim 12, wherein, The processing module is configured to store the audio files of each AI task type and their corresponding labels in a storage container corresponding to each AI task type.

14. The apparatus of any one of claims 10 to 13, wherein, The processing module is further configured to distribute the audio files corresponding to each AI task type to an AI model corresponding to each AI task type, and process the audio files based on the AI model.

15. A communication device, characterized by Comprise: one or more processors; The communication device is configured to execute the method of any one of claims 1 to 9.

16. A storage medium, the storage medium storing instructions, wherein, When the instructions are executed on the communication device, the communication device is caused to execute the method of any one of claims 1 to 9.

17. A program product, characterized by Comprise: a computer program, when executed by a communication device, causes the communication device to execute the method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Abnormal event detection method and device and electronic equipment

    CN113838478A

  • Coding and decoding method and device and storage medium

    CN118160286A

  • Coding and decoding method and device and storage medium

    CN118202406A

  • Communication method and apparatus

    WO2024169600A1