Data processing method, encoder, decoder, and communication system
By introducing metadata into the bitstream and employing intra-frame and inter-frame predictive compression and candidate feature similarity weighted summation methods, the problem of low audio feature encoding and decoding efficiency is solved, and efficient audio feature decoding and machine hearing tasks are supported.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING XIAOMI MOBILE SOFTWARE CO LTD
- Filing Date
- 2025-01-22
- Publication Date
- 2026-07-30
AI Technical Summary
Existing encoding and decoding technologies are inefficient in the process of encoding and decoding audio features, especially in automated tasks, where they cannot effectively transmit and decode audio features, resulting in the loss of critical information.
By introducing metadata into the bitstream, the decoder determines the decoding features of audio frames based on the metadata, and uses intra-frame prediction and inter-frame prediction compression methods, combined with candidate feature similarity and weighted summation, to improve decoding efficiency and accuracy.
It achieves efficient decoding of audio features, suitable for machine hearing tasks such as automatic speech recognition, automatic speaker verification, and emotion recognition, improving the flexibility and accuracy of data processing.
Smart Images

Figure CN2025074079_30072026_PF_FP_ABST
Abstract
Description
Data processing methods, encoders, decoders, and communication systems Technical Field
[0001] This disclosure relates to the field of communication technology, and in particular to a data processing method, encoder, decoder and communication system. Background Technology
[0002] With the rapid development of encoding / decoding and communication technologies, data used for AI tasks can be encoded into encoded data and then transmitted, ensuring transmission efficiency. Furthermore, features for AI models can be extracted from the data, encoded, and transmitted, reducing the amount of data during communication. Summary of the Invention
[0003] This disclosure provides a communication method, encoder, decoder, and communication system to provide a method for improving the efficiency of decoding audio feature codes.
[0004] In a first aspect, embodiments of this disclosure provide a data processing method, the method being executed by a decoder, the method comprising:
[0005] Receive a bitstream, the bitstream including metadata associated with a first audio frame;
[0006] The audio features of the first audio frame after decoding are determined based on the metadata.
[0007] Secondly, this disclosure also provides a data processing method, which is executed by an encoder, and the method includes:
[0008] Send a bitstream, the bitstream including metadata associated with the first audio frame;
[0009] The metadata is used to determine the audio features of the first audio frame after decoding.
[0010] Thirdly, this disclosure also provides a data processing method, including:
[0011] The encoder sends a bitstream, which includes metadata associated with the first audio frame;
[0012] The decoder receives the bit stream;
[0013] The decoder determines the audio features of the first audio frame after decoding based on the metadata.
[0014] Fourthly, embodiments of this disclosure also provide a decoding apparatus, including:
[0015] A transceiver module is used to receive a bit stream, the bit stream including metadata associated with a first audio frame;
[0016] The processing module is used to determine the audio features of the decoded first audio frame based on the metadata.
[0017] Fifthly, embodiments of this disclosure also provide an encoding device, including:
[0018] A transceiver module is used to send a bit stream, the bit stream including metadata associated with a first audio frame;
[0019] The metadata is used to determine the audio features of the first audio frame after decoding.
[0020] Sixthly, embodiments of this disclosure provide a decoder, the decoder comprising:
[0021] One or more processors;
[0022] The decoder is used to execute the data processing method described in the first aspect of the embodiments of this disclosure.
[0023] In a seventh aspect, embodiments of this disclosure provide an encoder, the encoder comprising:
[0024] One or more processors;
[0025] The encoder is used to execute the data processing method described in the second aspect of the embodiments of this disclosure.
[0026] Eighthly, embodiments of this disclosure also provide a communication system, including an encoder and a decoder;
[0027] The decoder is configured to implement the data processing method described in the first aspect, and the encoder is configured to implement the data processing method described in the second aspect.
[0028] Ninthly, embodiments of this disclosure also provide a storage medium storing instructions that, when executed on a communication device, cause the communication device to perform the data processing method as described in the first aspect of this disclosure, or to perform the data processing method as described in the second aspect of this disclosure.
[0029] In a tenth aspect, embodiments of this disclosure also provide a program product, including at least one of a program and instructions, wherein when the program or instructions are executed by a communication device, they implement the data processing method described in the first aspect or the data processing method described in the second aspect.
[0030] In this embodiment of the disclosure, the bitstream received by the decoder includes at least one first piece of audio and video information. The first piece of information includes audio feature encoding and metadata. The metadata is used to reproduce the audio features of the corresponding audio frame in the decoder. By transmitting the first piece of information of the audio frame, this disclosure enables the decoder to adaptively decode the audio feature encoding through the metadata, which can improve data processing, especially the decoding efficiency.
[0031] Additional aspects and advantages of embodiments of this disclosure will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of this disclosure. Attached Figure Description
[0032] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings required for the description of the embodiments are introduced below. The following drawings are only some embodiments of this disclosure and do not impose specific limitations on the protection scope of this disclosure.
[0033] Figure 1 is a schematic diagram of the architecture of the communication system provided in an embodiment of this disclosure;
[0034] Figure 2A is an interactive schematic diagram of the data processing method provided in an embodiment of this disclosure;
[0035] Figure 2B is a schematic diagram of encoding audio features according to an embodiment of this disclosure;
[0036] Figure 2C is a schematic diagram of the encoder processing flow according to an embodiment of the present disclosure;
[0037] Figure 2D is a schematic flowchart of the process for determining the second feature provided in an embodiment of this disclosure;
[0038] Figure 2E is a schematic flowchart of the process for determining the second feature provided in an embodiment of this disclosure;
[0039] Figure 2F is an interactive schematic diagram of the data processing method provided in the embodiments of this disclosure;
[0040] Figure 3A is a schematic flowchart illustrating a data processing method applied to an encoder according to an embodiment of this disclosure;
[0041] Figure 3B is a schematic flowchart illustrating a data processing method applied to an encoder according to an embodiment of this disclosure;
[0042] Figure 4A is a schematic flowchart illustrating a data processing method applied to a decoder according to an embodiment of this disclosure;
[0043] Figure 4B is a schematic flowchart illustrating a data processing method applied to a decoder according to an embodiment of this disclosure;
[0044] Figure 5 is a flowchart illustrating the data processing method according to an embodiment of this disclosure;
[0045] Figure 6A is a schematic diagram of the structure of the decoding device proposed in an embodiment of this disclosure;
[0046] Figure 6B is a schematic diagram of the structure of the encoding device proposed in an embodiment of this disclosure;
[0047] Figure 7 is a schematic diagram of the structure of the communication device proposed in an embodiment of this disclosure;
[0048] Figure 8 is a schematic diagram of the chip structure proposed in an embodiment of this disclosure. Detailed Implementation
[0049] This disclosure presents a data processing method, an encoder, a decoder, and a communication system.
[0050] In a first aspect, embodiments of this disclosure propose a data processing method, executed by a decoder, the method comprising:
[0051] Receive a bitstream, the bitstream including metadata associated with a first audio frame;
[0052] The audio features of the first audio frame after decoding are determined based on the metadata.
[0053] In the above embodiments, the problem of how the decoder obtains the decoded audio features based on the bitstream is solved, ensuring that the bitstream received by the decoder can determine the decoded audio features based on metadata, ensuring the accuracy of obtaining audio features based on the received bitstream, and improving decoding efficiency.
[0054] In conjunction with some embodiments of the first aspect, in some embodiments, the audio features of the decoded first audio frame are determined based on the metadata, including any one of the following:
[0055] Based on the metadata, the encoded audio features of the first audio frame are decoded to obtain the decoded audio features of the first audio frame, wherein the bitstream further includes the encoded audio features of the first audio frame;
[0056] Based on the metadata and the first feature, the audio features of the decoded first audio frame are determined, wherein the first feature is determined by the first information in the metadata.
[0057] In the above embodiments, two methods are provided to obtain the features of the first audio frame after encoding based on metadata decoding. When the bitstream also includes the audio features after encoding the first audio frame, the audio features after encoding the first audio frame are decoded based on the metadata to obtain the decoded audio features. When the bitstream does not include the audio features after encoding the first audio frame, the first feature is determined first through the first information in the metadata, and the audio features after decoding the first audio frame are determined based on the metadata and the first feature. The above embodiments can obtain the decoded audio features based on the metadata according to different bitstream conditions, thereby improving the flexibility of decoding.
[0058] In conjunction with some embodiments of the first aspect, in some embodiments, when the metadata is used to indicate the intra-frame prediction method, the audio features of the first audio frame are subjected to intra-frame prediction compression based on the intra-frame prediction method.
[0059] When the metadata is used to indicate the index of the second audio frame and the compression coefficient of the intra-frame prediction, the audio features of the first audio frame are subjected to inter-frame prediction compression based on the audio features of the second audio frame and the compression coefficient of the inter-frame prediction.
[0060] In the above embodiments, the decoder can determine the compression method of the audio features of the first audio frame based on the different metadata indications.
[0061] In conjunction with some embodiments of the first aspect, the metadata is used to indicate an intra-frame predictive compression method;
[0062] Decoding the encoded audio features of the first audio frame based on the metadata includes:
[0063] The audio features encoded from the first audio frame are decoded based on the intra-frame prediction compression method.
[0064] In the above embodiments, when the audio features of the first audio frame are encoded by intra-frame predictive compression, the encoded audio features of the first audio frame can be decompressed by using the intra-frame predictive compression method indicated by the metadata through reverse operation, thereby obtaining the encoded audio features of the first audio frame. This embodiment clarifies the intra-frame predictive compression method through metadata, thereby helping the decoder to accurately obtain the audio features through the reverse process.
[0065] In conjunction with some embodiments of the first aspect, in some embodiments, the first information is an index of the second audio frame, and the metadata also includes the compression coefficient of the inter-frame prediction of the second audio frame;
[0066] Determining the audio features of the decoded first audio frame based on the metadata and the first feature includes:
[0067] The audio features of the decoded second audio frame are determined based on the index of the second audio frame and used as the first feature; the decoding time of the second audio frame is earlier than the decoding time of the first audio frame.
[0068] Based on the first feature and the compression coefficient of the inter-frame prediction, the audio features of the decoded first audio frame are determined.
[0069] In the above embodiments, the metadata includes first information specifically the index of the second audio frame, and the metadata also includes the compression coefficient of the inter-frame prediction of the second audio frame. Thus, the decoder can determine the decoded audio features of the second audio frame from each of the obtained decoded audio features based on the index of the second audio frame, and use it as the first feature to reverse determine the decoded audio features of the first audio frame using the compression coefficient based on the first feature and the inter-frame prediction.
[0070] In the above embodiments, by recording the index of the first feature and the coding coefficients of inter-frame prediction in the metadata, the process of inter-frame prediction compression can be clearly defined, thereby helping the decoder to accurately obtain audio features through the reverse process.
[0071] In conjunction with some embodiments of the first aspect, in some embodiments, the metadata is generated by a compression method corresponding to the candidate feature with the highest similarity to the audio feature;
[0072] Different candidate features correspond to different compression methods.
[0073] In the above embodiments, unlike related technologies that determine which predictive compression mode to use based on the parity of audio frames, this embodiment can generate multiple candidate features based on multiple compression methods, and generate metadata for the compression method corresponding to the candidate feature with the highest similarity to the audio feature among multiple candidate features, laying the foundation for the decoder to decode the feature with the least distortion.
[0074] In conjunction with some embodiments of the first aspect, each candidate feature is obtained by compressing the audio features of the first audio frame using a compression method;
[0075] Each compression method corresponds to an intra-frame prediction compression method, or to the audio features of a second audio frame and the compression coefficients of inter-frame prediction.
[0076] In the above embodiments, each candidate feature is obtained by compressing the audio features of the first audio frame using a compression method. The compression method can be an intra-frame prediction compression method, or it can be an inter-frame prediction compression method using the audio features of a second audio frame and the compression coefficient of inter-frame prediction. This expands the range of candidate features, thereby obtaining the first feature with the least distortion and higher accuracy.
[0077] In some embodiments, audio features and candidate features include feature elements of the same dimension;
[0078] The similarity is obtained by weighted summation of the differences between the audio features and candidate features in each dimension of the feature elements;
[0079] In this context, for any given dimension, the weight of that dimension is positively correlated with its importance.
[0080] In the above embodiments, by weighted summation of the differences between the feature elements of audio features and candidate features in each dimension, the similarity between features can be accurately measured. Furthermore, different dimensions are assigned weights, which are related to the importance of that dimension in subsequent machine hearing tasks, thereby obtaining audio features that are more suitable for machine hearing tasks.
[0081] In conjunction with some embodiments of the first aspect, in some embodiments, similarity is characterized in any of the following ways:
[0082] Euclidean distance;
[0083] Cosine similarity;
[0084] Jaccard similarity coefficient;
[0085] Hamming distance.
[0086] In the above embodiments, the similarity between features can be obtained in a variety of ways, which improves the flexibility of similarity calculation.
[0087] In conjunction with some embodiments of the first aspect, in some embodiments, the audio features are used for machine hearing tasks.
[0088] In conjunction with some embodiments of the first aspect, in some embodiments, the machine hearing task includes at least one of the following:
[0089] Automatic speech recognition;
[0090] Automatic speaker verification;
[0091] Emotion recognition, and
[0092] Audio event classification.
[0093] In the above embodiments, the types of machine hearing tasks are expanded, thereby ensuring the comprehensiveness of the featured application scenarios.
[0094] Secondly, embodiments of this disclosure provide a data processing method, executed by an encoder, the method comprising:
[0095] Send a bitstream, the bitstream including metadata associated with the first audio frame;
[0096] The metadata is used to determine the audio features of the first audio frame after decoding.
[0097] In conjunction with some embodiments of the second aspect, in some embodiments, before sending the bit stream, any one of the following is further included:
[0098] The audio features of the first audio frame are compressed using an intra-frame prediction method to obtain the encoded audio features of the first audio frame. The bitstream also includes the encoded audio features of the first audio frame.
[0099] Determine the compression coefficient based on the audio features of the second audio frame and the inter-frame prediction, and perform inter-frame prediction compression on the audio features of the first audio frame.
[0100] In conjunction with some embodiments of the second aspect, in some embodiments, when the metadata is used to indicate the intra-frame prediction method, the audio features of the first audio frame are subjected to intra-frame prediction compression based on the intra-frame prediction method.
[0101] When the metadata is used to indicate the index of the second audio frame and the compression coefficient of the intra-frame prediction, the audio features of the first audio frame are subjected to inter-frame prediction compression based on the audio features of the second audio frame and the compression coefficient of the inter-frame prediction.
[0102] In conjunction with some embodiments of the second aspect, in some embodiments, the metadata is generated by a compression method corresponding to the candidate feature with the highest similarity to the audio feature;
[0103] Different candidate features correspond to different compression methods.
[0104] In conjunction with some embodiments of the second aspect, in some embodiments, each of the candidate features is obtained by compressing the audio features of the first audio frame using a compression method;
[0105] Each compression method corresponds to an intra-frame prediction compression method, or to the audio features of a second audio frame and the compression coefficients of inter-frame prediction.
[0106] In conjunction with some embodiments of the second aspect, in some embodiments, the audio features and candidate features include the same number of feature elements of the same dimension;
[0107] The similarity is obtained by weighted summation of the differences between the audio features and candidate features in each dimension of the feature elements;
[0108] In this context, for any given dimension, the weight of that dimension is positively correlated with its importance.
[0109] In conjunction with some embodiments of the second aspect, in some embodiments, similarity is characterized in any of the following ways:
[0110] Euclidean distance;
[0111] Cosine similarity;
[0112] Jaccard similarity coefficient;
[0113] Hamming distance.
[0114] In conjunction with some embodiments of the second aspect, in some embodiments, audio features are used for machine hearing tasks.
[0115] In conjunction with some embodiments of the second aspect, in some embodiments, the machine hearing task includes at least one of the following:
[0116] Automatic speech recognition;
[0117] Automatic speaker verification;
[0118] Emotion recognition, and
[0119] Audio event classification.
[0120] Thirdly, this disclosure also provides a data processing method, including:
[0121] The encoder sends a bitstream, which includes metadata associated with the first audio frame;
[0122] The decoder receives the bit stream;
[0123] The decoder determines the audio features of the first audio frame after decoding based on the metadata.
[0124] Fourthly, embodiments of this disclosure also provide a decoding apparatus, including:
[0125] A transceiver module is used to receive a bit stream, the bit stream including metadata related to a first audio frame;
[0126] The processing module is used to determine the audio features of the decoded first audio frame based on the metadata.
[0127] Fifthly, embodiments of this disclosure also provide an encoding device, including:
[0128] A transceiver module is used to send a bit stream, the bit stream including metadata related to a first audio frame;
[0129] The metadata is used to determine the audio features of the first audio frame after decoding.
[0130] Sixthly, embodiments of this disclosure provide a decoder, the decoder comprising:
[0131] One or more processors;
[0132] The decoder is used to execute the data processing method described in the first aspect of the embodiments of this disclosure.
[0133] In a seventh aspect, embodiments of this disclosure provide an encoder, the encoder comprising:
[0134] One or more processors;
[0135] The encoder is used to execute the data processing method described in the second aspect of the embodiments of this disclosure.
[0136] Eighthly, embodiments of this disclosure also provide a system including an encoder and a decoder;
[0137] The decoder is configured to implement the data processing method described in the first aspect, and the encoder is configured to implement the data processing method described in the second aspect.
[0138] Ninthly, embodiments of this disclosure also provide a storage medium storing instructions that, when executed on a communication device, cause the communication device to perform the data processing method as described in the first aspect of this disclosure, or to perform the data processing method as described in the second aspect of this disclosure.
[0139] In a tenth aspect, embodiments of this disclosure provide a program product that, when executed by a communication device, causes the communication device to perform the method as described in an optional implementation of the first or second aspect.
[0140] In one aspect, embodiments of this disclosure provide a computer program that, when run on a computer, causes the computer to perform the methods described in an optional implementation of the first or second aspect.
[0141] In a twelfth aspect, embodiments of this disclosure provide a chip or chip system. The chip or chip system includes processing circuitry configured to perform the method described in an optional implementation of the first or second aspect above.
[0142] It is understood that the aforementioned encoding device, decoding device, encoder, decoder, communication system, storage medium, program product, computer program, chip or chip system are all used to execute the data processing method proposed in the embodiments of this disclosure. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects in the corresponding methods, and will not be repeated here.
[0143] This disclosure provides a data processing method, an encoder, a decoder, and a communication system. In some embodiments, the terms "data processing method" and "signal transmission method," "wireless frame transmission method," etc., can be used interchangeably, as can the terms "information processing system," "communication system," etc.
[0144] This disclosure is not exhaustive, but merely illustrative of some embodiments, and is not intended to limit the scope of protection of this disclosure. Unless otherwise specified, each step in a particular embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a particular embodiment can also be implemented as an independent embodiment, and the order of the steps in a particular embodiment can be arbitrarily interchanged. Furthermore, the optional implementation methods in a particular embodiment can be arbitrarily combined; moreover, the embodiments can be arbitrarily combined, for example, some or all steps of different embodiments can be arbitrarily combined, and a particular embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.
[0145] In each of the disclosed embodiments, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of the embodiments are consistent and can be referenced by each other. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.
[0146] The terminology used in the embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure.
[0147] In the embodiments disclosed herein, "multiple" refers to two or more.
[0148] In some embodiments, the terms “at least one of A or B, at least one of A and B”, “one or more”, “a plurality of”, “multiple”, etc., may be used interchangeably.
[0149] In some embodiments, the notation "at least one of A and B", "A and / or B", "A in one case, B in another", "in response to one case A, in response to another case B", etc., may include the following technical solutions depending on the situation: in some embodiments, A (execute A regardless of whether there is a branch B); in some embodiments, B (execute B regardless of whether there is a branch A); in some embodiments, execution is selected from A and B (A and B are selectively executed); in some embodiments, both A and B are executed. The same applies when there are more branches such as A, B, C, etc.
[0150] In some embodiments, the notation "A or B" may include the following technical solutions, depending on the situation: in some embodiments, A (execute A regardless of whether a branch B exists); in some embodiments, B (execute B regardless of whether a branch A exists); in some embodiments, execution is selected from A and B (A and B are selectively executed). The same applies when there are more branches such as A, B, and C.
[0151] The prefixes "first," "second," etc., used in the embodiments of this disclosure are merely for distinguishing different descriptive objects and do not impose restrictions on the position, order, priority, quantity, or content of the descriptive objects. The description of the descriptive objects is found in the claims or the context of the embodiments, and the use of prefixes should not constitute unnecessary restrictions. For example, if the descriptive object is a "field," the ordinal numbers preceding "field" in "first field" and "second field" do not restrict the position or order of the "fields." "First" and "second" do not restrict whether the "fields" they modify are in the same message, nor do they restrict the order of "first field" and "second field." Similarly, if the descriptive object is a "level," the ordinal numbers preceding "level" in "first level" and "second level" do not restrict the priority between "levels." Furthermore, the number of descriptive objects is not limited by ordinal numbers and can be one or more. For example, in "first device," the number of "devices" can be one or more. Furthermore, the objects modified by different prefixes can be the same or different. For example, if the object being described is "device", then "first device" and "second device" can be the same device or different devices, and their types can be the same or different. Similarly, if the object being described is "information", then "first information" and "second information" can be the same information or different information, and their content can be the same or different.
[0152] In some embodiments, “including A,” “containing A,” “for indicating A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.
[0153] In some embodiments, terms such as "time / frequency" and "time-frequency domain" refer to the time domain and / or frequency domain.
[0154] In some embodiments, terms such as “in response to…”, “in response to determining…”, “in the case of…”, “when…”, “when…”, “if…”, etc. can be used interchangeably. These descriptions all refer to the device making a corresponding action under certain objective circumstances. They do not necessarily limit the time, nor do they require the device to make a judgment action when implementing it, nor do they mean that there must be other limitations.
[0155] In some embodiments, the terms “greater than,” “greater than or equal to,” “not less than,” “more than,” “more than or equal to,” “not less than,” “higher than,” “higher than or equal to,” “not lower than,” and “above” can be used interchangeably, as can the terms “less than,” “less than or equal to,” “not greater than,” “less than,” “less than or equal to,” “not more than,” “lower than,” “lower than or equal to,” “not higher than,” and “below”.
[0156] In some embodiments, devices, etc., may be interpreted as physical or virtual, and their names are not limited to those described in the embodiments. Terms such as “device,” “equipment,” “circuit,” “network element,” “network function,” “network device,” “function,” “node,” “unit,” “section,” “system,” “network,” “chip,” “chip system,” “entity,” and “subject” are interchangeable.
[0157] In some embodiments, "network" can be interpreted as devices included in a network (e.g., access network devices, core network devices, etc.).
[0158] In addition, terms such as "uplink" and "downlink" can be replaced with terms corresponding to inter-terminal communication (e.g., "side"). For example, uplink channel and downlink channel can be replaced with side channel, and uplink link and downlink link can be replaced with side link.
[0159] In some embodiments, "link" can mean "connection" or "link"; in various embodiments, "connection" and "link" can be used interchangeably.
[0160] In some embodiments, the acquisition of data, information, etc., may comply with the laws and regulations of the country where the location is situated.
[0161] In some embodiments, data, information, etc., may be obtained with the user's consent.
[0162] Furthermore, each element, each row, or each column in the table of this disclosure can be implemented as an independent embodiment, and any combination of any element, any row, or any column can also be implemented as an independent embodiment.
[0163] Figure 1 is a schematic diagram of the architecture of a communication system according to an embodiment of the present disclosure.
[0164] As shown in Figure 1, the communication system 100 includes an encoding end 101 and a decoding end 102.
[0165] In some embodiments, the encoding end 101 can be any electronic device with processing capabilities, such as a server or a terminal. The decoding end 102 can be any electronic device with processing capabilities, such as a server or a terminal.
[0166] In some embodiments, the encoding end 101 can encode (or compress) the raw data. The raw data may be the audio features of an audio frame.
[0167] In some embodiments, the encoding end 101 compresses the original data to meet transmission or storage needs, and the output of the encoding end 101 may be a bitstream. For example, in scenarios where the network or bus with limited transmission bandwidth cannot meet the real-time transmission requirements of the original data, the encoding end 101 compresses the original data. As another example, in scenarios where storage devices with limited storage space cannot meet the storage requirements of the original data, the encoding end 101 compresses the original data. The bitstream includes metadata for at least one audio frame. For any audio frame, the metadata of the audio frame is used to determine the audio characteristics of the decoded audio frame in the decoder.
[0168] In some embodiments, the decoding end 102 can receive an encoded dataset sent by the encoding end 101, which may be transmitted in the form of a bitstream. The decoding end 102 can decode the dataset to obtain metadata corresponding to each audio frame, use the metadata to decode the encoded audio features, determine the audio features of the corresponding audio frame, and distribute the decoded audio features to various AI tasks for processing.
[0169] In some embodiments, the terminal may be a user equipment (UE), including, but not limited to, at least one of the following: mobile phone, wearable device, Internet of Things device, car with communication function, smart car, tablet computer, computer with wireless transceiver function, virtual reality (VR) terminal device, augmented reality (AR) terminal device, wireless terminal device in industrial control, wireless terminal device in self-driving, wireless terminal device in remote medical surgery, wireless terminal device in smart grid, wireless terminal device in transportation safety, wireless terminal device in smart city, and wireless terminal device in smart home.
[0170] It is understood that the communication system described in this disclosure is for the purpose of more clearly illustrating the technical solutions of this disclosure, and does not constitute a limitation on the technical solutions proposed in this disclosure. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions proposed in this disclosure are also applicable to similar technical problems.
[0171] The following embodiments of this disclosure can be applied to the communication system 100 shown in FIG1, or to some of the main bodies, but are not limited thereto. The main bodies shown in FIG1 are illustrative. The communication system may include all or some of the main bodies in FIG1, or may include other main bodies outside of FIG1. The number and form of each main body are arbitrary. Each main body may be physical or virtual. The connection relationship between the main bodies is illustrative. The main bodies may not be connected or may be connected. The connection can be in any way, it can be a direct connection or an indirect connection, it can be a wired connection or a wireless connection.
[0172] The embodiments disclosed herein can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G new radio (NR), 6th generation mobile communication system (6G), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New radio access (NX), Future generation radio access (FX), Global System for Mobile communications (GSM), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi), IEEE 802.16 (WiMAX), and IEEE 802.20, Ultra-Wideband (UWB), Bluetooth (a registered trademark), Public Land Mobile Network (PLMN) networks, Device-to-Device (D2D) systems, Machine-to-Machine (M2M) systems, Internet of Things (IoT) systems, Vehicle-to-Everything (V2X) systems, systems utilizing other data processing methods, and next-generation systems built upon them, etc. Furthermore, multiple systems can be combined (e.g., a combination of LTE or LTE-A with 5G).
[0173] Sound waves can be divided into several frequency bands based on their different frequencies. Sound waves with frequencies below 20 Hz are called infrasound; sound waves with frequencies between 20 Hz and 20 kHz are called audible sound waves; and sound waves with frequencies above 20 kHz are called ultrasound.
[0174] Ultrasound frequencies typically used in medical diagnosis range from 1 to 5 MHz. Ultrasound is characterized by its good directionality, strong penetrating power, ease of obtaining relatively concentrated sound energy, and long propagation distance in water. It can be used for ranging, speed measurement, cleaning, welding, and stone removal. It can also be applied to ultrasonic welding, ultrasonic chemistry, ultrasonic cleaning, ultrasonic processing (drilling, carving, polishing, etc.), ultrasonic therapy, ultrasonic surgery, ultrasonic beauty treatments, ultrasonic motors, and ultrasonic levitation.
[0175] Infrasound can freely travel through areas virtually inaccessible to light and radio waves, such as the ocean and underground strata. Due to its properties, it can be used to explore deeply buried mineral deposits, measure the distribution of hot and cold air masses in the stratosphere, and inspect operating machinery for potential hazards. It can also be used to predict natural phenomena such as tsunamis, storms, volcanic eruptions, and geomagnetic storms. Therefore, infrasound can be used to detect weather patterns, earthquakes, and forecast typhoons and tsunamis.
[0176] The human ear can perceive sound waves in the frequency range of 20Hz-20kHz. Current lossy audio coding algorithms are designed based on this frequency range and use masking effects to make distortion less perceptible to the human ear. However, in increasingly automated fields, computers are used to "listen" to audio data, such as automated inspection in automated manufacturing plants. Sensors collect audio data during product manufacturing. Computers analyze the audio data to determine if the product is defective and take action. For subsequent processing, traditional lossy audio coding algorithms are unsuitable because valid audio frequency components are lost.
[0177] Furthermore, audio signals in the medical field, such as brainwaves (brainwaves are electrical oscillations produced by the activity of nerve cells in the human brain. Because these oscillations appear as waves on scientific instruments, they are called brainwaves) have a different frequency range compared to the human ear. Based on frequency, brainwaves can be divided into five categories: beta waves (conscious 14-30Hz), alpha waves (bridged conscious 8-14Hz), theta waves (subconscious 4-8Hz), delta waves (subconscious 4Hz or lower), and gamma waves (focused on something (30Hz or higher)). The combination of these conscious states shapes a person's internal and external behavior, emotions, and learning performance. Compressing brainwave signals using traditional lossy audio encoding algorithms also results in the loss of audio components.
[0178] Therefore, in the field of increasingly automated control, there is a need for a new lossless or near-lossless sound encoding algorithm to effectively store sound data and transmit it to a computer for analysis.
[0179] The Advanced Audio Coding (AAC) compression algorithm used in the Moving Pictures Experts Group (MPEG) standard employs a psychoacoustic model to achieve lossy compression. The encoding and decoding block diagram is shown below. This algorithm considers the importance of audio components based on the sensitive frequency bands of the human ear and masking effects, compressing signal components that are not noticeable to the human ear. For machine listeners, this compression will result in the loss of valid audio components.
[0180] According to "ISO / IEC 14496-3:2005 / Amd 3:2006 (Scalable to Lossless Coding)," MPEG-4 SLS (Scalable to Lossless) or MPEG-4 Scalable to Lossless is an extension of the MPEG-4 Part 3 (MPEG-4 Audio) standard that allows lossless audio compression that can be scaled down to lossy using MPEG-4 common audio coding methods (e.g., variants of AAC).
[0181] The MPEG-4 SLS codec utilizes a lossless coding method based on the integer modified discrete cosine transform (IntMDCT). IntMDCT spectral data is encoded using two complementary layers: the core MPEG-4 AAC layer and the Lossless Enhanced (LLE) layer. The core MPEG-4 AAC layer generates an AAC-compliant bitstream at a predefined bit rate that constitutes the minimum rate / quality unit of the lossless bitstream. The Lossless Enhanced (LLE) layer uses a bit-plane coding method to produce a fine-grained portion that can be scaled to the lossless portion of the lossless bitstream.
[0182] The core layer AAC encoder in MPEG-4 SLS follows the information-rich AAC coding specification described in [reference needed]. The encoded information in the core AAC bitstream is then removed from the IntMDCT spectral data through an error mapping process; the resulting IntMDCT spectral residuals are then encoded in the LLE encoder. This error mapping process also attempts to preserve the probability distribution skew of the original IntMDCT coefficients, which approximates a Lapacian distribution, in the IntMDCT residuals, so that they can be encoded very efficiently by the entropy encoder used in the LLE layer.
[0183] In particular, for high sampling rate (96 kHz and above) inputs, the performance of this scalable system is further enhanced by a technique known as oversampling. In this way, the LLE encoder can operate at a preferred longer transform length, while the AAC core codec can operate at a more suitable, lower sampling rate. For example, with an oversampling factor (osf) of 2, the AAC core can operate at 48 kHz, while the LLE encoder operates at 96 kHz, resulting in a frame length twice that of the AAC core (i.e., 2048 samples). In this case, the lower 1024 IntMDCT spectral value can be used as an approximation of the MDCT spectral value required by the AAC encoder. The error mapping process remains effective. In this case, the quantized AAC spectrum is mapped to the lower portion of the oversampled IntMDCT spectrum.
[0184] The MPEG-4 SLS codec provides a non-core mode for applications that only require lossless quality. This is achieved by simply disabling the AAC core used in the MPEG-4 SLS codec. Tests have shown that without the AAC core, the lossless compression performance of MPEG-4 SLS improves by 1% to 5%, and because the AAC encoder and decoder do not need to be implemented in the MPEG-4 SLS non-core mode codec, its computational complexity and implementation cost are also significantly reduced.
[0185] The 3GPP Enhanced Voice Services (EVS) compression algorithm uses a multi-core algorithm to achieve efficient compression of different types of audio signals. This algorithm uses a signal analysis module to analyze and classify audio frames. After this module, speech frames are compressed by a coding core based on linear precoding (LP), music signals are compressed by a frequency domain coding core, and silent frames are compressed by an inactive signal coding / CNG core. This algorithm considers the importance of audio components based on the sensitive frequency bands of the human ear and masking effects, and compresses signal components that are not obvious to the human ear. For machine monitors, compression based on the sensitive frequency bands of the human ear and masking effects will lose effective audio components.
[0186] In distributed speech recognition (DSR), the speech recognition front-end (FE) and back-end (BE) are separate. The front-end is located on the terminal device, such as a cellular phone, while the back-end is located in the telecommunications network.
[0187] In this way, the front end can capture high-quality speech and convert it in real time into feature extraction parameters or feature vectors, providing a relevant and compact representation of speech information for recognition. The feature vectors are transmitted as a bitstream with strong error protection to the back end—a high-performance speech recognition engine that performs the actual recognition task.
[0188] The recognition results are transmitted back to the terminal via the network. The concept of distributed speech recognition provides a framework for extracting high-quality feature vectors that are unaffected by speech coding or transmission effects. It has low computational requirements on the terminal and fully utilizes high-performance, multi-user speech recognition engines and other resources within the network.
[0189] In wireless networks, especially in bandwidth-constrained wireless networks, it is necessary to compress audio characteristics to reduce bandwidth usage.
[0190] Recommendation BS.2127, formally known as Audio Definition Model (ADM) Renderer for Advanced Sound Systems, specifies a baseline renderer for use with the audio-related metadata defined in Recommendation ITU-R BS.2051-2 and Recommendation ITU-R BS.2076-1, specifically the Audio Definition Model (ADM), including for program exchange. The audio renderer, based on provided content metadata and local context metadata, transforms a set of audio signals with related metadata into an audio signal and metadata with different configurations.
[0191] The overall architecture consists of several core components and processing steps. There are three types of data: input metadata, target environment, and audio channels. During target environment behavior initialization, the user can select a speaker layout from the speaker layout specified in "Recommendation ITU-R BS.2051-2". The rendering itself is divided into sub-components (object renderer, HOA (High-Order Stereo) renderer, and DirectSpeakers renderer) based on the project type (typeDefinition).
[0192] The ADM structure contains different levels describing the content and attributes of metadata: audioProgramme, audioContent, and audioObject. The ADM structure also uses different formats to establish the relationship between content and transmission / storage channels: audioPackFormat, audioChannelFormat, and audioBlockFormat.
[0193] MPEG metadata is a subset of ADM. However, ADM only defines metadata for rendering systems, which is insufficient for metadata types used for machine listening, such as sensor type and location, temperature, humidity, air pressure, etc.
[0194] When transmitting bitstreams, the related technologies do not include metadata for reproducing audio features, which leads to low efficiency in decoding audio features.
[0195] Based on this, embodiments of this disclosure provide a data processing method, which receives a bitstream, the bitstream including first information of at least one audio frame, the first information including metadata and audio feature encoding; wherein, for any audio frame, the metadata of the audio frame is used to reproduce the audio features of the audio frame in a decoder; for each audio frame, the audio feature encoding of the audio frame is decoded according to the metadata of the audio frame to obtain the audio features of the audio frame; since the bitstream includes metadata for re-encoding the audio, the decoder can decode the audio features more quickly, thereby improving data processing efficiency.
[0196] Figure 2A is an interactive schematic diagram of a data processing method according to an embodiment of the present disclosure. As shown in Figure 2A, the method includes:
[0197] In step S2101, the encoder obtains the audio features of at least one audio frame.
[0198] In some embodiments, mono or multi-channel audio data is acquired by sensors. This audio data can be used for machine hearing tasks and optionally human hearing (monitoring) tasks. The audio data is segmented into at least one audio frame, typically 10-50 milliseconds in length. Feature extraction is then performed on the segmented audio frames to obtain their audio features.
[0199] In some embodiments, audio features can be traditional audio features, such as Mel-Frequency Cepstral Coefficients (MFCCs), or features generated during the development of an artificial intelligence deep learning model. In sound processing, MFCC represents the short-term power spectrum of sound based on the linear cosine transform of the logarithmic power spectrum at a nonlinear Mel-frequency scale. MFCC combines two analyses: cepstral analysis and Mel-frequency analysis. Cepstral analysis aims to extract the signal envelope carrying the most relevant information. It uses the inverse Fourier transform (IFT) and a low-pass filter (LPF) to extract the coefficients representing the signal envelope of the logarithmic power spectrum. Mel-frequency analysis treats the signal as an object processed by the human auditory system, passing the signal spectrum through a Mel filter that simulates the filtering process of the human auditory system.
[0200] In some embodiments, audio features typically have multiple dimensions. Taking MFCC as an example, if a 32nd-order Mel filter is used, the obtained MFCC features are 32-dimensional, that is, they include 32 feature elements. It can be understood that the dimensions of the obtained MFCC features will also be different when using Mel filters of different orders.
[0201] In some embodiments, audio features can be used for machine hearing tasks. Machine hearing tasks according to embodiments of this disclosure may include automatic speech recognition, automatic speaker verification, emotion recognition, and audio event classification, etc.
[0202] In step S2102, the encoder performs encoding to obtain a bit stream.
[0203] In some embodiments, the encoder encodes the audio features of the audio frame to obtain the encoded audio features.
[0204] In some embodiments, the encoder may compress the audio features of the audio frame using either intra-frame predictive compression mode or inter-frame predictive compression mode.
[0205] In some embodiments, intra-frame prediction compression refers to a method of compressing the audio features of an audio frame by considering only the data within the audio features themselves, without considering the relationships between the audio features of other frames. Taking MFCC features as an example, this means considering the relationships between the 32-dimensional feature elements in the MFCC features of an audio frame and compressing the MFCC features of that audio frame. Intra-frame prediction compression reduces redundant information by utilizing the correlation of the data (feature elements) within the audio features themselves, thereby achieving the purpose of compression.
[0206] In some embodiments, inter-frame predictive compression refers to a method of compressing the audio features of an audio frame by considering the correlation between the audio features of one audio frame and the audio features of other audio frames. For example, for the 10th audio frame in time sequence, the audio features of the 10th audio frame are compressed by analyzing the relationship between the audio features of this audio frame and the audio features of the 8th frame. Inter-frame predictive compression reduces redundant information by utilizing the correlation between multiple audio features, thereby achieving the purpose of compression.
[0207] In some embodiments, the audio features of odd-numbered frames can be compressed using an intra-frame prediction compression mode, and the audio features of even-numbered frames can be compressed using an inter-frame prediction mode. Alternatively, the audio features of odd-numbered frames can be compressed using an inter-frame prediction compression mode, and the audio features of even-numbered frames can be compressed using an intra-frame prediction mode.
[0208] Please refer to Figure 2B, which exemplarily illustrates a schematic diagram of encoding audio features according to an embodiment of this disclosure. The audio feature of the nth frame is represented as v(n). It should be noted that n is a positive integer in each embodiment of this application. For ease of description, as shown in the figure, the odd-numbered frame is the first frame, and the audio feature of the first frame is compressed using the intra-frame prediction compression mode. The even-numbered frame is the second frame, and the audio feature of the second frame is compressed using the inter-frame prediction compression mode. The audio feature is an MFCC feature with 32-dimensional feature elements. In the figure, V(1)
[0031] represents the 31st-dimensional feature element in the audio feature V(1) of the first frame. The input for inter-frame prediction compression of the even-numbered frame is the audio feature of the adjacent previous frame after compression and decompression and the audio feature of the even-numbered frame. Therefore, the audio feature of the first frame needs to be decompressed after compression to obtain the intra-frame decompressed feature of the odd-numbered frame, i.e., V. d (1) The outputs of intra-frame prediction compression and inter-frame prediction compression obtained by the encoder in this embodiment of the present disclosure are multiplexed into a bitstream and sent to the decoder.
[0209] In some embodiments, if the encoder determines to perform intra-frame prediction compression on the audio features of the first audio frame based on an intra-frame prediction method to obtain the encoded audio features of the first audio frame, then the bitstream also includes the encoded audio features of the first audio frame.
[0210] In some embodiments, where metadata is used to indicate an intra-prediction method, the audio features of the first audio frame are subjected to intra-prediction compression based on the intra-prediction method.
[0211] In some embodiments, the encoder determines the compression coefficient based on the audio features of the second audio frame and the inter-frame prediction, and performs inter-frame prediction compression on the audio features of the first audio frame. Then, the bitstream only needs to transmit the metadata of the first audio frame, which can be used by the decoder to determine the audio features of the decoded first audio frame based on the metadata.
[0212] In some embodiments, where metadata is used to indicate the index of the second audio frame and the compression factor of the intra-frame prediction, the audio features of the first audio frame are subjected to inter-frame prediction compression based on the audio features of the second audio frame and the compression factor of the inter-frame prediction.
[0213] In some embodiments, whether an audio frame uses intra-frame predictive compression or inter-frame predictive compression is no longer fixed by the parity of the audio frame, but is determined based on some criteria.
[0214] In some embodiments, for an audio frame, multiple candidate features related to the audio frame can be obtained. Each candidate feature uniquely corresponds to a compression method, and each compression method corresponds to an intra-frame prediction compression method, or to the audio features of a second audio frame and the compression coefficients of inter-frame prediction. Then, based on some criteria, the candidate feature with the least distortion after decompression is determined from among the candidate features. It is understood that by selecting the candidate feature with the least distortion as the encoded audio feature of the audio frame, the audio feature obtained by decoding on the decoder side will be closer to the audio feature of the audio frame than other candidate features. Therefore, the metadata of the audio frame in this embodiment can be generated by the compression method corresponding to the candidate feature with the highest similarity to the audio feature.
[0215] It should be understood that there are many specific intra-predictive compression methods for intra-frame prediction compression modes. Different intra-predictive compression methods will produce different results when compressing the same audio feature. Therefore, the compression control parameter corresponding to the intra-predictive compression mode is the intra-predictive compression method of the audio feature of the audio frame.
[0216] The intra-frame prediction compression method in this disclosure can employ Differential Pulse Code Modulation (DPCM), Discrete Cosine Transform (DCT) coding, or intra-frame prediction compression methods in the H.26x series coding, etc. It is understood that the input for compression using the intra-frame prediction compression method is the audio features of the audio frame, and the output is the compressed audio features. In some cases, the compressed audio features are also referred to as the encoded audio features.
[0217] In some embodiments, using audio features from different reference frames to perform inter-frame prediction compression on the audio features of the current frame will result in different inter-frame compression results. Furthermore, even if the audio features of the same reference frame are used, the features obtained after compression by different intra-frame prediction compression methods will be different, which will also result in different inter-frame compression results. In addition, when different inter-frame prediction compression coefficients are used to perform inter-frame prediction compression on the intra-frame compression features of the same reference frame, different inter-frame compression results will also be obtained. Therefore, in the embodiments of this disclosure, for the inter-frame prediction compression mode, the metadata includes the audio features of a second audio frame and the inter-frame prediction compression coefficients.
[0218] In this embodiment, a unique index can be set for each audio frame. Since the encoding order of each audio frame by the encoder and the decoding order of each audio frame by the decoder are the same, when the decoder decodes the audio features of a first audio frame encoded with inter-frame predictive compression, it must have already obtained the decoded audio features of a second audio frame encoded with inter-frame predictive compression. Therefore, based on the index of the second audio frame in the metadata, the decoded audio features of the second audio frame can be obtained.
[0219] In some embodiments, the coding coefficients for inter-frame prediction may include scaling factors and weighting coefficients. When an audio feature is compressed for inter-frame prediction based on a unique first feature, the first feature is compressed based on the scaling factor (in some cases, the two can be multiplied, and in other cases, they can be divided), which is the compressed audio feature. When an audio feature is compressed for inter-frame prediction based on multiple first features, different weighting coefficients need to be configured for each first feature. It can be understood that the sum of the weighting coefficients configured for all first features is 1. That is, after the first features are weighted and summed based on the weighting coefficients, the summed feature is scaled by the scaling factor, which is the compressed audio feature.
[0220] For example, for audio frame n, the candidate features associated with that audio frame may include:
[0221] Candidate feature 1, which corresponds to intra-frame prediction compression method 1, is a feature obtained by compressing and then decompressing the audio features of audio frame n using intra-frame prediction compression method 1.
[0222] Candidate feature 2, which corresponds to intra-frame prediction compression method 1, is a feature obtained by compressing and then decompressing the audio features of audio frame n using intra-frame prediction compression method 2.
[0223] Candidate feature 3 corresponds to the index of audio frame n-1 and the inter-frame prediction coding coefficient a. This candidate feature 3 is obtained by performing inter-frame prediction compression on the audio features of audio frame n using the intra-frame decompression features of audio frame n-1 and the inter-frame prediction coding coefficient a. The intra-frame decompression features of the audio frame refer to the features obtained by compressing and then decompressing the audio features of the audio frame using the intra-frame prediction compression method.
[0224] Candidate feature 4 corresponds to the index of audio frame n-2 and the inter-frame prediction coding coefficient b. This candidate feature 4 is obtained by performing inter-frame prediction compression on the audio features of audio frame n using the intra-frame decompression features of audio frame n-2 and the inter-frame prediction coding coefficient b.
[0225] Please refer to Figure 2C, which exemplarily illustrates the processing flow diagram of the encoder according to an embodiment of this disclosure. For ease of description, the figure describes the encoding process of the audio features of the first and second audio frames. The input of the encoder is the audio feature V(1) of the first audio frame and the audio feature V(2) of the second audio frame. The audio feature V(1) is only applicable to the intra-frame prediction compression mode. After compression by the intra-frame prediction compression method 1, the encoded audio feature of the first audio frame is obtained and multiplexed into the bitstream. It should be noted that in some embodiments, although the audio feature V(1) is only applicable to the intra-frame prediction compression mode, considering the existence of multiple optional intra-frame prediction compression methods, the metadata of the first audio frame can also be generated and transmitted to the bitstream. The metadata of the first audio frame indicates the intra-frame prediction compression method 1.
[0226] For the encoding process of audio feature V(2), multiple candidate features can be obtained first. In this embodiment, the candidate features include:
[0227] The features obtained by compressing audio features V(2) using intra-frame prediction compression are called intra-frame prediction compressed features V. e (2); and
[0228] Intra-frame decompression feature V d (1) The features obtained through inter-frame prediction compression are called inter-frame prediction compression features V. e_inter (2).
[0229] Then, the intra-frame prediction compression feature V e (2) Perform decompression processing to obtain intra-frame decompression features V d (2); For inter-frame prediction compression features V e_inter (2) The features obtained by decompression are called inter-frame decompression features V. d_inter (2).
[0230] The above V is evaluated using some standards. d (2) and V d_inter (2) Perform analysis (understandably, the analysis needs to be combined with audio feature V(2)), and determine the feature with less distortion than audio feature V(2) from the two. Use the feature before decompression corresponding to this feature as the feature after encoding the second audio frame. For example, if V d (2) Compared to audio feature V(2), which has the least distortion, then V e (2) The encoded audio features V of the second audio frame ex (2). Further based on V ex (2) Construct metadata using the corresponding compression method, and combine the metadata and V ex(2) The data is multiplexed into the bitstream and then sent to the decoder. At this point, the bitstream includes encoded audio features and metadata. The encoded audio features are also known as V. e (2) is the intra-frame predictive compression coding of the second frame;
[0231] For example, if V d_inter (2) Compared to audio feature V(2), which has the least distortion, then V e inter (2) Encode the audio V of the second audio frame ex (2) Further based on V ex (2) The corresponding compression mode constructs metadata. At this time, the bitstream includes metadata, which indicates the index of the first audio frame.
[0232] In some embodiments, the present disclosure can determine the degree of distortion of a candidate feature relative to an audio feature by calculating the similarity between the candidate feature after decompression and the audio feature. The present disclosure does not limit the specific method of similarity calculation; similarity can be characterized by Euclidean distance, cosine similarity, or Jaccard similarity coefficient between features.
[0233] In some embodiments, the formula for the Euclidean distance between two features A(n) and Ad(n) is:
[0234] d(A(n),A d (n))=sqrt(Σ(A(n)[i]-A d (n)[i]) 2 )
[0235] Where A(n)[i] is the i-th dimension feature element of feature A(n), Ad(n)[i] is the i-th dimension feature element of feature Ad(n), and the summation Σ is performed on all dimensions. In this formula, each term (A(n)[i]-Ad(n)[i]) 2 The square of the difference between the feature elements of the i-th dimension is represented by the sum of the squared differences of all dimensions, and the square root of the sum is taken to obtain the Euclidean distance between the two features.
[0236] In some embodiments, the importance of different dimensions of a feature varies. For example, different dimensions correspond to different frequency values, and different frequencies play different roles in machine hearing tasks. Therefore, when calculating the differences between different dimensions, a weighting coefficient alpha is used. The higher the weighting coefficient alpha, the greater the importance of that dimension. The formula is as follows:
[0237] d w (A(n),A d(n))=sqrt(Σα[i]*(A(n)[i]-A d (n)[i]) 2 )
[0238] Where α[i] represents the weighting coefficient of the i-th dimension of the feature.
[0239] In some embodiments, metadata may be encoded and input into the bitstream; in other embodiments, metadata may be directly input into the bitstream.
[0240] Please refer to Figure 2D, which exemplarily illustrates a flowchart of determining a second feature according to an embodiment of this disclosure. As shown in the figure, this embodiment involves two features that are compared with the audio feature for similarity, namely, the intra-frame decompression feature V obtained by compressing and decompressing the audio feature V(2) of the second audio frame using an intra-frame predictive compression mode. d (2), and the inter-frame prediction compression feature V e_inter (2) Inter-frame decompression features V obtained by decompression d_inter (2) Calculate the intra-frame decompression features V respectively. d (2) The similarity result between the audio feature V(2) and the inter-frame decompression feature V d_inter (2) The similarity result 2 between the audio feature V(2) and the result 1 is compared with the result 2. The feature before decompression (i.e., the candidate feature) with the larger similarity is taken as the second feature. The metadata is constructed based on the compression method of generating the second feature and transmitted to the bit stream. On the other hand, if the second feature is obtained based on intra-frame prediction compression, then the second feature, i.e. the encoded audio feature V, is taken as the second feature. ex (2) Transmitted into the bitstream, if the second feature is obtained based on inter-frame prediction compression, then the metadata record is used for audio feature V. ex (2) The corresponding reference frame index, which is the index of the first audio frame.
[0241] Please refer to Figure 2E, which exemplarily illustrates a flowchart of determining a second feature according to another embodiment of this disclosure. Compared to the embodiment shown in Figure 2D, this embodiment involves more features that need to be compared with audio features for similarity:
[0242] The audio features V(2) of the second audio frame are obtained by performing intra-frame prediction compression and decompression using intra-frame prediction compression method 1. d_method1 (2); It can be understood that the candidate feature corresponding to this feature is: the feature V generated by intra-frame prediction compression of audio feature V(2) using intra-frame prediction compression method 1. e_method1 (2);
[0243] The audio feature V(2) is obtained by performing intra-frame prediction compression and then decompression using intra-frame prediction compression method 2. d_method2 (2); It can be understood that the candidate feature corresponding to this feature is: the feature V generated by intra-frame prediction compression of audio feature V(2) using intra-frame prediction compression method 2. e_method2 (2);
[0244] Based on the intra-frame decompression features of reference frame 1, the audio feature V(2) is obtained by performing inter-frame predictive compression and then decompression. d_inter1 (2); It can be understood that the candidate feature corresponding to this feature is the feature V that performs inter-frame prediction compression on audio feature V(2) based on the intra-frame decompression feature of reference frame 1. e_inter1 (2);
[0245] The inter-frame decompression feature V is obtained by performing inter-frame predictive compression and then decompression on audio feature V(2) based on the intra-frame decompression features of reference frame 2. d_inter2 (2); It can be understood that the candidate feature corresponding to this feature is the feature V obtained by performing inter-frame prediction compression on audio feature V(2) based on the intra-frame decompression feature of reference frame 2. e_inter2 (2);
[0246] Based on the intra-frame decompression features of reference frame 3, the audio feature V(2) is obtained by performing inter-frame predictive compression and then decompression. d_inter3 (2); It can be understood that the candidate feature corresponding to this feature is the feature V obtained by performing inter-frame prediction compression on audio feature V(2) based on the intra-frame decompression feature of reference frame 3. e_inter3 (2);
[0247] Based on the intra-frame decompression features of reference frame 4, the audio feature V(2) is obtained by performing inter-frame predictive compression and then decompression. d_inter4 (2); It can be understood that the candidate feature corresponding to this feature is the feature V obtained by performing inter-frame prediction compression on audio feature V(2) based on the intra-frame decompression feature of reference frame 4. e_inter4 (2);
[0248] The similarity between each of the above features and the audio feature V(2) is calculated to obtain results 1 to 6. Then, the results 1 to 6 are compared, and the candidate feature with the largest similarity is taken as the second feature. Based on the predicted compression mode and compression control parameters for generating the second feature, the metadata is transmitted to the bitstream. On the other hand, if the predicted compression mode corresponding to the second feature is the intra-frame predicted compression mode, then the second feature, i.e., the encoded audio feature V, is used. ex (2) When transmitted to the bitstream, if the predicted compression mode corresponding to the second feature is the inter-frame predicted compression mode, then the metadata record Vex (2) The index of the corresponding reference frame.
[0249] Step S2103: The encoder sends the bit stream.
[0250] In this embodiment of the disclosure, after the encoder obtains the bit stream, it can send the bit stream, and the encoder can then receive the bit stream.
[0251] Step S2104: The decoder receives the bit stream.
[0252] Step S2105: The decoder decodes the bitstream.
[0253] In some embodiments, the bitstream includes metadata associated with the first audio frame. The decoder determines the decoded audio features of the first audio frame based on the metadata.
[0254] In some embodiments, the audio features of the decoded first audio frame are determined based on metadata, including any one of the following:
[0255] Based on the metadata, the encoded audio features of the first audio frame are decoded to obtain the audio features of the first audio frame. The bitstream also includes the encoded audio features of the first audio frame.
[0256] Based on metadata and a first feature, the audio features of the decoded first audio frame are determined, wherein the first feature is determined by first information in the metadata.
[0257] In some embodiments, the decoder parses the bitstream to obtain metadata associated with the first audio frame and the encoded audio features of the first audio frame. The metadata of the audio frame indicates an intra-frame predictive compression method. The decoder can then reverse process the encoded audio features of the first audio frame according to the intra-frame predictive compression method to obtain the decoded audio features.
[0258] In some embodiments, the decoder parses the bitstream to obtain metadata associated with the first audio frame, the metadata of which indicates the index of the second audio frame and the compression coefficient of the intra-frame prediction. The decoder can obtain the decoded audio features of the second audio frame based on the index of the second audio frame, and then combine the compression coefficient of the inter-frame prediction indicated by the metadata to obtain the decoded audio features.
[0259] It should be noted that the embodiments disclosed herein involve audio features before encoding by the encoder, and also involve audio features obtained by decoding by the decoder. Encoding audio features by the encoder may cause data loss, and decoding audio features obtained by the decoder may also cause data loss. Therefore, the audio features before encoding by the encoder and the audio features obtained by decoding by the decoder may not be completely the same.
[0260] Alternatively, in the embodiments of this disclosure, both the encoder and the decoder are lossless processes during encoding and decoding. Therefore, the audio features before encoding by the encoder are the same as the audio features obtained by the decoder.
[0261] The audio features decoded in this embodiment can be used for training or inference of machine hearing tasks. The machine hearing tasks in this embodiment can include automatic speech recognition, automatic speaker verification, emotion recognition, and audio event classification, etc.
[0262] Please refer to Figure 2F, which exemplarily illustrates the interactive schematic diagram of the data processing method of this disclosure embodiment. As shown in the figure, for ease of description, this embodiment is described using the encoding and decoding process of the audio features of the first frame and the second frame. The audio feature V(1) of the first frame is compressed using the intra-frame predictive compression mode. The compression result is sent to the decoder as the encoded audio feature of the first frame via a bitstream. The decoder decodes the encoded audio feature of the first frame using the intra-frame predictive decompression method to obtain the decoded audio feature of the first frame. When determining the predictive compression mode of the second frame, the compressed audio feature of the first frame is decompressed to obtain the intra-frame decompressed feature V of the first frame. d (1) Perform intra-frame prediction compression on the audio features V(2) of the second frame to obtain the intra-frame prediction compressed features V. e (2) Further compress the intra-frame prediction feature V e (2) Decompress the data to obtain the intra-frame decompression feature V of the second frame. d (2) Utilizing the intra-frame decompression feature V of the first frame d (1) Perform inter-frame prediction compression on the audio features V(2) of the second frame to obtain the inter-frame compression features V of the second frame. e_inter (2) Then, the inter-frame compression feature V e_inter (2) Decompress the data to obtain the inter-frame decompression features V. d_inter (2) Calculate V d (2) and V d_inter (2) Calculate the similarity between each feature and V(2) to determine the feature V with the highest similarity. ex (2)(V e (2) or V e_inter (2)), based on feature V ex (2) Generate metadata based on the predicted compression mode and compression control parameters. If the predicted compression mode corresponding to the second feature is the intra-frame predicted compression mode, then the second feature, i.e., the encoded audio feature V, will be... ex (2) When transmitted to the bitstream, if the predicted compression mode corresponding to the second feature is the inter-frame predicted compression mode, then the metadata record Vex (2) The index of the corresponding reference frame, that is, the index of the first frame.
[0263] The decoder determines whether to perform intra-frame prediction decompression or inter-frame prediction decompression based on the metadata.
[0264] If intra-frame prediction decompression is used, then based on the intra-frame prediction compression method included in the metadata, the inverse operation is performed on feature V. ex (2) Decode the audio features V(2) after decoding.
[0265] If inter-frame prediction decompression is used, the decoded V(1) is obtained by using the index of the record in the metadata, and the decoded audio features V(2) are obtained by performing the inverse operation based on the coding coefficients of inter-frame prediction and the decoded V(1).
[0266] The data processing method involved in the embodiments of this disclosure may include at least one of steps S2101 to S2105. For example, steps S2104 to S2105 may be implemented as independent embodiments, but are not limited thereto.
[0267] In some embodiments, step S2101 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0268] In some embodiments, step S2102 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0269] In some embodiments, step S2103 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0270] In some embodiments, step S2104 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0271] In some embodiments, step S2105 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0272] In some embodiments, other alternative implementations may be described before or after the specification corresponding to FIG2A.
[0273] In some embodiments, the names of information, etc., are not limited to the names described in the embodiments. Terms such as "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "bit", "data", "program", and "chip" can be used interchangeably.
[0274] In some embodiments, terms such as “moment,” “point in time,” “time,” and “time location” can be used interchangeably, as can terms such as “duration,” “segment,” “time window,” “window,” and “time.”
[0275] In some embodiments, terms such as wireless access scheme and waveform can be used interchangeably.
[0276] In some embodiments, terms such as "certain," "preset," "default," "set," "indicated," "a certain," "any," and "first" can be used interchangeably. "Certain A," "preset A," "default A," "set A," "indicated A," "a certain A," "any A," and "first A" can be interpreted as A pre-defined in a protocol or the like, or as A obtained through setting, configuration, or instruction, or as specific A, a certain A, any A, or first A, but are not limited thereto.
[0277] In some embodiments, the determination or judgment can be made by a value represented by 1 bit (0 or 1), or by a true or false value (boolean), or by a comparison of numerical values (e.g., a comparison with a predetermined value), but is not limited thereto.
[0278] In some embodiments, "not expecting to receive" can be interpreted as not receiving on time domain resources and / or frequency domain resources, or as not performing subsequent processing on the data after receiving it; "not expecting to send" can be interpreted as not sending, or as sending but not expecting the receiver to respond to the sent content.
[0279] Figure 3A is a schematic flowchart of a data processing method according to an embodiment of the present disclosure, applied to an encoder. As shown in Figure 3A, the present disclosure relates to a data processing method, which includes:
[0280] Step S3101: The encoder obtains the audio features of the first audio frame.
[0281] The optional implementation of step S3101 can be found in step S2101 of Figure 2A and other related parts in the embodiments involved in Figure 2A, which will not be repeated here.
[0282] In step S3102, the encoder performs encoding to obtain a bit stream.
[0283] The optional implementation of step S3102 can be found in step S2102 of Figure 2A and other related parts in the embodiments involved in Figure 2A, which will not be repeated here.
[0284] Step S3103: The encoder sends the bit stream.
[0285] The optional implementation of step S3103 can be found in step S2103 of Figure 2A and other related parts in the embodiments involved in Figure 2A, which will not be repeated here.
[0286] The feature indication method involved in the embodiments of this disclosure may include at least one of steps S3101 to S3103. For example, step S3101 may be implemented as a standalone embodiment, step S3102 may be implemented as a standalone embodiment, step S3103 may be implemented as a standalone embodiment, or at least two steps may be combined, but it is not limited thereto.
[0287] In some embodiments, step S3101 is optional, step S3102 is optional, and step S3103 is optional. In different embodiments, one or more of these steps may be omitted or substituted. However, this is not a limitation.
[0288] Figure 3B is a schematic flowchart illustrating a data processing method according to an embodiment of the present disclosure, applied to an encoder. As shown in Figure 3B, this disclosure relates to a feature indication method, which includes:
[0289] Step S3201: The encoder sends the bit stream.
[0290] The optional implementation of step S3201 can be found in step S2103 of Figure 2A, step S3103 of Figure 3A, and other related parts in the embodiments involved in Figures 2A and 3A, which will not be repeated here.
[0291] Figure 4A is a flowchart illustrating a data processing method according to an embodiment of the present disclosure, applied to a decoder. As shown in Figure 4A, this disclosure relates to a data processing method, which includes:
[0292] Step S4101: The decoder receives the bit stream.
[0293] The optional implementation of step S4101 can be found in step S2104 of Figure 2A and other related parts in the embodiment involved in Figure 2A, which will not be repeated here.
[0294] Step S4102: The decoder decodes the bitstream.
[0295] Optional implementations of step S4102 can be found in step S2105 of Figure 2A and other related parts in the embodiments involved in Figure 2A, which will not be repeated here.
[0296] Figure 4B is a flowchart illustrating a data processing method according to an embodiment of the present disclosure, applied to a decoder. As shown in Figure 4B, this embodiment of the present disclosure relates to a data processing method, which includes:
[0297] In step S4201, the decoder receives a bitstream, which includes metadata associated with the first audio frame.
[0298] The optional implementation of step S4201 can be found in step S2104 of Figure 2A and other related parts in the embodiment involved in Figure 2A, which will not be repeated here.
[0299] Step S4202: Determine the audio features of the decoded first audio frame based on the metadata.
[0300] Optional implementations of step S4202 can be found in step S2105 of Figure 2A and other related parts in the embodiments involved in Figure 2A, which will not be repeated here.
[0301] Figure 5 is a flowchart illustrating a data processing method according to an embodiment of the present disclosure. As shown in Figure 5, the present disclosure relates to a data processing method, which includes:
[0302] Step S5101: The encoder sends the bit stream.
[0303] In some embodiments, the encoder may compress the audio features of the audio frame using either intra-frame predictive compression mode or inter-frame predictive compression mode.
[0304] In some embodiments, if the encoder determines to perform intra-predictive compression on the audio features of the first audio frame based on an intra-predictive method to obtain the encoded audio features of the first audio frame, then the bitstream includes metadata related to the first audio frame and the encoded audio features of the first audio frame. Accordingly, the metadata is used to indicate the intra-predictive method.
[0305] In some embodiments, if the encoder determines to perform inter-frame prediction compression on the audio features of the first audio frame based on an inter-frame prediction method, then the bitstream includes metadata related to the first audio frame.
[0306] In some embodiments, the encoder performs inter-frame prediction compression on the audio features of the first audio frame based on the audio features of the second audio frame and the compression coefficient of the inter-frame prediction. Correspondingly, elements are defined to indicate the index of the second audio frame and the compression coefficient of the intra-frame prediction.
[0307] In some embodiments, metadata is generated by compression methods corresponding to candidate features with the highest similarity to audio features. Different candidate features correspond to different compression methods, and each compression method corresponds to an intra-frame prediction compression method, or to the compression coefficients of audio features and inter-frame predictions of a second audio frame.
[0308] In some embodiments, audio features and candidate features include the same number of feature elements in each dimension. The similarity is obtained by weighted summation of the differences between the feature elements of the audio features and candidate features in each dimension. Furthermore, for any dimension, the weight of the dimension is positively correlated with the importance of the dimension.
[0309] Step S5102: The decoder receives the bit stream.
[0310] In some embodiments, the bitstream includes metadata associated with the first audio frame. The decoder determines the audio characteristics of the decoded first audio frame based on the metadata.
[0311] In some embodiments, when the encoder compresses the audio features of the first audio frame using intra-frame prediction, the bitstream also includes the compressed audio features of the first audio frame.
[0312] In step S5103, the decoder determines the audio features of the decoded first audio frame based on the metadata.
[0313] In some embodiments, the decoder parses the bitstream to obtain metadata associated with the first audio frame and the encoded audio features of the first audio frame. The metadata of the audio frame indicates an intra-frame predictive compression method. The decoder can then reverse process the encoded audio features of the first audio frame according to the intra-frame predictive compression method to obtain the decoded audio features.
[0314] In some embodiments, the decoder parses the bitstream to obtain metadata associated with the first audio frame. The metadata of the audio frame indicates the index of the second audio frame and the compression coefficient of the intra-frame prediction. The decoder can obtain the decoded audio features of the second audio frame from the decoded audio features of each audio frame according to the index of the second audio frame, and then perform reverse processing in combination with the compression coefficient of the inter-frame prediction indicated by the metadata to obtain the decoded audio features.
[0315] In some embodiments, the above methods may include the methods of the embodiments of the communication system side, encoder side, decoder side, etc., which will not be described in detail here.
[0316] In some embodiments, the steps and their optional implementations in other embodiments described before or after this embodiment, as well as other related parts in the specification, can be referred to, and will not be repeated here.
[0317] This disclosure also proposes an apparatus (also referred to as a communication device, etc.) for implementing any of the above methods. For example, an apparatus is proposed that includes units or modules for implementing the steps performed by the terminal in any of the above methods. Furthermore, another apparatus is proposed that includes units or modules for implementing the steps performed by a network device (e.g., an access network device, a core network functional node, a core network device, etc.) in any of the above methods.
[0318] It should be understood that the division of units or modules in the above device is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the units or modules in the device can be implemented by a processor calling software: for example, the device includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of the units or modules in the above device. The processor can be, for example, a general-purpose processor, such as a Central Processing Unit (CPU) or a microprocessor, and the memory can be internal or external to the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits. The functionality of some or all of the units or modules can be achieved through the design of these hardware circuits, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC). The functionality of some or all of the units or modules is achieved through the design of the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a programmable logic device (PLD). Taking a field-programmable gate array (FPGA) as an example, it can include a large number of logic gates. The connection relationships between the logic gates are configured through configuration files, thereby achieving the functionality of some or all of the units or modules. All units or modules of the above device can be implemented entirely through processor-called software, entirely through hardware circuits, or partially through processor-called software with the remaining parts implemented through hardware circuits.
[0319] In this embodiment, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction read and execute capabilities, such as a Central Processing Unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. The logical relationships of the aforementioned hardware circuits are fixed or reconfigurable. For example, the processor is a hardware circuit implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a Neural Network Processing Unit (NPU), a Tensor Processing Unit (TPU), or a Deep Learning Processing Unit (DPU).
[0320] Figure 6A is a schematic diagram of the decoding device proposed in an embodiment of this disclosure. The decoding device is used to execute the data processing method executed by any of the decoders described above. In some embodiments, as shown in Figure 6A, the decoding device may include at least one of a transceiver module 6001 and a processing module 6002.
[0321] In some embodiments, the transceiver module 6001 is used to receive bit streams.
[0322] In some embodiments, the bitstream includes metadata associated with the first audio frame.
[0323] In some embodiments, the processing module is used to determine the audio features of the decoded first audio frame based on metadata. In some embodiments, the processing module can be interchanged with the determining module or the processor, and the transceiver module can be interchanged with the sending module or the transceiver.
[0324] Figure 6B is a schematic diagram of the structure of the encoding device proposed in an embodiment of this disclosure. The encoding device is used to perform the data processing method executed by the encoder described above. In some embodiments, as shown in Figure 6B, the encoding device may include a transceiver module 6101.
[0325] In some embodiments, the transceiver module 6101 described above is used to send a bit stream.
[0326] Optionally, the transceiver module 1101 is used to perform at least one of the transceiver steps (such as step 2103, step 3103, but not limited thereto) executed by the encoder in any of the above methods, which will not be described in detail here.
[0327] The encoding device 6100 described above may include a processing module, which is used to execute at least one of the communication steps (such as steps 2101, S2102, S3101, but not limited thereto) executed by the encoder in any of the above methods, which will not be described in detail here.
[0328] In some embodiments, the processing module can be interchanged with the determining module and the processor, and the transceiver module can be interchanged with the sending module and the transceiver.
[0329] Figure 7 is a schematic diagram of the structure of the communication device 7000 proposed in an embodiment of this disclosure. The communication device 7000 can be a network device (e.g., access network device, core network device, etc.), a terminal (e.g., user equipment, etc.), a chip, chip system, or processor that supports the network device in implementing any of the above methods, or a chip, chip system, or processor that supports the terminal in implementing any of the above methods. The communication device 7000 can be used to implement the methods described in the above method embodiments; for details, please refer to the descriptions in the above method embodiments.
[0330] As shown in Figure 7, the communication device 7000 is used to execute any of the above methods. In some embodiments, the communication device 7000 includes one or more processors 7001. The processor 7001 may be a general-purpose processor or a special-purpose processor, such as a baseband processor or a central processing unit. The baseband processor may be used to process communication protocols and communication data, and the central processing unit may be used to control communication devices (e.g., base stations, baseband chips, terminal devices, terminal device chips, DUs or CUs, etc.), execute programs, and process program data. Optionally, the communication device 7000 is used to execute any of the above methods. Optionally, one or more processors 7001 are used to invoke instructions to cause the communication device 7000 to execute any of the above methods.
[0331] In some embodiments, the communication device 7000 further includes one or more transceivers 7002. When the communication device 7000 includes one or more transceivers 7002, the transceiver 7002 performs at least one of the communication steps such as sending and / or receiving in the above method (e.g., steps S2103, S3103, S3201, S4101, but not limited thereto), and the processor 7001 performs at least one of other steps (e.g., steps S2101, S2102, S3202, S4102, but not limited thereto). In optional embodiments, the transceiver may include a receiver and / or a transmitter, which may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, transceiver circuit, interface circuit, interface, etc., can be used interchangeably; the terms transmitter, transmitting unit, transmitter, transmitting circuit, etc., can be used interchangeably; the terms receiver, receiving unit, receiver, receiving circuit, etc., can be used interchangeably.
[0332] In some embodiments, the communication device 7000 further includes one or more memories 7003 for storing data and / or instructions. Optionally, one or more processors 7001 are used to invoke instructions stored in the memory 7003 to cause the communication device 7000 to perform any of the above methods. Optionally, all or part of the memory 7003 may also be located outside the communication device 7000. In an optional embodiment, the communication device 7000 may include one or more interface circuits 7004. Optionally, the interface circuit 7004 is connected to the memory 7002, and the interface circuit 7004 can be used to receive data and / or instructions from the memory 7002 or other devices, and can be used to send data and / or instructions to the memory 7002 or other devices. For example, the interface circuit 7004 can read data and / or instructions stored in the memory 7002 and send the data and / or instructions to the processor 7001.
[0333] The communication device 7000 described in the above embodiments may be a network device or a terminal, but the scope of the communication device 7000 described in this disclosure is not limited thereto, and the structure of the communication device 7000 may not be limited by FIG. 7. The communication device may be a standalone device or a part of a larger device. For example, the communication device may be: (1) a standalone integrated circuit IC, or chip, or chip system or subsystem; (2) a collection of one or more ICs, optionally, the IC collection may also include storage components for storing data, programs and / or instructions; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, terminal device, smart terminal device, cellular phone, wireless device, handheld device, mobile unit, vehicle device, network device, cloud device, artificial intelligence device, etc.; (6) others, etc.
[0334] Figure 8 is a schematic diagram of the structure of the chip 8000 proposed in an embodiment of this disclosure. For cases where the communication device 7000 can be a chip or a chip system, the schematic diagram of the chip 8000 shown in Figure 8 can be referenced, but is not limited thereto.
[0335] Chip 8000 includes one or more processors 8001. Chip 8000 is used to perform any of the above methods.
[0336] In some embodiments, chip 8000 further includes one or more interface circuits 8002. Optionally, terms such as interface circuit, interface, and transceiver pin can be used interchangeably. In some embodiments, chip 8000 further includes one or more memories 8003 for storing data and / or instructions. Optionally, all or part of the memories 8003 may be located outside of chip 8000. Optionally, interface circuit 8002 is connected to memory 8003, and interface circuit 8002 can be used to receive data and / or instructions from memory 8003 or other devices, and interface circuit 8002 can be used to send data and / or instructions to memory 8003 or other devices. For example, interface circuit 8002 can read data and / or instructions stored in memory 8003 and send the data and / or instructions to processor 8001.
[0337] In some embodiments, the interface circuit 8002 performs at least one of the communication steps such as sending and / or receiving in the above-described method (e.g., steps S2103, S3103, S3201, and S4101, but not limited thereto). The interface circuit 8002 performing the communication steps such as sending and / or receiving in the above-described method refers, for example, to the interface circuit 8002 performing data and / or instruction interaction between the processor 8001, the chip 8000, the memory 8003, or the transceiver device. In some embodiments, the processor 8001 performs at least one of other steps (e.g., steps S2101, S2102, S3202, and S4102, but not limited thereto).
[0338] The modules and / or devices described in the various embodiments, such as virtual devices, physical devices, and chips, can be combined or separated arbitrarily as needed. Optionally, some or all steps can also be performed collaboratively by multiple modules and / or devices, which is not limited here.
[0339] This disclosure also proposes a storage medium storing instructions that, when executed on a communication device, cause the communication device to perform any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but not limited thereto; it may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but not limited thereto; it may also be a temporary storage medium.
[0340] This disclosure also proposes a program product, including a program and / or instructions, which, when executed by a communication device, cause the communication device to perform any of the above methods. Optionally, the program product is a computer program product. Optionally, the program product is stored on the storage medium.
[0341] This disclosure also proposes a computer program that, when run on a computer, causes the computer to perform any of the above methods.