Feature indication method and apparatus, and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING XIAOMI MOBILE SOFTWARE CO LTD
- Filing Date
- 2024-10-12
- Publication Date
- 2026-06-16
Smart Images

Figure CN122228634A_ABST
Abstract
Description
Feature indication method, apparatus and storage medium Technical Field
[0001] This disclosure relates to the field of encoding and decoding technology, and in particular to feature indication methods, apparatus and storage media. Background Technology
[0002] With the rapid development of encoding / decoding and communication technologies, data used for AI (Artificial Intelligence) tasks can be encoded and transmitted efficiently. Furthermore, features for AI models can be extracted from the data, encoded, and transmitted, reducing the amount of data during communication.
[0003] Summary of the Invention
[0004] This disclosure solves the problem that the decoder cannot know whether the features included in the data stream are compressed, ensuring that the decoder can determine whether the included features in the data stream it receives are compressed, and thus determine whether to further decompress the data stream after decoding, thereby ensuring the accuracy of feature acquisition based on the received data stream.
[0005] This disclosure provides a feature indication method, apparatus, and storage medium.
[0006] According to a first aspect of the present disclosure, a feature indication method is provided, the method being executed by a decoder, the method comprising:
[0007] Receive data stream;
[0008] Determine whether the features in the decoded data stream are compressed, wherein the features are used for different AI tasks.
[0009] According to a second aspect of the present disclosure, a feature indication method is provided, the method being performed by an encoder, the method comprising:
[0010] Send data stream;
[0011] The data stream is used to determine whether the decoded features are compressed, wherein different features are used for different AI tasks.
[0012] According to a third aspect of the embodiments of this disclosure, a feature indication method is provided, the method comprising:
[0013] The encoder sends a data stream;
[0014] The decoder receives the data stream;
[0015] The decoder determines whether features in the decoded data stream are compressed, wherein the features are used for different AI tasks.
[0016] According to a fourth aspect of the present disclosure, a feature indicating device is provided, comprising:
[0017] The transceiver module is used to receive data streams;
[0018] The processing module is used to determine whether the features in the decoded data stream are compressed, wherein the features are used for different AI tasks.
[0019] According to a fifth aspect of the present disclosure, a feature indicating device is provided, comprising:
[0020] The transceiver module is used to send a data stream; the data stream is used to determine whether the decoded features are compressed, wherein different features are used for different AI tasks.
[0021] According to a sixth aspect of the embodiments of this disclosure, a decoder is provided, comprising:
[0022] One or more processors;
[0023] The decoder is used to perform any of the methods described in the first aspect.
[0024] According to a seventh aspect of the embodiments of this disclosure, an encoder is provided, comprising:
[0025] One or more processors;
[0026] The encoder is used to perform any of the methods described in the second aspect.
[0027] According to an eighth aspect of the embodiments of this disclosure, a system is provided, comprising:
[0028] An encoder and a decoder, wherein the decoder is configured to implement the feature indication method of the first aspect, and the encoder is configured to implement the feature indication method of the second aspect.
[0029] According to a ninth aspect of the present disclosure, a storage medium is provided that stores instructions that, when executed on a communication device, cause the communication device to perform the method as described in any one of the first or second aspects. Attached Figure Description
[0030] The accompanying drawings, which are included to provide a further understanding of the embodiments of this disclosure and form part of this disclosure, illustrate exemplary embodiments of this disclosure and, together with their descriptions, serve to explain the embodiments of this disclosure and do not constitute an improper limitation of the embodiments of this disclosure. In the drawings:
[0031] Figure 1 is a schematic diagram of the architecture of an encoding / decoding system according to an embodiment of the present disclosure;
[0032] Figure 2A is an interactive schematic diagram of a feature indication method according to an embodiment of the present disclosure;
[0033] Figure 2B is a flowchart illustrating a feature indication method according to an embodiment of the present disclosure;
[0034] Figure 2C is a schematic flowchart illustrating a feature indication method according to an embodiment of the present disclosure;
[0035] Figure 2D is a flowchart illustrating a feature indication method according to an embodiment of the present disclosure;
[0036] Figure 2E is a flowchart illustrating a feature indication method according to an embodiment of the present disclosure;
[0037] Figure 2F is a flowchart illustrating a feature indication method according to an embodiment of the present disclosure;
[0038] Figure 2G is a schematic flowchart illustrating a feature indication method according to an embodiment of the present disclosure;
[0039] Figure 3A is a flowchart illustrating a feature indication method according to an embodiment of the present disclosure;
[0040] Figure 3B is a flowchart illustrating a feature indication method according to an embodiment of the present disclosure;
[0041] Figure 4A is a flowchart illustrating a feature indication method according to an embodiment of the present disclosure;
[0042] Figure 4B is a flowchart illustrating a feature indication method according to an embodiment of the present disclosure;
[0043] Figure 5 is a flowchart illustrating a feature indication method according to an embodiment of the present disclosure;
[0044] Figure 6 is a flowchart illustrating a feature indication method according to an embodiment of the present disclosure;
[0045] Figure 7A is a schematic diagram of the feature indication device proposed in an embodiment of this disclosure;
[0046] Figure 7B is a schematic diagram of the feature indication device proposed in an embodiment of this disclosure;
[0047] Figure 8A is a schematic diagram of the structure of the communication device proposed in an embodiment of this disclosure;
[0048] Figure 8B is a schematic diagram of the chip structure proposed in an embodiment of this disclosure. Detailed Implementation
[0049] This disclosure provides a feature indication method, apparatus, and storage medium.
[0050] According to a first aspect of the present disclosure, a feature indication method is provided, the method being executed by a decoder, the method comprising:
[0051] Receive data stream;
[0052] Determine whether the features in the decoded data stream are compressed, wherein the features are used for different AI tasks.
[0053] In the above embodiments, the problem of the decoder being unable to know whether the features included in the data stream are compressed is solved, ensuring that the data stream received by the decoder can determine whether the included features are compressed, and thus determine whether to further decompress the data stream after decoding, thereby ensuring the accuracy of obtaining features based on the received data stream.
[0054] In conjunction with some embodiments of the first aspect, in some embodiments, determining whether the features in the decoded data stream are compressed includes:
[0055] Whether the feature is compressed is determined based on the first information in the decoded data stream.
[0056] In the above embodiments, the decoder determines whether the features included in the data stream are compressed, and then decodes the data stream to obtain the features, ensuring the accuracy of the obtained features.
[0057] In conjunction with some embodiments of the first aspect, in some embodiments, determining whether the feature is compressed based on the first information in the decoded data stream includes:
[0058] If the first information indicates that the feature is compressed, it is determined that the feature included in the data stream is compressed;
[0059] If the first information indicates that the feature is not compressed, it is determined that the feature included in the data stream is not compressed.
[0060] In the above embodiments, the feature is determined based on the indication of the first information to ensure the accuracy of the indication through the first information.
[0061] In conjunction with some embodiments of the first aspect, in some embodiments, when the first information indicates that the feature is compressed, the method further includes:
[0062] The data stream is decoded to obtain the compressed features;
[0063] The compressed features are decompressed to obtain the decompressed features.
[0064] In the above embodiments, when the first information indicates that the feature is compressed, the compressed feature needs to be decompressed to ensure the accuracy of the obtained feature, thereby ensuring the accuracy of subsequent AI tasks based on the feature.
[0065] In conjunction with some embodiments of the first aspect, in some embodiments, the feature includes at least one of the following:
[0066] Machine characteristics;
[0067] Auditory characteristics.
[0068] In conjunction with some embodiments of the first aspect, in some embodiments, the machine features include at least one of the following:
[0069] ER (Emotion Recognition) features;
[0070] ASR (Automatic Speech Recognition) features;
[0071] ASV (Automatic Speaker Verification) feature;
[0072] AEC (Audio event classification) features.
[0073] In the above embodiments, the types of features included in the data stream are expanded to ensure feature diversity, thereby ensuring the reliability of subsequent feature-based processing.
[0074] In conjunction with some embodiments of the first aspect, in some embodiments, the data stream includes at least one flag bit, different flag bits correspond to different AI tasks, and the first information on the flag bit is used to indicate whether the feature is compressed.
[0075] In conjunction with some embodiments of the first aspect, in some embodiments,
[0076] The first information includes a first value, which indicates that the feature is compressed; or,
[0077] The first information includes a second value, which indicates that the feature has not been compressed.
[0078] In the above embodiments, a flag bit is used to indicate whether the corresponding feature is compressed, ensuring the accuracy of the indication of each feature, and thus ensuring the accuracy of the features obtained by subsequent decoding.
[0079] In conjunction with some embodiments of the first aspect, in some embodiments, the data used for monitoring includes at least one of the following:
[0080] Training data, which is used to train the model;
[0081] Validation data, which is used to optimize the model;
[0082] Test data, which is used to test the model.
[0083] In conjunction with some embodiments of the first aspect, in some embodiments, the data used for monitoring includes raw data and tag data, wherein the tag data is used to classify the raw data.
[0084] In the above embodiments, the types of data included in the data used for monitoring are expanded, thereby ensuring the comprehensiveness of the data.
[0085] Secondly, embodiments of this disclosure provide a feature indication method, the method being executed by an encoder, the method comprising:
[0086] Send data stream;
[0087] The data stream is used to determine whether the decoded features are compressed, wherein different features are used for different AI tasks.
[0088] In conjunction with some embodiments of the second aspect, in some embodiments, the first information in the data stream indicates whether the feature is compressed.
[0089] In conjunction with some embodiments of the second aspect, in some embodiments, when the first information indicates that the feature is compressed, it is determined that the feature included in the data stream is compressed;
[0090] If the first information indicates that the feature is not compressed, it is determined that the feature included in the data stream is not compressed.
[0091] In conjunction with some embodiments of the second aspect, in some embodiments, the data stream further includes data for monitoring, the features being extracted based on the data.
[0092] In conjunction with some embodiments of the second aspect, in some embodiments, the first feature includes at least one of the following:
[0093] Machine characteristics;
[0094] Auditory characteristics.
[0095] In conjunction with some embodiments of the second aspect, in some embodiments, the machine features include at least one of the following:
[0096] ER characteristics;
[0097] ASR characteristics;
[0098] ASV features;
[0099] AEC characteristics.
[0100] In conjunction with some embodiments of the second aspect, in some embodiments, the data stream includes at least one flag bit, different flag bits correspond to different AI tasks, and the first information on the flag bit is used to indicate whether the feature is compressed.
[0101] In conjunction with some embodiments of the second aspect, in some embodiments, the first information includes a first value, which is used to indicate that the feature is compressed; or,
[0102] The first information includes a second value, which indicates that the feature has not been compressed.
[0103] In conjunction with some embodiments of the second aspect, in some embodiments, the data used for monitoring includes at least one of the following:
[0104] Training data, which is used to train the model;
[0105] Validation data, which is used to update the parameters of the model;
[0106] Test data, which is used to test the accuracy of the model.
[0107] In conjunction with some embodiments of the second aspect, in some embodiments, the data used for monitoring includes raw data and tag data, wherein the tag data is used to classify the raw data.
[0108] Thirdly, embodiments of this disclosure provide a feature indication method, the method comprising:
[0109] The encoder sends a data stream;
[0110] The decoder receives the data stream;
[0111] The decoder determines whether there are compressed features in the decoded data stream, wherein different compressed features are used for different AI tasks.
[0112] Fourthly, embodiments of this disclosure provide a feature indication device, which includes at least one of a transceiver module and a processing module; wherein the feature indication device is used to execute an optional implementation of the first aspect.
[0113] Fifthly, embodiments of this disclosure provide a feature indication device, which includes at least one of a transceiver module and a processing module; wherein the feature indication device is used to execute an optional implementation of the second aspect.
[0114] Sixthly, embodiments of this disclosure provide a decoder, including:
[0115] One or more processors;
[0116] The decoder is used to perform the method described in any one of the first aspects.
[0117] In a seventh aspect, embodiments of this disclosure provide an encoder, including:
[0118] One or more processors;
[0119] The encoder is used to perform the method described in any one of the second aspects.
[0120] Eighthly, embodiments of this disclosure provide a storage medium storing information that, when the information is executed on a communication device, causes the communication device to perform the method as described in any one of the first and second aspects.
[0121] Ninthly, embodiments of this disclosure provide a program product that, when executed by a communication device, causes the communication device to perform the method as described in either the first or second aspect.
[0122] In a tenth aspect, embodiments of this disclosure provide a computer program that, when run on a communication device, causes the communication device to perform the method described in either the first or second aspect.
[0123] Eleventhly, embodiments of this disclosure provide a chip or chip system. The chip or chip system includes processing circuitry configured to perform the methods described in either the first or second aspect.
[0124] It is understood that the encoders, decoders, storage media, program products, computer programs, chips, or chip systems described above are all used to perform the methods proposed in the embodiments of this disclosure. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here.
[0125] This disclosure provides a feature indication method, apparatus, and storage medium. In some embodiments, the terms "feature indication method" and "information feature indication method" can be used interchangeably, as can the terms "feature indication apparatus" and "information feature indication apparatus" and "indication apparatus," and the terms "information processing system" and "encoding / decoding system" can be used interchangeably.
[0126] This disclosure is not exhaustive, but merely illustrative of some embodiments, and is not intended to limit the scope of protection of this disclosure. Unless otherwise specified, each step in a particular embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a particular embodiment can also be implemented as an independent embodiment, and the order of the steps in a particular embodiment can be arbitrarily interchanged. Furthermore, the optional implementation methods in a particular embodiment can be arbitrarily combined; moreover, the embodiments can be arbitrarily combined, for example, some or all steps of different embodiments can be arbitrarily combined, and a particular embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.
[0127] In each of the disclosed embodiments, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of the embodiments are consistent and can be referenced by each other. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.
[0128] The terminology used in the embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure.
[0129] In this embodiment of the disclosure, unless otherwise stated, elements expressed in the singular form, such as "a," "an," "the," "the," "the," "the," "the," "the," "this," etc., can mean "one and only one," or "one or more," "at least one," etc. For example, when using articles such as "a," "an," "the," etc. in translation, the noun following the article can be understood as either a singular expression or a plural expression.
[0130] In the embodiments of this disclosure, "multiple" refers to two or more.
[0131] In some embodiments, the terms “at least one of”, “one or more”, “a plurality of”, “multiple”, etc., may be used interchangeably.
[0132] In some embodiments, the notation "at least one of A and B", "A and / or B", "A in one case, B in another", "in response to one case A, in response to another case B", etc., may include the following technical solutions depending on the situation: in some embodiments, A (execute A regardless of B); in some embodiments, B (execute B regardless of A); in some embodiments, execution is selected from A and B (A and B are selectively executed); in some embodiments, A and B (both A and B are executed). The same applies when there are more branches such as A, B, C, etc.
[0133] In some embodiments, the notation "A or B" may include the following technical solutions, depending on the situation: in some embodiments, A (execution of A regardless of B); in some embodiments, B (execution of B regardless of A); in some embodiments, execution is selected from A and B (A and B are selectively executed). The same applies when there are more branches such as A, B, C, etc.
[0134] The prefixes "first," "second," etc., used in the embodiments of this disclosure are merely for distinguishing different descriptive objects and do not impose restrictions on the position, order, priority, quantity, or content of the descriptive objects. The description of the descriptive objects is found in the claims or the context of the embodiments, and the use of prefixes should not constitute unnecessary restrictions. For example, if the descriptive object is a "field," the ordinal numbers preceding "field" in "first field" and "second field" do not restrict the position or order of the "fields." "First" and "second" do not restrict whether the "fields" they modify are in the same message, nor do they restrict the order of "first field" and "second field." Similarly, if the descriptive object is a "level," the ordinal numbers preceding "level" in "first level" and "second level" do not restrict the priority between "levels." Furthermore, the number of descriptive objects is not limited by ordinal numbers and can be one or more. For example, in "first device," the number of "devices" can be one or more. Furthermore, the objects modified by different prefixes can be the same or different. For example, if the object being described is "device", then "first device" and "second device" can be the same device or different devices, and their types can be the same or different. Similarly, if the object being described is "information", then "first information" and "second information" can be the same information or different information, and their content can be the same or different.
[0135] In some embodiments, “including A,” “containing A,” “for indicating A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.
[0136] In some embodiments, terms such as "time / frequency" and "time-frequency domain" refer to the time domain and / or frequency domain.
[0137] In some embodiments, the terms “in response to…”, “in response to determining…”, “in the case of…”, “when…”, “if…”, “if…”, etc., can be used interchangeably.
[0138] In some embodiments, the terms “greater than,” “greater than or equal to,” “not less than,” “more than,” “more than or equal to,” “not less than,” “higher than,” “higher than or equal to,” “not lower than,” and “above” can be used interchangeably, as can the terms “less than,” “less than or equal to,” “not greater than,” “less than,” “less than or equal to,” “not more than,” “lower than,” “lower than or equal to,” “not higher than,” and “below”.
[0139] In some embodiments, the apparatus and device may be interpreted as physical or virtual, and their names are not limited to the names recorded in the embodiments. In some cases, they may also be understood as "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "body", etc.
[0140] In some embodiments, "network" can be interpreted as devices included in the network, such as access network devices, core network devices, etc.
[0141] In some embodiments, "access network device (AN device)" may also be referred to as "radio access network device (RAN device)," "base station (BS)," "radio base station," or "fixed station." In some embodiments, it may also be understood as "node," "access point," "transmission point (TP)," "reception point (RP)," "transmission / reception point (TRP)," "panel," "antenna panel," "antenna array," "cell," "macro cell," "small cell," "femto cell," "pico cell," "sector," "cell group," "serving cell," "carrier," "component carrier," or "bandwidth part (BWP)."
[0142] In some embodiments, "encoder" or "encoder device" may be referred to as "user equipment (encoder)," "user encoder," "mobile station (MS)," "mobile encoder (MT)," "subscriber station," "mobile unit," "subscriber unit," "wireless unit," "remote unit," "mobile device," "wireless device," "wireless communication device," "remote device," "mobile subscriber station," "access terminal," "mobile encoder," "wireless terminal," "remote terminal," "handset," "user agent," "mobile client," "client," etc.
[0143] In some embodiments, the acquisition of data, information, etc., may comply with the laws and regulations of the country where the location is situated.
[0144] In some embodiments, data, information, etc., may be obtained with the user's consent.
[0145] Furthermore, each element, each row, or each column in the table of this disclosure can be implemented as an independent embodiment, and any combination of any element, any row, or any column can also be implemented as an independent embodiment.
[0146] Figure 1 is a schematic diagram of the architecture of an encoding / decoding system according to an embodiment of the present disclosure. As shown in Figure 1, the method provided in this embodiment can be applied to an encoding / decoding system 100, which may include an encoder 101 and a decoder 102. It should be noted that the encoding / decoding system 100 may also include other devices, and this disclosure does not limit the devices included in the encoding / decoding system 100.
[0147] In some embodiments, both encoder 101 and decoder 102 are located in a terminal. In some embodiments, the terminal can be various devices. For example, it includes at least one of, but is not limited to, mobile phones, wearable devices, Internet of Things devices, automobiles with communication capabilities, smart cars, tablets, computers with wireless transceiver capabilities, virtual reality (VR) encoder devices, augmented reality (AR) encoder devices, wireless encoder devices in industrial control, wireless encoder devices in self-driving, wireless encoder devices in remote medical surgery, wireless encoder devices in smart grids, wireless encoder devices in transportation safety, wireless encoder devices in smart cities, and wireless encoder devices in smart homes.
[0148] It is understood that the encoding and decoding system described in the embodiments of this disclosure is for the purpose of more clearly illustrating the technical solutions of the embodiments of this disclosure, and does not constitute a limitation on the technical solutions proposed in the embodiments of this disclosure. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions proposed in the embodiments of this disclosure are also applicable to similar technical problems.
[0149] The following embodiments of this disclosure can be applied to the encoding / decoding system 100 shown in FIG1, or to some of the main components, but are not limited thereto. The main components shown in FIG1 are illustrative. The encoding / decoding system may include all or some of the main components in FIG1, or may include other main components outside of FIG1. The number and form of each main component are arbitrary. Each main component may be physical or virtual. The connection relationship between the main components is illustrative. The main components may not be connected or may be connected. The connection can be in any way, it can be a direct connection or an indirect connection, it can be a wired connection or a wireless connection.
[0150] The embodiments disclosed herein can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G new radio (NR), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New radio access (NX), Future generation radio access (FX), Global System for Mobile communications (GSM), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi), IEEE 802.16 (WiMAX), and IEEE 802.20, Ultra-Wideband (UWB), Bluetooth (Bluetooth, a registered trademark), Public Land Mobile Network (PLMN) networks, Device-to-Device (D2D) systems, Machine-to-Machine (M2M) systems, Internet of Things (IoT) systems, Vehicle-to-Everything (V2X) systems, systems utilizing other characteristic indication methods, and next-generation systems extended from them, etc. Furthermore, multiple systems can be combined (e.g., a combination of LTE or LTE-A with 5G).
[0151] Figure 2A is an interactive schematic diagram of a feature indication method according to an embodiment of the present disclosure. As shown in Figure 2A, the present disclosure relates to a feature indication method, which includes:
[0152] Step S2101: The encoder acquires data for monitoring.
[0153] In some embodiments, the data used for monitoring is used for user monitoring. Optionally, user monitoring includes judging whether the data meets the requirements. In some embodiments, the data used for monitoring is also used for model processing. Optionally, model processing includes model training, model parameter updating, model usage, etc., which are not limited in this disclosure.
[0154] It should be noted that the embodiments disclosed herein do not limit the functionality of the model. The function of the model can be understood as classifying items, predicting data, and identifying data, etc., and the embodiments disclosed herein do not limit this.
[0155] In some embodiments, the data used for monitoring includes raw data and label data, with the label data used to classify the raw data. Alternatively, the raw data can be understood as being used for model processing, and the label data as being used to adjust the model's parameters.
[0156] In some embodiments, the data used for monitoring includes at least one of the following:
[0157] (1) Training data, which is used to train the model.
[0158] In some embodiments, model training refers to adjusting the parameters of the model to ensure that the model has the ability to recognize results that match the training data.
[0159] In some embodiments, the training data includes first raw data and first label data. The model is used to process the first raw data to obtain first predicted data, and then adjust the parameters of the model based on the difference between the first label data and the first predicted data, so that the model has the function of processing the first raw data to obtain the first label data.
[0160] (2) Validation data, which is used to optimize the model.
[0161] In this embodiment of the disclosure, the verification data is used to further update the parameters of the model so that the model's functionality is more accurate.
[0162] In some embodiments, the verification data includes second raw data and second label data. The model is used to process the second raw data to obtain second predicted data, and then optimize the parameters of the model based on the difference between the second label data and the second predicted data, so that the model has the function of processing the second raw data to obtain the second label data.
[0163] (3) Test data, which is used to test the model.
[0164] In this embodiment of the disclosure, the test data is used to verify the functionality of the model in order to determine whether the model's functionality meets the requirements.
[0165] In some embodiments, the test data includes third raw data and third labeled data. The model is used to process the third raw data to obtain third predicted data, and then compare the third labeled data with the third predicted data to determine whether the model is accurate based on the difference between the third labeled data and the third predicted data.
[0166] In some embodiments, the data used for monitoring is audio data. In this disclosure, the model can be processed based on the audio data. For example, speech recognition or classification of the audio data can be performed; however, this disclosure does not limit the scope of the application.
[0167] In step S2102, the encoder extracts features from the data used for monitoring.
[0168] In some embodiments, different features are used for different AI tasks. Optionally, the AI task includes at least one of emotion recognition, automatic speech recognition, automatic speech verification, or audio event classification. In some embodiments, the first feature is used for subsequent model processing. Alternatively, it can be understood that the first feature includes multiple types, with different kinds of first features used for different model processing.
[0169] In some embodiments, the features include at least one of the following:
[0170] (1) Machine characteristics.
[0171] In some embodiments, the machine feature is an audio feature used for the AI model, ensuring consistency between the machine feature acquired by the encoder and the audio feature required by the AI model for the backend task. This machine feature includes various types, with different types used for processing different models.
[0172] Optionally, the machine features include at least one of the following:
[0173] 1. ER (Emotion Recognition) features.
[0174] Optionally, the ER feature is used for an emotion recognition task. Optionally, the ER feature is used to train an emotion recognition model. Alternatively, it can be understood that the emotion recognition model is used for an emotion recognition task. Optionally, the emotion recognition task is used to identify emotions in audio data. Optionally, the emotion includes anger, happiness, joy, etc., and this disclosure does not limit the emotion.
[0175] 2. ASR (Automatic Speech Recognition) features.
[0176] Optionally, the ASR feature is used for an automatic speech recognition task. Optionally, the ASR feature is used to train an automatic speech recognition model. Alternatively, it can be understood that the automatic speech recognition model is used for an automatic speech recognition task. Optionally, the automatic speech recognition task is used to automatically recognize audio data. Optionally, automatically recognizing audio data includes recognizing text in the audio data.
[0177] 3. ASV (Automatic Voice Verification) feature.
[0178] Optionally, the ASV feature is used for an automatic speech verification task. Alternatively, the ASV feature is used to train an automatic speech verification model. Or, it can be understood that the automatic speech verification model is used for an automatic speech verification task. Optionally, the automatic speech verification task is used to verify audio data.
[0179] 4. AEC (Audio Event Classification) features.
[0180] Optionally, the AEC feature is used for an audio event classification task. Alternatively, the AEC feature is used to train an audio event classification model. Or, it can be understood that the audio event classification model is used for an audio event classification task.
[0181] (2) Auditory characteristics.
[0182] In some embodiments, the auditory feature is a feature acquired based on the auditory characteristics of the human ear. The auditory feature is decoded by a decoder, and the decoded receiver is typically the human ear.
[0183] It should be noted that for each type of feature, the encoder may extract some types of features, while other types of features may not be extracted. In other words, the features include some types of features from the above-mentioned multiple types, but do not include some types of features.
[0184] In step S2103, the encoder compresses the features.
[0185] In this embodiment of the disclosure, after feature extraction is performed on the data used for monitoring to obtain features, the amount of feature data may increase, and the features also need to be sent to the decoder. In order to reduce the amount of data sent between the encoder and the decoder, the features can also be compressed.
[0186] In some embodiments, if the encoder compresses the extracted features, it generates first information indicating that the features have been compressed.
[0187] In step S2104, the encoder encodes the data to obtain a data stream.
[0188] In some embodiments, the encoder encodes at least one of the data to be monitored, the compressed features, and the first information to obtain a data stream.
[0189] In some embodiments, if the feature includes multiple types, it can also indicate whether each type of feature exists. Optionally, identification information included in the first information in the data stream can indicate whether the corresponding feature is compressed.
[0190] In some embodiments, each piece of first information corresponds to a flag bit in the data stream, and different flag bits correspond to different AI tasks. In some embodiments, the data stream includes at least one flag bit, with different flag bits corresponding to different AI tasks, and the first information on the flag bit is used to indicate whether the feature is compressed. Optionally, the first information includes a first value, which indicates that the feature is compressed; or, the first information includes a second value, which indicates that the feature is not compressed. For example, if the first value is 1 and the second value is 0, a first information value of 1 indicates that the feature is compressed, and a first information value of 0 indicates that the feature is not compressed.
[0191] Optionally, if the features obtained by feature extraction of the data used for monitoring include ER (Emotion Recognition) features, ASV (Automatic Voice Verification) features, and AEC (Audio Event Classification) features, the flag bits may include flag bit A, flag bit B, and flag bit C. Flag bit A corresponds to the ER (Emotion Recognition) feature, flag bit B corresponds to the ASV (Automatic Voice Verification) feature, and flag bit C corresponds to the AEC (Audio Event Classification) feature. If the first information of flag bit A and flag bit B is 1, and the first information of flag bit C is 0, it indicates that the ER (Emotion Recognition) feature and the ASV (Automatic Voice Verification) feature are compressed, while the AEC (Audio Event Classification) feature is not compressed.
[0192] It should be noted that this embodiment is described using the example of whether the first information indicates whether the feature is compressed. In another embodiment, the first information can also be used to indicate whether the data used for monitoring is compressed. Optionally, the first information corresponding to the flag bit includes a first value, which indicates that the data used for monitoring is compressed; or, the first information corresponding to the flag bit includes a second value, which indicates that the data used for monitoring is not compressed. For example, if the first value is 1 and the second value is 0, when the first information corresponding to the flag bit is 1, it indicates that the data used for monitoring is compressed; when the first information corresponding to the flag bit is 0, it indicates that the data used for monitoring is not compressed.
[0193] For example, flags include indicators to show whether the features of the ER task are compressed, such as ER_compressed_feature_exist_flag (ER compressed feature existence flag). A flag value of 1 indicates that the compressed features of the ER task exist in the bitstream, and a flag value of 0 indicates that the compressed features of the ER task do not exist in the bitstream.
[0194] For example, flags include indicators to show whether the features of the ASR task are compressed, such as ASR_compressed_feature_exist_flag (ASR compressed feature existence flag). ASR_compressed_feature_exist_flag being 1 indicates that the features of the ASR task are compressed, and ASR_compressed_feature_exist_flag being 0 indicates that the features of the ASR task are not compressed.
[0195] For example, flags include indicators to show whether the features of an ASV task are compressed, such as ASV_compressed_feature_exist_flag (ASV compressed feature existence flag). A value of 1 for ASV_compressed_feature_exist_flag indicates that the ASV task is compressed, while a value of 0 indicates that the features of the ASV task are not compressed.
[0196] For example, flags are used to indicate whether the features of an AEC task are compressed, such as AEC_compressed_feature_exist_flag (AEC compressed feature existence flag). A value of 1 for AEC_compressed_feature_exist_flag indicates that the features of the AEC task are compressed, while a value of 0 indicates that the features of the AEC task are not compressed.
[0197] It should be noted that the features obtained by human hearing encoding in this embodiment can be understood as data used for listening, or it can be considered that the data used for listening has not changed after passing through the human hearing module. Therefore, the human_listening_compressed_data_exist_flag (human hearing compressed data existence flag) can be used to indicate whether the data used for listening has been compressed.
[0198] Step S2105: The encoder sends a data stream.
[0199] In this embodiment of the disclosure, after the encoder obtains the data stream, it can send the data stream, and the encoder can then receive the data stream.
[0200] In some embodiments, the first information may also be referred to as edge information or descriptive information, and this disclosure does not limit this.
[0201] Step S2106: The decoder receives the data stream.
[0202] Step S2107: The decoder decodes the data stream.
[0203] In step S2108, the decoder determines whether the features in the decoded data stream have been compressed.
[0204] In some embodiments, the decoder decodes the data stream to obtain at least one of the following: first information, features, or data for listening, including the flag bits.
[0205] In some embodiments, the decoder determines whether a feature is compressed based on first information in the decoded data stream.
[0206] In some embodiments, where the first information indicating feature is compressed, the features included in the data stream are compressed.
[0207] Optionally, if the first information indicating feature is compressed, the decoder decodes the data stream to obtain the compressed feature, and decompresses the compressed feature to obtain the decompressed feature.
[0208] In some embodiments, where the first information indicating feature is not compressed, the features included in the data stream are not compressed.
[0209] It should be noted that the embodiments disclosed herein involve data and features used for monitoring before the encoder encodes them, as well as data and features used for monitoring obtained by the decoder. Encoding the data and features used for monitoring by the encoder may cause data loss, and the data and features used for monitoring obtained by the decoder may also cause data loss. Therefore, the data used for monitoring before the encoder encodes them and the data used for monitoring obtained by the decoder may not be completely the same, and the features before the encoder encodes them and the features obtained by the decoder may not be completely the same.
[0210] Alternatively, in this embodiment of the invention, both the encoder and the decoder are lossless processes during encoding and decoding. Therefore, the data used for monitoring before encoding by the encoder is the same as the data used for monitoring after decoding by the decoder, and the features before encoding by the encoder are the same as the features after decoding by the decoder.
[0211] It should be noted that the features obtained in the embodiments of this disclosure are used for AI tasks. In some embodiments, the AI task can also be understood as AI processing. Optionally, the AI task includes ER (emotion recognition) tasks, ASV (automatic voice verification) tasks, and AEC (audio event classification) tasks, etc., and the embodiments of this disclosure are not limited thereto.
[0212] In some embodiments, after the decoder obtains the features, it sends the obtained features to an AI device, which then performs AI processing based on the received features.
[0213] It should be noted that the embodiments disclosed herein are illustrated using an encoder and a decoder as examples. In another embodiment, modules may be used for illustration.
[0214] Referring to Figure 2B, the data used for monitoring includes at least one of the training set, validation set, and test set. In this embodiment of the disclosure, it includes a transmission module, a local process module, and a back-end AI task module. Optionally, the transmission module is located in the encoder, the local process module and the back-end module are located in the decoder, or the back-end module is located in a server or other device. This embodiment of the disclosure does not limit the scope of the disclosure.
[0215] Optionally, the transmission module is used to compress the raw data to meet transmission or storage needs, and the output of the transmission module is a bitstream. Optionally, the first application scenario of this transmission module includes scenarios where the network or bus with limited transmission bandwidth cannot meet the real-time transmission of data used for monitoring. Optionally, the second application scenario of this transmission module includes scenarios where storage devices with limited storage space cannot meet the storage needs of the raw data.
[0216] Optionally, the input to this local process module is a bitstream. Optionally, the local process module decodes and converts the bitstream. The format conversion is used to adapt the data format of the backend module of the backend task.
[0217] Optionally, the backend module is the user of the data being monitored. The backend module includes AI models corresponding to multiple AI tasks. The AI model set is trained, validated, and tested using the data.
[0218] In some embodiments, as shown in Figure 2C, the transmission module includes an encoding module and the transmission of encoded bitstream. The encoding module includes Feature Extraction and Feature Encoding. Feature Extraction extracts features of a specific task from the input data. Feature Encoding compresses the task's features, i.e., the output of Feature Extraction, to reduce the amount of task feature data.
[0219] In some embodiments, as shown in Figure 2D, the backend module includes a local process stage and a backup AI tasks set. Optionally, the backup AI tasks set includes ASR, ASV, ER, AEC, etc. The local process stage includes a decoder module and a format conversion module. The decoder is used to decode the bitstream, and the decoded data is used for training, validation, and testing of multiple backend tasks, such as ER, ASR, ASV, and AEC. Optionally, the output of the format conversion includes audio data and labels.
[0220] In some embodiments, as shown in Figure 2E, the bitstream includes four features: ASR, ASV, ER, and AEC. The decoder decodes the ER compression feature to obtain the ER decompression feature, decodes the ASR compression feature to obtain the ASR decompression feature, decodes the ASV compression feature to obtain the ASV decompression feature, and decodes the AEC compression feature to obtain the AEC decompression feature. Then, the ER decompression feature, ASR decompression feature, ASV decompression feature, and AEC decompression feature are format converted to obtain ER, ASR, ASV, and AEC.
[0221] Optionally, embodiments of this disclosure may also add identification information to indicate whether the corresponding features are compressed. For example, as shown in Figure 2F, ER_compressed_feature_exist_flag, ASR_compressed_feature_exist_flag, ASV_compressed_feature_exist_flag, and AEC_compressed_feature_exist_flag are all 1, and human_listening_compressed_data_exist_flag is 0. This indicates that the features of the ER task, the ASR task, the ASV task, the AEC task, and the data used for listening are all compressed. For example, as shown in Figure 2G, if ER_compressed_feature_exist_flag is 1, and ASR_compressed_feature_exist_flag, ASV_compressed_feature_exist_flag, AEC_compressed_feature_exist_flag, and human_listening_compressed_data_exist_flag are all 0, it means that the features of the ER task are compressed, while the features of the ASR task, ASV task, AEC task, and the data used for listening are not compressed.
[0222] In some embodiments, the names of information, etc., are not limited to the names described in the embodiments. Terms such as "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "symbol", "codebook", "codeword", "codepoint", "bit", "data", "program", and "chip" can be used interchangeably.
[0223] In some embodiments, the terms "uplink", "uplink", and "physical uplink" can be used interchangeably, as can the terms "downlink", "downlink", and "physical downlink", as well as the terms "sidelink", "sidelink", "sidelink communication", "sidelink communication", "direct connection", "direct link", "direct communication", and "direct link communication".
[0224] In some embodiments, “get,” “obtain,” “get,” “receive,” “transmit,” “bidirectional transmission,” and “send and / or receive” can be used interchangeably and can be interpreted as receiving from other entities, obtaining from protocols, obtaining from higher layers, obtaining through self-processing, or autonomous implementation, among other meanings.
[0225] In some embodiments, terms such as “send,” “transmit,” “report,” “distribute,” “transfer,” “bidirectional transmission,” “send and / or receive” can be used interchangeably.
[0226] In some embodiments, terms such as “moment,” “point in time,” “time,” and “time location” can be used interchangeably, as can terms such as “duration,” “segment,” “time window,” “window,” and “time.”
[0227] In some embodiments, terms such as "certain," "preset," "default," "set," "indicated," "a certain," "any," and "first" can be used interchangeably. "Certain A," "preset A," "default A," "set A," "indicated A," "a certain A," "any A," and "first A" can be interpreted as A pre-defined in a protocol or the like, or as A obtained through setting, configuration, or instruction, or as specific A, a certain A, any A, or first A, but are not limited thereto.
[0228] The feature indication method involved in the embodiments of this disclosure may include at least one of steps S2101 to S2107. For example, step S2101 can be implemented as an independent embodiment, step S2102 can be implemented as an independent embodiment, step S2103 can be implemented as an independent embodiment, step S2104 can be implemented as an independent embodiment, step S2105 can be implemented as an independent embodiment, step S2106 can be implemented as an independent embodiment, step S2107 can be implemented as an independent embodiment, steps S2101 and S2102 can be implemented as independent embodiments, and steps S2103 and S2107 can be implemented as independent embodiments. Steps 2104, 2105, and 2106 can be implemented as independent embodiments, as can steps S2101, S2102, S2103, and S2104, 2105, and S2106, and are implemented as independent embodiments, but are not limited thereto.
[0229] In some embodiments, step S2101 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0230] In some embodiments, step S2102 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0231] In some embodiments, step S2103 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0232] In some embodiments, step S2104 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0233] In some embodiments, step S2105 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0234] In some embodiments, step S2106 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0235] In some embodiments, step S2107 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0236] In some embodiments, steps S2101 and S2102 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0237] In some embodiments, steps S2103 and S2104 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0238] In some embodiments, steps S2105 and S2106 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0239] In some embodiments, other alternative implementations may be described before or after the specification corresponding to FIG2A.
[0240] Figure 3A is a flowchart illustrating a feature indication method according to an embodiment of the present disclosure, applied to an encoder. As shown in Figure 3A, the present disclosure relates to a feature indication method, which includes:
[0241] Step S3101: The encoder acquires data for monitoring.
[0242] The optional implementation of step S3101 can be found in step S2101 of Figure 2A and other related parts in the embodiments involved in Figure 2A, which will not be repeated here.
[0243] In step S3102, the encoder extracts features from the data used for monitoring.
[0244] The optional implementation of step S3102 can be found in step S2102 of Figure 2A and other related parts in the embodiments involved in Figure 2A, which will not be repeated here.
[0245] Step S3103: The encoder compresses the features.
[0246] The optional implementation of step S3103 can be found in step S2103 of Figure 2A and other related parts in the embodiments involved in Figure 2A, which will not be repeated here.
[0247] In step S3104, the encoder encodes the data to obtain a data stream.
[0248] The optional implementation of step S3104 can be found in step S2104 of Figure 2A and other related parts in the embodiments involved in Figure 2A, which will not be repeated here.
[0249] Step S3105: The encoder sends a data stream.
[0250] The optional implementation of step S3105 can be found in step S2105 of Figure 2A and other related parts in the embodiment involved in Figure 2A, which will not be repeated here.
[0251] The feature indication method involved in the embodiments of this disclosure may include at least one of steps S3101 to S3105. For example, step S3101 may be implemented as a standalone embodiment, step S3102 may be implemented as a standalone embodiment, step S3103 may be implemented as a standalone embodiment, step S3104 may be implemented as a standalone embodiment, step S3105 may be implemented as a standalone embodiment, or at least two steps may be combined, but it is not limited thereto.
[0252] In some embodiments, steps S3101, S3102, S3103, S3104, and S3105 are optional, and one or more of these steps may be omitted or substituted in different embodiments. However, this is not a limitation.
[0253] Figure 3B is a flowchart illustrating a feature indication method according to an embodiment of the present disclosure, applied to an encoder. As shown in Figure 3B, the present disclosure relates to a feature indication method, which includes:
[0254] Step S3201: The encoder sends a data stream.
[0255] The optional implementation of step S3201 can be found in step S2105 of Figure 2A, step S3105 of Figure 3A, and other related parts in the embodiments involved in Figures 2A and 3A, which will not be repeated here.
[0256] Figure 4A is a flowchart illustrating a feature indication method according to an embodiment of the present disclosure, applied to a decoder. As shown in Figure 4A, this embodiment of the present disclosure relates to a feature indication method, which includes:
[0257] Step S4101: The decoder receives the data stream.
[0258] The optional implementation of step S4101 can be found in step S2106 of Figure 2A and other related parts in the embodiment involved in Figure 2A, which will not be repeated here.
[0259] In step S4102, the decoder decodes the data stream.
[0260] The optional implementation of step S4102 can be found in step S2107 of Figure 2A and other related parts in the embodiments involved in Figure 2A, which will not be repeated here.
[0261] In step S4103, the decoder determines whether the features in the decoded data stream have been compressed.
[0262] The optional implementation of step S4103 can be found in step S2108 of Figure 2A and other related parts in the embodiment involved in Figure 2A, which will not be repeated here.
[0263] Figure 4B is a flowchart illustrating a feature indication method according to an embodiment of the present disclosure, applied to a decoder. As shown in Figure 4B, the present disclosure relates to a feature indication method, which includes:
[0264] Step S4201: The decoder receives the data stream.
[0265] The optional implementation of step S4201 can be found in step S2106 of Figure 2A and other related parts in the embodiment involved in Figure 2A, which will not be repeated here.
[0266] In step S4202, the decoder determines whether the features in the decoded data stream have been compressed.
[0267] Optional implementations of step S4202 can be found in step S2107 of Figure 2A and other related parts in the embodiments involved in Figure 2A, which will not be repeated here.
[0268] Figure 5 is a flowchart illustrating a feature indication method according to an embodiment of the present disclosure. As shown in Figure 5, the present disclosure relates to a feature indication method, which includes:
[0269] Step S5101: The encoder sends a data stream.
[0270] The optional implementation of step S5101 can be found in step S2105 of Figure 2A, step S3105 of Figure 3A, and other related parts in the embodiments involved in Figures 2A and 3A, which will not be repeated here.
[0271] Step S5102: The decoder receives the data stream.
[0272] The optional implementation of step S5102 can be found in step S2106 of Figure 2A, step S3101 of Figure 4A, and other related parts in the embodiments involved in Figures 2A and 4A, which will not be repeated here.
[0273] In step S5103, the decoder determines whether the features in the decoded data stream have been compressed.
[0274] The optional implementation of step S5103 can be found in step S2107 of Figure 2A, step S3102 of Figure 4A, and other related parts in the embodiments involved in Figures 2A and 4A, which will not be repeated here.
[0275] In some embodiments, the above methods may include the methods of the embodiments of the communication system side, encoder side, decoder side, etc., which will not be described again here.
[0276] Figure 6 is a flowchart illustrating a feature indication method according to an embodiment of the present disclosure. As shown in Figure 6, the present disclosure relates to a feature indication method, which includes:
[0277] In step S6101, the flag bit in the bitstream is used to indicate whether the feature is compressed.
[0278] In some embodiments, the flag bit includes multiple types. Optionally, it includes ER_compressed_feature_exist_flag, ASR_compressed_feature_exist_flag, ASV_compressed_feature_exist_flag, and AEC_compressed_feature_exist_flag.
[0279] Here, ER_compressed_feature_exist_flag is a 1-bit flag indicating whether the features of the ER task are compressed. A value of 1 indicates that the features of the ER task are compressed, and a value of 0 indicates that the features of the ER task are not compressed.
[0280] ASR_compressed_feature_exist_flag: A 1-bit flag indicating whether the features of the ASR task are compressed. A value of 1 indicates that the features of the ASR task are compressed, and a value of 0 indicates that the features of the ASR task are not compressed.
[0281] ASV_compressed_feature_exist_flag: A 1-bit flag indicating whether the features of an ASV task are compressed. A value of 1 indicates that the features of the ASV task are compressed, and a value of 0 indicates that the features of the ASV task are not compressed.
[0282] AEC_compressed_feature_exist_flag: A 1-bit flag indicating whether the features of the AEC task are compressed. A value of 1 indicates that the features of the AEC task are compressed, and a value of 0 indicates that the features of the AEC task are not compressed.
[0283] In the embodiments disclosed herein, some or all of the steps and their optional implementations may be arbitrarily combined with some or all of the steps in other embodiments, or may be arbitrarily combined with the optional implementations in other embodiments.
[0284] This disclosure also provides embodiments of an apparatus for implementing any of the above methods. For example, an apparatus is provided that includes units or modules for implementing the steps performed by the encoder in any of the above methods. Furthermore, another apparatus is provided that includes units or modules for implementing the steps performed by the decoder in any of the above methods.
[0285] It should be understood that the division of units or modules in the above device is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the units or modules in the device can be implemented by a processor calling software: for example, the device includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of the units or modules in the above device. The processor can be, for example, a general-purpose processor, such as a Central Processing Unit (CPU) or a microprocessor, and the memory can be internal or external to the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits. The functionality of some or all of the units or modules can be achieved through the design of these hardware circuits, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC). The functionality of some or all of the units or modules is achieved through the design of the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a programmable logic device (PLD). Taking a field-programmable gate array (FPGA) as an example, it can include a large number of logic gates. The connection relationships between the logic gates are configured through configuration files, thereby achieving the functionality of some or all of the units or modules. All units or modules of the above device can be implemented entirely through processor-called software, entirely through hardware circuits, or partially through processor-called software with the remaining parts implemented through hardware circuits.
[0286] In this embodiment, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction read and execute capabilities, such as a Central Processing Unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. The logical relationships of the aforementioned hardware circuits are fixed or reconfigurable. For example, the processor is a hardware circuit implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a Neural Network Processing Unit (NPU), a Tensor Processing Unit (TPU), or a Deep Learning Processing Unit (DPU).
[0287] Figure 7A is a schematic diagram of the feature indication device proposed in an embodiment of this disclosure. As shown in Figure 7A, the feature indication device 7100 may include at least one of a transceiver module 7101, a processing module 7102, etc. In some embodiments, the transceiver module 7102 is used to receive a data stream. The processing module 7102 is used to determine whether the features in the decoded data stream are compressed, wherein the features are used for different AI tasks. Optionally, the transceiver module 7101 is used to perform at least one of the communication steps such as sending and / or receiving performed by the encoder in any of the above methods (e.g., step S2101, but not limited thereto), which will not be described in detail here. Optionally, the processing module is used to perform at least one of the other steps performed by the encoder in any of the above methods, which will not be described in detail here.
[0288] Optionally, the processing module 7102 is used to perform at least one of the communication steps, such as the processing performed by the encoder in any of the above methods, which will not be described in detail here.
[0289] Figure 7B is a schematic diagram of the feature indication device proposed in an embodiment of this disclosure. As shown in Figure 7B, the feature indication device 7200 may include at least one of a transceiver module 7201, a processing module 7202, etc. In some embodiments, the transceiver module 7202 is used to send a data stream; the data stream is used to determine whether the decoded feature is compressed, wherein different features are used for different AI tasks. Optionally, the transceiver module is used to perform at least one of the communication steps such as sending and / or receiving performed by the decoder in any of the above methods, which will not be described in detail here.
[0290] Optionally, the processing module 7202 is used to perform at least one of the communication steps, such as the processing performed by the decoder in any of the above methods, which will not be described in detail here.
[0291] In some embodiments, the transceiver module may include a transmitting module and / or a receiving module, which may be separate or integrated. Optionally, the transceiver module may be interchangeable with a transceiver.
[0292] In some embodiments, the processing module may be a single module or may include multiple sub-modules. Optionally, the multiple sub-modules may each perform all or part of the steps required by the processing module. Optionally, the processing module may be interchangeable with a processor.
[0293] In some embodiments, the transceiver module may include a transmitting module and / or a receiving module, which may be separate or integrated. Optionally, the transceiver module may be interchangeable with a transceiver.
[0294] In some embodiments, the processing module may be a single module or may include multiple sub-modules. Optionally, the multiple sub-modules may each perform all or part of the steps required by the processing module. Optionally, the processing module may be interchangeable with a processor.
[0295] Figure 8A is a schematic diagram of the structure of the communication device 8100 proposed in an embodiment of this disclosure. The communication device 8100 can be a second device, an encoder, a decoder, or a chip, chip system, or processor that supports the second device, encoder, or decoder in implementing any of the above methods. The communication device 8100 can be used to implement the methods described in the above method embodiments; for details, please refer to the descriptions in the above method embodiments.
[0296] As shown in Figure 8A, the communication device 8100 includes one or more processors 8101. The processor 8101 can be a general-purpose processor or a dedicated processor, such as a baseband processor or a central processing unit (CPU). The baseband processor can be used to process communication protocols and communication data, while the CPU can be used to control the feature indication device, execute programs, and process program data. The communication device 8100 is used to execute any of the above methods.
[0297] In some embodiments, the communication device 8100 further includes one or more memories 8102 for storing instructions. Optionally, all or part of the memories 8102 may also be located outside the communication device 8100.
[0298] In some embodiments, the communication device 8100 further includes one or more transceivers 8103. When the communication device 8100 includes one or more transceivers 8103, the transceivers 8103 perform at least one of the communication steps such as sending and / or receiving in the above-described method.
[0299] In some embodiments, a transceiver may include a receiver and / or a transmitter, which may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, transceiver circuit, etc., may be used interchangeably; the terms transmitter, transmitting unit, transmitter, transmitting circuit, etc., may be used interchangeably; and the terms receiver, receiving unit, receiver, receiving circuit, etc., may be used interchangeably.
[0300] In some embodiments, the communication device 8100 may include one or more interface circuits 8104. Optionally, the interface circuit 8104 is connected to the memory 8102, and the interface circuit 8104 can be used to receive signals from the memory 8102 or other devices, and can be used to send signals to the memory 8102 or other devices. For example, the interface circuit 8104 can read instructions stored in the memory 8102 and send the instructions to the processor 8101.
[0301] The communication device 8100 described in the above embodiments may be a second device, a decoder, or an encoder, but the scope of the communication device 8100 described in this disclosure is not limited thereto, and the structure of the communication device 8100 may not be limited by FIG8A. The communication device may be a standalone device or may be part of a larger device. For example, the communication device may be: (1) a standalone integrated circuit IC, or chip, or chip system or subsystem; (2) a collection of one or more ICs, optionally, the IC collection may also include storage components for storing data and programs; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, encoder device, smart encoder device, cellular phone, wireless device, handheld device, mobile unit, vehicle device, decoder, cloud device, artificial intelligence device, etc.; (6) others, etc.
[0302] Figure 8B is a schematic diagram of the structure of chip 8200 according to an embodiment of this disclosure. For cases where the communication device 8100 can be a chip or a chip system, please refer to the schematic diagram of chip 8200 shown in Figure 8B, but it is not limited thereto.
[0303] Chip 8200 includes one or more processors 8201, which are used to perform any of the above methods.
[0304] In some embodiments, chip 8200 further includes one or more interface circuits 8202. Optionally, the interface circuit 8202 is connected to memory 8203, and the interface circuit 8202 can be used to receive signals from memory 8203 or other devices, and the interface circuit 8202 can be used to send signals to memory 8203 or other devices. For example, the interface circuit 8202 can read instructions stored in memory 8203 and send the instructions to processor 8201.
[0305] In some embodiments, the interface circuit 8202 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processor 8201 performs at least one of the other steps.
[0306] In some embodiments, the terms interface circuit, interface, transceiver pin, transceiver, etc., can be used interchangeably.
[0307] In some embodiments, chip 8200 further includes one or more memories 8203 for storing instructions. Optionally, all or part of the memories 8203 may be located outside of chip 8200.
[0308] This disclosure also proposes a storage medium storing instructions that, when executed on a communication device 8100, cause the communication device 8100 to perform any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but not limited thereto; it may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but not limited thereto; it may also be a temporary storage medium.
[0309] This disclosure also provides a program product that, when executed by the communication device 8100, causes the communication device 8100 to perform any of the above methods. Optionally, the program product is a computer program product.
[0310] This disclosure also proposes a computer program that, when run on a computer, causes the computer to perform any of the above methods.
Claims
1. A feature indication method, characterized in that, The method is executed by a decoder, and the method includes: Receive data stream; Determine whether the features in the decoded data stream are compressed, wherein the features are used for different AI tasks.
2. The method according to claim 1, characterized in that, Determining whether the features in the decoded data stream are compressed includes: Whether the feature is compressed is determined based on the first information in the decoded data stream.
3. The method according to claim 2, characterized in that, When the first information indicates that the feature is compressed, the feature included in the data stream is compressed; If the first information indicates that the feature is not compressed, the feature included in the data stream is not compressed.
4. The method according to claim 3, characterized in that, When the first information indicates that the feature is compressed, the method further includes: The data stream is decoded to obtain the compressed features; The compressed features are decompressed to obtain the decompressed features.
5. The method according to any one of claims 1 to 4, characterized in that, The data stream also includes data for monitoring, and the features are extracted based on the data for monitoring.
6. The method according to any one of claims 1 to 5, characterized in that, The features used for monitoring include at least one of the following: Machine characteristics; Auditory characteristics.
7. The method according to claim 6, characterized in that, The machine features include at least one of the following: Emotion recognition ER features; Automatic Speech Recognition (ASR) Features; Automatic voice verification of ASV features; Audio event classification AEC features.
8. The method according to any one of claims 1 to 7, characterized in that, The data stream includes at least one flag bit, with different flag bits corresponding to different AI tasks. The first information on the flag bit is used to indicate whether the feature is compressed.
9. The method according to claim 8, characterized in that, The first information includes a first value, which indicates that the feature is compressed; or, The first information includes a second value, which indicates that the feature has not been compressed.
10. The method according to any one of claims 4 to 9, characterized in that, The data used for monitoring includes at least one of the following: Training data, which is used to train the model; Validation data, which is used to optimize the model; Test data, which is used to test the model.
11. The method according to any one of claims 4 to 9, characterized in that, The data used for monitoring includes raw data and tag data, and the tag data is used to classify the raw data.
12. A feature indication method, characterized in that, The method is executed by an encoder, and the method includes: Send data stream; The data stream is used to determine whether the decoded features are compressed, wherein different features are used for different AI tasks.
13. The method according to claim 12, characterized in that, The first piece of information in the data stream indicates whether the feature is compressed.
14. The method according to claim 13, characterized in that, If the first information indicates that the feature is compressed, it is determined that the feature included in the data stream is compressed; If the first information indicates that the feature is not compressed, it is determined that the feature included in the data stream is not compressed.
15. The method according to any one of claims 12 to 14, characterized in that, The data stream also includes data for monitoring, and the features are extracted based on the data.
16. The method according to any one of claims 12 to 15, characterized in that, The first feature includes at least one of the following: Machine characteristics; Auditory characteristics.
17. The method according to claim 16, characterized in that, The machine features include at least one of the following: Emotion recognition ER features; Automatic Speech Recognition (ASR) Features; Automatic voice verification of ASV features; Audio event classification AEC features.
18. The method according to any one of claims 12 to 17, characterized in that, The data stream includes at least one flag bit, with different flag bits corresponding to different AI tasks. The first information on the flag bit is used to indicate whether the feature is compressed.
19. The method according to claim 18, characterized in that, The first information includes a first value, which indicates that the feature is compressed; or, The first information includes a second value, which indicates that the feature has not been compressed.
20. The method according to any one of claims 15 to 17, characterized in that, The data used for monitoring includes at least one of the following: Training data, which is used to train the model; Validation data, which is used to update the parameters of the model; Test data, which is used to test the accuracy of the model.
21. The method according to any one of claims 15 to 17, characterized in that, The data used for monitoring includes raw data and tag data, and the tag data is used to classify the raw data.
22. A feature indication method, characterized in that, The method includes: The encoder sends a data stream; The decoder receives the data stream; The decoder determines whether there are compressed features in the decoded data stream, wherein different compressed features are used for different AI tasks.
23. An encoding device, characterized in that, The encoding device includes: The transceiver module is used to receive data streams; The processing module is used to determine whether there are compressed features in the decoded data stream, wherein the compressed features are used for different AI tasks.
24. A decoding device, characterized in that, The decoding device includes: The transceiver module is used to send a data stream; the data stream is used to determine whether there are compressed features after decoding, wherein different features are used for different AI tasks.
25. A decoder, characterized in that, The decoder includes: One or more processors; The processor is used to execute the feature indication method according to any one of claims 1 to 11.
26. An encoder, characterized in that, The encoder includes: One or more processors; The processor is configured to execute the feature indication method according to any one of claims 12 to 21.
27. A system, characterized in that, It includes an encoder and a decoder, wherein the decoder is configured to implement the feature indication method according to any one of claims 1 to 11, and the encoder is configured to implement the feature indication method according to any one of claims 12 to 21.
28. A storage medium storing instructions, characterized in that, When the instruction is executed on the communication device, the communication device performs the feature indication method as described in any one of claims 1 to 11, or performs the feature indication method as described in any one of claims 12 to 21.
29. A computer program product, characterized in that, When the computer program product is run on a communication device, it causes the communication device to perform the feature indication method as described in any one of claims 1 to 21.