Information indication method, encoding end, decoding end, apparatus, and communication system

By introducing an information indication method between the encoding and decoding ends, and using flag bits in the bitstream to judge and decode metadata encoding, the problem of unknown metadata encoding in the bitstream is solved, ensuring the accuracy of data and features in AI tasks, and is applicable to a variety of AI tasks and scenarios.

WO2026156640A1PCT designated stage Publication Date: 2026-07-30BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING XIAOMI MOBILE SOFTWARE CO LTD
Filing Date
2025-01-23
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

The inability to determine whether the bitstream contains metadata encoding makes it impossible for the decoding end to accurately identify and decode the metadata, affecting the accuracy of data and feature acquisition for AI tasks.

Method used

An information indication method is introduced between the encoding and decoding ends. The presence of metadata encoding is determined by the flag bits in the bit stream, and the corresponding decoder is used to decode according to different flag bits, so as to ensure the accuracy and diversity of metadata encoding.

Benefits of technology

It enables accurate judgment and decoding of metadata encoding in bitstreams at the decoding end, ensuring accurate acquisition of data and features in AI tasks, and is applicable to AI tasks of different types and application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025074431_30072026_PF_FP_ABST
    Figure CN2025074431_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to an information indication method, an encoding end, a decoding end, and a communication system. The information indication method comprises: receiving a bitstream; and determining whether the bitstream comprises metadata coding, wherein the metadata coding is used for determining decoded metadata, and the decoded metadata is used for an AI task. By means of the method, decoding efficiency of an audio feature by the decoding end can be further improved, and the accuracy of obtaining data on the basis of the received bitstream can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Information indication method, encoding end, decoding end, device and communication system Technical Field

[0001] This disclosure relates to the field of communication technology, and in particular to an information indication method, encoding end, decoding end, apparatus and communication system. Background Technology

[0002] With the rapid development of encoding / decoding and communication technologies, data used for AI tasks can be encoded into encoded data and then transmitted, ensuring transmission efficiency. Furthermore, features for AI models can be extracted from the data, encoded, and transmitted, reducing the amount of data during communication. Summary of the Invention

[0003] This disclosure provides an information indication method, an encoding end, a decoding end, an apparatus, and a communication system, which solves the problem of not being able to know whether the bit stream contains metadata encoding. It ensures that the bit stream received by the decoding end can determine whether it contains metadata encoding, thereby determining the decoder to decode the metadata encoding, and ensuring the accuracy of data and / or features obtained based on the received bit stream.

[0004] In a first aspect, embodiments of this disclosure provide an information indication method, the method being executed by a decoding end, the decoding end including at least one decoder, different decoders corresponding to different AI tasks, the method including:

[0005] Receive bit stream;

[0006] Determine whether the bitstream includes metadata encoding;

[0007] The metadata encoding is used to determine the decoded metadata, which is then used for AI tasks.

[0008] Secondly, this disclosure also provides an information indication method, which is executed by an encoding end, and the method includes:

[0009] Send bit stream;

[0010] The bitstream is used to determine whether it includes metadata encoding;

[0011] The metadata encoding is used to determine the decoded metadata, which is then used for AI tasks.

[0012] Thirdly, embodiments of this disclosure also provide a signal indication method, including:

[0013] The encoding end sends the bit stream;

[0014] The decoding end receives the bit stream;

[0015] The decoding end determines whether the bitstream includes metadata encoding. The metadata encoding is used to determine the decoded metadata, which is used for AI tasks.

[0016] Fourthly, embodiments of this disclosure also provide a decoding apparatus, including:

[0017] The transceiver module is used to receive bit streams;

[0018] The processing module is used to determine whether there is metadata encoding in the decoded bitstream. The metadata encoding is used to determine the decoded metadata, and the decoded metadata is used for AI tasks.

[0019] Fifthly, embodiments of this disclosure also provide an encoding device, including:

[0020] The transceiver module is used to send bit streams;

[0021] The bitstream is used to determine whether it includes metadata encoding;

[0022] The metadata encoding is used to determine the decoded metadata, which is then used for AI tasks.

[0023] Sixthly, embodiments of this disclosure provide a decoding terminal, which includes:

[0024] One or more processors;

[0025] The decoding end is used to execute the signal indication method described in the first aspect of the embodiments of this disclosure.

[0026] Seventhly, embodiments of this disclosure provide an encoding terminal, which includes:

[0027] One or more processors;

[0028] The encoding end is used to execute the signal indication method described in the second aspect of the embodiments of this disclosure.

[0029] Eighthly, embodiments of this disclosure also provide a communication device for performing the signal indication method described in the first or second aspect.

[0030] In a ninth aspect, embodiments of this disclosure also provide a communication system, including an encoding end and a decoding end;

[0031] The decoding end is configured to implement the signal indication method described in the first aspect, and the encoding end is configured to implement the signal indication method described in the second aspect.

[0032] In a tenth aspect, embodiments of this disclosure also provide a storage medium storing instructions that, when executed on a communication device, cause the communication device to perform a signal indication method as described in the first aspect of embodiments of this disclosure, or to perform a signal indication method as described in the second aspect of embodiments of this disclosure.

[0033] Eleventhly, embodiments of this disclosure also provide a program product, including at least one of a program and instructions, wherein when the program or instructions are executed by a communication device, they implement the signal indication method described in the first aspect or the signal indication method described in the second aspect.

[0034] In this embodiment of the disclosure, the decoding end determines whether the received bitstream includes metadata encoding; wherein the metadata encoding is used to determine the decoded metadata, and the decoded metadata is used for AI tasks. This solves the problem of not being able to know whether the bitstream includes metadata encoding, ensuring that the bitstream received by the decoding end can be determined to include metadata encoding, and ensuring the accuracy of data and / or features obtained based on the received bitstream.

[0035] Additional aspects and advantages of embodiments of this disclosure will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of this disclosure. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings required for the description of the embodiments are introduced below. The following drawings are only some embodiments of this disclosure and do not impose specific limitations on the protection scope of this disclosure.

[0037] Figure 1 is a schematic diagram of the architecture of the communication system provided in an embodiment of this disclosure;

[0038] Figure 2 is an interactive schematic diagram of the signal indication method provided in an embodiment of this disclosure;

[0039] Figure 3A is a schematic flowchart illustrating a signal indication method applied to an encoding end according to an embodiment of this disclosure;

[0040] Figure 3B is a flowchart illustrating a signal indication method applied to an encoding end according to an embodiment of this disclosure;

[0041] Figure 4A is a schematic flowchart illustrating a signal indication method applied to the decoding end according to an embodiment of this disclosure;

[0042] Figure 4B is a schematic flowchart illustrating a signal indication method applied to the decoding end according to an embodiment of this disclosure;

[0043] Figure 5 is a flowchart illustrating a signal indication method according to an embodiment of this disclosure;

[0044] Figure 6A is a schematic diagram of the structure of the decoding device proposed in an embodiment of this disclosure;

[0045] Figure 6B is a schematic diagram of the structure of the encoding device proposed in an embodiment of this disclosure;

[0046] Figure 7 is a schematic diagram of the structure of the communication device proposed in an embodiment of this disclosure;

[0047] Figure 8 is a schematic diagram of the chip structure proposed in an embodiment of this disclosure. Detailed Implementation

[0048] This disclosure presents a communication method, communication device, and communication system.

[0049] In a first aspect, embodiments of this disclosure propose an information indication method, the method being executed by a decoding end, the decoding end including at least one decoder, different decoders corresponding to different AI tasks, the method comprising:

[0050] Receive bit stream;

[0051] Determine whether the bitstream includes metadata encoding;

[0052] The metadata encoding is used to determine the decoded metadata, which is then used for AI tasks.

[0053] In the above embodiments, the decoding end determines whether the received bitstream includes metadata encoding. The metadata encoding is used to determine the decoded metadata, which is then used for AI tasks. This solves the problem of not being able to determine whether the bitstream contains metadata encoding, ensuring that the bitstream received by the decoding end can be used to determine whether it contains metadata encoding, and guaranteeing the accuracy of data and / or features acquired based on the received bitstream.

[0054] In conjunction with some embodiments of the first aspect, in some embodiments, determining whether the decoded bitstream includes metadata encoding includes:

[0055] Based on the first information in the bitstream, determine whether the bitstream includes metadata encoding.

[0056] In the above embodiments, the decoding end determines whether the data stream includes metadata encoding based on the first information, and then decodes the data stream to obtain the metadata encoding, ensuring the accuracy of the obtained metadata encoding.

[0057] In conjunction with some embodiments of the first aspect, in some embodiments, when the first information indicates the presence of metadata encoding, the bit stream includes the metadata encoding;

[0058] If the first information indicates that there is no metadata encoding, the bitstream does not include the metadata encoding.

[0059] In the above embodiments, if the first information indicates whether the data stream has metadata encoding, it can be intuitively determined whether the corresponding metadata encoding can be obtained when decoding the data stream, thus ensuring the accuracy of the obtained metadata encoding.

[0060] In conjunction with some embodiments of the first aspect, in some embodiments, the first information includes a first value, the first value being used to indicate the presence of the metadata encoding; or,

[0061] The first information includes a second value, which indicates that the metadata encoding does not exist.

[0062] In the above embodiments, the first information of the first flag bit indicates whether metadata encoding exists, ensuring the accuracy of the indication of each metadata encoding, and thus ensuring the accuracy of the features obtained by subsequent decoding.

[0063] In conjunction with some embodiments of the first aspect, in some embodiments, the decoding end includes at least one decoder, the bit stream includes at least one first flag bit, each first flag bit corresponds to one decoder, and each first flag bit includes the first information;

[0064] When the first information indicated by the second flag bit in the at least one first flag bit is present, the metadata encoding is decoded by the decoder corresponding to the second flag bit, and the metadata encoding decoded by different decoders is applied to different AI tasks.

[0065] In the above embodiments, different first flag bits correspond to different decoders, so that the existence of metadata encoding to be decoded by the corresponding decoder can be determined based on the first information of a first flag bit. If the first information indicated by the second flag bit in at least one of the first flag bits indicates that the metadata encoding exists, the metadata encoding is decoded by the decoder corresponding to the second flag bit, thereby ensuring the accuracy of the metadata obtained by subsequent decoding.

[0066] In conjunction with some embodiments of the first aspect, in some embodiments, the decoding end includes at least one first decoder, and the first decoder includes at least one second decoder;

[0067] When the first information indicated by the second flag bit indicates the existence of metadata encoding, decoding the metadata encoding through the decoder corresponding to the second flag bit includes:

[0068] If the first information indicating metadata encoding exists as indicated by the second flag bit corresponding to the first decoder, the first metadata is obtained by decoding the metadata encoding through the first decoder.

[0069] The second information in the first metadata is used to determine whether the first metadata should be decoded by the second decoder.

[0070] In the above embodiments, the decoder architecture at the decoding end includes two levels. One level includes at least one first decoder, and the other level includes at least one second decoder included in each first decoder. In this disclosure, during decoding, if the first information indicated by the second flag bit corresponding to the first decoder indicates that the metadata encoding exists, the first metadata is obtained by decoding the metadata encoding through the first decoder. Then, the second information in the first metadata is used to confirm whether the first metadata is decoded through the second decoder, thereby realizing ordered decoding.

[0071] In conjunction with some embodiments of the first aspect, in some embodiments, the second information includes a third flag bit corresponding to each of the second decoders, and the step of confirming whether to decode the first metadata using the second information in the first metadata includes:

[0072] When the third flag bit corresponding to the second decoder indicates that metadata encoding exists, the first metadata is decoded by the second decoder to obtain the second metadata, and the second metadata is used as the decoded metadata.

[0073] If the third flag corresponding to the second decoder indicates that the metadata encoding does not exist, the first metadata is used as the decoded metadata.

[0074] In the above embodiments, the second information also includes a third flag bit corresponding to each second decoder included in the corresponding first decoder. The third flag bit is used to indicate whether there is metadata encoding of the corresponding second decoder. Thus, if the third flag bit corresponding to the second decoder indicates that metadata encoding exists, the first metadata is further decoded by the second decoder to obtain the second metadata, and the second metadata is used as the decoded metadata. If the third flag bit corresponding to the second decoder indicates that metadata encoding does not exist, the first metadata is used as the decoded metadata, thus realizing step-by-step decoding.

[0075] In conjunction with some embodiments of the first aspect, in some embodiments, when the second metadata is applied to a portion of the AI ​​model of the AI ​​task corresponding to the second decoder, the second metadata includes a fourth flag bit, different fourth flag bits correspond to different AI models, and the fourth flag bit is used to indicate whether there is metadata applied to the corresponding AI model;

[0076] When the second metadata is applied to all AI models of the AI ​​task corresponding to the second decoder, the second metadata does not include the fourth flag bit.

[0077] In the above embodiment, if the second metadata obtained by the second decoder is used for all AI models of the AI ​​task, the fourth flag bit can be used to save the space occupied by the metadata.

[0078] In conjunction with some embodiments of the first aspect, in some embodiments, the metadata encoding decoded by different first decoders is applied to different types of AI tasks;

[0079] Metadata encoding decoded by different second decoders is applied to AI tasks in different application scenarios;

[0080] or

[0081] Metadata encoding decoded by different first decoders is applied to AI tasks in different application scenarios;

[0082] The metadata encoding decoded by different second decoders is applied to different types of AI tasks. In conjunction with some embodiments of the first aspect, in some embodiments, the types of AI tasks include at least one of the following:

[0083] Emotion recognition (ER) task;

[0084] Automatic Speech Recognition (ASR) task;

[0085] Automatic voice verification of ASV tasks;

[0086] Audio event classification AEC task.

[0087] In the above embodiments, the types of metadata encoding included in the data stream are expanded to ensure the diversity of metadata encoding, thereby ensuring the reliability and diversity of subsequent processing based on metadata encoding characteristics.

[0088] In conjunction with some embodiments of the first aspect, in some embodiments, the application scenarios of the AI ​​model involved in the AI ​​task include at least one of the following:

[0089] Inference scenarios for AI models;

[0090] Training scenarios for AI models.

[0091] In conjunction with some embodiments of the first aspect, in some embodiments, when the first metadata includes second metadata with different quantization bits, the third flag bit corresponding to the second decoder indicates the presence of metadata encoding;

[0092] The step of obtaining decoded metadata by decoding the first metadata through the second decoder includes:

[0093] When the third flag bit corresponding to the second decoder indicates that metadata encoding exists, the second decoder decodes the first metadata to obtain the second metadata with the corresponding number of quantization bits.

[0094] In conjunction with some embodiments of the first aspect, in some embodiments, at least one of the bit stream and the second information further includes a fifth flag bit, the fifth flag bit being used to indicate whether there is a metadata encoding with the same quantization bit length and the same value, the metadata encoding with the same quantization bit length and the same value being used for multiple AI tasks.

[0095] In the above embodiments, by setting the fifth flag bit, when there is a metadata encoding with the same number of quantization bits and the same value, and the metadata encoding is used for multiple AI tasks, the encoder does not need to repeatedly transmit the metadata encoding, thus improving transmission efficiency.

[0096] Secondly, embodiments of this disclosure propose an information indication method, which is executed by an encoding end, and the method includes:

[0097] Send bit stream;

[0098] The bitstream is used to determine whether it includes metadata encoding;

[0099] The metadata encoding is used to determine the decoded metadata, which is then used for AI tasks.

[0100] In conjunction with some embodiments of the second aspect, in some embodiments, the bitstream includes first information used to determine whether the bitstream includes metadata encoding.

[0101] In conjunction with some embodiments of the second aspect, in some embodiments, when the first information indicates the presence of metadata encoding, the bitstream includes the metadata encoding;

[0102] If the first information indicates that there is no metadata encoding, the bitstream does not include the metadata encoding.

[0103] In conjunction with some embodiments of the second aspect, in some embodiments, the bit stream includes at least one first flag bit, each first flag bit corresponds to a decoder, and each first flag bit includes the first information;

[0104] When the first information indicated by the first flag bit indicates the existence of metadata encoding, the metadata encoding is decoded by the decoder corresponding to the first flag bit, and the metadata encoding decoded by different decoders is applied to different AI tasks.

[0105] In conjunction with some embodiments of the second aspect, in some embodiments the bitstream includes a plurality of first flag bits, different first flag bits corresponding to different decoders, and the first information on the first flag bits is used to indicate whether there is metadata encoding decoded by the corresponding decoder.

[0106] In conjunction with some embodiments of the second aspect, the types of AI tasks described in some embodiments include at least one of the following:

[0107] Emotion recognition (ER) task;

[0108] Automatic Speech Recognition (ASR) task;

[0109] Automatic voice verification of ASV tasks;

[0110] Audio event classification AEC task.

[0111] In conjunction with some embodiments of the second aspect, in some embodiments, the application scenarios of the AI ​​model involved in the AI ​​task include at least one of the following:

[0112] Inference scenarios for AI models;

[0113] Training scenarios for AI models.

[0114] In conjunction with some embodiments of the second aspect, in some embodiments, the AI ​​task is performed by multiple AI models;

[0115] In the case where the metadata encoding applied to the AI ​​task is applied by not all of the multiple AI models, the data stream includes a first flag bit corresponding to the first decoder; the first decoder corresponds to the machine task.

[0116] When the metadata encoding applied to the AI ​​task is applied by all of the multiple AI models, the data stream does not include the first flag bit corresponding to the first decoder;

[0117] If the value of the metadata encoding in the bitstream changes, the bitstream includes a first flag bit corresponding to the decoder that decodes the metadata encoding;

[0118] If the value of the metadata encoding in the bitstream remains unchanged, the bitstream does not include a first flag bit corresponding to the decoder that decodes the metadata encoding;

[0119] If the transmission period of the metadata encoding is met at the current moment and the value of the metadata encoding has not changed, the bit stream includes a first flag bit corresponding to the decoder that decodes the metadata encoding.

[0120] In conjunction with some embodiments of the second aspect, in some embodiments,

[0121] If the value of the metadata encoding changes, the bitstream includes a first flag bit corresponding to the decoder that decodes the metadata encoding;

[0122] If the value of the metadata encoding remains unchanged, the bitstream does not include the first flag bit corresponding to the decoder that decodes the metadata encoding;

[0123] If the transmission period of the metadata encoding is met at the current moment and the value of the metadata encoding has not changed, the bit stream includes a first flag bit corresponding to the decoder that decodes the metadata encoding.

[0124] Thirdly, this disclosure also provides an information indication method, including:

[0125] The encoding end sends the bit stream;

[0126] The decoding end receives the bit stream.

[0127] The decoding end determines whether the bitstream includes metadata encoding;

[0128] The metadata encoding is used to determine the decoded metadata, which is then used for AI tasks.

[0129] Fourthly, embodiments of this disclosure also provide a decoding apparatus, including:

[0130] The transceiver module is used to receive bit streams;

[0131] The processing module is used to determine whether the bitstream includes metadata encoding;

[0132] The metadata encoding is used to determine the decoded metadata, which is then used for AI tasks.

[0133] Fifthly, embodiments of this disclosure also provide an encoding device, including:

[0134] The transceiver module is used to send a bit stream, which is used to determine whether it includes metadata encoding;

[0135] The metadata encoding is used to determine the decoded metadata, which is then used for AI tasks.

[0136] Sixthly, embodiments of this disclosure provide a decoding terminal, which includes:

[0137] One or more processors;

[0138] The decoding end is used to execute the signal indication method described in the first aspect of the embodiments of this disclosure.

[0139] Seventhly, embodiments of this disclosure provide an encoding terminal, which includes:

[0140] One or more processors;

[0141] The encoding end is used to execute the signal indication method described in the second aspect of the embodiments of this disclosure.

[0142] Eighthly, embodiments of this disclosure also provide a system including an encoding end and a decoding end;

[0143] The decoding end is configured to implement the signal indication method described in the first aspect, and the encoding end is configured to implement the signal indication method described in the second aspect.

[0144] Ninthly, embodiments of this disclosure also provide a storage medium storing instructions that, when executed on a communication device, cause the communication device to perform a signal indication method as described in the first aspect of this disclosure, or to perform a signal indication method as described in the second aspect of this disclosure.

[0145] In a tenth aspect, embodiments of this disclosure provide a program product that, when executed by a communication device, causes the communication device to perform the method as described in an optional implementation of the first or second aspect.

[0146] In one aspect, embodiments of this disclosure provide a computer program that, when run on a computer, causes the computer to perform the methods described in an optional implementation of the first or second aspect.

[0147] In a twelfth aspect, embodiments of this disclosure provide a chip or chip system. The chip or chip system includes processing circuitry configured to perform the method described in an optional implementation of the first or second aspect above.

[0148] This disclosure provides an information indication method, an encoding end, a decoding end, and a communication system. In some embodiments, the terms "information indication method" and "signal transmission method," "wireless frame transmission method," etc., can be used interchangeably, as can the terms "information processing system," "communication system," etc.

[0149] This disclosure is not exhaustive, but merely illustrative of some embodiments, and is not intended to limit the scope of protection of this disclosure. Unless otherwise specified, each step in a particular embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a particular embodiment can also be implemented as an independent embodiment, and the order of the steps in a particular embodiment can be arbitrarily interchanged. Furthermore, the optional implementation methods in a particular embodiment can be arbitrarily combined; moreover, the embodiments can be arbitrarily combined, for example, some or all steps of different embodiments can be arbitrarily combined, and a particular embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.

[0150] In each of the disclosed embodiments, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of the embodiments are consistent and can be referenced by each other. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.

[0151] The terminology used in the embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure.

[0152] In the embodiments disclosed herein, "multiple" refers to two or more.

[0153] In some embodiments, the terms “at least one of A or B, at least one of A and B”, “one or more”, “a plurality of”, “multiple”, etc., may be used interchangeably.

[0154] In some embodiments, the notation "at least one of A and B", "A and / or B", "A in one case, B in another", "in response to one case A, in response to another case B", etc., may include the following technical solutions depending on the situation: in some embodiments, A (execute A regardless of whether there is a branch B); in some embodiments, B (execute B regardless of whether there is a branch A); in some embodiments, execution is selected from A and B (A and B are selectively executed); in some embodiments, both A and B are executed. The same applies when there are more branches such as A, B, C, etc.

[0155] In some embodiments, the notation "A or B" may include the following technical solutions, depending on the situation: in some embodiments, A (execute A regardless of whether a branch B exists); in some embodiments, B (execute B regardless of whether a branch A exists); in some embodiments, execution is selected from A and B (A and B are selectively executed). The same applies when there are more branches such as A, B, and C.

[0156] The prefixes "first," "second," etc., used in the embodiments of this disclosure are merely for distinguishing different descriptive objects and do not impose restrictions on the position, order, priority, quantity, or content of the descriptive objects. The description of the descriptive objects is found in the claims or the context of the embodiments, and the use of prefixes should not constitute unnecessary restrictions. For example, if the descriptive object is a "field," the ordinal numbers preceding "field" in "first field" and "second field" do not restrict the position or order of the "fields." "First" and "second" do not restrict whether the "fields" they modify are in the same message, nor do they restrict the order of "first field" and "second field." Similarly, if the descriptive object is a "level," the ordinal numbers preceding "level" in "first level" and "second level" do not restrict the priority between "levels." Furthermore, the number of descriptive objects is not limited by ordinal numbers and can be one or more. For example, in "first device," the number of "devices" can be one or more. Furthermore, the objects modified by different prefixes can be the same or different. For example, if the object being described is "device", then "first device" and "second device" can be the same device or different devices, and their types can be the same or different. Similarly, if the object being described is "information", then "first information" and "second information" can be the same information or different information, and their content can be the same or different.

[0157] In some embodiments, “including A,” “containing A,” “for indicating A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.

[0158] In some embodiments, terms such as "time / frequency" and "time-frequency domain" refer to the time domain and / or frequency domain.

[0159] In some embodiments, terms such as “in response to…”, “in response to determining…”, “in the case of…”, “when…”, “when…”, “if…”, etc. can be used interchangeably. These descriptions all refer to the device making a corresponding action under certain objective circumstances. They do not necessarily limit the time, nor do they require the device to make a judgment action when implementing it, nor do they mean that there must be other limitations.

[0160] In some embodiments, the terms “greater than,” “greater than or equal to,” “not less than,” “more than,” “more than or equal to,” “not less than,” “higher than,” “higher than or equal to,” “not lower than,” and “above” can be used interchangeably, as can the terms “less than,” “less than or equal to,” “not greater than,” “less than,” “less than or equal to,” “not more than,” “lower than,” “lower than or equal to,” “not higher than,” and “below”.

[0161] In some embodiments, devices, etc., may be interpreted as physical or virtual, and their names are not limited to those described in the embodiments. Terms such as “device,” “equipment,” “circuit,” “network element,” “network function,” “network device,” “function,” “node,” “unit,” “section,” “system,” “network,” “chip,” “chip system,” “entity,” and “subject” are interchangeable.

[0162] In some embodiments, "network" can be interpreted as devices included in a network (e.g., access network devices, core network devices, etc.).

[0163] In addition, terms such as "uplink" and "downlink" can be replaced with terms corresponding to inter-terminal communication (e.g., "side"). For example, uplink channel and downlink channel can be replaced with side channel, and uplink link and downlink link can be replaced with side link.

[0164] In some embodiments, "link" can mean "connection" or "link"; in various embodiments, "connection" and "link" can be used interchangeably.

[0165] In some embodiments, the acquisition of data, information, etc., may comply with the laws and regulations of the country where the location is situated.

[0166] In some embodiments, data, information, etc., may be obtained with the user's consent.

[0167] Furthermore, each element, each row, or each column in the table of this disclosure can be implemented as an independent embodiment, and any combination of any element, any row, or any column can also be implemented as an independent embodiment.

[0168] Figure 1 is a schematic diagram of the architecture of a communication system according to an embodiment of the present disclosure.

[0169] As shown in Figure 1, the communication system 100 includes an encoding end 101 and a decoding end 102.

[0170] In some embodiments, the encoding end 101 can be any electronic device with processing capabilities, such as a server or a terminal. The decoding end 102 can be any electronic device with processing capabilities, such as a server or a terminal.

[0171] In some embodiments, the encoding end 101 may encode (or compress) the raw data. The raw data may be at least one of the audio features and metadata of the audio frame.

[0172] In some embodiments, the encoding end 101 compresses the original data to meet transmission or storage needs. The output of the encoding end 101 can be a bitstream, also known as a data stream. For example, in scenarios where the network or bus with limited transmission bandwidth cannot meet the real-time transmission requirements of the original data, the encoding end 101 compresses the original data. As another example, in scenarios where storage devices with limited storage space cannot meet the storage requirements of the original data, the encoding end 101 compresses the original data. The bitstream is used to determine whether it includes metadata encoding, and the metadata encoding is used to determine the decoded metadata, which is used for AI tasks.

[0173] In some embodiments, the decoding end 102 can receive an encoded dataset sent by the encoding end 101, and the encoded dataset can be transmitted in the form of a bit stream. The decoding end 102 can decode the dataset, determine whether it includes metadata encoding, and distribute different decoded metadata to different AI tasks for processing.

[0174] In some embodiments, the terminal may be a user equipment (UE), including, but not limited to, at least one of the following: mobile phone, wearable device, Internet of Things device, car with communication function, smart car, tablet computer, computer with wireless transceiver function, virtual reality (VR) terminal device, augmented reality (AR) terminal device, wireless terminal device in industrial control, wireless terminal device in self-driving, wireless terminal device in remote medical surgery, wireless terminal device in smart grid, wireless terminal device in transportation safety, wireless terminal device in smart city, and wireless terminal device in smart home.

[0175] It is understood that the communication system described in this disclosure is for the purpose of more clearly illustrating the technical solutions of this disclosure, and does not constitute a limitation on the technical solutions proposed in this disclosure. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions proposed in this disclosure are also applicable to similar technical problems.

[0176] The following embodiments of this disclosure can be applied to the communication system 100 shown in FIG1, or to some of the main bodies, but are not limited thereto. The main bodies shown in FIG1 are illustrative. The communication system may include all or some of the main bodies in FIG1, or may include other main bodies outside of FIG1. ​​The number and form of each main body are arbitrary. Each main body may be physical or virtual. The connection relationship between the main bodies is illustrative. The main bodies may not be connected or may be connected. The connection can be in any way, it can be a direct connection or an indirect connection, it can be a wired connection or a wireless connection.

[0177] The embodiments disclosed herein can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G new radio (NR), 6th generation mobile communication system (6G), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New radio access (NX), Future generation radio access (FX), Global System for Mobile communications (GSM), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi), IEEE 802.16 (WiMAX), and IEEE 802.20, Ultra-Wideband (UWB), Bluetooth (a registered trademark), Public Land Mobile Network (PLMN) networks, Device-to-Device (D2D) systems, Machine-to-Machine (M2M) systems, Internet of Things (IoT) systems, Vehicle-to-Everything (V2X) systems, systems utilizing other data processing methods, and next-generation systems built upon them, etc. Furthermore, multiple systems can be combined (e.g., a combination of LTE or LTE-A with 5G).

[0178] Sound waves can be divided into several frequency bands based on their different frequencies. Sound waves with frequencies below 20 Hz are called infrasound; sound waves with frequencies between 20 Hz and 20 kHz are called audible sound waves; and sound waves with frequencies above 20 kHz are called ultrasound.

[0179] Ultrasound frequencies typically used in medical diagnosis range from 1 to 5 MHz. Ultrasound is characterized by its good directionality, strong penetrating power, ease of obtaining relatively concentrated sound energy, and long propagation distance in water. It can be used for ranging, speed measurement, cleaning, welding, and stone removal. It can also be applied to ultrasonic welding, ultrasonic chemistry, ultrasonic cleaning, ultrasonic processing (drilling, carving, polishing, etc.), ultrasonic therapy, ultrasonic surgery, ultrasonic beauty treatments, ultrasonic motors, and ultrasonic levitation.

[0180] Infrasound can freely travel through areas virtually inaccessible to light and radio waves, such as the ocean and underground strata. Due to its properties, it can be used to explore deeply buried mineral deposits, measure the distribution of hot and cold air masses in the stratosphere, and inspect operating machinery for potential hazards. It can also be used to predict natural phenomena such as tsunamis, storms, volcanic eruptions, and geomagnetic storms. Therefore, infrasound can be used to detect weather patterns, earthquakes, and forecast typhoons and tsunamis.

[0181] The human ear can perceive sound waves in the frequency range of 20Hz-20kHz. Current lossy audio coding algorithms are designed based on this frequency range and use masking effects to make distortion less perceptible to the human ear. However, in increasingly automated fields, computers are used to "listen" to audio data, such as automated inspection in automated manufacturing plants. Sensors collect audio data during product manufacturing. Computers analyze the audio data to determine if the product is defective and take action. For subsequent processing, traditional lossy audio coding algorithms are unsuitable because valid audio frequency components are lost.

[0182] Furthermore, audio signals in the medical field, such as brainwaves (brainwaves are electrical oscillations produced by the activity of nerve cells in the human brain. Because these oscillations appear as waves on scientific instruments, they are called brainwaves) have a different frequency range compared to the human ear. Based on frequency, brainwaves can be divided into five categories: beta waves (conscious 14-30Hz), alpha waves (bridged conscious 8-14Hz), theta waves (subconscious 4-8Hz), delta waves (subconscious 4Hz or lower), and gamma waves (focused on something (30Hz or higher)). The combination of these conscious states shapes a person's internal and external behavior, emotions, and learning performance. Compressing brainwave signals using traditional lossy audio encoding algorithms also results in the loss of audio components.

[0183] Therefore, in the field of increasingly automated control, there is a need for a new lossless or near-lossless sound encoding algorithm to effectively store sound data and transmit it to a computer for analysis.

[0184] The Advanced Audio Coding (AAC) compression algorithm used in the Moving Pictures Experts Group (MPEG) standard employs a psychoacoustic model to achieve lossy compression. The encoding and decoding block diagram is shown below. This algorithm considers the importance of audio components based on the sensitive frequency bands of the human ear and masking effects, compressing signal components that are not noticeable to the human ear. For machine listeners, this compression will result in the loss of valid audio components.

[0185] According to "ISO / IEC 14496-3:2005 / Amd 3:2006 (Scalable to Lossless Coding)," MPEG-4 SLS (Scalable to Lossless) or MPEG-4 Scalable to Lossless is an extension of the MPEG-4 Part 3 (MPEG-4 Audio) standard that allows lossless audio compression that can be scaled down to lossy using MPEG-4 common audio coding methods (e.g., variants of AAC).

[0186] The MPEG-4 SLS codec utilizes a lossless coding method based on the integer modified discrete cosine transform (IntMDCT). The IntMDCT spectral data is encoded using two complementary layers: the core MPEG-4 AAC layer and the Lossless Enhanced (LLE) layer. The core MPEG-4 AAC layer generates an AAC-compliant bitstream at a predefined bit rate that constitutes the minimum rate / quality unit of the lossless bitstream. The Lossless Enhanced (LLE) layer uses a bit-plane coding method to generate a fine-grained portion that can be scaled to the lossless bitstream.

[0187] The core layer AAC encoder in MPEG-4 SLS follows the information-rich AAC coding specification described in [reference needed]. The encoded information in the core AAC bitstream is then removed from the IntMDCT spectral data through an error-mapping process; the resulting IntMDCT spectral residuals are then encoded at the LLE encoder. This error-mapping process also attempts to preserve the probability distribution skew of the original IntMDCT coefficients, which approximates a Lapacian distribution, in the IntMDCT residuals, so that they can be encoded very efficiently by the entropy encoder used in the LLE layer.

[0188] In particular, for high sampling rate (96 kHz and above) inputs, the performance of this scalable system is further enhanced by a technique known as oversampling. In this way, the LLE encoder can operate at a preferred longer transform length, while the AAC core encoder / decoder can operate at a more suitable, lower sampling rate. For example, with an oversampling factor (osf) of 2, the AAC core can operate at 48 kHz, while the LLE encoder operates at 96 kHz, resulting in a frame length twice that of the AAC core (i.e., 2048 samples). In this case, the lower 1024 IntMDCT spectral value can be used as an approximation of the MDCT spectral value required by the AAC encoder. The error mapping process remains effective. In this case, the quantized AAC spectrum is mapped to the lower portion of the oversampled IntMDCT spectrum.

[0189] The MPEG-4 SLS codec provides a non-core mode for applications that only require lossless quality. This is achieved by simply disabling the AAC core used in the MPEG-4 SLS codec. Tests have shown that without the AAC core, the lossless compression performance of MPEG-4 SLS improves by 1% to 5%, and because the AAC encoder and decoder do not need to be implemented in the non-core mode of the MPEG-4 SLS codec, its computational complexity and implementation cost are also significantly reduced.

[0190] The 3GPP Enhanced Voice Services (EVS) compression algorithm uses a multi-core algorithm to achieve efficient compression of different types of audio signals. This algorithm uses a signal analysis module to analyze and classify audio frames. After this module, speech frames are compressed by a coding core based on linear precoding (LP), music signals are compressed by a frequency domain coding core, and silent frames are compressed by an inactive signal coding / CNG core. This algorithm considers the importance of audio components based on the sensitive frequency bands of the human ear and masking effects, and compresses signal components that are not obvious to the human ear. For machine monitors, compression based on the sensitive frequency bands of the human ear and masking effects will lose effective audio components.

[0191] In distributed speech recognition (DSR), the speech recognition front-end (FE) and back-end (BE) are separate. The front-end is located on the terminal device, such as a cellular phone, while the back-end is located in the telecommunications network.

[0192] In this way, the front end can capture high-quality speech and convert it in real time into feature extraction parameters or feature vectors, providing a relevant and compact representation of speech information for recognition. The feature vectors are transmitted as a bitstream with strong error protection to the back end—a high-performance speech recognition engine that performs the actual recognition task.

[0193] The recognition results are transmitted back to the terminal via the network. The concept of distributed speech recognition provides a framework for extracting high-quality feature vectors that are unaffected by speech coding or transmission effects. It has low computational requirements on the terminal and fully utilizes high-performance, multi-user speech recognition engines and other resources within the network.

[0194] In wireless networks, especially in bandwidth-constrained wireless networks, it is necessary to compress audio characteristics to reduce bandwidth usage.

[0195] Recommendation BS.2127, formally known as Audio Definition Model (ADM) Renderer for Advanced Sound Systems, specifies a baseline renderer for use with the audio-related metadata defined in Recommendation ITU-R BS.2051-2 and Recommendation ITU-R BS.2076-1, specifically the Audio Definition Model (ADM), including for program exchange. The audio renderer, based on provided content metadata and local context metadata, transforms a set of audio signals with related metadata into an audio signal and metadata with different configurations.

[0196] The overall architecture consists of several core components and processing steps. There are three types of data: input metadata, target environment, and audio channels. During target environment behavior initialization, the user can select a speaker layout from the speaker layout specified in "Recommendation ITU-R BS.2051-2". The rendering itself is divided into sub-components (object renderer, HOA (High-Order Stereo) renderer, and DirectSpeakers renderer) based on the project type (typeDefinition).

[0197] The ADM structure contains different levels describing the content and attributes of metadata: audioProgramme, audioContent, and audioObject. The ADM structure also uses different formats to establish the relationship between content and transmission / storage channels: audioPackFormat, audioChannelFormat, and audioBlockFormat.

[0198] MPEG metadata is a subset of ADM. However, ADM only defines metadata for rendering systems, which is insufficient for metadata types used for machine listening, such as sensor type and location, temperature, humidity, air pressure, etc.

[0199] Both ADM (Audio Definition Model) and MPEG-H object metadata contain very limited amounts of metadata and are unable to transmit the metadata required for AI tasks. Furthermore, the metadata in ADM and MPEG-H object metadata is designed for rendering algorithms used by the human ear, not for machine listening applications; therefore, their metadata formats differ significantly. Thus, it is necessary to design a metadata structure suitable for machine listening encoding.

[0200] When transmitting bitstreams, the related technologies do not include metadata for reproducing audio features, which leads to low efficiency in decoding audio features.

[0201] Based on this, the present disclosure provides an information indication method, including: receiving a bit stream; determining whether the bit stream includes metadata encoding; wherein the metadata encoding is used to determine the decoded metadata, and the decoded metadata is used for AI tasks; this solves the problem of not being able to know whether the bit stream includes metadata encoding, ensuring that the bit stream received by the decoding end can be determined to include metadata encoding, and ensuring the accuracy of data acquisition based on the received bit stream.

[0202] Figure 2 is an interactive schematic diagram of an information indication method according to an embodiment of the present disclosure. As shown in Figure 2, the method includes:

[0203] Step S2101: The encoding end obtains data for the audio task.

[0204] In some embodiments, the data disclosed herein for audio tasks includes at least human ear listening data and machine listening data, and optionally also metadata. Since each type is possible, an presence flag is added to improve transmission and storage efficiency.

[0205] Step S2102: The encoding end performs encoding to obtain a bit stream.

[0206] In some embodiments, the encoder encodes human ear listening data and machine listening data, as well as optional metadata, to obtain a bitstream.

[0207] In some embodiments, if the bitstream includes multiple types of data, the presence or absence of each type of data can be indicated. Optionally, the presence or absence of the corresponding data can be indicated by identification information included in the first information of the bitstream.

[0208] In some embodiments, each piece of first information corresponds to a flag bit in the bitstream, and different flag bits correspond to different data. In some embodiments, the bitstream includes at least one flag bit, with different flag bits corresponding to different data, and the first information on the flag bit is used to indicate whether the data exists. Optionally, the first information includes a first value, which is used to indicate the presence of a feature; or, the first information includes a second value, which is used to indicate the absence of a feature. For example, if the first value is 1 and the second value is 0, when the first information is 1, it indicates that the data exists; when the first information is 0, it indicates that the data does not exist.

[0209] Optionally, if the bitstream includes human ear monitoring data, machine monitoring data, and metadata encoding, the flag bits may include flag bit A, flag bit B, and flag bit C, where flag bit A corresponds to the human ear monitoring data encoding, flag bit B corresponds to the machine monitoring data encoding, and flag bit C corresponds to the metadata encoding. If the first information of flag bits A and B is 1, and the first information of flag bit C is 0, it indicates that the human ear monitoring data encoding and machine monitoring data encoding exist, but the metadata encoding does not exist.

[0210] For example, flags include identifiers indicating whether human listening data encoding exists in the bitstream, such as `human_listening_data_exist_flag`. Optionally, `human_listening_data_exist_flag` being 1 indicates that human listening data encoding exists in the bitstream, and `human_listening_data_exist_flag` being 0 indicates that human listening data encoding does not exist in the bitstream.

[0211] For example, flags include identifiers indicating whether machine listening data encoding exists in the bitstream, such as `machine_listening_data_exist_flag`. Optionally, `machine_listening_data_exist_flag` being 1 indicates that machine listening data encoding exists in the bitstream, and `machine_listening_data_exist_flag` being 0 indicates that machine listening data encoding does not exist in the bitstream.

[0212] For example, flags include identifiers indicating whether metadata encoding exists in the bitstream, such as metadata_exist_flag. Optionally, metadata_exist_flag being 1 indicates that metadata encoding exists in the bitstream, and metadata_exist_flag being 0 indicates that metadata encoding does not exist in the bitstream.

[0213] In some embodiments, metadata encodings only need to be restored or transmitted when their values ​​change; when their values ​​have not changed, the flag corresponding to the metadata encoding is set to 0.

[0214] In some embodiments, if the value of the metadata encoding in the bitstream changes, the bitstream includes a first flag bit corresponding to the decoder that decodes the metadata encoding.

[0215] In some embodiments, if the value of the metadata encoding in the bitstream remains unchanged, the bitstream does not include a first flag bit corresponding to the decoder that decodes the metadata encoding.

[0216] In some embodiments, if the current time conforms to the transmission period of the metadata encoding and the value of the metadata encoding has not changed, the bit stream includes a first flag bit corresponding to the decoder that decodes the metadata encoding.

[0217] In some embodiments, the metadata encoding is transmitted at fixed time intervals (e.g., half a second) regardless of whether the value of the metadata encoding changes, which makes it easier to randomly access the bit stream.

[0218] In some embodiments, if the data encoding for human ear listening exists in the bitstream, the decoding end will use a human ear listening decoder to encode and decode the human ear listening data into data that can be heard by the human ear.

[0219] In some embodiments, if the machine listening data encoding exists in the bitstream, the decoding end will use a machine listening decoder to decode the machine listening encoding into data available for machine tasks.

[0220] In some embodiments, if metadata encoding exists in the bitstream, the decoding end will use a metadata decoder to decode the metadata encoding into metadata that can be used for machine listening and human listening.

[0221] In some embodiments, different metadata is used for different AI tasks, and the decoding end includes at least one decoder, with different decoders corresponding to different AI tasks. By encoding metadata and decoding it with different decoders, metadata applicable to the corresponding AI task can be obtained. For example, if metadata encoding A is decoded with a decoder used for AI task 1, the obtained metadata will be used for AI task 1; if metadata encoding A is decoded with a decoder used for AI task 2, the obtained metadata will be used for AI task 2.

[0222] In some embodiments, the type of AI task includes at least one of the following:

[0223] Emotion Recognition (ER) task; Emotion recognition refers to identifying emotions in audio data. Optionally, the emotion includes anger, happiness, joy, etc., and this disclosure does not limit the emotion.

[0224] Automatic Speech Recognition (ASR) task; Automatic speech recognition refers to the automatic recognition of audio data, optionally including the recognition of text in the audio data;

[0225] Automatic Speaker Verification (ASV) task; Automatic Speaker Verification refers to determining whether a speaker is a pre-registered target speaker by analyzing the speaker's voice characteristics.

[0226] Audio event classification (AEC) task.

[0227] In some embodiments, the application scenarios of the AI ​​model involved in the AI ​​task include at least one of the following:

[0228] In the inference scenario of AI models, training of AI models refers to adjusting the parameters of AI models to ensure that the AI ​​models have the ability to recognize results that match the training data.

[0229] The training scenario of an AI model and the reasoning of an AI model refer to the process of using an AI model to predict or make decisions based on new and unseen data.

[0230] In some embodiments, the bitstream includes a plurality of first flag bits, different first flag bits correspond to different decoders, and the first information on the first flag bits is used to indicate whether there is metadata encoding decoded by the corresponding decoder.

[0231] Since each AI task may not be necessary, a first flag bit was added to improve transmission and storage efficiency. That is, a single codebase supports multiple backend tasks. For a specific application scenario or a specific backend task, the bitstream only contains the metadata encoding of the corresponding AI task, while the metadata of tasks not used in that application scenario only contains the first flag bit.

[0232] Step S2103: The encoding end sends the bit stream.

[0233] In this embodiment of the disclosure, after the encoding end obtains the bit stream, it can send the bit stream, and the subsequent encoding end can receive the bit stream.

[0234] Step S2104: The decoding end receives the bit stream.

[0235] Step S2105: The decoding end decodes the bit stream.

[0236] In some embodiments, the decoding end determines whether the bitstream includes metadata encoding.

[0237] In some embodiments, the decoding end determines whether the bitstream includes metadata encoding based on the first information in the bitstream.

[0238] In some embodiments, the decoding end determines that the bitstream includes the metadata encoding based on the presence of the first information; the decoding end determines that the bitstream does not include the metadata encoding based on the absence of the first information.

[0239] In some embodiments, the bitstream includes a plurality of first flag bits, different first flag bits correspond to different decoders, and the first information on the first flag bits is used to indicate whether there is metadata encoding decoded by the corresponding decoder.

[0240] For example, if the human_listening_data_exist_flag is 1, the decoding end determines that the human ear listening data encoding exists in the bitstream. The decoding end will then use the human ear listening decoder to encode and decode the audio data into audio features that can be heard by the human ear.

[0241] The decoding end determines that the machine listening data encoding exists in the bitstream based on machine_listening_data_exist_flag being 1. The decoding end will then use the machine task decoder to decode the machine listening data encoding into data that can be used by the machine task.

[0242] The decoding end determines that the metadata encoding exists in the bitstream based on the fact that metadata_data_exist_flag is 1. The decoding end will use a metadata encoder to decode the metadata encoding into metadata that can be used by machine tasks and human listening.

[0243] In some embodiments, the decoding end includes at least one decoder, the bit stream includes at least one first flag bit, each first flag bit corresponds to one decoder, and each first flag bit includes the first information.

[0244] When the first information indicated by the second flag bit in at least one of the first flag bits indicates the existence of metadata encoding, the metadata encoding is decoded by the decoder corresponding to the second flag bit. The metadata encoding decoded by different decoders is applied to different AI tasks. In some embodiments, the decoding end determines that the metadata encoding decoded by the decoder corresponding to AI model training exists in the bitstream based on the first information of the second flag bit nn_training_metadata_exist_flag being 1. Thus, the decoding end can use the decoder corresponding to AI model training to decode the metadata encoding and obtain the metadata used for AI model training.

[0245] In some embodiments, the decoding end determines that the metadata encoding decoded by the decoder corresponding to the AI ​​model inference exists in the bitstream based on the first information of the second flag bit nn_inference_metadata_exist_flag being 1. Thus, the decoding end can use the decoder corresponding to the AI ​​model inference to decode the metadata encoding and obtain the metadata used for AI model inference.

[0246] As can be seen from the above embodiments, the metadata is further subdivided into more metadata based on the type of AI task, on the basis of AI model training and inference scenarios. Correspondingly, the decoder type is also subdivided.

[0247] For example, if er_training_metadata_exist_flag is 1, it means that the metadata encoding corresponding to the training scenario of the AI ​​model for the ER task exists in the bitstream. The decoding end will use the decoder corresponding to the training scenario of the AI ​​model for the ER task to decode the metadata encoding and obtain the metadata used in the training scenario of the AI ​​model for the ER task.

[0248] In some embodiments, metadata can be initially categorized into metadata for training and inference scenarios of AI models, and further categorized according to the type of AI task, such as metadata for emotion recognition (ER) tasks, metadata for automatic speech recognition (ASR) tasks, metadata for automatic speech verification (ASV) tasks, metadata for audio event classification (AEC) tasks, etc. In this case, based on the first information in the bitstream, and considering that the bitstream includes metadata encoding, a decoder for decoding the metadata encoding is determined, including:

[0249] In some embodiments, the metadata may be first classified into metadata for training the AI ​​model and metadata for inference of the AI ​​model. Correspondingly, the decoder is also classified into a decoder corresponding to the training scenario of the AI ​​model and a decoder corresponding to the inference scenario of the AI ​​model.

[0250] In some embodiments, the decoders are hierarchical, with the metadata decoder being the highest-level decoder, and the next highest-level decoders including the decoder corresponding to the training scenario of the AI ​​model (hereinafter referred to as the model training decoder) and the decoder corresponding to the inference scenario of the AI ​​model (hereinafter referred to as the model inference decoder).

[0251] For example, the first flag includes an identifier indicating whether there is a metadata encoding decoded by the decoder corresponding to the training of the AI ​​model, such as nn_training_metadata_exist_flag. Optionally, nn_training_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the model training decoder exists in the bitstream, and nn_training_metadata_exist_flag being 0 indicates that the metadata encoding decoded by the model training decoder does not exist in the bitstream.

[0252] For example, the first flag includes an identifier indicating whether there is a metadata encoding decoded by the decoder corresponding to the inference used for the AI ​​model, such as nn_inference_metadata_exist_flag. Optionally, nn_inference_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the model inference decoder exists in the bitstream, and nn_inference_metadata_exist_flag being 0 indicates that the metadata encoding decoded by the model inference decoder does not exist in the bitstream.

[0253] Furthermore, metadata for training scenarios used in AI models can include:

[0254] Metadata of the training scenario for the AI ​​model used in the emotion recognition (ER) task;

[0255] Metadata of the training scenario for the AI ​​model used in the Automatic Speech Recognition (ASR) task;

[0256] Metadata of the training scenario for the AI ​​model used in the Automatic Speech Verification (ASV) task;

[0257] Metadata of the training scenario for the AI ​​model used in the Audio Event Classification (AEC) task.

[0258] Accordingly, the model training decoder further includes:

[0259] The decoder corresponding to the training scenario of the AI ​​model used for the emotion recognition (ER) task (hereinafter referred to as the ER training decoder);

[0260] The decoder corresponding to the training scenario of the AI ​​model for the Automatic Speech Recognition (ASR) task (hereinafter referred to as the ASR training decoder);

[0261] The decoder corresponding to the training scenario of the AI ​​model for the Automatic Speech Verification (ASV) task (hereinafter referred to as the ASV training decoder);

[0262] The decoder corresponding to the training scenario of the AI ​​model for the Audio Event Classification (AEC) task (hereinafter referred to as the AEC training decoder).

[0263] In some embodiments, the ER training decoder, ASR training decoder, ASV training decoder, and AEC training decoder belong to the next level encoder of the model training decoder.

[0264] In some embodiments, when the first information indicated by the second flag bit indicates the existence of metadata encoding, decoding the metadata encoding through the decoder corresponding to the second flag bit includes:

[0265] If the first information indicating metadata encoding exists as indicated by the second flag bit corresponding to the first decoder, the first metadata is obtained by decoding the metadata encoding through the first decoder.

[0266] The second information in the first metadata is used to determine whether the first metadata should be decoded by the second decoder.

[0267] In some embodiments, the second information includes a third flag bit corresponding to each of the second decoders, and the step of determining whether to decode the first metadata using the second information in the first metadata includes:

[0268] When the third flag bit corresponding to the second decoder indicates that metadata encoding exists, the first metadata is decoded by the second decoder to obtain the second metadata, and the second metadata is used as the decoded metadata.

[0269] If the third flag corresponding to the second decoder indicates that the metadata encoding does not exist, the first metadata is used as the decoded metadata. For example, the third flag includes an identifier indicating whether a metadata encoding decoded by the decoder corresponding to the training scenario of the AI ​​model for the emotion recognition (ER) task exists, such as er_training_metadata_exist_flag. Optionally, er_training_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the ER training decoder exists in the bitstream, and er_training_metadata_exist_flag being 0 indicates that the metadata encoding decoded by the ER training decoder does not exist in the bitstream.

[0270] For example, the third flag includes an identifier indicating whether metadata encoding decoded by the decoder exists corresponding to the training scenario of the AI ​​model for the Automatic Speech Recognition (ASR) task, such as `asr_training_metadata_exist_flag`. Optionally, `asr_training_metadata_exist_flag` being 1 indicates that the metadata encoding decoded by the ASR training decoder exists in the bitstream, and `asr_training_metadata_exist_flag` being 0 indicates that the metadata encoding decoded by the ASR training decoder does not exist in the bitstream.

[0271] For example, the third flag includes an identifier indicating whether metadata encoding decoded by the decoder exists corresponding to the training scenario of the AI ​​model for the Automatic Speech Verification (ASV) task, such as `asv_training_metadata_exist_flag`. Optionally, `asv_training_metadata_exist_flag` being 1 indicates that the metadata encoding decoded by the ASV training decoder exists in the bitstream, and `asv_training_metadata_exist_flag` being 0 indicates that the metadata encoding decoded by the ASV training decoder does not exist in the bitstream.

[0272] For example, the third flag includes an identifier indicating whether there is a metadata encoding corresponding to the decoder decoding of the training scenario of the AI ​​model for the Audio Event Classification (AEC) task, such as `aec_training_metadata_exist_flag`. Optionally, `aec_training_metadata_exist_flag` being 1 indicates that the metadata encoding decoded by the AEC training decoder exists in the bitstream, and `aec_training_metadata_exist_flag` being 0 indicates that the metadata encoding corresponding to the AEC training decoder decoding does not exist in the bitstream.

[0273] Similarly, metadata for inference scenarios used in AI models includes:

[0274] Metadata of the inference scenario for the AI ​​model used in emotion recognition (ER) tasks;

[0275] Metadata of the inference scenario for the AI ​​model used in the Automatic Speech Recognition (ASR) task;

[0276] Metadata of the inference scenario for the AI ​​model used in the automated speech verification ASV task;

[0277] Metadata of the inference scenario for the AI ​​model used in the Audio Event Classification (AEC) task.

[0278] Accordingly, the model inference decoder further includes:

[0279] The decoder corresponding to the inference scenario of the AI ​​model used for emotion recognition (ER) task (hereinafter referred to as the ER inference decoder);

[0280] The decoder corresponding to the inference scenario of the AI ​​model for the Automatic Speech Recognition (ASR) task (hereinafter referred to as the ASR inference decoder);

[0281] The decoder corresponding to the inference scenario of the AI ​​model for the Automatic Speech Verification (ASV) task (hereinafter referred to as the ASV inference decoder);

[0282] The decoder corresponding to the inference scenario of the AI ​​model for the Audio Event Classification (AEC) task (hereinafter referred to as the AEC inference decoder).

[0283] In some embodiments, the ER inference decoder, ASR inference decoder, ASV inference decoder, and AEC inference decoder are second decoders of the model inference decoder.

[0284] For example, the third flag includes an identifier, such as er_inference_metadata_exist_flag, indicating whether a metadata encoding decoded by the decoder exists corresponding to the inference scenario of the AI ​​model for the emotion recognition (ER) task. Optionally, er_inference_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the ER inference decoder exists in the bitstream, and er_inference_metadata_exist_flag being 0 indicates that the metadata encoding decoded by the ER inference decoder does not exist in the bitstream.

[0285] For example, the third flag includes an identifier indicating whether metadata encoding decoded by the decoder exists corresponding to the inference scenario of the AI ​​model for the Automatic Speech Recognition (ASR) task, such as asr_inference_metadata_exist_flag. Optionally, asr_inference_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the ASR inference decoder exists in the bitstream, and asr_inference_metadata_exist_flag being 0 indicates that the metadata encoding decoded by the ASR inference decoder does not exist in the bitstream.

[0286] For example, the third flag includes an identifier, such as asv_inference_metadata_exist_flag, indicating whether a metadata encoding decoded by the decoder exists corresponding to the inference scenario of the AI ​​model for the Automatic Speech Verification (ASV) task. Optionally, asv_inference_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the ASV inference decoder exists in the bitstream, and asv_inference_metadata_exist_flag being 0 indicates that the metadata encoding decoded by the ASV inference decoder does not exist in the bitstream.

[0287] For example, the third flag includes an identifier indicating whether there is a metadata encoding corresponding to the decoder decoding of the inference scenario of the AI ​​model for the audio event classification AEC task, such as aec_inference_metadata_exist_flag. Optionally, aec_inference_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the AEC inference decoder exists in the bitstream, and aec_inference_metadata_exist_flag being 0 indicates that the metadata encoding corresponding to the AEC inference decoder does not exist in the bitstream.

[0288] In some embodiments, the metadata can be segmented into metadata applicable to different types of AI tasks. Accordingly, the metadata decoder includes decoders for different AI tasks, for example:

[0289] The decoder corresponding to the emotion recognition (ER) task (hereinafter referred to as the ER decoder);

[0290] The decoder corresponding to the Automatic Speech Recognition (ASR) task (hereinafter referred to as the ASR decoder);

[0291] The decoder corresponding to the Automatic Voice Verification (ASV) task (hereinafter referred to as the ASV decoder);

[0292] The decoder corresponding to the Audio Event Classification (AEC) task (hereinafter referred to as the AEC decoder).

[0293] In some embodiments, the ER decoder, ASR decoder, ASV decoder, and AEC decoder are decoders at the next level below the metadata decoder.

[0294] For example, the first flag includes an identifier, such as er_metadata_exist_flag, indicating whether there is a metadata encoding decoded by the decoder corresponding to the emotion recognition (ER) task. Optionally, er_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the ER decoder exists in the bitstream, and er_metadata_exist_flag being 0 indicates that the metadata encoding decoded by the ER decoder does not exist in the bitstream.

[0295] For example, the first flag includes an identifier indicating whether metadata encoding decoded by the decoder corresponding to the Automatic Speech Recognition (ASR) task exists, such as asr_metadata_exist_flag. Optionally, asr_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the ASR decoder exists in the bitstream, and asr_metadata_exist_flag being 0 indicates that the metadata encoding decoded by the ASR decoder does not exist in the bitstream.

[0296] For example, the first flag includes an identifier, such as asv_metadata_exist_flag, indicating whether metadata encoding decoded by the decoder corresponding to the Automatic Speech Verification (ASV) task exists. Optionally, asv_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the ASV decoder exists in the bitstream, and asv_metadata_exist_flag being 0 indicates that the metadata encoding decoded by the ASV decoder does not exist in the bitstream.

[0297] For example, the first flag includes an identifier, such as aec_metadata_exist_flag, indicating whether metadata encoding decoded by the decoder corresponding to the Audio Event Classification (AEC) task exists. Optionally, aec_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the AEC decoder exists in the bitstream, and aec_metadata_exist_flag being 0 indicates that the metadata encoding decoded by the AEC decoder does not exist in the bitstream.

[0298] Furthermore, in this embodiment of the disclosure, the metadata for each type of AI task is further divided into metadata for inference scenarios and training scenarios, respectively. Therefore, the decoder corresponding to each type of AI model may further include decoders corresponding to the training and inference scenarios of the AI ​​model for the corresponding type of AI task, for example:

[0299] The ER decoder may further include two second decoders: an ER training metadata decoder and an ER inference metadata decoder;

[0300] The ASR decoder may further include two second decoders: an ASR training metadata decoder and an ASR inference metadata decoder;

[0301] The ASV decoder may further include two second decoders: an ASV training metadata decoder and an ASV inference metadata decoder.

[0302] The AEC decoder may further include two second decoders: an AEC training metadata decoder and an AEC inference metadata decoder.

[0303] For example, if the metadata encoding indicating the ER decoder exists, the third flag may also include a flag training_metadata_exist_flag indicating the existence of metadata encoding decoded by the sub-encoder corresponding to the inference scenario based on the ER task. Optionally, training_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the ER training metadata decoder exists in the bitstream. The third flag may also include a flag inference_metadata_exist_flag indicating the existence of metadata encoding decoded by the encoder corresponding to the inference scenario based on the ER task. Optionally, inference_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the ER training metadata decoder exists in the bitstream.

[0304] For example, if the metadata encoding indicating the ASR decoder exists, the third flag may also include a flag training_metadata_exist_flag indicating the existence of metadata encoding decoded by the sub-encoder corresponding to the inference scenario based on the ASR task. Optionally, training_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the ASR training metadata decoder exists in the bitstream. The third flag may also include a flag inference_metadata_exist_flag indicating the existence of metadata encoding decoded by the sub-encoder corresponding to the inference scenario based on the ASR task. Optionally, inference_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the ASR training metadata decoder exists in the bitstream.

[0305] For example, if the metadata encoding indicating the ASV decoder exists, the third flag may also include a flag training_metadata_exist_flag indicating the existence of metadata encoding decoded by the sub-encoder corresponding to the inference scenario based on the ASV task. Optionally, training_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the ASV training metadata decoder exists in the bitstream. The third flag may also include a flag inference_metadata_exist_flag indicating the existence of metadata encoding decoded by the sub-encoder corresponding to the inference scenario based on the ASV task. Optionally, inference_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the ASV training metadata decoder exists in the bitstream.

[0306] For example, if the metadata encoding indicating that the AEC decoder is present exists, the third flag may also include a flag training_metadata_exist_flag indicating the existence of metadata encoding decoded by the sub-encoder corresponding to the inference scenario based on the AEC task. Optionally, training_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the AEC training metadata decoder exists in the bitstream. The third flag may also include a flag inference_metadata_exist_flag indicating the existence of metadata encoding decoded by the sub-encoder corresponding to the inference scenario based on the AEC task. Optionally, inference_metadata_exist_flag being 1 indicates that the metadata encoding decoded by the AEC training metadata decoder exists in the bitstream.

[0307] In some embodiments, when the third flag bit corresponding to the second decoder indicates that metadata encoding exists, the second metadata is obtained by decoding the first metadata through the second decoder, and the second metadata is used as the decoded metadata.

[0308] In some embodiments, if the third flag corresponding to the second decoder indicates that the metadata encoding does not exist, the first metadata is used as the decoded metadata.

[0309] In some embodiments, an AI task can be performed by multiple AI models, and the encoder can predetermine whether metadata will be applied to all AI models. For example, if a speech recognition ASR training task is performed by AI model 1 and AI model 2, AI model 1 will use metadata 1 for the training scenario, while AI model 2 will not use metadata 1 for the training scenario. Therefore, adding a fourth flag related to metadata 1 can improve storage and transmission efficiency. For example, the fourth flag may include asr_training_metadata1_exist_flag and asr_training_metadata2_exist_flag. It should be understood that asr_training_metadata1_exist_flag being 1 indicates the existence of metadata 1 for model 1 in the ASR training task, and asr_training_metadata2_exist_flag being 1 indicates the existence of metadata 2 for model 2 in the ASR training task.

[0310] For example, if the speech recognition ASR inference task is performed by AI Model 1 and AI Model 2, AI Model 1 will use metadata 1 for the inference scenario, while AI Model 2 will not. Therefore, adding a flag for metadata 1 can improve storage and transmission efficiency. For example, the first flag could include asr_inference_metadata1_exist_flag and asr_inference_metadata2_exist_flag. It should be understood that asr_inference_metadata1_exist_flag being 1 indicates the existence of metadata 1 for Model 1 in the ASR inference task, and asr_inference_metadata2_exist_flag being 1 indicates the existence of metadata 2 for Model 2 in the ASR inference task.

[0311] For example, if the speech recognition inference (ASR) task is performed by AI model 1 and AI model 2, and both model 1 and model 2 use metadata 3 for training and inference scenarios, then the fourth flag bit related to metadata 3 can be omitted to save one bit.

[0312] In some embodiments, when the first metadata includes second metadata with different quantization bits, the third flag bit corresponding to the second decoder indicates that metadata encoding exists;

[0313] The step of obtaining decoded metadata by decoding the first metadata through the second decoder includes:

[0314] When the third flag bit corresponding to the second decoder indicates that metadata encoding exists, the second decoder decodes the first metadata to obtain the second metadata with the corresponding number of quantization bits.

[0315] In this embodiment of the disclosure, there is a situation where the same metadata exists in multiple AI tasks, but with different quantization precisions. Therefore, different quantization bits are used to quantize the metadata encoding. For example, metadata A is a 16-bit integer in the decoder of the corresponding ASR task, while metadata A is a 24-bit integer in the decoder of the corresponding ASV task.

[0316] In some embodiments, the same metadata is sometimes applied to multiple AI tasks, and the values ​​of these metadata are also the same. In this case, a common metadata field and a flag can be added to further reduce the redundancy of metadata encoding. Specifically, at least one of the bitstream and the second information also includes a fifth flag, which is used to indicate whether there is metadata encoding with the same number of quantization bits and the same value used for multiple AI tasks.

[0317] For example, the fifth flag could be Common_metadata_flag. Optionally, Common_metadata1_exist_flag being 1 indicates the existence of metadata codes with the same quantization bit depth and the same value, which are used for multiple AI tasks; Common_metadata_exist_flag being 0 indicates the absence of metadata codes with the same quantization bit depth and the same value, which are used for multiple AI tasks.

[0318] The decoder can determine that there is metadata encoding with the same quantization bit depth and the same value multiplexed by multiple decoders based on Common_metadata_flag being 1. At the same time, it can determine which decoders to decode the metadata encoding based on the first or third flag indicating the existence of metadata encoding. For example, if the third flags asr_training_metadata_exist_flag, asv_training_metadata_exist_flag, er_training_metadata_exist_flag, and aec_training_metadata_exist_flag are 1, then it is determined that the metadata encoding will be decoded by the decoder corresponding to the AI ​​model training for ASR, the AI ​​model training for ASV, the AI ​​model training for ER, and the AI ​​model training for AEC, respectively.

[0319] It should be noted that the embodiments disclosed herein involve audio features and metadata before encoding at the encoding end, as well as audio features and metadata obtained by decoding at the decoding end. Encoding audio features and metadata at the encoding end may cause data loss, and the audio features and metadata obtained by decoding at the decoding end may also cause data loss. Therefore, the audio features and metadata before encoding at the encoding end may not be completely the same as the audio features and metadata obtained by decoding at the decoding end.

[0320] Alternatively, in this embodiment of the invention, both the encoding end and the decoding end are lossless processes during encoding and decoding. Therefore, the audio features and metadata before encoding at the encoding end are the same as the audio features and metadata obtained by decoding at the decoding end, and the metadata before encoding at the encoding end is the same as the metadata obtained by decoding at the decoding end.

[0321] The audio features and metadata obtained by decoding according to the embodiments of this disclosure can be used for training or inference of machine hearing tasks.

[0322] The signal indication method involved in the embodiments of this disclosure may include at least one of steps S2101 to S2105. For example, steps S2104 to S2105 may be implemented as independent embodiments, but are not limited thereto.

[0323] In some embodiments, step S2101 is optional, and one or more of these steps may be omitted or substituted in different embodiments.

[0324] In some embodiments, step S2102 is optional, and one or more of these steps may be omitted or substituted in different embodiments.

[0325] In some embodiments, step S2103 is optional, and one or more of these steps may be omitted or substituted in different embodiments.

[0326] In some embodiments, step S2104 is optional, and one or more of these steps may be omitted or substituted in different embodiments.

[0327] In some embodiments, step S2105 is optional, and one or more of these steps may be omitted or substituted in different embodiments.

[0328] In some embodiments, other optional implementations described before or after the specification corresponding to FIG2 may be referred to.

[0329] In some embodiments, the names of information, etc., are not limited to the names described in the embodiments. Terms such as "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "bit", "data", "program", and "chip" can be used interchangeably.

[0330] In some embodiments, terms such as “moment,” “point in time,” “time,” and “time location” can be used interchangeably, as can terms such as “duration,” “segment,” “time window,” “window,” and “time.”

[0331] In some embodiments, terms such as wireless access scheme and waveform can be used interchangeably.

[0332] In some embodiments, terms such as "certain," "preset," "default," "set," "indicated," "a certain," "any," and "first" can be used interchangeably. "Certain A," "preset A," "default A," "set A," "indicated A," "a certain A," "any A," and "first A" can be interpreted as A pre-defined in a protocol or the like, or as A obtained through setting, configuration, or instruction, or as specific A, a certain A, any A, or first A, but are not limited thereto.

[0333] In some embodiments, the determination or judgment can be made by a value represented by 1 bit (0 or 1), or by a true or false value (boolean), or by a comparison of numerical values ​​(e.g., a comparison with a predetermined value), but is not limited thereto.

[0334] In some embodiments, "not expecting to receive" can be interpreted as not receiving on time domain resources and / or frequency domain resources, or as not performing subsequent processing on the data after receiving it; "not expecting to send" can be interpreted as not sending, or as sending but not expecting the receiver to respond to the sent content.

[0335] Figure 3A is a flowchart illustrating a signal indication method according to an embodiment of the present disclosure, applied to an encoding end. As shown in Figure 3A, the present disclosure relates to a signal indication method, which includes:

[0336] Step S3101: The encoding end obtains the audio features of at least one audio frame.

[0337] The optional implementation of step S3101 can be found in step S2101 of Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.

[0338] Step S3102: The encoding end performs encoding to obtain a bit stream.

[0339] The optional implementation of step S3102 can be found in step S2102 of Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.

[0340] Step S3103: The encoding end sends the bit stream.

[0341] The optional implementation of step S3103 can be found in step S2103 of Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.

[0342] The feature indication method involved in the embodiments of this disclosure may include at least one of steps S3101 to S3103. For example, step S3101 may be implemented as a standalone embodiment, step S3102 may be implemented as a standalone embodiment, step S3103 may be implemented as a standalone embodiment, or at least two steps may be combined, but it is not limited thereto.

[0343] In some embodiments, step S3101 is optional, step S3102 is optional, and step S3103 is optional. In different embodiments, one or more of these steps may be omitted or substituted. However, this is not a limitation.

[0344] Figure 3B is a flowchart illustrating a signal indication method according to an embodiment of the present disclosure, applied at the encoding end. As shown in Figure 3B, this disclosure relates to a signal indication method, which includes:

[0345] Step S3201: The encoding end sends the bit stream.

[0346] The optional implementation of step S3201 can be found in step S2103 of Figure 2, step S3103 of Figure 3A, and other related parts in the embodiments involved in Figures 2 and 3A, which will not be repeated here.

[0347] Figure 4A is a flowchart illustrating a signal indication method according to an embodiment of the present disclosure, applied to a decoding end. As shown in Figure 4A, the present disclosure relates to a signal indication method, which includes:

[0348] Step S4101: The decoding end receives the bit stream.

[0349] The optional implementation of step S4101 can be found in step S2104 of Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.

[0350] Step S4102: The decoding end decodes the bit stream.

[0351] The optional implementation of step S4102 can be found in step S2105 of Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.

[0352] Figure 4B is a flowchart illustrating a signal indication method according to an embodiment of the present disclosure, applied to a decoding end. As shown in Figure 4B, the present disclosure relates to a signal indication method, which includes:

[0353] Step S4201: The decoding end receives the bit stream.

[0354] The optional implementation of step S4201 can be found in step S2104 of Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.

[0355] Step S4202: Determine whether the bit stream includes metadata encoding.

[0356] The optional implementation of step S4202 can be found in step S2105 of Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.

[0357] Figure 5 is a flowchart illustrating a signal indication method according to an embodiment of the present disclosure. As shown in Figure 5, the present disclosure relates to a signal indication method, which includes:

[0358] Step S5101: The encoding end sends the bit stream.

[0359] In some embodiments, the bitstream includes first information, which is used to determine whether the bitstream includes metadata encoding.

[0360] In some embodiments, if the first information indicates that metadata encoding exists, the bitstream includes metadata encoding; if the first information indicates that metadata encoding does not exist, the bitstream does not include metadata encoding.

[0361] In some embodiments, the bitstream includes at least one first flag bit, each first flag bit corresponds to a decoder, and each first flag bit includes first information;

[0362] When the first information indicated by the first flag bit indicates that the metadata encoding exists, the metadata encoding is decoded by the decoder corresponding to the first flag bit. The metadata encoding decoded by different decoders is applied to different AI tasks.

[0363] In some embodiments, when the value of the metadata encoding changes, the bitstream includes a first flag bit corresponding to the decoder that decodes the metadata encoding;

[0364] If the value of the metadata encoding remains unchanged, the bitstream does not include the first flag bit corresponding to the decoder that decodes the metadata encoding;

[0365] If the transmission period of the metadata encoding is met at the current moment and the value of the metadata encoding has not changed, the bit stream includes a first flag bit corresponding to the decoder that decodes the metadata encoding. Step S5102: The decoding end receives the bit stream.

[0366] Step S5103: The decoding end determines whether the bit stream includes metadata encoding.

[0367] In some embodiments, the bitstream includes first information, which is used to determine whether the bitstream includes metadata encoding.

[0368] In some embodiments, if the first information indicates that metadata encoding exists, the bitstream includes metadata encoding; if the first information indicates that metadata encoding does not exist, the bitstream does not include metadata encoding.

[0369] In some embodiments, the bitstream includes at least one first flag bit, each first flag bit corresponds to a decoder, and each first flag bit includes first information;

[0370] When the first information indicated by the first flag bit indicates that the metadata encoding exists, the metadata encoding is decoded by the decoder corresponding to the first flag bit. The metadata encoding decoded by different decoders is applied to different AI tasks.

[0371] In some embodiments, the decoding end includes at least one decoder, the bit stream includes at least one first flag bit, each first flag bit corresponds to one decoder, and each first flag bit includes the first information;

[0372] When the first information indicated by the second flag bit in the at least one first flag bit is present, the metadata encoding is decoded by the decoder corresponding to the second flag bit, and the metadata encoding decoded by different decoders is applied to different AI tasks.

[0373] In some embodiments, the decoding end includes at least one first decoder, and the first decoder includes at least one second decoder;

[0374] When the first information indicated by the second flag bit indicates the existence of metadata encoding, decoding the metadata encoding through the decoder corresponding to the second flag bit includes:

[0375] If the first information indicating metadata encoding exists as indicated by the second flag bit corresponding to the first decoder, the first metadata is obtained by decoding the metadata encoding through the first decoder.

[0376] The second information in the first metadata is used to determine whether the first metadata should be decoded by the second decoder.

[0377] In some embodiments, the second information includes a third flag bit corresponding to each of the second decoders, and the step of determining whether to decode the first metadata using the second information in the first metadata includes:

[0378] When the third flag bit corresponding to the second decoder indicates that metadata encoding exists, the first metadata is decoded by the second decoder to obtain the second metadata, and the second metadata is used as the decoded metadata.

[0379] If the third flag corresponding to the second decoder indicates that the metadata encoding does not exist, the first metadata is used as the decoded metadata.

[0380] In some embodiments, when the second metadata is applied to a portion of the AI ​​models of the AI ​​task corresponding to the second decoder, the second metadata includes a fourth flag bit, with different fourth flag bits corresponding to different AI models, and the fourth flag bit is used to indicate whether there is metadata applied to the corresponding AI model; when the second metadata is applied to all AI models of the AI ​​task corresponding to the second decoder, the second metadata does not include the fourth flag bit.

[0381] In some embodiments, the metadata encoding decoded by different first decoders is applied to different types of AI tasks;

[0382] Metadata encoding decoded by different second decoders is applied to AI tasks in different application scenarios;

[0383] or

[0384] Metadata encoding decoded by different first decoders is applied to AI tasks in different application scenarios;

[0385] Different second decoders decode metadata encodings for different types of AI tasks

[0386] In some embodiments, the type of AI task includes at least one of the following:

[0387] Emotion recognition (ER) task;

[0388] Automatic Speech Recognition (ASR) task;

[0389] Automatic voice verification of ASV tasks;

[0390] Audio event classification AEC task.

[0391] In some embodiments, the application scenarios of the AI ​​task include at least one of the following:

[0392] Inference scenarios for AI models;

[0393] Training scenarios for AI models.

[0394] In some embodiments, when the first metadata includes second metadata with different quantization bits, the third flag bit corresponding to the second decoder indicates that metadata encoding exists;

[0395] The step of obtaining decoded metadata by decoding the first metadata through the second decoder includes:

[0396] When the third flag bit corresponding to the second decoder indicates that metadata encoding exists, the second decoder decodes the first metadata to obtain the second metadata with the corresponding number of quantization bits.

[0397] In some embodiments, at least one of the bitstream and the second information further includes a fifth flag bit, which indicates whether there is a metadata encoding with the same quantization bit length and the same value, wherein the metadata encoding with the same quantization bit length and the same value is used for multiple AI tasks. In some embodiments, the above method may include the methods of the embodiments described above on the communication system side, encoding end side, decoding end side, etc., which will not be repeated here.

[0398] In some embodiments, the steps and their optional implementations in other embodiments described before or after this embodiment, as well as other related parts in the specification, can be referred to, and will not be repeated here.

[0399] This disclosure also proposes an apparatus (also referred to as a communication device, etc.) for implementing any of the above methods. For example, an apparatus is proposed that includes units or modules for implementing the steps performed by the terminal in any of the above methods. Furthermore, another apparatus is proposed that includes units or modules for implementing the steps performed by a network device (e.g., an access network device, a core network functional node, a core network device, etc.) in any of the above methods.

[0400] It should be understood that the division of units or modules in the above device is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the units or modules in the device can be implemented by a processor calling software: for example, the device includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of the units or modules in the above device. The processor can be, for example, a general-purpose processor, such as a Central Processing Unit (CPU) or a microprocessor, and the memory can be internal or external to the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits. The functionality of some or all of the units or modules can be achieved through the design of these hardware circuits, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC). The functionality of some or all of the units or modules is achieved through the design of the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a programmable logic device (PLD). Taking a field-programmable gate array (FPGA) as an example, it can include a large number of logic gates. The connection relationships between the logic gates are configured through a configuration file, thereby achieving the functionality of some or all of the units or modules. All units or modules of the above device can be implemented entirely through processor-called software, entirely through hardware circuits, or partially through processor-called software with the remaining parts implemented through hardware circuits.

[0401] In this embodiment, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction read and execute capabilities, such as a Central Processing Unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. The logical relationships of the aforementioned hardware circuits are fixed or reconfigurable. For example, the processor is a hardware circuit implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a Neural Network Processing Unit (NPU), a Tensor Processing Unit (TPU), or a Deep Learning Processing Unit (DPU).

[0402] Figure 6A is a schematic diagram of the decoding device proposed in an embodiment of this disclosure. The decoding device is used to perform the signal indication method executed by any of the decoding terminals described above. In some embodiments, as shown in Figure 6A, the decoding device may include at least one of a transceiver module 6001 and a processing module 6002.

[0403] In some embodiments, the transceiver module 6001 is used to receive bit streams.

[0404] In some embodiments, the bitstream includes first information of at least one audio frame, the first information including metadata and audio feature encoding; wherein, for any audio frame, the metadata of the audio frame is used to reproduce the audio features of the audio frame at the decoding end.

[0405] In some embodiments, the processing module 6002 is used to determine whether the bitstream includes metadata encoding. In some embodiments, the processing module can be interchanged with the determining module or the processor, and the transceiver module can be interchanged with the sending module or the transceiver.

[0406] Figure 6B is a schematic diagram of the structure of the encoding device proposed in an embodiment of this disclosure. The encoding device is used to perform the signal indication method executed by the encoding end described above. In some embodiments, as shown in Figure 6B, the encoding device may include a transceiver module 6101.

[0407] In some embodiments, the transceiver module 6101 described above is used to send a bit stream.

[0408] Optionally, the transceiver module 6101 is used to execute at least one of the transceiver steps (such as step 2103, step 3103, but not limited thereto) executed by the encoding end in any of the above methods, which will not be described in detail here.

[0409] The above-mentioned encoding device may include a processing module, which is used to execute at least one of the communication steps (such as steps S2101, S2102, S3101, but not limited thereto) executed by the encoding end in any of the above methods, which will not be described in detail here.

[0410] In some embodiments, the processing module can be interchanged with the determining module and the processor, and the transceiver module can be interchanged with the sending module and the transceiver.

[0411] Figure 7 is a schematic diagram of the structure of the communication device 7000 proposed in an embodiment of this disclosure. The communication device 7000 can be a network device (e.g., access network device, core network device, etc.), a terminal (e.g., user equipment, etc.), a chip, chip system, or processor that supports the network device in implementing any of the above methods, or a chip, chip system, or processor that supports the terminal in implementing any of the above methods. The communication device 7000 can be used to implement the methods described in the above method embodiments; for details, please refer to the descriptions in the above method embodiments.

[0412] As shown in Figure 7, the communication device 7000 is used to execute any of the above methods. In some embodiments, the communication device 7000 includes one or more processors 7001. The processor 7001 may be a general-purpose processor or a special-purpose processor, such as a baseband processor or a central processing unit. The baseband processor may be used to process communication protocols and communication data, and the central processing unit may be used to control communication devices (e.g., base stations, baseband chips, terminal devices, terminal device chips, DUs or CUs, etc.), execute programs, and process program data. Optionally, the communication device 7000 is used to execute any of the above methods. Optionally, one or more processors 7001 are used to invoke instructions to cause the communication device 7000 to execute any of the above methods.

[0413] In some embodiments, the communication device 7000 further includes one or more transceivers 7002. When the communication device 7000 includes one or more transceivers 7002, the transceiver 7002 performs at least one of the communication steps such as sending and / or receiving in the above method (e.g., steps S2103, S3103, S3201, S4101, but not limited thereto), and the processor 7001 performs at least one of other steps (e.g., steps S2101, S2102, S3202, S4102, but not limited thereto). In optional embodiments, the transceiver may include a receiver and / or a transmitter, which may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, transceiver circuit, interface circuit, interface, etc., can be used interchangeably; the terms transmitter, transmitting unit, transmitter, transmitting circuit, etc., can be used interchangeably; the terms receiver, receiving unit, receiver, receiving circuit, etc., can be used interchangeably.

[0414] In some embodiments, the communication device 7000 further includes one or more memories 7003 for storing data and / or instructions. Optionally, one or more processors 7001 are used to invoke instructions stored in the memory 7003 to cause the communication device 7000 to perform any of the above methods. Optionally, all or part of the memory 7003 may also be located outside the communication device 7000. In an optional embodiment, the communication device 7000 may include one or more interface circuits 7004. Optionally, the interface circuit 7004 is connected to the memory 7002, and the interface circuit 7004 can be used to receive data and / or instructions from the memory 7002 or other devices, and can be used to send data and / or instructions to the memory 7002 or other devices. For example, the interface circuit 7004 can read data and / or instructions stored in the memory 7002 and send the data and / or instructions to the processor 7001.

[0415] The communication device 7000 described in the above embodiments may be a network device or a terminal, but the scope of the communication device 7000 described in this disclosure is not limited thereto, and the structure of the communication device 7000 may not be limited by FIG. 7. The communication device may be a standalone device or a part of a larger device. For example, the communication device may be: (1) a standalone integrated circuit IC, or chip, or chip system or subsystem; (2) a collection of one or more ICs, optionally, the IC collection may also include storage components for storing data, programs and / or instructions; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, terminal device, smart terminal device, cellular phone, wireless device, handheld device, mobile unit, vehicle device, network device, cloud device, artificial intelligence device, etc.; (6) others, etc.

[0416] Figure 8 is a schematic diagram of the structure of the chip 8000 proposed in an embodiment of this disclosure. For cases where the communication device 7000 can be a chip or a chip system, the schematic diagram of the chip 8000 shown in Figure 8 can be referenced, but is not limited thereto.

[0417] Chip 8000 includes one or more processors 8001. Chip 8000 is used to perform any of the above methods.

[0418] In some embodiments, chip 8000 further includes one or more interface circuits 8002. Optionally, terms such as interface circuit, interface, and transceiver pin can be used interchangeably. In some embodiments, chip 8000 further includes one or more memories 8003 for storing data and / or instructions. Optionally, all or part of the memories 8003 may be located outside of chip 8000. Optionally, interface circuit 8002 is connected to memory 8003, and interface circuit 8002 can be used to receive data and / or instructions from memory 8003 or other devices, and interface circuit 8002 can be used to send data and / or instructions to memory 8003 or other devices. For example, interface circuit 8002 can read data and / or instructions stored in memory 8003 and send the data and / or instructions to processor 8001.

[0419] In some embodiments, the interface circuit 8002 performs at least one of the communication steps such as sending and / or receiving in the above-described method (e.g., steps S2103, S3103, S3201, and S4101, but not limited thereto). The interface circuit 8002 performing the communication steps such as sending and / or receiving in the above-described method refers, for example, to the interface circuit 8002 performing data and / or instruction interaction between the processor 8001, the chip 8000, the memory 8003, or the transceiver device. In some embodiments, the processor 8001 performs at least one of other steps (e.g., steps S2101, S2102, S3202, and S4102, but not limited thereto).

[0420] The modules and / or devices described in the various embodiments, such as virtual devices, physical devices, and chips, can be combined or separated arbitrarily as needed. Optionally, some or all steps can also be performed collaboratively by multiple modules and / or devices, which is not limited here.

[0421] This disclosure also proposes a storage medium storing instructions that, when executed on a communication device, cause the communication device to perform any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but not limited thereto; it may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but not limited thereto; it may also be a temporary storage medium.

[0422] This disclosure also proposes a program product, including a program and / or instructions, which, when executed by a communication device, cause the communication device to perform any of the above methods. Optionally, the program product is a computer program product. Optionally, the program product is stored on the storage medium.

[0423] This disclosure also proposes a computer program that, when run on a computer, causes the computer to perform any of the above methods.

Claims

1. An information indication method, characterized in that, The method is executed by the decoding end, and the method includes: Receive bit stream; Determine whether the bitstream includes metadata encoding; The metadata encoding is used to determine the decoded metadata, which is then used for AI tasks.

2. The method according to claim 1, characterized in that, The determination of whether the decoded bitstream includes metadata encoding includes: Based on the first information in the bitstream, determine whether the bitstream includes metadata encoding.

3. The method according to claim 3, characterized in that, The first information includes a first value, which indicates the existence of the metadata encoding; or, The first information includes a second value, which indicates that the metadata encoding does not exist.

4. The method according to claim 2 or 3, characterized in that, The decoding end includes at least one decoder, and the bit stream includes at least one first flag bit, each first flag bit corresponding to one decoder, and each first flag bit including the first information; When the first information indicated by the second flag bit in the at least one first flag bit is present, the metadata encoding is decoded by the decoder corresponding to the second flag bit, and the metadata encoding decoded by different decoders is applied to different AI tasks.

5. The method according to claim 4, characterized in that, The decoding end includes at least one first decoder, and the first decoder includes at least one second decoder; When the first information indicated by the second flag bit indicates the existence of metadata encoding, decoding the metadata encoding through the decoder corresponding to the second flag bit includes: If the first information indicating metadata encoding exists as indicated by the second flag bit corresponding to the first decoder, the first metadata is obtained by decoding the metadata encoding through the first decoder. The second information in the first metadata is used to determine whether the first metadata should be decoded by the second decoder.

6. The method according to claim 5, characterized in that, The second information includes a third flag bit corresponding to each of the second decoders. The step of determining whether to decode the first metadata using the second information in the first metadata includes: When the third flag bit corresponding to the second decoder indicates that metadata encoding exists, the first metadata is decoded by the second decoder to obtain the second metadata, and the second metadata is used as the decoded metadata. If the third flag corresponding to the second decoder indicates that the metadata encoding does not exist, the first metadata is used as the decoded metadata.

7. The method according to claim 6, characterized in that, When the second metadata is applied to a portion of the AI ​​model of the AI ​​task corresponding to the second decoder, the second metadata includes a fourth flag bit, different fourth flag bits correspond to different AI models, and the fourth flag bit is used to indicate whether there is metadata applied to the corresponding AI model; When the second metadata is applied to all AI models of the AI ​​task corresponding to the second decoder, the second metadata does not include the fourth flag bit.

8. The method according to any one of claims 5-7, characterized in that, The metadata encoding decoded by different first decoders is applied to different types of AI tasks; Metadata encoding decoded by different second decoders is applied to AI tasks in different application scenarios; or Metadata encoding decoded by different first decoders is applied to AI tasks in different application scenarios; The metadata encoding decoded by different second decoders is applied to different types of AI tasks.

9. The method according to claim 8, characterized in that, The types of AI tasks include at least one of the following: Emotion recognition (ER) task; Automatic Speech Recognition (ASR) task; Automatic voice verification of ASV tasks; Audio event classification AEC task.

10. The method according to claim 8 or 9, characterized in that, The application scenarios for the AI ​​task include at least one of the following: Inference scenarios for AI models; Training scenarios for AI models.

11. The method according to claim 6, characterized in that, When the first metadata includes second metadata with different quantization bits, the third flag bit corresponding to the second decoder indicates that metadata encoding exists; The step of obtaining decoded metadata by decoding the first metadata through the second decoder includes: When the third flag bit corresponding to the second decoder indicates that metadata encoding exists, the second decoder decodes the first metadata to obtain the second metadata with the corresponding quantization bits.

12. The method according to claim 6, characterized in that, At least one of the bitstream and the second information further includes a fifth flag bit, which is used to indicate whether there is a metadata encoding with the same quantization bit length and the same value, the metadata encoding with the same quantization bit length and the same value is used for multiple AI tasks.

13. An information indication method, characterized in that, The method is executed by the encoding end, and the method includes: Send a bit stream, which is used to determine whether it includes metadata encoding; The metadata encoding is used to determine the decoded metadata, which is then used for AI tasks.

14. The method according to claim 13, characterized in that, The bitstream includes first information, which is used to determine whether the bitstream includes metadata encoding.

15. The method according to claim 14, characterized in that, If the first information indicates the presence of metadata encoding, the bitstream includes the metadata encoding; If the first information indicates that there is no metadata encoding, the bitstream does not include the metadata encoding.

16. The method according to claim 14 or 15, characterized in that, The bitstream includes at least one first flag bit, each first flag bit corresponds to a decoder, and each first flag bit includes the first information; When the first information indicated by the first flag bit indicates the existence of metadata encoding, the metadata encoding is decoded by the decoder corresponding to the first flag bit, and the metadata encoding decoded by different decoders is applied to different AI tasks.

17. The method according to claim 16, characterized in that, If the value of the metadata encoding changes, the bitstream includes a first flag bit corresponding to the decoder that decodes the metadata encoding; If the value of the metadata encoding remains unchanged, the bitstream does not include the first flag bit corresponding to the decoder that decodes the metadata encoding; If the transmission period of the metadata encoding is met at the current moment and the value of the metadata encoding has not changed, the bit stream includes a first flag bit corresponding to the decoder that decodes the metadata encoding.

18. An information indication method, characterized in that, include: The encoding end sends the bit stream; The decoding end receives the bit stream; The decoding end determines whether the decoded bitstream includes metadata encoding; The metadata encoding is used to determine the decoded metadata, which is then used for AI tasks.

19. A decoding device, characterized in that, The decoding device includes: The transceiver module is used to receive bit streams; The processing module is used to determine whether there is metadata encoding in the decoded bitstream. The metadata encoding is used to determine the decoded metadata, and the decoded metadata is used for AI tasks.

20. An encoding device, characterized in that, The encoding device includes: The transceiver module is used to send bit streams; The bitstream is used to determine whether it includes metadata encoding; The metadata encoding is used to determine the decoded metadata, which is then used for AI tasks.

21. A decoding terminal, characterized in that, The decoding end includes: One or more processors; The processor is used to execute the information indication method according to any one of claims 1 to 12.

22. An encoding terminal, characterized in that, The encoding end includes: One or more processors; The processor is used to execute the information indication method according to any one of claims 13 to 17.

23. A communication system, characterized in that, It includes an encoding end and a decoding end, wherein the decoding end is configured to implement the feature indication method according to any one of claims 1 to 12, and the encoding end is configured to implement the information indication method according to any one of claims 13 to 17.

24. A storage medium storing instructions, characterized in that, When the instruction is executed on the communication device, the communication device performs the feature indication method as described in any one of claims 1 to 12, or the information indication method as described in any one of claims 13 to 17.

25. A computer program product, characterized in that, When the computer program product is run on a communication device, it causes the communication device to perform the information indication method as described in any one of claims 1 to 17.