Audio signal processing method and apparatus and storage medium

By collaboratively selecting the appropriate processing mode for the first and second devices during audio signal processing, the latency and complexity issues of the audio stream signal on the extended reality device side are resolved, achieving balanced computing power and an immersive audio experience between devices.

WO2025199759A1PCT designated stage Publication Date: 2025-10-02BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/083889
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-26
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

In the existing technology, the audio reception, rendering and playback process on the extended reality device side has problems with large latency and high complexity, especially in the processing and distribution of audio bitstream signals, which makes it impossible to achieve real-time binaural rendering and immersive audio experience.

Method used

The audio signal is collaboratively processed by the first device and the second device, a target processing mode is selected from multiple processing modes according to the decision indication information, and the processing tasks of the audio signal are reasonably distributed to ensure balanced sharing of computing power between devices, including operations such as bitstream decoding, rendering processing and encoding.

Benefits of technology

It optimizes the audio experience service requirements on the extended reality device side, reduces latency, improves processing efficiency, and ensures real-time rendering of audio signals and immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024083889_02102025_PF_FP_ABST
    Figure CN2024083889_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides an audio signal processing method and apparatus, a device and a storage medium. The method is executed by a first device, and comprises: determining a target processing mode from among a plurality of processing modes on the basis of decision indication information; processing an audio signal in the target processing mode to obtain a target audio signal; and sending the target audio signal to a second device, wherein the plurality of processing modes are different on the basis of processing functions completed by the first device as a carrier. The optimal working mode meeting audio experience service requirements can be selected on the basis of the decision indication information.
Need to check novelty before this filing date? Find Prior Art

Description

Audio signal processing method, device and storage medium Technical Field

[0001] The present disclosure relates to the field of communication technologies, and in particular to an audio signal processing method, device, and storage medium. Background Art

[0002] In order to complete the reception, rendering and playback of audio on terminal devices, especially extended reality (XR) devices, it is necessary to determine a specific working mode. In other words, how to reasonably distribute the entire decoding and rendering process of the received audio code stream signal from decoding to binaural rendering to each processing unit. The working mode proposed by related technologies has the problems of large latency and high complexity.

[0003] Summary of the Invention

[0004] The present disclosure provides an audio signal processing method, apparatus, communication equipment, communication system, and storage medium.

[0005] According to a first aspect of an embodiment of the present disclosure, an audio signal processing method is proposed, which is executed by a first device. The method includes: determining a target processing mode from multiple processing modes based on decision indication information; processing the audio signal in the target processing mode to obtain a target audio signal; and sending the target audio signal to a second device. The multiple processing modes are different processing functions performed based on the first device as a carrier.

[0006] In the above method, the first device can determine the target processing mode of the audio signal based on the first information, and can determine the appropriate processing mode of the audio signal based on the indication information. By processing the audio signal and sending it to the second device, the optimal working mode that meets the audio experience service requirements can be achieved.

[0007] According to a second aspect of an embodiment of the present disclosure, a method for processing an audio signal is proposed, which is executed by a second device. The method includes: receiving a target audio signal sent by a first device, where the target audio signal is obtained by the first device processing the audio signal in a target processing mode, where the target processing mode is determined from multiple processing modes based on decision indication information; processing the target audio signal in the target processing mode; the multiple processing modes are different processing functions performed based on the first device as a carrier.

[0008] In the above method, the second device can receive the target audio signal and process the target audio signal, thereby achieving an optimal working mode that meets the audio experience service requirements.

[0009] According to a third aspect of an embodiment of the present disclosure, a first device is provided, comprising a processing module configured to determine a target processing mode from a plurality of processing modes based on decision indication information; the processing module further configured to process an audio signal in the target processing mode to obtain a target audio signal; and a transceiver module configured to transmit the target audio signal to a second device. The plurality of processing modes are based on different processing functions performed by the first device as a carrier.

[0010] According to the fourth aspect of an embodiment of the present disclosure, a second device is proposed, including a transceiver module for receiving a target audio signal sent by a first device, where the target audio signal is obtained by the first device processing the audio signal in a target processing mode, where the target processing mode is determined from multiple processing modes based on decision indication information; a processing module for processing the target audio signal in the target processing mode; the multiple processing modes are different processing functions performed based on the first device as a carrier.

[0011] According to the fifth aspect of the embodiments of the present disclosure, a communication device is proposed, which includes: a transceiver; a memory; and a processor, which is connected to the transceiver and the memory respectively, and is configured to control the wireless signal reception and transmission of the transceiver by executing computer-executable instructions on the memory, and is capable of executing the method of any embodiment of the first aspect or the second aspect.

[0012] According to a sixth aspect of an embodiment of the present disclosure, a storage medium is proposed, which stores instructions. When the instructions are executed on a communication device, the communication device executes a method as in any one of the first and second aspects.

[0013] According to the seventh aspect of the embodiments of the present disclosure, a communication system is proposed, including a first device and a second device, wherein the first device is configured to implement the method of the embodiment of the first aspect, and the second device is configured to implement the method of the embodiment of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The above and / or additional aspects and advantages of the present disclosure will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0015] FIG1 is a schematic diagram of the architecture of a communication system provided by an embodiment of the present disclosure;

[0016] FIG2 is an interactive schematic diagram of an audio signal processing method provided by an embodiment of the present disclosure;

[0017] FIG3 is a flowchart of some audio signal processing methods provided by embodiments of the present disclosure;

[0018] FIG4 is a flowchart of other audio signal processing methods provided by embodiments of the present disclosure;

[0019] FIG5 is an interactive diagram of an audio signal processing method provided by an embodiment of the present disclosure;

[0020] FIG6 is a schematic diagram of a decision process provided by an embodiment of the present disclosure;

[0021] FIG7 is a flow chart of a method for determining a working mode according to an embodiment of the present disclosure;

[0022] FIG8a is a schematic structural diagram of a first device provided by an embodiment of the present disclosure;

[0023] FIG8b is a schematic structural diagram of a second device provided by an embodiment of the present disclosure;

[0024] FIG9a is a schematic structural diagram of a communication device provided by an embodiment of the present disclosure;

[0025] FIG9 b is a schematic structural diagram of a chip provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0026] The embodiments of the present disclosure provide an audio signal processing method and apparatus, a communication device, a communication system, and a storage medium.

[0027] In a first aspect, an embodiment of the present disclosure proposes an audio signal processing method, which is executed by a first device, and the method includes: determining a target processing mode from multiple processing modes based on decision indication information; processing the audio signal in the target processing mode to obtain a target audio signal; and sending the target audio signal to a second device; the multiple processing modes are different processing functions performed based on the first device as a carrier.

[0028] In the above embodiment, the first device can determine the target processing mode of the audio signal based on the decision indication information, and can determine the appropriate processing mode for the audio signal based on the indication information. By processing the audio signal and sending it to the second device, the optimal working mode that meets the audio experience service requirements can be achieved.

[0029] In combination with some embodiments of the first aspect, in some embodiments, the method further includes: determining decision indication information based on first information, where the first information is used to represent at least one of the delay requirements, accuracy characteristics, and listener freedom characteristics of the audio signal.

[0030] In combination with some embodiments of the first aspect, in some embodiments, the first information includes at least one of the following: a first parameter, the first parameter is used to identify the degree of freedom, and the degree of freedom is the information of the listener's position movement and / or posture change received by the sensor configured by the first device and / or the second device; a second parameter, the second parameter is used to identify the audio-visual scene and / or rendering accuracy, the audio-visual scene is the number of audio and video included in the audio signal, and the rendering accuracy is the accuracy required for rendering the audio signal; a third parameter, the third parameter is used to identify the background clarity of the audio signal; a fourth parameter, the fourth parameter is used to identify the delay, and the delay is the time between the second device sending the first parameter to the first device and the second device receiving the target audio signal sent by the first device.

[0031] In the above embodiment, the current application scenario can be determined by determining the first information, so as to facilitate determining a processing mode suitable for the current scenario based on the first information.

[0032] In combination with some embodiments of the first aspect, in some embodiments, the target processing mode includes at least one of the following: a first processing mode, in which the first device performs bitstream decoding, rendering processing, and stereo encoding on the audio signal, and after the first device sends the target audio signal to the second device, the second device performs stereo decoding and playback on the target audio signal; a second processing mode, in which the first device performs bitstream decoding, downmixing and / or format conversion, and bitstream encoding on the audio signal, and after the first device sends the target audio signal to the second device, the second device performs bitstream decoding, rendering, and playback on the target audio signal; a third processing mode, in which the first device performs bitstream decoding, first rendering processing, and intermediate variable encoding on the audio signal, and after the first device sends the target audio signal to the second device, the second device performs intermediate variable decoding, second rendering processing, and playback on the target audio signal.

[0033] In the above embodiment, by determining the processing mode, a reasonable division of labor in audio processing between the first device and the second device can be achieved, ensuring that the computing power sharing between the first device and the second device remains balanced.

[0034] In combination with some embodiments of the first aspect, in some embodiments, determining the target processing mode based on the decision indication information includes: determining the processing mode in which the decision indication information satisfies preset conditions as the target processing mode; or, receiving indication information, and determining the target processing mode based on the indication information, the indication information including the target processing mode determined by the user of the first device based on the decision indication information.

[0035] In the above embodiment, the decision indication information can be determined by the first information, and the target processing mode can be determined according to the decision indication information, so that the appropriate working mode can be selected according to the effective indication information.

[0036] In combination with some embodiments of the first aspect, in some embodiments, determining the processing mode in which the decision indication information meets the preset conditions as the target processing mode includes: when the decision indication information meets the first condition, determining the target processing mode as the first processing mode; when the decision indication information meets the second condition, determining the target processing mode as the second processing mode; when the decision indication information meets the third condition, determining the target processing mode as the third processing mode.

[0037] In the above embodiment, it is possible to determine a suitable target processing mode according to the decision indication information.

[0038] In combination with some embodiments of the first aspect, in some embodiments, based on the first information, determining the decision indication information includes: when the first parameter in the first information indicates that the degree of freedom is 0 and / or the fourth parameter indicates that the delay is low delay, determining the first parameter to be a first preset value; when the first parameter indicates that the degree of freedom is greater than or equal to 3 and less than 6 and / or the fourth parameter indicates that the delay is medium delay, determining the first parameter to be a second preset value; when the first parameter indicates that the degree of freedom is 6 and / or the fourth parameter indicates that the delay is high delay, determining the first parameter value to be a third preset value; when the second parameter in the first information indicates that the audio-visual scene is a single audio-visual scene and the rendering accuracy is high precision, determining the second parameter to be a fourth preset value; when the second parameter indicates that the audio-visual scene is a single audio-visual scene and the rendering accuracy is medium precision, determining the second parameter to be a fifth preset value; when the second parameter indicates that the audio-visual scene is a single audio-visual scene and the rendering accuracy is When the precision is low, the second parameter is determined to be the sixth preset value; when the second parameter indicates that the sound-image scene is a multi-sound-image scene and the rendering precision is high precision, the second parameter is determined to be the seventh preset value; when the second parameter indicates that the sound-image scene is a multi-sound-image scene and the rendering precision is medium precision, the second parameter is determined to be the eighth preset value; when the second parameter indicates that the sound-image scene is a multi-sound-image scene and the rendering precision is low precision, the second parameter is determined to be the ninth preset value; when the third parameter in the first information indicates that the background clarity is a pure background, the third parameter is determined to be the tenth preset value; when the third parameter indicates that the background clarity is a noisy background, the third parameter is determined to be the eleventh preset value; when the third parameter indicates that the background clarity is a medium background, the third parameter is determined to be the twelfth preset value; based on the preset function, with the first parameter, the second parameter, and the third parameter as independent variables, the function value of the preset function is determined as the decision indication information.

[0039] In the above embodiment, by determining the value of the parameter, the parameter index can be quantified, which facilitates determining the decision indication information according to the value of the parameter.

[0040] In combination with some embodiments of the first aspect, in some embodiments, based on a preset function, with the first parameter, the second parameter, and the third parameter as independent variables, determining the function value of the preset function as decision indication information includes: setting weight factors for the first parameter, the second parameter, and the third parameter respectively; and performing weighted summation of the first parameter, the second parameter, the third parameter and the weight factor to obtain decision indication information.

[0041] In the above embodiment, the decision indicator information may be determined according to the weight of the parameter and the value of the parameter.

[0042] In combination with some embodiments of the first aspect, in some embodiments, processing the audio signal in a target processing mode to obtain a target audio signal includes: when the target processing mode is a first processing mode, performing bitstream decoding, rendering processing, and stereo encoding on the audio signal to obtain a binaural audio signal as the target audio signal; when the target processing mode is a second processing mode, performing bitstream decoding, downmixing, and bitstream encoding on the audio signal to obtain an encoded audio signal as the target audio signal; when the target processing mode is a third processing mode, performing bitstream decoding, first rendering processing, and intermediate variable encoding on the audio signal to obtain a pre-rendered audio signal as the target audio signal.

[0043] In the above embodiment, the audio signal can be processed according to the target processing mode to obtain the target audio signal, and a suitable working mode can be selected to process the audio signal to ensure that the computing power sharing between the first device and the second device remains balanced.

[0044] In the second aspect, an embodiment of the present disclosure proposes an audio signal processing method, which is executed by a second device. The method includes receiving a target audio signal sent by a first device, where the target audio signal is obtained by the first device processing the audio signal in a target processing mode, and the target processing mode is determined from multiple processing modes based on decision indication information; processing the target audio signal in the target processing mode; the multiple processing modes are different processing functions performed based on the first device as a carrier.

[0045] In the above embodiment, the second device can receive the target audio signal and process the target audio signal, thereby achieving an optimal working mode that meets the audio experience service requirements.

[0046] In combination with some embodiments of the second aspect, in some embodiments, the first information includes at least one of the following: a first parameter, the first parameter is used to identify the degree of freedom, the degree of freedom being information about the listener's position movement and / or posture change received by the sensor configured by the first device and / or the second device; a second parameter, the second parameter is used to identify the audio-visual scene and / or rendering accuracy, the audio-visual scene is the number of audio and video included in the audio signal, and the rendering accuracy is the accuracy required for rendering the audio signal; a third parameter, the third parameter is used to identify the background clarity of the audio signal; a fourth parameter, the fourth parameter is used to identify the delay, the delay being the time between the second device sending the first parameter to the first device and the second device receiving the target audio signal sent by the first device.

[0047] In the above embodiment, the current application scenario can be determined by determining the first information, so as to facilitate determining a processing mode suitable for the current scenario based on the first information.

[0048] In combination with some embodiments of the second aspect, in some embodiments, the target processing mode includes at least one of the following: a first processing mode, in which the first device performs bitstream decoding, rendering processing, and stereo encoding on the audio signal, and after the first device sends the target audio signal to the second device, the second device performs stereo decoding and playback on the target audio signal; a second processing mode, in which the first device performs bitstream decoding, downmixing and / or format conversion, and bitstream encoding on the audio signal, and after the first device sends the target audio signal to the second device, the second device performs bitstream decoding, rendering, and playback on the target audio signal; a third processing mode, in which the first device performs bitstream decoding, first rendering processing, and intermediate variable encoding on the audio signal, and after the first device sends the target audio signal to the second device, the second device performs intermediate variable decoding, second rendering processing, and playback on the target audio signal.

[0049] In the above embodiment, by determining the processing mode, a reasonable division of labor in audio processing between the first device and the second device can be achieved, ensuring that the computing power tasks of the first device and the second device remain balanced.

[0050] In combination with some embodiments of the second aspect, in some embodiments, when the decision indication information meets the first condition, the target processing mode is the first processing mode; when the decision indication information meets the second condition, the target processing mode is the second processing mode; when the decision indication information meets the third condition, the target processing mode is the third processing mode, and the decision indication information is determined based on the first information.

[0051] In the above embodiment, it is possible to determine a suitable target processing mode according to the decision indication information.

[0052] In combination with some embodiments of the second aspect, in some embodiments, the decision indication information is the function value of a preset function, and the independent variables of the preset function are the first parameter, the second parameter, and the third parameter in the first information, wherein, when the first parameter indicates that the degree of freedom is 0 and / or the fourth parameter indicates that the delay is low, the first parameter is the first preset value; when the first parameter indicates that the degree of freedom is greater than or equal to 3 and less than 6 and / or the fourth parameter indicates that the delay is medium, the first parameter is the second preset value; when the first parameter indicates that the degree of freedom is 6 and / or the fourth parameter indicates that the delay is high, the first parameter value is the third preset value; when the second parameter indicates that the audio-visual scene is a single audio-visual scene and the rendering accuracy is high precision, the second parameter is the fourth preset value; when the second parameter indicates that the audio-visual scene is a single audio-visual scene and the rendering accuracy is medium precision. When the second parameter indicates that the sound and image scene is a single sound and image scene and the rendering precision is low precision, the second parameter is the sixth preset value; when the second parameter indicates that the sound and image scene is a multi-sound and image scene and the rendering precision is high precision, the second parameter is the seventh preset value; when the second parameter indicates that the sound and image scene is a multi-sound and image scene and the rendering precision is medium precision, the second parameter is the eighth preset value; when the second parameter indicates that the sound and image scene is a multi-sound and image scene and the rendering precision is low precision, the second parameter is the ninth preset value; when the third parameter indicates that the background clarity is a pure background, the third parameter is the tenth preset value; when the third parameter indicates that the background clarity is a noisy background, the third parameter is the eleventh preset value; when the third parameter indicates that the background clarity is a medium background, the third parameter is the twelfth preset value.

[0053] In the above embodiment, by determining the value of the parameter, the parameter index can be quantified, which facilitates determining the decision indication information according to the value of the parameter.

[0054] In combination with some embodiments of the second aspect, in some embodiments, processing the target audio signal in the target processing mode includes: when the target processing mode is the first processing mode, stereo decoding and playing the target audio signal; when the target processing mode is the second processing mode, bitstream decoding, rendering and playing the target audio signal; when the target processing mode is the third processing mode, intermediate variable decoding, second rendering processing and playing the target audio signal.

[0055] In the above embodiment, the audio signal can be processed according to the target processing mode to obtain the target audio signal, and a suitable working mode can be selected to process the audio signal to ensure that the computing power sharing between the first device and the second device remains balanced.

[0056] In a third aspect, embodiments of the present disclosure provide a first device comprising a processing module configured to determine a target processing mode from multiple processing modes based on decision indication information; the processing module further configured to process an audio signal in the target processing mode to obtain a target audio signal; and a transceiver module configured to transmit the target audio signal to a second device. The multiple processing modes are based on different processing functions performed by the first device as a carrier.

[0057] In a fourth aspect, an embodiment of the present disclosure proposes a second device, comprising a transceiver module for receiving a target audio signal sent by a first device, where the target audio signal is obtained by the first device processing the audio signal in a target processing mode, where the target processing mode is determined from the multiple processing modes based on the decision indication information; a processing module for processing the target audio signal in the target processing mode; the multiple processing modes are different processing functions performed based on the first device as a carrier.

[0058] In a fifth aspect, an embodiment of the present disclosure proposes a communication device, which includes: wherein the processor is used to call instructions to enable the communication device to execute the method described in the optional implementation manner of the first aspect and the second aspect of the embodiment of the present disclosure.

[0059] In the sixth aspect, an embodiment of the present disclosure proposes a communication system, which includes: a first device and a second device; wherein the first device is configured to execute the method described in the first aspect and the optional implementation of the first aspect, and the second device is configured to execute the method described in the second aspect and the optional implementation of the second aspect.

[0060] In the seventh aspect, an embodiment of the present disclosure proposes a storage medium, wherein the computer storage medium stores computer-executable instructions; after the computer-executable instructions are executed by the processor, the method described in the first aspect, the optional implementation of the first aspect, the second aspect, and the optional implementation of the second aspect can be executed.

[0061] In an eighth aspect, an embodiment of the present disclosure proposes a computer program product, characterized in that it includes a computer program, and when the computer program is executed by a processor, it implements the method of any one of the embodiments of the first and second aspects of the present disclosure.

[0062] It is understandable that the first device, the second device, the communication device, the communication system, and the storage medium are all used to perform the method proposed in the embodiment of the present disclosure. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method and will not be repeated here.

[0063] The present disclosure provides a communication method and apparatus, a communication device, a communication system, and a storage medium. In some embodiments, the terms "communication method" and "information processing method" and "communication method" are interchangeable; the terms "apparatus" and "terminal" and "network device" and "communication device" are interchangeable; and the terms "information processing system" and "communication system" are interchangeable.

[0064] The embodiments of the present disclosure are not exhaustive and are merely illustrative of some embodiments, and are not intended to be a specific limitation on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment can be arbitrarily exchanged. In addition, the optional implementation methods in a certain embodiment can be arbitrarily combined; in addition, the embodiments can be arbitrarily combined. For example, some or all steps of different embodiments can be arbitrarily combined, and a certain embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.

[0065] In each embodiment of the present disclosure, unless otherwise specified or provided for by logic, the terms and / or descriptions between the embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form a new embodiment based on their inherent logical relationships.

[0066] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure.

[0067] In the embodiments of the present disclosure, unless otherwise specified, elements expressed in the singular, such as "a", "an", "the", "above", "said", "the", "the", etc., may mean "one and only one", or "one or more", "at least one", etc. For example, when using articles such as "a", "an", "the" in English in translation, the noun following the article may be understood as a singular expression or a plural expression.

[0068] In the embodiments of the present disclosure, “plurality” refers to two or more.

[0069] In some embodiments, the terms "at least one of", "at least one of", "at least one of", "one or more", "a plurality of", "multiple", etc. can be used interchangeably.

[0070] In the embodiments of the present disclosure, descriptions such as “at least one of A, B, C…”, “A and / or B and / or C…”, etc. include the situation where any one of A, B, C… exists alone, and also include any combination of any multiple of A, B, C…, and each situation can exist alone; for example, “at least one of A, B, C” includes the situation where A exists alone, B exists alone, C exists alone, the combination of A and B, the combination of A and C, the combination of B and C, and the combination of A, B, and C; for example, A and / or B includes the situation where A exists alone, B exists alone, and the combination of A and B.

[0071] In some embodiments, descriptions such as "in one case A, in another case B," or "in response to one case A, in response to another case B," may include the following technical solutions depending on the situation: executing A independently of B (in some embodiments, A); executing B independently of A (in some embodiments, B); selectively executing A and B (in some embodiments, selecting between A and B); and executing both A and B (in some embodiments, A and B). The same applies when there are more branches, such as A, B, and C.

[0072] The prefixes such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different description objects and do not constitute any restriction on the position, order, priority, quantity or content of the description objects. For the statement of the description object, please refer to the description in the context of the claims or embodiments, and no unnecessary restriction should be constituted due to the use of prefixes. For example, if the description object is a "field", the ordinal number before the "field" in the "first field" and the "second field" does not limit the position or order between the "fields". "First" and "second" do not limit whether the "fields" they modify are in the same message, nor do they limit the order of the "first field" and the "second field". For another example, if the description object is a "level", the ordinal number before the "level" in the "first level" and the "second level" does not limit the priority between the "levels". For another example, the number of description objects is not limited by the ordinal number and can be one or more. Taking "first device" as an example, the number of "devices" can be one or more. In addition, the objects modified by different prefixes can be the same or different. For example, if the description object is "device", then the "first device" and the "second device" can be the same device or different devices, and their types can be the same or different; for another example, if the description object is "information", then the "first information" and the "second information" can be the same information or different information, and their contents can be the same or different.

[0073] In some embodiments, “including A,” “comprising A,” “used to indicate A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.

[0074] In some embodiments, terms such as "in response to...", "in response to determining...", "in the case of...", "at the time of...", "when...", "if...", "if...", etc. can be used interchangeably.

[0075] In some embodiments, terms such as "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not less than", and "above" can be replaced with each other, and terms such as "less than", "less than or equal to", "not greater than", "less than", "less than or equal to", "not more than", "lower than", "lower than or equal to", "not higher than", and "below" can be replaced with each other.

[0076] In some embodiments, devices, etc. can be interpreted as physical or virtual, and their names are not limited to the names recorded in the embodiments. Terms such as "device", "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", and "subject" can be used interchangeably.

[0077] In some embodiments, the terms "access network device (AN device)", "radio access network device (RAN device)", "base station (BS)", "radio base station" "fixed station", "node", "access point", "transmission point (TP)", "reception point (RP)", "transmission / reception point (TRP)", "panel", "antenna panel", "antenna array", "cell", "macro cell", "small cell", "femto cell", "pico cell", "sector", "cell group", "carrier", "component carrier", "bandwidth part (BWP)" and the like may be used interchangeably.

[0078] In some embodiments, the terms "terminal", "terminal device", "user equipment (UE)", "user terminal", "mobile station (MS)", "mobile terminal (MT)", subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, etc. can be used interchangeably.

[0079] In some embodiments, the access network device, the core network device, or the network device can be replaced by a terminal. For example, the various embodiments of the present disclosure can also be applied to a structure in which the communication between the access network device, the core network device, or the network device and the terminal is replaced by communication between multiple terminals (for example, it can also be called device-to-device (D2D), vehicle-to-everything (V2X), etc.). In this case, it can also be set as a structure in which the terminal has all or part of the functions of the access network device. In addition, language such as "uplink" and "downlink" can also be replaced by language corresponding to communication between terminals (for example, "side"). For example, uplink channels, downlink channels, etc. can be replaced by side channels, and uplinks, downlinks, etc. can be replaced by side links.

[0080] In some embodiments, the terminal may be replaced by an access network device, a core network device, or a network device. In this case, the access network device, the core network device, or the network device may have a structure that has all or part of the functions of the terminal.

[0081] In some embodiments, the names of information, etc. are not limited to the names described in the embodiments, and terms such as "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "symbol", "codeword", "codebook", "codeword", "codepoint", "bit", "data", "program", and "chip" can be used interchangeably.

[0082] In some embodiments, terms such as "uplink", "uplink", "physical uplink" can be interchangeable with each other, and terms such as "downlink", "downlink", "physical downlink" can be interchangeable with each other, and terms such as "side", "sidelink", "side communication", "sidelink communication", "direct connection", "direct link", "direct communication", "direct link communication" can be interchangeable with each other.

[0083] In some embodiments, the terms "downlink control information (DCI)", "downlink (DL) assignment", "DL DCI", "uplink (UL) grant", "UL DCI" and the like may be used interchangeably.

[0084] In some embodiments, terms such as "physical downlink shared channel (PDSCH)" and "DL data" can be used interchangeably, and terms such as "physical uplink shared channel (PUSCH)" and "UL data" can be used interchangeably.

[0085] In some embodiments, the terms "radio", "wireless", "radio access network (RAN)", "access network (AN)", "RAN-based" and the like may be used interchangeably.

[0086] In some embodiments, the terms "synchronization signal (SS)", "synchronization signal block (SSB)", "reference signal (RS)", "pilot", "pilot signal" and the like can be used interchangeably.

[0087] In some embodiments, terms such as "moment", "time point", "time", and "time position" can be replaced with each other, and terms such as "duration", "period", "time window", "window", and "time" can be replaced with each other.

[0088] In some embodiments, "obtain", "get", "obtain", "receive", "transmit", "bidirectional transmission", "send and / or receive" can be interchangeable, and can be interpreted as receiving from other entities, obtaining from a protocol, obtaining by self-processing, autonomous implementation, etc.

[0089] In some embodiments, terms such as "send", "transmit", "report", "download", "transmit", "bidirectional transmission", "send and / or receive" can be used interchangeably.

[0090] In some embodiments, "predetermined" and "preset" can be interpreted as pre-specified in a protocol, etc., or can be interpreted as a pre-set action performed by a device, etc.

[0091] In some embodiments, determining may be interpreted as judging, calculating, computing, processing, deriving, investigating, searching, looking up, retrieving, ascertaining, receiving, transmitting, inputting, outputting, accessing, resolving, selecting, choosing, establishing, comparing, “assuming,” “expecting,” “considering,” broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, assigning, and the like, but is not limited thereto.

[0092] In some embodiments, the determination or judgment can be performed by a value represented by 1 bit (0 or 1), or by a true or false value (Boolean value) represented by true or false, or by comparison of numerical values ​​(for example, comparison with a predetermined value), but is not limited thereto.

[0093] In some embodiments, "network" can be interpreted as devices included in the network (eg, access network equipment, core network equipment, etc.).

[0094] In some embodiments, "not expecting to receive" can be interpreted as not receiving on time domain resources and / or frequency domain resources, or as not performing subsequent processing on the data after receiving it; "not expecting to send" can be interpreted as not sending, or as sending but not expecting the recipient to respond to the content sent.

[0095] In some embodiments, obtaining data, information, etc. may comply with the laws and regulations of the country where the data is obtained.

[0096] In some embodiments, data, information, etc. may be obtained after obtaining the user's consent. To solve the above problems, the present disclosure proposes a communication method and apparatus, a communication device, a communication system, and a storage medium.

[0097] The first generation of mobile communication technology (1G) began in the 1980s. 1G was the first generation of wireless cellular technology and an analog mobile communication network. The upgrade from 1G to 2G shifted mobile phones from analog to digital communication. my country adopted the Global System for Mobile Communications (GSM) network standard, using voice codecs such as Adaptive Multi-Rate (AMR), Enhanced Full Rate (EFR), Full Rate (FR), and Half Rate (HR) to provide single-channel narrowband voice services.

[0098] The 3G mobile communications system was proposed by the International Telecommunication Union (ITU) for international mobile communications in 2000. Its speech codec uses AMR-WB to provide single-channel wideband voice services. 4G is a significant improvement on 3G technology, using an all-IP approach for both data and voice, providing real-time HD / HD+Voice services. The EVS codec ensures high-quality compression and reconstruction of both voice and audio.

[0099] With the development of 4G / 5G and AR / VR / MR technologies, the immersive audio services provided by XR devices are a brand-new audio experience. Typical XR devices, such as AR glasses, are subject to strict restrictions on their appearance, structure, and weight for ease of wear. Therefore, to complete audio rendering and playback on the XR device, the entire decoding and rendering process of the received audio stream signal, from decoding to binaural rendering, needs to be rationally distributed among various processing units. AR glasses can be divided into the following types based on their functional allocation structure: 5G Standalone AR UE, 5G Edge-Dependent AR UE, 5G Wireless Tethered AR UE, and 5G Wired Tethered AR UE.

[0100] Among them, 5G edge-dependent AR user equipment and 5G wireless-constrained AR user equipment are the two target devices to be studied in the technology of this invention.

[0101] For the two aforementioned XR devices, audio processing is appropriately distributed between the Cloud / Edge Device and the XR device, ultimately outputting rendered binaural signals to the user on the XR device. Related technologies include the following audio processing distribution methods.

[0102] Solution 1: Existing technology Solution 1 uses a terminal device (such as a smartphone) to perform decoding and rendering processing. The rendering processing is to render the decoded output signal based on the head rotation information (3 degrees of freedom) output by the head tracking device or the head rotation information and the user's latest position information (6 degrees of freedom). The rendered binaural signal is encoded and sent to the playback device (such as headphones). After receiving the code stream signal, the playback device decodes and outputs the binaural signal. The user obtains an immersive audio experience service through the played binaural signal.

[0103] The problem with the method proposed in the above-mentioned solution 1 is that it takes a certain delay for the information output by the head tracking device to be transmitted to the terminal device. The delay exceeding the threshold makes it impossible for the binaural rendering signal to track the user's head rotation and / or position information in real time, resulting in the binaural rendering signal received by the user not being completely consistent with the signal that the user should hear at the current moment.

[0104] Solution 2: Existing technology Solution 2 directly transmits the code stream signal to the playback device (such as headphones) after the cloud terminal device (such as a smartphone) receives it. The playback device renders the decoded output signal based on the head rotation information (3 degrees of freedom) or the head rotation information and the user's latest position information (6 degrees of freedom) output by the head tracking device. The rendered binaural signal is played through a speaker to obtain an immersive audio experience service.

[0105] The problem with the method proposed in the above-mentioned solution 2 is that the playback device (such as headphones) is responsible for decoding and rendering processing, which requires the playback device to have powerful computing power and storage resources, that is, it has high requirements on the battery capacity of the playback device.

[0106] Solution 3: Existing technology Solution 3 uses a terminal device (such as a smart phone) to first decode the received bit stream signal and then perform a first-level rendering process (such as rendering process 1 in the figure). The first-level rendering process is based on the information output by the head tracking device that changes slowly over time to render the decoded signal. The signal after the first-level rendering process is encoded to obtain a bit stream signal and sent to the playback device (such as headphones). The playback device receives the bit stream signal and first decodes it and then performs a second-level rendering process on the output signal based on the real-time information output by the head tracking device. The rendered binaural signal is played back through the speaker to obtain an immersive audio service experience.

[0107] The problem with the method proposed in Solution 3 above is that the two-stage rendering processing method increases the overall complexity of the system solution from the received bitstream signal to the acquisition of binaural signals. How to balance the computing power sharing ratio of the two-stage rendering also has a significant impact on the performance of the end-to-end solution.

[0108] In order to solve the above problems, this solution proposes an audio signal processing method, which can select a suitable working mode according to the indication information. The specific content of the method is as follows.

[0109] FIG1 is a schematic diagram illustrating an architecture of a communication system according to an embodiment of the present disclosure. As shown in FIG1 , a communication system 100 may include a first device 101 and a second device 102 .

[0110] In some embodiments, the first device may be a device that receives an audio signal. For example, the first device may be a terminal, including but not limited to a mobile phone, a computer, a tablet, a conference system device, an AR / VR device, an automobile, etc. For example, the first device may receive an audio signal from a network device. For example, the audio signal from an incoming call on a mobile phone may be received from a network device. Optionally, the first device may process the received audio signal. Optionally, the first device may provide the processed audio signal or the unprocessed audio signal to the second device.

[0111] In some embodiments, the second device may be a device that outputs a target audio signal. For example, the second device may be a terminal, such as an XR device, specifically, augmented reality (AR) glasses. The second device may receive an audio signal from the first device. The audio signal may be an audio signal processed by the first device, or may be an unprocessed audio signal from the first device. Optionally, the second device may process the audio signal received from the first device. In some embodiments, the terminal includes, for example, a mobile phone, a wearable device, an Internet of Things device, a car with communication function, a smart car, a tablet computer, a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal device in industrial control, a wireless terminal device in self-driving, a wireless terminal device in remote medical surgery, a wireless terminal device in a smart grid, a wireless terminal device in transportation safety, a wireless terminal device in a smart city, and at least one of a wireless terminal device in a smart home, but is not limited thereto.

[0112] In some embodiments, the communication system may further include a third device. Optionally, the third device may be a device that provides an audio signal to the first device, such as a network device.

[0113] In some embodiments, a "network device" may be an access network device, such as a base station. Optionally, a network device may also be a broad network device including a core network device. In some embodiments, an access network device is, for example, a node or device that connects a terminal to a wireless network. The access network device may include at least one of an evolved NodeB (eNB), a next generation evolved NodeB (ng-eNB), a next generation NodeB (gNB), a node B (NB), a home nodeB (HNB), a home evolved nodeB (HeNB), a wireless backhaul device, a radio network controller (RNC), a base station controller (BSC), a base transceiver station (BTS), a base band unit (BBU), a mobile switching center, a base station in a 6G communication system, an open RAN, a cloud RAN, a base station in other communication systems, and an access node in a wireless fidelity (WiFi) system, but is not limited thereto.

[0114] In some embodiments, the technical solution of the present disclosure can be applied to the Open RAN architecture. In this case, the interfaces between or within the access network devices involved in the embodiments of the present disclosure can be transformed into internal interfaces of the Open RAN, and the processes and information interactions between these internal interfaces can be implemented through software or programs.

[0115] In some embodiments, the access network device can be composed of a centralized unit (CU) and a distributed unit (DU), where the CU can also be called a control unit. The CU-DU structure can be used to split the protocol layer of the access network device, with the functions of some protocol layers centrally controlled by the CU, and the functions of the remaining part or all of the protocol layers distributed in the DU, which is centrally controlled by the CU, but is not limited to this.

[0116] In some embodiments, a core network device may be a single device comprising one or more network elements, or may be a plurality of devices or a group of devices, each comprising all or part of one or more network elements. A network element may be virtual or physical. The core network may include, for example, at least one of an Evolved Packet Core (EPC), a 5G Core Network (5GCN), and a Next Generation Core (NGC).

[0117] In some embodiments, the above-mentioned one or more network elements may include, for example, AMF, UPF, MME, etc., and may also include other network elements, such as Policy Control Function (PCF), Application Function (AF), Network Application Function (NAF), Application Layer Authentication and Key Management Anchor Function (AAnF), Bootstrapping Server Functionality (BSF), Session Management Function (SMF), etc.

[0118] It can be understood that the communication system described in the embodiment of the present disclosure is for the purpose of more clearly illustrating the technical solution of the embodiment of the present disclosure, and does not constitute a limitation on the technical solution proposed in the embodiment of the present disclosure. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solution proposed in the embodiment of the present disclosure is also applicable to similar technical problems.

[0119] The following embodiments of the present disclosure may be applied to the communication system 100 shown in Figure 1, or a portion thereof, but are not limited thereto. The entities shown in Figure 1 are illustrative only. The communication system may include all or part of the entities shown in Figure 1, or may include other entities outside of Figure 1. The number and form of the entities may be arbitrary. The connection relationship between the entities is illustrative only. The entities may be connected or disconnected, and the connection may be in any manner, including direct or indirect, wired or wireless.

[0120] The embodiments of the present disclosure can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G new radio (NR), future radio access (FRA), new radio access technology (RAT), new radio (NR), new radio access (NX), future generation radio access (FX), Global System for Mobile communications (GSM (registered trademark)), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, Ultra-WideBand (UWB), Bluetooth (registered trademark), Public Land Mobile Network (PLMN) networks, Device-to-Device (D2D) systems, Machine-to-Machine (M2M) systems, Internet of Things (IoT) systems, Vehicle-to-Everything (V2X), systems utilizing other communication methods, and next-generation systems based on and extending these methods. Furthermore, multiple systems may be combined (for example, a combination of LTE or LTE-A with 5G).

[0121] FIG2 is an interactive diagram illustrating an audio signal processing method according to an embodiment of the present disclosure. As shown in FIG2 , the present disclosure embodiment relates to a communication method for use in a communication system 100. The communication system 100 may include a first device 101 and a second device 102. In some embodiments, the method includes:

[0122] Step 2101: The first device determines decision indication information based on the first information.

[0123] In some embodiments, the first information is used to represent at least one of a delay requirement, an accuracy characteristic, and a listener freedom characteristic of the audio signal.

[0124] In some embodiments, the first information includes at least one of the following:

[0125] A first parameter, the first parameter is used to identify a degree of freedom, where the degree of freedom is information about a listener's position movement and / or posture change received by a sensor configured by the first device and / or the second device;

[0126] A second parameter, the second parameter is used to identify an audiovisual scene and / or rendering precision, where the audiovisual scene is the number of audio and video included in the audio signal, and the rendering precision is the precision required for rendering the audio signal;

[0127] A third parameter is used to identify the background clarity of the audio signal;

[0128] The fourth parameter is used to identify the delay, which is the time between the second device sending the first parameter to the first device and the second device receiving the target audio signal sent by the first device.

[0129] In the above embodiment, for example, the first parameter may include 0 degrees of freedom, 3 degrees of freedom, 3+ degrees of freedom, 6 degrees of freedom, etc. The numerical value of the first parameter can determine the requirement for delay. Generally, the larger the degree of freedom, the lower the required delay. For example, in the 6-degree-of-freedom scenario, the lowest possible delay is required, so that users can obtain an audio service experience that matches their own posture and rotation information.

[0130] In the above embodiment, the fourth parameter is used to identify the delay, where the delay refers to the time required between the first device and the second device receiving the user posture, rotation and other information of the first parameter and outputting the rendered binaural signal, among which the delay of the above-mentioned first processing mode is the largest, the delay of the second processing mode is the smallest, and the delay of the third processing mode is in the middle.

[0131] In the above embodiment, the second parameter may be the sound and image accuracy. For example, the sound and image accuracy may be divided into two dimensions, wherein the first dimension is the sound and image scene, such as a single sound and image scene and a multi-sound and image scene; the second dimension is the rendering accuracy, such as high precision, medium precision and low precision.

[0132] In some embodiments, the decision indication information may be a recommended parameter for identifying different processing modes to be selected. Alternatively, the decision indication parameter may be a comprehensive indicator determined by the first device using the first information.

[0133] In some embodiments, determining decision indication information based on the first information includes:

[0134] When the first parameter in the first information indicates that the degree of freedom is 0 and / or the fourth parameter indicates that the delay is low, the first parameter is determined to be the first preset value; when the first parameter indicates that the degree of freedom is greater than or equal to 3 and less than 6 and / or the fourth parameter indicates that the delay is medium, the first parameter is determined to be the second preset value; when the first parameter indicates that the degree of freedom is 6 and / or the fourth parameter indicates that the delay is high, the first parameter value is determined to be the third preset value; when the second parameter in the first information indicates that the audio-visual scene is a single audio-visual scene and the rendering accuracy is high precision, the second parameter is determined to be the fourth preset value; when the second parameter indicates that the audio-visual scene is a single audio-visual scene and the rendering accuracy is medium precision, the second parameter is determined to be the fifth preset value; when the second parameter indicates that the audio-visual scene is a single audio-visual scene and the rendering accuracy is high precision. When the precision is low, the second parameter is determined to be a sixth preset value; when the second parameter indicates that the sound-image scene is a multi-sound-image scene and the rendering precision is high, the second parameter is determined to be a seventh preset value; when the second parameter indicates that the sound-image scene is a multi-sound-image scene and the rendering precision is medium, the second parameter is determined to be an eighth preset value; when the second parameter indicates that the sound-image scene is a multi-sound-image scene and the rendering precision is low, the second parameter is determined to be a ninth preset value; when the third parameter in the first information indicates that the background clarity is a pure background, the third parameter is determined to be a tenth preset value; when the third parameter indicates that the background clarity is a noisy background, the third parameter is determined to be an eleventh preset value; when the third parameter indicates that the background clarity is a medium background, the third parameter is determined to be a twelfth preset value;

[0135] Based on a preset function, with the first parameter, the second parameter and the third parameter as independent variables, a function value of the preset function is determined as decision indication information.

[0136] In other words, the decision indication information can be obtained by a preset function with the first parameter, the second parameter and the third parameter as independent variables. For example, the preset function can be a weighted sum function. For example, when the decision indication information is represented by Y, it can be expressed as Y=Function(X1, X2, X3)=aX1+bX2+cX3, where a, b, and c are the weight factors of the first parameter, the second parameter and the third parameter respectively, and X1, X2, and X3 are the parameter values ​​of the first parameter, the second parameter and the third parameter, and their parameter values ​​can be one of the above-mentioned first to twelfth preset values.

[0137] In the above embodiment, the specific values ​​of the first preset value to the twelfth preset value can be determined according to actual conditions, and the present disclosure does not limit this.

[0138] In some embodiments, based on a preset function, with the first parameter, the second parameter, and the third parameter as independent variables, determining the function value of the preset function as the decision indication information includes:

[0139] Setting weight factors for the first parameter, the second parameter, and the third parameter, respectively. For example, the weight factors for the first parameter, the second parameter, and the third parameter can be set based on actual conditions, which is not limited in the present disclosure.

[0140] A weighted sum is performed on the first parameter, the second parameter, the third parameter and the weight factor to obtain decision indication information, for example, Y=aX1+bX2+cX3.

[0141] Step 2102: The first device determines a target processing mode.

[0142] In some embodiments, the first device may determine the target processing mode from a plurality of processing modes according to the decision indication information.

[0143] In some embodiments, the multiple processing modes are based on different processing functions performed by the first device as a carrier. In other words, in the multiple processing modes, the processing functions performed by the first device as a carrier are different. In other words, the first device performs different processing functions in different processing modes. In other words, the first device and / or the second device are carriers for processing audio signals, and in different processing modes, the first device performs different processing functions, and accordingly, the second device performs different processing functions.

[0144] In some embodiments, the target processing mode may be one or more selected from a plurality of processing modes, and the plurality of processing modes may include at least one of the following, that is, the target processing mode may include at least one of the following:

[0145] In the first processing mode, the first device performs code stream decoding, rendering processing, and stereo encoding on the audio signal. After the first device sends the target audio signal to the second device, the second device performs stereo decoding and playback on the target audio signal.

[0146] The second processing mode is that the first device performs code stream decoding, downmixing and / or format conversion, and code stream encoding on the audio signal. After the first device sends the target audio signal to the second device, the second device decodes the code stream, renders, and plays the target audio signal;

[0147] The third processing mode is that the first device performs code stream decoding, first rendering processing and intermediate variable encoding on the audio signal. After the first device sends the target audio signal to the second device, the second device performs intermediate variable decoding, second rendering processing and playback on the target audio signal.

[0148] In some embodiments, a single-image scenario can be based on object-based sound capture, where one object corresponds to one image and one channel. In a single-image scenario, due to the small number of channels, decoding and rendering are relatively simple. Preferably, the second processing mode can be used.

[0149] In some embodiments, a multi-sound image scene may include more than three sound images. For example, a multi-sound image scene may be based on channel signals for sound acquisition, where one sound image may correspond to multiple channels. Due to the large number of channels in a multi-sound image scene, the decoding and rendering schemes are relatively complex. Preferably, the first processing mode may be used in this case.

[0150] In some embodiments, for a multi-sound image scene with fewer sound images, for example, 2 or 3 sound images, when the sound is collected in MASA format, the decoding and rendering complexity is moderate, so the third processing mode can be selected.

[0151] In some embodiments, for example, when the rendering precision is low precision, the second processing mode may be adopted, when the rendering precision is high precision, the first processing mode may be adopted, and when the rendering precision is medium precision, the third processing mode may be adopted.

[0152] In the above embodiment, the background clarity may include a pure background, a noisy background, a medium background, etc. For example, the background clarity may be determined by the signal-to-noise ratio. For example, a larger signal-to-noise ratio indicates a higher background clarity. When the signal-to-noise ratio exceeds a certain threshold, it may be determined that the current background is a pure background.

[0153] For example, when the background clarity is a pure background, fewer input channels are used, and the second processing mode can be used at this time; when the background clarity is a noisy background, more input channels are used, and the first processing mode can be used at this time; when the background clarity is a medium background, the third processing mode can be used.

[0154] In some embodiments, the first device may determine a target processing mode for processing the audio signal based on the decision indication information, including an automatic determination mode and a semi-automatic determination mode (or in other words, a non-automatic determination mode or a manual determination mode). The following details methods for the first device to determine the target processing mode in the two modes:

[0155] 1. Automatic determination mode (or first-level processing mode)

[0156] In the automatic determination mode, the first device may automatically determine the target processing mode according to the decision indication information.

[0157] For example, the first device may determine the decision indication information of each processing mode, ie, the Y value, and determine the maximum Y value from multiple Y values. The processing mode corresponding to the maximum Y value is the target processing mode. The present disclosure does not limit other methods for automatically determining the target processing mode.

[0158] In the above embodiment, the target processing mode can be directly determined by the first device based on the decision indication information, that is, the first device can directly determine the target processing mode. For example, when Y is less than or equal to 3, the first processing mode is selected; when Y is greater than 3 and less than 8, the third processing mode is selected; when Y is greater than or equal to 8, the second processing mode is selected.

[0159] 2. Semi-automatic determination mode (or in other words, second-level processing mode, non-automatic determination mode or manual determination mode):

[0160] In the semi-automatic determination mode, the first device may receive instruction information and determine the target processing mode according to the instruction information, where the instruction information includes the target processing mode determined by the user of the first device according to the decision instruction information.

[0161] For example, the first device provides some or all of the processing modes and their Y values ​​to the user, or the first device may select at least one candidate processing mode based on the Y value and provide it to the user. After the user makes a decision based on the information provided by the first device, the user feeds back indication information to the first device. The indication information includes the processing mode decided by the user. The first device determines the target processing mode based on the user's feedback.

[0162] In the above embodiment, the target processing mode can be manually determined by the user. For example, the first device can determine at least one candidate processing mode based on the decision indication information, and display the candidate processing modes and the decision indication information corresponding to each candidate processing mode to the user, that is, the Y value. The user can make a selection based on the pushed information. At this time, it can be determined that the processing mode selected by the user is the target processing mode.

[0163] In some embodiments, determining the target processing mode according to the decision indication information includes:

[0164] When the decision indication information satisfies a first condition, determining the target processing mode to be the first processing mode;

[0165] When the decision indication information satisfies the second condition, determining the target processing mode to be the second processing mode;

[0166] When the decision indication information satisfies the third condition, the target processing mode is determined to be the third processing mode.

[0167] For example, the first condition may be that Y is less than or equal to a first threshold. When the decision indication information Y value is less than or equal to the first threshold, the target processing mode can be determined to be the first processing mode. When the Y value is small, it indicates that the current decoding and rendering scheme is relatively complex. At this time, the first processing mode can be selected to meet the scheme requirements.

[0168] For another example, the second condition may be that Y refers to a second threshold value that is greater than or equal to the second threshold value. When the decision indication information is greater than or equal to the second threshold value, the target processing mode can be determined to be the second processing mode. When the Y value is large, it indicates that the decoding and rendering schemes are relatively simple, and the second processing mode can be used to reduce the latency while meeting the scheme requirements.

[0169] For another example, the third condition may be that Y refers to a value greater than a first threshold value and less than a second threshold value, wherein the first threshold value and the second threshold value may be specified based on actual conditions, which is not limited in the present disclosure.

[0170] Step 2103: The first device processes the audio signal in a target processing mode to obtain a target audio signal.

[0171] In some embodiments, processing the audio signal in the target processing mode to obtain the target audio signal includes:

[0172] When the target processing mode is the first processing mode, performing code stream decoding, rendering processing, and stereo encoding on the audio signal to obtain a binaural audio signal as the target audio signal;

[0173] When the target processing mode is the second processing mode, performing code stream decoding, down-mixing, and code stream encoding on the audio signal to obtain an encoded audio signal as the target audio signal;

[0174] When the target processing mode is the third processing mode, the audio signal is subjected to code stream decoding, first rendering processing, and intermediate variable encoding to obtain a pre-rendered audio signal as the target audio signal.

[0175] Step 2104: The first device sends the target audio signal to the second device.

[0176] In some embodiments, the first device and the second device may jointly complete the audio signal processing. After completing its own processing process, the first device may send the processed target audio signal to the second device.

[0177] Step 2105: The second device processes the target audio signal in a target processing mode.

[0178] In some embodiments, the second device processing the target audio signal in the target processing mode includes:

[0179] When the target processing mode is the first processing mode, stereo decoding and playing the target audio signal;

[0180] When the target processing mode is the second processing mode, decoding, rendering and playing the target audio signal;

[0181] When the target processing mode is the third processing mode, intermediate variable decoding, second rendering processing, and playback are performed on the target audio signal.

[0182] The communication method involved in the embodiment of the present disclosure may include at least one of steps 2101 to 2105. For example, step 2102 may be implemented as an independent embodiment, and steps 2101+2102+2103+2104+2105 may be implemented as independent embodiments, but are not limited thereto.

[0183] In this embodiment or example, unless there is any contradiction, each step can be independent, arbitrarily combined or exchanged in order, the optional methods or optional examples can be arbitrarily combined, and can be arbitrarily combined with any steps of other embodiments or other examples.

[0184] FIG3 is a flow chart of an audio signal processing method according to an embodiment of the present disclosure. As shown in FIG3 , the embodiment of the present disclosure relates to an audio signal processing method for a first device 101, the method comprising:

[0185] Step 3101: Determine a target processing mode for processing the audio signal based on the decision indication information.

[0186] For optional implementations of step 3101 , reference may be made to the optional implementations of step 2101 and step 2102 in FIG. 2 and other related parts in the embodiments involved in the steps of FIG. 2 .

[0187] Step 3102: Process the audio signal in a target processing mode to obtain a target audio signal.

[0188] For optional implementations of step 3102, reference may be made to the optional implementations of step 2103 in FIG. 2 and other related parts of the embodiments involved in the steps of FIG. 2 .

[0189] Step 3103: Send the target audio signal to the second device.

[0190] For optional implementations of step 3103 , reference may be made to the optional implementations of step 2104 in FIG. 2 and other related parts of the embodiments involved in the steps of FIG. 2 .

[0191] In some embodiments, the first device may send the target audio signal to the second device, but is not limited thereto. The first device may also send the target audio signal to other entities.

[0192] The communication method involved in the embodiment of the present disclosure may include at least one of steps 3101 to 3103. For example, step 3101 may be implemented as an independent embodiment, and steps 3101+3102+3103 may be implemented as independent embodiments, but are not limited thereto.

[0193] In this embodiment or example, unless there is any contradiction, each step can be independent, arbitrarily combined or exchanged in order, the optional methods or optional examples can be arbitrarily combined, and can be arbitrarily combined with any steps of other embodiments or other examples.

[0194] FIG4 is a flow chart of an audio signal processing method according to an embodiment of the present disclosure. As shown in FIG4 , the embodiment of the present disclosure relates to an audio signal processing method for a second device 102, the method comprising:

[0195] Step 4101: Receive a target audio signal.

[0196] For optional implementations of step 4101 , reference may be made to the optional implementations of step 2104 in FIG. 2 , step 3103 in FIG. 3 , and other related parts in the embodiments involved in the steps of FIG. 2 and FIG. 3 .

[0197] In some embodiments, the first device receives the target audio signal sent by the second device, but is not limited thereto and may also receive the target audio signal sent by other entities.

[0198] In some embodiments, the first device obtains a target audio signal specified by a protocol.

[0199] In some embodiments, the first device obtains the target audio signal from an upper layer(s).

[0200] In some embodiments, the first device performs processing to obtain the target audio signal.

[0201] Step 4102: Process the target audio signal in a target processing mode.

[0202] For optional implementations of step 4102, reference may be made to the optional implementations of step 2105 in FIG. 2 and other related parts of the embodiments involved in the steps of FIG. 2 .

[0203] The communication method involved in the embodiment of the present disclosure may include at least one of steps 4101 to 4102. For example, step 4102 may be implemented as an independent embodiment, and steps 4101+4102 may be implemented as independent embodiments, but are not limited thereto.

[0204] In this embodiment or example, unless there is any contradiction, each step can be independent, arbitrarily combined or exchanged in order, and the optional methods or optional examples can be arbitrarily combined and can be arbitrarily combined with other embodiments or examples.

[0205] FIG5 is a flow chart of an audio signal processing method according to an embodiment of the present disclosure. As shown in FIG5 , the embodiment of the present disclosure relates to an audio signal processing method for a communication system including a first device and a second device. The method includes:

[0206] Step 5101: The first device determines a target processing mode for processing an audio signal based on decision indication information.

[0207] For optional implementations of step 5101, reference may be made to the optional implementations of step 2101 and step 2102 in FIG. 2 , step 3101 in FIG. 3 , and other related parts in the embodiments involved in the steps of FIG. 2 and FIG. 3 .

[0208] Step 5102: The first device processes the audio signal in a target processing mode to obtain a target audio signal.

[0209] For optional implementations of step 5102, reference may be made to the optional implementations of step 2103 in FIG. 2 , step 3102 in FIG. 3 , and other related parts in the embodiments involved in the steps of FIG. 2 and FIG. 3 .

[0210] Step 5103: The first device sends the target audio signal to the second device.

[0211] For optional implementations of step 5103, reference may be made to the optional implementations of step 2104 in FIG. 2 , step 3103 in FIG. 3 , step 4101 in FIG. 4 , and other related parts in the embodiments involved in the steps of FIG. 2 , FIG. 3 , and FIG. 4 .

[0212] Step 5104: The second device processes the target audio signal in the target processing mode.

[0213] For optional implementations of step 5104, reference may be made to the optional implementations of step 2105 in FIG. 2 , step 4102 in FIG. 4 , and other related parts in the embodiments involved in the steps of FIG. 2 and FIG. 4 .

[0214] The method shown in the embodiment of the present disclosure relates to a method for processing audio signals of a terminal device.

[0215] The specific contents of this method are as follows.

[0216] This example provides three audio processing modes to choose from on the receiving end:

[0217] Mode 1: The terminal device completes decoding, rendering, and stereo encoding, and the playback device completes stereo decoding and playback.

[0218] Mode 2: The terminal device completes decoding, (downmixing) and encoding, and the playback device completes decoding, rendering and playback.

[0219] Mode 3: The terminal device completes decoding, pre-rendering, and intermediate variable encoding, and the playback device completes intermediate variable decoding, rendering, and playback.

[0220] "Decision indication information" is output based on the degree of freedom (DOF), image precision, ambience fidelity and other information selected by the terminal device user. The terminal device (automatic selection mode) or the user (manual selection mode) selects one of the three modes as a solution based on the "decision indication information".

[0221] The full text of the method is as follows.

[0222] In this example, the receiving end determines the decoding rendering mode based on the "decision indication information". The decision input information and decision basis of the decision indication information are as follows:

[0223] The judgment input information includes degrees of freedom, and the degrees of freedom include 0 degrees of freedom, 3 degrees of freedom, 3+ degrees of freedom, and 6 degrees of freedom.

[0224] The number of degrees of freedom (DOF) determines the latency requirement for the solution. Latency here refers to the time between the end device and playback device receiving sensor-generated user posture, rotation, and other information and outputting rendered binaural signals. Generally, the greater the degree of freedom, the lower the latency requirement. For example, in a 6-DOF scenario, the lowest possible latency is required to ensure that users receive an audio experience that matches their posture and rotation information.

[0225] Among them, the delay of mode 1 solution is the longest, the delay of mode 2 solution is the shortest, and the delay of mode 3 solution is in the middle.

[0226] The decision input information also includes audio and video accuracy. For example, the audio and video accuracy can be divided into: the first dimension is: single audio and video scene and multiple audio and video scenes; the second dimension is: high accuracy, medium accuracy, and low accuracy.

[0227] In a single-audio scenario, because the number of channels is small (for example, sound is collected based on object signals), the decoding and rendering solutions are relatively simple, so the second mode solution can be preferred.

[0228] In multi-audio scenarios, because there are many channels (for example, using sound acquisition based on channel signals (5.1.4)), the decoding and rendering solutions are relatively complex, so Mode 1 solution is preferred.

[0229] For multi-sound image scenes with fewer (such as 2 or 3) sound images (such as using MASA format to collect sound), the decoding and rendering complexity is moderate, so Mode 3 solution can be selected.

[0230] In the second dimension, when the required accuracy is low, the mode 2 solution can be selected, when the accuracy is high, the mode 1 solution can be selected, and when the accuracy is medium, the mode 3 solution can be selected.

[0231] The judgment input information also includes background clarity, which can be divided into: clean background, noisy background, and medium background.

[0232] When the background is clean and the number of input channels is small, the mode 2 solution can be selected. When the background is noisy and the number of input channels is large, the mode 1 solution can be selected. When the background is medium, the mode 3 solution can be selected.

[0233] When a single parameter is used as the basis for decision making, the principles of decision making are summarized as follows:

[0234] The decision flow chart is shown in FIG6 , where the decision indication information is the specified mode X.

[0235] The three available modes provided by the terminal device and the playback device are as follows:

[0236] Mode 1: The terminal device completes decoding, rendering and stereo encoding of the input code stream signal, and the playback device completes stereo decoding and playback of the binaural signal code stream.

[0237] Mode 2: The terminal device completes decoding, (downmixing) and encoding, and the playback device completes decoding + rendering and playback.

[0238] Mode 3: The terminal device completes decoding, pre-rendering, and intermediate variable encoding, and the playback device completes intermediate variable decoding, rendering, and playback.

[0239] The method proposed in this example is described in detail below through embodiments.

[0240] The three parameters of the indicator information decider in the terminal device (degrees of freedom, audio and video accuracy, and background clarity) are each assigned a weighting factor. The indicator information decider outputs the decision indicator information as the sum of the weighting factors. The mode component of the mode X solution is then selected based on the sum of the weighting factors. For example, the weighting factors can be set as follows.

[0241] Then the final weight value is Y=Function(X1, X2, X3);

[0242] Wherein, X1 refers to the parameter value of the degree of freedom parameter, X2 refers to the parameter value of the sound image accuracy parameter, X3 refers to the parameter value of the background clarity parameter, and Y refers to decision indication information.

[0243] For example, if Y is less than or equal to 3, then the mode 1 solution is selected; if Y is greater than 3 and less than 8, then the mode 3 solution is selected; if Y is greater than or equal to 8, then the mode 2 solution is selected.

[0244] For example, the scenario in which the user currently holds the terminal device is that he is sitting in a public welfare training classroom facing the podium and listening to the training teacher talking about the public welfare training content. At the same time, he calls a remote user sitting at home on the other end to share the content of the public welfare training speech. The usage requirement of the remote user on the other end is to be able to clearly understand the training content of the training teacher. As shown in Figure 7, at this time, the input of the remote user to the indication information decider is: 0 degrees of freedom, high precision of single audio and video scene, and pure background.

[0245] At this time, the decision indicator shows that the solution of selection mode 1 is selected.

[0246] In summary, the above examples of the present disclosure can determine the decision indication information based on the indication information decision maker, and based on the decision indication information, the terminal device or user selects one of the three modes to form a solution, thereby determining the appropriate working mode based on the indication information.

[0247] FIG8a is a schematic diagram of the structure of the first device 101 proposed in an embodiment of the present disclosure. As shown in FIG8a , the first device 101 includes: a processing module 8101, which is used to determine a target processing mode from multiple processing modes based on decision indication information; the processing module is also used to process the audio signal in the target processing mode to obtain a target audio signal; optionally, the processing module is used to execute at least one of the processing steps (such as step 2101, step 2102, etc., but not limited thereto) executed by the first device 101 in any of the above methods, which will not be repeated here. Among them, the multiple processing modes are different processing functions performed based on the first device as a carrier.

[0248] The first device 101 also includes: a transceiver module 8102, which is used to send the target audio signal to the second device; optionally, the above-mentioned transceiver module is used to execute at least one of the sending and / or receiving steps (for example, step 2103) performed by the first device 101 in any of the above methods, which will not be repeated here.

[0249] FIG8b is a schematic diagram of the structure of the second device 102 proposed in an embodiment of the present disclosure. As shown in FIG8b , the second device 102 includes: a transceiver module 8201 for receiving a target audio signal sent by the first device, where the target audio signal is obtained by the first device processing the audio signal in a target processing mode, and the target processing mode is determined from a plurality of processing modes according to the decision indication information; optionally, the transceiver module can be used to execute at least one of the steps of sending and / or receiving (e.g., step 2103) executed by the second device 102 in any of the above methods, which will not be described in detail here. Among them, the multiple processing modes are different processing functions performed based on the first device as the carrier.

[0250] The second device 102 also includes: a processing module 8202, which is used to process the target audio signal in a target processing mode; optionally, the above-mentioned processing module can be used to execute at least one of the processing steps (such as step 2104) performed by the second device 102 in any of the above methods, which will not be repeated here.

[0251] As shown in Figure 9a, the communication device 9100 includes one or more processors 9101. The processor 9101 can be a general-purpose processor or a dedicated processor, for example, a baseband processor or a central processing unit. The baseband processor can be used to process communication protocols and communication data, and the central processing unit can be used to control the communication device (such as a base station, baseband chip, terminal device, terminal device chip, DU or CU, etc.), execute programs, and process program data. The processor 9101 is used to call instructions to enable the communication device 9100 to perform any of the above methods.

[0252] In some embodiments, the communication device 9100 further includes one or more memories 9102 for storing instructions. Optionally, all or part of the memories 9102 may be located outside the communication device 9100.

[0253] In some embodiments, the communication device 9100 further includes one or more transceivers 9103. When the communication device 9100 includes one or more transceivers 9103, the communication steps such as sending and receiving in the above method are performed by the transceiver 9103, and the other steps are performed by the processor 9101.

[0254] In some embodiments, a transceiver may include a receiver and a transmitter, which may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, and transceiver circuit may be used interchangeably; the terms transmitter, transmitting unit, transmitter, and transmitting circuit may be used interchangeably; and the terms receiver, receiving unit, receiver, and receiving circuit may be used interchangeably.

[0255] Optionally, the communication device 9100 further includes one or more interface circuits 9104, which are connected to the memory 9102. The interface circuits 9104 can be used to receive signals from the memory 9102 or other devices, and can be used to send signals to the memory 9102 or other devices. For example, the interface circuits 9104 can read instructions stored in the memory 9102 and send the instructions to the processor 9101.

[0256] The communication device 9100 described in the above embodiments may be a network device or a terminal, but the scope of the communication device 9100 described in the present disclosure is not limited thereto, and the structure of the communication device 9100 may not be limited by FIG. 9a. The communication device may be an independent device or may be part of a larger device. For example, the communication device may be: 1) an independent integrated circuit IC, or a chip, or a chip system or subsystem; (2) a collection of one or more ICs, optionally, the above IC collection may also include a storage component for storing data or programs; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, a terminal device, an intelligent terminal device, a cellular phone, a wireless device, a handheld device, a mobile unit, an in-vehicle device, a network device, a cloud device, an artificial intelligence device, etc.; (6) others, etc.

[0257] FIG9b is a schematic diagram of the structure of a chip 9200 according to an embodiment of the present disclosure. If the communication device 9100 can be a chip or a chip system, reference can be made to the schematic diagram of the structure of the chip 9200 shown in FIG9b , but the present disclosure is not limited thereto.

[0258] The chip 9200 includes one or more processors 9201, and the processor 9201 is used to call instructions so that the chip 9200 executes any of the above methods.

[0259] In some embodiments, chip 9200 further includes one or more interface circuits 9202, which are connected to memory 9203. Interface circuits 9202 can be used to receive signals from memory 9203 or other devices, and can be used to send signals to memory 9203 or other devices. For example, interface circuit 9202 can read instructions stored in memory 9203 and send the instructions to processor 9201. Optionally, the terms interface circuit, interface, transceiver pin, and transceiver are interchangeable.

[0260] In some embodiments, the chip 9200 further includes one or more memories 9203 for storing instructions. Alternatively, all or part of the memories 9203 may be located outside the chip 9200.

[0261] The present disclosure also proposes a storage medium having instructions stored thereon, which, when executed on the communication device 9100, causes the communication device 9100 to execute any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but is not limited thereto and may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but is not limited thereto and may also be a temporary storage medium.

[0262] The present disclosure also provides a program product, which, when executed by the communication device 9100, enables the communication device 9100 to perform any of the above methods. Optionally, the program product is a computer program product.

[0263] The present disclosure also proposes a computer program, which, when executed on a computer, causes the computer to perform any one of the above methods.

[0264] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer programs. When the computer program is loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer program can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a high-density digital video disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).

[0265] The correspondences shown in the tables of the present disclosure can be configured or predefined. The values ​​of the information in each table are merely examples and can be configured to other values, which are not limited by the present disclosure. When configuring the correspondences between information and parameters, it is not necessarily required to configure all the correspondences shown in each table. For example, in the tables of the present disclosure, the correspondences shown in certain rows may not be configured. For another example, appropriate deformation adjustments can be made based on the above tables, such as splitting, merging, etc. The names of the parameters shown in the titles of the above tables may also adopt other names that can be understood by the communication device, and the values ​​or representations of the parameters may also adopt other values ​​or representations that can be understood by the communication device. When implementing the above tables, other data structures may also be used, such as arrays, queues, containers, stacks, linear lists, pointers, linked lists, trees, graphs, structures, classes, heaps, hash tables or hash tables, etc.

[0266] The predefined in the present disclosure may be understood as defined, predefined, stored, pre-stored, pre-negotiated, pre-configured, solidified, or pre-burned.

[0267] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0268] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0269] The above description is merely a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.

Claims

1. A method for processing an audio signal, characterized in that: The method is performed by a first device, and includes: determining a target processing mode from a plurality of processing modes according to the decision indication information; Processing the audio signal in the target processing mode to obtain a target audio signal; sending the target audio signal to a second device; The multiple processing modes are based on different processing functions performed by the first device as a carrier.

2. The method according to claim 1, characterized in that The method further comprises: The decision indication information is determined based on first information, where the first information is used to represent at least one of a delay requirement, an accuracy characteristic, and a listener freedom characteristic of the audio signal.

3. The method according to claim 2, characterized in that The first information includes at least one of the following: a first parameter, where the first parameter is used to identify a degree of freedom, where the degree of freedom is information about a listener's position movement and / or posture change received by a sensor configured by the first device and / or the second device; a second parameter, the second parameter being used to identify an audiovisual scene and / or rendering precision, the audiovisual scene being the number of audio and video included in the audio signal, and the rendering precision being the precision required for rendering the audio signal; a third parameter, the third parameter being used to identify the background clarity of the audio signal; A fourth parameter is used to identify a delay, where the delay is the time between when the second device sends the first parameter to the first device and when the second device receives the target audio signal sent by the first device.

4. The method according to any one of claims 1 to 3, characterized in that The target processing mode includes at least one of the following: a first processing mode, wherein the first device performs code stream decoding, rendering processing, and stereo encoding on the audio signal, and after the first device sends the target audio signal to the second device, the second device performs stereo decoding and playback on the target audio signal; a second processing mode, wherein the first device performs code stream decoding, downmixing and / or format conversion, and code stream encoding on the audio signal, and after the first device sends the target audio signal to the second device, the second device performs code stream decoding, rendering, and playback on the target audio signal; The third processing mode is that the first device performs code stream decoding, first rendering processing, and intermediate variable encoding on the audio signal. After the first device sends the target audio signal to the second device, the second device performs intermediate variable decoding, second rendering processing, and plays the target audio signal.

5. The method according to any one of claims 1 to 4, characterized in that Determining the target processing mode according to the decision indication information includes: Determine the processing mode in which the decision indication information satisfies a preset condition as the target processing mode; or Instruction information is received, and the target processing mode is determined according to the instruction information, where the instruction information includes the target processing mode determined by the user of the first device according to the decision instruction information.

6. The method according to claim 5, characterized in that Determining the processing mode in which the decision indication information satisfies a preset condition as the target processing mode includes: When the decision indication information satisfies a first condition, determining that the target processing mode is a first processing mode; When the decision indication information satisfies a second condition, determining that the target processing mode is the second processing mode; When the decision indication information satisfies a third condition, the target processing mode is determined to be a third processing mode.

7. The method according to any one of claims 2 to 5, characterized in that The determining the decision indication information based on the first information includes: When the first parameter in the first information indicates that the degree of freedom is 0 and / or the fourth parameter indicates that the delay is low, the first parameter is determined to be a first preset value; when the first parameter indicates that the degree of freedom is greater than or equal to 3 and less than 6 and / or the fourth parameter indicates that the delay is medium, the first parameter is determined to be a second preset value; when the first parameter indicates that the degree of freedom is 6 and / or the fourth parameter indicates that the delay is high, the first parameter value is determined to be a third preset value; When the second parameter in the first information indicates that the audiovisual scene is a single audiovisual scene and the rendering precision is high precision, the second parameter is determined to be a fourth preset value; when the second parameter indicates that the audiovisual scene is a single audiovisual scene and the rendering precision is medium precision, the second parameter is determined to be a fifth preset value; when the second parameter indicates that the audiovisual scene is a single audiovisual scene and the rendering precision is low precision, the second parameter is determined to be a fifth preset value. when the second parameter indicates that the sound-image scene is a multi-sound-image scene and the rendering precision is high precision, the second parameter is determined to be the seventh preset value; when the second parameter indicates that the sound-image scene is a multi-sound-image scene and the rendering precision is medium precision, the second parameter is determined to be the eighth preset value; when the second parameter indicates that the sound-image scene is a multi-sound-image scene and the rendering precision is low precision, the second parameter is determined to be the ninth preset value; When the third parameter in the first information indicates that the background definition is a pure background, the third parameter is determined to be a tenth preset value; when the third parameter indicates that the background definition is a noisy background, the third parameter is determined to be an eleventh preset value; when the third parameter indicates that the background definition is a medium background, the third parameter is determined to be a twelfth preset value; Based on a preset function, with the first parameter, the second parameter, and the third parameter as independent variables, a function value of the preset function is determined as the decision indication information.

8. The method according to claim 7, characterized in that The step of determining, based on a preset function and taking the first parameter, the second parameter, and the third parameter as independent variables, a function value of the preset function as the decision indication information includes: Setting weight factors for the first parameter, the second parameter, and the third parameter respectively; Performing a weighted summation on the first parameter, the second parameter, the third parameter, and the weight factor to obtain the decision indication information.

9. The method according to any one of claims 1 to 8, characterized in that Processing the audio signal in the target processing mode to obtain a target audio signal includes: When the target processing mode is the first processing mode, performing code stream decoding, rendering processing, and stereo encoding on the audio signal to obtain a binaural audio signal as the target audio signal; When the target processing mode is the second processing mode, performing code stream decoding, down-mixing, and code stream encoding on the audio signal to obtain an encoded audio signal as the target audio signal; When the target processing mode is the third processing mode, the audio signal is subjected to code stream decoding, first rendering processing, and intermediate variable encoding to obtain a pre-rendered audio signal as the target audio signal.

10. An audio signal processing method, characterized in that: The method is performed by a second device, and includes: receiving a target audio signal sent by a first device, where the target audio signal is obtained by the first device processing an audio signal in a target processing mode, where the target processing mode is determined from a plurality of processing modes according to decision indication information; processing the target audio signal in the target processing mode; The multiple processing modes are based on different processing functions performed by the first device as a carrier.

11. The method according to claim 10, characterized in that The first information includes at least one of the following: a first parameter, where the first parameter is used to identify a degree of freedom, where the degree of freedom is information about a listener's position movement and / or posture change received by a sensor configured by the first device and / or the second device; a second parameter, the second parameter being used to identify an audiovisual scene and / or rendering precision, the audiovisual scene being the number of audio and video included in the audio signal, and the rendering precision being the precision required for rendering the audio signal; a third parameter, the third parameter being used to identify the background clarity of the audio signal; A fourth parameter is used to identify a delay, where the delay is a time between when the first device receives the first parameter from the second device and when the first device sends the target audio signal to the second device.

12. The method according to claim 10 or 11, characterized in that The target processing mode includes at least one of the following: a first processing mode, wherein the first device performs code stream decoding, rendering processing, and stereo encoding on the audio signal, and after the first device sends the target audio signal to the second device, the second device performs stereo decoding and playback on the target audio signal; a second processing mode, wherein the first device performs code stream decoding, downmixing or format conversion, and code stream encoding on the audio signal, and after the first device sends the target audio signal to the second device, the second device performs code stream decoding, rendering, and playback on the target audio signal; The third processing mode is that the first device performs code stream decoding, first rendering processing, and intermediate variable encoding on the audio signal. After the first device sends the target audio signal to the second device, the second device performs intermediate variable decoding, second rendering processing, and plays the target audio signal.

13. The method according to any one of claims 10 to 12, characterized in that When the decision indication information satisfies a first condition, the target processing mode is the first processing mode; when the decision indication information satisfies a second condition, the target processing mode is the second processing mode; When the decision indication information satisfies a third condition, the target processing mode is a third processing mode, and the decision indication information is determined based on the first information.

14. The method according to claim 13, characterized in that The decision indication information is the function value of a preset function, and the independent variables of the preset function are the first parameter, the second parameter, and the third parameter in the first information. Among them, when the first parameter indicates that the degree of freedom is 0 and / or the fourth parameter indicates that the delay is low, the first parameter is a first preset value; when the first parameter indicates that the degree of freedom is greater than or equal to 3 and less than 6 and / or the fourth parameter indicates that the delay is medium, the first parameter is a second preset value; when the first parameter indicates that the degree of freedom is 6 and / or the fourth parameter indicates that the delay is high, the first parameter value is a third preset value; When the second parameter indicates that the sound image scene is a single sound image scene and the rendering precision is high precision, the second parameter is a fourth preset value; when the second parameter indicates that the sound image scene is a single sound image scene and the rendering precision is medium precision, the second parameter is a fifth preset value; when the second parameter indicates that the sound image scene is a single sound image scene and the rendering precision is low precision, the second parameter is a sixth preset value; when the second parameter indicates that the sound image scene is a multi-sound image scene and the rendering precision is high precision, the second parameter is a seventh preset value; when the second parameter indicates that the sound image scene is a multi-sound image scene and the rendering precision is medium precision, the second parameter is an eighth preset value; when the second parameter indicates that the sound image scene is a multi-sound image scene and the rendering precision is low precision, the second parameter is a ninth preset value; When the third parameter indicates that the background clarity is a pure background, the third parameter is the tenth preset value; when the third parameter indicates that the background clarity is a noisy background, the third parameter is the eleventh preset value; when the third parameter indicates that the background clarity is a medium background, the third parameter is the twelfth preset value.

15. The method according to any one of claims 10 to 14, characterized in that Processing the target audio signal in the target processing mode includes: When the target processing mode is the first processing mode, performing stereo decoding and playing on the target audio signal; When the target processing mode is the second processing mode, decoding, rendering and playing the target audio signal; When the target processing mode is the third processing mode, intermediate variable decoding, second rendering processing, and playback are performed on the target audio signal.

16. A first device, characterized in that: include: A processing module, configured to determine a target processing mode from a plurality of processing modes according to the decision indication information; The processing module is further configured to process the audio signal in the target processing mode to obtain a target audio signal; a transceiver module, configured to send the target audio signal to a second device; The multiple processing modes are based on different processing functions performed by the first device as a carrier.

17. A second device, characterized in that: include: a transceiver module, configured to receive a target audio signal sent by a first device, wherein the target audio signal is obtained by the first device processing an audio signal in a target processing mode, wherein the target processing mode is determined from a plurality of processing modes according to the decision indication information; a processing module, configured to process the target audio signal in the target processing mode; The multiple processing modes are based on different processing functions performed by the first device as a carrier.

18. A communication device, wherein: include: transceiver; Memory; A processor is connected to the transceiver and the memory respectively, and is configured to control the wireless signal reception and transmission of the transceiver by executing computer-executable instructions on the memory, and can implement any one of the methods of claims 1-15.

19. A computer storage medium, wherein: The computer storage medium stores computer-executable instructions; after the computer-executable instructions are executed by the processor, the method according to any one of claims 1 to 15 can be implemented.

20. A communication system, characterized in that: The method comprises a first device and a second device, wherein the first device is configured to execute the method according to any one of claims 1 to 9, and the second device is configured to execute the method according to any one of claims 10 to 15.

Citation Information

Patent Citations

  • Audio signal rendering method and device

    CN114067810A

  • Adaptive audio transmission and rendering

    CN115701777A

  • Audio signal processing method and device

    CN116830600A

  • XR object rendering method, communication device and system

    CN117411944A

  • Method and apparatus for transmitting 3D XR media data

    US20220028172A1