Audio processing method, electronic device, and computer readable storage medium

By performing semantic analysis of call audio data and dynamically adjusting audio processing strategies, the problem of poor call quality in noisy environments is solved, and a better call experience and audio clarity is achieved.

WO2025180410A1PCT designated stage Publication Date: 2025-09-04HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/079322
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-29
Filing Date
2025-02-26
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

When making a call in a noisy environment, the call quality is poor and the user cannot clearly hear the other party's voice.

Method used

By performing semantic analysis of call audio data, the uplink audio processing strategies of electronic devices are dynamically adjusted, such as noise reduction intensity, audio gain, and noise reduction scenarios, to improve call quality.

Benefits of technology

Improve call quality, enable users to have a better call experience, reduce noise interference, and enhance the clarity of audio data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025079322_04092025_PF_FP_ABST
    Figure CN2025079322_04092025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides an audio processing method, an electronic device, and a computer readable storage medium. The method comprises: during a call on a first electronic device and a second electronic device, performing semantic analysis on call audio data to obtain a target semantic analysis result, wherein the call audio data comprises call audio data from the second electronic device; and on the basis of the target semantic analysis result, adjusting an audio processing strategy, wherein the audio processing strategy comprises an uplink audio processing strategy, and the uplink audio processing strategy is used for processing call audio data collected by the first electronic device. According to embodiments of the present application, during a call, by performing semantic analysis on call audio data from a peer side, an uplink audio processing strategy of an electronic device can be dynamically adjusted, so that the call quality can be improved, and call experience of users can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Audio processing method, electronic device, and computer-readable storage medium

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on February 29, 2024, with application number 202410233300.8 and application name “Audio Processing Method, Electronic Device and Computer-readable Storage Medium”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of terminal technology, and in particular to an audio processing method, an electronic device, and a computer-readable storage medium. Background Art

[0003] As the processing capabilities of mobile phones, tablets and other terminal devices become increasingly powerful, audio processing technologies based on audio noise reduction have also been widely used in mobile phones, tablets and other terminal devices.

[0004] When users use their mobile phones to make daily calls, one or both parties may be in a noisy environment (such as a shopping mall, high-speed rail station, airport, etc. with a lot of traffic), which makes it difficult for one party to hear what the other party is saying, resulting in poor call quality. Summary of the Invention

[0005] The embodiments of the present application disclose an audio processing method, an electronic device, and a computer-readable storage medium. During a call, by performing semantic analysis on the call audio data from the other end, the uplink audio processing strategy of the electronic device can be dynamically adjusted, thereby improving the call quality and allowing users to have a better call experience.

[0006] The first aspect discloses an audio processing method, which can be applied to a first electronic device, or to a module (e.g., a processor) in the first electronic device, or to a logic module or software that can realize all or part of the functions of the first electronic device. The following description is given by taking the application to the first electronic device as an example. The communication method may include: during a call between the first electronic device and the second electronic device, performing semantic analysis on the call audio data to obtain a target semantic analysis result; the call audio data includes call audio data from the second electronic device; adjusting the audio processing strategy based on the target semantic analysis result; wherein the audio processing strategy includes an uplink audio processing strategy, and the uplink audio processing strategy is used to process the call audio data collected by the first electronic device.

[0007] In an embodiment of the present application, during a call, the first electronic device can perform semantic analysis on the call audio data from the second electronic device (such as "Your voice is so quiet", "It's too noisy over there, I can't hear you clearly") to obtain a semantic analysis result. Based on the semantic analysis result, the first electronic device can understand the actual call experience of the user of the other end (the second electronic device) or the call audio quality of the local end (the first electronic device), such as whether the noise on the local end is louder, whether the sound on the local end is quieter, etc. Therefore, the first electronic device can dynamically adjust the uplink audio processing strategy of the local end based on the semantic analysis result (such as selecting a higher level of noise reduction intensity when the other end feedbacks that the noise on the local end is louder). Afterwards, the first electronic device can process the call audio data collected by the first electronic device through the adjusted uplink audio processing strategy, which can make the audio quality of the call audio data sent to the second electronic device better, so that the user on the other end can get a better call experience.

[0008] With reference to the first aspect, in a possible implementation, the uplink audio processing strategy includes one or more of uplink noise reduction strength, uplink audio gain, and uplink noise reduction scenario.

[0009] In an embodiment of the present application, for the uplink audio processing strategy, multiple options such as uplink noise reduction intensity, uplink audio gain, and uplink noise reduction scenario may be included. Among them, the uplink noise reduction intensity can be used to adjust the degree of noise elimination, and can be used to meet the noise reduction needs of different users and various scenarios. For example, a weak noise reduction intensity can eliminate a small part of the noise, a medium noise reduction intensity can eliminate most of the noise, and a strong noise reduction intensity can completely eliminate the noise. The uplink audio gain can be used to adjust the amplitude of the audio data transmitted to the second electronic device side, that is, the volume can be adjusted. By adjusting the uplink audio gain, the user on the other end can hear more clearly. As for the noise reduction scenario, different noise reduction scenarios can correspond to different noise reduction models. By adjusting the appropriate noise reduction scenario, the best noise reduction effect can be obtained, and the impact on the target voice can be minimized while reducing the noise.

[0010] In combination with the first aspect, in a possible implementation, the call audio data further includes call audio data collected by the first electronic device.

[0011] In an embodiment of the present application, the call audio data also includes the call audio data collected by the first electronic device, that is, the call audio data of the local end. In this way, the first electronic device can obtain a more complete and accurate semantic analysis result based on the conversation context (the conversation between the local user and the opposite user), thereby ensuring that the uplink noise reduction strategy can be adjusted more accurately.

[0012] With reference to the first aspect, in a possible implementation, the audio processing strategy further includes a downlink audio processing strategy.

[0013] With reference to the first aspect, in a possible implementation, the downlink audio processing strategy includes one or more of downlink noise reduction strength, downlink audio gain, and downlink noise reduction scenario.

[0014] In an embodiment of the present application, the first electronic device can also dynamically adjust the downlink audio processing strategy based on the semantic analysis results of the call audio data, such as the noise reduction scene, noise reduction intensity, etc., so that the local user can obtain a better call experience.

[0015] In combination with the first aspect, in a possible implementation, adjusting the audio processing strategy based on the target semantic analysis result includes: when the target semantic analysis result is one of multiple preset semantic analysis results, adjusting the audio processing strategy based on the adjustment strategy corresponding to the target semantic result; the multiple preset semantic analysis results are all configured with corresponding adjustment strategies, and the adjustment strategy is used to indicate the parameters in the audio processing strategy that need to be adjusted and the adjustment method.

[0016] In an embodiment of the present application, multiple preset semantic analysis results and adjustment strategies corresponding to each preset semantic analysis result can be configured in advance. In this way, it can be ensured that the first electronic device can efficiently and accurately adjust the uplink audio processing strategy based on the semantic analysis results of the call audio data.

[0017] In combination with the first aspect, in a possible implementation, the method further includes: initiating positioning to obtain a first positioning result; and adjusting an uplink noise reduction scenario based on the first positioning result, where different noise reduction scenarios correspond to different noise reduction models.

[0018] In the embodiment of the present application, the uplink noise reduction scenario in which the local end is currently located can be accurately determined through positioning, so that the most appropriate noise reduction model can be used for noise reduction processing, thereby ensuring the best noise reduction effect.

[0019] In combination with the first aspect, in a possible implementation, the method further includes: receiving first indication information from the second electronic device, the first indication information being used to indicate an audio quality problem of the first electronic device; and adjusting an uplink audio processing strategy based on the first indication information.

[0020] In an embodiment of the present application, the first electronic device may further adjust the uplink audio processing strategy based on the first indication information from the second electronic device to improve the quality of the call audio, so that the opposite-end user may obtain a better call experience.

[0021] In combination with the first aspect, in a possible implementation, adjusting the uplink audio processing strategy based on the first indication information includes: adjusting the uplink audio processing strategy based on the adjustment strategy corresponding to the first indication information; the first indication information includes multiple value situations, and the multiple value situations are all configured with corresponding adjustment strategies, and the adjustment strategy is used to indicate the parameters in the audio processing strategy that need to be adjusted and the adjustment method.

[0022] In an embodiment of the present application, multiple possible scenarios can be determined in advance, each scenario can correspond to a value of indication information, and a corresponding adjustment strategy can be pre-configured for each value of indication information. In this way, it can be ensured that the first electronic device can efficiently and accurately adjust the uplink audio processing strategy based on the first indication information.

[0023] In combination with the first aspect, in a possible implementation, before adjusting the audio processing strategy based on the target semantic analysis result, the method also includes: performing semantic analysis on the sixth audio data to obtain a sixth semantic analysis result; the sixth audio data includes call audio data from the second electronic device and / or call audio data collected by the first electronic device; when the sixth semantic analysis result is one of multiple preset semantic analysis results, turning on the smart call mode; adjusting the audio processing strategy based on the target semantic analysis result includes: when the smart call mode is turned on, adjusting the audio processing strategy based on the target semantic analysis result.

[0024] In the embodiment of the present application, since in most cases the first electronic device is usually in an environment with little or no noise, the smart call mode can be turned off by default each time the first electronic device makes a call. In this case, the first electronic device does not need to perform noise reduction and other processing. During the call, the first electronic device can perform semantic analysis on the call audio data from the second electronic device. When it is determined based on the semantic analysis that there is a problem with the call audio quality (such as the noise on the local end is louder or the sound on the local end is quieter), the first electronic device can turn on the smart call mode again and perform noise reduction and other processing. It can be seen that in this way, the first electronic device can turn on the smart call mode for noise reduction and other processing only when it is determined based on the semantic analysis results that specific conditions are met. In this way, meaningless noise reduction processing of the first electronic device can be avoided as much as possible, and the overall power consumption of the first electronic device can be reduced.

[0025] In combination with the first aspect, in a possible implementation, before adjusting the audio processing strategy based on the target semantic analysis result, the method also includes: receiving a first request message from the second electronic device, the first request message being used to request activation of the smart call mode; activating the smart call mode based on the first request message; and adjusting the audio processing strategy based on the target semantic analysis result including: adjusting the audio processing strategy based on the target semantic analysis result when the smart call mode is activated.

[0026] In combination with the first aspect, in a possible implementation, the method further includes: receiving a second request message from the second electronic device, the second request message being used to request adjustment of the uplink audio processing strategy, the second request message including parameters in the uplink audio processing strategy that need to be adjusted and corresponding parameter values; adjusting the uplink audio processing strategy based on the second request message.

[0027] In an embodiment of the present application, the first electronic device can also directly enable the smart call mode based on the request message from the second electronic device, or directly adjust the uplink audio processing strategy, which can improve the flexibility of enabling the smart call mode and the flexibility of adjusting the uplink audio processing strategy. For example, the second electronic device can provide a corresponding control for requesting the other party to enable the smart call mode, and a control for requesting the other party to adjust the uplink audio processing strategy. In this way, the user of the second electronic device can request the other party to enable the smart call mode or adjust the uplink audio processing strategy based on the actual call experience, so as to improve the call quality and obtain a better call experience.

[0028] In combination with the first aspect, in a possible embodiment, the method also includes: displaying a first user interface, the first user interface including an audio processing configuration control; displaying a second user interface in response to a user operation acting on the audio processing configuration control; the second user interface including an upstream audio processing configuration area, the upstream audio processing configuration area including one or more of a noise reduction intensity configuration area, an audio gain configuration area, and a noise reduction scene configuration area, the noise reduction intensity configuration area including multiple noise reduction intensity options, the audio gain configuration area including multiple audio gain options, and the noise reduction scene configuration area including multiple noise reduction scene options; in response to a user operation acting on the first option, changing the first option to a selected state and enabling the audio processing strategy corresponding to the first option; the first option is any option in the second user interface.

[0029] In the embodiment of the present application, the first electronic device may further provide a configuration interface corresponding to the uplink audio processing strategy, which may facilitate manual adjustment by the user and may further improve the flexibility of adjusting the uplink audio processing strategy.

[0030] In combination with the first aspect, in a possible implementation, the method also includes: receiving fourth audio data from the second electronic device; processing the fourth audio data based on the current downlink audio processing strategy to obtain fifth audio data; analyzing the fifth audio data to determine the audio quality problem of the fifth audio data; and adjusting the downlink audio processing strategy based on the audio quality problem of the fifth audio data.

[0031] In an embodiment of the present application, the first electronic device can also analyze the audio data processed by the local downlink audio processing strategy, and then adjust the downlink audio processing strategy based on the analysis results, so as to ensure the accuracy of the downlink audio processing strategy and enable the local user to obtain a better call experience.

[0032] The second aspect discloses an audio processing method, which can be applied to a first electronic device, or to a module (e.g., a processor) in the first electronic device, or to a logic module or software that can implement all or part of the functions of the first electronic device. The following description is given using the application to the first electronic device as an example. The communication method may include: during a call between the first electronic device and the second electronic device, performing semantic analysis on the call audio data to obtain a target semantic analysis result; the call audio data includes call audio data collected by the first electronic device; adjusting the audio processing strategy based on the target semantic analysis result; the audio processing strategy includes a downlink audio processing strategy, and the downlink audio processing strategy is used to process the call audio data from the second electronic device.

[0033] With reference to the second aspect, in a possible implementation, the downlink audio processing strategy includes one or more of downlink noise reduction strength, downlink audio gain, and downlink noise reduction scenario.

[0034] In combination with the second aspect, in a possible implementation, the audio processing strategy further includes an uplink audio processing strategy.

[0035] With reference to the second aspect, in a possible implementation, the uplink audio processing strategy includes one or more of uplink noise reduction strength, uplink audio gain, and uplink noise reduction scenario.

[0036] In combination with the second aspect, in a possible implementation, the call audio data further includes call audio data from the second electronic device.

[0037] It should be noted that the technical solution of the second aspect of the present application may correspond to or be similar to the technical solution of the first aspect, and the relevant beneficial effects can refer to the beneficial effects of the first aspect.

[0038] The third aspect discloses an audio processing method, which can be applied to a first electronic device, or to a module (e.g., a processor) in the first electronic device, or to a logic module or software that can realize all or part of the functions of the first electronic device. The following description is given using the application to the first electronic device as an example. The communication method may include: during a call between the first electronic device and the second electronic device, performing semantic analysis on the call audio data to obtain a target semantic analysis result; the call audio data includes call audio data from the second electronic device; adjusting the audio processing strategy based on the target semantic analysis result; wherein the audio processing strategy includes a downlink audio processing strategy, and the downlink audio processing strategy is used to process the call audio data from the second electronic device.

[0039] In combination with the third aspect, in a possible implementation, the downlink audio processing strategy includes one or more of downlink noise reduction strength, downlink audio gain, and downlink noise reduction scenario.

[0040] In combination with the third aspect, in a possible implementation, the audio processing strategy further includes an uplink audio processing strategy.

[0041] In combination with the third aspect, in a possible implementation, the uplink audio processing strategy includes one or more of uplink noise reduction strength, uplink audio gain, and uplink noise reduction scenario.

[0042] In combination with the third aspect, in a possible implementation, the call audio data also includes call audio data collected by the first electronic device.

[0043] It should be noted that the technical solution of the third aspect of the present application may correspond to or be similar to the technical solution of the first aspect, and the relevant beneficial effects can refer to the beneficial effects of the first aspect.

[0044] The fourth aspect discloses an audio processing method, which can be applied to a first electronic device, or to a module (for example, a processor) in the first electronic device, or to a logic module or software that can realize all or part of the functions of the first electronic device. The following description is given by taking the application to the first electronic device as an example. The communication method may include: during a call between the first electronic device and the second electronic device, performing semantic analysis on the call audio data to obtain a target semantic analysis result; the call audio data includes the call audio data collected by the first electronic device; adjusting the audio processing strategy based on the target semantic analysis result; the audio processing strategy includes an uplink audio processing strategy, and the uplink audio processing strategy is used to process the call audio data collected by the first electronic device.

[0045] In combination with the fourth aspect, in a possible implementation, the uplink audio processing strategy includes one or more of uplink noise reduction strength, uplink audio gain, and uplink noise reduction scenario.

[0046] In combination with the fourth aspect, in a possible implementation, the audio processing strategy also includes a downlink audio processing strategy.

[0047] In combination with the fourth aspect, in a possible implementation, the downlink audio processing strategy includes one or more of downlink noise reduction strength, downlink audio gain, and downlink noise reduction scenario.

[0048] In combination with the fourth aspect, in a possible implementation, the call audio data also includes call audio data from the second electronic device.

[0049] It should be noted that the technical solution of the fourth aspect of the present application may correspond to or be similar to the technical solution of the first aspect, and the relevant beneficial effects can refer to the beneficial effects of the first aspect.

[0050] The fifth aspect discloses an electronic device, which may be a first electronic device, comprising a processor and a communication interface; the communication interface is used to receive and send data; the processor calls a computer program or computer instruction stored in a memory to implement the method provided in the above-mentioned first aspect and any possible implementation of the first aspect, or implements the method provided in the above-mentioned second aspect and any possible implementation of the second aspect, or implements the method provided in the above-mentioned third aspect and any possible implementation of the third aspect, or implements the method provided in the above-mentioned fourth aspect and any possible implementation of the fourth aspect.

[0051] As a possible implementation, the communication device disclosed in the fifth aspect may include one or more processors.

[0052] Optionally, the communication device disclosed in the fifth aspect above also includes one or more memories.

[0053] The sixth aspect discloses a computer-readable storage medium having a computer program or computer instructions stored thereon. When the computer program or computer instructions are executed, the method provided in the first aspect and any possible implementation of the first aspect is implemented, or the method provided in the second aspect and any possible implementation of the second aspect is implemented, or the method provided in the third aspect and any possible implementation of the third aspect is implemented, or the method provided in the fourth aspect and any possible implementation of the fourth aspect is implemented.

[0054] The seventh aspect discloses a chip, including a processor for executing a program stored in a memory. When the program is executed, the chip executes the method provided in the above-mentioned first aspect and any possible implementation of the first aspect, or executes the method provided in the above-mentioned second aspect and any possible implementation of the second aspect, or executes the method provided in the above-mentioned third aspect and any possible implementation of the third aspect, or executes the method provided in the above-mentioned fourth aspect and any possible implementation of the fourth aspect.

[0055] As a possible implementation, the memory is located outside the chip.

[0056] The eighth aspect discloses a computer program product, which includes computer program code. When the computer program code is run, the method provided in the first aspect and any possible implementation of the first aspect is executed, or the method provided in the second aspect and any possible implementation of the second aspect is executed, or the method provided in the third aspect and any possible implementation of the third aspect is executed, or the method provided in the fourth aspect and any possible implementation of the fourth aspect is executed.

[0057] It should be understood that the implementation and beneficial effects of the above-mentioned multiple aspects or any possible implementation methods of the present application can be referenced to each other. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0059] FIG1 is a schematic structural diagram of a communication system provided in an embodiment of the present application;

[0060] FIG2 is a schematic diagram of the hardware structure of an electronic device 100 provided in an embodiment of the present application;

[0061] FIG3 is a schematic diagram of the software structure of an electronic device 100 provided in an embodiment of the present application;

[0062] 4A to 4H are schematic diagrams of scenarios in which some electronic devices 100 provide smart call modes and related audio processing configurations according to embodiments of the present application;

[0063] 5A to 5D are schematic diagrams of scenarios in which an electronic device 100 intelligently adjusts an audio processing strategy according to embodiments of the present application;

[0064] 6A to 6C are schematic diagrams of scenarios in which some electronic devices 100 provide noise reduction scene selection and intelligently adjust audio processing strategies based on positioning results, provided in embodiments of the present application;

[0065] 7A to 7F are schematic diagrams of scenarios in which some electronic devices 100 provide a smart call mode and uplink and downlink-related audio processing configurations according to embodiments of the present application;

[0066] 8A to 8E are schematic diagrams of scenarios in which some electronic devices 200 request the electronic device 100 to enable the smart call mode and adjust the uplink audio processing strategy, provided by embodiments of the present application;

[0067] FIG9 is a flow chart of an audio processing method provided in an embodiment of the present application;

[0068] FIG10 is a flow chart of another audio processing method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0069] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in the text is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0070] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.

[0071] The term "user interface (UI)" in the following embodiments of this application refers to a medium interface for interaction and information exchange between an application (APP) or an operating system (OS) and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. The user interface is a source code written in a specific computer language such as Java and extensible markup language (XML). The interface source code is parsed and rendered on an electronic device and finally presented as content that the user can recognize. The commonly used form of user interface is a graphical user interface (GUI), which refers to a user interface related to computer operations that is displayed in a graphical manner. It can be a visual interface element such as text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, widgets, etc. displayed on the display screen of an electronic device.

[0072] Audio processing technologies such as intelligent noise reduction and intelligent compensation have been widely used in various scenarios in daily life, such as calls, recordings, live broadcasts, etc. For example, during calls, live broadcasts, etc., the local electronic device may be in a noisy environment. At this time, the collected external audio may include not only the target voice, but also environmental noise (such as traffic noise, keyboard tapping, noisy voices, etc.). Therefore, the local electronic device can perform noise reduction processing on the collected external audio in real time to reduce / eliminate the environmental noise, and then send the noise-reduced audio to the other electronic device so that the other user can clearly hear the target voice. In addition, after receiving the audio, the other electronic device can also perform noise reduction processing, compensation processing (such as signal band loss compensation, packet loss signal compensation, etc.) to obtain better call quality, etc.

[0073] The present application provides an audio processing method. Among them, in an audio scenario such as a voice call or a video call, the first electronic device can perform semantic analysis on the received audio data from the second electronic device, and dynamically adjust the audio processing strategy based on the semantic analysis result. For example, if the semantic analysis result of a certain segment of audio data from the second electronic device (such as "It's so noisy on your side", "The noise on your side is so loud", etc.) by the first electronic device is "The current noise is relatively large", then the first electronic device can adjust the current noise reduction intensity, such as changing the noise reduction intensity from "weak" to "medium". After that, if the semantic analysis result of a certain segment of audio data from the second electronic device (such as "It's still so noisy on your side", "The noise on your side is still so loud", etc.) by the first electronic device is still "The current noise is relatively large", then the first electronic device can continue to adjust the current noise reduction intensity, such as changing the noise reduction intensity from "medium" to "strong". Another example, if the semantic analysis result of a certain segment of audio data from the second electronic device (such as "Your voice is so small", "I can't hear what you're saying", etc.) by the first electronic device is "The current voice is relatively small", then the first electronic device can increase the current audio gain, such as changing the audio gain from x1 decibels (dB) to x2 dB, where x1 < x2. Another example, if the semantic analysis result of a certain segment of audio data from the second electronic device (such as "Your voice is so loud", "You can speak a little quieter", etc.) by the first electronic device is "The current voice is relatively large", then the first electronic device can decrease the current audio gain, such as changing the audio gain from x2 dB to x1 dB. It can be seen that in the above method, for audio scenarios such as voice calls and video calls, the first electronic device can dynamically adjust the audio processing strategy (such as noise reduction intensity, audio gain, etc.) based on the semantic analysis result of the audio data, which can improve the audio quality in scenarios such as calls and live broadcasts, and thus improve the user's usage experience.

[0074] In another embodiment, in an audio scenario such as a voice call or video call between a first electronic device and a second electronic device, the second electronic device can perform real-time analysis of the audio data from the first electronic device and then return the analysis results to the first electronic device. The first electronic device can then dynamically adjust the audio processing strategy based on the analysis results. For example, the second electronic device can analyze the audio data from the first electronic device based on a voice quality detection algorithm to obtain a real-time voice quality score. If the voice quality score is greater than a quality score threshold (e.g., 6, where the voice quality score is 1-10), the first electronic device can perform no processing. If the voice quality score is less than or equal to the quality score threshold, the first electronic device can determine the reason why the voice quality score is lower than the quality score threshold based on the corresponding audio data (e.g., the audio amplitude is too low, the audio data contains a lot of noise, etc.), and then return the reason to the first electronic device. The first electronic device can then dynamically adjust the audio processing strategy based on the reason. For example, if the reason why the voice quality score is lower than the quality score threshold is that the audio data contains a lot of noise, that is, it indicates that the noise on the first electronic device side is higher, then the first electronic device can adjust the current noise reduction strength, such as adjusting the noise reduction strength from "weak" to "medium."

[0075] In another embodiment, since the noise sources in different scenarios (such as subway stations, offices, construction sites, etc.) may be different, multiple audio processing scenarios can be pre-defined, and different noise reduction models can be configured for different audio processing scenarios. In audio scenarios such as voice calls and video calls between a first electronic device and a second electronic device, the first electronic device can determine the current audio processing scenario based on the positioning results, and then automatically help the user select the noise reduction model corresponding to the audio processing scenario for processing, which can ensure better audio quality. In addition, the first electronic device can also provide a corresponding configuration interface to facilitate the user to manually select the corresponding audio processing scenario. For example, the first electronic device can display a corresponding audio processing configuration interface in response to user input, and the audio processing configuration interface includes multiple audio processing scenario options. The user can select the corresponding audio processing scenario option by touching or clicking. Afterwards, the first electronic device can use the noise reduction model corresponding to the audio processing scenario option selected by the user for processing, which can ensure better audio quality.

[0076] Please refer to FIG1 , which exemplarily shows a structural diagram of a communication system provided in an embodiment of the present application.

[0077] The communication system may include an electronic device 100 and an electronic device 200. The electronic device 100 and the electronic device 200 may perform an audio call, such as a voice call, a video call, etc. The electronic device 100 may be a first electronic device, and the electronic device 200 may be a second electronic device.

[0078] The electronic device 100 and the electronic device 200 may be equipped with Or portable electronic devices with other operating systems, such as mobile phones, tablet computers, wearable devices (such as smart watches, smart bracelets, etc.), etc., can also be non-portable electronic devices such as augmented reality (AR) devices, virtual reality (VR) devices, laptop computers with touch-sensitive surfaces or touch panels, desktop computers with touch-sensitive surfaces or touch panels. Electronic devices 100 and 200 can also be chips or processing systems in these devices. The embodiments of the present application do not limit the types of electronic devices 100 and 200.

[0079] Among them, in audio scenarios such as voice calls and video calls between the electronic device 100 and the electronic device 200, the electronic device 100 can perform semantic analysis on the audio data from the electronic device 200, and dynamically adjust the audio processing strategy based on the semantic analysis results. The electronic device 100 can also dynamically adjust the audio processing strategy based on the analysis results returned by the electronic device 200. The electronic device 100 can also determine the current audio processing scenario based on the positioning results, and then automatically adopt the noise reduction model corresponding to the audio processing scenario. Similarly, the electronic device 200 can perform semantic analysis on the audio data from the electronic device 100, and dynamically adjust the audio processing strategy based on the semantic analysis results. The electronic device 200 can also dynamically adjust the audio processing strategy based on the analysis results returned by the electronic device 100. The electronic device 200 can also determine the current audio processing scenario based on the positioning results, and then automatically adopt the noise reduction model corresponding to the audio processing scenario.

[0080] It should be noted that the system architecture and business scenarios (or application scenarios) described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application.

[0081] The structure of the electronic device 100 involved in this application is introduced below.

[0082] FIG2 exemplarily shows a schematic diagram of the hardware structure of the electronic device 100 provided in an embodiment of the present application.

[0083] As shown in Figure 2, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0084] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0085] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0086] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0087] Processor 110 may also include a memory for storing instructions and data. In some examples, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces processor 110 latency, and thus improves system efficiency.

[0088] USB interface 130 is an interface that complies with USB standards and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. USB interface 130 can be used to connect a charger to charge electronic device 100, transfer data between electronic device 100 and peripheral devices, or connect headphones to play audio.

[0089] The charging management module 140 is used to receive charging input from a charger. The charger can be a wireless charger or a wired charger. While charging the battery 142, the charging management module 140 can also power the electronic device through the power management module 141.

[0090] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to provide power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160.

[0091] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.

[0092] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.

[0093] The mobile communication module 150 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low-noise amplifier (LNA), and the like. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, filter and amplify the received electromagnetic waves, and transmit them to the modem processor for demodulation. The mobile communication module 150 can also amplify the signals modulated by the modem processor and convert them into electromagnetic waves for radiation via the antenna 1.

[0094] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the electronic device 100. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0095] The electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing that connects the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering.

[0096] The display screen 194 is used to display images, videos, etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194 , where N is a positive integer greater than 1.

[0097] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.

[0098] The ISP is used to process data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then transmitted to the ISP for processing and converted into an image visible to the naked eye.

[0099] The camera 193 is used to capture still images or videos. In some embodiments, the electronic device 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.

[0100] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.

[0101] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU can enable intelligent cognitive applications in electronic device 100, such as image recognition, face recognition, speech recognition, and text comprehension.

[0102] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.

[0103] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0104] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.

[0105] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some examples, the audio module 170 can be set in the processor 110, or some functional modules of the audio module 170 can be set in the processor 110. The speaker 170A, also known as the "speaker", is used to convert audio electrical signals into sound signals. The receiver 170B, also known as the "earpiece", is used to convert audio electrical signals into sound signals. The microphone 170C, also known as the "microphone" or "microphone", is used to convert sound signals into electrical signals. The headphone jack 170D is used to connect wired headphones.

[0106] The sensor module 180 may include a pressure sensor, a gyro sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, and the like.

[0107] Buttons 190 include a power button, a volume button, etc. Motor 191 can generate vibration prompts. Indicator 192 can be an indicator light that can be used to indicate charging status, power changes, messages, missed calls, notifications, etc.

[0108] The SIM card interface 195 is used to connect a SIM card. A SIM card can be connected to and disconnected from the electronic device 100 by inserting or removing it from the SIM card interface 195. The electronic device 100 may support one or N SIM card interfaces, where N is a positive integer greater than 1. The electronic device 100 interacts with the network through the SIM card to implement functions such as calls and data communications. In some examples, the electronic device 100 uses an eSIM, or embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.

[0109] The electronic device 100 may be equipped with Hongmeng system ( OS) or other operating systems, such as mobile phones, tablet computers, laptop computers, smart watches, smart bracelets, etc. The embodiment of the present application does not limit the specific type of the electronic device 100.

[0110] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. Taking the system as an example, the software structure of the electronic device 100 is exemplarily described.

[0111] FIG3 is a software structure block diagram of the electronic device 100 according to an embodiment of the present application.

[0112] The layered architecture divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, The system is divided into four layers, from top to bottom: application layer, application framework layer, Android runtime and system library, and kernel layer.

[0113] The application layer can include a series of application packages.

[0114] As shown in FIG3 , the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, and short message.

[0115] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0116] As shown in FIG3 , the application framework layer may include a window manager, a content provider, a view system, a telephony manager, a resource manager, a notification manager, an activity manager, and the like.

[0117] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.

[0118] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.

[0119] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.

[0120] The phone manager is used to provide communication functions of the electronic device 100, such as management of call status (including answering, hanging up, etc.).

[0121] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.

[0122] The Notification Manager allows applications to display notification information in the status bar (such as the pull-down notification bar). It can be used to convey notification-type messages and can disappear automatically after a short stay without user interaction. For example, the Notification Manager is used to notify the completion of downloads, message reminders, etc. The Notification Manager can also be used to display notifications in the form of icons or scrolling text in the status bar at the top of the system, such as notifications from applications running in the background, or notifications that appear on the screen in the form of dialog windows. For example, text messages can be displayed in the status bar, prompts can be sounded, electronic devices can vibrate, indicator lights can flash, etc.

[0123] The Activity Manager is responsible for managing activities, starting, switching, and scheduling components in the system, as well as managing and scheduling applications. The Activity Manager can be called by upper-level applications to open corresponding activities.

[0124] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for scheduling and management of the Android system.

[0125] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.

[0126] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.

[0127] The system library can include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.

[0128] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.

[0129] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0130] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0131] A 2D graphics engine is a drawing engine for 2D drawings.

[0132] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.

[0133] The hardware structure and software structure of the electronic device 200 may also be as shown in FIG. 2 and FIG. 3 . The embodiment of the present application does not specifically limit the hardware structure and software structure of the electronic device 200 .

[0134] The following describes the usage scenarios involved in the embodiments of the present application in conjunction with the user interface of the electronic device 100.

[0135] 4A to 4H are schematic diagrams showing scenarios in which the electronic device 100 provides a smart call mode and related audio processing configurations.

[0136] As shown in Figure 4A, the electronic device 100 can display a user interface 410. The user interface 410 displays a page with application icons. The page may include multiple application icons (for example, a clock application icon, a calendar application icon, a gallery application icon, a memo application icon, a file management application icon, etc.). A page indicator may also be displayed below the above-mentioned multiple application icons to indicate the positional relationship between the currently displayed page and other pages. There are multiple tray icons (for example, a camera application icon, an address book application icon, a phone application icon 411, and a message application icon) below the page indicator. The tray icon remains displayed when the page is switched. It should be understood that the embodiments of the present application do not limit the content displayed on the user interface 410. In response to a user operation acting on the above-mentioned application icon or tray icon, such as a touch operation, the electronic device 100 can start the application corresponding to the application icon or the tray icon.

[0137] For example, in response to a user operation on the phone application icon 411, the electronic device 100 may open the phone application and display the user interface 420 shown in FIG4B . The user interface 420 may include a dial pad, a number display area, a dialing control (such as the dialing control 421), etc. The user may enter the number to be dialed through the dial pad. Then, in response to a user operation on the dialing control 421, such as a touch operation, the electronic device 100 may dial the corresponding number. After the call is connected, the user interface 430 shown in FIG4C may be displayed.

[0138] As shown in Figure 4C , in response to a user operation of sliding downward on the status bar at the top of user interface 430, electronic device 100 may display user interface 440 as shown in Figure 4D . For example, user interface 440 may be a notification information interface. It should be understood that notification information interfaces may also be accessed through other user interfaces, and this embodiment of the present application is not limited thereto. User interface 440 may include a call status card 441 and a smart call mode card 445. Call status card 441 may be used to display call-related information, such as call time and call number. Call status card 441 may also include a mute control 442, a hang-up control 443, and a hands-free control 444. In response to a user operation on mute control 442, electronic device 100 may turn the mute function on or off. When the mute function is on, the other party can hear the current caller's voice; when the mute function is off, the other party cannot hear the current caller's voice. In response to a user operation on hang-up control 443, electronic device 100 may hang up the current call. In response to user operations on the hands-free control 444, the electronic device 100 can turn off or on the hands-free function. When the hands-free function is turned on, the voice of the other party can be played out through the speaker at a louder volume. When the hands-free function is turned off, the voice of the other party can be played out through the earpiece at a lower volume.

[0139] The smart call mode card 445 may include relevant information about the smart call mode, such as whether the smart call mode is turned on (such as currently in the off state), relevant prompt information for intelligently adjusting the audio processing strategy, controls for turning on or off the smart call mode, etc. As shown in Figure 4D, in response to a user operation acting on the expansion control 446, such as a touch operation, the electronic device 100 can display a user interface 440 as shown in Figure 4E. It can be understood that due to the limited screen size of the electronic device 100, usually, in the notification information interface, some cards with more content may only present part of the content, and the electronic device 100 can respond to the user operation of the user acting on the display control corresponding to a card to expand the complete content of the card. In other words, the user can view the complete content of the card by touching the user operation of the expansion control.

[0140] As shown in FIG4E , user interface 440 may display the full contents of a smart call mode card 445. Smart call mode card 445 may also display an audio processing configuration control 447, an enable smart call mode control 448, and a hide control 449. Audio processing configuration control 447 can be used to open the audio processing configuration interface, where it can be used to configure audio processing parameters such as noise reduction intensity and audio gain. Enable smart call mode control 448 can be used to enable smart call mode. Hide control 449 can be used to hide the full contents of smart call mode card 445, presenting only a portion of the contents. This is the opposite of the expand control.

[0141] As shown in FIG4E , in response to a user operation on the smart call mode activation control 448, the electronic device 100 may activate the smart call mode and display the user interface 440 shown in FIG4F . The user interface 440 shown in FIG4F displays the prompt "Enabled, Optimizing Call Quality," indicating that the smart call mode is activated. Furthermore, the user interface 440 shown in FIG4F may also display a control 450 for deactivating the smart call mode. The deactivation control 450 may be used to deactivate the smart call mode.

[0142] As shown in Figure 4F, in response to a user operation on the audio processing configuration control 447, such as a touch operation, the electronic device 100 may display a user interface 460, i.e., an audio processing configuration interface, as shown in Figure 4G. User interface 460 may include a configuration area for noise reduction intensity and an area for audio gain. The noise reduction intensity configuration area can be used to adjust the noise reduction intensity. For example, the noise reduction intensity may include three levels: weak, medium, and strong, with the currently selected noise reduction intensity being weak. The noise reduction intensity here may be the noise reduction intensity that the electronic device 100 applies to the collected external audio. The stronger the noise reduction intensity, the greater the ability to reduce / eliminate ambient noise. In some embodiments, each time the electronic device makes a call, if the user does not change the audio processing configuration, the default noise reduction intensity may be weak. The audio gain may include multiple options, such as 0dB, 3dB, 5dB, 10dB, 12dB, 15dB, etc. The currently selected audio gain is 5dB. The audio gain here may be the audio gain of the electronic device 100 for audio processing of the collected external audio. The stronger the audio gain, the louder the sound that the other end can hear.

[0143] The return control 461 can be used to trigger the electronic device 100 to return to the previous user interface, that is, to display the user interface 440 shown in Figure 4F.

[0144] As shown in Figure 4G, in response to a user operation acting on the optional control 462 corresponding to the noise reduction intensity, such as a touch operation, the electronic device 100 can display a user interface 460 as shown in Figure 4H. In the user interface 460, the optional control 462 corresponding to the noise reduction intensity is in a selected state, indicating that the current electronic device 100 has switched the noise reduction intensity to medium. It should be understood that in the embodiment of the present application, each noise reduction intensity can correspond to a noise reduction model, and different noise reduction intensities correspond to different noise reduction models. When switching the noise reduction intensity, the electronic device 100 can switch the noise reduction model used accordingly. Among them, for the same noise reduction model (such as a neural network model with the same structure), different noise reduction parameters can be understood as different noise reduction models in the embodiment of the present application, that is, if any one of the noise reduction model and the model parameter is different, it can be regarded as a different noise reduction model. Exemplarily, the noise reduction model corresponding to weak noise reduction strength may be noise reduction model 1 (model 1 + model parameter 1), the noise reduction model corresponding to medium noise reduction strength may be noise reduction model 2 (model 1 + model parameter 2), and the noise reduction model corresponding to strong noise reduction strength may be noise reduction model 3 (model 1 + model parameter 3). Noise reduction model 1, noise reduction model 2, and noise reduction model 3 may use the same model 1, but have different model parameters. As another example, the noise reduction model corresponding to weak noise reduction strength may be noise reduction model 1 (model 1 + model parameter 1), the noise reduction model corresponding to medium noise reduction strength may be noise reduction model 2 (model 2 + model parameter 2), and the noise reduction model corresponding to strong noise reduction strength may be noise reduction model 3 (model 3 + model parameter 3). Noise reduction model 1, noise reduction model 2, and noise reduction model 3 may use different models, and the model parameters may be different.

[0145] It should be understood that the embodiments of the present application do not limit the level of noise reduction intensity, and can also be divided into more or fewer noise reduction levels. Similarly, the embodiments of the present application do not limit the options for audio gain, and can also include more or fewer audio gain options.

[0146] As can be seen from the above audio processing configuration interface, users can select different noise reduction intensities and audio gains based on their actual needs to improve call quality for the other end user. For example, if the other end user complains that the noise is too loud during a call, if the currently selected noise reduction level is weak, a higher noise reduction intensity level, such as medium or strong, can be selected. For another example, if the other end user complains that the voice is too quiet during a call, if the currently selected audio gain is 5dB, a higher audio gain level, such as 10dB or 15dB, can be selected.

[0147] It is understood that in some cases, the noise reduction intensity mentioned above may also refer to the noise reduction intensity used by the electronic device 100 to reduce the audio received from other electronic devices (such as the electronic device 200). The stronger the noise reduction intensity, the greater the ability to reduce / eliminate noise in the received audio. Similarly, the audio gain mentioned above may also refer to the audio gain used by the electronic device 100 to reduce the audio received from other electronic devices (such as the electronic device 200). The stronger the audio gain, the louder the sound that can be heard by the local end.

[0148] The above briefly introduces the scenario in which the user manually adjusts the audio processing strategy, but in an embodiment of the present application, in audio scenarios such as voice calls and video calls, the electronic device 100 can also automatically adjust the audio processing strategy, which can improve the adjustment efficiency of the audio processing strategy and improve the user experience. For example, the electronic device 100 can perform semantic analysis on the received audio from the other end, and then intelligently adjust the audio processing strategy based on the results of the semantic analysis. For another example, the electronic device 100 can receive audio quality issues fed back by the other end electronic device (such as the electronic device 200), and then automatically adjust the audio processing strategy based on the feedback from the other end.

[0149] 5A to 5D are schematic diagrams illustrating scenarios in which the electronic device 100 provides an intelligent call mode and intelligently adjusts an audio processing strategy.

[0150] After turning on the smart call mode, the electronic device 100 can perform real-time semantic analysis on the audio received from the other end of the call. If the semantic analysis result of a certain section of audio data (such as "It's so noisy over there", "The noise over there is so loud", etc.) is "The current noise is relatively loud", then the electronic device 100 can adjust the current noise reduction intensity, such as adjusting the noise reduction intensity from "weak" to "medium". If the semantic analysis result of a certain section of audio data (such as "Your voice is so soft", "I can't hear what you are saying", etc.) is "The current sound is relatively soft", then the electronic device 100 can increase the current audio gain, such as adjusting the audio gain from 5dB to 10dB. In the embodiment of the present application, semantic analysis can also be called semantic understanding.

[0151] As shown in Figure 5A, assuming that at the 10th second of a call, electronic device 100 determines based on the semantic analysis results that the current noise level on the local end is relatively high, electronic device 100 can automatically adjust the noise reduction intensity from "weak" to "medium." Furthermore, smart call mode card 445 can include relevant prompt information, such as display area 451 of smart call mode card 445, to inform the user that the noise reduction intensity has been automatically adjusted.

[0152] As shown in Figure 5B, assuming that at the 20th second of the call, the electronic device 100's semantic analysis result of the audio received from the other end of the call (such as "Your voice is so soft", "I can't hear what you are saying", etc.) is "The current voice is relatively soft", then the electronic device 100 can automatically adjust the audio gain from 3dB to 10dB, and relevant prompt information can be displayed in the display area 452 in the smart call mode card 445 to inform the user that the audio gain has been automatically adjusted.

[0153] The following describes a method for automatically adjusting the audio processing strategy based on feedback from the peer electronic device.

[0154] After the smart call mode is turned on, when the electronic device 100 receives feedback on audio quality issues from the other electronic device (such as electronic device 200), the electronic device 100 can adjust the current audio processing strategy based on the audio quality issues fed back by the other end. For example, if the other end feedback is "the current noise is relatively loud", the electronic device 100 can adjust the current noise reduction strength, such as adjusting the noise reduction strength from "weak" to "medium". If the other end feedback is "the current sound is relatively quiet", the electronic device 100 can increase the current audio gain, such as adjusting the audio gain from 5dB to 10dB.

[0155] As shown in Figure 5C, assuming that in the 10th second of the call, the electronic device 100 receives feedback from the other electronic device about an audio quality problem ("the current noise is relatively loud"), then the electronic device 100 can automatically adjust the noise reduction intensity from "weak" to "medium", and relevant prompt information can be displayed in the display area 453 in the smart call mode card 445 to inform the user that the noise reduction intensity has been automatically adjusted.

[0156] As shown in Figure 5D, assuming that at the 20th second of the call, the electronic device 100 receives feedback from the other electronic device about an audio quality problem ("the current sound is relatively low"), then the electronic device 100 can automatically adjust the audio gain from 3dB to 10dB, and relevant prompt information can be displayed in the display area 454 in the smart call mode card 445 to inform the user that the audio gain has been automatically adjusted.

[0157] 6A to 6C exemplarily illustrate scenarios in which the electronic device 100 provides noise reduction scenario selection and intelligently adjusts the audio processing strategy based on positioning results.

[0158] It is understandable that in daily life, users often need to use the call function in different scenarios, and the noise sources in different scenarios may be different. Therefore, if the electronic device 100 uses the same noise reduction model for all scenarios, it may result in poor noise reduction effect in some scenarios. In the embodiment of the present application, in order to provide users with a better call experience, a noise reduction model can be configured for each of a variety of different scenarios (such as offices, shopping malls, train stations, subway stations, etc.) to achieve better noise reduction effect.

[0159] As shown in Figure 6A, in the user interface 460, a configuration area for noise reduction scenes can also be displayed. The configuration area for noise reduction scenes can be used to adjust the noise reduction scenes. For example, the noise reduction scenes can include general, office, shopping mall, subway station, railway station, construction site, and so on. Normally, you can select the "general" noise reduction scene. The noise reduction model corresponding to the "general" noise reduction scene can meet the usage requirements of most scenes, but the noise reduction effect may be weaker than the dedicated noise reduction model in a specific scene. It can be understood that the noise reduction scenes here are only exemplary. In other embodiments of the present application, more or fewer noise reduction scenes may be included.

[0160] Assuming the user is currently in a shopping mall, in order to obtain a better noise reduction effect, the user can switch the noise reduction scene to the shopping mall. As shown in Figure 6A, in response to a user operation, such as a touch operation, on the selection control 463 corresponding to the shopping mall, the electronic device 100 can display the user interface 460 shown in Figure 6B. In the user interface 460, the selection control 463 corresponding to the shopping mall is in the selected state, indicating that the electronic device 100 has switched the noise reduction scene to the shopping mall, and the noise reduction model can also be switched from the noise reduction model corresponding to the general scene to the noise reduction model corresponding to the shopping mall scene.

[0161] In order to improve user experience, the electronic device 100 can also adjust the noise reduction scene based on the real-time positioning results during a call, that is, switch the noise reduction model of the corresponding scene based on the real-time positioning results.

[0162] As shown in FIG6C , assuming that at the 20th second of a call, electronic device 100 determines based on positioning results that the user is currently located in a shopping mall, the noise reduction scene can be automatically adjusted to the shopping mall, switching to the noise reduction model corresponding to the shopping mall scene. Furthermore, display area 455 of smart call mode card 445 can display a relevant prompt to inform the user that the noise reduction scene has been automatically adjusted.

[0163] 7A to 7F are schematic diagrams of scenarios in which the electronic device 100 provides a smart call mode, and uplink and uplink-related audio processing configurations.

[0164] It is understandable that for audio scenarios such as voice calls and video calls, they are generally bidirectional, that is, audio can be received from the other electronic device, and external audio collected by the local end can be sent to the other end. The local electronic device sending the external audio collected by the local electronic device to the other electronic device can be called uplink, and the local electronic device receiving the audio sent by the other electronic device can be called downlink. Therefore, for the audio processing function of the electronic device 100 (such as the noise reduction function), it can be divided into uplink audio processing and downlink audio processing. Uplink audio processing can be the processing of external audio collected by the electronic device 100, and downlink audio processing can be the processing of audio received from the other electronic device (such as electronic device 200). Uplink processing is mainly to reduce / eliminate the noise in the external audio collected by the local end, and to adjust the audio amplitude so that the user on the other end can hear the target voice more clearly. Downlink processing is mainly to reduce / eliminate the noise in the audio received from the other electronic device, and to adjust the audio amplitude so that the user on the local end can hear the target voice more clearly.

[0165] Typically, the two parties on a call are in different environments (eg, user A is in a quiet room, and user B is in a crowded shopping mall). Therefore, the electronic device 100 can provide functions for uplink audio processing and downlink audio processing respectively.

[0166] As shown in Figure 7A, in response to a user operation on the smart call mode control 448, the electronic device 100 can turn on the smart call mode and display the user interface 440 shown in Figure 7B. In the user interface 440 shown in Figure 7B, the display area 471 displays the prompt message "Enabled (uplink + downlink), optimizing call quality", indicating that the smart call mode has been turned on and both the uplink audio processing and downlink audio processing functions are turned on. When the uplink audio processing function is turned on, the mute control 442 can be in a lit state. When the downlink audio processing function is turned on, the hands-free control 444 can be in a lit state.

[0167] As shown in FIG7B , in response to a user operation on hands-free control 444, such as a double-click, electronic device 100 may display user interface 440 as shown in FIG7C . In user interface 440 shown in FIG7C , hands-free control 444 is not illuminated, indicating that electronic device 100 has disabled the downlink audio processing function. Furthermore, the prompt message only indicates that the uplink audio processing function is enabled.

[0168] As shown in FIG7C , in response to a user operation on mute control 442, such as a double-click, electronic device 100 may display user interface 440 as shown in FIG7D . In user interface 440 shown in FIG7D , mute control 442 is not illuminated, indicating that electronic device 100 has disabled the uplink audio processing function. At this time, since both the uplink audio processing function and the downlink audio processing function are disabled, the smart call mode is disabled.

[0169] It is understood that by performing another user operation on the mute control 442, such as double-clicking, the electronic device 100 can re-enable the uplink audio processing function. Similarly, by performing another user operation on the hands-free control 444, such as double-clicking, the electronic device 100 can re-enable the downlink audio processing function.

[0170] As shown in Figure 7D, in response to a user operation, such as a touch operation, on the audio processing configuration control 447, the electronic device 100 may display a user interface 460 as shown in Figure 7E. In the user interface 460, an uplink audio processing configuration area 472 and a downlink audio processing configuration area 473 may be displayed. Both the uplink audio processing configuration area 472 and the downlink audio processing configuration area 473 may include a configuration area for noise reduction intensity, a configuration area for audio gain, and a configuration area for noise reduction scenarios.

[0171] The user can configure the audio processing strategy for uplink audio and downlink audio through the user interface 460 shown in Figure 7E. For example, assuming that the shopping mall where the electronic device 100 is located has a loud noise, the user can set the uplink noise reduction scene to shopping mall and the noise reduction intensity to medium, as shown in Figure 7F.

[0172] It is understood that in addition to the user manually setting the uplink and downlink audio processing strategies, during a call on the electronic device 100, the electronic device 100 can also automatically adjust the uplink and downlink noise reduction strength, audio gain, and noise reduction scenario based on semantic analysis or feedback from the other electronic device. The electronic device 100 can also adjust the uplink noise reduction scenario based on positioning results, which will not be described in detail here.

[0173] 8A to 8E exemplarily illustrate a scenario in which the electronic device 200 requests the electronic device 100 to enable the smart call mode and adjust the uplink audio processing strategy.

[0174] Please refer to Figure 8A. In the user interface 470 shown in Figure 8A, a call status card and a smart call mode card may be displayed. Among them, the call status card can be used to display call-related information, such as call time, call number, etc. The smart call mode card can display an audio processing configuration control (local end), a control for turning on the smart call mode, a peer audio processing configuration control 471, and a hidden control. The audio processing configuration control (local end) can be used to open the local audio processing configuration interface, and the peer audio processing configuration control 471 can be used to open the peer audio processing configuration interface. The control for turning on the smart call mode can be used to turn on the smart call mode. The user interface 470 can be a notification information interface of the electronic device 200.

[0175] In response to a user operation, such as a touch operation, on the peer audio processing configuration control 471, the electronic device 200 may display a user interface 480 as shown in FIG8B . User interface 480 may display two areas: area 481 for requesting the peer to enable smart call mode, and area 482 for requesting the peer to adjust noise reduction intensity, audio gain, and noise reduction scenarios. Area 481 may include a send request control 4811, which may be used to trigger the electronic device 200 to send a request message to the peer to enable smart call mode.

[0176] In response to a user operation (such as a touch operation) on the send request control 4811, the electronic device 200 may send a request message (a first request message) to the electronic device 100, requesting to enable the smart call mode. Accordingly, the electronic device 100 may receive the request message from the electronic device 200, and then, based on the request message, the electronic device 100 may pop up a prompt message box, which may display a prompt message, and the prompt message is used to prompt the other end user to request to enable the smart call mode. For example, the prompt message may be "the other end requests to enable the smart call mode." For example, assuming that the electronic device 100 is currently displaying a call interface, after the electronic device 100 receives the request message from the electronic device 200, a prompt message box may pop up on the call interface, as shown in FIG8C . The electronic device 100 may display a user interface 430, which may display a prompt message box 431. The prompt message box 431 displays the prompt message "the other end requests to enable the smart call mode", as well as corresponding confirmation controls 4311 and rejection controls 4312. Among them, the confirmation control 4311 can be used to confirm / accept the request sent by the other party, and the rejection control 4312 can be used to reject the request sent by the other party.

[0177] Exemplarily, in response to a user operation (such as a touch operation) acting on the determination control 4311, the electronic device 100 can enable the smart call mode. Opening the notification information interface of the electronic device 100 can display the user interface shown in Figure 4F. In response to a user operation (such as a touch operation) acting on the rejection control 4312, the electronic device 100 can do nothing, or the electronic device 100 can send an instruction message to the electronic device 200, which is used to indicate the rejection of the request of the electronic device 200, that is, to indicate the rejection of the smart call mode.

[0178] As shown in Figure 8D, in area 482, the user can select noise reduction intensity, audio gain, and noise reduction scenarios, and the user can send a request to the electronic device 100 through the electronic device 200 to request that the electronic device 100 adjust the uplink audio processing strategy. For example, the user can touch or click the selection control 463 corresponding to the shopping mall and select the noise reduction scenario as the shopping mall. The user can then click or touch the control 4822. In response to this operation, the electronic device 200 can send a corresponding request message to the electronic device 100, requesting adjustment of the uplink audio processing strategy. The request message can carry the audio processing parameters to be adjusted and the corresponding parameter values, such as {noise reduction scenario: shopping mall}. Accordingly, the electronic device 100 can receive the request message from the electronic device 200. Afterwards, the electronic device 100 can pop up a prompt message box based on the request message. The prompt message box can display a prompt message, which can be used to prompt the other end user to request adjustment of the uplink audio processing strategy. For example, assuming that the electronic device 100 currently displays a call interface, after the electronic device 100 receives a request message from the electronic device 200, a prompt information box may pop up on the call interface. As shown in FIG8E , the electronic device 100 may display a user interface 430, which may display a prompt information box 432. The prompt information box 432 displays the prompt message "The other party requests to switch the noise reduction scene to the shopping mall", as well as the corresponding confirmation control 4321 and rejection control 4322. Among them, the confirmation control 4311 can be used to confirm / accept the request sent by the other party, and the rejection control 4312 can be used to reject the request sent by the other party.

[0179] For example, in response to a user operation (such as a touch operation) on the determination control 4321, the electronic device 100 may switch the uplink noise reduction scene to the shopping mall. In response to a user operation (such as a touch operation) on the rejection control 4322, the electronic device 100 may not process the request, or the electronic device 100 may send an instruction to the electronic device 200, which is used to indicate the rejection of the request of the electronic device 200, that is, to indicate the rejection of switching the uplink noise reduction scene to the shopping mall.

[0180] It is understood that the user interfaces of Figures 8A-8E are merely exemplary. For example, in some possible implementations, the electronic device 200 may further provide a control for requesting the peer end to disable the smart call mode. For another example, in some possible implementations, the electronic device 200 may further provide configuration of a corresponding downlink audio processing strategy for the peer end, which is not limited in this embodiment of the present application.

[0181] In some possible implementations, in audio scenarios such as voice calls and video calls, the electronic device 100 can, during the process of processing audio (such as call audio), present in real time on the user interface visual graphics (such as time domain graphs, spectrum graphs, time-spectrum graphs, etc.) of the processed original audio and processed audio (such as audio after noise reduction, audio after gain adjustment, etc.), as well as other visual graphics that can present the audio processing effects.

[0182] In the embodiments of the present application, the audio processed by the electronic device may be audio collected from the outside world, such as audio collected from the outside world through a microphone during a call. The audio processed by the electronic device may also be audio received from other electronic devices, such as audio sent by the other electronic device during a call.

[0183] It can be seen that through the above method, users can view the visual graphics of the original audio and the processed audio in real time. By comparing the visual graphics of the original audio and the processed audio, users can intuitively understand the processing effect and processing capabilities of audio processing such as noise reduction of electronic devices.

[0184] For example, during a call, the electronic device can collect audio in real time through a microphone, and can perform real-time noise reduction processing on the collected audio, and can present a time domain graph of the original collected audio and the audio after noise reduction processing in real time on the user interface. The user can know how much ambient noise is filtered through the time domain graph of the original collected audio and the audio after noise reduction processing, and can intuitively feel / understand the real-time noise reduction effect of the electronic device.

[0185] The following is an exemplary description of the processing flow of the technical solution provided in the embodiment of the present application. Please refer to Figure 9, which is a flow chart of an audio processing method disclosed in the embodiment of the present application. As shown in Figure 9, the method may include but is not limited to the following steps:

[0186] 901. During a call between electronic device 100 and electronic device 200, electronic device 200 sends first audio data to electronic device 100.

[0187] For example, assume that user A holds electronic device 100 and user B holds electronic device 200. When user A needs to contact user B, user A and user B can conduct a call (such as a voice call, video call, etc.) through electronic devices 100 and 200. It is understood that the call here can be a cellular call or a call function provided by various communication applications, social applications, etc., and the embodiments of the present application are not limited thereto.

[0188] During a call between electronic device 100 and electronic device 200, electronic device 100 can collect external sounds, such as external ambient sounds and the voice of user A, through its own microphone in real time, and can send corresponding audio data to electronic device 200. Accordingly, electronic device 200 can receive audio data from electronic device 100. Similarly, electronic device 200 can also collect external sounds, such as external ambient sounds and the voice of user B, through its own microphone in real time, and can send corresponding audio data (such as first audio data) to electronic device 100. Accordingly, electronic device 100 can receive audio data from electronic device 200.

[0189] 902. The electronic device 100 performs semantic analysis based on the first audio data to obtain a first semantic analysis result.

[0190] The electronic device 100 can receive audio data from the electronic device 200 in real time and can perform semantic analysis in real time, so as to dynamically adjust the current uplink audio processing strategy based on the semantic analysis result.

[0191] For example, the electronic device 100 may receive first audio data from the electronic device 200, and then the electronic device 100 may perform semantic analysis on the first audio data to obtain a first semantic analysis result. For example, the first audio data may be "It's so noisy over there," "The noise over there is so loud," "It's a bit noisy over there, I can't hear clearly," "It's still so noisy over there," "The noise over there is still so loud," etc., and the first semantic analysis result corresponding to these audio data may be "The external noise on this end is currently quite loud." For another example, the first audio data may be "Your voice is so soft," "I can't hear what you are saying," "Can you speak louder?", "Your voice is still a little soft," etc., and the corresponding first semantic analysis result may be "The current sound on this end is relatively soft." For another example, the first audio data may be "Your voice is so loud," "Can you speak more quietly?", "Your voice is too loud," etc., and the corresponding first semantic analysis result may be "The current sound on this end is relatively loud."

[0192] In some possible implementations, the electronic device 100 may perform semantic analysis on the first audio data only when the smart call mode is enabled, so as to dynamically adjust the current uplink audio processing strategy based on the semantic analysis results. That is, when the smart call mode is disabled, the electronic device 100 may not perform any processing. In other possible implementations, the electronic device 100 may also perform semantic analysis on the first audio data when the smart call mode is disabled, so as to determine whether to enable the smart call mode based on the semantic analysis results. Furthermore, after enabling the smart call mode, the electronic device 100 may continue to perform semantic analysis on subsequent audio data, so as to dynamically adjust the current uplink audio processing strategy based on the semantic analysis results. For example, during an initial call, the smart call mode is disabled by default. Subsequently, the electronic device 100 may automatically enable the smart call mode if the semantic analysis results indicate "external noise on the current local end is relatively loud," "the current local end sound is relatively quiet," or "the current local end sound is relatively loud," so as to perform audio processing such as uplink noise reduction and adjusting the uplink audio amplitude. In this way, during a call, the intelligent call mode can be automatically turned on based on the semantic analysis results, which can meet usage needs in a timely manner. While ensuring call quality, it can also reduce the overall power consumption of electronic devices.

[0193] It is understandable that when the smart call mode is turned on, the electronic device 100 can perform corresponding processing based on the audio processing configuration, such as performing noise reduction processing based on the configured noise reduction intensity, and adjusting the audio amplitude based on the configured audio gain.

[0194] It should be noted that the embodiment of the present application does not specifically limit the semantic analysis model used by the electronic device 100, and it can be a deep neural network model or other semantic analysis model. In addition, it should be understood that for audio data, the semantic analysis model can first convert the audio data into text, that is, extract the content of the target voice in the audio data, and then perform semantic analysis based on the text (that is, analyze and understand the text). In other words, the semantic analysis model can include the function of converting speech to text and the function of semantic analysis.

[0195] 903. The electronic device 100 adjusts the uplink audio processing strategy based on the first semantic analysis result.

[0196] After obtaining the first semantic analysis result, the electronic device 100 may adjust the currently used uplink audio processing strategy based on the first semantic analysis result.

[0197] It is understandable that the conversation between user A and user B usually includes a lot of content that does not involve audio quality or audio effects (such as noise, sound volume, etc.), but only the semantic analysis results corresponding to the conversation content related to audio quality will affect the audio processing strategy. Therefore, the electronic device 100 can adjust the currently used uplink audio processing strategy based on the first semantic analysis result when the first semantic analysis result is related to the audio quality. In other words, the electronic device 100 can adjust the currently used uplink audio processing strategy based on the first semantic analysis result when it is confirmed based on the first semantic analysis result that the audio effect of the call needs to be improved. For example, if the first semantic analysis result is "the external noise on this end is relatively large at the moment", "the sound on this end is relatively small at the moment", or "the sound on this end is relatively loud at the moment", it indicates that the audio effect of the call is poor, and the electronic device 100 can determine that the audio effect of the call needs to be improved.

[0198] For example, if the first semantic analysis result is "the current external noise on this end is relatively loud", it indicates that the current uplink noise reduction strength is insufficient, and the electronic device 100 can select a higher uplink noise reduction strength. For example, if the current uplink noise reduction strength is "weak", the electronic device 100 can adjust the uplink noise reduction strength to "medium" or "strong". If the current uplink noise reduction strength is "medium", the electronic device 100 can adjust the uplink noise reduction strength to "strong". In the embodiment of the present application, a corresponding noise reduction model can be configured for each uplink noise reduction strength. When a certain uplink noise reduction strength is selected, the collected external audio can be processed by the noise reduction model corresponding to the uplink noise reduction strength. The embodiment of the present application does not specifically limit the noise reduction model used by the electronic device 100, which can be a deep neural network model or other noise reduction model. It should be noted that for some users, during a call, the best listening experience is when the target voice is mixed with a little ambient sound. Therefore, by providing multiple noise reduction strengths, the flexibility of noise reduction can be improved, the noise reduction needs of different users can be met, and the users can get a better user experience.

[0199] For example, if the first semantic analysis result is "the current local audio is relatively low," it indicates that the current uplink audio gain is insufficient, and the electronic device 100 can select a higher uplink audio gain. For example, if the current uplink audio gain is 5dB, the electronic device 100 can adjust the uplink audio gain to a higher value, such as 8dB, 10dB, or 12dB.

[0200] As another example, if the first semantic analysis result is "the current sound on this end is relatively loud", it indicates that the current uplink audio gain is relatively large, and the electronic device 100 can select a smaller uplink audio gain. For example, if the current uplink audio gain is 10dB, the electronic device 100 can adjust the uplink audio gain to a smaller audio gain such as 8dB or 5dB. It is understandable that when a certain uplink audio gain (such as 5dB) is set, the electronic device 100 can perform corresponding gain processing (such as 5dB gain processing) on ​​the collected external audio, and then send the gain-processed audio data to the electronic device 200.

[0201] It is understandable that in some cases, the first semantic analysis result may also be "the external noise on the current end is relatively loud, and the current sound on the local end is relatively quiet". In this case, the electronic device 100 can adjust the uplink noise reduction intensity and the uplink audio gain at the same time.

[0202] It should be noted that, in a possible implementation, some semantic analysis results and corresponding adjustment strategies can be pre-configured. After that, the electronic device 100 can compare the first semantic analysis result with the pre-configured semantic analysis result. If the first semantic analysis result is the same as a pre-configured semantic analysis result, the corresponding adjustment strategy can be adopted. Among them, these pre-configured semantic analysis results can represent the situation where the audio effect of the call needs to be improved, and the adjustment strategy can be used to indicate the parameters in the audio processing strategy that need to be adjusted (such as noise reduction intensity, audio gain, etc.) and the adjustment method (such as selecting a higher-level parameter). Exemplarily, the pre-configured semantic analysis results may include "the current external noise on this end is relatively large", "the current sound on this end is relatively small", and "the current sound on this end is relatively large". Among them, the adjustment strategy corresponding to "the current external noise on this end is relatively large" can be to select a higher-level uplink noise reduction intensity, the adjustment strategy corresponding to "the current sound on this end is relatively small" can be to select a higher-level uplink audio gain, and the adjustment strategy corresponding to "the current sound on this end is relatively large" can be to select a lower-level uplink audio gain. In this way, when the first semantic analysis result is one of multiple preset semantic results, the electronic device 100 can adjust the uplink audio processing strategy based on the first semantic analysis result, that is, adjust the uplink audio processing strategy based on the adjustment strategy corresponding to the first semantic analysis result.

[0203] In the above implementation, the electronic device 100 mainly performs semantic analysis on the audio data from the electronic device 200, and then dynamically adjusts the uplink audio processing strategy based on the semantic analysis results. In addition to dynamically adjusting the uplink audio processing strategy based on semantic analysis, in an embodiment of the present application, the electronic device 100 can also analyze the audio data from the electronic device 200 (such as the first audio data), determine the noise level in the corresponding audio data, and then dynamically adjust the uplink audio processing strategy based on the real-time noise level. Exemplarily, the electronic device 100 can analyze the first audio data and determine the ratio of the noise to the target voice in the first audio data, such as the amplitude ratio of the noise to the target voice. If the amplitude ratio of the noise to the target voice is less than a threshold value 1 (such as 0.5), it indicates that the noise on the other end (such as the electronic device 200 or the user B side) is relatively small, and the uplink noise reduction intensity can be adjusted to weak. If the amplitude ratio of the noise to the target voice is greater than or equal to a threshold value 1 (such as 0.5) and less than a threshold value 2 (such as 1.5), it indicates that the noise on the other end is relatively large, and the uplink noise reduction intensity can be adjusted to medium. If the amplitude ratio of the noise to the target voice is greater than or equal to a threshold value 2, it indicates that the noise on the other end is very large, and the uplink noise reduction intensity can be adjusted to strong. Because the larger the amplitude ratio of the noise to the target voice, the noisier the environment on the other end is, and the noisier the environment on the other end is, the more affected the call of the user on the other end is, and the more likely it is that the user cannot hear clearly. Therefore, when the environment of the other end is noisier, the electronic device 100 can adopt a higher level of noise reduction intensity, so that the noise in the audio data sent to the other end can be reduced as much as possible, thereby reducing the impact on the user of the other end.

[0204] 904. The electronic device 200 sends first indication information to the electronic device 100, where the first indication information is used to indicate an audio quality problem of the electronic device 100.

[0205] During a call between electronic device 100 and electronic device 200, electronic device 200 can monitor the call quality of electronic device 100 in real time. When there is an audio quality problem with electronic device 100, electronic device 200 can send a first indication message to electronic device 100. The first indication message can be used to indicate the audio quality problem of electronic device 100.

[0206] Specifically, in a possible implementation, the electronic device 200 can score the audio data (such as the third audio data) from the electronic device 100 based on the voice quality detection algorithm, and obtain a real-time voice quality score. If the voice quality score is greater than the quality score threshold (such as 6, the voice quality score is 1 to 10), it indicates that the call quality of the electronic device 100 is good, and the electronic device 200 can do no processing. On the contrary, if the voice quality score is less than or equal to the quality score threshold, it indicates that the call quality of the electronic device 100 is poor. The electronic device 200 can determine the reason why the voice quality score is lower than the quality score threshold based on the corresponding audio data (such as the third audio data) (such as the current external noise on the local end is relatively large, the current sound on the local end is relatively small, etc.). Afterwards, the electronic device 200 can send a first indication information to the electronic device 100. The first indication information can be used to indicate the reason why the voice quality score is lower than the quality score threshold, that is, the audio quality problem existing in the third audio data. For example, assuming that the electronic device 200 scores the third audio data from the electronic device 100 and obtains a voice quality score of 5, the electronic device 200 can determine that the voice quality score is less than the quality score threshold (such as 6). After that, the electronic device 200 can analyze the third audio data to determine the reason why the voice quality score is lower than the quality score threshold. For example, by analyzing the third audio data, the electronic device 200 can determine that the reason why the voice quality score is lower than the quality score threshold is that "the third audio data includes too much ambient noise." It should be noted that the embodiment of the present application does not limit the way in which the electronic device 200 analyzes the audio quality problem existing in the third audio data. For example, the electronic device 200 can calculate the average amplitude of the third audio data and judge whether the sound is loud or small based on the average amplitude.

[0207] In the embodiments of the present application, there are multiple ways to implement the first indication information. For example, the first indication information can be text content that directly indicates the audio effect, such as "The current external noise on this end is relatively loud," "The current sound on this end is relatively quiet," or "The current sound on this end is relatively loud," or the first indication information can be an identifier of the audio effect, such as the identifier "0" can indicate "The current external noise on this end is relatively loud," the identifier "1" can indicate "The current sound on this end is relatively quiet," and the identifier "2" can indicate "The current sound on this end is relatively loud."

[0208] In the embodiment of the present application, a (dedicated) communication channel can be established between the electronic device 100 and the electronic device 200, and the electronic device 200 can send the first indication information through the communication channel.

[0209] It should be understood that the above-mentioned method for determining the audio quality problem of the electronic device 100 is merely illustrative and is not limited in the embodiments of the present application. For example, in some possible implementations, the voice quality detection algorithm used by the electronic device 200 can score various dimensions of the audio data from the electronic device 100, and can directly determine whether the corresponding audio data includes a lot of noise, and whether the audio amplitude is large or small.

[0210] 905. The electronic device 100 adjusts the uplink audio processing strategy based on the first indication information.

[0211] After the electronic device 100 receives the first indication information from the electronic device 200, the electronic device 100 may adjust the currently used uplink audio processing strategy based on the first indication information.

[0212] For example, if the first indication information indicates that "external noise at the current local end is relatively loud," it indicates that the current uplink noise reduction strength is insufficient, and the electronic device 100 may select a higher uplink noise reduction strength. For example, if the current uplink noise reduction strength is "weak," the electronic device 100 may adjust the uplink noise reduction strength to "medium" or "strong." If the current uplink noise reduction strength is "medium," the electronic device 100 may adjust the uplink noise reduction strength to "strong."

[0213] For another example, if the first indication information indicates "the current local audio is relatively low," it indicates that the current uplink audio gain is insufficient, and the electronic device 100 can select a higher uplink audio gain. For example, if the current uplink audio gain is 5dB, the electronic device 100 can adjust the uplink audio gain to a higher audio gain, such as 8dB, 10dB, or 12dB.

[0214] For another example, if the first indication information indicates "the current local audio is relatively loud," it indicates that the current uplink audio gain is relatively large, and the electronic device 100 can select a smaller uplink audio gain. For example, if the current uplink audio gain is 10dB, the electronic device 100 can adjust the uplink audio gain to a smaller audio gain, such as 8dB or 5dB.

[0215] It should be noted that, in a possible implementation, a corresponding adjustment strategy can be configured for each value of the first indication information. Afterwards, when the electronic device 100 receives the first indication information from the electronic device 200, it can directly adopt the corresponding adjustment strategy. The adjustment strategy can be used to indicate the parameters in the audio processing strategy that need to be adjusted (such as noise reduction intensity, audio gain, etc.) and the adjustment method (such as selecting a higher-level parameter). For example, when the first indication information is text content that directly represents the audio effect, it can include values ​​such as "the current external noise on this end is relatively large", "the current sound on this end is relatively small", and "the current sound on this end is relatively large". The adjustment strategy corresponding to "the current external noise on this end is relatively large" can be to select a higher-level uplink noise reduction intensity, the adjustment strategy corresponding to "the current sound on this end is relatively small" can be to select a higher-level uplink audio gain, and the adjustment strategy corresponding to "the current sound on this end is relatively large" can be to select a lower-level uplink audio gain. In this way, the electronic device 100 can adjust the uplink audio processing strategy based on the adjustment strategy corresponding to the first indication information.

[0216] In the above implementation, the electronic device 200 can feedback quality problems or indication information to the electronic device 100, and then the electronic device 100 can adjust the uplink audio processing strategy based on the quality problems or indication information. In this way, the electronic device 100 usually adjusts the uplink audio processing strategy based on the corresponding configuration, such as configuring a corresponding adjustment strategy for each value of the first indication information. In addition to this implementation, in an embodiment of the present application, the electronic device 100 can also receive a request message from the electronic device 200 and adjust the uplink audio processing strategy based on the request message. For example, the electronic device 200 can provide a user interface for configuring the audio processing of the opposite end, and the user can select one or more configurations such as noise reduction scene, audio gain and noise reduction intensity based on the user interface. After that, the user can click or touch the corresponding confirmation control (control 4822 as shown in Figure 8D). In response to this operation, the electronic device 200 can send a second request message to the electronic device 100. The second request message is used to request adjustment of the uplink audio processing strategy. The second request message includes the parameters in the uplink audio processing strategy that need to be adjusted (such as noise reduction scene) and the corresponding parameter values ​​(shopping mall). Correspondingly, the electronic device 100 may receive a second request message from the electronic device 200, after which the electronic device 100 may pop up a corresponding prompt information box, which is used to prompt the user of the uplink audio processing strategy requested by the opposite end to be adjusted. The prompt information box may also include corresponding confirmation controls and rejection controls, the confirmation control may be used to confirm or receive the opposite end's request, and the rejection control may be used to reject the opposite end's request. In some possible implementations, after the electronic device 100 receives the second request message from the electronic device 200, it may adjust the uplink audio processing strategy directly based on the second request message, such as switching the noise reduction scenario from the current "general" to "shopping mall." For the specific user interface corresponding to the above description, reference may be made to the user interfaces shown in FIG8A-FIG8E.

[0217] In some possible implementations, when the smart call mode of the electronic device 100 is not enabled, the electronic device 200 may further send a first request message to the electronic device 100. The first request message may be used to request that the smart call mode be enabled. Accordingly, the electronic device 100 may receive the first request message from the electronic device 200. Afterwards, the electronic device 100 may pop up a corresponding prompt message box, which prompts the user to request that the other party enable the smart call mode. The prompt message box may also include corresponding confirmation controls and rejection controls. The confirmation control may be used to confirm or receive the request from the other party, and the rejection control may be used to reject the request from the other party. In some possible implementations, after receiving the first request message from the electronic device 200, the electronic device 100 may directly enable the smart call mode based on the first request message. In embodiments of the present application, there may be multiple situations in which the electronic device 200 sends the first request message, the following examples of which are illustrative. First, the electronic device 200 may respond to a user operation and then send the first request message to the electronic device 100. As shown in FIG8B , the electronic device 200 may send the first request message to the electronic device 100 in response to a user operation on the send request control 4811. Second, the electronic device 200 can score the audio data (such as the third audio data) from the electronic device 100 based on the voice quality detection algorithm, and obtain a real-time voice quality score. If the voice quality score is less than or equal to the quality score threshold (such as 6), it indicates that the audio quality of the audio data from the electronic device 100 is poor. In this case, the electronic device 200 can send a first request message to the electronic device 100.

[0218] 906. The electronic device 100 adjusts the uplink noise reduction scenario based on the first positioning result.

[0219] In an embodiment of the present application, a variety of uplink noise reduction scenarios / audio processing scenarios are provided, and different uplink noise reduction scenarios correspond to different noise reduction models. Therefore, in order to obtain a better uplink noise reduction effect, the electronic device 100 can dynamically adjust the uplink noise reduction scenario based on the first positioning result. Among them, the first positioning result can be a positioning result obtained during the call between the electronic device 100 and the electronic device 200, or it can be a positioning result obtained before the electronic device 100 and the electronic device 200 have a call. For example, when the electronic device 100 and the electronic device 200 just have a call, the electronic device 100 can initiate positioning (such as satellite positioning), and the obtained positioning result can be used as the first positioning result. For another example, some time before the electronic device 100 and the electronic device 200 have a call (such as the first few seconds), the electronic device 100 initiates positioning and obtains a positioning result, which can be used as the first positioning result.

[0220] For example, during a call between electronic device 100 and electronic device 200, assuming the currently set uplink noise reduction scenario is "General," if the first positioning result is "XX Shopping Mall, District B, City A," electronic device 100 can switch the uplink noise reduction scenario to "Shopping Mall." Similarly, if the first positioning result is "XX Subway Station," electronic device 100 can switch the uplink noise reduction scenario to "Subway Station."

[0221] It is understandable that since the position of the electronic device 100 may change in real time, during the call between the electronic device 100 and the electronic device 200, the electronic device 100 can periodically (such as 30 seconds) initiate positioning to facilitate timely adjustment of the noise reduction scene.

[0222] It can be understood that in the above implementation, the electronic device 100 can adjust the uplink noise reduction scene based on the positioning result. In another possible implementation, the electronic device 100 can also adjust the uplink noise reduction scene based on semantic analysis. Specifically, the electronic device 100 can obtain the audio data of the local end and the audio data from the electronic device 200 in real time, and then perform real-time semantic analysis, and can dynamically adjust the uplink noise reduction scene based on the semantic analysis results. For example, assuming that the conversation between user A and user B includes (user B: where are you, user A: I am in the mall), in this case, the electronic device 100 performs semantic analysis on the audio data of the local end and the audio data from the electronic device 200, and the obtained semantic analysis result can be "the current local end is in the mall", so the electronic device 100 can switch the uplink noise reduction scene to the mall.

[0223] It can be understood that the three processing methods of dynamically adjusting the uplink audio processing strategy based on semantic analysis, dynamically adjusting the uplink audio processing strategy based on the first indication information, and dynamically adjusting the uplink noise reduction scenario based on the positioning results can be parallel technical means and have no dependency on each other. Therefore, in some possible implementations, the electronic device 100 can adopt any one or more of these three methods for processing, and the embodiments of the present application are not limited here. For example, the electronic device 100 can dynamically adjust the uplink audio processing strategy based only on semantic analysis. For another example, the electronic device can dynamically adjust the uplink audio processing strategy based only on the first indication information. For another example, the electronic device 100 can dynamically adjust the uplink noise reduction scenario based only on the positioning results.

[0224] The embodiment of the present application does not limit the execution order of steps 901 to 903, steps 904 to 905, and step 906.

[0225] In the embodiment of the present application, audio processing may include uplink audio processing and downlink audio processing. FIG9 above mainly illustrates the case where the electronic device 100 includes only the uplink audio processing function. The downlink audio processing is similar to the uplink audio processing and can be referenced to each other. The following briefly illustrates the case where the electronic device 100 includes only the downlink audio processing function.

[0226] Please refer to Figure 10, which is a flowchart of another audio processing method disclosed in an embodiment of the present application. As shown in Figure 10, the method may include but is not limited to the following steps:

[0227] 1001. During a call between the electronic device 100 and the electronic device 200, the electronic device 100 performs semantic analysis based on the second audio data to obtain a second semantic analysis result.

[0228] For example, assuming that user A holds electronic device 100 and user B holds electronic device 200, when user A needs to contact user B, user A and user B can make a call (such as a voice call, video call, etc.) through electronic device 100 and electronic device 200.

[0229] During a call between electronic device 100 and electronic device 200, electronic device 100 can collect external sounds, such as external ambient sounds and the voice of user A, through its own microphone in real time, and can send corresponding audio data (such as second audio data) to electronic device 200. Accordingly, electronic device 200 can receive audio data from electronic device 100. Similarly, electronic device 200 can also collect external sounds, such as external ambient sounds and the voice of user B, through its own microphone in real time, and can send corresponding audio data to electronic device 100. Accordingly, electronic device 100 can receive audio data from electronic device 200.

[0230] In an embodiment of the present application, the electronic device 100 can perform semantic analysis on the collected external audio in real time, so as to dynamically adjust the current downlink audio processing strategy based on the semantic analysis results.

[0231] Exemplarily, the electronic device 100 may obtain the second audio data collected by itself, and then the electronic device 100 may perform semantic analysis on the second audio data to obtain a second semantic analysis result. For example, the second audio data may be "It's so noisy over there", "The noise over there is so loud", "It's a bit noisy over there, I can't hear clearly", "It's still so noisy over there", "The noise over there is still so loud", etc., and the second semantic analysis results corresponding to these audio data may be "The current noise on the other end is relatively loud". For another example, the second audio data may be "Your voice is so soft", "I can't hear what you are saying", "Can you speak louder", "Your voice is still a little soft", etc., and the corresponding second semantic analysis results may be "The current voice on the other end is relatively soft". For another example, the second audio data may be "Your voice is so loud", "Can you speak more quietly", "Your voice is too loud", etc., and the corresponding second semantic analysis results may be "The current voice on the other end is relatively loud".

[0232] In some possible implementations, the electronic device 100 may perform semantic analysis on the second audio data only when the smart call mode is enabled, to dynamically adjust the current downlink audio processing strategy based on the semantic analysis results. In other words, when the smart call mode is disabled, the electronic device 100 may not perform any processing. In other possible implementations, the electronic device 100 may also perform semantic analysis on the second audio data when the smart call mode is disabled, to determine whether to enable the smart call mode based on the semantic analysis results. Furthermore, after enabling the smart call mode, the electronic device 100 may continue to perform semantic analysis on subsequent collected audio data, to dynamically adjust the current downlink audio processing strategy based on the semantic analysis results. For example, the electronic device 100 may automatically enable the smart call mode when the semantic analysis results indicate "the current peer noise is relatively loud," "the current peer voice is relatively quiet," or "the current peer voice is relatively loud," to perform audio processing such as downlink noise reduction and audio amplitude adjustment. In this manner, during a call, the smart call mode can be automatically enabled based on the semantic analysis results, promptly meeting user needs while ensuring call quality and reducing the overall power consumption of the electronic device.

[0233] 1002. The electronic device 100 adjusts the downlink audio processing strategy based on the second semantic analysis result.

[0234] After obtaining the second semantic analysis result, the electronic device may adjust the currently used downlink audio processing strategy based on the second semantic analysis result.

[0235] It is understandable that the conversation between user A and user B usually includes a lot of content that does not involve audio quality or audio effects (such as noise, sound volume, etc.), but only the semantic analysis results corresponding to the conversation content related to audio quality will affect the audio processing strategy. Therefore, the electronic device 100 can adjust the currently used uplink audio processing strategy based on the second semantic analysis result when the second semantic analysis result is related to the audio quality. In other words, the electronic device 100 can adjust the currently used uplink audio processing strategy based on the second semantic analysis result when it is confirmed based on the second semantic analysis result that the audio effect of the call needs to be improved. For example, if the second semantic analysis result is "the current noise on the other end is relatively loud", "the current sound on the other end is relatively quiet", or "the current sound on the other end is relatively loud", etc., it indicates that the audio effect of the call is poor, and the electronic device 100 can determine that the audio effect of the call needs to be improved.

[0236] For example, if the second semantic analysis result is "the current noise on the other end is relatively loud", it indicates that the current downlink noise reduction strength is insufficient, and the electronic device 100 can select a larger downlink noise reduction strength. For example, if the current downlink noise reduction strength is "weak", the electronic device 100 can adjust the downlink noise reduction strength to "medium" or "strong". If the current downlink noise reduction strength is "medium", the electronic device 100 can adjust the downlink noise reduction strength to "strong". In the embodiment of the present application, a corresponding noise reduction model can be configured for each downlink noise reduction strength. When a certain downlink noise reduction strength is selected, the audio data received from the electronic device 200 can be processed by the noise reduction model corresponding to the downlink noise reduction strength.

[0237] For example, if the second semantic analysis result is "the current peer voice is relatively quiet," it indicates that the current downlink audio gain is insufficient, and the electronic device 100 can select a higher downlink audio gain. For example, if the current downlink audio gain is 5dB, the electronic device 100 can adjust the downlink audio gain to a higher value, such as 8dB, 10dB, or 12dB.

[0238] As another example, if the second semantic analysis result is "the current sound of the other end is relatively loud", it indicates that the current downlink audio gain is relatively large, and the electronic device 100 can select a smaller downlink audio gain. For example, if the current downlink audio gain is 10dB, the electronic device 100 can adjust the downlink audio gain to a smaller audio gain such as 8dB or 5dB. It is understandable that when a certain downlink audio gain (such as 5dB) is set, the electronic device 100 can perform corresponding downlink gain processing (such as 5dB gain processing) on ​​the audio received from the electronic device 200, and then play the corresponding audio through the earpiece or speaker.

[0239] It should be noted that in one possible implementation, some semantic analysis results and corresponding adjustment strategies can be pre-configured. The electronic device 100 can then compare the second semantic analysis result with the pre-configured semantic analysis results. If the second semantic analysis result matches a pre-configured semantic analysis result, the corresponding adjustment strategy can be adopted. These pre-configured semantic analysis results can represent situations where the audio quality of a call needs to be improved, and the adjustment strategy can be used to indicate the parameters in the audio processing strategy that need to be adjusted and the adjustment method. For example, the pre-configured semantic analysis results may include "the current peer noise is relatively loud," "the current peer voice is relatively quiet," and "the current peer voice is relatively loud." The adjustment strategy corresponding to "the current peer noise is relatively loud" can be to select a higher level of downlink noise reduction intensity, the adjustment strategy corresponding to "the current peer voice is relatively quiet" can be to select a higher level of downlink audio gain, and the adjustment strategy corresponding to "the current peer voice is relatively loud" can be to select a lower level of downlink audio gain. Thus, when the second semantic analysis result is one of multiple pre-set semantic results, the electronic device 100 can adjust the downlink audio processing strategy based on the second semantic analysis result, that is, adjust the downlink audio processing strategy based on the adjustment strategy corresponding to the second semantic analysis result.

[0240] 1003. The electronic device 200 sends the fourth audio data to the electronic device 100.

[0241] During a call between electronic device 100 and electronic device 200, electronic device 200 may collect external sounds through a microphone, obtain corresponding fourth audio data, and then send the fourth audio data to electronic device 100. Accordingly, electronic device 100 may receive the fourth audio data from electronic device 200.

[0242] 1004. The electronic device 100 performs audio processing on the fourth audio data to obtain fifth audio data.

[0243] After receiving the fourth audio data, the electronic device 100 can process the fourth audio data according to the current downlink audio processing strategy to obtain the fifth audio data. In other words, the fifth audio data is the audio data obtained by the electronic device 100 performing noise reduction processing, gain processing, etc. according to the current downlink audio processing strategy.

[0244] 1005. The electronic device 100 analyzes the fifth audio data and determines an audio quality problem of the fifth audio data.

[0245] During a call between the electronic device 100 and the electronic device 200, the electronic device 100 can process the audio data from the electronic device 200 in real time through the downlink audio processing strategy of the local end, and can monitor the call quality of the processed audio data (such as the fifth audio data) in real time. When there is an audio quality problem in the processed audio data, the electronic device 100 can adjust the currently used downlink audio processing strategy based on the audio quality problem.

[0246] Specifically, in one possible implementation, the electronic device 100 can score the fifth audio data based on a voice quality detection algorithm and obtain a real-time voice quality score. If the voice quality score is greater than a quality score threshold (e.g., 6, where the voice quality score is 1 to 10), it indicates that the electronic device 100 can perform no processing after processing the audio from the electronic device 200. If the voice quality score is less than or equal to the quality score threshold, it indicates that the call quality of the electronic device 200 is poor. The electronic device 100 can determine the reason why the voice quality score is lower than the quality score threshold based on the corresponding audio data (e.g., the current external noise is relatively loud, the current sound is relatively quiet, etc.). For example, assuming that the electronic device 100 scores the fifth audio data from the electronic device 200 and obtains a voice quality score of 5, the electronic device 100 can determine that the voice quality score is lower than the quality score threshold (e.g., 6). Afterwards, the electronic device 100 can analyze the fifth audio data to determine the reason why the voice quality score is lower than the quality score threshold. For example, by analyzing the fifth audio data, the electronic device 100 can determine that the reason why the voice quality score is lower than the quality score threshold is that "the fifth audio data includes too much ambient noise."

[0247] 1006. The electronic device 100 adjusts the downlink audio processing strategy based on the audio quality problem of the fifth audio data.

[0248] After determining the audio quality problem of the fifth audio data, the electronic device 100 may adjust the currently used downlink audio processing strategy based on the audio quality problem of the fifth audio data.

[0249] For example, if the audio quality problem of the fifth audio data is "loud external noise", it indicates that the current downlink noise reduction strength is insufficient, and the electronic device 100 can select a higher downlink noise reduction strength. For example, if the current downlink noise reduction strength is "weak", the electronic device 100 can adjust the downlink noise reduction strength to "medium" or "strong". If the current downlink noise reduction strength is "medium", the electronic device 100 can adjust the downlink noise reduction strength to "strong".

[0250] For another example, if the audio quality problem of the fifth audio data is "relatively low sound", it indicates that the current downlink audio gain is insufficient, and the electronic device 100 can select a larger downlink audio gain. For example, if the current downlink audio gain is 5dB, the electronic device 100 can adjust the downlink audio gain to a larger audio gain, such as 8dB, 10dB, or 12dB.

[0251] For another example, if the audio quality problem of the fifth audio data is "the current sound is relatively loud," it indicates that the current downlink audio gain is large, and the electronic device 100 can select a smaller downlink audio gain. For example, if the current downlink audio gain is 10dB, the electronic device 100 can adjust the downlink audio gain to a smaller audio gain, such as 8dB or 5dB.

[0252] 1007. The electronic device 100 adjusts the downlink noise reduction scenario based on the second positioning result.

[0253] In an embodiment of the present application, a variety of downlink noise reduction scenarios / audio processing scenarios are provided, and different downlink noise reduction scenarios correspond to different noise reduction models. Therefore, in order to obtain a better noise reduction effect, the electronic device 100 can dynamically adjust the downlink noise reduction scenario based on the second positioning result. Among them, the second positioning result can be a positioning result obtained during the call between the electronic device 100 and the electronic device 200, or it can be a positioning result obtained before the electronic device 100 and the electronic device 200 have a call. For example, when the electronic device 100 and the electronic device 200 just have a call, the electronic device 200 can initiate positioning (such as satellite positioning), and the obtained positioning result can be used as the second positioning result, and the second positioning result can be sent to the electronic device 100. For another example, some time before the electronic device 100 and the electronic device 200 have a call (such as the first few seconds), the electronic device 200 initiates positioning and obtains a positioning result, which can be used as the second positioning result.

[0254] It is understandable that since the position of the electronic device 200 may change in real time, during the call between the electronic device 100 and the electronic device 200, the electronic device 200 can initiate positioning periodically (such as 30 seconds) and can send the positioning results to the electronic device 100 so that the electronic device 100 can adjust the downlink noise reduction scenario in time.

[0255] In another possible implementation, the electronic device 100 can also adjust the downlink noise reduction scene based on semantic analysis. Specifically, the electronic device 100 can obtain the audio data of the local end and the audio data from the electronic device 200 in real time, and then perform real-time semantic analysis, and can dynamically adjust the downlink noise reduction scene based on the semantic analysis results. For example, assuming that the conversation between user A and user B includes (user A: where are you, user B: I am in the mall), in this case, the electronic device 100 performs semantic analysis on the audio data of the local end and the audio data from the electronic device 200, and the obtained semantic analysis result can be "the other end is currently in the mall", so the electronic device 100 can switch the downlink noise reduction scene to the mall.

[0256] The embodiment of the present application does not limit the execution order of steps 1001 to 1003, steps 1004 to 1005, and step 1006.

[0257] The above examples illustrate uplink audio processing and downlink audio processing, respectively. However, it should be understood that in some possible implementations, electronic devices can support both uplink and downlink audio processing, and this is not a limitation of the present application. Furthermore, for the same noise reduction scenario and at the same noise reduction intensity, the noise reduction model used for both uplink and downlink can be the same.

[0258] In the above processing flow, for uplink audio processing and downlink audio processing, the electronic device 100 can automatically adjust the noise reduction strength, audio gain, and noise reduction scenario based on the semantic analysis results, and can also automatically adjust the noise reduction scenario based on positioning. In addition, it also supports users to manually adjust the noise reduction strength, audio gain, and noise reduction scenario. In this way, the most appropriate noise reduction model and corresponding audio gain can be selected according to the actual situation, which can improve call quality and enable users to have a better call experience.

[0259] It should be understood that the transmission in the embodiments of the present application can be direct transmission or indirect transmission. Direct transmission means that a device or module sends information / data directly to a corresponding device or module, and indirect transmission means that a device or module sends information / data to a corresponding device or module through another device or module.

[0260] Obviously, the embodiments described above are only some of the embodiments of this application, and not all of them. Reference to "embodiments" herein means that the specific features, structures, or characteristics described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it represent an independent or alternative embodiment that is mutually exclusive with other embodiments. Those skilled in the art will understand, both explicitly and implicitly, that the embodiments described herein may be combined with other embodiments. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application. The terms "including," "having," and any variations thereof, as used herein, are intended to cover non-exclusive inclusions. For example, a description may include a series of steps or units, or alternatively, steps or units not listed, or other steps or units inherent to the process, method, product, or device. It will be understood that, in some embodiments, the equal sign in the above-mentioned conditional judgments may be greater than or less than. For example, the above-mentioned conditional judgment regarding a threshold being greater than, less than, or equal to may be modified to mean that the threshold is greater than, equal to, or less than, without limitation. It can also be understood that, for an architecture with multiple devices or modules, if one of the devices or modules generates information and another device or module uses the information, there may be multiple ways for the other device to obtain the information. For example, the device or module that generates the information may send the information directly to the device or module that uses the information (equivalent to direct sending), or the device or module that generates the information may send the information to the device or module that uses the information through other devices or modules (equivalent to indirect sending).

[0261] It will be appreciated that only the parts relevant to the present application, not all of the contents, are shown in the accompanying drawings. It will be appreciated that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the various operations (or steps) as sequential processes, many of the operations therein can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the various operations can be rearranged as long as it is logical. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0262] As used in this specification, the terms "component," "module," "system," "unit," and the like are used to refer to computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a unit can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or distributed between two or more computers. In addition, these units can be executed from various computer-readable media having various data structures stored thereon. For example, a unit can communicate through local and / or remote processes based on signals having one or more data packets (e.g., data from a second unit interacting with another unit in a local system, a distributed system, and / or a network. For example, the Internet interacts with other systems via signals).

[0263] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of this application. It should be understood that the above description is only the specific implementation methods of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of this application should be included in the scope of protection of this application.

Claims

1. An audio processing method, applied to a first electronic device, characterized in that: include: During a call between the first electronic device and the second electronic device, performing semantic analysis on call audio data to obtain a target semantic analysis result; the call audio data includes call audio data from the second electronic device; Adjusting the audio processing strategy based on the target semantic analysis result; The audio processing strategy includes an uplink audio processing strategy, and the uplink audio processing strategy is used to process the call audio data collected by the first electronic device.

2. The method according to claim 1, characterized in that The uplink audio processing strategy includes one or more of uplink noise reduction strength, uplink audio gain, and uplink noise reduction scenario.

3. The method according to claim 1 or 2, characterized in that The adjusting the audio processing strategy based on the target semantic analysis result includes: In the case where the target semantic analysis result is one of a plurality of preset semantic analysis results, The audio processing strategy is adjusted based on the adjustment strategy corresponding to the target semantic result; the multiple preset semantic analysis results are each configured with a corresponding adjustment strategy, and the adjustment strategy is used to indicate the parameters in the audio processing strategy that need to be adjusted and the adjustment method.

4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: Initiate positioning and obtain the first positioning result; The uplink noise reduction scenario is adjusted based on the first positioning result, and different noise reduction scenarios correspond to different noise reduction models.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: receiving first indication information from the second electronic device, where the first indication information is used to indicate an audio quality problem of the first electronic device; Adjust the uplink audio processing strategy based on the first indication information.

6. The method according to claim 5, characterized in that The adjusting the uplink audio processing strategy based on the first indication information includes: Adjust the uplink audio processing strategy based on the adjustment strategy corresponding to the first indication information; the first indication information includes multiple value situations, and the multiple value situations are all configured with corresponding adjustment strategies, and the adjustment strategy is used to indicate the parameters in the audio processing strategy that need to be adjusted and the adjustment method.

7. The method according to any one of claims 1 to 6, characterized in that Before adjusting the audio processing strategy based on the target semantic analysis result, the method further includes: Performing semantic analysis on the sixth audio data to obtain a sixth semantic analysis result; the sixth audio data includes call audio data from the second electronic device and / or call audio data collected by the first electronic device; When the sixth semantic analysis result is one of a plurality of preset semantic analysis results, enabling the smart call mode; The adjusting the audio processing strategy based on the target semantic analysis result includes: When the smart call mode is turned on, the audio processing strategy is adjusted based on the target semantic analysis result.

8. The method according to any one of claims 1 to 6, characterized in that Before adjusting the audio processing strategy based on the target semantic analysis result, the method further includes: receiving a first request message from the second electronic device, where the first request message is used to request to enable a smart call mode; Enable smart call mode based on the first request message; The adjusting the audio processing strategy based on the target semantic analysis result includes: When the smart call mode is turned on, the audio processing strategy is adjusted based on the target semantic analysis result.

9. The method according to any one of claims 1 to 8, characterized in that The method further comprises: receiving a second request message from the second electronic device, where the second request message is used to request adjustment of the uplink audio processing policy, and the second request message includes parameters in the uplink audio processing policy that need to be adjusted and corresponding parameter values; Adjust the uplink audio processing strategy based on the second request message.

10. The method according to any one of claims 1 to 9, characterized in that The method further comprises: displaying a first user interface, the first user interface including audio processing configuration controls; In response to a user operation on the audio processing configuration control, a second user interface is displayed; the second user interface includes an upstream audio processing configuration area, the upstream audio processing configuration area includes one or more of a noise reduction intensity configuration area, an audio gain configuration area, and a noise reduction scene configuration area, the noise reduction intensity configuration area includes a plurality of noise reduction intensity options, the audio gain configuration area includes a plurality of audio gain options, and the noise reduction scene configuration area includes a plurality of noise reduction scene options; In response to a user operation on a first option, the first option is changed to a selected state, and an audio processing strategy corresponding to the first option is enabled; the first option is any option in the second user interface.

11. An electronic device, characterized in that: The method comprises a processor and a communication interface; the communication interface is used to receive and send data; the processor calls a computer program or computer instruction stored in a memory to implement the method according to any one of claims 1 to 10.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program or computer instructions, and the computer program or computer instructions are executed by a processor to implement the method according to any one of claims 1 to 10.

13. A computer program product, characterized in that The computer program product comprises computer program codes or computer instructions, and when the computer program codes or computer instructions are executed, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Method for adjusting volume automatically, volume adjusting apparatus and electronic apparatus

    CN104335559A

  • Voice signal processing method, device, and mobile terminal

    CN105825854A

  • Intelligent video call method and system

    CN110351515A

  • Call tone quality adjusting method and electronic equipment

    CN115665318A

  • Audio processing method, electronic equipment and computer readable storage medium

    CN119182849A