Audio processing method, electronic device, and computer readable storage medium
By performing semantic analysis on call audio data and dynamically adjusting audio processing strategies, the problem of poor call quality in noisy environments was solved, resulting in a better call experience and audio clarity.
Patent Information
- Application Number
- PCT/CN2025/079322
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-29
- Filing Date
- 2025-02-26
- Publication Date
- 2026-01-29
AI Technical Summary
When making a call in a noisy environment, the call quality is poor, and the user cannot clearly hear the other party's voice.
By performing semantic analysis on call audio data, the uplink audio processing strategies of electronic devices can be dynamically adjusted, such as noise reduction intensity, audio gain, and noise reduction scenarios, to improve call quality.
It improves call quality, giving users a better call experience, reduces noise interference, and enhances the clarity of audio data.
Smart Images

Figure CN2025079322_29012026_PF_FP_ABST
Abstract
Description
Audio processing methods, electronic devices and computer-readable storage media
[0001] This application claims priority to Chinese Patent Application No. 202410233300.8, filed on February 29, 2024, entitled "Audio Processing Method, Electronic Device and Computer-Readable Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of terminal technology, and in particular to audio processing methods, electronic devices, and computer-readable storage media. Background Technology
[0003] As the processing power of mobile phones, tablets and other terminal devices becomes stronger, audio processing technologies, mainly audio noise reduction, have been widely used in these devices.
[0004] During daily phone calls, one or both parties may be in a noisy environment (such as a shopping mall, high-speed rail station, or airport with many people), which makes it difficult for one party to hear what the other party is saying, resulting in poor call quality. Summary of the Invention
[0005] This application discloses an audio processing method, an electronic device, and a computer-readable storage medium. During a call, by performing semantic analysis on the call audio data from the other end, the uplink audio processing strategy of the electronic device can be dynamically adjusted, thereby improving call quality and providing users with a better call experience.
[0006] The first aspect discloses an audio processing method, which can be applied to a first electronic device, a module (e.g., a processor) within the first electronic device, or a logic module or software capable of implementing all or part of the functions of the first electronic device. The following description uses an application to a first electronic device as an example. The communication method may include: performing semantic analysis on call audio data during a conversation between the first electronic device and a second electronic device to obtain a target semantic analysis result; the call audio data includes call audio data from the second electronic device; adjusting an audio processing strategy based on the target semantic analysis result; wherein the audio processing strategy includes an uplink audio processing strategy, which is used to process the call audio data collected by the first electronic device.
[0007] In this embodiment, during a call, the first electronic device can perform semantic analysis on the call audio data (such as "Your voice is too soft," or "It's too noisy on your end, I can't hear you") from the second electronic device to obtain semantic analysis results. Based on the semantic analysis results, the first electronic device can understand the actual call experience of the user at the other end (the second electronic device) or the call audio quality at its own end (the first electronic device), such as whether the noise at its own end is too loud or whether the volume at its own end is too low. Therefore, the first electronic device can dynamically adjust its uplink audio processing strategy based on the semantic analysis results (such as selecting a higher level of noise reduction intensity if the other end reports that the noise at its own end is too loud). Afterwards, the first electronic device can process the call audio data collected by itself using the adjusted uplink audio processing strategy, which can improve the audio quality of the call audio data sent to the second electronic device, thus providing a better call experience for the user at the other end.
[0008] In conjunction with the first aspect, in one possible implementation, the uplink audio processing strategy includes one or more of uplink noise reduction intensity, uplink audio gain, and uplink noise reduction scenario.
[0009] In this embodiment, the uplink audio processing strategy can include multiple options such as uplink noise reduction intensity, uplink audio gain, and uplink noise reduction scenario. Uplink noise reduction intensity can be used to adjust the degree of noise elimination, meeting the noise reduction needs of different users and various scenarios. For example, weak noise reduction intensity can eliminate a small portion of noise, medium noise reduction intensity can eliminate most noise, and strong noise reduction intensity can completely eliminate noise. Uplink audio gain can be used to adjust the amplitude of the audio data transmitted to the second electronic device, i.e., to adjust the volume. Adjusting the uplink audio gain can help the user at the other end hear more clearly. Regarding the noise reduction scenario, different noise reduction scenarios can correspond to different noise reduction models. By adjusting the appropriate noise reduction scenario, the best noise reduction effect can be obtained, minimizing the impact on the target human voice while reducing noise.
[0010] In conjunction with the first aspect, in one possible implementation, the call audio data also includes call audio data collected by the first electronic device.
[0011] In this embodiment of the application, the call audio data also includes the call audio data collected by the first electronic device, that is, the call audio data of the local end. In this way, the first electronic device can obtain more complete and accurate semantic analysis results based on the dialogue context (the dialogue between the local user and the remote user), thereby ensuring that the uplink noise reduction strategy can be adjusted more accurately.
[0012] In conjunction with the first aspect, in one possible implementation, the audio processing strategy also includes a downlink audio processing strategy.
[0013] In conjunction with the first aspect, in one possible implementation, the downlink audio processing strategy includes one or more of downlink noise reduction intensity, downlink audio gain, and downlink noise reduction scenario.
[0014] In this embodiment, the first electronic device can also dynamically adjust the downlink audio processing strategy based on the semantic analysis results of the call audio data, such as noise reduction scenarios and noise reduction intensity, so that the user on this end can have a better call experience.
[0015] In conjunction with the first aspect, in one possible implementation, adjusting the audio processing strategy based on the target semantic analysis result includes: when the target semantic analysis result is one of a plurality of preset semantic analysis results, adjusting the audio processing strategy based on the adjustment strategy corresponding to the target semantic result; each of the plurality of preset semantic analysis results is configured with a corresponding adjustment strategy, and the adjustment strategy is used to indicate the parameters and adjustment method in the audio processing strategy that needs to be adjusted.
[0016] In this embodiment, multiple preset semantic analysis results and corresponding adjustment strategies for each preset semantic analysis result can be configured in advance. This ensures that the first electronic device can efficiently and accurately adjust the uplink audio processing strategy based on the semantic analysis results of the call audio data.
[0017] In conjunction with the first aspect, in one possible implementation, the method further includes: initiating positioning to obtain a first positioning result; adjusting the uplink noise reduction scenario based on the first positioning result, wherein different noise reduction scenarios correspond to different noise reduction models.
[0018] In this embodiment of the application, the uplink noise reduction scenario in which the local terminal is located can be accurately determined by positioning, so that the most suitable noise reduction model can be used for noise reduction processing, thereby ensuring the best noise reduction effect.
[0019] In conjunction with the first aspect, in one possible implementation, the method further includes: receiving first indication information from the second electronic device, the first indication information being used to indicate an audio quality problem of the first electronic device; and adjusting an uplink audio processing strategy based on the first indication information.
[0020] In this embodiment of the application, the first electronic device may also adjust the uplink audio processing strategy based on the first instruction information from the second electronic device in order to improve the quality of the call audio and enable the other end user to have a better call experience.
[0021] In conjunction with the first aspect, in one possible implementation, adjusting the uplink audio processing strategy based on the first indication information includes: adjusting the uplink audio processing strategy based on the adjustment strategy corresponding to the first indication information; the first indication information includes multiple value cases, each of which is configured with a corresponding adjustment strategy, and the adjustment strategy is used to indicate the parameters and adjustment method in the audio processing strategy that needs to be adjusted.
[0022] In this embodiment of the application, multiple possible scenarios can be determined in advance, each scenario can correspond to a value of an indication information, and a corresponding adjustment strategy can be pre-configured for each value of the indication information. In this way, it can be ensured that the first electronic device can efficiently and accurately adjust the uplink audio processing strategy based on the first indication information.
[0023] In conjunction with the first aspect, in one possible implementation, before adjusting the audio processing strategy based on the target semantic analysis result, the method further includes: performing semantic analysis on the sixth audio data to obtain a sixth semantic analysis result; the sixth audio data includes call audio data from the second electronic device and / or call audio data collected by the first electronic device; if the sixth semantic analysis result is one of a plurality of preset semantic analysis results, activating an intelligent call mode; the adjustment of the audio processing strategy based on the target semantic analysis result includes: adjusting the audio processing strategy based on the target semantic analysis result when the intelligent call mode is activated.
[0024] In this embodiment, since the first electronic device is usually in an environment with low or almost no noise in most cases, the smart call mode can be turned off by default each time the first electronic device makes a call. In this case, the first electronic device does not need to perform noise reduction or other processing. During the call, the first electronic device can perform semantic analysis on the call audio data from the second electronic device. When the semantic analysis determines that there is a problem with the call audio quality (such as high noise or low volume at the local end), the first electronic device can turn on the smart call mode again to perform noise reduction or other processing. It can be seen that in this way, the first electronic device can only turn on the smart call mode to perform noise reduction or other processing when the semantic analysis results determine that specific conditions are met. In this way, meaningless noise reduction processing by the first electronic device can be avoided as much as possible, and the overall power consumption of the first electronic device can be reduced.
[0025] In conjunction with the first aspect, in one possible implementation, before adjusting the audio processing strategy based on the target semantic analysis result, the method further includes: receiving a first request message from the second electronic device, the first request message being used to request the activation of a smart call mode; activating the smart call mode based on the first request message; the adjustment of the audio processing strategy based on the target semantic analysis result includes: adjusting the audio processing strategy based on the target semantic analysis result when the smart call mode is activated.
[0026] In conjunction with the first aspect, in one possible implementation, the method further includes: receiving a second request message from the second electronic device, the second request message being used to request adjustment of an uplink audio processing strategy, the second request message including parameters in the uplink audio processing strategy to be adjusted and corresponding parameter values; and adjusting the uplink audio processing strategy based on the second request message.
[0027] In this embodiment, the first electronic device can also directly activate the smart call mode based on a request message from the second electronic device, or directly adjust the uplink audio processing strategy. This improves the flexibility of activating the smart call mode and adjusting the uplink audio processing strategy. For example, the second electronic device can provide corresponding controls for requesting the other end to activate the smart call mode and for requesting the other end to adjust the uplink audio processing strategy. In this way, the user on the second electronic device can request the other end to activate the smart call mode or adjust the uplink audio processing strategy based on their actual call experience, thereby improving call quality and providing a better call experience.
[0028] In conjunction with the first aspect, in one possible implementation, the method further includes: displaying a first user interface, the first user interface including an audio processing configuration control; displaying a second user interface in response to a user operation applied to the audio processing configuration control; the second user interface including an uplink audio processing configuration area, the uplink audio processing configuration area including one or more of a noise reduction intensity configuration area, an audio gain configuration area, and a noise reduction scene configuration area, the noise reduction intensity configuration area including multiple noise reduction intensity options, the audio gain configuration area including multiple audio gain options, and the noise reduction scene configuration area including multiple noise reduction scene options; in response to a user operation applied to a first option, changing the first option to a selected state and enabling the audio processing strategy corresponding to the first option; the first option is any option in the second user interface.
[0029] In this embodiment of the application, the first electronic device may also provide a configuration interface corresponding to the uplink audio processing strategy, which can facilitate manual adjustment by the user and further improve the flexibility of the uplink audio processing strategy adjustment.
[0030] In conjunction with the first aspect, in one possible implementation, the method further includes: receiving fourth audio data from the second electronic device; processing the fourth audio data based on a current downlink audio processing strategy to obtain fifth audio data; analyzing the fifth audio data to determine audio quality problems of the fifth audio data; and adjusting the downlink audio processing strategy based on the audio quality problems of the fifth audio data.
[0031] In this embodiment of the application, the first electronic device can also analyze the audio data processed by the local downlink audio processing strategy, and then adjust the downlink audio processing strategy based on the analysis results, which can ensure the accuracy of the downlink audio processing strategy and enable the local user to have a better call experience.
[0032] The second aspect discloses an audio processing method, which can be applied to a first electronic device, a module (e.g., a processor) within the first electronic device, or a logic module or software capable of implementing all or part of the functions of the first electronic device. The following description uses an application to a first electronic device as an example. The communication method may include: performing semantic analysis on call audio data during a conversation between the first electronic device and a second electronic device to obtain a target semantic analysis result; the call audio data includes call audio data collected by the first electronic device; adjusting an audio processing strategy based on the target semantic analysis result; the audio processing strategy includes a downlink audio processing strategy used to process the call audio data from the second electronic device.
[0033] In conjunction with the second aspect, in one possible implementation, the downlink audio processing strategy includes one or more of the following: downlink noise reduction intensity, downlink audio gain, and downlink noise reduction scenario.
[0034] In conjunction with the second aspect, in one possible implementation, the audio processing strategy also includes an uplink audio processing strategy.
[0035] In conjunction with the second aspect, in one possible implementation, the uplink audio processing strategy includes one or more of the following: uplink noise reduction intensity, uplink audio gain, and uplink noise reduction scenario.
[0036] In conjunction with the second aspect, in one possible implementation, the call audio data also includes call audio data from the second electronic device.
[0037] It should be noted that the technical solution of the second aspect of this application may correspond to or be similar to the technical solution of the first aspect, and the relevant beneficial effects can be referred to the beneficial effects of the first aspect.
[0038] The third aspect discloses an audio processing method, which can be applied to a first electronic device, a module (e.g., a processor) within the first electronic device, or a logic module or software capable of implementing all or part of the functions of the first electronic device. The following description uses an application to a first electronic device as an example. The communication method may include: performing semantic analysis on call audio data during a conversation between the first electronic device and a second electronic device to obtain a target semantic analysis result; the call audio data includes call audio data from the second electronic device; adjusting an audio processing strategy based on the target semantic analysis result; wherein the audio processing strategy includes a downlink audio processing strategy used to process the call audio data from the second electronic device.
[0039] In conjunction with the third aspect, in one possible implementation, the downlink audio processing strategy includes one or more of downlink noise reduction intensity, downlink audio gain, and downlink noise reduction scenario.
[0040] In conjunction with the third aspect, in one possible implementation, the audio processing strategy also includes an uplink audio processing strategy.
[0041] In conjunction with the third aspect, in one possible implementation, the uplink audio processing strategy includes one or more of uplink noise reduction intensity, uplink audio gain, and uplink noise reduction scenario.
[0042] In conjunction with the third aspect, in one possible implementation, the call audio data also includes call audio data collected by the first electronic device.
[0043] It should be noted that the technical solution of the third aspect of this application may correspond to or be similar to the technical solution of the first aspect, and the relevant beneficial effects can be referred to the beneficial effects of the first aspect.
[0044] The fourth aspect discloses an audio processing method, which can be applied to a first electronic device, a module (e.g., a processor) within the first electronic device, or a logic module or software capable of implementing all or part of the functions of the first electronic device. The following description uses an application to a first electronic device as an example. The communication method may include: performing semantic analysis on call audio data during a conversation between the first electronic device and a second electronic device to obtain a target semantic analysis result; the call audio data includes call audio data collected by the first electronic device; adjusting an audio processing strategy based on the target semantic analysis result; the audio processing strategy includes an uplink audio processing strategy used to process the call audio data collected by the first electronic device.
[0045] In conjunction with the fourth aspect, in one possible implementation, the uplink audio processing strategy includes one or more of uplink noise reduction intensity, uplink audio gain, and uplink noise reduction scenario.
[0046] In conjunction with the fourth aspect, in one possible implementation, the audio processing strategy further includes a downlink audio processing strategy.
[0047] In conjunction with the fourth aspect, in one possible implementation, the downlink audio processing strategy includes one or more of downlink noise reduction intensity, downlink audio gain, and downlink noise reduction scenario.
[0048] In conjunction with the fourth aspect, in one possible implementation, the call audio data also includes call audio data from the second electronic device.
[0049] It should be noted that the technical solution of the fourth aspect of this application may correspond to or be similar to the technical solution of the first aspect, and the relevant beneficial effects can be referred to the beneficial effects of the first aspect.
[0050] The fifth aspect discloses an electronic device, which may be a first electronic device, comprising a processor and a communication interface; the communication interface is used to receive and transmit data; the processor invokes computer programs or computer instructions stored in a memory to implement the methods provided in the first aspect and any possible embodiments thereof, or to implement the methods provided in the second aspect and any possible embodiments thereof, or to implement the methods provided in the third aspect and any possible embodiments thereof, or to implement the methods provided in the fourth aspect and any possible embodiments thereof.
[0051] As one possible implementation, the communication device disclosed in the fifth aspect above may include one or more processors.
[0052] Optionally, the communication device disclosed in the fifth aspect above further includes one or more memories.
[0053] The sixth aspect discloses a computer-readable storage medium storing a computer program or computer instructions that, when executed, implement the methods provided in the first aspect and any possible embodiments thereof, or implement the methods provided in the second aspect and any possible embodiments thereof, or implement the methods provided in the third aspect and any possible embodiments thereof, or implement the methods provided in the fourth aspect and any possible embodiments thereof.
[0054] The seventh aspect discloses a chip including a processor for executing a program stored in a memory, wherein when the program is executed, the chip performs the methods provided in the first aspect and any possible embodiments thereof, or performs the methods provided in the second aspect and any possible embodiments thereof, or performs the methods provided in the third aspect and any possible embodiments thereof, or performs the methods provided in the fourth aspect and any possible embodiments thereof.
[0055] As one possible implementation, the memory is located outside the chip.
[0056] The eighth aspect discloses a computer program product comprising computer program code that, when executed, causes the methods provided in the first aspect and any possible embodiments thereof to be performed, or causes the methods provided in the second aspect and any possible embodiments thereof to be performed, or causes the methods provided in the third aspect and any possible embodiments thereof to be performed, or causes the methods provided in the fourth aspect and any possible embodiments thereof to be performed.
[0057] It should be understood that the implementation and beneficial effects of the above-mentioned aspects or any possible implementation methods of this application can be referred to each other. Attached Figure Description
[0058] The accompanying drawings are provided to more clearly illustrate the technical solutions of the embodiments of this application. The drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0059] Figure 1 is a schematic diagram of the structure of a communication system provided in an embodiment of this application;
[0060] Figure 2 is a schematic diagram of the hardware structure of an electronic device 100 provided in an embodiment of this application;
[0061] Figure 3 is a schematic diagram of the software structure of an electronic device 100 provided in an embodiment of this application;
[0062] Figures 4A to 4H are schematic diagrams illustrating scenarios in which some electronic devices 100 provided in the embodiments of this application provide intelligent call modes and related audio processing configurations;
[0063] Figures 5A to 5D are schematic diagrams of scenarios where electronic devices 100 intelligently adjust audio processing strategies according to embodiments of this application;
[0064] Figures 6A to 6C are schematic diagrams of some electronic devices 100 provided in the embodiments of this application, which provide noise reduction scene selection and intelligent adjustment of audio processing strategies based on positioning results;
[0065] Figures 7A to 7F are schematic diagrams illustrating scenarios in which some electronic devices 100 provided in the embodiments of this application provide intelligent call mode, as well as uplink and uplink-related audio processing configurations;
[0066] Figures 8A to 8E are schematic diagrams of scenarios provided in the embodiments of this application, in which an electronic device 200 requests an electronic device 100 to enable the smart call mode and adjust the uplink audio processing strategy;
[0067] Figure 9 is a flowchart illustrating an audio processing method provided in an embodiment of this application;
[0068] Figure 10 is a flowchart illustrating another audio processing method provided in an embodiment of this application. Detailed Implementation
[0069] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.
[0070] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0071] In the following embodiments of this application, the term "user interface (UI)" refers to the medium interface through which an application (APP) or operating system (OS) interacts and exchanges information with the user. It realizes the conversion between the internal form of information and the form that the user can accept. The user interface is source code written in a specific computer language such as Java or Extensible Markup Language (XML). The interface source code is parsed and rendered on the electronic device, ultimately presenting content that the user can recognize. A common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be visible interface elements such as text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets displayed on the screen of an electronic device.
[0072] Intelligent noise reduction and intelligent compensation audio processing technologies have been widely applied in various scenarios in daily life, such as calls, recordings, and live streaming. For example, during calls or live streaming, the local electronic device may be in a noisy environment. In this case, the acquired external audio may include not only the target voice but also ambient noise (such as traffic noise, keyboard typing, and other background noise). Therefore, the local electronic device can perform noise reduction processing on the acquired external audio in real time to reduce / eliminate ambient noise, and then send the noise-reduced audio to the other electronic device so that the user can clearly hear the target voice. Furthermore, after receiving the audio, the other electronic device can also perform noise reduction and compensation processing (such as signal frequency band loss compensation and packet loss compensation) to obtain better call quality.
[0073] This application provides an audio processing method. Among them, the first electronic device can perform semantic analysis on the received audio data from the second electronic device in audio scenarios such as voice calls and video calls, and dynamically adjust the audio processing strategy based on the semantic analysis result. For example, if the semantic analysis result of a certain segment of audio data from the second electronic device by the first electronic device (such as "It's so noisy on your side", "The noise on your side is so loud", etc.) is "The current noise is relatively large", then the first electronic device can adjust the current noise reduction intensity, such as changing the noise reduction intensity from "weak" to "medium". After that, if the semantic analysis result of a certain segment of audio data from the second electronic device by the first electronic device (such as "It's still so noisy on your side", "The noise on your side is still so loud", etc.) is still "The current noise is relatively large", then the first electronic device can continue to adjust the current noise reduction intensity, such as changing the noise reduction intensity from "medium" to "strong". Another example, if the semantic analysis result of a certain segment of audio data from the second electronic device by the first electronic device (such as "Your voice is too low", "I can't hear what you're saying", etc.) is "The current voice is relatively low", then the first electronic device can increase the current audio gain, such as changing the audio gain from x1 decibels (dB) to x2 dB, where x1 < x2. Another example, if the semantic analysis result of a certain segment of audio data from the second electronic device by the first electronic device (such as "Your voice is too loud", "You can speak more softly", etc.) is "The current voice is relatively large", then the first electronic device can decrease the current audio gain, such as changing the audio gain from x2 dB to x1 dB. It can be seen that in the above method, for audio scenarios such as voice calls and video calls, the first electronic device can dynamically adjust the audio processing strategy (such as noise reduction intensity, audio gain, etc.) based on the semantic analysis result of the audio data, which can improve the audio quality in scenarios such as calls and live broadcasts, thereby improving the user experience.
[0074] In another implementation, during audio scenarios such as voice or video calls between the first and second electronic devices, the second electronic device can perform real-time analysis of the audio data from the first electronic device and then return the analysis results to the first electronic device. The first electronic device can dynamically adjust its audio processing strategy based on these results. For example, the second electronic device can analyze the audio data from the first electronic device using a voice quality detection algorithm to obtain a real-time voice quality score. If the voice quality score is greater than a quality score threshold (e.g., 6, where 1-10 is the threshold), the first electronic device may not perform any processing. If the voice quality score is less than or equal to the threshold, the first electronic device can determine the reason for the voice quality score being lower than the threshold based on the corresponding audio data (e.g., low audio amplitude, excessive noise in the audio data), and then return the reason to the first electronic device. The first electronic device can then dynamically adjust its audio processing strategy based on this reason. For instance, if the reason for the voice quality score being lower than the threshold is excessive noise in the audio data, indicating that the noise level on the first electronic device's side is high, the first electronic device can adjust the current noise reduction intensity, such as changing it from "weak" to "medium".
[0075] In another implementation, since noise sources may differ in different scenarios (such as subway stations, offices, construction sites, etc.), multiple audio processing scenarios can be predefined, and different noise reduction models can be configured for different audio processing scenarios. In audio scenarios such as voice calls and video calls between the first and second electronic devices, the first electronic device can determine the current audio processing scenario based on the positioning results, and then automatically help the user select the noise reduction model corresponding to that audio processing scenario for processing, ensuring better audio quality. Furthermore, the first electronic device can also provide a corresponding configuration interface, allowing the user to manually select the corresponding audio processing scenario. For example, the first electronic device can respond to user input and display the corresponding audio processing configuration interface, which includes multiple audio processing scenario options. The user can select the corresponding audio processing scenario option by touching or clicking. Then, the first electronic device can use the noise reduction model corresponding to the user-selected audio processing scenario option for processing, ensuring better audio quality.
[0076] Please refer to Figure 1, which exemplarily shows a schematic diagram of the structure of a communication system provided in an embodiment of this application.
[0077] The communication system may include electronic device 100 and electronic device 200. Electronic device 100 and electronic device 200 can conduct audio calls, such as voice calls and video calls. Electronic device 100 can be a first electronic device, and electronic device 200 can be a second electronic device.
[0078] Electronic device 100 and electronic device 200 may be equipped with Alternatively, it can be a portable electronic device with other operating systems, such as a mobile phone, tablet computer, wearable device (such as a smartwatch, smart bracelet, etc.), etc., or it can be an augmented reality (AR) device, a virtual reality (VR) device, a laptop computer with a touch-sensitive surface or touch panel, a desktop computer with a touch-sensitive surface or touch panel, or other non-portable electronic devices. Electronic device 100 and electronic device 200 can also be chips or processing systems from the aforementioned devices. This application embodiment does not limit the types of electronic device 100 and electronic device 200.
[0079] In audio scenarios such as voice and video calls between electronic devices 100 and 200, electronic device 100 can perform semantic analysis on the audio data from electronic device 200 and dynamically adjust its audio processing strategy based on the semantic analysis results. Electronic device 100 can also dynamically adjust its audio processing strategy based on the analysis results returned by electronic device 200. Electronic device 100 can also determine the current audio processing scenario based on the location results and then automatically adopt the noise reduction model corresponding to that audio processing scenario. Similarly, electronic device 200 can perform semantic analysis on the audio data from electronic device 100 and dynamically adjust its audio processing strategy based on the semantic analysis results. Electronic device 200 can also dynamically adjust its audio processing strategy based on the analysis results returned by electronic device 100. Electronic device 200 can also determine the current audio processing scenario based on the location results and then automatically adopt the noise reduction model corresponding to that audio processing scenario.
[0080] It should be noted that the system architecture and business scenarios (or application scenarios) described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application.
[0081] The structure of the electronic device 100 involved in this application is described below.
[0082] Figure 2 illustrates a schematic diagram of the hardware structure of an electronic device 100 provided in an embodiment of this application.
[0083] As shown in Figure 2, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.
[0084] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0085] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0086] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.
[0087] The processor 110 may also include a memory for storing instructions and data. In some examples, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or is recurring. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0088] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback.
[0089] The charging management module 140 receives charging input from a charger, which can be a wireless charger or a wired charger. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.
[0090] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, external memory, display 194, camera 193, and wireless communication module 160, etc.
[0091] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0092] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.
[0093] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1.
[0094] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0095] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connecting the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering.
[0096] The display screen 194 is used to display images, videos, etc. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.
[0097] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0098] The ISP is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, converting it into an image visible to the naked eye.
[0099] Camera 193 is used to capture still images or videos. In some embodiments, electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0100] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.
[0101] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0102] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.
[0103] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0104] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0105] Audio module 170 is used to convert digital audio information into analog audio signal output, and also to convert analog audio input into digital audio signal. Audio module 170 can also be used for encoding and decoding audio signals. In some examples, audio module 170 may be located in processor 110, or some functional modules of audio module 170 may be located in processor 110. Speaker 170A, also called a "loudspeaker," is used to convert audio electrical signals into sound signals. Receiver 170B, also called a "handset," is used to convert audio electrical signals into sound signals. Microphone 170C, also called a "microphone" or "microphone," is used to convert sound signals into electrical signals. Headphone jack 170D is used to connect wired headphones.
[0106] The sensor module 180 may include pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, proximity sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, bone conduction sensors, etc.
[0107] Buttons 190 include a power button, volume buttons, etc. Motor 191 can generate vibration feedback. Indicator 192 can be an indicator light, used to indicate charging status, battery level changes, and also to indicate messages, missed calls, notifications, etc.
[0108] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and detach from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The electronic device 100 interacts with the network through the SIM card to achieve functions such as calls and data communication. In some examples, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be removed from it.
[0109] Electronic device 100 may be equipped with HarmonyOS ( Electronic devices using an OS (or other operating system), such as mobile phones, tablets, laptops, smartwatches, smart bracelets, etc. This application does not limit the specific type of electronic device 100.
[0110] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses a layered architecture. Taking the system as an example, the software structure of electronic device 100 is illustrated.
[0111] Figure 3 is a software structure block diagram of an electronic device 100 according to an embodiment of this application.
[0112] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, [the following is omitted as the text is incomplete and likely refers to a specific implementation or feature]. The system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0113] The application layer can include a series of application packages.
[0114] As shown in Figure 3, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, and SMS.
[0115] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0116] As shown in Figure 3, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, activity manager, etc.
[0117] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0118] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0119] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0120] The phone manager is used to provide communication functions for electronic device 100. For example, it manages call status (including connection and disconnection).
[0121] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0122] The notification manager allows applications to display notifications in the status bar (such as the pull-down notification bar). It can be used to convey informational messages and can disappear automatically after a short pause without user interaction. For example, the notification manager can be used to notify users of download completion or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.
[0123] The Activity Manager is responsible for managing activities, including starting, switching, and scheduling components in the system, as well as managing and scheduling applications. The Activity Manager can be called by upper-level applications to open the corresponding activities.
[0124] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.
[0125] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0126] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0127] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0128] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0129] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0130] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0131] A 2D graphics engine is a graphics engine for 2D drawing.
[0132] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0133] The hardware and software structures of the electronic device 200 can also be as shown in Figures 2 and 3. This application embodiment does not specifically limit the hardware and software structures of the electronic device 200.
[0134] The following describes the usage scenarios involved in the embodiments of this application in conjunction with the user interface of the electronic device 100.
[0135] Figures 4A to 4H exemplarily illustrate scenarios where electronic device 100 provides a smart call mode and related audio processing configurations.
[0136] As shown in Figure 4A, the electronic device 100 can display a user interface 410. The user interface 410 displays a page with application icons. This page may include multiple application icons (e.g., clock application icon, calendar application icon, gallery application icon, memo application icon, file management application icon, etc.). Below these multiple application icons, a page indicator may also be displayed to indicate the positional relationship between the currently displayed page and other pages. Below the page indicator are multiple tray icons (e.g., camera application icon, contacts application icon, phone application icon 411, messaging application icon). The tray icons remain displayed when switching pages. It should be understood that the content displayed on the user interface 410 is not limited in this embodiment. In response to a user operation, such as a touch operation, applied to the application icon or tray icon, the electronic device 100 can open the application corresponding to that application icon or tray icon.
[0137] For example, in response to a user action on the phone application icon 411, the electronic device 100 can open the phone application and display the user interface 420 as shown in FIG4B. The user interface 420 may include a dial pad, a number display area, and dial controls (such as dial control 421). The user can enter the number to be dialed through the dial pad, and then, in response to a user action on the dial control 421, such as a touch operation, the electronic device 100 can dial the corresponding number, and after the call is connected, the user interface 430 as shown in FIG4C can be displayed.
[0138] As shown in Figure 4C, in response to a user operation of swiping down on the status bar at the top of the user interface 430, the electronic device 100 can display the user interface 440 shown in Figure 4D. For example, the user interface 440 can be a notification information interface. It should be understood that the notification information interface can also be opened through other user interfaces, and this embodiment of the application is not limited thereto. The user interface 440 may include a call status card 441 and a smart call mode card 445. The call status card 441 can be used to display call-related information, such as call time and call number. The call status card 441 may also include a mute control 442, a hang-up control 443, and a hands-free control 444. In response to a user operation on the mute control 442, the electronic device 100 can turn the mute function on or off. When the mute function is on, the other end of the call can hear the local end's voice; when the mute function is off, the other end cannot hear the local end's voice. In response to a user operation on the hang-up control 443, the electronic device 100 can hang up the current call. In response to a user operation on the hands-free control 444, the electronic device 100 can turn the hands-free function on or off. When the hands-free function is on, the sound from the other end can be played through the speaker, and the volume is relatively loud. When the hands-free function is off, the sound from the other end can be played through the earpiece, and the volume is relatively low.
[0139] The smart call mode card 445 may include information related to the smart call mode, such as whether the smart call mode is enabled (e.g., currently disabled), prompts related to intelligent audio processing strategy adjustment, and controls for enabling or disabling the smart call mode. As shown in Figure 4D, in response to a user operation, such as a touch operation, on the expansion control 446, the electronic device 100 can display the user interface 440 as shown in Figure 4E. It is understood that, due to the limited screen size of the electronic device 100, in the notification information interface, some cards with a lot of content may only display a portion of their content. The electronic device 100 can respond to a user operation on the display control corresponding to a card to expand the complete content of that card. That is, the user can view the complete content of the card by touching the expansion control.
[0140] As shown in Figure 4E, the user interface 440 displays the complete content of the Smart Call Mode card 445. The Smart Call Mode card 445 displays an audio processing configuration control 447, an enable Smart Call Mode control 448, and a hide control 449. The audio processing configuration control 447 is used to open the audio processing configuration interface and configure relevant audio processing parameters, such as noise reduction intensity and audio gain. The enable Smart Call Mode control 448 is used to enable Smart Call Mode. The hide control 449 hides the complete content of the Smart Call Mode card 445, displaying only a portion of the content; this is the opposite of the expand control.
[0141] As shown in Figure 4E, in response to a user operation on the Smart Call Mode Activation Control 448, the electronic device 100 can activate Smart Call Mode and display the user interface 440 shown in Figure 4F. The user interface 440 in Figure 4F displays the message "Activated, optimizing call quality," indicating that Smart Call Mode is activated. Furthermore, the user interface 440 in Figure 4F may also display a Smart Call Mode Deactivation Control 450. The Smart Call Mode Deactivation Control 450 can be used to deactivate Smart Call Mode.
[0142] As shown in Figure 4F, in response to a user operation, such as a touch operation, on the audio processing configuration control 447, the electronic device 100 can display the user interface 460 shown in Figure 4G, i.e., the audio processing configuration interface. The user interface 460 can display a configuration area for noise reduction intensity and a configuration area for audio gain. The noise reduction intensity configuration area can be used to adjust the noise reduction intensity. For example, the noise reduction intensity can include three levels: weak, medium, and strong, with the currently selected noise reduction intensity being weak. Here, the noise reduction intensity can refer to the noise reduction intensity used by the electronic device 100 to process acquired external audio. The stronger the noise reduction intensity, the stronger the ability to reduce / eliminate ambient noise. In some embodiments, the default noise reduction intensity can be weak each time the electronic device makes a call if the user does not change the audio processing configuration. The audio gain can include multiple options, such as 0dB, 3dB, 5dB, 10dB, 12dB, 15dB, etc., with the currently selected audio gain being 5dB. The audio gain here can be the audio gain of the electronic device 100 for audio processing of the external audio it collects. The stronger the audio gain, the louder the sound can be heard at the other end.
[0143] The return control 461 can be used to trigger the electronic device 100 to return to the previous level user interface, that is, to display the user interface 440 shown in Figure 4F.
[0144] As shown in Figure 4G, in response to a user operation, such as a touch operation, on the selection control 462 corresponding to the noise reduction intensity, the electronic device 100 can display the user interface 460 shown in Figure 4H. In the user interface 460, the selection control 462 corresponding to the noise reduction intensity is selected, indicating that the electronic device 100 has switched the noise reduction intensity to medium. It should be understood that in this embodiment, each noise reduction intensity can correspond to a noise reduction model, and different noise reduction intensities correspond to different noise reduction models. When switching the noise reduction intensity, the electronic device 100 can correspondingly switch the noise reduction model used. For the same noise reduction model (such as a neural network model with the same structure), different noise reduction parameters can be understood as different noise reduction models in this embodiment; that is, any difference in either the noise reduction model or the model parameters can be considered a different noise reduction model. For example, the denoising model corresponding to weak denoising intensity can be denoising model 1 (model 1 + model parameter 1), the denoising model corresponding to medium denoising intensity can be denoising model 2 (model 1 + model parameter 2), and the denoising model corresponding to strong denoising intensity can be denoising model 3 (model 1 + model parameter 3). Noise denoising models 1, 2, and 3 can all use the same model 1, but their model parameters can be different. As another example, the denoising model corresponding to weak denoising intensity can be denoising model 1 (model 1 + model parameter 1), the denoising model corresponding to medium denoising intensity can be denoising model 2 (model 2 + model parameter 2), and the denoising model corresponding to strong denoising intensity can be denoising model 3 (model 3 + model parameter 3). Noise denoising models 1, 2, and 3 can use different models, and their model parameters can be different.
[0145] It should be understood that the embodiments of this application do not limit the classification of noise reduction intensity levels, and may also be divided into more or fewer noise reduction levels. Similarly, the embodiments of this application do not limit the audio gain options, and may also include more or fewer audio gain options.
[0146] As can be seen, based on the aforementioned audio processing configuration interface, users can select different noise reduction intensities and audio gains according to their actual needs to improve the call quality for the other end. For example, if the other end reports excessive noise during a call, and the currently selected noise reduction level is weak, a higher noise reduction intensity can be selected, such as medium or strong. As another example, if the other end reports insufficient volume during a call, and the currently selected audio gain is 5dB, a higher audio gain can be selected, such as 10dB or 15dB.
[0147] It is understandable that, in some cases, the aforementioned noise reduction strength can also refer to the noise reduction strength of the electronic device 100 in processing the audio received from other electronic devices (such as electronic device 200). The stronger the noise reduction strength, the stronger the ability to reduce / eliminate noise in the received audio. Similarly, the aforementioned audio gain can also refer to the audio gain of the electronic device 100 in processing the audio received from other electronic devices (such as electronic device 200). The stronger the audio gain, the louder the sound that can be heard at this end.
[0148] The above briefly describes the scenario where users manually adjust audio processing strategies. However, in this embodiment, in audio scenarios such as voice calls and video calls, the electronic device 100 can also automatically adjust the audio processing strategy, which can improve the efficiency of audio processing strategy adjustment and enhance the user experience. For example, the electronic device 100 can perform semantic analysis on the received audio from the other end, and then intelligently adjust the audio processing strategy based on the results of the semantic analysis. As another example, the electronic device 100 can receive audio quality issues reported by the other end electronic device (such as electronic device 200), and then automatically adjust the audio processing strategy based on the feedback from the other end.
[0149] Figures 5A to 5D exemplarily illustrate scenarios where electronic device 100 provides an intelligent call mode and intelligently adjusts audio processing strategies.
[0150] After activating the intelligent call mode, the electronic device 100 can perform real-time semantic analysis on the received audio from the other end of the call. If the semantic analysis result for a certain segment of audio data (such as "It's so noisy over there," "The noise is so loud over there") is "The current noise is relatively loud," then the electronic device 100 can adjust the current noise reduction intensity, such as changing the noise reduction intensity from "weak" to "medium." If the semantic analysis result for a certain segment of audio data (such as "Your voice is so soft," "I can't hear what you're saying," etc.) is "The current voice is relatively soft," then the electronic device 100 can increase the current audio gain, such as increasing the audio gain from 5dB to 10dB. In this embodiment, semantic analysis can also be referred to as semantic understanding.
[0151] As shown in Figure 5A, assuming that at the 10th second (s) of the call, the electronic device 100 determines that the local noise level is relatively high based on the semantic analysis results, the electronic device 100 can automatically adjust the noise reduction intensity from "weak" to "medium". Furthermore, the smart call mode card 445 can include relevant prompts, such as displaying relevant prompts in the display area 451 of the smart call mode card 445, to inform the user that the noise reduction intensity has been automatically adjusted.
[0152] As shown in Figure 5B, assuming that at the 20th second of the call, the semantic analysis result of the audio received by the electronic device 100 from the other end of the call (such as "Your voice is too soft", "I can't hear what you are saying") is "The current voice is too soft", then the electronic device 100 can automatically adjust the audio gain from 3dB to 10dB, and the display area 452 in the smart call mode card 445 can display relevant prompt information to inform the user that the audio gain has been automatically adjusted.
[0153] The following describes a method for automatically adjusting audio processing strategies based on feedback from peer electronic devices.
[0154] After enabling intelligent call mode, when electronic device 100 receives audio quality feedback from the other end of the call (e.g., electronic device 200), it can adjust its current audio processing strategy based on this feedback. For example, if the other end reports "the current noise is too loud," electronic device 100 can adjust the current noise reduction intensity, such as changing it from "weak" to "medium." If the other end reports "the current volume is too low," electronic device 100 can increase the current audio gain, such as increasing it from 5dB to 10dB.
[0155] As shown in Figure 5C, assuming that at the 10th second of the call, the electronic device 100 receives feedback from the other electronic device regarding audio quality issues ("the current noise is relatively loud"), the electronic device 100 can automatically adjust the noise reduction intensity from "weak" to "medium", and the display area 453 in the smart call mode card 445 can display relevant prompt information to inform the user that the noise reduction intensity has been automatically adjusted.
[0156] As shown in Figure 5D, assuming that at the 20th second of the call, the electronic device 100 receives feedback from the other electronic device regarding an audio quality problem ("the current volume is relatively low"), then the electronic device 100 can automatically adjust the audio gain from 3dB to 10dB, and the display area 454 in the smart call mode card 445 can display relevant prompt information to inform the user that the audio gain has been automatically adjusted.
[0157] Figures 6A to 6C exemplarily illustrate scenarios where electronic device 100 provides noise reduction scene selection and intelligently adjusts audio processing strategies based on positioning results.
[0158] It is understandable that users frequently use the call function in different scenarios in daily life, and the noise sources in different scenarios may be different. Therefore, if the electronic device 100 uses the same noise reduction model for all scenarios, the noise reduction effect may be poor in some scenarios. In this embodiment of the application, in order to provide users with a better call experience, a noise reduction model can be configured for a variety of different scenarios (such as office, shopping mall, train station, subway station, etc.) to achieve a better noise reduction effect.
[0159] As shown in Figure 6A, the user interface 460 can also display a configuration area for noise reduction scenes. This configuration area can be used to adjust the noise reduction scene. For example, noise reduction scenes can include general, office, shopping mall, subway station, train station, construction site, etc. Typically, a "general" noise reduction scene can be selected. The noise reduction model corresponding to a "general" noise reduction scene can meet the needs of most scenarios, but its noise reduction effect may be weaker than a dedicated noise reduction model for a specific scenario. It is understood that the noise reduction scenes described here are merely illustrative; in other embodiments of this application, more or fewer noise reduction scenes may be included.
[0160] Assuming the user is currently in a shopping mall, to achieve better noise reduction, the user can switch the noise reduction scene to the shopping mall setting. As shown in Figure 6A, in response to a user operation, such as a touch operation, on the selection control 463 corresponding to the shopping mall setting, the electronic device 100 can display the user interface 460 shown in Figure 6B. In the user interface 460, the selection control 463 corresponding to the shopping mall setting is selected, indicating that the electronic device 100 has switched the noise reduction scene to the shopping mall setting, and the noise reduction model can also be switched from the noise reduction model corresponding to the general scene to the noise reduction model corresponding to the shopping mall scene.
[0161] To improve user experience, the electronic device 100 can also adjust the noise reduction scene based on real-time location results during a call, that is, switch the noise reduction model for the corresponding scene based on real-time location results.
[0162] As shown in Figure 6C, assuming that at the 20th second of the call, the electronic device 100 determines that it is currently in a shopping mall based on the location result, it can automatically adjust the noise reduction scene to the shopping mall and switch to the noise reduction model corresponding to the shopping mall scene. Furthermore, the display area 455 of the intelligent call mode card 445 can display relevant prompts to inform the user that the noise reduction scene has been automatically adjusted.
[0163] Figures 7A to 7F exemplarily illustrate a scenario where electronic device 100 provides a smart call mode, as well as uplink and uplink-related audio processing configurations.
[0164] Understandably, audio scenarios such as voice and video calls are generally bidirectional, meaning the device can receive audio from the other end's electronic device and also send external audio collected by itself to the other end. The sending of external audio collected by the local electronic device to the other end's electronic device is called uplink, and the receiving of audio from the other end's electronic device is called downlink. Therefore, the audio processing function (such as noise reduction) of electronic device 100 can be divided into uplink audio processing and downlink audio processing. Uplink audio processing can process external audio collected by electronic device 100, while downlink audio processing can process audio received from the other end's electronic device (such as electronic device 200). Uplink processing mainly aims to reduce / eliminate noise in the external audio collected by the local device and adjust the audio amplitude so that the other end user can hear the target voice more clearly. Downlink processing mainly aims to reduce / eliminate noise in the received audio from the other end's electronic device and adjust the audio amplitude so that the local user can hear the target voice more clearly.
[0165] Typically, the two parties in a call are in different environments (e.g., user A is in a quiet indoor space, while user B is in a crowded shopping mall). Therefore, electronic device 100 can provide functions for uplink audio processing and downlink audio processing respectively.
[0166] As shown in Figure 7A, in response to a user operation on the Smart Call Mode Activation Control 448, the electronic device 100 can activate the Smart Call Mode and display the user interface 440 shown in Figure 7B. In the user interface 440 shown in Figure 7B, the display area 471 shows the message "Activated (Uplink + Downlink), optimizing call quality," indicating that the Smart Call Mode has been activated, and both uplink and downlink audio processing functions are enabled. When the uplink audio processing function is enabled, the mute control 442 can be illuminated. When the downlink audio processing function is enabled, the hands-free control 444 can be illuminated.
[0167] As shown in Figure 7B, in response to a user operation on the hands-free control 444, such as a double-click, the electronic device 100 can display the user interface 440 shown in Figure 7C. In the user interface 440 shown in Figure 7C, the hands-free control 444 is in a de-illuminated state, indicating that the electronic device 100 has disabled the downlink audio processing function. Furthermore, the prompt message only indicates that the uplink audio processing function is enabled.
[0168] As shown in Figure 7C, in response to a user operation on the mute control 442, such as a double-click, the electronic device 100 can display the user interface 440 shown in Figure 7D. In the user interface 440 shown in Figure 7D, the mute control 442 is in a de-illuminated state, indicating that the electronic device 100 has turned off the uplink audio processing function. At this time, since both the uplink and downlink audio processing functions are turned off, the smart call mode is in a disabled state.
[0169] Understandably, by performing a user operation on the mute control 442 again, such as double-clicking, the electronic device 100 can re-enable the uplink audio processing function. Similarly, by performing a user operation on the hands-free control 444 again, such as double-clicking, the electronic device 100 can re-enable the downlink audio processing function.
[0170] As shown in Figure 7D, in response to a user operation, such as a touch operation, on the audio processing configuration control 447, the electronic device 100 can display a user interface 460 as shown in Figure 7E. The user interface 460 may display an uplink audio processing configuration area 472 and a downlink audio processing configuration area 473. Both the uplink audio processing configuration area 472 and the downlink audio processing configuration area 473 may include a configuration area for noise reduction intensity, a configuration area for audio gain, and a configuration area for noise reduction scenarios.
[0171] Users can configure the audio processing strategies for uplink and downlink audio through the user interface 460 shown in Figure 7E. For example, assuming the shopping mall environment where the electronic device 100 is located is noisy, the user can set the uplink noise reduction scenario to the shopping mall and the noise reduction intensity to medium, as shown in Figure 7F.
[0172] Understandably, besides users manually setting uplink and downlink audio processing strategies, during a call, electronic device 100 can also automatically adjust uplink and downlink noise reduction intensity, audio gain, and noise reduction scenarios based on semantic analysis or feedback from the other electronic device. Electronic device 100 can also adjust the uplink noise reduction scenario based on location results, which will not be elaborated upon here.
[0173] Figures 8A to 8E illustrate exemplary scenarios where electronic device 200 requests electronic device 100 to enable smart call mode and adjust uplink audio processing strategies.
[0174] Please refer to Figure 8A. The user interface 470 shown in Figure 8A can display a call status card and a smart call mode card. The call status card can display call-related information, such as call time and call number. The smart call mode card can display an audio processing configuration control (local end), a smart call mode activation control, a remote audio processing configuration control 471, and a hide control. The audio processing configuration control (local end) can be used to open the local audio processing configuration interface, and the remote audio processing configuration control 471 can be used to open the remote audio processing configuration interface. The smart call mode activation control can be used to activate smart call mode. The user interface 470 can be the notification information interface of the electronic device 200.
[0175] In response to a user operation, such as a touch operation, on the audio processing configuration control 471 at the other end, the electronic device 200 can display a user interface 480 as shown in FIG8B. The user interface 480 may display two areas: area 481 is for requesting the other end to enable Smart Call Mode, and area 482 is for requesting the other end to adjust noise reduction intensity, audio gain, and noise reduction scene configuration. Area 481 may include a request control 4811, which can be used to trigger the electronic device 200 to send a request message to the other end of the call requesting the activation of Smart Call Mode.
[0176] In response to a user operation (such as a touch operation) acting on the request control 4811, electronic device 200 can send a request message (first request message) to electronic device 100, which requests the activation of smart call mode. Correspondingly, electronic device 100 can receive the request message from electronic device 200, and then electronic device 100 can pop up a prompt message box based on the request message. The prompt message box can display a prompt message indicating that the other end user has requested the activation of smart call mode. For example, the prompt message could be "The other end requests the activation of smart call mode". For instance, assuming electronic device 100 is currently displaying a call interface, after receiving the request message from electronic device 200, the call interface can pop up a prompt message box, as shown in Figure 8C. Electronic device 100 can display a user interface 430, which can display a prompt message box 431. The prompt message box 431 displays the prompt message "The other end requests the activation of smart call mode", as well as a corresponding confirm control 4311 and reject control 4312. The confirm control 4311 can be used to confirm / accept the request sent by the other end, and the reject control 4312 can be used to reject the request sent by the other end.
[0177] For example, in response to a user operation (such as a touch operation) acting on the determining control 4311, the electronic device 100 can activate the smart call mode. Opening the notification information interface of the electronic device 100 can display the user interface shown in Figure 4F. In response to a user operation (such as a touch operation) acting on the reject control 4312, the electronic device 100 may either take no action, or the electronic device 100 may send an instruction message to the electronic device 200 indicating that the request from the electronic device 200 is rejected, that is, indicating that the smart call mode is rejected.
[0178] As shown in Figure 8D, in area 482, the user can select noise reduction intensity, audio gain, and noise reduction scene, etc. Furthermore, the user can send a request from electronic device 200 to electronic device 100 to request adjustment of the uplink audio processing strategy. For example, the user can select the noise reduction scene as "shopping mall" by touching or clicking the selection control 463 corresponding to the shopping mall. Then, the user can click or touch control 4822. In response to this operation, electronic device 200 can send a corresponding request message to electronic device 100. This request message requests adjustment of the uplink audio processing strategy and can carry the audio processing parameters to be adjusted and their corresponding values, such as "{noise reduction scene: shopping mall}". Correspondingly, electronic device 100 can receive the request message from electronic device 200. Then, electronic device 100 can pop up a prompt message box based on the request message, displaying prompt information that can be used to inform the user of the requested adjustment of the uplink audio processing strategy. For example, assuming that electronic device 100 is currently displaying a call interface, after receiving a request message from electronic device 200, a prompt message box can pop up on the call interface, as shown in Figure 8E. Electronic device 100 can display a user interface 430, which can display a prompt message box 432. The prompt message box 432 displays the prompt message "The other party requests to switch the noise reduction scene to shopping mall," as well as corresponding confirm control 4321 and reject control 4322. Among them, the confirm control 4311 can be used to confirm / accept the request sent by the other party, and the reject control 4312 can be used to reject the request sent by the other party.
[0179] For example, in response to a user action (such as a touch operation) acting on the determining control 4321, the electronic device 100 can switch the uplink noise reduction scenario to a shopping mall. In response to a user action (such as a touch operation) acting on the reject control 4322, the electronic device 100 may not take any action, or the electronic device 100 may send an indication message to the electronic device 200 to indicate that the request of the electronic device 200 is rejected, that is, to indicate that the switching of the uplink noise reduction scenario to the shopping mall is rejected.
[0180] It is understood that the user interfaces in Figures 8A-8E are merely illustrative. For example, in some possible implementations, the electronic device 200 may also provide controls for requesting the peer to disable the smart call mode. Furthermore, in some possible implementations, the electronic device 200 may also provide configuration of the peer's corresponding downlink audio processing strategy; however, this application embodiment does not limit this.
[0181] In some possible implementations, during audio scenarios such as voice calls and video calls, the electronic device 100 can display in real time on the user interface the original audio and the processed audio (such as noise-reduced audio, audio after gain adjustment, etc.) in a visual graph (such as time domain graph, spectrum graph, time spectrum graph, etc.), as well as other visual graphs that can show the audio processing effect.
[0182] In this embodiment, the audio processed by the electronic device can be audio collected from the outside world, such as audio collected from the outside world through a microphone during a call. The audio processed by the electronic device can also be audio received from other electronic devices, such as audio sent by the other electronic device during a call.
[0183] As can be seen, through the above method, users can view the visualization of the original audio and the processed audio in real time. By comparing the visualization of the original audio and the processed audio, users can intuitively understand the processing effect and processing capability of audio processing such as noise reduction of electronic devices.
[0184] For example, during a call, an electronic device can capture audio in real time through a microphone, perform real-time noise reduction on the captured audio, and display a time-domain graph of the original captured audio and the noise-reduced audio in real time on the user interface. Users can see how much ambient noise has been filtered through the time-domain graph of the original captured audio and the noise-reduced audio, and can intuitively feel / understand the real-time noise reduction effect of the electronic device.
[0185] The processing flow of the technical solution provided in the embodiments of this application will be described by way of example below. Please refer to Figure 9, which is a schematic flowchart of an audio processing method disclosed in an embodiment of this application. As shown in Figure 9, the method may include, but is not limited to, the following steps:
[0186] 901. During a conversation between electronic device 100 and electronic device 200, electronic device 200 sends first audio data to electronic device 100.
[0187] For example, suppose user A holds electronic device 100 and user B holds electronic device 200. When user A needs to contact user B, user A and user B can communicate (such as voice calls, video calls, etc.) through electronic devices 100 and 200. It is understood that the call here can be a cellular call, or a call function provided by various communication applications, social applications, etc., and this application embodiment does not limit it.
[0188] During a conversation between electronic device 100 and electronic device 200, electronic device 100 can use its microphone to collect external sounds in real time, such as ambient sounds and user A's voice, and can send the corresponding audio data to electronic device 200. Correspondingly, electronic device 200 can receive audio data from electronic device 100. Similarly, electronic device 200 can also use its microphone to collect external sounds in real time, such as ambient sounds and user B's voice, and can send the corresponding audio data (such as the first audio data) to electronic device 100. Correspondingly, electronic device 100 can receive audio data from electronic device 200.
[0189] 902. Electronic device 100 performs semantic analysis based on the first audio data to obtain the first semantic analysis result.
[0190] Electronic device 100 can receive audio data from electronic device 200 in real time and perform semantic analysis in real time, so as to dynamically adjust the current uplink audio processing strategy based on the semantic analysis results.
[0191] For example, electronic device 100 can receive first audio data from electronic device 200. Then, electronic device 100 can perform semantic analysis on the first audio data to obtain a first semantic analysis result. For instance, the first audio data could be phrases like "It's so noisy over there," "The noise over there is so loud," "It's a bit noisy over there, I can't hear you clearly," "It's still so noisy over there," or "The noise over there is still so loud." The corresponding first semantic analysis result could be "The external noise is quite loud at this point." As another example, the first audio data could be phrases like "Your voice is so soft," "I can't hear what you're saying," "Could you speak louder?" or "Your voice is still a bit soft." The corresponding first semantic analysis result could be "The sound is relatively soft at this point." Yet another example, the first audio data could be phrases like "Your voice is so loud," "Could you speak softer?" or "Your voice is too loud." The corresponding first semantic analysis result could be "The sound is relatively loud at this point."
[0192] In some possible implementations, the electronic device 100 may perform semantic analysis on the first audio data only when the smart call mode is enabled, so as to dynamically adjust the current uplink audio processing strategy based on the semantic analysis results. That is, when the smart call mode is disabled, the electronic device 100 may not perform any processing. In other possible implementations, even when the smart call mode is disabled, the electronic device 100 may still perform semantic analysis on the first audio data, so as to determine whether to enable the smart call mode based on the semantic analysis results. Furthermore, after enabling the smart call mode, the electronic device 100 may continue to perform semantic analysis on subsequent audio data, so as to dynamically adjust the current uplink audio processing strategy based on the semantic analysis results. For example, during the initial call, the smart call mode is disabled by default. Subsequently, the electronic device 100 may automatically enable the smart call mode when the semantic analysis results are "the external noise at the current end is relatively loud," "the current volume at the current end is relatively low," or "the current volume at the current end is relatively loud," so as to perform audio processing such as uplink noise reduction and adjustment of uplink audio amplitude. In this way, during a call, the intelligent call mode can be automatically activated based on the semantic analysis results, which can meet the usage needs in a timely manner, and reduce the overall power consumption of electronic devices while ensuring call quality.
[0193] Understandably, when the smart call mode is enabled, the electronic device 100 can perform corresponding processing based on the audio processing configuration, such as noise reduction processing based on the configured noise reduction intensity, and adjusting the audio amplitude based on the configured audio gain.
[0194] It should be noted that the semantic analysis model used by the electronic device 100 in this application embodiment is not specifically limited, and can be a deep neural network model or other semantic analysis models. Furthermore, it should be understood that, for audio data, the semantic analysis model can first convert the audio data into text, that is, extract the content spoken by the target human voice from the audio data, and then perform semantic parsing based on the text (i.e., analyze and understand the text). In other words, the semantic analysis model can include speech-to-text functionality and semantic parsing functionality.
[0195] 903. Electronic device 100 adjusts the uplink audio processing strategy based on the first semantic analysis result.
[0196] After obtaining the first semantic analysis result, the electronic device 100 can adjust the currently used uplink audio processing strategy based on the first semantic analysis result.
[0197] It is understandable that conversations between user A and user B typically include content that does not relate to audio quality or effects (such as noise, volume, etc.). However, only the semantic analysis results corresponding to conversation content related to audio quality will affect the audio processing strategy. Therefore, electronic device 100 can adjust its current uplink audio processing strategy based on the first semantic analysis result if the result is related to audio quality. Alternatively, electronic device 100 can adjust its current uplink audio processing strategy based on the first semantic analysis result if it confirms that the audio quality of the call needs to be improved. For example, if the first semantic analysis result is "external noise is high," "local volume is low," or "local volume is high," all of these indicate poor audio quality, and electronic device 100 can determine that the audio quality needs to be improved.
[0198] For example, if the first semantic analysis result is "the current external noise is relatively high," it indicates that the current uplink noise reduction intensity is insufficient, and the electronic device 100 can select a higher uplink noise reduction intensity. For instance, if the current uplink noise reduction intensity is "weak," the electronic device 100 can adjust the uplink noise reduction intensity to "medium" or "strong," and if the current uplink noise reduction intensity is "medium," the electronic device 100 can adjust the uplink noise reduction intensity to "strong." In this embodiment, a corresponding noise reduction model can be configured for each uplink noise reduction intensity. When a certain uplink noise reduction intensity is selected, the collected external audio can be processed by the noise reduction model corresponding to that uplink noise reduction intensity. This embodiment does not specifically limit the noise reduction model used by the electronic device 100; it can be a deep neural network model or other noise reduction models. It should be noted that for some users, the listening experience is best when the target voice is mixed with a little ambient noise during a call. Therefore, by providing multiple noise reduction intensities, the flexibility of noise reduction can be improved, meeting the noise reduction needs of different users and providing a better user experience.
[0199] For another example, if the first semantic analysis result is "the current local sound is relatively low", it indicates that the current uplink audio gain is insufficient, and the electronic device 100 can select a larger uplink audio gain. For instance, if the current uplink audio gain is 5dB, the electronic device 100 can adjust the uplink audio gain to a larger audio gain such as 8dB, 10dB, or 12dB.
[0200] For another example, if the first semantic analysis result is "the current local sound is relatively loud," it indicates that the current uplink audio gain is relatively high, and electronic device 100 can select a smaller uplink audio gain. For instance, if the current uplink audio gain is 10dB, electronic device 100 can adjust the uplink audio gain to a smaller audio gain such as 8dB or 5dB. It is understandable that when a certain uplink audio gain (such as 5dB) is set, electronic device 100 can perform corresponding gain processing (such as 5dB gain processing) on the acquired external audio, and then send the gain-processed audio data to electronic device 200.
[0201] Understandably, in some cases, the first semantic analysis result may also be "the external noise at the current local end is relatively loud, and the sound at the current local end is relatively quiet". In this case, the electronic device 100 can adjust the uplink noise reduction intensity and the uplink audio gain at the same time.
[0202] It should be noted that, in one possible implementation, some semantic analysis results and corresponding adjustment strategies can be pre-configured. Then, the electronic device 100 can compare the first semantic analysis result with the pre-configured semantic analysis results. If the first semantic analysis result is the same as a pre-configured semantic analysis result, the corresponding adjustment strategy can be applied. These pre-configured semantic analysis results can represent situations where the audio quality of the call needs to be improved, and the adjustment strategy can be used to indicate the parameters (such as noise reduction intensity, audio gain, etc.) and adjustment methods (such as selecting a higher-level parameter) in the audio processing strategy that need to be adjusted. For example, the pre-configured semantic analysis results may include "currently, the external noise at the local end is relatively high," "currently, the local volume is relatively low," and "currently, the local volume is relatively high." The adjustment strategy corresponding to "currently, the external noise at the local end is relatively high" could be selecting a higher-level uplink noise reduction intensity; the adjustment strategy corresponding to "currently, the local volume is relatively low" could be selecting a higher-level uplink audio gain; and the adjustment strategy corresponding to "currently, the local volume is relatively high" could be selecting a lower-level uplink audio gain. In this way, when the first semantic analysis result is one of multiple preset semantic results, the electronic device 100 can adjust the uplink audio processing strategy based on the first semantic analysis result, that is, adjust the uplink audio processing strategy based on the adjustment strategy corresponding to the first semantic analysis result.
[0203] In the above implementation, electronic device 100 mainly performs semantic analysis on the audio data from electronic device 200, and then dynamically adjusts the uplink audio processing strategy based on the semantic analysis results. Besides dynamically adjusting the uplink audio processing strategy based on semantic analysis, in this embodiment, electronic device 100 can also analyze the audio data (such as the first audio data) from electronic device 200 to determine the noise level in the corresponding audio data, and then dynamically adjust the uplink audio processing strategy based on the real-time noise level. For example, electronic device 100 can analyze the first audio data to determine the ratio of noise to the target human voice in the first audio data, such as the amplitude ratio of noise to the target human voice. If the amplitude ratio of noise to the target human voice is less than threshold 1 (e.g., 0.5), it indicates that the noise at the other end (e.g., electronic device 200 or user B) is relatively low, and the uplink noise reduction intensity can be adjusted to weak. If the amplitude ratio of noise to the target human voice is greater than or equal to threshold 1 (e.g., 0.5) and less than threshold 2 (e.g., 1.5), it indicates that the noise at the other end is relatively high, and the uplink noise reduction intensity can be adjusted to medium. If the amplitude ratio of noise to the target human voice is greater than or equal to threshold 2, it indicates that the noise at the other end is very high, and the uplink noise reduction intensity can be adjusted to strong. This is because the larger the amplitude ratio of noise to the target human voice, the noisier the environment at the other end, and the more noisy the environment at the other end, the greater the impact on the call of the user at the other end, and the more likely they are to have difficulty hearing clearly. Therefore, when the environment of the other end is noisier, the electronic device 100 can use a higher level of noise reduction, so as to minimize the noise in the audio data sent to the other end, thereby reducing the impact on the other end user.
[0204] 904. Electronic device 200 sends a first indication message to electronic device 100, the first indication message being used to indicate an audio quality problem of electronic device 100.
[0205] During a call between electronic device 100 and electronic device 200, electronic device 200 can monitor the call quality of electronic device 100 in real time. When electronic device 100 has an audio quality problem, electronic device 200 can send a first indication message to electronic device 100. The first indication message can be used to indicate the audio quality problem of electronic device 100.
[0206] Specifically, in one possible implementation, electronic device 200 can score the audio data (such as third audio data) from electronic device 100 based on a voice quality detection algorithm to obtain a real-time voice quality score. If the voice quality score is greater than a quality score threshold (e.g., 6, voice quality scores range from 1 to 10), it indicates that the call quality of electronic device 100 is good, and electronic device 200 can do nothing. Conversely, if the voice quality score is less than or equal to the quality score threshold, it indicates that the call quality of electronic device 100 is poor, and electronic device 200 can determine the reason for the voice quality score being lower than the quality score threshold based on the corresponding audio data (such as third audio data) (e.g., the external noise is loud at the local end, the local sound is low, etc.). After that, electronic device 200 can send a first indication information to electronic device 100. The first indication information can be used to indicate the reason for the voice quality score being lower than the quality score threshold, that is, the audio quality problem existing in the third audio data. For example, suppose electronic device 200 scores the third audio data from electronic device 100 and obtains a speech quality score of 5. Electronic device 200 can determine that this speech quality score is less than a quality score threshold (e.g., 6). Then, electronic device 200 can analyze the third audio data to determine the reason why the speech quality score is lower than the quality score threshold. For example, electronic device 200 can determine that the reason for the speech quality score being lower than the quality score threshold is that "the third audio data contains too much ambient noise." It should be noted that the embodiments of this application do not limit the method by which electronic device 200 analyzes the audio quality problems existing in the third audio data. For example, electronic device 200 can calculate the average amplitude of the third audio data and determine whether the sound is too loud or too soft based on the average amplitude.
[0207] In this application embodiment, the first indication information can be implemented in various ways. For example, the first indication information can be text content that directly indicates the audio effect, such as "the external noise on this end is relatively loud", "the sound on this end is relatively quiet", or "the sound on this end is relatively loud". Alternatively, the first indication information can be an identifier of the audio effect, such as identifier "0" indicating "the external noise on this end is relatively loud", identifier "1" indicating "the sound on this end is relatively quiet", and identifier "2" indicating "the sound on this end is relatively loud".
[0208] In this embodiment of the application, a (dedicated) communication channel can be established between electronic device 100 and electronic device 200, and electronic device 200 can send first instruction information through the communication channel.
[0209] It should be understood that the above-described method for determining audio quality problems of electronic device 100 is merely illustrative and is not limited in this application. For example, in some possible implementations, the voice quality detection algorithm used by electronic device 200 can score various dimensions of the audio data from electronic device 100, and can directly determine whether the corresponding audio data includes a lot of noise, and whether the audio amplitude is large or small, etc.
[0210] 905. Electronic device 100 adjusts the uplink audio processing strategy based on the first instruction information.
[0211] After receiving the first instruction information from the electronic device 200, the electronic device 100 can adjust the currently used uplink audio processing strategy based on the first instruction information.
[0212] For example, if the first indication information indicates that "the current external noise is relatively high", it means that the current uplink noise reduction intensity is insufficient, and the electronic device 100 can select a higher uplink noise reduction intensity. For instance, if the current uplink noise reduction intensity is "weak", the electronic device 100 can adjust the uplink noise reduction intensity to "medium" or "strong", and if the current uplink noise reduction intensity is "medium", the electronic device 100 can adjust the uplink noise reduction intensity to "strong".
[0213] For another example, if the first indication message indicates "the current local volume is relatively low", it means that the current uplink audio gain is insufficient, and the electronic device 100 can select a larger uplink audio gain. For example, if the current uplink audio gain is 5dB, the electronic device 100 can adjust the uplink audio gain to a larger audio gain such as 8dB, 10dB, or 12dB.
[0214] For another example, if the first indication message indicates "the current local volume is relatively loud", it means that the current uplink audio gain is relatively high, and the electronic device 100 can select a smaller uplink audio gain. For example, if the current uplink audio gain is 10dB, the electronic device 100 can adjust the uplink audio gain to a smaller audio gain such as 8dB or 5dB.
[0215] It should be noted that, in one possible implementation, a corresponding adjustment strategy can be configured for each value of the first indication information. Then, when the electronic device 100 receives the first indication information from the electronic device 200, it can directly adopt the corresponding adjustment strategy. The adjustment strategy can be used to indicate the parameters (such as noise reduction intensity, audio gain, etc.) and adjustment method (such as selecting a higher-level parameter) in the audio processing strategy that needs adjustment. For example, when the first indication information is text content directly representing the audio effect, it can include values such as "external noise is currently high," "local sound is currently low," and "local sound is currently high." The adjustment strategy corresponding to "external noise is currently high" can be selecting a higher-level uplink noise reduction intensity; the adjustment strategy corresponding to "local sound is currently low" can be selecting a higher-level uplink audio gain; and the adjustment strategy corresponding to "local sound is currently high" can be selecting a lower-level uplink audio gain. In this way, the electronic device 100 can adjust the uplink audio processing strategy based on the adjustment strategy corresponding to the first indication information.
[0216] In the above implementation, electronic device 200 can provide feedback on quality problems or indication information to electronic device 100. Electronic device 100 can then adjust its uplink audio processing strategy based on the quality problems or indication information. In this approach, electronic device 100 typically adjusts the uplink audio processing strategy based on corresponding configurations, such as configuring a corresponding adjustment strategy for each value of the first indication information. Besides this implementation, in this embodiment, electronic device 100 can also receive request messages from electronic device 200 and adjust its uplink audio processing strategy based on these request messages. For example, electronic device 200 can provide a user interface with peer audio processing configurations. Users can select one or more configurations such as noise reduction scene, audio gain, and noise reduction intensity based on this user interface. Then, users can click or touch the corresponding confirmation control (control 4822 as shown in Figure 8D). In response to this operation, electronic device 200 can send a second request message to electronic device 100. The second request message requests adjustment of the uplink audio processing strategy. The second request message includes parameters (such as noise reduction scene) in the uplink audio processing strategy to be adjusted and the corresponding parameter values (e.g., shopping mall). Accordingly, electronic device 100 can receive a second request message from electronic device 200. Then, electronic device 100 can display a corresponding prompt message box, which prompts the user to adjust the uplink audio processing strategy requested by the other end. This prompt message box may also include corresponding confirm and reject controls. The confirm control can be used to confirm or accept the request from the other end, and the reject control can be used to reject the request from the other end. In some possible implementations, after receiving the second request message from electronic device 200, electronic device 100 can directly adjust the uplink audio processing strategy based on the second request message, such as switching the noise reduction scenario from the current "General" to "Shopping Mall". For the specific user interface corresponding to the above description, please refer to the user interfaces shown in Figures 8A-8E.
[0217] In some possible implementations, when the smart call mode of electronic device 100 is not enabled, electronic device 200 can also send a first request message to electronic device 100. The first request message can be used to request to enable the smart call mode. Correspondingly, electronic device 100 can receive the first request message from electronic device 200. Afterwards, electronic device 100 can pop up a corresponding prompt message box. The prompt message box is used to prompt the user that the other party requests to enable the smart call mode. The prompt message box can also include corresponding confirm and reject controls. The confirm control can be used to confirm or accept the request from the other party, and the reject control can be used to reject the request from the other party. In some possible implementations, after receiving the first request message from electronic device 200, electronic device 100 can directly enable the smart call mode based on the first request message. In this application embodiment, the situation in which electronic device 200 sends the first request message can include a variety of cases, and several examples are provided below. First, electronic device 200 can respond to user operation and then send the first request message to electronic device 100. As shown in FIG8B, electronic device 200 can respond to user operation acting on the send request control 4811 and send the first request message to electronic device 100. The second method is that electronic device 200 can score the audio data (such as the third audio data) from electronic device 100 based on a voice quality detection algorithm to obtain a real-time voice quality score. If the voice quality score is less than or equal to the quality score threshold (such as 6), it indicates that the audio quality of the audio data from electronic device 100 is poor. In this case, electronic device 200 can send a first request message to electronic device 100.
[0218] 906. Electronic device 100 adjusts the uplink noise reduction scenario based on the first positioning result.
[0219] In this embodiment, multiple uplink noise reduction scenarios / audio processing scenarios are provided, with different noise reduction models corresponding to different uplink noise reduction scenarios. Therefore, to obtain better uplink noise reduction effects, electronic device 100 can dynamically adjust the uplink noise reduction scenario based on the first positioning result. The first positioning result can be a positioning result obtained during a call between electronic device 100 and electronic device 200, or a positioning result obtained before the call between electronic device 100 and electronic device 200. For example, when electronic device 100 and electronic device 200 first begin a call, electronic device 100 can initiate positioning (such as satellite positioning), and the obtained positioning result can be used as the first positioning result. As another example, some time before the call between electronic device 100 and electronic device 200 (such as the first few seconds), electronic device 100 initiated positioning and obtained a positioning result, which can be used as the first positioning result.
[0220] For example, during a call between electronic device 100 and electronic device 200, assuming the current uplink noise reduction scenario is set to "General", if the first location result is "XX Shopping Mall in District B, City A", then electronic device 100 can switch the uplink noise reduction scenario to "Shopping Mall". Similarly, if the first location result is "XX Subway Station", then electronic device 100 can switch the uplink noise reduction scenario to "Subway Station".
[0221] Understandably, since the location of electronic device 100 may change in real time, during the communication between electronic device 100 and electronic device 200, electronic device 100 can periodically (e.g., every 30 seconds) initiate location tracking to adjust the noise reduction scenario in a timely manner.
[0222] It is understandable that in the above implementation, electronic device 100 can adjust the uplink noise reduction scenario based on the positioning result. In another possible implementation, electronic device 100 can also adjust the uplink noise reduction scenario based on semantic analysis. Specifically, electronic device 100 can acquire local audio data and audio data from electronic device 200 in real time, then perform real-time semantic analysis, and dynamically adjust the uplink noise reduction scenario based on the semantic analysis result. For example, suppose the dialogue between user A and user B includes (user B: Where are you?, user A: I'm in the mall). In this case, electronic device 100 performs semantic analysis on its local audio data and audio data from electronic device 200, and the semantic analysis result can be "Currently, the local device is located in the mall". Therefore, electronic device 100 can switch the uplink noise reduction scenario to the mall.
[0223] It is understood that the three processing methods—dynamically adjusting the uplink audio processing strategy based on semantic analysis, dynamically adjusting the uplink audio processing strategy based on the first indication information, and dynamically adjusting the uplink noise reduction scenario based on the positioning results—can be parallel technical means, without any dependency on each other. Therefore, in some possible implementations, the electronic device 100 can use any one or more of these three methods for processing, and this application embodiment does not limit this. For example, the electronic device 100 can dynamically adjust the uplink audio processing strategy solely based on semantic analysis. Another example is that the electronic device can dynamically adjust the uplink audio processing strategy solely based on the first indication information. Yet another example is that the electronic device 100 can dynamically adjust the uplink noise reduction scenario solely based on the positioning results.
[0224] The execution order of steps 901-903, 904-905, and 906 is not limited in the embodiments of this application.
[0225] In this embodiment, audio processing may include uplink audio processing and downlink audio processing. Figure 9 above mainly illustrates the case where the electronic device 100 only includes uplink audio processing functionality. Downlink audio processing is similar to uplink audio processing and can be referred to accordingly. The following is a brief description of the case where the electronic device 100 only includes downlink audio processing functionality.
[0226] Please refer to Figure 10, which is a schematic flowchart of another audio processing method disclosed in an embodiment of this application. As shown in Figure 10, the method may include, but is not limited to, the following steps:
[0227] 1001. During a conversation between electronic device 100 and electronic device 200, electronic device 100 performs semantic analysis based on the second audio data to obtain the second semantic analysis result.
[0228] For example, suppose user A holds electronic device 100 and user B holds electronic device 200. When user A needs to contact user B, user A and user B can communicate through electronic device 100 and electronic device 200 (such as voice call, video call, etc.).
[0229] During a conversation between electronic device 100 and electronic device 200, electronic device 100 can use its microphone to collect external sounds in real time, such as ambient sounds and user A's voice, and can send the corresponding audio data (such as second audio data) to electronic device 200. Correspondingly, electronic device 200 can receive audio data from electronic device 100. Similarly, electronic device 200 can also use its microphone to collect external sounds in real time, such as ambient sounds and user B's voice, and can send the corresponding audio data to electronic device 100. Correspondingly, electronic device 100 can receive audio data from electronic device 200.
[0230] In this embodiment of the application, the electronic device 100 can perform semantic analysis on the collected external audio in real time, so as to dynamically adjust the current downlink audio processing strategy based on the semantic analysis results.
[0231] For example, electronic device 100 can acquire second audio data it collects. Then, electronic device 100 can perform semantic analysis on the second audio data to obtain a second semantic analysis result. For instance, the second audio data could be phrases like "It's so noisy over there," "The noise over there is so loud," "It's a bit noisy over there, I can't hear you clearly," "It's still so noisy over there," or "The noise over there is still so loud." The corresponding second semantic analysis result for these audio data could be "The noise at the other end is relatively loud." As another example, the second audio data could be phrases like "Your voice is so soft," "I can't hear what you're saying," "Could you speak louder?" or "Your voice is still a bit soft." The corresponding second semantic analysis result could be "The voice at the other end is relatively soft." And as yet another example, the second audio data could be phrases like "Your voice is so loud," "Could you speak softer?" or "Your voice is too loud." The corresponding second semantic analysis result could be "The voice at the other end is relatively loud."
[0232] In some possible implementations, the electronic device 100 may perform semantic analysis on the second audio data only when the smart call mode is enabled, so as to dynamically adjust the current downlink audio processing strategy based on the semantic analysis results. That is, when the smart call mode is disabled, the electronic device 100 may not perform any processing. In other possible implementations, even when the smart call mode is disabled, the electronic device 100 may still perform semantic analysis on the second audio data, so as to determine whether to enable the smart call mode based on the semantic analysis results. Furthermore, after enabling the smart call mode, the electronic device 100 may continue to perform semantic analysis on subsequently acquired audio data, so as to dynamically adjust the current downlink audio processing strategy based on the semantic analysis results. For example, the electronic device 100 may automatically enable the smart call mode when the semantic analysis result is "the noise at the other end is relatively high," "the sound at the other end is relatively low," or "the sound at the other end is relatively high," so as to perform downlink noise reduction, adjust audio amplitude, and other audio processing. In this way, the smart call mode can be automatically enabled based on the semantic analysis results during the call, which can meet the usage needs in a timely manner, ensure call quality, and reduce the overall power consumption of the electronic device.
[0233] 1002. Electronic device 100 adjusts downlink audio processing strategy based on the second semantic analysis result.
[0234] After obtaining the second semantic analysis result, the electronic device can adjust the currently used downlink audio processing strategy based on the second semantic analysis result.
[0235] It is understandable that conversations between user A and user B typically include content that does not relate to audio quality or effects (such as noise, volume, etc.). However, only the semantic analysis results corresponding to conversation content related to audio quality will affect the audio processing strategy. Therefore, electronic device 100 can adjust its current uplink audio processing strategy based on the second semantic analysis result if the result is related to audio quality. Alternatively, electronic device 100 can adjust its current uplink audio processing strategy based on the second semantic analysis result if it confirms that the audio quality of the call needs to be improved. For example, if the second semantic analysis result is "the noise at the other end is relatively high," "the volume at the other end is relatively low," or "the volume at the other end is relatively high," all of these indicate poor audio quality, and electronic device 100 can determine that the audio quality of the call needs to be improved.
[0236] For example, if the second semantic analysis result is "the current noise at the peer end is relatively large", it indicates that the current downlink noise reduction intensity is insufficient, and the electronic device 100 can select a larger downlink noise reduction intensity. For instance, if the current downlink noise reduction intensity is "weak", the electronic device 100 can adjust the downlink noise reduction intensity to "medium" or "strong", and if the current downlink noise reduction intensity is "medium", the electronic device 100 can adjust the downlink noise reduction intensity to "strong". In this embodiment, a corresponding noise reduction model can be configured for each downlink noise reduction intensity. When a certain downlink noise reduction intensity is selected, the received audio data from the electronic device 200 can be processed by the noise reduction model corresponding to that downlink noise reduction intensity.
[0237] For another example, if the second semantic analysis result is "the current sound from the other end is relatively low", it indicates that the current downlink audio gain is insufficient, and the electronic device 100 can select a larger downlink audio gain. For instance, if the current downlink audio gain is 5dB, the electronic device 100 can adjust the downlink audio gain to a larger audio gain such as 8dB, 10dB, or 12dB.
[0238] For another example, if the second semantic analysis result is "the current sound from the other end is relatively loud," it indicates that the current downlink audio gain is relatively high, and electronic device 100 can select a smaller downlink audio gain. For instance, if the current downlink audio gain is 10dB, electronic device 100 can adjust the downlink audio gain to a smaller audio gain such as 8dB or 5dB. It is understandable that when a certain downlink audio gain (such as 5dB) is set, electronic device 100 can perform corresponding downlink gain processing (such as 5dB gain processing) on the audio received from electronic device 200, and then play the corresponding audio through a handset or speaker.
[0239] It should be noted that, in one possible implementation, some semantic analysis results and corresponding adjustment strategies can be pre-configured. Then, the electronic device 100 can compare the second semantic analysis result with the pre-configured semantic analysis results. If the second semantic analysis result is the same as a pre-configured semantic analysis result, the corresponding adjustment strategy can be applied. These pre-configured semantic analysis results can represent situations where the audio quality of the call needs to be improved, and the adjustment strategy can be used to indicate the parameters and adjustment methods in the audio processing strategy that need to be adjusted. For example, the pre-configured semantic analysis results may include "currently, the noise at the other end is relatively high," "currently, the sound at the other end is relatively low," and "currently, the sound at the other end is relatively high." The adjustment strategy corresponding to "currently, the noise at the other end is relatively high" can be to select a higher level of downlink noise reduction intensity; the adjustment strategy corresponding to "currently, the sound at the other end is relatively low" can be to select a higher level of downlink audio gain; and the adjustment strategy corresponding to "currently, the sound at the other end is relatively high" can be to select a lower level of downlink audio gain. Thus, when the second semantic analysis result is one of multiple pre-configured semantic results, the electronic device 100 can adjust the downlink audio processing strategy based on the second semantic analysis result, that is, adjust the downlink audio processing strategy based on the adjustment strategy corresponding to the second semantic analysis result.
[0240] 1003. Electronic device 200 sends fourth audio data to electronic device 100.
[0241] During a conversation between electronic device 100 and electronic device 200, electronic device 200 can collect external sounds through its microphone to obtain corresponding fourth audio data, which it can then send to electronic device 100. Correspondingly, electronic device 100 can receive the fourth audio data from electronic device 200.
[0242] 1004. Electronic device 100 performs audio processing on the fourth audio data to obtain the fifth audio data.
[0243] After receiving the fourth audio data, the electronic device 100 can process the fourth audio data according to the current downlink audio processing strategy to obtain the fifth audio data. That is to say, the fifth audio data is the audio data obtained by the electronic device 100 through noise reduction processing, gain processing, etc., according to the current downlink audio processing strategy.
[0244] 1005. Electronic device 100 analyzes the fifth audio data to determine the audio quality problems of the fifth audio data.
[0245] During a call between electronic device 100 and electronic device 200, electronic device 100 can process audio data from electronic device 200 in real time using its own downlink audio processing strategy, and can monitor the call quality of the processed audio data (such as the fifth audio data) in real time. If there is an audio quality problem in the processed audio data, electronic device 100 can adjust the currently used downlink audio processing strategy based on the audio quality problem.
[0246] Specifically, in one possible implementation, electronic device 100 can score the fifth audio data based on a voice quality detection algorithm to obtain a real-time voice quality score. If the voice quality score is greater than a quality score threshold (e.g., 6, where 1 to 10 is a threshold), it indicates that electronic device 100 does not need to process the audio from electronic device 200. If the voice quality score is less than or equal to the quality score threshold, it indicates that the call quality of electronic device 200 is poor. Electronic device 100 can determine the reason why the voice quality score is lower than the quality score threshold based on the corresponding audio data (e.g., high ambient noise, low volume). For example, assuming electronic device 100 scores the fifth audio data from electronic device 200 and obtains a voice quality score of 5, electronic device 100 can determine that the voice quality score is less than the quality score threshold (e.g., 6). Then, electronic device 100 can analyze the fifth audio data to determine the reason why the voice quality score is lower than the quality score threshold. For example, electronic device 100 can determine that the reason for the voice quality score being lower than the quality score threshold is that "the fifth audio data contains too much ambient noise."
[0247] 1006. Electronic device 100 Adjusts downlink audio processing strategy based on audio quality issues of fifth audio data.
[0248] After determining that there is an audio quality problem with the fifth audio data, the electronic device 100 can adjust the currently used downlink audio processing strategy based on the audio quality problem of the fifth audio data.
[0249] For example, if the audio quality problem of the fifth audio data is "external noise is relatively large", it indicates that the current downlink noise reduction intensity is insufficient, and the electronic device 100 can select a greater downlink noise reduction intensity. For instance, if the current downlink noise reduction intensity is "weak", the electronic device 100 can adjust the downlink noise reduction intensity to "medium" or "strong", and if the current downlink noise reduction intensity is "medium", the electronic device 100 can adjust the downlink noise reduction intensity to "strong".
[0250] For another example, if the audio quality problem of the fifth audio data is "the sound is too low," it indicates that the current downlink audio gain is insufficient, and the electronic device 100 can select a higher downlink audio gain. For instance, if the current downlink audio gain is 5dB, the electronic device 100 can adjust the downlink audio gain to a higher audio gain such as 8dB, 10dB, or 12dB.
[0251] For another example, if the audio quality problem of the fifth audio data is "the current sound is too loud", it indicates that the current downlink audio gain is too high, and the electronic device 100 can select a lower downlink audio gain. For example, if the current downlink audio gain is 10dB, the electronic device 100 can adjust the downlink audio gain to a smaller audio gain such as 8dB or 5dB.
[0252] 1007. Electronic device 100 adjusts downlink noise reduction scenario based on second positioning result.
[0253] In this embodiment, multiple downlink noise reduction scenarios / audio processing scenarios are provided, with different noise reduction models corresponding to different downlink noise reduction scenarios. Therefore, to obtain better noise reduction effects, electronic device 100 can dynamically adjust the downlink noise reduction scenario based on the second positioning result. The second positioning result can be a positioning result obtained during a call between electronic device 100 and electronic device 200, or a positioning result obtained before the call between electronic device 100 and electronic device 200. For example, when electronic device 100 and electronic device 200 first begin a call, electronic device 200 can initiate positioning (such as satellite positioning), and the obtained positioning result can be used as the second positioning result, which can then be sent to electronic device 100. As another example, some time before the call between electronic device 100 and electronic device 200 (such as the first few seconds), electronic device 200 may have initiated positioning and obtained a positioning result, which can be used as the second positioning result.
[0254] It is understandable that, since the location of electronic device 200 may change in real time, during the communication between electronic device 100 and electronic device 200, electronic device 200 can periodically (e.g., every 30 seconds) initiate location positioning and send the location results to electronic device 100 so that electronic device 100 can adjust the downlink noise reduction scenario in a timely manner.
[0255] In another possible implementation, electronic device 100 can also adjust the downlink noise reduction scenario based on semantic analysis. Specifically, electronic device 100 can acquire local audio data and audio data from electronic device 200 in real time, then perform real-time semantic analysis, and dynamically adjust the downlink noise reduction scenario based on the semantic analysis results. For example, suppose the dialogue between user A and user B includes (user A: Where are you?, user B: I'm in the mall). In this case, electronic device 100 performs semantic analysis on its local audio data and audio data from electronic device 200, and the semantic analysis result can be "the other party is currently in the mall". Therefore, electronic device 100 can switch the downlink noise reduction scenario to the mall.
[0256] The execution order of steps 1001-1003, 1004-1005, and 1006 is not limited in this embodiment.
[0257] The above descriptions of uplink and downlink audio processing are exemplary. However, it should be understood that in some possible implementations, the electronic device may support both uplink and downlink audio processing simultaneously, and this application embodiment does not limit this. Furthermore, for the same noise reduction scenario, under the same noise reduction intensity, the noise reduction models used for uplink and downlink can be the same.
[0258] In the above processing flow, for both uplink and downlink audio processing, the electronic device 100 can automatically adjust the noise reduction intensity, audio gain, and noise reduction scenario based on semantic analysis results, and can also automatically adjust the noise reduction scenario based on location. Furthermore, it also supports users manually adjusting the noise reduction intensity, audio gain, and noise reduction scenario. This allows for the selection of the most suitable noise reduction model and corresponding audio gain based on the actual situation, improving call quality and providing users with a better call experience.
[0259] It should be understood that the transmission in the embodiments of this application can be direct or indirect. Direct transmission means that one device or module directly sends information / data to the corresponding device or module, while indirect transmission means that one device or module sends information / data to the corresponding device or module through other devices or modules.
[0260] Obviously, the embodiments described above are merely some embodiments of this application, and not all embodiments. The term "embodiment" as used herein means that a specific feature, structure, or characteristic described in connection with an embodiment can be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily indicate the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will understand, explicitly and implicitly, that the embodiments described herein can be combined with other embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application. The terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, it may include a series of steps or units, or optionally, steps or units not listed, or optionally other steps or units inherent to these processes, methods, products, or devices. It is understood that in some embodiments, the equality sign of the above conditional judgment can be greater than or less than one end. For example, the above conditional judgment of a threshold being greater than, less than, or equal to can be changed to a conditional judgment of the threshold being greater than or equal to, or less than, without limitation herein. It is also understandable that, for an architecture with multiple devices or modules, if one device or module generates a piece of information and another device or module uses that information, there are multiple ways for the other device to obtain that information. For example, the device or module that generated the information may send the information directly to the device or module that used the information (equivalent to direct sending), or the device or module that generated the information may send the information to the device or module that used the information through other devices or modules (equivalent to indirect sending).
[0261] It is understood that the accompanying drawings show only the parts relevant to this application and not all of them. It should be understood that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged, as long as it is logically sound. The process can be terminated when its operations are completed, but it may also have additional steps not included in the drawings. The process may correspond to a method, function, procedure, subroutine, subprogram, etc.
[0262] The terms “component,” “module,” “system,” “unit,” etc., used in this specification are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a unit can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, a thread of execution, a program, and / or distributed between two or more computers. Furthermore, these units can be executed from various computer-readable media on which various data structures are stored. For example, a unit can communicate via local and / or remote processes based on signals having one or more data packets (e.g., data from a second unit interacting with another unit between a local system, a distributed system, and / or a network; for example, the Internet interacting with other systems via signals).
[0263] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above description is only a specific embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of this application should be included within the scope of protection of this application.
Claims
1. An audio processing method, applied to a first electronic device, characterized in that, include: During a call between the first electronic device and the second electronic device, semantic analysis is performed on the call audio data to obtain the target semantic analysis result; the call audio data includes call audio data from the second electronic device; Adjust the audio processing strategy based on the target semantic analysis results; The audio processing strategy includes an uplink audio processing strategy, which is used to process the call audio data collected by the first electronic device.
2. The method according to claim 1, characterized in that, The uplink audio processing strategy includes one or more of the following: uplink noise reduction intensity, uplink audio gain, and uplink noise reduction scenario.
3. The method according to claim 1 or 2, characterized in that, The adjustment of the audio processing strategy based on the target semantic analysis results includes: When the target semantic analysis result is one of a plurality of preset semantic analysis results. The audio processing strategy is adjusted based on the adjustment strategy corresponding to the target semantic result; each of the multiple preset semantic analysis results is configured with a corresponding adjustment strategy, and the adjustment strategy is used to indicate the parameters and adjustment methods in the audio processing strategy that need to be adjusted.
4. The method according to any one of claims 1-3, characterized in that, The method further includes: Initiate location tracking and obtain the first location result; Based on the first positioning result, the uplink noise reduction scenario is adjusted, and different noise reduction scenarios correspond to different noise reduction models.
5. The method according to any one of claims 1-4, characterized in that, The method further includes: Receive first indication information from the second electronic device, the first indication information being used to indicate an audio quality problem of the first electronic device; Adjust the uplink audio processing strategy based on the first indication information.
6. The method according to claim 5, characterized in that, The adjustment of the uplink audio processing strategy based on the first indication information includes: The uplink audio processing strategy is adjusted based on the adjustment strategy corresponding to the first indication information; the first indication information includes multiple value cases, each of which is configured with a corresponding adjustment strategy, and the adjustment strategy is used to indicate the parameters and adjustment methods in the audio processing strategy that need to be adjusted.
7. The method according to any one of claims 1-6, characterized in that, Before adjusting the audio processing strategy based on the target semantic analysis results, the method further includes: Semantic analysis is performed on the sixth audio data to obtain the sixth semantic analysis result; the sixth audio data includes call audio data from the second electronic device and / or call audio data collected by the first electronic device; If the sixth semantic analysis result is one of several preset semantic analysis results, the intelligent call mode is activated; The adjustment of the audio processing strategy based on the target semantic analysis results includes: When the intelligent call mode is enabled, the audio processing strategy is adjusted based on the target semantic analysis results.
8. The method according to any one of claims 1-6, characterized in that, Before adjusting the audio processing strategy based on the target semantic analysis results, the method further includes: Receive a first request message from the second electronic device, the first request message being used to request the activation of smart call mode; Intelligent call mode is activated based on the first request message; The adjustment of the audio processing strategy based on the target semantic analysis results includes: When the intelligent call mode is enabled, the audio processing strategy is adjusted based on the target semantic analysis results.
9. The method according to any one of claims 1-8, characterized in that, The method further includes: Receive a second request message from the second electronic device. The second request message is used to request adjustment of the uplink audio processing strategy. The second request message includes parameters in the uplink audio processing strategy that need to be adjusted and the corresponding parameter values. Adjust the uplink audio processing strategy based on the second request message.
10. The method according to any one of claims 1-9, characterized in that, The method further includes: Display a first user interface, which includes audio processing configuration controls; In response to a user operation on the audio processing configuration control, a second user interface is displayed; the second user interface includes an uplink audio processing configuration area, which includes one or more of a noise reduction intensity configuration area, an audio gain configuration area, and a noise reduction scene configuration area. The noise reduction intensity configuration area includes multiple noise reduction intensity options, the audio gain configuration area includes multiple audio gain options, and the noise reduction scene configuration area includes multiple noise reduction scene options. In response to a user action on the first option, the first option is changed to a selected state, and the audio processing strategy corresponding to the first option is enabled; the first option is any option in the second user interface.
11. An electronic device, characterized in that, It includes a processor and a communication interface; the communication interface is used to receive and send data; the processor calls a computer program or computer instructions stored in memory to implement the method as described in any one of claims 1-10.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program or computer instructions that are executed by a processor to implement the method as described in any one of claims 1-10.
13. A computer program product, characterized in that, The computer program product includes computer program code or computer instructions, which, when executed, implement the method described in any one of claims 1-10.