Call quality detection method and related device
By detecting uplink network speed and call audio status, the system can identify stuttering in voice or video calls, perform network optimization, resolve poor call quality issues caused by network fluctuations, and improve user experience.
Patent Information
- Application Number
- CN202410598692.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-14
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-05-14
AI Technical Summary
During voice or video calls, network fluctuations can cause poor call quality, resulting in stuttering and silence, which negatively impacts the user experience.
By detecting uplink network speed and whether there is sound during the call, setting speed and audio loudness thresholds, and accumulating parameters, it is possible to determine whether there is a pause in the call, and perform network optimization when a pause is detected.
Accurately determine if there is a pause in the call, reduce false alarms, improve user experience, and enhance network quality.
Smart Images

Figure CN121001101A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of terminal, and in particular, to a call quality detection method and related device. BACKGROUND
[0002] Some applications in an electronic device can support voice call function or video call function, and a user can use the applications to make voice call or video call.
[0003] However, in the scenario of voice call or video call, the electronic device can cause poor call quality due to network fluctuation and other factors, for example, application freezing and other problems, thereby affecting user experience. SUMMARY
[0004] The call quality detection method and related device provided by the embodiments of the present application can determine that the application freezes in the call process when the electronic device detects that the uplink network rate or the downlink network rate continuously falls below the rate threshold and the call has sound, and the electronic device can optimize the network, thereby reducing the freezing of the application and improving user experience.
[0005] In a first aspect, the call quality detection method provided by the embodiments of the present application comprises:
[0006] The electronic device enters a call; in a first time period in the call, if the network rate in the call is less than a rate threshold, a first parameter is accumulated by a first value, and the first parameter is used to indicate the number of times of low rate in the call; in the first time period, if the audio loudness in the call is greater than an audio loudness threshold, a second parameter is accumulated by a second value, and the second parameter is used to indicate the number of times of sound in the call; in the first time period, if the first parameter is greater than a first threshold and the second parameter is greater than a second threshold, it is determined that the call freezes. In this way, it can be determined that freezing occurs in the call process, and the electronic device can optimize the network, thereby reducing the freezing of the call and improving user experience.
[0007] In a possible implementation, the method further comprises: in the first time period, if the first parameter is less than or equal to the first threshold and / or the second parameter is less than or equal to the second threshold, it is determined that the call scene does not freeze. In this way, by comprehensively considering the speed of the network rate and whether the call has sound, it can be more accurately determined whether the current call freezes and whether the network rate optimization needs to be further performed. The situation of false judgment of call freezing when the user does not speak and the call is silent is reduced.
[0008] In a possible implementation, the method further includes: setting the first parameter to a first initial value if the network rate is greater than the rate threshold in the first time period. In this way, if the network rate is greater than the rate threshold, it indicates that the network rate of the current call is high, the call quality is good, and no lag occurs, so the first parameter does not need to be accumulated and counted again, and therefore the first parameter can be set to the first initial value.
[0009] In a possible implementation, the method further includes: setting the second parameter to a second initial value if the audio loudness is less than the audio loudness threshold in the first time period. In this way, if the audio loudness is less than the audio loudness threshold, it indicates that the current call is silent, which can be caused by the user not speaking during the call, and cannot be used to determine whether the call has lag, so the second parameter does not need to be accumulated and counted again, and the second parameter can be set to the second initial value.
[0010] In a possible implementation, the network rate at the first time point is an average value calculated based on N network rates recorded in the electronic device before the first time point, and the first time point is within the first time period, and N is a positive integer. In this way, the average rate is used to determine the network rate at the current time point, which can reflect the overall trend of the network rate in a period of time (N network rates), and the data obtained is more stable and is not affected by fluctuations in individual unstable data, so that the determination result is more accurate and the robustness of the algorithm is improved.
[0011] In a possible implementation, the audio loudness at the second time point is an average value calculated based on M root mean square (RMS) values of calls recorded in the electronic device before the second time point, and the second time point is within the first time period, and M is a positive integer. In this way, the average value calculated based on the M root mean square (RMS) values of calls is used to determine the audio loudness at the current time point, which can reflect the overall trend of the audio loudness in a period of time (M root mean square (RMS) values of calls), and the data obtained is more stable and is not affected by fluctuations in individual unstable data, so that the determination result is more accurate and the robustness of the algorithm is improved.
[0012] In a possible implementation, the electronic device includes an audio digital signal processor (ADSP) and a target module, the ADSP includes a detection point of the target module, and the target module is configured to acquire M root mean square (RMS) values of calls from the call based on the detection point and record the M root mean square (RMS) values of calls. In this way, the target module can acquire the root mean square (RMS) values of the adjacent M calls at the current time point in time based on the detection point, so that newer RMS values can be acquired, thereby facilitating real-time detection of whether the call is silent, obtaining RMS values that are more in line with the current time, and making the determination more accurate.
[0013] In a possible implementation, after determining that the call scene is stuck, the method further includes: setting the target value as a third value, the third value being used to instruct the electronic device to perform network optimization processing. In this way, by setting the target value as the third value, the network optimization processing is performed, so that the network quality can be improved and the user experience can be improved.
[0014] In a second aspect, an embodiment of the present application provides a control device for call quality detection. The device can be an electronic device, or a chip or chip system in the electronic device. The device can include a processing unit. The processing unit is configured to implement any method related to processing performed by the first aspect or any possible implementation of the first aspect. When the device is an electronic device, the processing unit can be a processor. The device can further include a storage unit, which can be a memory. The storage unit is configured to store instructions. The processing unit executes the instructions stored in the storage unit, so that the electronic device implements the method described in the first aspect or any possible implementation of the first aspect. When the device is a chip or chip system in the electronic device, the processing unit can be a processor. The processing unit executes the instructions stored in the storage unit, so that the electronic device implements the method described in the first aspect or any possible implementation of the first aspect. The storage unit can be a storage unit (for example, a register, a cache, etc.) in the chip, or a storage unit (for example, a read-only memory, a random access memory, etc.) outside the chip in the electronic device.
[0015] For example, the processing unit is configured to enter the call, and configured to accumulate the first parameter by a first value, and configured to accumulate the second parameter by a second value, and configured to determine that the call scene is stuck.
[0016] In a possible implementation, the processing unit is configured to determine that the call scene is not stuck, if the first parameter is less than or equal to a first threshold value, and / or the second parameter is less than or equal to a second threshold value.
[0017] In a possible implementation, the processing unit is configured to set the first parameter as a first initial value, if the network rate is greater than a rate threshold value.
[0018] In a possible implementation, the processing unit is configured to set the second parameter as a second initial value, if the audio loudness is less than an audio loudness threshold value.
[0019] In a possible implementation, the network rate at the first time is an average value calculated based on N network rates recorded in the electronic device before the first time, the first time is within a first time period, and N is a positive integer.
[0020] In a possible implementation, the audio loudness at the second time point is an average value calculated based on the root mean square (RMS) values of the M conversations recorded in the electronic device before the second time point, the second time point is within the first time period, and M is a positive integer.
[0021] In a possible implementation, the electronic device includes an audio digital signal processor (ADSP) and a target module, the ADSP includes a detection point of the target module, and the target module is configured to obtain the root mean square (RMS) values of the M conversations from the conversation based on the detection point and record the RMS values of the M conversations.
[0022] In a possible implementation, the processing unit is configured to set the target value as the third value.
[0023] In a third aspect, an embodiment of the present application provides an electronic device, including one or more processors and a memory, the memory being coupled to the one or more processors, and the memory being configured to store computer program codes, the computer program codes including computer instructions, and the one or more processors being configured to invoke the computer instructions to cause the electronic device to perform the method described in the first aspect or any possible implementation of the first aspect.
[0024] In a fourth aspect, the present application provides a chip or a chip system, which is applied to an electronic device, and includes one or more processors and a communication interface, the communication interface and the at least one processor are interconnected through a line, and the one or more processors are configured to invoke computer instructions to cause the electronic device to perform the method described in the first aspect or any possible implementation of the first aspect. The communication interface in the chip can be an input / output interface, a pin, or a circuit, etc.
[0025] In a possible implementation, the chip or the chip system described in the present application further includes at least one memory, and the at least one memory stores instructions. The memory can be a storage unit inside the chip, such as a register, a cache, etc., or a storage unit of the chip (such as a read-only memory, a random access memory, etc.).
[0026] In a fifth aspect, an embodiment of the present application provides a computer readable storage medium, which includes computer instructions, and when the computer instructions run on an electronic device, cause the electronic device to perform the method described in the first aspect or any possible implementation of the first aspect.
[0027] In a sixth aspect, an embodiment of the present application provides a computer program product, which includes computer program codes, and when the computer program codes run on an electronic device, cause the electronic device to perform the method described in the first aspect or any possible implementation of the first aspect.
[0028] It should be understood that the second aspect to the sixth aspect of the present application correspond to the technical solutions of the first aspect of the present application, and the beneficial effects obtained by each aspect and the corresponding feasible implementation manners are similar, which will not be repeated. BRIEF DESCRIPTION OF DRAWINGS
[0029] Figure 1 A structural schematic diagram of an electronic device provided for an embodiment of the present application is shown in FIG. 1.
[0030] Figure 2 A software structural schematic diagram of an electronic device provided for an embodiment of the present application is shown in FIG. 2.
[0031] Figure 3 A flowchart of detecting application lag based on uplink rate provided for an embodiment of the present application is shown in FIG. 3.
[0032] Figure 4 A flowchart of detecting application lag based on uplink rate and whether the call has sound provided for an embodiment of the present application is shown in FIG. 4.
[0033] Figure 5 An array structure schematic diagram of cyclically storing RMS values provided for an embodiment of the present application is shown in FIG. 5.
[0034] Figure 6 A logic schematic diagram of judging application lag provided for an embodiment of the present application is shown in FIG. 6.
[0035] Figure 7 A VDM module call detection point schematic diagram provided for an embodiment of the present application is shown in FIG. 7.
[0036] Figure 8 A schematic diagram of a call quality detection method provided for an embodiment of the present application is shown in FIG. 8.
[0037] Figure 9 A structural schematic diagram of a chip provided for an embodiment of the present application is shown in FIG. 9. DETAILED DESCRIPTION
[0038] In order to clearly describe the technical solutions of the embodiments of the present application, the following briefly introduces some terms and technologies involved in the embodiments of the present application:
[0039] 1. Terms
[0040] In the embodiments of the present application, the same items or similar items with basically the same functions and effects are distinguished by using “first”, “second”, etc. For example, the first chip and the second chip are only used to distinguish different chips, and do not limit the sequence. Those skilled in the art can understand that “first”, “second”, etc. do not limit the quantity and execution sequence, and “first”, “second”, etc. also do not necessarily mean different.
[0041] It should be noted that the terms "exemplary" and "for example" are used herein to mean "serving as an example, instance, or illustration," and not "preferred" or "advantageous over other examples." The usage of these terms in this application is not intended to convey any preference or advantage for the embodiments or examples described with such terms.
[0042] In the embodiments of the present application, "at least one" means one or more, and "multiple" means two or more. The "and / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the following three cases: A exists alone, A and B exist together, and B exists alone, wherein A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items, including any combination of single item or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, wherein a, b, and c can be single or multiple.
[0043] 2. Electronic device
[0044] The electronic device of the embodiments of the present application can also be any form of terminal device. For example, the electronic device can include a mobile phone, a tablet computer, a palm computer, a notebook computer, a mobile internet device (MID), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a wireless terminal in self driving, a wireless terminal in remote medical surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, a cellular phone, a cordless phone, a session initiation protocol (SIP) phone, a wireless local loop (WLL) station, a personal digital assistant (PDA), a handheld device with wireless communication function, a computing device or other processing device connected to a wireless modem, a vehicle-mounted device, a wearable device, an electronic device in a 5G network, or an electronic device in a future evolved public land mobile network (PLMN), and the like. The embodiments of the present application are not limited thereto.
[0045] By way of example and not limitation, in the embodiments of the present application, the electronic device can also be a wearable device. The wearable device can also be referred to as a wearable smart device, which is a general term for devices that are designed and developed by applying wearable technology to daily wear, such as glasses, gloves, watches, clothing, and shoes. The wearable device is a portable device that is directly worn on the body or integrated into the clothes or accessories of the user. The wearable device is not only a hardware device, but also has powerful functions through software support and data interaction and cloud interaction. The general wearable smart device includes a device with full functions and large size, which can realize complete or partial functions without relying on a smart phone, such as a smart watch or smart glasses, and a device that focuses on a certain application function and needs to be used in cooperation with other devices, such as a smart phone, such as various smart wristbands and smart jewelry for monitoring vital signs.
[0046] In addition, in the embodiments of the present application, the electronic device can also be an electronic device in an internet of things (IoT) system. The IoT is an important component of future information technology development, and its main technical feature is to connect objects through communication technology and network, so as to realize the intelligent network of man-machine interconnection and object-object interconnection.
[0047] The electronic device in the embodiments of the present application can also be referred to as a user equipment (UE), a mobile station (MS), a mobile terminal (MT), an access terminal, a subscriber unit, a subscriber station, a mobile station, a mobile terminal, a remote station, a remote terminal, a mobile device, a user terminal, a terminal, a wireless communication device, a user agent, or a user equipment, etc.
[0048] In the embodiments of the present application, the electronic device or each network device includes a hardware layer, an operating system layer running on the hardware layer, and an application layer running on the operating system layer. The hardware layer includes central processing unit (CPU), memory management unit (MMU), and memory (also known as main memory) and other hardware. The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux operating system, Unix operating system, Android operating system, iOS operating system, or windows operating system, etc. The application layer includes browsers, address books, word processing software, instant messaging software, etc.
[0049] Exemplary, Figure 1 A structural schematic diagram of the electronic device is shown.
[0050] The electronic device can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyro sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0051] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device can include more or fewer components than the illustration, or combine certain components, or split certain components, or different component arrangements. The illustrated components can include hardware, software, or a combination of software and hardware.
[0052] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices, or can be integrated in one or more processors. The controller can generate operation control signals according to instruction operation codes and timing signals, complete the control of fetching instructions and executing instructions.
[0053] The processor 110 can also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can hold instructions or data that the processor 110 has just used or is using repeatedly. If the processor 110 needs to use the instructions or data again, it can be directly called from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thereby improving the efficiency of the system. For example, in the embodiments of the present application, the processor 110 can be used to detect the uplink rate, detect whether the call has sound, etc.
[0054] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a USB interface, etc.
[0055] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a structural limitation on the electronic device. In other embodiments of the present application, the electronic device can also use different interface connection methods or a combination of multiple interface connection methods in the above embodiments.
[0056] The internal memory 121 can be used to store computer executable program codes, the executable program codes including instructions. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, application programs required by at least one function, and the like. The data storage area can store data created during use of the electronic device, and the like. In addition, the internal memory 121 can include a high-speed random access memory, and can further include a non-volatile memory such as at least one of a magnetic disk storage device, a flash memory device, a universal flash storage (UFS), and the like. The processor 110 performs various function applications and data processing of the electronic device by running instructions stored in the internal memory 121 and / or instructions stored in a memory disposed in the processor. For example, in the embodiments of the present application, the internal memory 121 can be used to store codes related to determining whether the application is stuck based on uplink rate detection and talk activity detection, and the like.
[0057] The electronic device can implement an audio function through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset interface 170D, and an application processor, and the like. For example, voice communication or video communication, and the like.
[0058] The audio module 170 is used to convert digital audio information into an analog audio signal output, and is also used to convert an analog audio input into a digital audio signal. The speaker 170A, also referred to as a "loudspeaker", is used to convert an audio electrical signal into a sound signal, and the electronic device can include one or N speakers 170A, N being a positive integer greater than one. The electronic device can listen to music, video communication, or voice communication, and the like through the speaker 170A. The receiver 170B, also referred to as a "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device is on the phone or receives a voice message, the receiver 170B can be used to listen to the voice by being close to the human ear. The microphone 170C, also referred to as a "microphone" or "microphone", is used to convert a sound signal into an electrical signal. The headset interface 170D is used to connect a wired headset.
[0059] Figure 2 The software architecture diagram of the electronic device of the embodiments of the present application is shown in FIG. 1. The layered architecture divides the software into several layers, each layer has a clear role and division of labor. The layers communicate with each other through a software interface. In some embodiments, the Android system is divided into five layers, from top to bottom, the application layer, the application framework layer, the hardware adaptation layer (hardware adaptation layer, HAL), and the kernel layer.
[0060] The application layer can also be referred to as the application layer, and the application layer can include a series of application packages. For example, the application layer can include a series of application packages such as a system application package, an application package, and the like. Figure 2As shown, the application package can include social networking, video, and other applications. Applications can include system applications and third-party applications.
[0061] The application framework layer, also known as the framework layer, provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The framework layer can include some predefined functions.
[0062] like Figure 2 As shown, the Framework layer may include a call management module, a quality of experience (QOE) module, a self-healing integration module, and an audio decoding module. In addition, the Framework layer may also include an activity manager, window manager, resource manager, notification manager, content provider, and view system (not shown in the figure), etc. For details, please refer to relevant technologies; further elaboration is omitted here.
[0063] The call management module can be used to request the audio and video resources needed for a call, handle downlink operations such as dialing, answering, and hanging up, and manage uplink status such as incoming call status and call waiting status. The call management module may include a Telephony module.
[0064] The quality experience module can be used to acquire network parameters and determine the quality of the network based on these parameters. These network parameters can include network speed, reference signal receiving power (RSRP), etc. For example, network speed can be used to determine call speed, and RSRP can be used to determine the quality of the network signal.
[0065] The integrated self-healing module can be used to obtain the cause value passed by the quality experience module and perform corresponding processing on the network based on the cause value, such as resetting the network.
[0066] The video decoder manager (VDM) is used to manage the acquisition, encoding, decoding, and rendering of audio data. It can also be called an audio decoding module or a VDM module. For ease of description, the VDM module will be used as an example below.
[0067] Understandably, the software architecture of electronic devices can also include other layers, such as the Android runtime and system libraries, which can be found in relevant technologies and will not be elaborated here.
[0068] The hardware abstraction layer is an abstract layer between the kernel layer and the Android runtime. The hardware abstraction layer can be a package of hardware drivers, providing a unified interface for the upper-layer application to call.
[0069] The kernel layer is a layer between hardware and software. In the embodiment of the present application, the kernel layer can include an audio digital signal processor (ADSP) driver, a microphone driver, a speaker driver, and the like.
[0070] The ADSP driver can be used to process audio data, for example, including calculation, conversion, compression, filtering, enhancement, decoding of audio data, and transmission and reception of audio data.
[0071] It should be noted that the present application only takes the Android system as an example to illustrate, and in other operating systems (such as Windows system, IOS system, etc.), as long as the functions of each functional module are similar to the embodiments of the present application, the scheme of the present application can also be implemented.
[0072] In the scenario of voice call or video call, the electronic device can cause poor call quality due to network fluctuations and other factors, for example, call lag, no sound, and the like, thereby affecting the user experience.
[0073] Therefore, the call quality detection method provided in the embodiments of the present application can determine that the application has call lag in the call process when the electronic device detects that the uplink network rate or the downlink network rate is continuously lower than the rate threshold and the call has sound, and the electronic device can optimize the network, thereby reducing the call lag of the application and improving the user experience.
[0074] In the embodiments of the present application, the uplink network rate can be referred to as the uplink rate, and the downlink network rate can be referred to as the downlink rate. For ease of description, the uplink rate is taken as an example for description in the subsequent embodiments of the present application.
[0075] Figure 3 A flowchart for detecting application lag based on the uplink rate is shown.
[0076] It can be understood that, Figure 3 The flowchart of the corresponding embodiment only shows the execution of one loop. In the actual execution process, the electronic device can cyclically execute Figure 3 Steps S303-S307 of the corresponding embodiment are not described again. When the application exits the voice call scenario or the video call scenario, the electronic device can stop executing Figure 3 The flow of the corresponding embodiment.
[0077] S301, entering an application.
[0078] It can be understood that the application can be any application in the electronic device, and embodiments of the present application are not limited.
[0079] In a possible implementation, the electronic device can record an application capable of supporting voice call and / or video call in a whitelist, and the electronic device can execute step S302 based on the application recorded in the whitelist.
[0080] For example, if the application currently running in the foreground of the electronic device is in the whitelist, it means that the application supports voice call function and / or video call function, and then the electronic device can execute step S302; otherwise, the subsequent steps are not continued.
[0081] S302, judging whether the application enters a voice call scene or a video call scene.
[0082] When the application enters the voice call scene or the video call scene, since the application needs to call the call management module to implement the related functions of voice call or video call, the call management module can be used to detect whether the application enters the voice call scene or the video call scene.
[0083] If the application enters the voice call scene or the video call scene, step S303 can be executed; otherwise, the subsequent steps are not continued.
[0084] S303, taking the average rate of the previous n seconds of the uplink network as the input to detect the uplink rate.
[0085] After the application enters the voice call scene or the video call scene, the quality experience module can detect the uplink rate of the audio data in the call process. In a possible implementation, the quality experience module can obtain the audio data through the transmission control protocol (TCP) and / or user datagram protocol (UDP) transmission protocol.
[0086] After obtaining the audio data, the quality experience module can take the audio data of the previous n seconds at the current time to calculate the average rate, for example, avgLowRate, and take the average rate avgLowRate as the uplink rate.
[0087] It can be understood that using the average rate to determine the network rate at the current time can reflect the overall trend of the network rate in a period of time (the previous n seconds), and the data obtained is more stable and is not affected by the fluctuation of individual unstable data, so that the judgment result is more accurate and the robustness of the algorithm is improved.
[0088] Of course, the uplink rate can also be calculated by other algorithms, for example, the rate of the previous n seconds can be summed, and the application embodiments are not limited.
[0089] After calculating the uplink rate, the quality experience module can execute step S304 to judge the uplink rate.
[0090] S304, whether the uplink rate is lower than the uplink rate threshold.
[0091] If the uplink rate avgLowRate is lower than the uplink rate threshold ulRateThreshold, that is, the uplink rate is less than the uplink rate threshold, it indicates that the current uplink rate is low, and the call quality of the uplink network is poor. Therefore, the quality experience module can execute step S305 to count the number of low rates of the uplink network.
[0092] If the uplink rate avgLowRate is not lower than the uplink rate threshold ulRateThreshold, that is, the uplink rate is greater than or equal to the uplink rate threshold, it indicates that the current uplink rate is high, and the call quality of the uplink network is good. Therefore, the quality experience module can execute step S308 to clear the low rate count of the uplink network.
[0093] S305, the low rate count of the uplink network is incremented by 1.
[0094] When the uplink rate is lower than the uplink rate threshold, it indicates that the call quality of the uplink network is poor. Therefore, the quality experience module can count the number of low rates of the uplink network, increment the low rate count of the uplink network by 1, and execute step S306.
[0095] Optionally, the low rate count of the uplink network can be incremented by 1 based on 0 or incremented by 1 based on a certain number, as long as the counting can be achieved, and the application embodiments are not limited.
[0096] Optionally, the low rate count of the uplink network can be incremented by 1 or incremented by other values, as long as the counting can be achieved, and the application embodiments are not limited.
[0097] S306, whether the low rate count of the uplink network is greater than a first threshold.
[0098] If the low rate count of the uplink network is greater than the first threshold, it indicates that the current application has appeared low rate for many times. Therefore, the quality experience module can execute step S307, otherwise, the quality experience module executes step S309, which indicates that the current application does not appear to be stuck.
[0099] The first threshold can be pre-set by the electronic device according to an empirical value, and the value of the first threshold is not limited in the application embodiments.
[0100] After the end of the current cycle, if the application is in a voice call scenario or a video call scenario, the electronic device can re-execute step S303, otherwise, the electronic device can end the execution Figure 3 Flow of the corresponding embodiment.
[0101] S307, determine that the application is stuck, and set a reason value.
[0102] It can be understood that the quality experience module can set different reason values for different network environments. For example, the corresponding reason value when the network drops can be represented by PS_DROP; the corresponding reason value when the network is slow or the network is stuck can be represented by PS_SLOW. The specific representation of the reason value, and the corresponding relationship between the network environment and the reason value, are not limited by the embodiments of the present application.
[0103] In the embodiments of the present application, when the quality experience module determines that the Android application package (APK) of the current application continuously appears in a low uplink rate during the running process, it can be considered that the application is stuck, and the quality experience module can set the corresponding reason value as QOE_APK_SUBREASON, which is used to indicate that the current application enters a low rate scenario. The reason value QOE_APK_SUBREASON can also be named by other fields, and the naming of the specific field is not limited by the embodiments of the present application.
[0104] The quality experience module can pass the reason value to the fusion self-healing module, and the fusion self-healing module can perform corresponding self-healing actions based on different reason values. The self-healing action can also be understood as a network optimization process, for example, including resetting the network, restarting the application, and the like. Thus, the network quality is improved, and the user experience is improved.
[0105] S308, set the uplink low rate count to 0.
[0106] If the uplink rate is greater than or equal to the uplink rate threshold, it means that the uplink network has good call quality, and the quality experience module can clear the low rate count of the uplink network and execute step S309.
[0107] It can be understood that the low rate count clearing can be that the quality experience module sets the low rate count to 0, or the quality experience module sets the low rate count to a negative number, such as -1, or the quality experience module sets the low rate count to other possible values, and the embodiments of the present application are not limited. As long as the value set by the quality experience module can identify the clearing state of the low rate count.
[0108] S309, determine that the application is not stuck.
[0109] When the quality experience module determines that the current application does not have a stall, the quality experience module can clear the previously set reason value QOE_APK_SUBREASON. In this way, the electronic device can accurately record the state of whether the application has a stall based on the reason value, so as to more reasonably perform network setting and maintenance, and improve the stability of the network.
[0110] In a possible scenario, when the application is in a voice call scenario or a video call scenario, if the user does not speak, the call scenario can be in a silent state, and at this time, the uplink rate detected by the quality experience module can be lower than the uplink rate threshold. For example, when the user does not speak during the call, the uplink rate detected by the quality experience module is 12 kilobytes (kb), and the uplink rate threshold value is 20 kb. If the uplink rate is less than the uplink rate threshold, the quality experience module determines that the current network environment is poor, and the application has a stall, so as to set the reason value. Then, the electronic device can optimize the network.
[0111] That is, the low uplink rate can be caused by poor network environment, or can be caused by the user not speaking and the small amount of audio data transmission. Therefore, if the determination is made only according to whether the uplink rate is less than the uplink rate threshold, the application stall cannot be accurately determined, and whether network rate optimization is needed. Because when the electronic device is in a silent call state in which the user does not speak, the quality experience module can misjudge that the application has a stall. Therefore, for this situation, the embodiments of the present application can increase the determination of whether the call is voiced on the basis of the corresponding embodiments. Figure 3
[0112] Figure 4 A flowchart for detecting application stall based on uplink rate and whether the call is voiced is shown.
[0113] It can be understood that the following steps S403-S408 are related steps for determining the uplink rate, steps S409-S413 are related steps for determining whether the call is voiced, and step S314 is a step of determining whether the application has a stall according to the uplink rate and whether the call is voiced.
[0114] Steps S401-S406 are similar to steps S301-S306 in the corresponding embodiments, and steps S407-S408 are similar to steps S308-S309 in the corresponding embodiments. The execution of each step can be referred to the related description in the corresponding embodiments, and will not be described herein. Figure 3 Steps S407-S408 are similar to steps S308-S309 in the corresponding embodiments. The execution of each step can be referred to the related description in the corresponding embodiments, and will not be described herein. Figure 3 Steps S407-S408 are similar to steps S308-S309 in the corresponding embodiments. The execution of each step can be referred to the related description in the corresponding embodiments, and will not be described herein. Figure 3 Steps S407-S408 are similar to steps S308-S309 in the corresponding embodiments. The execution of each step can be referred to the related description in the corresponding embodiments, and will not be described herein.
[0115] S401, enter the application.
[0116] S402, determine whether the application enters a voice call scenario or a video call scenario.
[0117] If the application enters a voice call scenario or a video call scenario, step S403 can be executed; otherwise, subsequent steps are not continued to be executed.
[0118] S403, take the average rate of the previous n seconds of the uplink network as input to detect the uplink rate.
[0119] After the quality experience module calculates the uplink rate, step S404 can be executed to determine the uplink rate.
[0120] S404, whether the uplink rate is lower than the uplink rate threshold.
[0121] If the uplink rate avgLowRate is lower than the uplink rate threshold ulRateThreshold, the quality experience module can execute step S405 to count the number of low rates of the uplink network.
[0122] If the uplink rate avgLowRate is not lower than the uplink rate threshold ulRateThreshold, the quality experience module can execute step S307 to clear the low rate count of the uplink network.
[0123] S405, the low rate count of the uplink network is incremented by 1.
[0124] When the uplink rate is lower than the uplink rate threshold, the quality experience module can increment the low rate count of the uplink network by 1 and execute step S406.
[0125] S406, whether the low rate count of the uplink network is greater than a first threshold.
[0126] If the low rate count of the uplink network is greater than the first threshold, the quality experience module can execute step S414, otherwise, the quality experience module executes step S408.
[0127] S407, the uplink low rate count is set to 0.
[0128] If the uplink rate is greater than or equal to the uplink rate threshold, the quality experience module can clear the low rate count of the uplink network and execute step S408.
[0129] S408, determine whether the application has no stuttering.
[0130] S409, take the average audio stream of the previous n seconds of the uplink as input to detect whether the uplink call has sound.
[0131] When the application enters a voice call scenario or a video call scenario, the audio decoding module can detect whether the call has sound.
[0132] In one possible implementation, a call detection point for the audio decoding module (VDM module) can be added to the ADSP driver to acquire audio data. For details on how to add a call detection point for the VDM module to the ADSP driver, please refer to [link / reference needed]. Figure 7 The relevant descriptions in the corresponding embodiments will not be repeated here.
[0133] In the ADSP driver, the call detection point of the VDM module can acquire audio data and report it to the VDM module in the Framework layer. For example, the call detection point of the VDM module can report audio data once every 1 second (s), but other time intervals can also be used, such as once every 0.5s or 2s. This application embodiment does not limit the time interval.
[0134] The VDM module can convert audio data into root mean square (RMS) values and store them in an array. These RMS values can represent the loudness of an audio track, and can also be simply referred to as audio loudness.
[0135] like Figure 5 As shown, the VDM module can store the RMS value corresponding to the audio into a circular array at 1-second intervals. Other time intervals, such as 0.5s or 2s, can also be used; this embodiment does not limit the storage. The array size can be, for example, 10, allowing it to store a maximum of 10 RMS values. Of course, the array size can be preset by the electronic device based on empirical values, as long as it can detect the presence of sound during audio calls based on the RMS values in the array. The specific array size is not limited in this embodiment.
[0136] Understandably, taking an array size of 10 as an example, when the VDM module's call detection point reports the audio data at the 11th second, the VDM module can discard the RMS value at the 1st second in the array and store the RMS value corresponding to the audio data at the 11th second in the array. In this way, the RMS value stored in the array can be in a relatively recent state, thus facilitating the VDM module to detect whether there is sound during the call in real time.
[0137] In the embodiments of the present application, it can be obtained through empirical tests that audio data of about 10s is relatively reliable. That is, the audio data of about 10s can detect whether the current audio is voiced. If too much audio data is saved in the array, on the one hand, discontinuity may occur between the audio data, affecting the audio voiced detection result; on the other hand, the audio data will occupy more memory space of the electronic device, and more audio data means that the calculation amount of the electronic device is also large, thereby affecting the running efficiency of the electronic device. If too little audio data is saved in the array, the detection of whether the audio call is voiced is not accurate enough, affecting the audio optimization effect. Therefore, the array size of about 10 in the embodiments of the present application is used as a relatively optimal array size.
[0138] It can be understood that the VDM module can also store a timestamp corresponding to each RMS value when storing the RMS value. In this way, it is convenient for subsequent judgment according to the timestamp whether there is a low rate in the voiced stage of the call, so as to optimize the low rate, improve the call quality and enhance the user experience.
[0139] When detecting whether the uplink call is voiced, the VDM module can take the RMS values of the previous m seconds of the current time in the array for calculation.
[0140] Since the network quality is poor for a period of time, the electronic device can use the audio data of the previous m seconds in the array to judge whether the application is stuck. However, in order to reduce the lag of judgment, m can also take a value less than 10, for example, 3s, indicating that the VDM module takes the RMS values of the previous 3 seconds of the current time in the array to detect whether the call is voiced. It can be understood that the specific value of m can be determined by different business modules according to their own business needs, and m and n can be the same or different, which is not limited in the embodiments of the present application.
[0141] It can be understood that using the average value to determine the audio loudness of the current time can reflect the overall trend of the audio loudness in a period of time (the previous m seconds), so that the data obtained is more stable and is not affected by the fluctuation of individual unstable data, thereby making the judgment result more accurate and improving the robustness of the algorithm.
[0142] Of course, the audio loudness can also be calculated by other algorithms, for example, summing the audio loudness of the previous m seconds, which is not limited in the embodiments of the present application.
[0143] After detecting the uplink audio, the audio decoding module can execute step S410 to judge the uplink call.
[0144] S410, judging whether the uplink call is voiced.
[0145] If the uplink call has sound, it indicates that the current uplink rate should be larger, for example, the uplink rate can be greater than the rate threshold, if the uplink rate is lower than the rate threshold, it can be judged that the application appears to be stuck, then step S411 can be executed to count the number of times of uplink network call with sound.
[0146] If the uplink call has no sound, it indicates that the current uplink rate is small, which can be caused by the user not speaking during the call, and the transmission amount of audio data is small, which should not be judged as application stuck, then step S413 can be executed to clear the uplink network call with sound count.
[0147] S411, the uplink network call with sound count is incremented by 1.
[0148] When the uplink call has sound, the audio decoding module can count the number of times of uplink network call with sound, for example, the uplink network call with sound count ulCallMuteCount is incremented by 1, and step S412 is executed.
[0149] Optionally, the call with sound count ulCallMuteCount can be incremented by 1 based on 0, or can be incremented by 1 based on a certain number, which is not limited in the embodiment of the application.
[0150] Optionally, the call with sound count ulCallMuteCount can be incremented by 1, or can be incremented by other values, as long as it can achieve counting, which is not limited in the embodiment of the application.
[0151] S412, whether the uplink network call with sound count is greater than a second threshold.
[0152] If the uplink network call with sound count ulCallMuteCount is greater than the second threshold, it indicates that the current application is not in the call silent state, then step S414 can be executed to judge whether the application is stuck; otherwise, the current loop ends.
[0153] The second threshold can be pre-set by the electronic device according to the experience value, and the specific value of the second threshold can be understood as that the second threshold and the first threshold can be the same or different, which is not limited in the embodiment of the application.
[0154] After the current loop ends, if the application is in a voice call scene or a video call scene, the electronic device can re-execute step S409, otherwise, the electronic device can end the execution Figure 4 corresponding to the flow of the embodiment.
[0155] S413, the uplink network call with sound count is set to 0.
[0156] If the uplink call is silent, the audio decoding module can clear the uplink network call voice count ulCallMuteCount.
[0157] It can be understood that the audio decoding module can set the call voice count to 0, or set the call voice count to a negative number such as -1, or set the call voice count to other possible values, and the embodiments of the present application are not limited thereto, as long as the set value can identify the clearing state of the call voice count.
[0158] S414, when the low rate count and the call voice count are satisfied at the same time, it is judged that the application is stuck, and the reason value is set.
[0159] In a possible implementation, when the uplink network low rate count is greater than or equal to the first threshold value, and the call voice count is greater than or equal to the second threshold value, the quality experience module can judge that the application is stuck, and then set the reason value. That is, when the uplink network low rate appears continuously, and the call voice is on, the quality experience module can judge that the application is stuck. If the uplink network low rate count is less than the first threshold value, and / or the call voice count is less than the second threshold value, the quality experience module can judge that the application is not stuck.
[0160] Optionally, the case that the uplink network low rate count is equal to the first threshold value can also be judged as not meeting the low rate condition. Similarly, the case that the call voice count is equal to the second threshold value can also be judged as not meeting the call voice condition, which can be set by the electronic device, and the embodiments of the present application are not limited thereto.
[0161] Figure 6 A logical diagram for judging application sticking is shown.
[0162] When the application enters an audio scene or a video scene, the quality experience module can calculate the number of times that the uplink rate is less than the uplink rate threshold. The audio decoding module can calculate the number of times that the call voice is on.
[0163] It can be understood that when the call voice is on, the audio loudness range is usually around -60, the smaller the call voice, the lower the audio loudness, and when the audio loudness reaches -999, it can be judged that the call voice is off.
[0164] Taking the case that the audio loudness and the uplink rate are collected every 1s, and the first threshold value and the second threshold value are both 3 as an example, when the uplink network low rate count is greater than or equal to the first threshold value, and the uplink network call voice count is greater than or equal to the second threshold value, it can also be understood that in 3 seconds, the uplink rate is in a low rate state when the call voice is on. The electronic device can judge that the application is stuck, and then the reason value can be set.
[0165] It should be noted that, Figure 4The execution sequence in the corresponding embodiment is not strictly executed according to the sequence of the steps. For example, the process of detecting the uplink rate and the process of detecting whether the call is voiced can be executed in parallel. It can also be understood that the process of detecting the uplink rate and the process of detecting whether the call is voiced can be executed in two threads respectively, and the process of detecting the uplink rate and the process of detecting whether the call is voiced are not distinguished in the order of execution, and the electronic device can execute the process of detecting the uplink rate first, or execute the process of detecting whether the call is voiced first, or execute the process of detecting the uplink rate and the process of detecting whether the call is voiced in parallel.
[0166] It can be understood that the execution of step S414 needs to be based on the detection result of the uplink rate and the detection result of whether the call is voiced. Therefore, if the electronic device first obtains the detection result of the uplink rate, it also needs to wait for the detection result of whether the call is voiced; if the electronic device first obtains the detection result of whether the call is voiced, it also needs to wait for the detection result of the uplink rate. The electronic device will execute step S414 only after obtaining the detection result of the uplink rate and the detection result of whether the call is voiced. Optionally, in the embodiment of the present application, the uplink rate threshold and the downlink rate threshold can be set to be the same or different, which is not limited in the embodiment of the present application.
[0167] Figure 7 A VDM module call detection point position schematic diagram is shown.
[0168] In the embodiment of the present application, the VDM module call detection point position can be added in the ADSP driver to obtain audio data. The ADSP driver can be used to transmit an audio data stream, and the audio data stream can be transmitted based on a speaker data transmission link or a microphone data transmission link.
[0169] (1) Explanation of the speaker data transmission link.
[0170] Taking the speaker data transmission link as an example, the ADSP driver can include a Down Mailbox module, a video quality manager (VQM) module, a streamRxPP port, a DevicRxPP 3A driver, a deviceRx receiving port, and a speaker driver.
[0171] The Down Mailbox module can be used for communication with a MODEM.
[0172] The VQM module can include a first VQM module (VQM 158a0x1) and a second VQM module (VQM 158a0x2). Wherein, 158a0 can be used to represent an audio dump node, the first VQM module can be used to calibrate the audio quality of the audio data, VQM 158a0x1 can be used to represent that the VQM is placed in the first position. The second VQM module can be used to detect the audio quality and determine whether the calibration of the first VQM module is effective, and VQM 158a0x2 can be used to represent that the VQM is placed in the second position. It can be understood that in some scenarios, the audio quality of the audio data collected by the microphone can be poor, for example, the audio noise is large, etc., and then the VQM module can calibrate the audio quality.
[0173] The streamRxPP port can be used to optimize the audio data according to the echo mark. In a possible call scenario, when the electronic device receives the audio data of the opposite electronic device and parses the audio data, the streamRxPP port can perform echo cancellation optimization processing on the audio data based on the echo mark in the audio data, thereby improving the user experience. Optionally, the electronic device can perform echo cancellation optimization processing, or can not perform echo cancellation optimization processing.
[0174] The deviceRx receiving port can be used to deliver the audio data obtained by the electronic device to the speaker driver, and then the speaker driver can play the audio data through the speaker hardware device to realize the call function.
[0175] The DeviceRxPP 3A driver can convert the audio data into audio data that can be recognized by the VQM module according to a certain protocol specification. In some scenarios, the DeviceRxPP 3A driver can also be understood as the DeviceRxPP 3A port.
[0176] In a possible implementation, the Down Mailbox module can communicate with the modem, receive audio data of the opposite electronic device, and deliver the audio data to the first VQM module. After calibrating the audio data, the first VQM module can deliver the calibrated audio data to the streamRxPP port. After optimizing the audio data according to the echo mark, the streamRxPP port can deliver the audio data to the DeviceRxPP 3A driver. The DeviceRxPP 3A driver can convert the audio data into audio data recognizable by the second VQM module according to a certain protocol specification, and deliver the audio data to the second VQM module. After detecting that the calibration of the first VQM module is effective, the second VQM module can deliver the audio data to the deviceRx receiving port. The deviceRx receiving port can deliver the audio data to the speaker driver, and the speaker driver can play the audio data through the speaker hardware device to realize the call function.
[0177] In the embodiments of the present application, a detection point of the VDM module can also be added in the ADSP driver. For example, the first detection point (VDM0 / 1 dotting) and the second detection point (VDM5 dotting) are included.
[0178] The VDM0, VDM 1, VDM 3, VDM 4 and VDM 5 can all be used to represent the audio dump node, which can be understood as a naming method of the VDM.
[0179] For example, the first detection point can be located between the Down Mailbox module and the first VQM module, and the second detection point can be located between the second VQM module and the deviceRx receiving port. It can be understood that the first detection point and the second detection point can also be arranged at other positions in the ADSP driver. The positions of the first detection point and the second detection point in the ADSP driver can be pre-set by the electronic device, and the embodiments of the present application are not limited.
[0180] In the process of delivering the audio data from the Down Mailbox module to the first VQM module, the first detection point can obtain the audio data. In the process of delivering the audio data from the second VQM module to the deviceRx receiving port, the second detection point can check whether the obtained audio data is consistent with the audio data obtained by the first detection point.
[0181] If the audio data obtained by the first detection point and the second detection point is consistent, it indicates that the audio data does not appear data loss in the transmission process; if the audio data obtained by the first detection point and the second detection point is inconsistent, it indicates that the audio data may appear data loss in the transmission process, and the second detection point can optimize the data, for example, the audio data obtained by the first detection point and the second detection point can be averaged to improve the accuracy of the audio data.
[0182] In the embodiment of the application, the first detection point and the second detection point in the ADSP drive can detect the audio data in real time and report the audio data to the VDM module of the Framework layer. Thus, the VDM module of the Framework layer can detect whether the call has sound based on the audio data.
[0183] (2) Description of the microphone data transmission link.
[0184] Taking the microphone data transmission link as an example, the ADSP drive can include a microphone drive, a deviceTx sending port, a DeviceTxPP 3A drive, a VQM module, a streamTxPP port and an Up Mailbox module.
[0185] The deviceTx sending port can be used to obtain the audio data reported by the microphone drive.
[0186] The DeviceTxPP 3A drive can convert the audio data into audio data that can be recognized by the VQM module according to a certain protocol specification. In some scenarios, the DeviceTxPP 3A drive can also be understood as the DeviceTxPP 3A port.
[0187] The VQM module can include a first VQM module (VQM 158a0×1) and a second VQM module (VQM 158a0×2). For details, refer to the related description in the above (1) description of the speaker data transmission link, which will not be repeated here.
[0188] The streamTxPP port can be used to mark the echo in the audio data with an echo mark. In possible call scenarios, the audio data collected by the microphone can have echo, and the streamTxPP port can mark the audio data with echo. When the opposite electronic device receiving the audio data analyzes the audio data, the echo mark can be used to identify the audio data with echo, and the echo can be optimized by de-echo processing, thereby improving the user experience.
[0189] The Up Mailbox module can be used for communication with the MODEM.
[0190] In a possible implementation, after the microphone driver collects the audio data, the audio data can be transmitted to the deviceTx sending port, which can transmit the audio data to the DeviceTxPP 3A driver. The DeviceTxPP 3A driver can convert the audio data into audio data recognizable by the first VQM module according to a certain protocol specification, and transmit the audio data to the first VQM module. After calibrating the audio data, the first VQM module can transmit the calibrated audio data to the streamTxPP port. After marking the echo in the audio data, the streamTxPP port can transmit the audio data to the second VQM module. After detecting that the calibration of the first VQM module is effective, the second VQM module can transmit the audio data to the Up Mailbox module, so that the Up Mailbox module can communicate with the modem and transmit the audio data.
[0191] In the embodiments of the application, the detection points of the VDM module can also be added in the ADSP driver. For example, the third detection point (VDM4 dotting) and the fourth detection point (VDM3 dotting) are included.
[0192] For example, the third detection point can be located between the deviceTx sending port and the DeviceTxPP 3A driver, and the fourth detection point can be located between the second VQM module and the Up Mailbox module. It can be understood that the third detection point and the fourth detection point can also be arranged at other positions in the ADSP driver. The specific positions of the third detection point and the fourth detection point in the ADSP driver can be set by the electronic device in advance, and the embodiments of the application are not limited.
[0193] In the process of transmitting the audio data from the deviceTx sending port to the DeviceTxPP 3A driver, the third detection point can obtain the audio data. In the process of transmitting the audio data from the second VQM module to the Up Mailbox module, the fourth detection point can check whether the obtained audio data is consistent with the audio data obtained by the third detection point.
[0194] If the audio data obtained by the third detection point and the fourth detection point is consistent, it indicates that there is no data loss in the transmission process of the audio data. If the audio data obtained by the third detection point and the fourth detection point is inconsistent, it indicates that there may be data loss in the transmission process of the audio data. Then, the second detection point can optimize the data, for example, the audio data obtained by the third detection point and the fourth detection point can be averaged to improve the accuracy of the audio data.
[0195] In the embodiment of the present application, the third detection point and the fourth detection point in the ADSP driving can detect the audio data in real time and report the audio data to the VDM module of the Framework layer. Thus, the VDM module of the Framework layer can detect whether the call is voiced based on the audio data.
[0196] The method of the embodiment of the present application is described in detail through specific embodiments. The following embodiments can be combined or implemented independently, and the same or similar concepts or processes can not be described in some embodiments.
[0197] Figure 8 The call quality detection method of the embodiment of the present application is shown. The method comprises:
[0198] S801, the electronic device enters a call.
[0199] In the embodiment of the present application, the call can include a call implemented based on a TCP protocol and / or a UDP protocol, for example, the call can include a voice call or a video call, etc.
[0200] Specifically, the electronic device enters the call can be detected, which can be referred to Figure 3 The related description of step S302 in the corresponding embodiment will not be repeated.
[0201] S802, in a first time period in the call, if a network rate in the call is less than a rate threshold, a first parameter is accumulated by a first value, and the first parameter is used to indicate a number of times that a low rate occurs in the call.
[0202] In the embodiment of the present application, the length of the first time period can be pre-set by the electronic device, for example, the first time period can be the length of time that the call is performed, or the first time period can be a preset period of time, for example, 3s, 10s, etc., and the length of the first time period is not limited in the embodiment of the present application.
[0203] The network rate can include an uplink network rate and a downlink network rate. If the network rate is the uplink network rate, the rate threshold can be understood as the uplink rate threshold ulRateThreshold in the Figure 3 If the network rate is the downlink network rate, the rate threshold can correspond to the downlink rate threshold.
[0204] The first parameter is used to indicate the number of times that the low rate occurs in the call. The first parameter can include the low rate count of the uplink network in the Figure 3 The first parameter can also include the low rate count of the downlink network.
[0205] The first value can include Figure 3In step S305, the first value is 1, and the low-rate count of the uplink network is increased by 1. It can be understood that the first value can also be set to other values, and the embodiments of the present application are not limited thereto.
[0206] It should be understood that if the network rate in the call is less than the rate threshold, it means that the network rate is low and the call quality is poor. Therefore, the number of low rates in the call can be counted, so as to facilitate subsequent determination of whether the call is stuck.
[0207] In step S803, if the audio loudness in the call is greater than the audio loudness threshold in the first time period, the second parameter is added by a second value, and the second parameter is used to indicate the number of times that the call is voiced.
[0208] In the embodiments of the present application, the audio loudness can be represented by the root mean square (RMS) value in the call, which can be used to represent the audio track loudness. The audio loudness can include the audio loudness of the uplink network and the audio loudness of the downlink network.
[0209] Since in the call, if the call is voiced, the audio loudness range can be around -60. It should be understood that the smaller the call voice, the lower the audio loudness. When the audio loudness reaches -999, it can be determined that the call is voiceless. Therefore, in the embodiments of the present application, the audio loudness threshold can be set to -999, and of course it can also be set to other smaller values, and the embodiments of the present application are not limited thereto.
[0210] The second parameter is used to indicate the number of times that the call is voiced. The second parameter can include Figure 4 the call voiced count of the uplink network in step S411, and the second parameter can also include the call voiced count of the downlink network.
[0211] The second value can include Figure 4 In step S411, the call voiced count of the uplink network is increased by 1. It can be understood that the second value can also be set to other values, and the first value and the second value can be the same or different, and the embodiments of the present application are not limited thereto.
[0212] It should be understood that if the audio loudness in the call is greater than the audio loudness threshold, it means that the call is voiced, and therefore the number of times that the call is voiced can be counted, so as to facilitate subsequent determination of whether the call is stuck.
[0213] In step S804, in the first time period, if the first parameter is greater than the first threshold and the second parameter is greater than the second threshold, it is determined that the call is stuck.
[0214] In the embodiments of the present application, the first parameter is greater than the first threshold, that is, the low rate count during the call is greater than the first threshold, which indicates that the low rate has occurred multiple times during the call. The second parameter is greater than the second threshold, that is, the voice count during the call is greater than the second threshold, which indicates that the call is voiced. Therefore, in the case that the call is voiced but the network rate is low, it can be determined that the call has a stutter.
[0215] It should be understood that the first parameter being greater than the first threshold can include that the first parameter is continuously greater than the first threshold within the first time period, or that the first parameter is not continuously greater than the first threshold within the first time period, which is not limited in the embodiments of the present application.
[0216] The call quality detection method provided in the embodiments of the present application can determine that a stutter occurs in the call process when the electronic device detects that the uplink network rate or the downlink network rate is continuously lower than the rate threshold and the call is voiced, and the electronic device can optimize the network to reduce the call stutter and improve the user experience.
[0217] Optionally, in the embodiments of the present application, Figure 8 Based on the corresponding embodiments, the method can further include: determining that the call scene does not have a stutter in the case that the first parameter is less than or equal to the first threshold and / or the second parameter is less than or equal to the second threshold within the first time period.
[0218] In the embodiments of the present application, the first parameter is less than or equal to the first threshold, that is, the low rate count during the call is less than or equal to the first threshold, which indicates that the low rate occurs less frequently during the call, and the call quality is better. The second parameter is less than or equal to the second threshold, that is, the voice count during the call is less than or equal to the second threshold, which indicates that the voiceless call occurs more frequently.
[0219] Since the voiceless call can be caused by the user not speaking during the call, the transmission amount of the audio data during the call is small. Therefore, in the case of the voiceless call and / or the good call quality, it can be determined that the call does not have a stutter. In this way, by comprehensively considering the speed of the network rate and whether the call is voiced, it can be more accurately determined whether the current call has a stutter and whether the network rate optimization needs to be further performed. The situation that the call stutter is misjudged when the voiceless call state occurs due to the user not speaking is reduced.
[0220] Optionally, in the embodiments of the present application, Figure 8 Based on the corresponding embodiments, the method can further include: setting the first parameter to a first initial value if the network rate is greater than the rate threshold within the first time period.
[0221] In the embodiments of the present application, the first initial value can include 0 in step S308, or Figure 3 0 in step S308.Figure 4 The first initial value can also be set to other values as long as the zero clearing state of the first parameter can be identified, and the specific value of the first initial value is not limited in the embodiments of the present application.
[0222] It can be understood that the network rate is greater than the rate threshold, indicating that the network rate of the current call is high, the call quality is good, and no stuttering occurs, so the first parameter does not need to be accumulated and counted again, and therefore the first parameter can be set to the first initial value.
[0223] Optionally, in the Figure 8 Based on the corresponding embodiments, the method can further include: if the audio loudness is less than the audio loudness threshold, setting the second parameter to a second initial value in the first time period.
[0224] In the embodiments of the present application, the second initial value can include Figure 4 The second initial value can also be set to other values as long as the zero clearing state of the second parameter can be identified, and the specific value of the second initial value is not limited in the embodiments of the present application.
[0225] It can be understood that the audio loudness is less than the audio loudness threshold, indicating that the current call is silent, which can be caused by the user not speaking during the call, and cannot be used to determine whether the call has stuttering, so the second parameter does not need to be accumulated and counted again, and the second parameter can be set to the second initial value. Optionally, in the Figure 8 Based on the corresponding embodiments, the network rate at the first time is an average value calculated based on N network rates recorded in the electronic device before the first time, and the first time is within the first time period, and N is a positive integer.
[0226] In the embodiments of the present application, the calculation method of the network rate can refer to Figure 3 The network rate at the first time can be understood as the average rate avgLowRate in the related description of step S302 of the corresponding embodiments, and will not be repeated here.
[0227] It can be understood that using the average rate to determine the network rate at the current time can reflect the overall trend of the network rate in a period of time (N network rates), so that the data obtained is more stable and is not affected by the fluctuation of individual unstable data, thereby making the judgment result more accurate and improving the robustness of the algorithm.
[0228] Optionally, in the Figure 8 Based on the corresponding embodiments, the audio loudness at the second time is an average value calculated based on M root mean square (RMS) values of calls recorded in the electronic device before the second time, and the second time is within the first time period, and M is a positive integer.
[0229] In the embodiments of the present application, the calculation method of the audio loudness can refer to Figure 4 The related description in step S409 of the corresponding embodiments, and the related description in Figure 5 The corresponding embodiments and Figure 6 The related description in the corresponding embodiments, and the related description in
[0230] It can be understood that the audio loudness at the current moment is determined by using the average value of the root mean square RMS values of M calls to reflect the overall trend of the audio loudness in a period of time (the root mean square RMS values of M calls), and the data obtained is more stable and is not affected by the fluctuation of individual unstable data, so that the judgment result is more accurate and the robustness of the algorithm is improved.
[0231] Optionally, in the Figure 8 On the basis of the corresponding embodiments, the electronic device comprises an audio digital signal processor ADSP and a target module, the ADSP comprises a detection point of the target module, and the target module is configured to obtain the root mean square RMS values of M calls from the calls based on the detection point and record the root mean square RMS values of M calls.
[0232] In the embodiments of the present application, the target module can be understood as an audio decoding module (VDM module), and the implementation method of obtaining the root mean square RMS value based on the detection point can refer to Figure 3 The related description in step S409 of the corresponding embodiments, and the related description in Figure 5 The corresponding embodiments and Figure 7 The related description in the corresponding embodiments, and the related description in
[0233] It can be understood that the target module can obtain the root mean square RMS values of the adjacent M calls at the current moment in time based on the detection point, so that the more recent RMS values can be obtained, thereby facilitating the real-time detection of whether the call is voiced, obtaining the RMS values more in line with the current moment, and making the judgment more accurate.
[0234] Optionally, in the Figure 8 On the basis of the corresponding embodiments, after determining that the call scene is stuck, the method can further comprise: setting a target value as a third value, the third value being used to instruct the electronic device to perform network optimization processing.
[0235] In the embodiments of the present application, the target value can be understood as Figure 3 The cause value in the corresponding embodiments, and the third value can be understood as QOE_APK_SUBREASON, which can be used to indicate that the current call enters a low-rate scene and can be used to instruct the electronic device to perform network optimization processing.
[0236] The network optimization processing can include resetting the network, restarting the application, and the like, and the specific electronic device can perform corresponding network optimization processing according to actual conditions, which is not limited in the embodiments of the present application.
[0237] By setting the target value as the third value, the network optimization processing is performed, so that the network quality can be improved and the user experience can be improved.
[0238] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to select authorization or refusal.
[0239] The above mainly introduces the scheme provided by the embodiments of the present application from the perspective of method. In order to realize the above functions, it contains the hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that the method steps of each example described in connection with the embodiments disclosed in the present application can be realized in the form of hardware or combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0240] The embodiments of the present application can divide the functions of the device implementing the method according to the above method examples, for example, each function module can be divided according to each function, or two or more functions can be integrated in one processing module. The integrated module can be realized in the form of hardware or software function module. It should be noted that the division of modules in the embodiments of the present application is illustrative, and is only a logical function division. Actual implementation can have another division method.
[0241] As Figure 9 The chip 900 includes one or more (including two) processors 901, a communication line 902, a communication interface 903 and a memory 904.
[0242] In some embodiments, the memory 904 stores the following elements: executable modules or data structures, or a subset thereof, or an extended set thereof.
[0243] The method described in the embodiments of the present application can be applied to the processor 901 or implemented by the processor 901. The processor 901 can be an integrated circuit chip having a signal processing capability. In the implementation process, each step of the above method can be completed by the integrated logic circuit or the instruction in the software form of the hardware in the processor 901. The processor 901 described above can be a general processor (for example, a microprocessor or a conventional processor), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, a discrete gate or transistor logic device, or a discrete hardware component. The processor 901 can implement or execute the disclosed processing-related methods, steps and logic block diagrams in the embodiments of the present application.
[0244] The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. Among them, the software module can be located in a mature storage medium in the field, such as a random access memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable read-only memory (EEPROM). The storage medium is located in the memory 904, and the processor 901 reads the information in the memory 904 and combines the hardware to complete the steps of the above method.
[0245] The processor 901, the memory 904 and the communication interface 903 can communicate through the communication line 902.
[0246] In the above embodiments, the instructions stored in the memory for the processor to execute can be implemented in the form of a computer program product. Among them, the computer program product can be written in the memory in advance, or downloaded and installed in the memory in the form of software.
[0247] The embodiments of the present application further provide a computer program product including one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from a website, a computer, a server or a data center of one site to a website, a computer, a server or a data center of another site through a wired (such as a coaxial cable, an optical fiber, a digital subscriber line (DSL) or a wireless (such as infrared, wireless, microwave, etc.)) mode. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, a data center, etc. including one or more available media sets. For example, the available media can include magnetic media (such as a floppy disk, a hard disk or a magnetic tape), optical media (such as a digital versatile disc (DVD)), or semiconductor media (such as a solid state disk (SSD)) and the like.
[0248] The embodiments of the present application further provide a computer-readable storage medium. The methods described in the above embodiments can be implemented in whole or in part by software, hardware, firmware or any combination thereof. The computer-readable medium can include a computer storage medium and a communication medium, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium accessible by a computer.
[0249] As a possible design, the computer-readable medium can include a compact disc read-only memory (CD-ROM), RAM, ROM, EEPROM or other optical disk storage; the computer-readable medium can include a magnetic disk storage or other magnetic disk storage device. Moreover, any connection line can also be appropriately referred to as a computer-readable medium. For example, if software is transmitted from a website, a server or other remote source using a coaxial cable, an optical fiber cable, a twisted pair, a DSL or a wireless technology (such as infrared, radio and microwave), the coaxial cable, the optical fiber cable, the twisted pair, the DSL or the wireless technology such as infrared, radio and microwave are included in the definition of the medium. As used herein, the disk and the optical disk include a compact disc (CD), a laser disc, an optical disc, a digital versatile disc (DVD), a floppy disk and a Blu-ray disc, wherein the disk is usually reproduced in a magnetic manner, and the optical disk is optically reproduced by laser.
[0250] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as a combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices generate a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application are described. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as a combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices generate a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application are described. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as a combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices to
Claims
1. A method of detecting voice quality, characterized by, The method comprises: The electronic device enters into a call; In a first time period in the call, if a network rate in the call is less than a rate threshold, a first parameter is accumulated by a first value, the first parameter is used to indicate a number of times that a low rate occurs in the call; In the first time period, if an audio loudness in the call is greater than an audio loudness threshold, a second parameter is accumulated by a second value, the second parameter is used to indicate a number of times that the call is voiced; In the first time period, if the first parameter is greater than a first threshold and the second parameter is greater than a second threshold, it is determined that the call scene has a stutter.
2. The method of claim 1, wherein, The method further comprises: In the first time period, if the first parameter is less than or equal to the first threshold and / or the second parameter is less than or equal to the second threshold, it is determined that the call scene does not have a stutter.
3. The method according to claim 1 or 2, characterized in that, The method further comprises: In the first time period, if the network rate is greater than the rate threshold, the first parameter is set to a first initial value.
4. The method according to any one of claims 1 to 3, characterized in that, The method further comprises: In the first time period, if the audio loudness is less than the audio loudness threshold, the second parameter is set to a second initial value.
5. The method according to any one of claims 1 to 4, characterized in that, A network rate at a first time is an average value calculated based on N network rates recorded in the electronic device before the first time, the first time is in the first time period, and N is a positive integer.
6. The method according to any one of claims 1 to 5, characterized in that, An audio loudness at a second time is an average value calculated based on M root mean square (RMS) values of the call recorded in the electronic device before the second time, the second time is in the first time period, and M is a positive integer.
7. The method of claim 6, wherein, The electronic device comprises an audio digital signal processor (ADSP) and a target module, the ADSP comprises a detection point of the target module, and the target module is used to acquire the M root mean square (RMS) values of the call from the call based on the detection point and record the M root mean square (RMS) values of the call.
8. The method according to any one of claims 1 to 7, characterized in that, After the call scene is determined to have a stutter, the method further comprises: A target value is set to a third value, the third value is used to indicate that the electronic device performs network optimization processing.
9. An electronic device, comprising: The electronic device comprises one or more processors and a memory; The memory is coupled to the one or more processors, and the memory is used to store computer program codes, the computer program codes comprise computer instructions, and the one or more processors invoke the computer instructions to enable the electronic device to perform the method in any one of claims 1-8.
10. A chip system, characterized by The chip system is applied to an electronic device, and the chip system comprises one or more processors, and the one or more processors are used to invoke computer instructions to enable the electronic device to perform the method in any one of claims 1-8.
11. A computer readable storage medium, characterized in that, The computer readable storage medium comprises computer instructions, and when the computer instructions run on an electronic device, the computer instructions enable the electronic device to perform the method in any one of claims 1-8.
12. A computer program product, characterised in that, The computer program product comprises computer program code which, when run on an electronic device, causes the electronic device to perform the method according to any one of claims 1-8.
Citation Information
Patent Citations
Audio playing method, medium, device and computing equipment
CN109524024A
Video live broadcast switching method and device, equipment and storage medium
CN114025204A
Network quality detection method and device
CN116232959A
Method and device for optimizing traffic call lagging, mobile terminal equipment and medium
CN117376945A
Data transmission rate control method and system, and user equipment
WO2021203829A1