VOIP evaluation system, method and device
By using a VoIP evaluation system and methodology, existing voice apps are used to simulate VoIP voice transmission and reception. The voice quality is evaluated using PESQ and ESOI metrics. This solves the problem of evaluating the audio electrical and acoustic metrics of VoIP voice software terminal devices, enables effective evaluation of microphones and speakers, and reduces costs.
Patent Information
- Application Number
- CN202411232167.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2026-03-10
AI Technical Summary
Existing VoIP software lacks high reliability in terminal device evaluation systems, leading to higher requirements for terminal devices with VoIP software installed. However, existing technologies are insufficient to effectively evaluate their audio electrical and acoustic properties, especially the performance of microphones and speakers.
This invention provides a VoIP evaluation system and method that simulates VoIP voice transmission and reception by testing the IP network communication between the host and the device under test, uses existing voice apps for sound pickup and playback evaluation, and employs PESQ and ESOI metrics to assess voice quality, thereby achieving hardware evaluation of the microphone and speaker and software evaluation of the sound pickup and playback process.
It enables effective evaluation of microphones and speakers of VoIP voice software terminal devices, reduces costs, adapts to different testing scenarios, and improves the reliability and efficiency of the evaluation system.
Smart Images

Figure CN121641073A_ABST
Abstract
Description
Technical Field
[0001] This application relates to evaluation technology for Voice over Internet Protocol (VOIP), and more particularly to a VOIP evaluation system, method and apparatus. Background Technology
[0002] With technological advancements, more and more people prefer using online chat tools for voice chat. This voice communication bypasses traditional telephone networks and is transmitted via the internet. VoIP technology converts voice into Internet Protocol (IP) data packets for transmission over IP networks. Its basic principle is as follows: the sending end digitizes the analog voice signal to obtain voice data, compresses the voice data using a compression algorithm to obtain a bitstream, and then packages the bitstream according to the Transmission Control Protocol / Internet Protocol (TCP / IP) standard to obtain IP data packets. These IP data packets are then sent to the receiving end via the IP network. The receiving end decapsulates and reassembles these IP data packets according to the TCP / IP standard to obtain the bitstream. After decompression, the voice data is obtained, and the analog voice signal is recovered from the voice data, thus achieving the goal of transmitting voice over the internet.
[0003] VoIP voice software (such as WeChat, QQ, and carrier camera applications (APPs) such as China Mobile's Hejiaqin and China Telecom's Tianyi Guanjia) is widely used among the population. Many businesses and brands have begun to focus their marketing efforts on this platform, which has led to increasingly higher requirements for terminal devices that have VoIP voice software installed. Therefore, a highly reliable VoIP evaluation system is of great significance and value in the implementation of VoIP technology. Summary of the Invention
[0004] This application provides a VoIP evaluation system, method, and apparatus to evaluate the hardware and software of the device under test, and to enable voice communication using existing voice apps, thereby reducing costs.
[0005] In a first aspect, this application provides a VoIP testing system, comprising: a test host and a device under test; wherein the test host is equipped with a voice application (APP), and the device under test includes a microphone and a speaker; the voice application and the device under test communicate via an Internet Protocol (IP) network.
[0006] In one possible implementation, the microphone picks up a first sound to obtain a first comparison audio, which is emitted by playing a first original audio; the device under test transmits the first comparison audio to a voice app; and the test host compares the first comparison audio with the first original audio to obtain the sound pickup evaluation result of the device under test.
[0007] The audio pickup evaluation results correspond to the audio pickup evaluation process when the device under test (DUT) acts as the transmitter (also known as the uplink evaluation process), that is,
[0008] The device under test simulates a voice transmitter in VoIP, and records the first sound (including human voice and possibly ambient sounds such as various noises) through a microphone to obtain the first comparison audio. The first comparison audio is digitized from an analog signal to obtain voice data. The voice data is then compressed using a voice compression algorithm to obtain a bitstream. The bitstream is then packaged according to the TCP / IP standard to obtain IP data packets. The IP data packets are then transmitted to the test host 301 via the IP network.
[0009] At this point, the test host simulates the voice receiver in VoIP. The voice APP receives IP data packets, and then the test host decapsulates and concatenates these IP data packets according to the TCP / IP standard to obtain the bitstream. After decompressing the bitstream, the voice data is obtained. Then, the voice data is digitally converted into analog signals to recover the first comparison audio.
[0010] The test host compares the first comparison audio with the first original audio (in the evaluation scenario, the first original audio can be acquired by the test host, selected by the test host from a pre-stored audio library, or generated by the test host through audio generation software, without specific limitations) to obtain the sound pickup evaluation result of the device under test. The evaluation result may include the perceptual evaluation of speech quality (PESQ) index and the extended short-time objective intelligence (ESTOI) index. The acquisition method can refer to the acquisition method of the corresponding index, without specific limitations.
[0011] In one possible implementation, the voice app transmits the second original audio to the device under test; the speaker plays the second original audio to emit a second sound; the test host records the second sound to obtain a second comparison audio; the test host compares the second comparison audio with the second original audio to obtain the sound playback evaluation result of the device under test.
[0012] The playback evaluation results correspond to the playback evaluation process when the device under test is used as the receiving end (also known as the downlink evaluation process), that is,
[0013] The test host simulates the voice transmitter in VoIP. It digitizes the second raw audio (including human voice; in the evaluation scenario, the second raw audio can be acquired by the test host, selected by the test host from a pre-stored audio library, or generated by the test host using audio generation software, without specific limitations) to obtain voice data. Then, it compresses the voice data using a voice compression algorithm to obtain a bitstream. The bitstream is then packaged according to the TCP / IP standard to obtain IP data packets. The IP data packets are then transmitted to the device under test via the IP network through the voice APP.
[0014] At this time, the device under test simulates the voice receiver in VoIP, receives IP data packets, and then the device under test decapsulates and concatenates these IP data packets according to the TCP / IP standard to obtain the bitstream. After decompressing the bitstream, the voice data is obtained. Then, the voice data is digitally converted into analog signals to recover the second original audio. The speaker plays the second comparison audio to produce the second sound.
[0015] To obtain the evaluation results, the second sound played out needs to be compared with the original sound. Therefore, a standard microphone (which can reduce noise interference and recover as much sound as possible from the speaker) is used to record the second sound to obtain the second comparison audio. The test host compares the second comparison audio with the second original audio to obtain the sound playback evaluation results of the device under test. These evaluation results may include PESQ and ESOI indicators, and the methods for obtaining them can refer to the methods for obtaining the corresponding indicators, without specific limitations.
[0016] This application embodiment provides a VoIP evaluation system that can perform hardware evaluation of the microphone and speaker of the device under test, as well as software evaluation of the sound pickup and playback processes. Furthermore, it can utilize existing voice apps to achieve voice communication, thereby reducing costs.
[0017] The first and second original audio recordings mentioned above each include human voices, which can be the sounds of a person simulating a dialogue scene.
[0018] Optionally, the first raw audio may also include noise, including steady noise (e.g., wind noise, white noise, etc.) and / or sudden noise (e.g., the sound of a door opening, the sound of a cup breaking, etc.).
[0019] Throughout the evaluation process, IP data packets were transmitted via a voice app. This voice app does not require separate deployment or design; existing VoIP voice software (such as WeChat, QQ, or carrier camera applications (e.g., China Mobile's Hejiaqin, China Telecom's Tianyi Guanjia)) can be used. This facilitates co-deployment with the test host. Even if the voice app's supported system is incompatible with the test host's operating system, it can be deployed on the test host using a virtual machine. Furthermore, using an existing voice app reduces physical cabling, thereby lowering costs.
[0020] In the audio pickup scenario evaluation, the microphone records the first sound, allowing for the evaluation of the microphone's frequency response and loudness. In the audio playback scenario evaluation, the speaker plays the second comparison audio, allowing for the evaluation of the speaker's frequency response and loudness. Therefore, the VoIP evaluation system of this application can also effectively evaluate hardware.
[0021] Optionally, the first sound can be emitted by playing the first original audio through an artificial mouth, which can make the first sound closer to a human voice to better simulate a call scenario.
[0022] Optionally, the second comparison audio can be obtained by recording a second sound using a standard microphone. This can reduce noise interference, recover the sound played by the speaker as much as possible, and avoid interference from other factors on the evaluation results.
[0023] In one possible implementation, the test host may further include: audio evaluation software, audio acquisition software, and audio driver software; wherein, the audio acquisition software is used to acquire or generate the aforementioned first raw audio (e.g., s0 (human voice) or s0 plus n0 (noise)) or second raw audio (e.g., x0); the audio driver software is used to mix different audios and provide an audio transmission channel; the audio evaluation software is used to obtain the pickup or call evaluation results of the device under test or to obtain the playback evaluation results of the device under test. Furthermore, the VoIP evaluation system also includes: a sound card, used to transmit the first raw audio or the second comparison audio.
[0024] In one possible implementation, the device under test includes an Internet of Things (IoT) camera or a smart screen.
[0025] Secondly, this application provides a VoIP evaluation method, which is applied to the VoIP evaluation system described in any one of the first aspects above. The method includes: recording a first sound through a microphone to obtain a first comparison audio, wherein the first sound is emitted by playing a first original audio; and comparing the first comparison audio with the first original audio to obtain a sound pickup evaluation result of the device under test.
[0026] This application embodiment provides a VoIP evaluation method that can achieve both hardware evaluation of the microphone of the device under test and software evaluation of the sound pickup process.
[0027] As described above, the first sound can be a first original audio signal played back (e.g., played back via an artificial mouth). This first original audio signal may include a human voice, and optionally, it may also include noise, including steady-state noise and / or burst noise, to accommodate different test cases.
[0028] The sound pickup evaluation results correspond to the sound pickup evaluation process when the device under test is the transmitter, that is,
[0029] The device under test simulates a voice transmitter in VoIP, recording a first sound (including human voice and possibly ambient sounds such as various noises) through a microphone to obtain a first comparison audio. The first comparison audio is then digitized from an analog signal to obtain voice data. The voice data is then compressed using a voice compression algorithm to obtain a bitstream. The bitstream is then packaged according to the TCP / IP standard to obtain IP data packets. The IP data packets are then transmitted to the test host via the IP network.
[0030] At this point, the test host simulates the voice receiver in VoIP, receives IP data packets, and then decapsulates and concatenates these IP data packets according to the TCP / IP standard to obtain the bitstream. After decompressing the bitstream, the voice data is obtained, and then the voice data is digitally converted into analog signals to recover the first comparison audio.
[0031] The test host compares the first comparison audio with the first original audio (in the evaluation scenario, the first original audio can be acquired by the test host, selected by the test host from a pre-stored audio library, or generated by the test host through audio generation software, without specific limitations) to obtain the sound pickup evaluation result of the device under test. The evaluation result may include the objective speech quality assessment index PESQ and the ESOI index. The acquisition method can refer to the acquisition method of the corresponding index, without specific limitations.
[0032] In the evaluation of sound pickup scenarios, the microphone's frequency response and loudness can be evaluated by recording the first sound. Therefore, the VoIP evaluation method of this application can also achieve effective hardware evaluation.
[0033] Thirdly, this application provides a VoIP evaluation method, which is applied to the VoIP evaluation system described in any one of the first aspects above. The method includes: playing a second original audio through a speaker to emit a second sound; recording the second sound to obtain a second comparison audio; and comparing the second comparison audio with the second original audio to obtain a sound playback evaluation result of the device under test.
[0034] This application embodiment provides a VoIP evaluation method that can achieve both hardware evaluation of the microphone and speaker of the device under test, and software evaluation of the sound pickup and playback processes.
[0035] As mentioned above, the second original audio includes human voice. The second comparison audio can be obtained by recording the second sound using a standard microphone.
[0036] The playback evaluation results correspond to the playback evaluation process when the device under test is used as the receiving end, that is,
[0037] The test host simulates the voice transmitter in VoIP. It digitizes the second raw audio (including human voice; in the evaluation scenario, the second raw audio can be acquired by the test host, selected by the test host from a pre-stored audio library, or generated by the test host using audio generation software, without specific limitations) to obtain voice data. Then, it compresses the voice data using a voice compression algorithm to obtain a bitstream. The bitstream is then packaged according to the TCP / IP standard to obtain IP data packets. The IP data packets are then transmitted to the device under test via the IP network.
[0038] At this time, the device under test simulates the voice receiver in VoIP, receives IP data packets, and then the device under test decapsulates and concatenates these IP data packets according to the TCP / IP standard to obtain the bitstream. After decompressing the bitstream, the voice data is obtained. Then, the voice data is digitally converted into analog signals to recover the second original audio. The speaker plays the second comparison audio to produce the second sound.
[0039] To obtain the evaluation results, the second sound played out needs to be compared with the original sound. Therefore, a standard microphone (which can reduce noise interference and recover as much sound as possible from the speaker) is used to record the second sound to obtain the second comparison audio. The test host compares the second comparison audio with the second original audio to obtain the sound playback evaluation results of the device under test. These evaluation results may include PESQ and ESOI indicators, and the methods for obtaining them can refer to the methods for obtaining the corresponding indicators, without specific limitations.
[0040] During the evaluation of the sound playback scenario, the second comparison audio is played by the speaker, which allows for the evaluation of the speaker's frequency response and loudness. It is evident that the VoIP evaluation method of this application can also achieve effective hardware evaluation.
[0041] Fourthly, this application provides a VoIP evaluation device, comprising: a recording module for recording a first sound through a microphone to obtain a first comparison audio, wherein the first sound is emitted by playing a first original audio; and an evaluation module for comparing the first comparison audio with the first original audio to obtain a sound pickup evaluation result of the device under test.
[0042] In one possible implementation, the first original audio includes human voice.
[0043] In one possible implementation, the first original audio also includes noise, which includes stationary noise and / or burst noise.
[0044] In one possible implementation, the first sound is emitted by playing the first original audio through an artificial mouth.
[0045] In one possible implementation, the system further includes: a playback module for playing a second original audio through a speaker to emit a second sound; the recording module for recording the second sound to obtain a second comparison audio; and the evaluation module for comparing the second comparison audio with the second original audio to obtain a sound playback evaluation result of the device under test.
[0046] In one possible implementation, the second original audio includes human voice.
[0047] In one possible implementation, the second comparison audio is obtained by recording the second sound using a standard microphone.
[0048] Fifthly, this application provides a terminal device, comprising: one or more processors; a memory for storing one or more programs; and when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the method as described in any one of the second to third aspects above.
[0049] Sixthly, this application provides a computer-readable storage medium including a computer program that, when executed on a computer, causes the computer to perform the method described in any one of the second to third aspects above.
[0050] In a seventh aspect, this application provides a computer program product, characterized in that the computer program product includes computer program code, which, when run on a computer, causes the computer to perform the method described in any one of the second to third aspects above. Attached Figure Description
[0051] Figure 1 This is a schematic diagram of the structure of the terminal device 100 according to an embodiment of this application;
[0052] Figure 2 This is a software structure block diagram of the terminal device 100 according to an embodiment of this application;
[0053] Figure 3 This is a schematic diagram of the structure of the VoIP evaluation system 300 according to an embodiment of this application;
[0054] Figure 4 This is a schematic diagram of the structure of the VoIP evaluation system according to an embodiment of this application;
[0055] Figure 5 This is a schematic diagram of the audio test of the IoT camera according to an embodiment of this application;
[0056] Figure 6 This is a schematic diagram of a multi-track audio file.
[0057] Figure 7 This is a schematic diagram of virtual lines;
[0058] Figure 8 This is an illustration of the Hejiaqin app;
[0059] Figure 9 A schematic diagram of 0-20kHz sine wave speech;
[0060] Figure 10 This is a schematic diagram of audio testing on a large screen according to an embodiment of this application;
[0061] Figure 11 A flowchart of process 1100 of the VoIP evaluation method provided in the embodiments of this application;
[0062] Figure 12 A flowchart of process 1200 of the VoIP evaluation method provided in the embodiments of this application;
[0063] Figure 13 This is a schematic diagram of the structure of the VOIP evaluation device 1300 of this application. Detailed Implementation
[0064] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0065] The terms "first," "second," etc., used in the specification, embodiments, claims, and drawings of this application are for distinguishing purposes only and should not be construed as indicating or implying relative importance or order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0066] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0067] In VoIP technology, the transmitting end digitizes the analog voice signal to obtain voice data. This voice data is then compressed using a voice compression algorithm to obtain a bitstream. The bitstream is then packaged according to TCP / IP standards to obtain IP data packets. These IP data packets are sent to the receiving end via the IP network. The receiving end decapsulates and concatenates these IP data packets according to TCP / IP standards to obtain the bitstream. After decompression of the bitstream, the voice data is obtained. The voice data is then used to reconstruct the analog voice signal, thus achieving the purpose of transmitting voice over the internet. VoIP voice software based on the aforementioned technology (e.g., WeChat, QQ, and operator camera apps (e.g., China Mobile's Hejiaqin, China Telecom's Tianyi Guanjia)) does not incur costs themselves; instead, it is charged based on the data usage for voice transmission, which is then collected by the network operator. VoIP voice software is widely used, and many businesses and brands are increasingly focusing their marketing efforts on this platform. Consequently, the requirements for terminal devices with VoIP voice software installed are gradually increasing. Therefore, a highly reliable VoIP evaluation system is needed to ensure the implementation of VoIP technology.
[0068] Based on this, this application provides a VoIP evaluation system. Before introducing the technical solution of this application, the application scenario of this application will be described first. This application can be applied to the evaluation of audio electrical and acoustic indicators of terminal devices with VoIP voice software installed, the terminal devices including microphones (MICs) and speakers. For example, in terms of hardware, the technical solution of this application can realize the evaluation of the frequency response (referred to as frequency response) of the MIC, and / or the frequency response and loudness (also referred to as volume) of the speaker; in terms of software, the technical solution of this application can realize the evaluation of uplink / downlink voice quality, as well as the evaluation of echo cancellation in single-talk or two-talk voice communication. In addition, the technical solution of this application can also realize other evaluations of audio electrical and acoustic indicators of terminal devices with VoIP voice software installed, which are not specifically limited.
[0069] Figure 1 This is a schematic diagram of the structure of the terminal device 100 according to an embodiment of this application. It should be understood that... Figure 1 The terminal device 100 shown is merely an example, and the terminal device 100 may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations. Figure 1 The various components shown can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.
[0070] Terminal device 100 may include: processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0071] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0072] The controller can serve as the central nervous system and command center of the terminal device 100. The controller can generate operation control signals based on the instruction opcode and timing signals to control the fetching and execution of instructions.
[0073] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0074] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0075] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C buses. The processor 110 can couple to the touch sensor 180K, charger, flash, camera 193, etc., through different I2C bus interfaces. For example, the processor 110 can couple to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface, thereby realizing the touch function of the terminal device 100.
[0076] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.
[0077] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0078] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface to enable music playback through Bluetooth headphones.
[0079] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to enable the shooting function of the terminal device 100. The processor 110 and the display screen 194 communicate via the DSI interface to enable the display function of the terminal device 100.
[0080] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, a wireless communication module 160, an audio module 170, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0081] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge terminal device 100, and can also be used for data transfer between terminal device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other terminal devices, such as AR devices.
[0082] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the terminal device 100. In other embodiments of this application, the terminal device 100 may also adopt different interface connection methods or a combination of multiple interface connection methods as described in the above embodiments.
[0083] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the terminal device 100. While charging the battery 142, the charging management module 140 can also supply power to the terminal device via the power management module 141.
[0084] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, external memory, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.
[0085] The wireless communication function of the terminal device 100 can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0086] Antennas 1 and 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.
[0087] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the terminal device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.
[0088] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.
[0089] The wireless communication module 160 can provide solutions for wireless communication applications on the terminal device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0090] In some embodiments, antenna 1 of terminal device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling terminal device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).
[0091] Terminal device 100 implements display functions through a GPU, display screen 194, and application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0092] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, terminal device 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0093] Terminal device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0094] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0095] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the terminal device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0096] A digital signal processor (DSP) is used to process digital signals. Besides digital image signals, it can also process other digital signals. For example, when terminal device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.
[0097] Video codecs are used to compress or decompress digital video. Terminal device 100 may support one or more video codecs. Thus, terminal device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.
[0098] NPU stands for Neural Network (NN) Computing Processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in terminal devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0099] The external storage interface 120 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the terminal device 100. The external storage card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.
[0100] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of terminal device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of terminal device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0101] Terminal device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0102] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0103] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The terminal device 100 can listen to music or make hands-free calls through the speaker 170A.
[0104] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the terminal device 100 answers a phone call or voice message, the receiver 170B can be brought close to the listener's ear to hear the voice.
[0105] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Terminal device 100 may be equipped with at least one microphone 170C. In some embodiments, terminal device 100 may be equipped with two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, terminal device 100 may be equipped with three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.
[0106] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.
[0107] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be disposed on display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Terminal device 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 194, terminal device 100 detects the intensity of the touch operation based on pressure sensor 180A. Terminal device 100 can also calculate the touch position based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example: when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.
[0108] The gyroscope sensor 180B can be used to determine the motion attitude of the terminal device 100. In some embodiments, the gyroscope sensor 180B can determine the angular velocity of the terminal device 100 around three axes (i.e., the x, y, and z axes). The gyroscope sensor 180B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the terminal device 100's shake, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the terminal device 100 through reverse movement, thus achieving image stabilization. The gyroscope sensor 180B can also be used in navigation and motion-sensing game scenarios.
[0109] The barometric pressure sensor 180C is used to measure air pressure. In some embodiments, the terminal device 100 calculates altitude using the air pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.
[0110] The magnetic sensor 180D includes a Hall sensor. The terminal device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip cover. In some embodiments, when the terminal device 100 is a flip phone, the terminal device 100 can detect the opening and closing of the flip cover using the magnetic sensor 180D. Then, based on the detected opening and closing state of the cover or the flip cover, features such as automatic flip unlocking can be set.
[0111] The 180E accelerometer can detect the magnitude of acceleration of the terminal device 100 in various directions (typically three axes). When the terminal device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the attitude of the terminal device, and can be applied to applications such as landscape / portrait switching and pedometers.
[0112] A distance sensor 180F is used to measure distance. The terminal device 100 can measure distance via infrared or laser. In some embodiments, during a shooting scene, the terminal device 100 can utilize the distance sensor 180F to measure distance for rapid focusing.
[0113] The proximity sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The terminal device 100 emits infrared light outward through the LED. The terminal device 100 uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the terminal device 100. When insufficient reflected light is detected, the terminal device 100 can determine that there is no object near the terminal device 100. The terminal device 100 may use the proximity sensor 180G to detect when a user holds the terminal device 100 close to their ear for a call, so as to automatically turn off the screen to save power. The proximity sensor 180G can also be used in holster mode and pocket mode for automatic unlocking and screen locking.
[0114] The ambient light sensor 180L is used to sense the ambient light intensity. The terminal device 100 can adaptively adjust the brightness of the display screen 194 based on the sensed ambient light intensity. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also work with the proximity sensor 180G to detect whether the terminal device 100 is in a pocket to prevent accidental touches.
[0115] The fingerprint sensor 180H is used to collect fingerprints. The terminal device 100 can use the collected fingerprint characteristics to achieve fingerprint unlocking, accessing application locks, taking photos with fingerprints, answering calls with fingerprints, etc.
[0116] Temperature sensor 180J is used to detect temperature. In some embodiments, terminal device 100 uses the temperature detected by temperature sensor 180J to execute a temperature handling strategy. For example, when the temperature reported by temperature sensor 180J exceeds a threshold, terminal device 100 reduces the performance of the processor located near temperature sensor 180J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is below another threshold, terminal device 100 heats battery 142 to prevent abnormal shutdown of terminal device 100 due to low temperature. In still other embodiments, when the temperature is below yet another threshold, terminal device 100 boosts the output voltage of battery 142 to prevent abnormal shutdown due to low temperature.
[0117] Touch sensor 180K, also known as a "touch panel," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of terminal device 100, in a different position than display screen 194.
[0118] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire vibration signals from the vibrating bone segments of the human vocal cords. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure signals. In some embodiments, the bone conduction sensor 180M can also be incorporated into headphones to form bone conduction headphones. The audio module 170 can parse the voice signals from the vibrating bone segments of the vocal cords acquired by the bone conduction sensor 180M to realize voice functionality. The application processor can parse heart rate information from the blood pressure signals acquired by the bone conduction sensor 180M to realize heart rate detection functionality.
[0119] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Terminal device 100 can receive button input and generate key signal inputs related to user settings and function control of terminal device 100.
[0120] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to touch operations performed on different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations performed on different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.
[0121] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.
[0122] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the terminal device 100. The terminal device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The terminal device 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the terminal device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the terminal device 100 and cannot be separated from the terminal device 100.
[0123] The software system of terminal device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of terminal device 100.
[0124] Figure 2 This is a software structure block diagram of the terminal device 100 according to an embodiment of this application.
[0125] The layered architecture of the terminal device 100 divides the software into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0126] The application layer can include a series of application packages.
[0127] like Figure 2 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, SMS, and voice apps (such as the VoIP voice software mentioned above (e.g., WeChat, QQ, operator camera apps (e.g., China Mobile's Hejiaqin, China Telecom's Tianyi Guanjia, etc.)).
[0128] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0129] like Figure 2 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0130] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0131] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0132] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0133] The phone manager is used to provide communication functions for terminal device 100. For example, it manages call status (including connection, hang-up, etc.).
[0134] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0135] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of download completion or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating the device, and flashing indicator lights.
[0136] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.
[0137] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0138] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0139] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0140] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0141] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0142] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0143] A 2D graphics engine is a graphics engine for 2D drawing.
[0144] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0145] The aforementioned terminal device 100 can be a smart terminal such as a mobile phone, smart screen, or tablet, or it can be an IoT device such as an Internet of Things (IoT) camera; this application does not specifically limit it in this regard. Based on this, the terminal device 100 can support the direct installation of the aforementioned voice app, or the configuration of drivers, components, etc., corresponding to the aforementioned voice app, to achieve communication with the same voice app on other devices.
[0146] Figure 3 This is a schematic diagram of the structure of the VoIP evaluation system 300 according to an embodiment of this application, as shown below. Figure 3 As shown, the VoIP evaluation system 300 includes a test host 301 and a device under test 302; wherein, the test host 301 is equipped with a voice application APP 301a, and the device under test 302 includes a microphone 302a and a speaker 302b; the voice APP 301a and the device under test 302 communicate through an IP network 303.
[0147] In one possible implementation, microphone 302a records a first sound to obtain a first comparison audio, the first sound being emitted by playing a first original audio; the device under test 302 transmits the first comparison audio to voice APP 301a; the test host 301 compares the first comparison audio with the first original audio to obtain the sound pickup evaluation result of the device under test 302.
[0148] The audio pickup evaluation results correspond to the audio pickup evaluation process when the Device Under Test (DUT) 302 is used as the transmitter (also known as the uplink evaluation process), that is,
[0149] The device under test 302 simulates the voice transmitter in VoIP. It records the first sound (including human voice and possibly ambient sounds such as various noises) through the microphone 302a to obtain the first comparison audio. The first comparison audio is digitized from an analog signal to obtain voice data. The voice data is then compressed using a voice compression algorithm to obtain a bitstream. The bitstream is then packaged according to the TCP / IP standard to obtain IP data packets. The IP data packets are then transmitted to the test host 301 via the IP network.
[0150] At this time, the test host 301 simulates the voice receiver in VoIP. The voice APP 301a receives IP data packets. Then, the test host 301 decapsulates and concatenates these IP data packets according to the TCP / IP standard to obtain the bitstream. After decompressing the bitstream, the voice data is obtained. Then, the voice data is digitally converted into analog signals to recover the first comparison audio.
[0151] The test host 301 compares the first comparison audio with the first original audio (in the evaluation scenario, the first original audio can be acquired by the test host, selected by the test host from a pre-stored audio library, or generated by the test host through audio generation software, without specific limitations) to obtain the sound pickup evaluation result of the device under test 302. The evaluation result may include the perceptual evaluation of speech quality (PESQ) index and the extended short-time objective intelligence (ESTOI) index. The acquisition method can refer to the acquisition method of the corresponding index, without specific limitations.
[0152] In one possible implementation, the voice app 301a transmits the second original audio to the device under test 302; the speaker 302b plays the second original audio to emit a second sound; the test host 301 records the second sound to obtain a second comparison audio; the test host 301 compares the second comparison audio with the second original audio to obtain the sound playback evaluation result of the device under test 302.
[0153] The playback evaluation results correspond to the playback evaluation process when the device under test 302 acts as the receiving end (also known as the downlink evaluation process), that is,
[0154] The test host 301 simulates the voice transmitter in VoIP. It digitizes the second original audio (including human voice; in the evaluation scenario, the second original audio can be acquired by the test host, selected by the test host from a pre-stored audio library, or generated by the test host through audio generation software, without specific limitations) to obtain voice data. Then, it compresses the voice data using a voice compression algorithm to obtain a bitstream. The bitstream is then packaged according to the TCP / IP standard to obtain IP data packets. The IP data packets are transmitted to the device under test 302 via the IP network through the voice APP 301a.
[0155] At this time, the device under test 302 simulates the voice receiver in VoIP, receives IP data packets, and then the device under test 302 decapsulates and concatenates these IP data packets according to the TCP / IP standard to obtain the bit stream. After decompressing the bit stream, the voice data is obtained. Then, the voice data is digitally converted into analog signals to recover the second original audio. The speaker 302b plays the second comparison audio to emit the second sound.
[0156] To obtain the evaluation results, the second sound played needs to be compared with the original sound. Therefore, a standard microphone (which can reduce noise interference and recover as much sound as possible from the speaker 302b) is used to record the second comparison audio. The test host 301 compares the second comparison audio with the second original audio to obtain the sound playback evaluation results of the device under test 302. These evaluation results may include PESQ and ESOI indices, and the methods for obtaining them can refer to the methods for obtaining the corresponding indices, without specific limitations.
[0157] The first and second original audio recordings mentioned above each include human voices, which can be the sounds of a person simulating a dialogue scene.
[0158] Optionally, the first raw audio may also include noise, including steady noise (e.g., wind noise, white noise, etc.) and / or sudden noise (e.g., the sound of a door opening, the sound of a cup breaking, etc.).
[0159] For example, the VoIP evaluation system 300 can evaluate sound pickup scenarios. For instance, in a quiet environment of 40dB, the first original audio contains human voices and steady-state noise. By comparing the first comparison audio with the first original audio, the PESQ and ESOI metrics under steady-state noise can be obtained. Alternatively, in a quiet environment of 40dB, the first original audio contains human voices and sudden noise. By comparing the first comparison audio with the first original audio, the PESQ and ESOI metrics under sudden noise can be obtained. Or, in a quiet environment of 40dB, the first original audio contains human voices, steady-state noise, and sudden noise. By comparing the first comparison audio with the first original audio, the PESQ and ESOI metrics under a mixture of steady-state noise and sudden noise can be obtained. Therefore, by mixing in different types of noise, the sound pickup performance of the device under test under different noise levels can be measured.
[0160] For example, the VoIP evaluation system 300 can evaluate playback scenarios. For instance, in a quiet environment of 40dB, if the second original audio contains human voices, and the second comparison audio is compared with the second original audio, playback PESQ and playback ESOI metrics can be obtained. Thus, the playback performance of the device under test can be measured in a quiet environment.
[0161] For example, the VoIP evaluation system 300 can evaluate two-way call scenarios. In a quiet environment of 40dB, the first original audio and the second original audio contain human voices (the two can be the same or different human voices, which is not specifically limited). Then, the loudness difference between the first comparison audio and the first original audio can be calculated to obtain the evaluation result of echo cancellation.
[0162] In addition to the examples above, other test cases can be designed to comprehensively evaluate the VoIP call performance of the device under test, and this application does not impose any specific limitations on this.
[0163] In the aforementioned evaluation process, IP data packets were transmitted using the Voice App 301a. This Voice App does not require separate deployment or design; instead, it can be existing VoIP voice software (e.g., WeChat, QQ, carrier camera applications (e.g., China Mobile's Hejiaqin, China Telecom's Tianyi Guanjia), etc.). This facilitates co-deployment with the test host. Even if the voice app's supported system is incompatible with the test host's operating system, it can be deployed on the test host using a virtual machine. Furthermore, using an existing voice app reduces physical cabling, thereby lowering costs.
[0164] During the sound pickup evaluation, the microphone 302a records the first sound, allowing for the evaluation of the microphone's frequency response and loudness. During the sound playback evaluation, the speaker 302b plays the second comparison audio, allowing for the evaluation of the speaker's frequency response and loudness. Therefore, the VoIP evaluation system of this application can also effectively evaluate hardware.
[0165] Optionally, the first sound can be emitted by playing the first original audio through an artificial mouth, which can make the first sound closer to a human voice to better simulate a call scenario.
[0166] Optionally, the second comparison audio can be obtained by recording a second sound using a standard microphone. This can reduce noise interference, recover the sound played by the speaker as much as possible, and avoid interference from other factors on the evaluation results.
[0167] This application embodiment provides a VoIP evaluation system that can perform hardware evaluation of the microphone and speaker of the device under test, as well as software evaluation of the sound pickup and playback processes. Furthermore, it can utilize existing voice apps to achieve voice communication, thereby reducing costs.
[0168] Figure 4 This is a schematic diagram of the structure of the VoIP evaluation system according to an embodiment of this application, as shown below. Figure 4 As shown, in Figure 3 Based on the VoIP evaluation system shown, the test host also includes: audio evaluation software, audio acquisition software, and audio driver software. The audio acquisition software is used to acquire or generate the first raw audio (e.g., s0 (human voice) or s0 plus n0 (noise)) or the second raw audio (e.g., x0). The audio driver software is used to mix different audios and provide an audio transmission channel. The audio evaluation software is used to obtain the pickup or call evaluation results of the device under test or the playback evaluation results of the device under test. Furthermore, the VoIP evaluation system also includes: a sound card for transmitting the first raw audio or the second comparison audio.
[0169] based on Figure 4 The evaluation process for the VoIP evaluation system shown in this application is as follows:
[0170] I. The device under test acts as the transmitter:
[0171] 1. The audio acquisition software acquires the first original audio, either human voice s0 or human voice s0 mixed with noise n0, and plays the first sound through audio driver software (e.g., audio virtual channel), sound card, and artificial mouth; (Sequence numbers 1, 2, 3, 4, 5, 6)
[0172] 2. The device under test picks up the first sound through the microphone to obtain the first comparison audio, and sends the first comparison audio to the voice APP via the network; (Sequence 7)
[0173] 3. The audio driver software acquires the first comparison audio from the voice APP and sends it to the audio acquisition software for recording as s1; (serial numbers 8 and 9)
[0174] 4. Audio evaluation software performs sound pickup and call evaluation on S0 and S1. (Item 10)
[0175] II. The device under test acts as the receiving end:
[0176] 1. The audio acquisition software obtains human voice x0 as the second raw audio, which is then sent to the voice APP via the audio driver software; (Sequence 1, 2)
[0177] 2. The voice app sends the second original audio file to the device under test via the network; (Item 3)
[0178] 3. The device under test plays the second original audio through a speaker to produce the second sound. The standard microphone picks up the second sound to obtain the second comparison audio, and sends it to the audio acquisition software via the sound card and audio driver software to be recorded as x1; (serial numbers 4, 5, 6, 7, 8, 9).
[0179] 4. Audio testing software is used to evaluate the playback performance of x0 and x1. (Item 10)
[0180] The following are several specific embodiments to illustrate... Figure 3 and Figure 4 The technical solutions of the embodiments shown will be described in detail.
[0181] Figure 5 This is a schematic diagram of the audio test of the IoT camera according to an embodiment of this application, as shown below. Figure 5 As shown, taking the mobile Hejiaqin APP and the mobile home security camera as examples, the voice APP is the Hejiaqin APP and the tested device is the IoT camera.
[0182] The hardware and software architecture includes:
[0183] 1. RME-FIREFACE UCX-II Sound Card
[0184] 2. Artificial mouth amplifier & standard microphone power supply
[0185] 3. AM3000 artificial mouthpiece
[0186] 4. Test host: PC (with Adobe Audition (audio capture software installed, such as...) Figure 6 ( Figure 6(As shown in the diagram for multi-track audio), ASIO LINK pro (audio driver software, such as...) Figure 7 ( Figure 7 (as shown in the diagram of the virtual bar), Android virtual machine and Hejiaqin app (e.g.) Figure 8 ( Figure 8 (This is a diagram of the Hejiaqin app)
[0187] 5. Device under test: IoT camera
[0188] 6. Network: Wireless mobile router
[0189] Based on the above hardware and software architecture, the evaluation process for this application is as follows:
[0190] I. The device under test acts as the transmitter:
[0191] 1. Adobe Audition uses multitrack audio to extract the first original audio mix0, either human voice s0 or human voice s0 mixed with noise n0, n1..., and plays it out as the first sound through ASIO LINK pro, RME-FIREFACE UCX-II sound card and AM3000 artificial mouth.
[0192] 2. The IoT camera picks up the first sound emitted by the AM3000 artificial mouth through the microphone to obtain the first comparison audio, and sends the first comparison audio to the Hejiaqin APP deployed on the PC through the network;
[0193] 3. ASIO LINK Pro picks up the first comparison audio from the Hejiaqin APP and sends it to Adobe Audition for recording as S1;
[0194] 4. The indicator evaluation software performs sound pickup and call evaluation on s0 and s1.
[0195] II. The device under test acts as the receiving end:
[0196] 1. Adobe Audition uses the human voice x0 as the second original audio, processes it through ASIO LINK Pro, and sends it to the Hejiaqin APP;
[0197] 2. The Hejiaqin APP sends the second original audio file to the IoT camera via the network;
[0198] 3. The IoT camera plays a second original audio through a speaker to produce a second sound. The standard microphone picks up the second sound to obtain a second comparison audio, and sends it to Adobe Audition via the RME-FIREFACE UCX-II sound card and ASIO LINK pro for recording as x1.
[0199] 4. The indicator evaluation software performs sound playback evaluation on x0 and x1.
[0200] The following test cases can be used in this embodiment:
[0201] I. Hardware Testing
[0202] like Figure 9 ( Figure 9 (A schematic diagram of 0-20kHz sinusoidal speech) is shown.
[0203] 1. When s0 is a sine wave audio frequency of 0-20kHz, comparing and evaluating s0 and s1 can test the frequency response of the microphone of the IoT camera.
[0204] 2. When x0 is a sine wave audio frequency of 0-20kHz, comparing and evaluating x0 and x1 can test the frequency response of the speaker of the IoT camera.
[0205] II. Software Algorithm Testing:
[0206] 1. The VoIP evaluation system 300 can evaluate sound pickup scenarios. For example, in a quiet environment of 40dB, the first original audio contains human voice and steady-state noise. By comparing the first comparison audio with the first original audio, the PESQ and ESOI metrics under steady-state noise can be obtained. Alternatively, in a quiet environment of 40dB, the first original audio contains human voice and sudden noise. By comparing the first comparison audio with the first original audio, the PESQ and ESOI metrics under sudden noise can be obtained. Or, in a quiet environment of 40dB, the first original audio contains human voice, steady-state noise, and sudden noise. By comparing the first comparison audio with the first original audio, the PESQ and ESOI metrics under mixed steady-state noise and sudden noise can be obtained. Therefore, by mixing in different types of noise, the sound pickup performance of the device under test under different noise levels can be measured.
[0207] 2. The VoIP evaluation system 300 can evaluate playback scenarios. For example, in a quiet environment of 40dB, the second original audio contains human voices. By comparing the second comparison audio with the second original audio, playback PESQ and playback ESOI indices can be obtained. Therefore, the playback performance of the device under test can be measured in a quiet environment.
[0208] 3. Based on the VoIP evaluation system 300, two-way call scenarios can be evaluated. In a quiet environment of 40dB, the first original audio and the second original audio contain human voices (the two can be the same or different human voices, which is not specifically limited). Then, the loudness difference between the first comparison audio and the first original audio can be calculated to obtain the evaluation result of echo cancellation.
[0209] Figure 10 This is a schematic diagram of the audio test of a large screen according to an embodiment of this application, as shown below. Figure 10 As shown, taking the WeChat Work APP and a large screen as examples, the voice APP is the WeChat Work APP, and the device under test is the large screen.
[0210] The hardware and software architecture includes:
[0211] 1. RME-FIREFACE UCX-II Sound Card
[0212] 2. Artificial mouth amplifier & standard microphone power supply
[0213] 3. AM3000 artificial mouthpiece
[0214] 4. Test host: PC (with Adobe Audition (audio capture software installed, such as...) Figure 6 As shown), ASIOLINK pro (audio driver software, such as...) Figure 7 (As shown), Android virtual machine and WeChat Work APP
[0215] 5. Device under test: Large screen
[0216] 6. Network: Wireless mobile router
[0217] Based on the above hardware and software architecture, the evaluation process for this application is as follows:
[0218] I. The device under test acts as the transmitter:
[0219] 1. Adobe Audition uses multitrack audio to extract the first original audio mix0, either human voice s0 or human voice s0 mixed with noise n0, n1..., and plays it out as the first sound through ASIO LINK pro, RME-FIREFACE UCX-II sound card and AM3000 artificial mouth.
[0220] 2. The large screen picks up the first sound emitted by the AM3000 artificial mouth through the microphone to obtain the first comparison audio, and sends the first comparison audio to the enterprise WeChat APP deployed on the PC via the network;
[0221] 3. ASIO LINK Pro picks up the first comparison audio from the WeChat Work APP and sends it to Adobe Audition for recording as S1;
[0222] 4. The indicator evaluation software performs sound pickup and call evaluation on s0 and s1.
[0223] II. The device under test acts as the receiving end:
[0224] 1. Adobe Audition uses the human voice x0 as the second original audio, processes it through ASIO LINK Pro, and sends it to the WeChat app for businesses;
[0225] 2. The WeChat Work app sends the second original audio file to the large screen via the network;
[0226] 3. The large screen plays the second original audio through the speakers to produce the second sound. The standard microphone picks up the second sound to obtain the second comparison audio, and sends it to Adobe Audition for recording as x1 via the RME-FIREFACE UCX-II sound card and ASIO LINK pro.
[0227] 4. The indicator evaluation software performs sound playback evaluation on x0 and x1.
[0228] The test cases in this embodiment can be referred to. Figure 5 The test cases shown in the embodiments will not be described again here.
[0229] Based on the aforementioned VoIP evaluation system Figure 11 This is a flowchart of process 1100 of the VoIP evaluation method provided in an embodiment of this application. Process 1100 can be executed by the VoIP evaluation system described above. Process 1100 is described as a series of steps or operations, and it should be understood that process 1100 can be executed in various orders and / or occur simultaneously, and is not limited to... Figure 11 The execution order is shown. Process 1100 may include:
[0230] Step 1101: Record the first sound through the microphone to obtain the first comparison audio.
[0231] As described above, the first sound can be a first original audio signal played back (e.g., played back via an artificial mouth). This first original audio signal may include a human voice, and optionally, it may also include noise, including steady-state noise and / or burst noise, to accommodate different test cases.
[0232] Step 1102: Compare the first comparison audio with the first original audio to obtain the sound pickup evaluation result of the device under test.
[0233] The sound pickup evaluation results correspond to the sound pickup evaluation process when the device under test is the transmitter, that is,
[0234] The device under test simulates a voice transmitter in VoIP, recording a first sound (including human voice and possibly ambient sounds such as various noises) through a microphone to obtain a first comparison audio. The first comparison audio is then digitized from an analog signal to obtain voice data. The voice data is then compressed using a voice compression algorithm to obtain a bitstream. The bitstream is then packaged according to the TCP / IP standard to obtain IP data packets. The IP data packets are then transmitted to the test host via the IP network.
[0235] At this point, the test host simulates the voice receiver in VoIP, receives IP data packets, and then decapsulates and concatenates these IP data packets according to the TCP / IP standard to obtain the bitstream. After decompressing the bitstream, the voice data is obtained, and then the voice data is digitally converted into analog signals to recover the first comparison audio.
[0236] The test host compares the first comparison audio with the first original audio (in the evaluation scenario, the first original audio can be acquired by the test host, selected by the test host from a pre-stored audio library, or generated by the test host through audio generation software, without specific limitations) to obtain the sound pickup evaluation result of the device under test. The evaluation result may include the objective speech quality assessment index PESQ and the ESOI index. The acquisition method can refer to the acquisition method of the corresponding index, without specific limitations.
[0237] In the evaluation of sound pickup scenarios, the microphone's frequency response and loudness can be evaluated by recording the first sound. Therefore, the VoIP evaluation method of this application can also achieve effective hardware evaluation.
[0238] This application embodiment provides a VoIP evaluation method that can achieve both hardware evaluation of the microphone of the device under test and software evaluation of the sound pickup process.
[0239] Figure 12 This is a flowchart of process 1200 of the VoIP evaluation method provided in an embodiment of this application. Process 1200 can be executed by the VoIP evaluation system described above. Process 1200 is described as a series of steps or operations, and it should be understood that process 1200 can be executed in various orders and / or occur simultaneously, and is not limited to... Figure 12 The execution order is shown. Process 1200 may include:
[0240] Step 1201: Play the second original audio through a speaker to produce a second sound.
[0241] As mentioned above, the second original audio includes human voices.
[0242] Step 1202: Record the second sound to obtain the second comparison audio.
[0243] The second comparison audio can be obtained by recording a second sound using a standard microphone.
[0244] Step 1203: Compare the second comparison audio with the second original audio to obtain the sound playback evaluation result of the device under test.
[0245] The playback evaluation results correspond to the playback evaluation process when the device under test is used as the receiving end, that is,
[0246] The test host simulates the voice transmitter in VoIP. It digitizes the second raw audio (including human voice; in the evaluation scenario, the second raw audio can be acquired by the test host, selected by the test host from a pre-stored audio library, or generated by the test host using audio generation software, without specific limitations) to obtain voice data. Then, it compresses the voice data using a voice compression algorithm to obtain a bitstream. The bitstream is then packaged according to the TCP / IP standard to obtain IP data packets. The IP data packets are then transmitted to the device under test via the IP network.
[0247] At this time, the device under test simulates the voice receiver in VoIP, receives IP data packets, and then the device under test decapsulates and concatenates these IP data packets according to the TCP / IP standard to obtain the bitstream. After decompressing the bitstream, the voice data is obtained. Then, the voice data is digitally converted into analog signals to recover the second original audio. The speaker plays the second comparison audio to produce the second sound.
[0248] To obtain the evaluation results, the second sound played out needs to be compared with the original sound. Therefore, a standard microphone (which can reduce noise interference and recover as much sound as possible from the speaker) is used to record the second sound to obtain the second comparison audio. The test host compares the second comparison audio with the second original audio to obtain the sound playback evaluation results of the device under test. These evaluation results may include PESQ and ESOI indicators, and the methods for obtaining them can refer to the methods for obtaining the corresponding indicators, without specific limitations.
[0249] During the evaluation of the sound playback scenario, the second comparison audio is played by the speaker, which allows for the evaluation of the speaker's frequency response and loudness. It is evident that the VoIP evaluation method of this application can also achieve effective hardware evaluation.
[0250] This application embodiment provides a VoIP evaluation method that can achieve both hardware evaluation of the microphone and speaker of the device under test, and software evaluation of the sound pickup and playback processes.
[0251] It should be noted that, in Figure 3 or Figure 4 In the VoIP evaluation system shown, Figure 11 and Figure 12 Both methods shown can be applied. These two methods correspond to the evaluation of the sound pickup scenario and the sound playback scenario, respectively. Therefore, these two methods can be applied separately in the VoIP evaluation system to evaluate one-way calls, or they can be applied simultaneously in the VoIP evaluation system to evaluate two-way calls. There are no specific limitations on this.
[0252] Figure 13 This is a schematic diagram of the structure of the VoIP testing device 1300 of this application, as shown below. Figure 13 As shown, the VoIP evaluation device 1300 of this embodiment can be applied to the VoIP evaluation system described above. The VoIP evaluation device 1300 may include: a recording module 1301, a playback module 1302, and an evaluation module 1303. Wherein,
[0253] The recording module 1301 is used to record a first sound through a microphone to obtain a first comparison audio, wherein the first sound is emitted by playing a first original audio; the evaluation module 1303 is used to compare the first comparison audio with the first original audio to obtain the sound pickup evaluation result of the device under test.
[0254] In one possible implementation, the first original audio includes human voice.
[0255] In one possible implementation, the first original audio also includes noise, which includes stationary noise and / or burst noise.
[0256] In one possible implementation, the first sound is emitted by playing the first original audio through an artificial mouth.
[0257] In one possible implementation, the playback module 1302 is used to play the second original audio through a speaker to emit a second sound; the recording module 1301 is further used to record the second sound to obtain a second comparison audio; and the evaluation module 1303 is further used to compare the second comparison audio with the second original audio to obtain the playback evaluation result of the device under test.
[0258] In one possible implementation, the second original audio includes human voice.
[0259] In one possible implementation, the second comparison audio is obtained by recording the second sound using a standard microphone.
[0260] The apparatus of this embodiment can be used to perform Figure 11 or Figure 12 The technical solutions of the method embodiments shown are similar in principle and in effect, and will not be described again here.
[0261] In implementation, each step of the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this application can be directly implemented by a hardware encoding processor, or implemented by a combination of hardware and software modules in the encoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0262] The memory mentioned in the above embodiments can be volatile memory or non-volatile memory, or may include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0263] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0264] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0265] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0266] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0267] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0268] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0269] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A VOIP evaluation system, characterized by, Comprising: a test host and a device under test; wherein the test host is loaded with a voice application program APP, and the device under test comprises a microphone and a loudspeaker; the voice APP and the device under test communicate through an Internet Protocol IP network; the microphone records a first sound to obtain a first comparison audio, the first sound being emitted by playing a first original audio; the device under test transmits the first comparison audio to the voice APP; the test host compares the first comparison audio with the first original audio to obtain a pickup evaluation result of the device under test; the voice APP transmits a second original audio to the device under test; the loudspeaker plays the second original audio to emit a second sound; the test host records the second sound to obtain a second comparison audio; the test host compares the second comparison audio with the second original audio to obtain a playback evaluation result of the device under test.
2. The system of claim 1, wherein, The first original audio and the second original audio respectively comprise human voice.
3. The system of claim 2, wherein, The first original audio further comprises noise, and the noise comprises stationary noise and / or burst noise.
4. The system of any one of claims 1-3, wherein, The first sound is emitted by playing the first original audio through an artificial mouth; and / or, the second comparison audio is obtained by recording the second sound through a standard microphone.
5. The system of any one of claims 1-4, wherein, The test host further comprises audio acquisition software and audio driving software; the audio acquisition software is used to acquire the first original audio or the second original audio; and the audio driving software is used to mix different audios and provide a transmission channel of the audio.
6. The system of any one of claims 1-5, wherein, The test host further comprises audio evaluation software, which is used to obtain the pickup evaluation result of the device under test or the playback evaluation result of the device under test.
7. The system of any one of claims 1-6, wherein, Further comprising: a sound card; The sound card is used to transmit the first original audio or the second comparison audio.
8. The system of any one of claims 1-7, wherein, The device under test comprises an Internet of Things IOT camera or a smart screen.
9. The system of any one of claims 1-8, wherein, The evaluation result comprises an objective speech quality assessment PESQ index and / or an intelligibility ESTOI index.
10. A method of VOIP evaluation, characterized by, The method is applied to the VOIP evaluation system of any one of claims 1-9, and the method comprises: recording a first sound through a microphone to obtain a first comparison audio, the first sound being emitted by playing a first original audio; comparing the first comparison audio with the first original audio to obtain a pickup evaluation result of the device under test.
11. The method of claim 10, wherein, The first original audio comprises human voice.
12. The method of claim 11, wherein, The first original audio further comprises noise, and the noise comprises stationary noise and / or burst noise.
13. The method according to any one of claims 10-12, characterized in that, The first sound is emitted by playing the first original audio through an artificial mouth.
14. The method according to any one of claims 10-13, characterized in that, Further comprising: playing a second original audio through a loudspeaker to emit a second sound; recording the second sound to obtain a second comparison audio; comparing the second comparison audio with the second original audio to obtain a playback evaluation result of the device under test.
15. The method of claim 14, wherein, The second original audio comprises human voice.
16. The method according to claim 14 or 15, characterized in that The second comparison audio is obtained by recording the second sound through a standard microphone.
17. A method of VOIP evaluation, characterized by, The method is applied to the VOIP evaluation system of any one of claims 1-9, and the method comprises: playing the second original audio through a speaker to emit a second sound; recording the second sound to obtain a second comparison audio; comparing the second comparison audio with the second original audio to obtain a playback evaluation result of the device under test.
18. The method of claim 17, wherein, The second original audio comprises human voice.
19. The method of claim 17 or 18, wherein, The second comparison audio is obtained by recording the second sound through a standard microphone.
20. A terminal device, comprising: comprise: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 10-19.
21. A computer-readable storage medium, characterized in that, The computer program comprises a computer program which, when executed on a computer, causes the computer to perform the method of any one of claims 10-19.
22. A computer program product, characterised in that, The computer program product comprises computer program code which, when executed on a computer, causes the computer to perform the method of any one of claims 10-19.