Test method, device, electronic device, and storage medium

By setting up the speech recognition service SDK and player SDK under test in different processes and adopting cross-process communication technology, the problem of insufficient accuracy in performance testing of speech recognition service SDK in noisy scenarios is solved, and more accurate and stable performance evaluation is achieved.

CN116150032BActive Publication Date: 2026-05-01BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2023-03-16
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies lack the accuracy and stability for performance testing of speech recognition service SDKs in noisy scenarios, making it impossible to accurately evaluate the performance of speech recognition service SDKs on different terminal devices.

Method used

By setting the speech recognition service SDK and the player SDK under test in different processes and using cross-process communication technology, the impact of internal noise playback and processing on the operation of the speech recognition service SDK is eliminated, and its performance data is obtained to determine the performance test results.

Benefits of technology

It improves the accuracy and stability of performance testing for the speech recognition service SDK, ensures the reliability and consistency of test data, and adapts to various scenarios and devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116150032B_ABST
    Figure CN116150032B_ABST
Patent Text Reader

Abstract

The present disclosure provides a test method and device, electronic equipment and storage medium. The present disclosure relates to the technical field of computers, specifically to the technical field of software testing. The specific implementation scheme is: in response to detecting a test instruction, running a test software, the test software integrating a to-be-tested speech recognition service software development kit (SDK) and a player SDK, the to-be-tested speech recognition service SDK and the player SDK belonging to different processes, the to-be-tested speech recognition service SDK and the player SDK being capable of cross-process communication, the player SDK being used for processing internal noise data; obtaining performance data of a process in which the to-be-tested speech recognition service SDK is located formed in a running process of the test software; and determining a performance test result of the to-be-tested speech recognition service SDK based on the performance data of the process in which the to-be-tested speech recognition service SDK is located. According to the scheme of the present disclosure, the accuracy of the performance test of the speech recognition service SDK can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Test methods, apparatus, electronic devices and storage media Technical Field

[0001] This disclosure relates to the field of computer technology, specifically the field of software testing technology. Background Technology

[0002] With the rapid development of artificial intelligence technology and breakthroughs in core technologies, voice-interactive smart devices, in-vehicle terminals, and mobile terminals such as smartphones all rely on voice software development kits (SDKs) for functional development and expansion. The performance of software integrating a speech recognition service SDK on a device directly impacts the user experience. Therefore, testers need to conduct full-scenario performance testing of the device or application's voice interaction to ensure its quality. However, the accuracy of speech recognition service SDK performance testing in noisy environments is a primary condition for R&D personnel to analyze their technical solutions and is one of the challenges that testers in the field of speech recognition need to overcome. Summary of the Invention

[0003] This disclosure provides a test method, apparatus, electronic device, and storage medium.

[0004] According to a first aspect of this disclosure, a testing method is provided, comprising:

[0005] In response to the detection of the test command, the test software is run. The test software integrates the speech recognition service SDK under test and the player SDK. The speech recognition service SDK under test and the player SDK belong to different processes. The speech recognition service SDK under test and the player SDK can communicate across processes. The player SDK is used to process internal noise data.

[0006] Obtain performance data of the process containing the tested speech recognition service SDK generated during the operation of the test software;

[0007] Based on the performance data of the process where the speech recognition service SDK under test is located, the performance test results of the speech recognition service SDK under test are determined.

[0008] According to a second aspect of this disclosure, a testing apparatus is provided, comprising:

[0009] The runtime module is used to run the test software in response to the detection of the test command. The test software integrates the speech recognition service SDK under test and the player SDK. The speech recognition service SDK under test and the player SDK belong to different processes. The speech recognition service SDK under test and the player SDK can communicate across processes. The player SDK is used to process internal noise data.

[0010] The first acquisition module is used to acquire the performance data of the process where the tested speech recognition service SDK is located during the operation of the test software.

[0011] The determination module is used to determine the performance test results of the speech recognition service SDK under test based on the performance data of the process where the speech recognition service SDK under test is located.

[0012] According to a third aspect of this disclosure, an electronic device is provided, comprising:

[0013] At least one processor;

[0014] Memory that is communicatively connected to at least one processor;

[0015] The memory stores instructions that can be executed by at least one processor to enable the at least one processor to perform the methods of any embodiment of this disclosure.

[0016] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform a method according to any embodiment of this disclosure.

[0017] According to a fifth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements a method according to any embodiment of this disclosure.

[0018] The solution disclosed herein can improve the accuracy of performance testing of speech recognition service SDK.

[0019] The above overview is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of this application will become readily apparent from the accompanying drawings and the following detailed description. Attached Figure Description

[0020] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments disclosed in this application and should not be construed as limiting the scope of this application.

[0021] Figure 1 is a schematic diagram of the architecture for performance testing of the speech recognition service SDK according to an embodiment of the present disclosure;

[0022] Figure 2 is a flowchart illustrating a test method according to an embodiment of the present disclosure;

[0023] Figure 3 is a schematic diagram of the architecture for performance testing of the speech recognition service SDK according to an embodiment of the present disclosure;

[0024] Figure 4 is a schematic diagram of the test processing when there is internal noise according to an embodiment of the present disclosure;

[0025] Figure 5 is a schematic diagram of the test process when there is no internal noise according to an embodiment of the present disclosure;

[0026] Figure 6 is a schematic diagram of the process of obtaining identification response information according to an embodiment of the present disclosure;

[0027] Figure 7 is a schematic diagram of the structure of a test apparatus according to an embodiment of the present disclosure;

[0028] Figure 8 is a schematic diagram of a test method according to an embodiment of the present disclosure;

[0029] Figure 9 is a schematic diagram of the structure of an electronic device used to implement the test method of the embodiments of this disclosure. Detailed Implementation

[0030] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0031] The terms "first," "second," and "third," etc., used in the embodiments, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.

[0032] In related technologies, the main functions of a voice interaction SDK include:

[0033] (1): Complete the exposure of external interfaces, including but not limited to the basic function interfaces of voice interaction, namely the voice recognition interface; complete the status return of the corresponding interface data through the observer model; complete the exposure of external public constant parameters through public methods and static variables.

[0034] (2): Organize important terminal data and communicate with the server via the network module. The speech recognition service SDK needs to process the audio during recognition; however, in noisy environments, the audio captured by the SDK through the microphone includes both audio played by the device itself and other audio. To improve recognition performance, noise reduction processing is required for the captured audio. The Acoustic Echo Cancellation (AEC) algorithm performs noise reduction processing using the audio captured by the microphone and the reference audio played by the device itself (also known as ref audio).

[0035] In related technologies, the end-to-end software performance testing of voice interaction in intelligent voice devices and applications mainly includes the following test schemes:

[0036] (1): Offline performance test: Since the performance resources occupied by the speech recognition service SDK are mainly in the core algorithm module integrated inside, in order to quickly evaluate the performance of the current version, the algorithm module can be run directly offline locally. By constructing virtual input to make the algorithm run normally, the information of the currently occupied central processing unit (CPU) and memory can be captured to reflect the overall performance of the SDK.

[0037] (2): Application Layer Performance Testing: Software testing for smart terminals typically requires control via a personal computer (PC), deploying a test suite on the PC. Its core comprises four modules: the first is the test control module (also known as the software control instruction module), used to manipulate the terminal device, responsible for sending scheduling instructions such as installing or uninstalling the test Android application package (APK) and pushing resource files to the designated terminal device; the second is the external noise control module, used for simulating external noise during actual use and recognizing and playing interactive dialogue; the third is the performance acquisition module, used to use built-in system tools to obtain corresponding CPU and memory data based on the test software process; the fourth is the data analysis and result display module, used to format and analyze the acquired performance data, ultimately generating a test report that can be displayed in Hyper Text Markup Language (HTML) and table (Microsoft Office Excel) formats to assist in locating problems and completing test tasks.

[0038] Figure 1 illustrates the architecture diagram for the performance test of the speech recognition service SDK. As shown in Figure 1, speech recognition interaction software, i.e., the business application software under test (hereinafter referred to as the business APP), is installed on a smart device that interacts directly with the PC. During normal interaction, the smart device responds to each recognition response. The speech recognition service SDK needs to obtain the reference audio data played by the device and perform algorithm noise reduction processing. While integrating the Automatic Speech Recognition (ASR) SDK, the business layer application software also needs to integrate and maintain the data transmission module, the internal noise data processing module, the signal processing module, and the recognition engine. When the business APP under test is running, the data transmission module obtains the external audio signal from the microphone and transmits it to the signal processing module. At the same time, it confirms whether there is a response playback according to the business instructions, controls the internal noise data processing module to generate the corresponding internal noise data and saves it in the device memory (also known as the hardware kernel). The player plays the data and notifies the data transmission module to obtain the internal noise data from the device memory and transmit it to the signal processing module. The signal processing module performs noise reduction and amplification algorithms on the two types of audio data, and then the data transmission module transmits it to the recognition engine or cloud recognition service for processing to obtain the final recognition result.

[0039] In related technologies, the ASR recognition SDK and the playback of internal noise, as well as the generation and transmission of data, are two modules coupled together by the business logic. The interaction between them affects the utilization of performance resources. In fact, the technologies and forms used for response playback vary on different smart terminals. Different internal noise data processing modules will inevitably affect the accuracy of the performance test data of the entire tested business APP. Since the original end-to-end speech recognition ASR SDK performance test scheme reflects the changes in the performance utilization of the business APP process, under this scheme, different performance test data may be obtained for the ASR recognition SDK with the same performance utilization, making it impossible to give accurate performance evaluation and analysis conclusions.

[0040] In related technologies, offline performance testing schemes can only reflect the computational load of the algorithm. However, due to differences in CPU scheduling across different systems, the performance consumption of the ASR SDK integrated on a real device cannot be assessed. Therefore, it is also impossible to evaluate whether the current recognition technology is suitable for running on a specific device.

[0041] In related technologies, application-layer performance testing schemes can obtain the software performance of an end-to-end integrated ASR (Automatic Speech Recognition) SDK. However, due to the influence of the internal noise data processing module, the performance consumption of the ASR SDK may vary across different terminal devices in actual applications. Therefore, without an evaluation of the performance of the speech recognition service SDK itself, performance test analysis conclusions may be misleading, hindering the expansion, iteration, and implementation of the technology.

[0042] To at least partially address one or more of the aforementioned problems and other potential issues, this disclosure proposes a performance testing method for a speech recognition service SDK. This method utilizes cross-process communication to eliminate internal noise playback and processing within the speech recognition service SDK, thereby making the performance test data more accurate and stable. Furthermore, it resolves the issue of performance data fluctuations during recognition testing, thus improving the accuracy and effectiveness of speech recognition service SDK performance testing.

[0043] This disclosure provides a testing method. Figure 2 is a flowchart illustrating the testing method according to an embodiment of this disclosure. This testing method can be applied to a testing device. The testing device is located in an electronic device. The electronic device includes, but is not limited to, fixed devices and / or mobile devices. For example, fixed devices include, but are not limited to, servers, which can be cloud servers or ordinary servers. For example, mobile devices include, but are not limited to, mobile phones, tablets, and vehicle terminals. In some possible implementations, the testing method can also be implemented by a processor calling computer-readable instructions stored in memory. As shown in Figure 2, the testing method includes:

[0044] S201: In response to the detection of a test command, run the test software; wherein, the test software integrates the speech recognition service SDK under test and the player SDK, the speech recognition service SDK under test and the player SDK belong to different processes, the speech recognition service SDK under test and the player SDK can perform cross-process communication, and the player SDK is used to process internal noise data;

[0045] S202: Obtain performance data of the process containing the tested speech recognition service SDK generated during the operation of the test software;

[0046] S203: Determine the performance test results of the speech recognition service SDK under test based on the performance data of the process where the speech recognition service SDK under test is located.

[0047] In this embodiment, the test software is also referred to as a test service APP. The test software may include a speech recognition service SDK module under test and a player SDK module. The speech recognition service SDK and the player SDK belong to different processes, and can communicate across processes using cross-process technology. The test software may be a test service APP that integrates the speech recognition service SDK, and this test service APP needs to use various functions of the speech recognition service SDK.

[0048] In this embodiment of the disclosure, the player SDK is used to process internal noise data, specifically including the transmission of internal noise information and the generation of internal noise data.

[0049] In this embodiment of the disclosure, the speech recognition service SDK under test and the player SDK can communicate across processes using cross-process communication technology. For example, this cross-process communication technology can be Remote Procedure Call (RPC) technology.

[0050] In this embodiment, the internal noise data refers to the internal noise generated by the audio data played by the device itself. External noise may include voice wake-up audio, voice command audio, background audio, and other audio. For example, in one scenario, a smart speaker is playing crosstalk, a mobile phone is playing music, and the user calls out to the smart speaker, "Xiaodu Xiaodu, what's the weather like today?" Here, the crosstalk played by the smart speaker is internal noise; the music played by the mobile phone, the voice wake-up audio "Xiaodu Xiaodu," and the voice command audio "what's the weather like today" are external noise.

[0051] In this embodiment of the disclosure, the performance test result of the speech recognition service SDK under test is determined based on the performance data of the process where the speech recognition service SDK under test resides. For example, the performance test result may be a stability test result. Or, for example, the performance test result may be an accuracy test result. The above are merely illustrative examples and are not intended to limit all possible information included in the performance test result; they are simply not exhaustive.

[0052] Figure 3 illustrates the second architecture diagram for performance testing of the speech recognition service SDK. As shown in Figure 3, the speech recognition SDK under test, such as the ASR SDK, is located in the process under test. This SDK may include a data transmission module, a signal processing module, and a local recognition engine; the signal processing module includes AEC noise reduction. The player SDK (also referred to as the SPEAK SDK) is located in the auxiliary process. This SDK may include an internal noise information transmission module and an internal noise data generation module. Since the ASR SDK and the player SDK are located in two separate processes, the influence of the player SDK on the operation of the ASR SDK is eliminated, thus making the performance test data more accurate and stable.

[0053] In this embodiment, the testing process includes: Step 1: The tester controls the PC-side test control module to install the test service APP on the device, and starts the test service APP and external noise control module according to the test scenario; Step 2: After the device-side test service APP runs successfully, the PC-side test control module controls the performance acquisition module to run and filters the process data of the ASR SDK on the test service APP obtained by the performance acquisition module; Step 3: After the test service APP obtains the external recognition speech audio, it calls the data transmission module to send the audio data received through the microphone to the signal processing module; the test service APP determines whether there is internal noise data; Step 4: If there is internal noise data, it performs cross-process communication with the internal noise information transmission module of the SPEAK SDK through RPC to obtain the internal noise data and transmit it to the signal processing module, which then performs AEC noise reduction algorithm processing; Step 5: If there is no internal noise data, it directly calls the signal processing module to perform algorithm processing; Step 6: ASR The SDK's data transmission module acquires the audio processed by the signal processing module. If it's offline mode recognition, it's passed to the local recognition engine; if it's online mode recognition, it's passed to the cloud recognition service for processing, obtaining the recognition result and response. Step 7: The ASR SDK's data transmission module transmits the recognition response information across processes via RPC to the SPEAK SDK's internal noise information transmission module for processing. It calls the internal noise data generation module to generate response data, controls the player to play, and writes the data to local memory. Step 8: When external recognition dialogue continues to play, the process jumps to step 4. Step 9: After the test time is reached, if other scenarios are detected, the process jumps to step 1, generating a new start command until all test scenarios are completed. Step 10: The test ends, and the test information is summarized through the data analysis and result display module.

[0054] The technical solution of this disclosure embodiment, in response to the detection of a test command, runs test software. The test software integrates a speech recognition service SDK under test and a player SDK. The speech recognition service SDK under test and the player SDK belong to different processes and can perform cross-process communication. The player SDK is used to process internal noise data. The system acquires performance data of the process containing the speech recognition service SDK under test generated during the operation of the test software. Based on the performance data of the process containing the speech recognition service SDK under test, the performance test result of the speech recognition service SDK under test is determined. By setting the speech recognition service SDK under test and the player SDK in different processes, the influence of internal noise playback and processing on the operation of the speech recognition service SDK can be eliminated, thereby improving the accuracy and effectiveness of speech recognition service SDK performance testing.

[0055] In some embodiments, the testing method may further include:

[0056] S204: Obtain test scenario information;

[0057] S205: Start the test software and external noise control module based on the test scenario information. The external noise control module is used to sense the noise level of the environment in which the terminal is located. The terminal is the terminal with the test software installed.

[0058] This disclosure does not limit the method of obtaining test scenario information. For example, testers can input test scenario information through a test interface on a PC. Alternatively, test scenario information can be obtained based on recorded test audio.

[0059] In this embodiment of the disclosure, when the test scenario involves an external noise environment, the external noise control module needs to be activated. The tester inputs test scenario information through the test interface on the PC; the test control module on the PC activates the test service APP on the terminal and the external noise control module on the PC based on the test scenario information.

[0060] In this embodiment of the disclosure, performance testing may include full-scene coverage testing and single-scene testing.

[0061] This single-scenario test can include: a quiet test scenario with no external or internal noise, an external noise test scenario with only external noise, and an internal noise test scenario with only internal noise. For example, a quiet test scenario with no external or internal noise includes a scenario where, in an indoor environment with no ambient noise, the smart speaker is in a waiting-to-wake state when it is not playing any audio files. Similarly, an external noise test scenario includes a scenario where, in an outdoor environment with ambient noise, the smart speaker is in a waiting-to-wake state when it is not playing any audio files. Finally, an internal noise test scenario includes a scenario where, in an indoor environment with no ambient noise, the smart speaker is in a waiting-to-wake state while playing other audio files.

[0062] In this embodiment of the disclosure, the full-scene coverage test refers to a complex multi-noise scenario that combines a single test scenario. For example, in an indoor environment with no ambient noise, when the smart speaker is playing an audio file and receives a wake-up command, the microphone simultaneously records the audio of the wake-up command and the audio played by the smart speaker. As another example, in an outdoor environment with ambient noise, when the smart speaker is playing an audio file and receives a wake-up command, the microphone simultaneously records the audio of the wake-up command, the audio played by the smart speaker, and the outdoor ambient noise.

[0063] In this way, by acquiring test scenario information and starting the test software and external noise control module based on the test scenario information, the diversity and versatility of the test method can be improved, making the test method applicable to multiple scenarios and multiple types of devices, thereby helping to improve the accuracy and effectiveness of the test method.

[0064] In some embodiments, S202 may include:

[0065] S202a: Based on the process identifier of the process where the speech recognition service SDK under test is located, filter out the performance data of the process where the speech recognition service SDK under test is located from the performance data of all processes.

[0066] In this embodiment of the disclosure, the performance acquisition module of the PC acquires the performance data of all processes.

[0067] In this embodiment of the disclosure, the performance acquisition module obtains device performance data through system-provided tools (such as Top and dumpsys). Top is used to obtain performance data of all software currently running on the device.

[0068] In this embodiment of the disclosure, the performance data of the process where the speech recognition service SDK is located is filtered out from the performance data of all processes based on the process identifier of the process where the speech recognition service SDK is located.

[0069] For example, the ASR SDK process is identified as 1, and the SPEAK SDK process is identified as 2. The performance acquisition module obtains the performance data of all processes through the system's built-in tool Top. Based on process identifier 1, it obtains the performance data of the process with process identifier 1, accurately filtering out the performance data of the process containing the tested speech recognition service SDK.

[0070] Thus, by accurately and quickly filtering out the performance data of the process containing the speech recognition service SDK under test from the performance data of all processes based on the process identifier of the process containing the speech recognition service SDK under test, the accuracy and efficiency of the test can be improved.

[0071] Figure 4 illustrates the test processing diagram when internal noise is present. As shown in Figure 4, the test method further includes: in response to acquiring the original audio data recorded by the recording device, when it is determined that internal noise exists, the speech recognition service SDK under test obtains the internal noise data by communicating with the player SDK, and performs noise reduction processing based on the original audio data and the internal noise data to obtain the target audio data; the speech recognition service SDK under test obtains the recognition response information of the target audio data; the speech recognition service SDK under test transmits the recognition response information across processes to the player SDK, so that the player SDK can generate response audio data based on the recognition response information, play the response audio corresponding to the response audio data, and write the response audio data to the terminal, which is the terminal with the test software installed.

[0072] In some implementations, the recording device may include a microphone, a sound card, and a microphone. The above is merely illustrative and is not intended to limit the scope of all possible recording devices; an exhaustive list is not provided here.

[0073] In some implementations, the raw audio data specifically refers to the unprocessed audio data recorded by the recording device.

[0074] In some implementations, this noise reduction process involves comparing the audio recorded by the recording device with the audio played by the device itself (which can also be referred to as ref audio), removing the ref audio from the audio recorded by the recording device, and obtaining the target audio. Here, ref audio refers to the reference audio, which the device generates from its own played audio.

[0075] In some implementations, the terminal needs to have recording permissions enabled. Taking a smart speaker as an example, when the smart speaker is powered on, it is in a waiting-to-wake state. The smart speaker will request microphone permissions and convey to the SDK that recording is possible. Furthermore, if microphone permissions are not enabled, the smart speaker will not be able to record audio.

[0076] In some implementations, the identification response information may include the device's response to a voice query command. For example, given the voice query command "nursery rhyme," the smart speaker responds by playing a nursery rhyme. This "playing nursery rhyme" is the smart speaker's response information.

[0077] In some implementations, the data transmission module in the speech recognition service SDK under test obtains internal noise data by communicating with the internal noise information transmission module in the player SDK, and then transmits the internal noise data to the signal processing module in the speech recognition service SDK under test; the signal processing module performs noise reduction processing based on the original audio data and the internal noise data to obtain the target audio data.

[0078] In some implementations, the identification response information is transmitted across processes to the internal noise information transmission module for processing, the response audio data generated by the internal noise data generation module in the player SDK is called, the response audio corresponding to the response audio data is played, and the response audio data is written to the terminal, which is the terminal where the test software is installed.

[0079] Thus, when internal noise is confirmed, the speech recognition service SDK under test obtains the internal noise data by communicating with the player SDK. Based on the original audio data and the internal noise data, it performs noise reduction processing to obtain the target audio data and acquires the recognition response information of the target audio data. The speech recognition service SDK then transmits the recognition response information across processes to the player SDK, which generates response audio data based on the recognition response information, plays the response audio corresponding to the response audio data, and writes the response audio data to the terminal. By processing the internal noise data separately, the impact of internal noise playback and processing on the operation of the speech recognition service SDK can be eliminated, resulting in more accurate and stable performance test data.

[0080] Figure 5 illustrates the test processing diagram when there is no internal noise. As shown in Figure 5, the test method further includes: in response to acquiring the original audio data recorded by the recording device, and assuming there is no internal noise, the speech recognition service SDK under test processes the original audio data to obtain the target audio data; the speech recognition service SDK under test acquires the recognition response information of the target audio data; the speech recognition service SDK under test transmits the recognition response information across processes to the player SDK, so that the player SDK can generate response audio data based on the recognition response information, play the response audio corresponding to the response audio data, and write the response audio data to the terminal, which is the terminal with the test software installed.

[0081] In some embodiments, the data transmission module in the speech recognition service SDK under test calls the signal processing module in the speech recognition service SDK under test to obtain the target audio data obtained by the signal processing module based on the original audio data.

[0082] Thus, assuming no internal noise is present, the tested speech recognition service SDK processes the raw audio data to obtain the target audio data. The SDK then acquires the recognition response information from the target audio data and transmits this information across processes to the player SDK. The player SDK then generates response audio data based on the recognition response information, plays the corresponding response audio, and writes the response audio data to the terminal. By determining whether internal noise exists in the raw audio and adjusting the processing method accordingly, the testing approach becomes more flexible, improving the flexibility and versatility of speech recognition SDK performance testing.

[0083] In some embodiments, the testing method includes: when the speech recognition service SDK under test can connect to the cloud service recognition device, sending target audio data to the cloud service recognition device and receiving recognition response information returned by the cloud service recognition device based on the target audio data; when the speech recognition service SDK under test cannot connect to the cloud service recognition device, sending target audio data to the local recognition engine and receiving recognition response information returned by the local recognition engine based on the target audio data.

[0084] Figure 6 illustrates the processing diagram for obtaining recognition response information. As shown in Figure 6, when the ASK SDK is detected to be connected to the network (i.e., online), it sends the target audio data to the cloud service recognition device and receives the recognition response information returned by the cloud service recognition device based on the target audio data. This allows for end-to-end verification of the ASR SDK's performance. When the ASK SDK is detected to be offline (i.e., not connected to the network), the target audio is converted into text using the local recognition engine within the ASR SDK, and response information is generated based on this text content. This allows for local verification of the ASR speech recognition SDK's performance.

[0085] This provides a universal testing method that allows for both offline and online states to coexist, enabling performance testing of voice SDKs in various environments and helping to improve the versatility and diversity of testing methods.

[0086] In some embodiments, S203 includes:

[0087] S203a: Obtain performance data of the process where the tested speech recognition service SDK resides under different scenario information;

[0088] S203b: Based on the performance data of the process where the speech recognition service SDK under test is located under different scenario information, determine the stability test data change information of the speech recognition service SDK under test under different scenario information;

[0089] S203c: Determine the performance test results of the speech recognition service SDK under test based on the stability test data change information of the SDK under test under different scenario information.

[0090] In some implementations, when the terminal is in an indoor environment with no ambient noise, the performance data of the process where the speech recognition service SDK under test is located is obtained under this scenario information; based on the performance data of the process where the speech recognition service SDK under test is located under this scenario information, the stability test data change information of the speech recognition service SDK under test under this scenario information is determined; based on the stability test data change information of the speech recognition service SDK under test under this scenario information, the performance test result of the speech recognition service SDK under test in an indoor environment with no ambient noise is determined.

[0091] In some implementations, when the terminal is in an indoor environment with ambient noise, the performance data of the process where the speech recognition service SDK under test is located is obtained under this environment information; based on the performance data of the process where the speech recognition service SDK under test is located under this environment information, the stability test data change information of the speech recognition service SDK under test under this environment information is determined; based on the stability test data change information of the speech recognition service SDK under test under this environment information, the performance test result of the speech recognition service SDK under test in an indoor environment with ambient noise is determined.

[0092] In some implementations, when the terminal is in an outdoor environment with ambient noise, the performance data of the process where the speech recognition service SDK under test is located is obtained under this environment information; based on the performance data of the process where the speech recognition service SDK under test is located under this environment information, the stability test data change information of the speech recognition service SDK under test under this environment information is determined; based on the stability test data change information of the speech recognition service SDK under test under this environment information, the performance test result of the speech recognition service SDK under test in an outdoor environment with ambient noise is determined.

[0093] In some embodiments, the overall performance test results of the tested speech recognition service SDK are obtained by statistically analyzing the performance test results of the tested speech recognition service SDK in different scenarios.

[0094] In this way, based on the performance data of the process where the tested speech recognition service SDK resides under different scenario information, the performance test results of the tested speech recognition service SDK under all scenarios can be obtained, which can significantly improve the stability and accuracy of speech SDK testing under all scenarios.

[0095] In this embodiment of the disclosure, the impact of internal noise playback and internal noise processing on the operation of the speech recognition service SDK is eliminated through cross-process communication, thereby making the performance test data more accurate and stable, solving the problem of fluctuation in test recognition performance data, and thus improving the accuracy and effectiveness of performance testing of the speech recognition service SDK.

[0096] It should be understood that the schematic diagrams shown in Figures 3, 4, 5 and 6 are merely exemplary and not restrictive, and are scalable. Those skilled in the art can make various obvious changes and / or substitutions based on the examples in Figures 3, 4, 5 and 6, and the resulting technical solutions still fall within the scope of the disclosure of the embodiments of this disclosure.

[0097] This disclosure provides a testing apparatus, as shown in FIG7. The testing apparatus may include: a running module 701, used to run testing software in response to detecting a test command. The testing software integrates a speech recognition service SDK under test and a player SDK. The speech recognition service SDK under test and the player SDK belong to different processes and can perform cross-process communication. The player SDK is used to process internal noise data; a first acquisition module 702, used to acquire performance data of the process where the speech recognition service SDK under test is located, generated during the operation of the testing software; and a determination module 703, used to determine the performance test result of the speech recognition service SDK under test based on the performance data of the process where the speech recognition service SDK under test is located.

[0098] In some embodiments, the testing apparatus may further include: a second acquisition module 704 (not shown in FIG. 7) for acquiring test scenario information; and a startup module 705 (not shown in FIG. 7) for starting the operation of the test software and the external noise control module based on the test scenario information, wherein the external noise control module is used to sense the noise status of the environment in which the terminal is located, and the terminal is a terminal on which the test software is installed.

[0099] In some embodiments, the first acquisition module 702 includes: a filtering submodule, which filters out the performance data of the process where the speech recognition service SDK is located from the performance data of all processes based on the process identifier of the process where the speech recognition service SDK is located.

[0100] In some embodiments, the testing apparatus may further include: a first processing module 706 (not shown in FIG. 7), configured to, in response to acquiring the original audio data recorded by the recording device, obtain the internal noise data by communicating with the player SDK when it is determined that internal noise exists, and perform noise reduction processing based on the original audio data and the internal noise data to obtain the target audio data; a third acquisition module 707 (not shown in FIG. 7), configured to allow the speech recognition service SDK under test to acquire the recognition response information of the target audio data; and a first writing module 708 (not shown in FIG. 7), configured to allow the speech recognition service SDK under test to transmit the recognition response information across processes to the player SDK, so that the player SDK can generate response audio data based on the recognition response information, play the response audio corresponding to the response audio data, and write the response audio data to a terminal, wherein the terminal is a terminal with the test software installed.

[0101] In some embodiments, the testing apparatus may further include: a second processing module 709 (not shown in FIG. 7), configured to, in response to acquiring the original audio data recorded by the recording device, process the speech recognition service SDK under test based on the original audio data to obtain target audio data, provided that no internal noise is determined; a third acquisition module 707 (not shown in FIG. 7), configured to allow the speech recognition service SDK under test to acquire recognition response information of the target audio data; and a second writing module 710 (not shown in FIG. 7), configured to allow the speech recognition service SDK under test to transmit the recognition response information across processes to the player SDK, so that the player SDK can generate response audio data based on the recognition response information, play the response audio corresponding to the response audio data, and write the response audio data to a terminal, wherein the terminal is a terminal with the test software installed.

[0102] In some embodiments, the third acquisition module 707 (not shown in FIG7) includes: a first acquisition submodule, configured to send target audio data to the cloud service recognition device and receive recognition response information returned by the cloud service recognition device based on the target audio data when the speech recognition service SDK under test can connect to the cloud service recognition device; and a second acquisition submodule, configured to send target audio data to the local recognition engine and receive recognition response information returned by the local recognition engine based on the target audio data when the speech recognition service SDK under test cannot connect to the cloud service recognition device.

[0103] In some embodiments, the determining module 703 includes: a third acquisition submodule, configured to acquire performance data of the process where the speech recognition service SDK under test resides under different scenario information; a first determining submodule, configured to determine the stability test data change information of the speech recognition service SDK under test under different scenario information based on the performance data of the process where the speech recognition service SDK under test resides under different scenario information; and a second determining submodule, configured to determine the performance test result of the speech recognition service SDK under test based on the stability test data change information of the speech recognition service SDK under test under different scenario information.

[0104] Those skilled in the art should understand that the functions of each processing module in the testing device of this disclosure embodiment can be understood with reference to the relevant description of the foregoing testing method. Each processing module in the testing device of this disclosure embodiment can be implemented by an analog circuit that implements the functions of this disclosure embodiment, or by running software that executes the functions of this disclosure embodiment on an electronic device.

[0105] The testing apparatus of this disclosure can improve the accuracy and effectiveness of performance testing of speech recognition service SDK.

[0106] This disclosure provides a schematic diagram of a test scenario, as shown in Figure 8.

[0107] As previously described, the testing methods provided in this disclosure are applied to electronic devices. Electronic devices are intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices.

[0108] Specifically, the electronic device may perform the following operations:

[0109] In response to the detection of the test command, the test software is run. The test software integrates the speech recognition service SDK under test and the player SDK. The speech recognition service SDK under test and the player SDK belong to different processes. The speech recognition service SDK under test and the player SDK can communicate across processes. The player SDK is used to process internal noise data.

[0110] Obtain performance data of the process containing the tested speech recognition service SDK generated during the operation of the test software;

[0111] Based on the performance data of the process where the speech recognition service SDK under test is located, the performance test results of the speech recognition service SDK under test are determined.

[0112] The performance data of the process containing the tested speech recognition service SDK can be stored on various forms of data storage devices.

[0113] The test instructions can be obtained from a data source. The data source can be various forms of data storage devices, such as laptops, desktop computers, workstations, personal digital assistants (PDAs), servers, blade servers, mainframes, and other suitable computers. The data source can also represent various forms of mobile devices, such as PDAs, cellular phones, smartphones, wearable devices, and other similar computing devices. Furthermore, the data source and the user terminal can be the same device.

[0114] It should be understood that the scene diagram shown in Figure 8 is merely illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on the examples in Figure 8, and the resulting technical solutions still fall within the scope of the disclosure of the embodiments of this disclosure.

[0115] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0116] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0117] Figure 9 illustrates a schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0118] As shown in Figure 9, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 can also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.

[0119] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0120] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as test methods. For example, in some embodiments, the test method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the test methods described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute the test method by any other suitable means (e.g., by means of firmware).

[0121] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0122] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0123] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM), flash memory, optical fiber, compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0124] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0125] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), and the Internet.

[0126] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0127] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0128] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A testing method, comprising: In response to the detection of a test command, the test software is run. The test software integrates the Software Development Kit (SDK) for the speech recognition service under test and the player SDK. The speech recognition service SDK and the player SDK belong to different processes. The speech recognition service SDK and the player SDK can communicate across processes. The player SDK is used to process internal noise data, which is the internal noise generated by the audio data played by the device itself. Obtain the performance data of the process containing the tested speech recognition service SDK generated during the operation of the test software; Based on the performance data of the process where the speech recognition service SDK under test is located, the performance test result of the speech recognition service SDK under test is determined.

2. The method according to claim 1, further comprising: Obtain test scenario information; The test software and external noise control module are started based on the test scenario information. The external noise control module is used to sense the noise status of the environment in which the terminal is located. The terminal is the terminal on which the test software is installed.

3. The method according to claim 1, wherein, The step of obtaining the performance data of the process where the speech recognition service SDK under test is located during the operation of the test software includes: filtering out the performance data of the process where the speech recognition service SDK under test is located from the performance data of all processes based on the process identifier of the process where the speech recognition service SDK under test is located.

4. The method according to claim 1, further comprising: In response to acquiring the original audio data collected by the recording device, if it is determined that there is internal noise, the speech recognition service SDK under test obtains the internal noise data by communicating with the player SDK, and performs noise reduction processing based on the original audio data and the internal noise data to obtain the target audio data; the speech recognition service SDK under test obtains the recognition response information of the target audio data; The speech recognition service SDK under test transmits the recognition response information across processes to the player SDK, so that the player SDK can generate response audio data based on the recognition response information, play the response audio corresponding to the response audio data, and write the response audio data to the terminal, which is the terminal with the test software installed.

5. The method according to claim 1, further comprising: In response to acquiring the raw audio data collected by the recording device, and assuming there is no internal noise, the speech recognition service SDK under test processes the raw audio data to obtain target audio data; the speech recognition service SDK under test acquires the recognition response information of the target audio data; The speech recognition service SDK under test transmits the recognition response information across processes to the player SDK, so that the player SDK can generate response audio data based on the recognition response information, play the response audio corresponding to the response audio data, and write the response audio data to the terminal, which is the terminal with the test software installed.

6. The method according to claim 4 or 5, wherein, The step of obtaining the recognition response information of the target audio data includes: when the speech recognition service SDK under test can connect to the cloud service recognition device, sending the target audio data to the cloud service recognition device and receiving the recognition response information returned by the cloud service recognition device based on the target audio data; when the speech recognition service SDK under test cannot connect to the cloud service recognition device, sending the target audio data to the local recognition engine and receiving the recognition response information returned by the local recognition engine based on the target audio data.

7. The method according to claim 1, wherein, The step of determining the performance test result of the speech recognition service SDK under test based on the performance data of the process where the speech recognition service SDK under test is located includes: obtaining the performance data of the process where the speech recognition service SDK under test is located under different scenario information; determining the stability test data change information of the speech recognition service SDK under test under different scenario information based on the performance data of the process where the speech recognition service SDK under test is located under different scenario information; and determining the performance test result of the speech recognition service SDK under test based on the stability test data change information of the speech recognition service SDK under test under different scenario information.

8. A testing apparatus, comprising: The running module is used to run the test software in response to the detection of the test command. The test software integrates the Software Development Kit (SDK) of the speech recognition service under test and the player SDK. The speech recognition service SDK and the player SDK belong to different processes and can communicate across processes. The player SDK is used to process internal noise data, which is the internal noise generated by the audio data played by the device itself. The first acquisition module is used to acquire the performance data of the process where the tested speech recognition service SDK is located during the operation of the test software. The determination module is used to determine the performance test result of the speech recognition service SDK under test based on the performance data of the process where the speech recognition service SDK under test is located.

9. The apparatus according to claim 8, further comprising: The second acquisition module is used to acquire test scenario information; The startup module is used to start the operation of the test software and the external noise control module based on the test scenario information. The external noise control module is used to sense the noise status of the environment in which the terminal is located. The terminal is the terminal on which the test software is installed.

10. The apparatus according to claim 8, wherein, The first acquisition module includes: a filtering submodule, which filters out the performance data of the process where the speech recognition service SDK under test is located from the performance data of all processes based on the process identifier of the process where the speech recognition service SDK under test is located.

11. The apparatus according to claim 8, further comprising: The first processing module is used to respond to the acquisition of the original audio data collected by the recording device. If it is determined that there is internal noise, the speech recognition service SDK under test obtains the internal noise data by communicating with the player SDK, and performs noise reduction processing based on the original audio data and the internal noise data to obtain the target audio data. The third acquisition module is used by the speech recognition service SDK under test to acquire the recognition response information of the target audio data; The first writing module is used by the speech recognition service SDK under test to transmit the recognition response information across processes to the player SDK, so that the player SDK can generate response audio data based on the recognition response information, play the response audio corresponding to the response audio data, and write the response audio data to the terminal, wherein the terminal is a terminal with the test software installed.

12. The apparatus according to claim 8, further comprising: The second processing module is used to respond to the acquisition of the original audio data collected by the recording device, and, if it is determined that there is no internal noise, the speech recognition service SDK under test processes the original audio data to obtain the target audio data. The third acquisition module is used by the speech recognition service SDK under test to acquire the recognition response information of the target audio data; The second writing module is used by the speech recognition service SDK under test to transmit the recognition response information across processes to the player SDK, so that the player SDK can generate response audio data based on the recognition response information, play the response audio corresponding to the response audio data, and write the response audio data to the terminal, wherein the terminal is the terminal with the test software installed.

13. The apparatus according to claim 11 or 12, wherein, The third acquisition module includes: a first acquisition submodule, configured to send the target audio data to the cloud service recognition device and receive recognition response information returned by the cloud service recognition device based on the target audio data when the speech recognition service SDK under test can connect to the cloud service recognition device; and a second acquisition submodule, configured to send the target audio data to the local recognition engine and receive recognition response information returned by the local recognition engine based on the target audio data when the speech recognition service SDK under test cannot connect to the cloud service recognition device.

14. The apparatus according to claim 8, wherein, The determining module includes: a third acquisition submodule, used to acquire performance data of the process where the speech recognition service SDK under test is located under different scenario information; a first determining submodule, used to determine the stability test data change information of the speech recognition service SDK under test under different scenario information based on the performance data of the process where the speech recognition service SDK under test is located under different scenario information; and a second determining submodule, used to determine the performance test result of the speech recognition service SDK under test based on the stability test data change information of the speech recognition service SDK under test under different scenario information.

15. An electronic device comprising: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.

17. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Performance test method, device, electronic device, and storage medium

    CN109298995A

  • Generating and attributing unique identifiers representing performance issues within a call stack

    US20210109844A1