Voice testing methods, apparatus, storage media, and computer equipment for the device
By using an automated voice testing method to obtain test requests and load dialect resource packages, and controlling the device to play target audio to execute instructions, the problem of low efficiency in manual testing is solved, and more efficient and accurate voice testing is achieved.
Patent Information
- Application Number
- CN202210139904.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-16
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-02-16
AI Technical Summary
The voice recognition testing of existing smart home products relies on manual testing, which is inefficient, easily affected by environmental noise, and its accuracy needs to be improved. In addition, the testing scenarios are not fully covered.
By obtaining test requests, determining text test cases and the dialect to be tested, loading the corresponding dialect resource package, controlling the device to play the target dialect audio, and automatically executing instructions to obtain test results, automated voice testing is achieved.
It improves the efficiency and accuracy of voice testing, reduces reliance on manual testing, and enhances test scenario coverage and noise resistance.
Smart Images

Figure CN114550694B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a voice testing method, apparatus, storage medium, and computer equipment for a device. Background Technology
[0002] For smart voice-controlled home products, the accuracy of dialect voice recognition is a crucial aspect of the product experience. Issues such as unresponsive voice recognition, ambiguous voice recognition, or voice recognition errors can significantly negatively impact the user experience. Currently, the R&D and verification process for smart voice-controlled home products typically employs manual testing. The manual testing process usually involves a user uttering the corresponding control command, the smart voice-controlled home product under test recognizing and executing the command, and the user manually confirming the correct recognition and execution of the command. This testing method has significant limitations. It heavily relies on manual testing and is constrained in terms of the number of tests, test duration, test efficiency, test scenario coverage, and test costs. It is also relatively inefficient and easily affected by environmental noise, which can influence the product's command recognition, and its accuracy needs improvement. Summary of the Invention
[0003] This application provides a method, apparatus, storage medium, and computer device for voice testing of a device, which can improve the accuracy and efficiency of voice testing of the device.
[0004] This application provides a method for testing the voice of a device, including:
[0005] Obtain the test request from the device under test, and determine the text test cases and the corresponding dialect to be tested based on the test request;
[0006] Load the dialect resource package corresponding to the dialect to be tested, wherein the dialect resource package includes text and the dialect audio of the dialect to be tested corresponding to the text;
[0007] Based on the dialect resource package, determine the target dialect audio corresponding to the text test case;
[0008] The system controls the playback of the target dialect audio to the device under test, so that the device under test can recognize the target dialect audio and execute the instructions corresponding to the target dialect audio, thereby obtaining the test results of the device under test for the target dialect.
[0009] This application embodiment also provides a voice testing device for a device, including:
[0010] The acquisition module is used to acquire the test request of the device under test, and determine the text test cases and the dialect to be tested corresponding to this test based on the test request.
[0011] The loading module is used to load the dialect resource package corresponding to the dialect to be tested, wherein the dialect resource package includes text and the dialect audio corresponding to the dialect to be tested;
[0012] The determination module is used to determine the target dialect audio corresponding to the text test case based on the dialect resource package;
[0013] The playback execution module is used to control the playback of the target dialect audio to the device under test, so that the device under test can recognize the target dialect audio and execute the instructions corresponding to the target dialect audio, so as to obtain the test result of the device under test for the target dialect.
[0014] This application also provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute the voice testing method of any of the above-described devices.
[0015] This application also provides a computer device, including a processor and a memory, wherein the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used in the steps of the voice testing method of any of the above-described devices.
[0016] The speech testing method, apparatus, storage medium, and computer equipment provided in this application determine the text test cases and the corresponding dialect to be tested based on the test request. A corresponding dialect resource package is loaded based on the dialect to be tested. The target dialect audio corresponding to the text test cases is determined based on the dialect resource package. The target dialect audio is then played to the device under test, enabling the device to recognize and execute the target dialect audio to obtain test results. This embodiment loads the corresponding dialect resource package based on the dialect to be tested in the test request, determines the target dialect audio corresponding to the text test cases, and automatically plays the target dialect audio to the device under test. This automatically realizes dialect speech testing of the device under test, improving the testing efficiency and accuracy compared to manual testing. Furthermore, determining the text test cases based on the test request, which are in text form, improves the efficiency and accuracy of test case preparation, further enhancing the testing efficiency and accuracy of the speech test of the device under test. Attached Figure Description
[0017] The following detailed description of the specific embodiments of the present application in conjunction with the accompanying drawings will make the technical solutions and other beneficial effects of the present application apparent.
[0018] Figure 1aThis is a schematic diagram illustrating an application scenario of the voice testing method for the device provided in this application embodiment.
[0019] Figure 1b This is a schematic diagram illustrating another application scenario of the voice testing method for the device provided in the embodiments of this application.
[0020] Figure 2 A flowchart illustrating the voice testing method for the device provided in this application embodiment.
[0021] Figure 3 Another flowchart illustrating the voice testing method for the device provided in this application embodiment.
[0022] Figure 4 This is another flowchart illustrating the voice testing method for the device provided in this application embodiment.
[0023] Figure 5 This is a schematic diagram illustrating the determination of whether dialect audio is distorted, provided as an embodiment of this application.
[0024] Figure 6 This is another flowchart illustrating the voice testing method for the device provided in this application embodiment.
[0025] Figure 7 This is a schematic diagram of the structure of the voice testing device for the device provided in the embodiments of this application.
[0026] Figure 8 Another structural schematic diagram of the voice testing device for the equipment provided in the embodiments of this application.
[0027] Figure 9 A schematic diagram of the structure of a computer device provided in an embodiment of this application.
[0028] Figure 10 Another structural schematic diagram of the computer device provided in the embodiments of this application. Detailed Implementation
[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0030] This application provides a method, apparatus, storage medium, and computer device for voice testing of a device. The voice testing apparatus for any device provided in this application can be integrated into a computer device, which includes devices such as terminals and / or servers. The terminal can include smartphones, tablets, wearable devices, robots, personal computers (PCs), etc. The server can be an independent physical server, a service node in a blockchain system, a server cluster consisting of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0031] like Figure 1a The diagram illustrates an application scenario provided in this application embodiment. A test engineer interacts with a computer device via voice. The computer device acquires multiple dialect audio samples of the dialect to be tested, recorded by the test engineer. These samples are then cleaned to obtain cleaned dialect audio. A dialect audio recognition method is used to identify the dialect audio, yielding the corresponding text. The text and the corresponding dialect audio are saved to obtain a dialect resource package. The test engineer also interacts with the computer device to trigger a test request. The computer device receives the test request and determines the text test cases and the corresponding dialect to be tested based on it. It loads the dialect resource package for the dialect to be tested, determines the target dialect audio corresponding to the text test cases based on the dialect resource package, and controls the playback of the target dialect audio to the device under test. This allows the device to recognize the target dialect audio and execute the corresponding instructions to obtain test results, which are then sent to the computer device.
[0032] In one embodiment, the computer device of this application may be a test server, which integrates different functional modules, such as a voice command recording module, a background noise and background human voice filtering module, a dialect resource package automatic selection module, and a data storage module. Figure 1aThe functions performed by the computer equipment are accomplished collaboratively by multiple modules, including a voice command recording module, a background noise and background human voice filtering module, a dialect resource package automatic selection module, and a data storage module (which can also be a database). It's important to note that the functions implemented by these modules can be performed on a single test server or multiple test servers.
[0033] like Figure 1b As shown, the voice command recording module communicates with the background noise and background voice filtering module, and with the dialect resource package automatic selection module. The data storage module communicates with the voice command recording module, the background noise and background voice filtering module, and the dialect resource package automatic selection module. The voice command recording module is used to interact with test engineers, such as recording dialect audio samples, providing a web interface for creating test plans and test cases, etc. The data storage module is used to save audio resource packages (dialect resource packages), text test cases, test results, error logs, and runtime logs. The background noise and background voice filtering module cleans the dialect audio samples to obtain cleaned dialect audio. Using dialect audio recognition methods, it identifies the dialect audio to obtain the corresponding text. The text and the corresponding dialect audio are saved to obtain the dialect resource package for the dialect to be tested. The dialect resource package automatic selection module loads the dialect resource package corresponding to the dialect to be tested in the test request, determines the target dialect audio corresponding to the text test case based on the dialect resource package, and controls the playback of the target dialect audio to the device under test. The test server synchronously / periodically pulls the test results from the device under test and saves them to the data storage module, while simultaneously returning the corresponding test results to the voice command recording module for display.
[0034] In one embodiment, the device under test (DUT) is replaced by a test module. The test module includes a test bench, the DUT, and a playback terminal. The DUT is connected to a computer for test result synchronization. The test bench hosts the DUT, and each area of the DUT is equipped with a playback terminal, which can be a simple host computer for playing target dialect audio, enabling machine-based voice testing instead of human voice testing. The test results of the DUT are temporarily stored on the disk of the simple host computer for the computer to synchronize / periodically retrieve and save to the data storage module.
[0035] The device to be tested can be a smart home product with voice control, such as a smart TV, smart refrigerator, smart washing machine, smart air conditioner, or smart curtains. It can also be a smart in-vehicle device, navigation system, or other device with integrated voice functionality. As long as the device has integrated voice functionality, it can be used as the device to be tested, and the voice testing method for the device in this application embodiment can be used to perform dialect voice testing.
[0036] The voice testing method, apparatus, computer-readable storage medium, and computer device of the device in the embodiments of this application will be described in detail below. It should be noted that the sequence numbers of the following embodiments are not intended to limit the preferred order of the embodiments.
[0037] Please refer to Figure 2 This is a flowchart illustrating the voice testing method for the device provided in an embodiment of this application. The voice testing method for this device is applied in computer devices, such as... Figure 1a As shown, the voice testing method for this device includes the following steps.
[0038] 101. Obtain the test request from the device under test. Based on the test request, determine the text test cases for this test, whether dialect testing is required for this test, and the corresponding dialect to be tested if dialect testing is required for this test.
[0039] Computer devices can be used to write and store test suites. Before testing the device under test, test engineers need to write complete test scenarios (test cases) covering basic functionalities such as various dialects, slow speech rates, fast speech rates, quiet environments, normal environments, noisy environments, pauses, and multiple commands. They also need to include abnormal test scenarios such as long sentences, short sentences, and multiple speakers, as well as test cases for each scenario. Test engineers can write debugging test suites, each containing test cases for the corresponding test scenarios, and save these suites to the computer device. The computer device then retrieves the test suites, which include test cases for various dialects, slow speech rates, fast speech rates, quiet environments, normal environments, noisy environments, pauses, and multiple commands, and can also include test cases for scenarios with long sentences, short sentences, and multiple speakers.
[0040] Computer equipment can also be used to create and distribute test plans. Test plans are typically matched to the test phase, the device under test (test product), and the specific modifications required for the device during this test. Test cases are selected from the test suite based on different test objectives. The computer equipment receives the selected test cases and creates a set of test cases, forming a test plan for this test.
[0041] When the test plan is triggered, such as by triggering a test plan control like the "Execute" button, by voice command, or by command, a test request for the device under test is generated. Alternatively, when the test plan reaches its corresponding execution time, a test request for the device under test is generated based on the test plan. The computer device obtains the test request from the device under test and retrieves the set of test cases from the corresponding test plan. This set of test cases includes multiple text test cases. The test plan also includes a data item for "Does dialect testing need to be performed?" for the text test cases in this test. If this data item is empty or absent, it means that dialect testing is not required in this test; otherwise, it means that dialect testing is required. The test plan also includes a data item for the dialect to be tested corresponding to the text test cases in this test. If this data item is not empty, the dialect to be tested is retrieved. For example, if the data item is "Chinese-Hakka", then the dialect to be tested is Hakka.
[0042] That is, the computer device obtains the test request from the device under test, and determines the text test cases for this test, whether dialect testing is required for this test, and the corresponding dialect to be tested if dialect testing is required for this test.
[0043] 102. If this test requires a dialect test, load the dialect resource package corresponding to the dialect to be tested. The dialect resource package includes text and the dialect audio of the dialect to be tested corresponding to the text.
[0044] The system queries the saved audio resource packages to find the dialect corresponding to the dialect to be tested. If the query is successful, the dialect resource package is loaded. If the query fails, a prompt is displayed and the test step is marked as abnormal.
[0045] This dialect resource package includes multiple texts and their corresponding audio recordings in the dialect to be tested. The dialect resource package is obtained and saved in advance, for example, on a computer device or in a data storage module.
[0046] The method for obtaining dialect resource packs will be described in detail below, and will not be described in detail here.
[0047] 103. Based on the dialect resource package, determine the target dialect audio corresponding to the text test cases.
[0048] Since the dialect resource package includes multiple texts and their corresponding audio recordings in the dialects to be tested, the texts in the dialect resource package are matched against the text test cases. The audio recordings of the corresponding dialects in the dialect resource package are then used as the target dialect audio recordings for each text test case. This process is repeated to determine the target dialect audio recordings for each text test case.
[0049] 104. Control the playback of target dialect audio to the device under test, so that the device under test can recognize the target dialect audio and execute the instructions corresponding to the target dialect audio, so as to obtain the test results of the device under test for the target dialect.
[0050] After obtaining the target dialect audio corresponding to each text test case, a playback command is generated, and the target dialect audio is played to the device under test according to the playback command. This can be done by generating a single playback command after obtaining the target dialect audio corresponding to one text test case, or by generating playback commands after obtaining the target dialect audio corresponding to all text test cases, and then playing the target dialect audio corresponding to multiple text test cases sequentially to the device under test.
[0051] exist Figure 1a In the application scenario shown, the device under test and the computer can be in the same environment, which facilitates playing target dialect audio from the computer to the device under test. Figure 1b In the application scenario shown, the device under test and the playback terminal can be in the same environment. After the test server obtains the target dialect audio, it sends a playback command to the playback terminal and sends the corresponding target dialect audio to the playback terminal so that the playback terminal can play the target dialect audio to the device under test.
[0052] The device under test identifies the target dialect audio and executes the corresponding instructions to obtain the test results for the target dialect. Figure 1a In the application scenario shown, the device under test sends the test results for the dialect under test to the computer device. Figure 1b In the application scenario shown, the device under test will temporarily store the test results for the dialect under test in the playback terminal, so that the computer device can synchronously / periodically retrieve the test results and save them to the data storage module.
[0053] In this embodiment, when dialect testing is required, the corresponding dialect resource package is loaded according to the dialect to be tested in the test request to determine the target dialect audio corresponding to the text test case. The target dialect audio is then automatically played to the device under test. In this way, dialect speech testing of the device under test is automatically realized. Compared with manual testing, this improves the testing efficiency and accuracy of dialect speech testing of the device under test. Moreover, determining the text test case according to the test request, which is in text form (compared to using voice test cases, which are prone to omissions), improves the efficiency and accuracy of test case preparation, further enhancing the testing efficiency and accuracy of dialect speech testing of the device under test.
[0054] 105. If dialect testing is not required for this test, load the general audio resource package, which includes the text and the general audio corresponding to the text.
[0055] The general audio resource package can include Mandarin audio resource packages corresponding to Chinese text, or standard English audio resource packages corresponding to English text, depending on the specific use case. For example, in a Chinese testing scenario, the general audio resource package is a Mandarin audio resource package, which includes Chinese text and its corresponding Mandarin audio (general audio). In an English testing scenario, the general audio resource package is a standard English audio resource package, which includes English text and its corresponding standard English audio (general audio).
[0056] 106. Based on the general audio resource package, determine the target general audio corresponding to the text test cases.
[0057] The text test cases are matched against text in a generic audio resource package, and the generic audio corresponding to the text in the generic audio resource package is used as the target generic audio for each text test case. This process is repeated to determine the target generic audio for each text test case.
[0058] 107. Control the playback of the target general audio to the device under test, so that the device under test can recognize the target general audio and execute the instructions corresponding to the target general audio, so as to obtain the test results of the device under test for the general audio.
[0059] The specific implementation steps are the same as in step 104, and you can refer to the description in step 104 for details.
[0060] In this embodiment, when dialect testing is required for the device under test, a corresponding dialect resource package is loaded based on the dialect to be tested in the test request. This determines the target dialect audio corresponding to the text test cases, and the target dialect audio is automatically played to the device under test. When dialect testing is not required for the current test, a general audio resource package is loaded to determine the target general audio corresponding to the text test cases, and the target general audio is automatically played to the device under test. This automatically performs voice testing on the device under test, improving the efficiency and accuracy of voice testing compared to manual testing. Furthermore, determining the text test cases based on the test request, in text format, improves the efficiency and accuracy of test case preparation, further enhancing the efficiency and accuracy of voice testing on the device under test.
[0061] Since the embodiments in this application mainly target the application scenario of dialect speech testing, the following embodiments will mainly use dialect speech testing as an example for explanation. The implementation steps of general audio testing and dialect speech testing are the same, and will not be repeated here.
[0062] Figure 3 This is another schematic flowchart of the voice testing method for a device provided in this application embodiment. The voice testing method for the device is applied to a computer device and includes the following steps.
[0063] 201. Obtain multiple recorded audio samples of the dialect to be tested.
[0064] The recorded audio samples of the dialect to be tested include audio samples under basic functional test points such as slow speech speed, fast speech speed, quiet environment, normal environment, noisy environment, sentence pauses, and multiple instructions, as well as audio samples under abnormal functional test points such as long sentences, short sentences, and multiple people speaking.
[0065] 202. The dialect audio samples are cleaned to obtain cleaned dialect audio.
[0066] Since the dialect audio samples are audio samples under different functional test points, it is necessary to clean the dialect audio samples.
[0067] In one embodiment, step 202 specifically includes: performing background noise reduction processing on the dialect audio sample to obtain processed dialect audio; using Mel-frequency cepstral coefficients to filter background human voices from the dialect audio to obtain filtered dialect audio. The filtered dialect audio is then used as the cleaned dialect audio.
[0068] Background noise reduction for dialect audio samples can be performed using the MMSE (Minimum Mean-Square Error) noise reduction algorithm or an improved version of the MMSE algorithm, or other audio noise reduction methods. Background voice filtering can also be achieved using other methods.
[0069] In one embodiment, the step of performing background noise reduction processing on dialect audio samples to obtain processed dialect audio includes: performing linear processing on the dialect audio samples, sorting the linearly processed dialect audio according to signal energy, and performing serial interference cancellation operation.
[0070] Linear processing involves partial decorrelation operations. The decorrelation operation, mathematically described as y = hx + n, involves eliminating h by multiplying both sides of the equation by the inverse matrix of h, which is obtained from channel estimation. Serial interference cancellation is necessary because the multi-channel nature of MIMO results in the receiver receiving signals from multiple transmitters. If the transmitters are sending the same data stream, the received data is correlated, and these received signals can be used for signal reliability recovery (though capacity improvement is not very significant). If the transmitters are not sending the same data stream (typically, the data streams are mixed, as seen in various space-time coding structures), multiple data streams will overlap at the receiver. Therefore, separating these data streams and restoring them independently becomes crucial, which is the purpose of interference cancellation. Interference removal: First, the received user signals are ranked by power strength. Only one user is detected at a time, and the user with the strongest power is demodulated first. Then, the interference from the reconstructed strongest user is subtracted from the total received signal. Next, the second strongest interference is reconstructed and canceled, and so on.
[0071] In one embodiment, the step of using Mel frequency cepstral coefficients to filter background human voice in dialect audio to obtain filtered dialect audio includes: splitting the dialect audio into frames, calculating the difference using short-time FFT, the dialect audio being divided into many frames, each frame of speech corresponding to a spectrum, which identifies the relationship between frequency and energy; then performing cepstral analysis to obtain peaks (representing the main frequency components of the dialect speech; i.e., the most relevant part of the human voice), these peaks are called formants, and the formants carry the sound identification attributes; then extracting the spectrum envelope based on the formants, this envelope is... It is a smooth curve connecting these resonant peaks (given logX[k], logH[k] and logE[k] are obtained, satisfying logX[k]=logH[k]+logE[k]). The envelope mainly consists of low-frequency components, while the high-frequency part mainly consists of spectral details. Superimposing the low-frequency and high-frequency parts gives the original spectral signal. That is, h[k] is the low-frequency part of x[k]. Therefore, passing x[k] through a low-pass filter will give h[k], which is the envelope of the spectrum. Finally, the envelope in the spectrum is removed, and the high frequency is retained, which gives the filtered dialect audio. The filtered dialect audio includes the recognized human voice.
[0072] 203. Using dialect audio recognition methods, dialect audio is recognized to obtain the corresponding text. The text and the corresponding dialect audio are saved to obtain a dialect resource package for the dialect to be tested.
[0073] The filtered dialect audio is just a segment of human voice; the text behind the voice needs to be understood and translated.
[0074] Dialect audio recognition methods are used to identify dialect audio and obtain the corresponding text. In one embodiment, a dialect recognition method can also be used to read dialect audio from a large database of dialects online, automatically identifying the dialect audio and obtaining the corresponding text. Multiple texts corresponding to multiple dialect audios are identified, and the multiple texts and their corresponding dialect audios are grouped into multiple lines of records, which are then saved to a computer device to form a dialect resource package for the dialect to be tested. Each text and its corresponding dialect audio form a single line of records. Thus, dialect audio recognition methods can be used to identify any dialect audio and obtain a dialect resource package corresponding to any dialect.
[0075] The method uses dialect audio recognition to identify dialect audio. If the recognition is successful, the successfully identified text and the corresponding dialect audio are saved. If the recognition fails, the failed text is discarded and the recognition log is updated.
[0076] 204. Obtain the test request from the device under test. Based on the test request, determine the text test cases for this test, whether dialect testing is required for this test, and the corresponding dialect to be tested if dialect testing is required for this test.
[0077] 205. If this test requires a dialect test, load the dialect resource package corresponding to the dialect to be tested. The dialect resource package includes text and the dialect audio of the dialect to be tested corresponding to the text.
[0078] 206. Based on the dialect resource package, determine the target dialect audio corresponding to the text test cases.
[0079] 207. Control the playback of target dialect audio to the device under test, so that the device under test can recognize the target dialect audio and execute the instructions corresponding to the target dialect audio, so as to obtain the test results of the device under test for the target dialect.
[0080] This embodiment further details how to obtain the dialect resource package. This embodiment does not simply save the recorded dialect audio samples as a dialect resource package; it also requires cleaning and recognition processing of the dialect audio samples to obtain the dialect audio text. The dialect resource package obtained in this embodiment includes text and the corresponding dialect audio, so that text test cases can be used in test requests to match the text test cases in the test requests using the text in the dialect resource package. Furthermore, the dialect audio samples are cleaned to improve the accuracy of the dialect resource package.
[0081] Figure 4 This is another schematic flowchart of the voice testing method for a device provided in this application embodiment. The voice testing method for this device is applied to a computer device and includes the following steps.
[0082] 301, Obtain multiple recorded audio samples of the dialect to be tested.
[0083] 302. Copy the dialect audio sample to obtain two dialect audio samples.
[0084] Multiple dialect audio samples were copied to obtain two dialect audio samples.
[0085] 303. One of the dialect audio samples is cleaned to obtain the cleaned dialect audio. For specific cleaning methods, please refer to the description above; they will not be repeated here.
[0086] 304. Compare the dialect audio with another dialect audio sample to determine whether the dialect audio is distorted.
[0087] The cleaned dialect audio is compared with another copied dialect audio sample to determine whether the dialect audio is distorted. If it is not distorted, the next step is performed. In one embodiment, step 304 above includes: obtaining the audio value of each frame in the dialect audio and the other dialect audio sample; determining a first average value of the audio value of each frame in the dialect audio and a second average value of the audio value of each frame in the other dialect audio sample; comparing the first average value and the second average value corresponding to each frame to obtain a comparison result; if the proportion of the comparison result between the dialect audio and the other dialect audio sample that exceeds a preset threshold exceeds a preset proportion, the dialect audio is determined to be distorted; otherwise, the dialect audio is determined not to be distorted.
[0088] The comparison process involves comparing the first average value and the second average value corresponding to each frame to obtain a comparison result. This includes subtracting the second average value from the first average value for each frame to obtain a difference, which is then used as the comparison result. The number of frames with differences exceeding a preset threshold is counted. If the ratio of this number to the total number of frames exceeds a preset ratio, the dialect audio is determined to be distorted; otherwise, the dialect audio is determined not to be distorted.
[0089] Please combine Figure 5 Let's understand steps 301 to 304 above.
[0090] If the dialect audio is distorted, proceed to step 305; if the dialect audio is not distorted, proceed to step 306.
[0091] 305, Discard the dialect audio. Discard the distorted dialect audio and update the recognition log.
[0092] 306. Using dialect audio recognition methods, dialect audio is recognized to obtain the corresponding text. The text and the corresponding dialect audio are saved to obtain a dialect resource package for the dialect to be tested.
[0093] 307. Obtain the test request from the device under test. Based on the test request, determine the text test cases for this test, whether dialect testing is required for this test, and the corresponding dialect to be tested if dialect testing is required for this test.
[0094] 308. If this test requires a dialect test, load the dialect resource package corresponding to the dialect to be tested. The dialect resource package includes text and the dialect audio of the dialect to be tested corresponding to the text.
[0095] 309. Based on the dialect resource package, determine the target dialect audio corresponding to the text test cases.
[0096] 310, Control the playback of target dialect audio to the device under test, so that the device under test can recognize the target dialect audio and execute the instructions corresponding to the target dialect audio, so as to obtain the test results of the device under test for the target dialect.
[0097] In this embodiment, the distortion of the dialect audio after cleaning is further determined to improve the accuracy of the dialect resource package.
[0098] In one embodiment, before obtaining the test request from the device under test, the method further includes: obtaining multiple recorded general audio samples, cleaning the general audio samples to obtain cleaned general audio; using a general audio recognition method to recognize the general audio to obtain the text corresponding to the general audio, and saving the text and the corresponding general audio to obtain a general audio resource package of the general audio.
[0099] Figure 6 This is a schematic flowchart of a voice testing method for a device provided in an embodiment of this application. The voice testing method for this device is applied to a computer device and includes the following steps.
[0100] 401, retrieves multiple recorded generic audio samples.
[0101] The recorded audio samples include those for basic functional test points such as slow speech rate, fast speech rate, quiet environment, normal environment, noisy environment, sentence pauses, and multiple commands, as well as those for abnormal functional test points such as long sentences, short sentences, and multiple people speaking. The general audio samples refer to Mandarin audio samples in Chinese test scenarios or standard English audio samples in English test scenarios.
[0102] 402, Copy the generic audio sample to obtain two generic audio samples.
[0103] 403. One of the general audio samples is cleaned to obtain the cleaned dialect audio. The specific cleaning method is described above and will not be repeated here.
[0104] 404. Compare the generic audio with another generic audio sample to determine if the generic audio is distorted. For the steps to determine if the generic audio is distorted, please refer to the steps for determining if dialect audio is distorted above; they will not be repeated here.
[0105] If the general audio is distorted, proceed to step 405; if the general audio is not distorted, proceed to step 406.
[0106] 405, Discard the generic audio. Discard the distorted generic audio and update the recognition log.
[0107] 406. Using a general audio recognition method, the general audio is recognized to obtain the text corresponding to the general audio. The text and the corresponding general audio are saved to obtain a general audio resource package of the general audio.
[0108] 407. Obtain the test request from the device under test. Based on the test request, determine the text test cases for this test, whether dialect testing is required for this test, and the corresponding dialect to be tested if dialect testing is required for this test.
[0109] 408. If dialect testing is not required for this test, load the general audio resource package, which includes text and the corresponding general audio.
[0110] 409. Based on the general audio resource package, determine the target general audio corresponding to the text test cases.
[0111] 410, Control the playback of the target general audio to the device under test, so that the device under test can recognize the target general audio and execute the instructions corresponding to the target general audio, so as to obtain the test results of the device under test for the general audio.
[0112] In one embodiment, when the test results of the device under test for the dialect under test or for general audio are obtained, the voice testing method of the device further includes: saving the test results of the device under test for the dialect under test; converting the test results into a preset format according to the configuration file to obtain a test report; and displaying the test report.
[0113] The test results cannot be directly displayed or the display effect is unsatisfactory. Therefore, the test results are converted into a preset format according to the configuration file for easy display. Furthermore, the test results need to be further summarized and then converted into the preset format according to the configuration file to obtain a test report.
[0114] In one embodiment, the test report can also be sent to the email address of the corresponding test engineer or relevant personnel via email or other means.
[0115] In the voice testing method for the device provided in this application embodiment, in the early stage of testing, it is necessary to prepare the recording scenario, obtain audio samples in the corresponding scenario, store the audio samples, and continuously update and maintain the audio samples. During the testing phase, test cases are created, a test plan is formulated, the test plan is executed (manually / timed, etc.), and test cases are automatically run while awaiting result feedback. In the test closing stage, test results are summarized, a test report is generated, the test report is automatically saved, and the test includes automatically sending emails, etc.
[0116] The above embodiments automatically perform voice testing on the device under test, which improves the testing efficiency and accuracy of voice testing compared to manual testing. Furthermore, the determination of text test cases based on test requests, which are in text form, improves the efficiency and accuracy of test case preparation, further enhancing the testing efficiency and accuracy of voice testing on the device under test.
[0117] Based on the method described in the above embodiments, this embodiment will further describe it from the perspective of the voice testing device of the device. The voice testing device of the device can be implemented as an independent entity or integrated into a computer device.
[0118] Please see Figure 7 , Figure 7 This application provides a specific description of a voice testing apparatus for a device applied in a computer device. The voice testing apparatus may include: an acquisition module 501, a loading module 502, a determination module 503, and a playback execution module 504.
[0119] The acquisition module 501 is used to acquire the test request of the device under test, and determine the text test cases and the dialect to be tested corresponding to this test based on the test request.
[0120] The loading module 502 is used to load the dialect resource package corresponding to the dialect to be tested, wherein the dialect resource package includes text and the dialect audio corresponding to the dialect to be tested;
[0121] The determination module 503 is used to determine the target dialect audio corresponding to the text test case based on the dialect resource package;
[0122] The playback execution module 504 is used to control the playback of the target dialect audio to the device under test, so that the device under test can recognize the target dialect audio and execute the instructions corresponding to the target dialect audio, so as to obtain the test result of the device under test for the target dialect.
[0123] In one embodiment, the loading module 502 is further configured to load a general audio resource package if dialect testing is not required in this test, the general audio resource package including text and general audio corresponding to the text; the determining module 503 is further configured to determine the target general audio corresponding to the text test case based on the general audio resource package; the playback execution module 504 is further configured to control the playback of the target general audio to the device under test, so that the device under test can identify the target general audio and execute the instructions corresponding to the target general audio, so as to obtain the test result of the device under test for the general audio.
[0124] In one embodiment, such as Figure 8 As shown, the voice testing device of the equipment also includes a cleaning module 505, a recognition module 506, and a storage module 507. Specifically, the acquisition module 501 is used to acquire multiple recorded dialect audio samples of the dialect to be tested; the cleaning module 505 is used to clean the dialect audio samples to obtain cleaned dialect audio; the recognition module 506 is used to recognize the dialect audio using a dialect audio recognition method to obtain the text corresponding to the dialect audio; and the storage module 507 is used to save the text and the corresponding dialect audio to obtain a dialect resource package of the dialect to be tested.
[0125] In one embodiment, the cleaning module 505 is specifically used to perform background noise reduction processing on the dialect audio sample to obtain processed dialect audio; and to perform background voice filtering on the dialect audio using Mel frequency cepstral coefficients to obtain filtered dialect audio.
[0126] In one embodiment, such as Figure 8 As shown, the voice testing device of the equipment also includes a copying module 508 and a distortion comparison module 509. Specifically, before performing the cleaning process on the dialect audio sample, the copying module 508 copies the dialect audio sample to obtain two copies; the cleaning module 505 specifically cleans one of the dialect audio samples to obtain the cleaned dialect audio; the distortion comparison module 509 compares the dialect audio with the other dialect audio sample to determine whether the dialect audio is distorted; and the recognition module 506, if the dialect audio is not distorted, uses a dialect audio recognition method to recognize the dialect audio to obtain the text corresponding to the dialect audio.
[0127] In one embodiment, the distortion comparison module 509 is specifically used to acquire the audio value of each frame in the dialect audio and the other dialect audio sample; determine a first average value of the audio value of each frame in the dialect audio and a second average value of the audio value of each frame in the other dialect audio sample; compare the first average value and the second average value corresponding to each frame to obtain a comparison result; if the proportion of the comparison result in the dialect audio and the other dialect audio sample that exceeds a preset threshold exceeds a preset proportion, then the dialect audio is determined to be distorted; otherwise, the dialect audio is determined not to be distorted.
[0128] In one embodiment, such as Figure 8As shown, the voice testing device of the equipment also includes a conversion and display module 510. The storage module 507 is further used to store the test results of the device under test for the dialect under test. The conversion and display module 510 is used to convert the test results into a preset format according to a configuration file to obtain a test report; and to display the test report.
[0129] In practice, the above modules can be implemented as independent entities or combined arbitrarily as the same or several entities. For the specific implementation of the above modules, please refer to the previous method embodiments. For the specific beneficial effects that can be achieved, please also refer to the beneficial effects in the previous method embodiments, which will not be repeated here.
[0130] In addition, embodiments of this application also provide a computer device, such as... Figure 9 As shown, the computer device 600 includes a processor 601 and a memory 602. The processor 601 and the memory 602 are electrically connected.
[0131] The processor 601 is the control center of the computer device 600. It connects various parts of the computer device through various interfaces and lines. By running or loading applications stored in the memory 602 and calling data stored in the memory 602, it performs various functions of the computer device and processes data, thereby monitoring the computer device as a whole.
[0132] In this embodiment, the processor 601 in the computer device 600 loads the instructions corresponding to the processes of one or more application programs into the memory 602 according to the following steps, and the processor 601 runs the application programs stored in the memory 602 to realize various functions, such as:
[0133] The system obtains a test request from the device under test, determines the text test cases and the corresponding dialect for this test based on the test request, loads the dialect resource package corresponding to the dialect under test, the dialect resource package includes text and the dialect audio corresponding to the text in the dialect under test, determines the target dialect audio corresponding to the text test cases based on the dialect resource package, and controls the playback of the target dialect audio to the device under test, so that the device under test can recognize the target dialect audio and execute the instructions corresponding to the target dialect audio, thereby obtaining the test results of the device under test for the dialect under test.
[0134] This computer device can implement the steps of any embodiment of the voice testing method for the device provided in the embodiments of this application. Therefore, it can achieve the beneficial effects that the voice testing method for any device provided in the embodiments of this invention can achieve. For details, please refer to the previous embodiments, which will not be repeated here.
[0135] Figure 10 A specific structural block diagram of a computer device provided in an embodiment of the present invention is shown. This computer device can be used to implement the voice testing method of the device provided in the above embodiments. The computer device includes the following modules / units.
[0136] RF circuit 710 is used to receive and transmit electromagnetic waves, converting electromagnetic waves into electrical signals and vice versa, thereby enabling communication with communication networks or other devices. RF circuit 710 may include various existing circuit elements used to perform these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, Subscriber Identity Module (SIM) cards, memory, etc. RF circuit 710 can communicate with various networks such as the Internet, corporate intranets, and wireless networks, or communicate with other devices via wireless networks. The aforementioned wireless networks may include cellular telephone networks, wireless local area networks (WLANs), or metropolitan area networks (MANs). The aforementioned wireless networks may use various communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging, and short messages, and any other suitable communication protocols, including those that have not yet been developed.
[0137] The memory 720 can be used to store software programs (computer programs) and modules, such as the program instructions / modules corresponding to those in the above embodiments. The processor 780 executes various functional applications and data processing by running the software programs and modules stored in the memory 720. The memory 720 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 720 may further include memory remotely located relative to the processor 780, and these remote memories can be connected to the computer device 700 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0138] The input unit 730 can be used to receive input digital or character information, and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. Specifically, the input unit 730 may include a touch-sensitive surface 731 and other input devices 732. The touch-sensitive surface 731, also known as a touch display screen (touchscreen) or touchpad, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch-sensitive surface 731), and drive the corresponding connection device according to a pre-set program. Optionally, the touch-sensitive surface 731 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 780, and can also receive and execute commands sent by the processor 780. In addition, the touch-sensitive surface 731 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 731, the input unit 730 may also include other input devices 732. Specifically, other input devices 732 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0139] Display unit 740 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of computer device 700. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Display unit 740 may include display panel 741, optionally configured as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), or other similar forms. Further, touch-sensitive surface 731 may cover display panel 741. When touch-sensitive surface 731 detects a touch operation on or near it, it transmits the information to processor 780 to determine the type of touch event. Subsequently, processor 780 provides corresponding visual output on display panel 741 according to the type of touch event. Although in the figures, touch-sensitive surface 731 and display panel 741 are implemented as two separate components to achieve input and output functions, it is understood that touch-sensitive surface 731 and display panel 741 can be integrated to achieve input and output functions.
[0140] The computer device 700 may also include at least one sensor 750, such as a light sensor, an orientation sensor, a proximity sensor, and other sensors. As a type of motion sensor, a gravity acceleration sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometers, taps), etc. Other sensors that the computer device 700 may also be equipped with, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0141] Audio circuitry 760, speaker 761, and microphone 762 provide an audio interface between the user and computer device 700. Audio circuitry 760 converts received audio data into electrical signals, which are then transmitted to speaker 761, where they are converted into sound signals for output. Conversely, microphone 762 collects sound signals, converts them into electrical signals, which are received by audio circuitry 760, converted back into audio data, and then processed by processor 780 before being transmitted via RF circuitry 710 to, for example, another computer device, or output to memory 720 for further processing. Audio circuitry 760 may also include an earphone jack to facilitate communication between peripheral headphones and computer device 700.
[0142] Computer device 700, through transmission module 770 (e.g., Wi-Fi module), can help users receive requests, send information, etc., providing users with wireless broadband internet access. Although transmission module 770 is shown in the figure, it is understood that it is not an essential component of computer device 700 and can be omitted as needed without changing the essence of the invention.
[0143] The processor 780 is the control center of the computer device 700. It connects to various parts of the mobile phone via various interfaces and lines. By running or executing software programs (computer programs) and / or modules stored in the memory 720, and by calling data stored in the memory 720, it performs various functions of the computer device 700 and processes data, thereby providing overall monitoring of the computer device. Optionally, the processor 780 may include one or more processing cores; in some embodiments, the processor 780 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 780.
[0144] The computer device 700 also includes a power supply 790 (such as a battery) that supplies power to various components. In some embodiments, the power supply may be logically connected to the processor 780 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 790 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0145] Although not shown, the computer device 700 also includes cameras (such as front-facing cameras and rear-facing cameras), Bluetooth modules, etc., which will not be described in detail here. Specifically, in this embodiment, the display unit of the computer device is a touch screen display, and the computer device also includes a memory and one or more programs (computer programs), one or more of which are stored in the memory and configured to be executed by one or more processors. One or more programs contain instructions for performing the voice testing method of the above-described device in any embodiment. For the beneficial effects achieved, please refer to the beneficial effects achieved in any embodiment of the voice testing method of the above-described device.
[0146] In practice, the above modules can be implemented as independent entities or combined in any way to be implemented as the same or several entities. For the specific implementation of the above modules, please refer to the previous method implementation examples, which will not be repeated here.
[0147] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by instructions (computer programs) or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. Therefore, embodiments of the present invention provide a storage medium storing multiple instructions that can be loaded by a processor to execute the steps of any embodiment of the voice testing method of the device provided in the embodiments of the present invention.
[0148] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0149] Since the instructions stored in the storage medium can execute the steps in any embodiment of the voice testing method of the device provided in the embodiments of the present invention, the beneficial effects that the voice testing method of any device provided in the embodiments of the present invention can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.
[0150] The above provides a detailed description of a voice testing method, apparatus, storage medium, and computer device for a device according to embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for testing the voice of a device, characterized in that, include: Obtain the test request from the device under test, and determine the text test cases and the corresponding dialect to be tested based on the test request; Load the dialect resource package corresponding to the dialect to be tested, wherein the dialect resource package includes text and the dialect audio of the dialect to be tested corresponding to the text; Based on the dialect resource package, determine the target dialect audio corresponding to the text test case; The system controls the playback of the target dialect audio to the device under test, so that the device under test can recognize the target dialect audio and execute the instructions corresponding to the target dialect audio, thereby obtaining the test results of the device under test for the target dialect. Obtaining the dialect resource package includes: Obtain multiple dialect audio samples of the dialect to be tested; Duplicate the dialect audio sample to obtain two copies of the dialect audio sample; One of the dialect audio samples was cleaned to obtain the cleaned dialect audio, including: Obtain the audio value of each frame from the dialect audio and another dialect audio sample; Determine a first average value of the audio values of each frame of the dialect audio, and determine a second average value of the audio values of each frame of the other dialect audio sample; The first average value and the second average value corresponding to each frame are compared to obtain a comparison result. If the proportion of the comparison result between the dialect audio and the other dialect audio sample exceeds a preset threshold, the dialect audio is determined to be distorted; otherwise, the dialect audio is determined to be undistorted. The comparison result includes: subtracting the second average value from the first average value corresponding to each frame to obtain a difference, and using the difference as the comparison result. If the dialect audio is not distorted, a dialect audio recognition method is used to recognize the dialect audio to obtain the text corresponding to the dialect audio. The text and the corresponding dialect audio are saved to obtain the dialect resource package of the dialect to be tested.
2. The voice testing method for the device according to claim 1, characterized in that, The step of cleaning the dialect audio samples to obtain cleaned dialect audio includes: The dialect audio samples are subjected to background noise reduction processing to obtain the processed dialect audio. The dialect audio is filtered for background human voice using Mel frequency cepstral coefficients to obtain the filtered dialect audio.
3. The voice testing method for the device according to claim 1, characterized in that, After the step of obtaining the test results of the device under test for the dialect under test, the method further includes: Save the test results of the device under test for the dialect under test; The test results are converted into a preset format according to the configuration file to obtain a test report; The test report is displayed.
4. The voice testing method for the device according to claim 1, characterized in that, Also includes: If dialect testing is not required for this test, load the general audio resource package, which includes the text and the general audio corresponding to the text; Based on the general audio resource package, determine the target general audio corresponding to the text test case; The system controls the playback of the target general audio to the device under test, so that the device under test can recognize the target general audio and execute the instructions corresponding to the target general audio, thereby obtaining the test results of the device under test for the general audio.
5. A voice testing device for an instrument, characterized in that, include: The acquisition module is used to acquire the test request of the device under test, and determine the text test cases and the dialect to be tested corresponding to this test based on the test request. The loading module is used to load the dialect resource package corresponding to the dialect to be tested, wherein the dialect resource package includes text and the dialect audio corresponding to the dialect to be tested; The determination module is used to determine the target dialect audio corresponding to the text test case based on the dialect resource package; The playback execution module is used to control the playback of the target dialect audio to the device under test, so that the device under test can recognize the target dialect audio and execute the instructions corresponding to the target dialect audio, so as to obtain the test result of the device under test for the target dialect; Obtaining the dialect resource package includes: Obtain multiple dialect audio samples of the dialect to be tested; Duplicate the dialect audio sample to obtain two copies of the dialect audio sample; One of the dialect audio samples was cleaned to obtain the cleaned dialect audio, including: Obtain the audio value of each frame from the dialect audio and another dialect audio sample; Determine a first average value of the audio values of each frame of the dialect audio, and determine a second average value of the audio values of each frame of the other dialect audio sample; The first average value and the second average value corresponding to each frame are compared to obtain a comparison result. If the proportion of the comparison result between the dialect audio and the other dialect audio sample exceeds a preset threshold, the dialect audio is determined to be distorted; otherwise, the dialect audio is determined to be undistorted. The comparison result includes: subtracting the second average value from the first average value corresponding to each frame to obtain a difference, and using the difference as the comparison result. If the dialect audio is not distorted, a dialect audio recognition method is used to recognize the dialect audio to obtain the text corresponding to the dialect audio. The text and the corresponding dialect audio are saved to obtain the dialect resource package of the dialect to be tested.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform steps in the voice testing method of the device according to any one of claims 1 to 4.
7. A computer device, characterized in that, The device includes a processor and a memory, the processor being electrically connected to the memory, the memory being used to store instructions and data, and the processor being used to execute the steps of the voice testing method of the device according to any one of claims 1 to 4.
Citation Information
Patent Citations
Audio processing method and device
CN112185410A
Speech synthesis effect evaluation method and device, computer equipment and storage medium
CN112669810A
Equipment performance test method and device, storage medium and electronic device
CN113595811A