Voice interface test method and device, electronic equipment and storage medium
Through the automated voice interface testing method, the problems of low efficiency and poor accuracy of traditional testing methods are solved, and more efficient and accurate test results are achieved.
Patent Information
- Application Number
- CN202510224539.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional voice interface testing methods rely on manual operations, resulting in low testing efficiency and poor accuracy.
Provide a voice interface testing method, which calculates the success rate and accuracy of the corpus recognition and analysis of the corpus to be tested by obtaining the corpus to be tested and the number of test times.
It improves testing efficiency and accuracy, reduces manual interference, and can intuitively reflect the voice resolution capabilities of the voice interface.
Smart Images

Figure CN120071894A_ABST
Abstract
Description
Background Art
[0002] With the development of intelligent vehicles, the voice interface has become an important way for drivers to interact with vehicle systems. To ensure the speech recognition accuracy, stability, and response speed of the voice interface, a comprehensive and efficient speech recognition test of the voice interface is required. However, traditional test methods often rely on manual operations, and testers need to perform speech recognition tests on the voice interface. Since there will be test interference during the manual operation process, the test accuracy of the voice interface is affected, and the test efficiency is low. Summary of the Invention
[0003] In view of this, the present invention provides a speech interface test method, device, electronic device, and storage medium, which at least partially solve the technical problems of low test efficiency and poor accuracy existing in the manual test of the speech interface in the prior art. The technical solution adopted by the present invention is as follows:
[0004] According to one aspect of the present application, a speech interface test method is provided, including:
[0005] In response to receiving a speech interface test request, obtaining a to-be-tested corpus corresponding to the speech interface test request and the number of times the corpus is to be tested;
[0006] Generating a test rule for the to-be-tested corpus according to the test conditions corresponding to the to-be-tested corpus;
[0007] Performing corpus recognition and parsing on the to-be-tested corpus according to the test rule of the to-be-tested corpus and the number of times the corpus is to be tested, and obtaining a corpus recognition and parsing result corresponding to each corpus recognition and parsing;
[0008] Comparing the corpus recognition and parsing result with the to-be-tested corpus to obtain the success rate and accuracy of the corpus recognition and parsing corresponding to the to-be-tested corpus.
[0009] In an exemplary embodiment of the present application, obtaining a to-be-tested corpus corresponding to the speech interface test request and the number of times the corpus is to be tested includes:
[0010] When receiving a speech interface test request and a corpus input by a user, determining the corpus input by the user as the to-be-tested corpus; when receiving a speech interface test request but not receiving a corpus input by the user, determining any corpus in a preset corpus library as the to-be-tested corpus;
[0011] And / or,
[0012] When receiving a speech interface test request and the number of test times input by the user, determining the number of test times input by the user as the number of times the corpus is to be tested; when receiving a speech interface test request but not receiving the number of test times input by the user, determining a preset number of test times as the number of times the corpus is to be tested.
[0013] In an exemplary embodiment of the present application, the corpus is a text corpus and / or a speech corpus;
[0014] Wherein, when the corpus only includes a text corpus, the corpus in the corpus is text corpus;
[0015] When the corpus only includes a speech corpus, the corpus in the corpus is speech corpus;
[0016] When the corpus includes a text corpus and a speech corpus, the corpus in the corpus includes at least text corpus, or also includes speech corpus and / or the text corpus corresponding to the speech corpus.
[0017] In an exemplary embodiment of the present application, when a speech interface test request is received but the number of test times input by the user is not received, determining the preset number of test times as the corpus test times includes:
[0018] When a speech interface test request is received but the number of test times input by the user is not received, obtaining the number of corpora to be tested;
[0019] When the number of corpora to be tested is greater than the preset test quantity threshold, determining the first number of test times as the corpus test times; when the number of corpora to be tested is less than or equal to the preset test quantity threshold, determining the second number of test times as the corpus test times; wherein, the first number of test times is less than the second number of test times.
[0020] In an exemplary embodiment of the present application, when the corpus to be tested is a single one, the test conditions corresponding to the corpus to be tested include the execution time of each test of the corpus to be tested and / or the average execution time of the corpus to be tested;
[0021] When the corpus to be tested is multiple ones, the test conditions corresponding to the corpus to be tested include the average execution time of each corpus to be tested and / or the overall execution time of all the corpora to be tested.
[0022] In an exemplary embodiment of the present application, according to the test rules of the corpus to be tested and the corpus test times, performing corpus recognition and parsing on the corpus to be tested to obtain the corpus recognition and parsing results corresponding to each corpus recognition and parsing, including:
[0023] According to the test rules of the corpus to be tested and the corpus test times, performing semantic recognition on the corpus to be tested several times in sequence, and obtaining the semantic recognition time corresponding to each semantic recognition in real time;
[0024] When the current semantic recognition time is less than the preset response time, the semantic intention information obtained by semantic recognition corresponding to the semantic recognition time is determined as the corpus recognition and parsing result corresponding to this semantic recognition; when the current semantic recognition time is greater than or equal to the preset response time, recognition timeout is determined as the corpus recognition and parsing result corresponding to this semantic recognition;
[0025] When the corpus recognition and parsing results corresponding to each semantic recognition within the continuous preset number of response times are all recognition timeouts, the test of the corpus to be tested is terminated.
[0026] In an exemplary embodiment of the present application, comparing the corpus recognition and parsing result with the corpus to be tested to obtain the corpus recognition and parsing success rate and corpus recognition and parsing accuracy corresponding to the corpus to be tested, including:
[0027] Obtain a plurality of corpus recognition and parsing results corresponding to the corpus to be tested;
[0028] Determine the corpus recognition and parsing result that is semantic intention information as the first corpus recognition and parsing result;
[0029] Compare each first corpus recognition and parsing result with the corpus to be tested, and when the matching degree between the first corpus recognition and parsing result and the semantic information corresponding to the corpus to be tested is greater than the preset matching degree threshold, determine the first corpus recognition and parsing result as the second corpus recognition and parsing result;
[0030] Determine the ratio of the number of second corpus recognition and parsing results corresponding to the corpus to be tested to the total number of corpus recognition and parsing results corresponding to the corpus to be tested as the corpus recognition and parsing success rate corresponding to the corpus to be tested;
[0031] Determine the average value of the matching degrees between a plurality of first corpus recognition and parsing results and the semantic information corresponding to the corpus to be tested as the corpus recognition and parsing accuracy corresponding to the corpus to be tested.
[0032] According to one aspect of the present application, there is provided a voice interface test device, including:
[0033] A request response module, configured to respond to receiving a voice interface test request, and obtain the corpus to be tested and the number of corpus tests corresponding to the voice interface test request;
[0034] A rule generation module, configured to generate a test rule for the corpus to be tested according to the test conditions corresponding to the corpus to be tested;
[0035] A test execution module, configured to perform corpus recognition and parsing on the corpus to be tested according to the test rule of the corpus to be tested and the number of corpus tests, and obtain the corpus recognition and parsing result corresponding to each corpus recognition and parsing;
[0036] A result analysis module for comparing the corpus recognition and parsing result with the corpus to be tested to obtain the success rate and accuracy of corpus recognition and parsing corresponding to the corpus to be tested.
[0037] According to one aspect of the present application, there is provided a non-transitory computer-readable storage medium storing at least one instruction or at least one program segment, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the foregoing voice interface testing method.
[0038] According to one aspect of the present application, there is provided an electronic device including a processor and the foregoing non-transitory computer-readable storage medium.
[0039] The present invention has at least the following beneficial effects:
[0040] The voice interface testing method of the present invention generates a test rule for the corpus to be tested according to the test conditions corresponding to the corpus to be tested. The test rule is expressed as a method of simulating user operations. According to the test rule of the corpus to be tested and the number of corpus tests, the corpus to be tested is subjected to at least one corpus recognition and parsing to obtain the corpus recognition and parsing result corresponding to each corpus recognition and parsing, and the obtained corpus recognition and parsing result is compared with the corpus to be tested in turn to obtain the success rate and accuracy of corpus recognition and parsing corresponding to the corpus to be tested. The voice parsing ability of the voice interface can be intuitively reflected through the success rate and accuracy of corpus recognition and parsing. Moreover, through the automated testing of the preset corpus to be tested, while reducing the manual workload, the testing efficiency is also improved, and the testing accuracy is also improved due to the absence of interference from manual testing. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0042] Figure 1 It is a flowchart of the voice interface testing method provided by the embodiment of the present invention;
[0043] Figure 2 It is a block diagram of the voice interface testing device provided by the embodiment of the present invention;
[0044] Figure 3 It is a functional logic diagram of the voice interface testing device provided by the embodiment of the present invention;
[0045] Figure 4Schematic diagram of input / output logic of the voice interface test method provided by the embodiment of the present invention;
[0046] Figure 5 It is a flowchart of the method for the voice interface test method provided by the embodiment of the present invention when performing corpus recognition and parsing on the to-be-tested corpus. Detailed implementation manners
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0048] The traditional test method for voice interfaces is to conduct manual tests by testers. Since testers may be affected by the external environment during the manual test process, and when the number of tests for voice interfaces is too large, it will also cause a physical burden on testers. Therefore, there may be test errors in the test results of voice interfaces, affecting the accuracy and test efficiency of the test results.
[0049] Therefore, in order to solve the problems of poor accuracy and low test efficiency of the traditional test method in the prior art, the voice interface test method described in the present invention is proposed, as Figure 1 shown, including:
[0050] Step S100, in response to receiving a voice interface test request, obtain the to-be-tested corpus corresponding to the voice interface test request and the number of corpus test times;
[0051] The to-be-tested corpus is the voice information for performing voice recognition tests on the voice interface. The to-be-tested corpus can be single or multiple. The number of corpus test times is the number of times for performing semantic parsing tests on each to-be-tested corpus. By performing multiple tests on each to-be-tested corpus, the semantic parsing performance of the voice interface for the same to-be-tested corpus is determined, and the voice recognition ability of the voice interface is judged by testing the semantic recognition results of the voice interface for the to-be-tested corpus.
[0052] Further, obtaining the to-be-tested corpus corresponding to the voice interface test request and the number of corpus test times in step S100 includes steps S110 - S140:
[0053] Step S110, in the case of receiving a voice interface test request and the corpus input by the user, determine the corpus input by the user as the to-be-tested corpus;
[0054] Step S120: When a voice interface test request is received but no corpus input by the user is received, any corpus in the preset corpus is determined as the corpus to be tested;
[0055] The corpus to be tested can be the corpus input by the user or the corpus in the preset corpus. When a voice interface test request for testing the voice interface is received, if the user also inputs a corpus, then this corpus is used as the corpus to be tested. If the user does not input a corpus, then any corpus is selected from the preset corpus as the corpus to be tested.
[0056] Among them, as a feasible embodiment, the preset corpus can be a text corpus and / or a voice corpus. The user can update the corpus in the corpus by performing operations such as adding / deleting / changing the corpus in the corpus. The format of the corpus files in the corpus can be WAV, MP3, TXT, etc., and the corpus should cover the instructions and expressions commonly used by the user (driver) to ensure the comprehensiveness and accuracy of the test.
[0057] In the case where the corpus only includes a text corpus, the corpus in the corpus is text corpus, that is, the corpus is text information; in the case where the corpus only includes a voice corpus, the corpus in the corpus is voice corpus, that is, the corpus is voice information; in the case where the corpus includes a text corpus and a voice corpus, the corpus in the corpus at least includes text corpus, or also includes voice corpus and / or the text corpus corresponding to the voice corpus.
[0058] The text corpus is preset through Excel (table) and TXT (text) editing, imported through the Demo App (auxiliary software tool), and the corresponding text corpus list is automatically parsed.
[0059] The voice corpus generates corresponding recording files through real-time recording with a recording software, or can also generate recording files through pre-recorded audio. One recording file corresponds to one voice corpus, and is stored in the corpus through the way of folder import, and the voice corpus in the corresponding folder is automatically recognized.
[0060] Correspondingly, the corpus input by the user can also be text corpus or voice corpus. The voice interface of this application includes an NLU interface (text natural language recognition interface) and an ASR interface (voice recognition interface). The NLU interface is text-tested through the text corpus, and the ASR interface is voice-tested through the voice corpus. The user can input text corpus by typing text, can input voice corpus by recording, or can also realize the input of the corpus by loading the text file storing the text corpus and the voice file storing the voice corpus.
[0061] Step S130: When receiving a voice interface test request and the number of tests input by the user, determine the number of tests input by the user as the corpus test times.
[0062] Step S140: When receiving a voice interface test request but not receiving the number of tests input by the user, determine the preset number of tests as the corpus test times.
[0063] The corpus test times can be specified by the user or the system. If the user inputs the number of tests, determine the number of tests input by the user as the corpus test times. If the user does not input the number of tests, determine the preset number of tests specified by the system as the corpus test times.
[0064] Further, when receiving a voice interface test request but not receiving the number of tests input by the user in Step S140, determining the preset number of tests as the corpus test times includes Steps S141 - S142:
[0065] Step S141: When receiving a voice interface test request but not receiving the number of tests input by the user, obtain the number of corpora to be tested.
[0066] When the user does not input the number of tests, in order to further improve the test efficiency of the voice interface, limit the corpus test times for further testing the same corpus to be tested based on the number of corpora to be tested.
[0067] Step S142: When the number of corpora to be tested is greater than the preset test quantity threshold, determine the first test times as the corpus test times; when the number of corpora to be tested is less than or equal to the preset test quantity threshold, determine the second test times as the corpus test times; where the first test times is less than the second test times.
[0068] If the number of corpora to be tested is greater than the preset test quantity threshold, it means that the number of corpora to be tested that need to be tested is too large. In order to shorten the test time of each voice interface, it is necessary to lower the corpus test times for each corpus to be tested. Since the semantic parsing process of each corpus to be tested is automated without manual intervention, therefore, lowering the corpus test times will not have a great impact on the test accuracy. On the contrary, if the number of corpora to be tested is too small, in order to further improve the test accuracy rate, it is necessary to increase the corpus test times for the corpus to be tested to balance the test time of different numbers of corpora to be tested.
[0069] In addition, the number of corpus tests can be limited by the semantic information included in the corpus to be tested. For example, when the amount of semantic information included in a corpus to be tested (which can be the execution information included in this corpus to be tested. For example, if the corpus to be tested is "Open the window, play music, turn off the air conditioner", that is, this corpus to be tested includes three semantic information) is too large, the first test number is determined as the corpus test number of this corpus to be tested. On the contrary, when the amount of semantic information included in the corpus to be tested is too small, the second test number is determined as the corpus test number of this corpus to be tested, realizing the dynamic adjustment of the corpus test number for each corpus to be tested, balancing the test time of different corpora to be tested, improving the test efficiency of the voice interface and shortening the overall test time at the same time.
[0070] Step S200: Generate a test rule for the corpus to be tested according to the test conditions corresponding to the corpus to be tested;
[0071] The test condition is the test item of the corpus to be tested. Among them, as a feasible embodiment, when the corpus to be tested is a single piece, the test conditions corresponding to the corpus to be tested include the execution time of each test of the corpus to be tested and / or the average execution time of the corpus to be tested; when the corpus to be tested is multiple pieces, the test conditions corresponding to the corpus to be tested include the average execution time of each corpus to be tested and / or the overall execution time of all corpora to be tested.
[0072] The test rule of the corpus to be tested is the test method of this corpus to be tested. According to the test conditions of each corpus to be tested, the corresponding test method is generated. The test method can be implemented through a preset script. For example, developers will pre - establish test scripts for different types of test corpora. When receiving the corpus to be tested, according to the type of the corpus to be tested (which can be the execution type obtained after semantic recognition), the corresponding test script is determined, or the function of the test script can be expanded according to actual needs, and the corresponding test rule is generated according to this test script to improve the comprehensiveness and accuracy of voice testing.
[0073] Step S300: Perform corpus recognition and parsing on the corpus to be tested according to the test rule of the corpus to be tested and the corpus test number, and obtain the corpus recognition and parsing result corresponding to each corpus recognition and parsing;
[0074] Performing corpus recognition and parsing on the corpus to be tested means performing semantic recognition on the corpus to be tested, parsing out the semantic information included in the corpus to be tested, and obtaining the corresponding corpus recognition and parsing result.
[0075] Furthermore, in step S300, performing corpus recognition and parsing on the corpus to be tested according to the test rule of the corpus to be tested and the corpus test number, and obtaining the corpus recognition and parsing result corresponding to each corpus recognition and parsing includes steps S310 - S330:
[0076] Step S310: According to the test rules of the to-be-tested corpus and the number of corpus test times, perform semantic recognition on the to-be-tested corpus several times in sequence, and obtain the semantic recognition time corresponding to each semantic recognition in real time.
[0077] The number of times of semantic recognition is the number of corpus test times of the to-be-tested corpus.
[0078] Step S320: When the current semantic recognition time is less than the preset response time, determine the semantic intention information obtained by the semantic recognition corresponding to this semantic recognition time as the corpus recognition and parsing result corresponding to this semantic recognition; when the current semantic recognition time is greater than or equal to the preset response time, determine that the recognition times out as the corpus recognition and parsing result corresponding to this semantic recognition.
[0079] When the voice interface is working, the semantic recognition time (voice parsing time, response time) is also particularly important. In order to avoid the user's waiting time being too long, it is necessary to require the semantic recognition time of the voice interface to be within a specific time period, that is, the voice interface must respond to the voice information input by the user within a specific time period. Therefore, the semantic recognition time is also a test item of the voice interface.
[0080] Obtain the semantic recognition time corresponding to each semantic recognition in real time. If the semantic recognition time of this semantic recognition is less than the preset response time, that is, within the preset response time, the voice interface returns the semantic parsing result, it indicates that the semantic recognition time of the voice interface in this semantic recognition meets the preset standard, and then determine the semantic parsing result (the semantic intention information obtained by the semantic recognition corresponding to this semantic recognition time) as the corpus recognition and parsing result corresponding to this semantic recognition; on the contrary, if the semantic recognition time of this semantic recognition is greater than or equal to the preset response time, it means that the semantic recognition time in this semantic recognition does not meet the preset standard, that is, the response times out, and then determine that the recognition times out as the corpus recognition and parsing result corresponding to this semantic recognition.
[0081] Step S330: When the corpus recognition and parsing results corresponding to each semantic recognition within the continuous preset number of response times are all recognition timeouts, terminate the test of the to-be-tested corpus.
[0082] If within the continuous preset number of response times, the voice interface times out in response, it means that there is an abnormality in the voice interface's parsing of the to-be-tested corpus, or there is a fault in the voice interface itself. Then stop the test of the to-be-tested corpus, directly proceed to the test of the next to-be-tested corpus, or send a warning signal to the user to remind the user to check the voice interface to shorten the test time of the to-be-tested corpus.
[0083] Such as Figure 5As shown, when performing corpus recognition and parsing on the test corpus to be processed, each test corpus to be processed corresponds to a simulated test task. After inputting a test corpus to be processed into the speech interface, the speech interface will perform semantic recognition, parsing, and generate a corresponding speech return result for the test corpus to be processed. The speech return result generally includes normal results, error results, no return (i.e., recognition timeout in this application), etc. When normal results and error results occur, they can be processed according to the expected situation results; when there is no return result, that is, when the corpus input times out, since the test method of this application defaults to a timeout mechanism (for example, if there is no result returned within two seconds, it will be defaulted that the test of this test corpus fails), therefore, the corpus recognition and parsing of this test corpus is re-performed. If the speech return result of each of the consecutive response times (for example, three times) of the re-performed corpus recognition and parsing is a no return result, then the corpus test of this test corpus will be ended and the corpus test of the next test corpus will be continued.
[0084] Step S400: Compare the corpus recognition and parsing result with the test corpus to be processed to obtain the success rate and accuracy of the corpus recognition and parsing corresponding to the test corpus to be processed.
[0085] The success rate of corpus recognition and parsing is the probability of successful corpus recognition and parsing for this test corpus to be processed, and the accuracy of corpus recognition and parsing is the accuracy rate of the corpus recognition and parsing result obtained after corpus recognition and parsing for this test corpus to be processed.
[0086] Furthermore, in step S400, comparing the corpus recognition and parsing result with the test corpus to be processed to obtain the success rate and accuracy of the corpus recognition and parsing corresponding to the test corpus to be processed includes steps S410 - S450:
[0087] Step S410: Obtain several corpus recognition and parsing results corresponding to the test corpus to be processed;
[0088] Step S420: Determine the corpus recognition and parsing result that is semantic intent information as the first corpus recognition and parsing result;
[0089] The first corpus recognition and parsing result is a corpus recognition and parsing result that is not a response timeout.
[0090] Step S430: Compare each first corpus recognition and parsing result with the test corpus to be processed. In the case where the matching degree between the first corpus recognition and parsing result and the semantic information corresponding to the test corpus to be processed is greater than the preset matching degree threshold, determine this first corpus recognition and parsing result as the second corpus recognition and parsing result;
[0091] The second corpus recognition and parsing result is the first corpus recognition and parsing result that meets the semantic matching degree standard.
[0092] Step S440: Determine the success rate of corpus recognition and parsing corresponding to the corpus to be tested as the ratio of the number of the second corpus recognition and parsing results corresponding to the corpus to be tested to the total number of the corpus recognition and parsing results corresponding to the corpus to be tested;
[0093] Step S450: Determine the precision of corpus recognition and parsing corresponding to the corpus to be tested as the average value of the matching degrees between several first corpus recognition and parsing results and the semantic information corresponding to the corpus to be tested.
[0094] After testing several corpora to be tested, the success rate of corpus recognition and parsing and the precision of corpus recognition and parsing corresponding to each corpus to be tested are obtained. By analyzing the success rate of corpus recognition and parsing and the precision of corpus recognition and parsing corresponding to several corpora to be tested, the automated performance testing of the voice interface can be realized (for example, by analyzing the success rate of corpus recognition and parsing corresponding to several corpora to be tested, the strength of the parsing ability of the voice interface, whether it meets the requirements of the current working environment and the direction that needs to be optimized, etc. can be obtained), so as to identify some problems that are not easily detected in the manual testing process, and the location of the problems can be intuitively found. The discovered problems can be output to the corresponding developers to solve. And since the corpora to be tested are preset in advance and there will be no deviation, the error interference caused by manual input of corpora is reduced, the manual testing workload is reduced, and the testing efficiency is improved.
[0095] As Figure 3 shown, it is the hierarchical construction diagram of the voice interface testing system, including UI (display interface), server (server), and voice SDK (voice parsing software development kit). The voice SDK performs semantic parsing on the input corpus to be tested and displays the obtained parsing results on the UI.
[0096] As Figure 4 shown, it is the testing block diagram of the voice interface. The user inputs information such as the corpus to be tested, the number of executions (the number of corpus tests), the voice configuration (the language to be tested), and the test performance conditions. Through the voice test of the corpus to be tested by the voice interface, information such as the parsing intention result (semantic recognition result), the execution intention result (the result of corresponding mechanism execution according to the semantic recognition result), the average time consumption (the average execution time of testing the corpus to be tested), the recognition rate (the precision of corpus recognition and parsing), and the success rate (the success rate of corpus recognition and parsing) is output.
[0097] In addition, after obtaining test data such as the success rate of corpus recognition and parsing, the accuracy of corpus recognition and parsing, the execution time of the corpus to be tested in each test, the average execution time of the corpus to be tested, and the overall execution time of all the corpora to be tested, it can be saved locally for the user to view at any time. These test data can also be integrated and output according to a preset format (such as file format output or table format output) to form a test report. The test report can include information such as test coverage, list of existing problems, and repair suggestions, so that the user can have an intuitive understanding of the execution performance of the voice interface to identify the problems and potential risks existing in the voice interface.
[0098] In addition, the present invention also provides a voice interface test device 100, as Figure 2 shown, including:
[0099] A request response module 110, configured to respond to receiving a voice interface test request, and obtain the corpus to be tested and the number of times of corpus testing corresponding to the voice interface test request;
[0100] Among them, the method for the request response module 110 to obtain the corpus to be tested and the number of times of corpus testing corresponding to the voice interface test request includes:
[0101] When receiving a voice interface test request and the corpus input by the user, determining the corpus input by the user as the corpus to be tested; when receiving a voice interface test request but not receiving the corpus input by the user, determining any corpus in the preset corpus library as the corpus to be tested; among them, when the corpus library only includes a text corpus library, the corpus in the corpus library is a text corpus; when the corpus library only includes a voice corpus library, the corpus in the corpus library is a voice corpus; when the corpus library includes a text corpus library and a voice corpus library, the corpus in the corpus library includes at least a text corpus, or may also include a voice corpus and / or the text corpus corresponding to the voice corpus;
[0102] When receiving a voice interface test request and the number of test times input by the user, determining the number of test times input by the user as the number of times of corpus testing; when receiving a voice interface test request but not receiving the number of test times input by the user, determining the preset number of test times as the number of times of corpus testing.
[0103] Further, the method for the request response module 110 to determine the preset number of test times as the number of times of corpus testing when receiving a voice interface test request but not receiving the number of test times input by the user includes:
[0104] When receiving a voice interface test request but not receiving the number of test times input by the user, obtaining the number of corpora to be tested;
[0105] When the quantity of the corpus to be tested is greater than the preset test quantity threshold, determine the first test number as the corpus test number; when the quantity of the corpus to be tested is less than or equal to the preset test quantity threshold, determine the second test number as the corpus test number; wherein, the first test number is less than the second test number.
[0106] A rule generation module 120, configured to generate a test rule for the corpus to be tested according to the test conditions corresponding to the corpus to be tested.
[0107] Wherein, when the corpus to be tested is a single one, the test conditions corresponding to the corpus to be tested include the execution time of each test of the corpus to be tested and / or the average execution time of the corpus to be tested.
[0108] When the corpus to be tested is multiple, the test conditions corresponding to the corpus to be tested include the average execution time of each corpus to be tested and / or the overall execution time of all the corpora to be tested.
[0109] A test execution module 130, configured to perform corpus recognition and parsing on the corpus to be tested according to the test rule of the corpus to be tested and the corpus test number, so as to obtain a corpus recognition and parsing result corresponding to each corpus recognition and parsing.
[0110] Furthermore, the method by which the test execution module 130 performs corpus recognition and parsing on the corpus to be tested according to the test rule of the corpus to be tested and the corpus test number, so as to obtain a corpus recognition and parsing result corresponding to each corpus recognition and parsing includes:
[0111] Perform semantic recognition on the corpus to be tested several times in sequence according to the test rule of the corpus to be tested and the corpus test number, and obtain the semantic recognition time corresponding to each semantic recognition in real time.
[0112] When the current semantic recognition time is less than the preset response time, determine the semantic intention information obtained by the semantic recognition corresponding to the semantic recognition time as the corpus recognition and parsing result corresponding to the semantic recognition; when the current semantic recognition time is greater than or equal to the preset response time, determine that the recognition times out as the corpus recognition and parsing result corresponding to the semantic recognition.
[0113] When the corpus recognition and parsing results corresponding to each semantic recognition within a continuous preset number of response times are all recognition timeouts, terminate the test of the corpus to be tested.
[0114] A result analysis module 140, configured to compare the corpus recognition and parsing result with the corpus to be tested, so as to obtain the corpus recognition and parsing success rate and the corpus recognition and parsing accuracy corresponding to the corpus to be tested.
[0115] Further, the method for the result analysis module 140 to compare the corpus recognition and parsing result with the corpus to be tested to obtain the corpus recognition and parsing success rate and accuracy corresponding to the corpus to be tested includes:
[0116] Obtain a plurality of corpus recognition and parsing results corresponding to the corpus to be tested;
[0117] Determine the corpus recognition and parsing result that is semantic intention information as the first corpus recognition and parsing result;
[0118] Compare each first corpus recognition and parsing result with the corpus to be tested. When the matching degree between the first corpus recognition and parsing result and the semantic information corresponding to the corpus to be tested is greater than the preset matching degree threshold, determine the first corpus recognition and parsing result as the second corpus recognition and parsing result;
[0119] Determine the ratio of the number of the second corpus recognition and parsing results corresponding to the corpus to be tested to the total number of the corpus recognition and parsing results corresponding to the corpus to be tested as the corpus recognition and parsing success rate corresponding to the corpus to be tested;
[0120] Determine the average value of the matching degrees between a plurality of first corpus recognition and parsing results and the semantic information corresponding to the corpus to be tested as the corpus recognition and parsing accuracy corresponding to the corpus to be tested.
[0121] The voice interface testing method of the present invention generates a test rule for the corpus to be tested according to the test conditions corresponding to the corpus to be tested. The test rule is expressed as a method for simulating user operations. According to the test rule of the corpus to be tested and the number of corpus tests, perform at least one corpus recognition and parsing on the corpus to be tested to obtain the corpus recognition and parsing results corresponding to each corpus recognition and parsing, and compare the obtained corpus recognition and parsing results with the corpus to be tested in sequence to obtain the corpus recognition and parsing success rate and accuracy corresponding to the corpus to be tested. The voice parsing ability of the voice interface can be intuitively reflected through the corpus recognition and parsing success rate and accuracy. Moreover, through the automated testing of the preset corpus to be tested, while reducing the manual workload, the testing efficiency is also improved. And due to the absence of interference from manual testing, the need for manual operations and interventions is reduced, and the testing accuracy is also improved.
[0122] The embodiment of the present invention also provides a computer program product, which includes program codes. When the program product runs on an electronic device, the program codes are used to cause the electronic device to execute the steps in the methods according to various exemplary embodiments of the present invention described above in this specification.
[0123] In addition, although the various steps of the methods in this disclosure are described in a specific order in the accompanying drawings, this is not a requirement or implication that these steps must be performed in that specific order, or that all of the shown steps must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0124] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0125] In an exemplary embodiment of this disclosure, there is also provided an electronic device capable of implementing the above method.
[0126] Those skilled in the relevant technical field can understand that various aspects of the present invention can be implemented as a system, method, or program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to herein as "circuitry", "module", or "system".
[0127] The electronic device according to this embodiment of the present invention. The electronic device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0128] The electronic device is presented in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: at least one of the above-mentioned processors, at least one of the above-mentioned memories, and a bus connecting different system components (including the memory and the processor).
[0129] Among them, the memory stores program code, and the program code can be executed by the processor, so that the processor executes the steps according to various exemplary embodiments of the present invention described in the "Exemplary Methods" section of this specification.
[0130] The memory may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) and / or a cache memory, and may further include a read-only memory (ROM).
[0131] The storage may also include a program / utilities having a set (at least one) of program modules, such program modules including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of these examples or some combination thereof may include an implementation of a network environment.
[0132] The bus may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus structures.
[0133] The electronic device may also communicate with one or more external devices (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device, and / or may communicate with any device that enables the electronic device to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be through an input / output (I / O) interface. Further, the electronic device may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter.
[0134] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to cause a computing device (which may be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0135] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above method of this specification is stored. In some possible embodiments, various aspects of the present invention may also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the "Exemplary Method" section above of this specification.
[0136] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0137] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0138] The program code contained on the readable medium may be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above.
[0139] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0140] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0141] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0142] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A voice interface testing method, characterized in that: The method comprises: In response to receiving a voice interface test request, obtaining a corpus to be tested and a number of corpus tests corresponding to the voice interface test request; Generating a test rule for the corpus to be tested according to the test conditions corresponding to the corpus to be tested; According to the test rules of the corpus to be tested and the number of corpus tests, the corpus recognition and analysis is performed on the corpus to be tested, and the corpus recognition and analysis results corresponding to each corpus recognition and analysis are obtained; The corpus recognition and parsing result is compared with the corpus to be tested to obtain the corpus recognition and parsing success rate and the corpus recognition and parsing accuracy corresponding to the corpus to be tested.
2. The method according to claim 1, characterized in that The obtaining of the to-be-tested corpus and the number of corpus tests corresponding to the voice interface test request includes: In the case of receiving a voice interface test request and a corpus input by a user, determining the corpus input by the user as the corpus to be tested; in the case of receiving a voice interface test request but not receiving the corpus input by the user, determining any corpus in a preset corpus as the corpus to be tested; and / or, When a voice interface test request and a test number input by the user are received, the test number input by the user is determined as the corpus test number; when a voice interface test request is received but the test number input by the user is not received, the preset test number is determined as the corpus test number.
3. The method according to claim 2, characterized in that The corpus is a text corpus and / or a speech corpus; Wherein, in the case where the corpus only includes a text corpus, the corpus in the corpus is a text corpus; In the case where the corpus only includes a speech corpus, the corpus in the corpus is a speech corpus; In the case where the corpus includes a text corpus and a speech corpus, the corpus in the corpus includes at least text corpus, or further includes speech corpus and / or text corpus corresponding to the speech corpus.
4. The method according to claim 2, characterized in that: In the case where a voice interface test request is received but a test number input by a user is not received, determining a preset test number as a corpus test number includes: When receiving a voice interface test request but not receiving a test number input by a user, obtaining the number of corpora to be tested; When the number of the corpus to be tested is greater than a preset test number threshold, the first test number is determined as the corpus test number; when the number of the corpus to be tested is less than or equal to the preset test number threshold, the second test number is determined as the corpus test number; wherein the first test number is less than the second test number.
5. The method according to claim 1, characterized in that In the case that the corpus to be tested is a single piece, the test condition corresponding to the corpus to be tested includes the execution time of each test of the corpus to be tested and / or the average execution time of the corpus to be tested; In the case that there are multiple pieces of the corpus to be tested, the test conditions corresponding to the corpus to be tested include the average execution time of each piece of the corpus to be tested and / or the overall execution time of all the corpus to be tested.
6. The method according to claim 1, characterized in that The step of performing corpus recognition and analysis on the corpus to be tested according to the test rules of the corpus to be tested and the number of corpus tests, and obtaining corpus recognition and analysis results corresponding to each corpus recognition and analysis, includes: According to the test rules of the corpus to be tested and the number of corpus tests, sequentially perform semantic recognition on the corpus to be tested for several times, and obtain the semantic recognition time corresponding to each semantic recognition in real time; When the current semantic recognition time is less than the preset response time, the semantic intention information obtained by the semantic recognition corresponding to the semantic recognition time is determined as the corpus recognition and parsing result corresponding to the semantic recognition; when the current semantic recognition time is greater than or equal to the preset response time, the recognition timeout is determined as the corpus recognition and parsing result corresponding to the semantic recognition; When the corpus recognition and parsing results corresponding to each semantic recognition within a continuous preset number of responses are all recognition timeouts, the test of the corpus to be tested is terminated.
7. The method according to claim 6, characterized in that The comparing the corpus recognition and parsing result with the corpus to be tested to obtain the corpus recognition and parsing success rate and corpus recognition and parsing accuracy corresponding to the corpus to be tested includes: Obtaining a plurality of corpus recognition and parsing results corresponding to the corpus to be tested; Determining the corpus recognition and parsing result that is the semantic intent information as the first corpus recognition and parsing result; Comparing each of the first corpus recognition and analysis results with the corpus to be tested, and determining the first corpus recognition and analysis result as the second corpus recognition and analysis result when the matching degree between the semantic information corresponding to the first corpus recognition and analysis result and the corpus to be tested is greater than a preset matching degree threshold; Determine the ratio of the number of the second corpus recognition and parsing results corresponding to the corpus to be tested to the total number of the corpus recognition and parsing results corresponding to the corpus to be tested as the corpus recognition and parsing success rate corresponding to the corpus to be tested; An average value of the matching degrees of a plurality of the first corpus recognition and parsing results and the semantic information corresponding to the corpus to be tested is determined as the corpus recognition and parsing accuracy corresponding to the corpus to be tested.
8. A voice interface testing device, characterized in that: include: A request response module, used for responding to a received voice interface test request, obtaining the corpus to be tested and the number of corpus tests corresponding to the voice interface test request; A rule generation module, used for generating a test rule for the corpus to be tested according to the test conditions corresponding to the corpus to be tested; A test execution module, used to perform corpus recognition and analysis on the corpus to be tested according to the test rules of the corpus to be tested and the number of corpus tests, and obtain corpus recognition and analysis results corresponding to each corpus recognition and analysis; The result analysis module is used to compare the corpus recognition and parsing result with the corpus to be tested to obtain the corpus recognition and parsing success rate and corpus recognition and parsing accuracy corresponding to the corpus to be tested.
9. A non-transitory computer-readable storage medium, wherein at least one instruction or at least one program is stored in the storage medium, characterized in that: The at least one instruction or the at least one program is loaded and executed by the processor to implement the voice interface testing method as described in any one of claims 1-7.
10. An electronic device, characterized in that: The invention comprises a processor and the non-transitory computer-readable storage medium as claimed in claim 9.