Intelligent Voice Interaction Testing Method, Device, Equipment and Storage Medium
By building a test database containing multiple types of text data and using programmable and test statistical tools, the problem of lack of unified testing standards for intelligent voice interaction systems is solved, and more scientific and accurate semantic understanding functional testing is achieved.
Patent Information
- Application Number
- CN202110627050.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-04
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2041-06-04
AI Technical Summary
The existing intelligent voice interaction system lacks a unified semantic understanding functional test standard and specification test data set, resulting in unsatisfactory test results, and the test data is separated from actual application scenarios, making it difficult to accurately evaluate product performance.
A test database is built, which contains text data of defined and undefined scenarios, including general, commonly used, chatty, meaningless text, etc. It conducts semantic understanding tests through programmable and test statistical tools, generates matching test data sets, and performs functional and performance testing of the system under test.
It realizes a more scientific and close-to-real scene semantic understanding functional test, improves the scientificity and accuracy of the test, and can comprehensively and in-depth evaluate the semantic understanding ability of the system under test.
Smart Images

Figure CN115438670B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent voice interaction technology, and in particular, to an intelligent voice interaction test method, device, equipment, and storage medium. Background Art
[0002] Intelligent voice interaction is widely used in many fields such as smart home, intelligent customer service, mobile terminals, in-vehicle terminals, as well as intelligent education, intelligent healthcare, intelligent office, service robots, etc., and has become one of the important ways of current human-computer interaction.
[0003] As intelligent voice interaction penetrates deeper into all aspects of production and life, it is necessary to uniformly standardize the system reference framework, basic technical requirements, Internet interface requirements, etc. of intelligent voice interaction. In this regard, the country has formulated basic national standards to support intelligent voice interaction systems, such as GB / T36464, GB / T 5271.29—2006, etc. On this basis, it is also necessary to use unified test methods and evaluation criteria to test the functions of intelligent voice interaction systems, so as to provide basic methods for testing products and services related to intelligent voice interaction.
[0004] An intelligent voice interaction system needs to at least have functions such as speech recognition and semantic understanding, so as to understand the intention of the speaker and then interact with the speaker. Among them, for the test of the semantic understanding function of the intelligent voice interaction system, there is no unified national standard, and there is no unified and standardized test method in the industry.
[0005] In addition, the test data of the semantic understanding function test schemes that have emerged in the industry are mostly temporarily collected data, and no complete test data set has been formed. The test data is detached from the product application scenario, and the semantic understanding function test implemented using these data cannot accurately test the performance or function of the product in actual applications, and the test effect is not ideal.
[0006] Moreover, in the existing semantic understanding function test solutions, the vocabulary or text specified in standard documents such as GB / T 5271.29—2006, GB / T 2312—1980, GB 18030, IETF RFC3629 F.Yergeau., UTF-8, a transformation format of ISO 10646, STD 63, RFC 3629, November, 2003, and GB / T 36464 are usually referred to, and the vocabulary or text that meets the above standards is selected for the semantic understanding function test. However, the above annotation documents only provide text standards, and cannot enable those skilled in the art to know which texts should be selected and how many of each type of text should be selected for the semantic understanding test. That is, the test data selected by the existing semantic understanding test solutions only simply selects the text data that can be used for testing, but cannot clarify the relationship between the test data and the test items or test scenarios, resulting in the test data being unable to well support the test requirements and it being difficult to obtain an ideal test effect. Summary of the Invention
[0007] Based on the above technical status quo, the present application proposes an intelligent voice interaction test method, device, equipment, and storage medium, which can realize intelligent voice interaction testing, especially can realize semantic understanding testing of the system under test.
[0008] An intelligent voice interaction test method includes:
[0009] Determine semantic understanding test items according to the functional requirements of the system under test and the semantic understanding test requirements;
[0010] For each semantic understanding test item, perform the following test processing respectively:
[0011] Select test data corresponding to the semantic understanding test item from the pre-constructed test database, and generate a test data set that matches the semantic understanding test item by using the selected test data;
[0012] Wherein, the test database includes text data of defined scenarios or services, and text data of undefined scenarios or services;
[0013] The text data of the defined scenarios or services includes general text data of the defined scenarios or services and common text data of the defined scenarios or services; among the general text data of the defined scenarios or services, the text data corresponding to each service is not less than 200 pieces; among the common text data of the defined scenarios or services, there are at least 3 real text data corresponding to each service;
[0014] The text data of the undefined scenarios or services includes general text data, common text data, chatting text data, and meaningless or illogical text data of undefined scenarios or services in the same field; for the general text data and the common text data of undefined scenarios or services in the same field, there are at least 3 real text data corresponding to each service; the number of the chatting text data is not less than 1000, and the average number of characters of each chatting text is not less than 5 characters; the number of the meaningless or illogical text data is not less than 100, and each one is not less than 5 characters.
[0015] The general text data, the common text data of the defined scenarios or services, the general text data of the undefined scenarios or services in the same field, and the common text data of the undefined scenarios or services in the same field are all composed of common text, special text, and abnormal text obtained by the system under test through speech recognition. Moreover, the length distribution of all text data belonging to the same type conforms to the set quantity distribution requirements.
[0016] The common text includes single-word or phrase text, short text, single-sentence text, dialogue text, paragraph text, and article text with intention representation, and the number of each type of text is not less than 5.
[0017] The special text includes sensitive information text, named entity text, special format text, specific language text, special character set encoding text, and special symbol text. Among them, the number of sensitive information text and named entity text is not less than 1000 respectively, and the number of special format text, specific language text, special character set encoding text, and special symbol text is not less than 5 respectively.
[0018] The abnormal text includes garbled text and unsupported language text, and the number of each type of text is not less than 5.
[0019] Through a programmable test tool, the test data in the test dataset are respectively input into the system under test to test the semantic understanding test items of the system under test.
[0020] Through a test statistics tool, the operation result of the system under test during the test of the semantic understanding test items is obtained, and based on the obtained operation result and the test content of the semantic understanding test items, the test result corresponding to the semantic understanding test items is determined.
[0021] An intelligent voice interaction test device includes:
[0022] A project selection unit for determining semantic understanding test items according to the functional requirements of the system under test and the semantic understanding test requirements.
[0023] A test data set generation unit, configured to select test data corresponding to semantic understanding test items from a pre - constructed test database, and generate a test data set matching the semantic understanding test items by using the selected test data;
[0024] Wherein, the test database includes text data of defined scenarios or services, and text data of undefined scenarios or services;
[0025] The text data of defined scenarios or services includes general text data of defined scenarios or services, and common text data of defined scenarios or services; in the general text data of defined scenarios or services, the text data corresponding to each service is not less than 200 pieces; in the common text data of defined scenarios or services, there are at least 3 pieces of real text data corresponding to each service;
[0026] The text data of undefined scenarios or services includes general text data of undefined scenarios or services in the same field, common text data of undefined scenarios or services in the same field, chat text data, and text data that is meaningless or illogical; in the general text data of undefined scenarios or services in the same field, and the common text data of undefined scenarios or services in the same field, there are at least 3 pieces of real text data corresponding to each service; the number of chat text data is not less than 1000 pieces, and the average number of characters in each chat text is not less than 5 characters; the text data that is meaningless or illogical is not less than 100 pieces, and each piece is not less than 5 characters;
[0027] The general text data of defined scenarios or services, the common text data of defined scenarios or services, the general text data of undefined scenarios or services in the same field, and the common text data of undefined scenarios or services in the same field are all composed of common text, special text, and abnormal text obtained by voice recognition of the system under test, and moreover, the length distribution of all text data belonging to the same type conforms to the set quantity distribution requirements;
[0028] The common text includes single - character or word text, short - text, single - sentence text, dialogue text, paragraph text, and article text with intention representation, and the number of each type of text is not less than 5 pieces;
[0029] The special text includes sensitive information text, named entity text, special - format text, specific - language text, special - character - set - encoding text, and special - symbol text. Among them, the number of sensitive information text and named entity text is not less than 1000 pieces respectively, and the number of special - format text, specific - language text, special - character - set - encoding text, and special - symbol text is not less than 5 pieces respectively;
[0030] The abnormal text includes garbled text and text in unsupported languages, and the number of each type of text is not less than 5;
[0031] A test processing unit, configured to input the test data in the test dataset into the system under test respectively through a programmable test tool, so as to perform tests on the semantic understanding test items of the system under test;
[0032] A test statistics unit, configured to obtain the operation result of the system under test during the test of the semantic understanding test item through a test statistics tool, and determine the test result corresponding to the semantic understanding test item based on the obtained operation result and the test content of the semantic understanding test item.
[0033] An intelligent voice interaction test device, comprising:
[0034] A memory and a processor;
[0035] The memory is connected to the processor and is used for storing computer programs;
[0036] The processor is configured to implement the above-mentioned intelligent voice interaction test method by running the program in the memory.
[0037] A storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned intelligent voice interaction test method is implemented.
[0038] When performing a semantic understanding function test on the system under test, the intelligent voice interaction test method provided by the embodiments of the present application can determine the semantic understanding test items according to the functional requirements of the system under test and the semantic understanding test requirements. At the same time, the embodiments of the present application also pre-construct a test database, and test data corresponding to the determined semantic understanding test items can be selected therefrom for the semantic understanding function test of the system under test. When the semantic understanding test items and the test data are determined, by inputting the test data into the system under test, it is possible to perform tests on the semantic understanding test items of the system under test, and then according to the operation result of the system under test during the test process and the test content of the semantic understanding test items, the test result corresponding to the semantic understanding test item can be determined.
[0039] Adopting the technical solution of the embodiment of the present application can realize the semantic understanding function test of the system to be measured. Moreover, the embodiment of the present application constructs a test database, which includes text data of defined scenarios or services, and text data of undefined scenarios or services; the text data of defined scenarios or services further includes general text data of defined scenarios or services and common text data of defined scenarios or services; the text data of undefined scenarios or services further includes general text data of undefined scenarios or services in the same field, common text data of undefined scenarios or services in the same field, chatting text data, and text data that is meaningless or illogical. At the same time, the present application also designs the specific text types, text quantities, length distributions, etc. of various text data in the above test database. The test data in the test database is text data related to various scenarios or services. Based on this database for semantic understanding function testing, test data matching the real working scenario or business scenario of the system to be measured can be selected from this database, and the selected test data is used to perform semantic understanding function testing on the system to be measured, which can make the test more scientific, closer to the real scenario, and the test effect is better.
[0040] At the same time, a sufficient number of text data can ensure that sufficient, comprehensive, and quantitatively balanced test data is provided for the semantic understanding function test, so as to be able to conduct a comprehensive and in-depth test on the semantic understanding function. Brief Description of the Drawings
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0042] Figure 1 It is a schematic flowchart of an intelligent voice interaction test method provided by an embodiment of the present application;
[0043] Figure 2 It is a schematic flowchart of another intelligent voice interaction test method provided by an embodiment of the present application;
[0044] Figure 3 It is a schematic structural diagram of an intelligent voice interaction test device provided by an embodiment of the present application;
[0045] Figure 4 It is a schematic structural diagram of an intelligent voice interaction test device provided by an embodiment of the present application. Detailed Description of the Embodiments
[0046] The technical solution of the embodiment of the present application is applicable to the intelligent voice interaction test scenario, and is specifically used for testing an intelligent voice interaction system or application, especially for testing the semantic understanding (sub) system in the intelligent voice interaction system or application.
[0047] The above intelligent voice interaction system or application includes, but is not limited to, the voice interaction system or voice interaction function module in artificial intelligence products such as intelligent voice home products, intelligent voice robots, etc.
[0048] The intelligent voice interaction test method proposed by the embodiment of the present application can be exemplarily applied to an intelligent voice interaction test system, device, equipment, etc., and the intelligent voice interaction test system, device, equipment, etc. can all be composed in a combination of hardware and software.
[0049] Preferably, the embodiment of the present application uses a programmable test tool, a test statistics tool, and a resource monitoring tool to implement the specific processing process of the intelligent voice interaction test method. Each of the above tools can also be called by the intelligent voice interaction test system, so as to implement the intelligent voice interaction test method proposed by the embodiment of the present application and complete the test of the intelligent voice interaction system.
[0050] Among them, the above programmable test tool should be able to call the open interface of the system under test, customize the tool configuration file, receive test data (voice data, text data) and input it into the system under test, perform functional tests and corresponding performance tests, and obtain the running result of the system under test in text form, etc.
[0051] The above test statistics tool should be able to automatically count and analyze the system running results of different test items, automatically compare the system running results with the standard result comparison file, etc. In addition, the test statistics tool can also record, analyze, and judge the running results of each test item according to the technical requirements of the system under test to form a test result.
[0052] The above resource monitoring tool should be able to monitor system resource parameters such as the memory, CPU, GPU, and handle count of the system under test.
[0053] Based on the above test tools, according to the function and performance requirements of the system under test and the application scenario, configure the corresponding software and hardware environment to form a test environment. In this test environment, use the programmable test tool and the test statistics tool to input the test data into the system under test in the online / offline state and obtain the running result. Furthermore, record, analyze, and judge the running results of each test item according to the technical requirements of the system under test to form a test result.
[0054] Specifically, according to the specific functions in the intelligent voice interaction application, the intelligent voice interaction test mainly includes the test of the speech recognition (sub) system of the intelligent voice interaction and the test of the semantic understanding (sub) system of the intelligent voice interaction.
[0055] There is still a lack of a national unified standard for the test of the semantic understanding part in the intelligent voice interaction system. Therefore, in the industry's intelligent voice interaction test, the test of the semantic understanding function of the system is relatively chaotic and random, and it is difficult to form a credible test result.
[0056] At the same time, the existing semantic understanding function test solutions in the industry have not formed a systematic and perfect test solution. Among them, the visible semantic understanding function test solutions in the industry have not formed a systematic test database. Usually, when test requirements are generated, data materials are collected temporarily and used for testing. Since the test data is detached from the product application scenario or business scenario, the test data for testing the product will be completely irrelevant to the actual working scenario or business scenario of the product, resulting in the obtained test results being seriously divorced from reality, with low test accuracy, poor scientificity, and unsatisfactory test effects.
[0057] Based on the above technical status quo, the embodiments of this application propose an intelligent voice interaction test method, which standardizes the general test items and general test methods for the semantic understanding (sub) system in the intelligent voice interaction test. By adopting the technical solutions of the embodiments of this application, it is possible for intelligent voice service providers, users, and third-party testing agencies to conduct scientific and credible tests on intelligent voice interaction applications, especially the semantic understanding (sub) system of intelligent voice interaction applications.
[0058] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts fall within the scope of protection of this application.
[0059] First, the references and terms involved in the test method for the semantic understanding (sub) system proposed in the embodiments of this application will be explained.
[0060] The research on the intelligent voice interaction test method proposed in the embodiments of this application refers to some normative documents specified in the industry or by the state. For the convenience of understanding, the documents referred to in the embodiments of this application are listed as follows. The terms and definitions in the embodiments of this application can refer to these documents to determine their actual meanings. Among them, for the referenced documents with a date, only the version corresponding to that date is applicable to the embodiments of this application; for the referenced documents without a date, their latest versions (including all amendments) are applicable to the embodiments of this application.
[0061] References:
[0062] [1]GB / T 5271.29—2006 Information Technology Vocabulary - Part 29: Artificial Intelligence Speech Recognition and Synthesis
[0063] [2]GB / T 2312—1980 Information Interchange Chinese Character Code Character Set - Basic Set
[0064] [3]GB 18030 Information Technology Chinese Coding Character Set
[0065] [4]IETF RFC3629 F.Yergeau., UTF-8, a transformation format of ISO10646, STD 63, RFC 3629, November, 2003
[0066] [5]GB / T 36464 (all parts) Information Technology Intelligent Speech Interaction System
[0067] Meanwhile, the embodiments of the present application also define the following terms:
[0068] Named entity, an entity with a specific or unique reference name.
[0069] Intention, the task or goal that the system needs to execute during the speech interaction process.
[0070] Next, a specific introduction to the test method for the semantic understanding (sub)-system for intelligent speech interaction testing proposed in the embodiments of the present application will be given.
[0071] See Figure 1 As shown, the embodiments of the present application propose the following intelligent speech interaction method:
[0072] S101. Determine the semantic understanding test items according to the functional requirements of the system under test and the semantic understanding test requirements.
[0073] Among them, the system under test mentioned above refers to the intelligent speech interaction system, specifically the semantic understanding (sub)-system in the intelligent speech interaction system.
[0074] The above-mentioned functional requirements for the system under test refer to the requirements for the semantic understanding functions that the system under test should be able to achieve. For example, the system under test should be able to understand the speaker's intention, should be able to recognize named entities and sensitive information from the speaker's speech text, and should be able to modify the text, etc.
[0075] The above semantic understanding test requirements refer to the relevant information about the system under test that needs to be determined through testing. For example, it is necessary to test whether the system under test has certain functions and test the performance of the system under test, etc.
[0076] It can be understood that when the functions that the system under test needs to implement and the semantic understanding test requirements are clear, the test items for the system under test can be determined. For example, test whether the system under test has the required functions and whether the system under test meets the semantic understanding performance requirements, etc.
[0077] The embodiments of the present application set test items in two aspects of function test and performance test for semantic understanding function test of the system under test.
[0078] Among them, the above function test items are mainly used to test whether the system under test provides various functions related to semantic understanding, specifically including intention understanding test, named entity recognition test, sensitive information discrimination test, semantic rejection recognition test, information retrieval test, text similarity calculation test, text modification test, semantic correction test, natural language generation test, logical reasoning test, dialogue guidance test, and multi-round conversation test related to context. In addition, it can also include similar semantic push, etc.
[0079] Based on the above various function test items, when testing a certain function test item of the system under test, it is specifically to test whether the system under test has the function corresponding to the function test item. For example, assuming that the named entity recognition test is performed on the system under test, it is mainly to test whether the system under test has the named entity recognition function.
[0080] The above performance test items are mainly used to test the various performances of the system under test related to semantic understanding, specifically including semantic understanding effect test, semantic understanding efficiency test, and system stability test.
[0081] When testing a certain performance test item of the system under test, it is specifically to test the performance of the system under test when executing the function corresponding to the function test item. For example, assuming that the semantic understanding effect test is performed on the system under test, it is mainly to test the semantic understanding effect of the system under test when executing the function corresponding to a certain function test item. For example, test the precision rate and recall rate when the system under test performs named entity recognition on test data, etc.
[0082] The specific test contents, test methods and other detailed contents of the above function test items and performance test items will be introduced separately in the subsequent embodiments.
[0083] Based on the functional test items and performance test items set in the embodiments of the present application, when determining the semantic understanding test according to the functional requirements of the system under test and the semantic understanding test requirements, the determined semantic understanding test items can include either functional test items or performance test items. The determined functional test items can be at least one of the above-mentioned intention understanding test, named entity recognition test, sensitive information discrimination test, semantic rejection test, information retrieval test, text similarity calculation test, text modification test, semantic correction test, natural language generation test, logical reasoning test, dialogue guidance test, and context-related multi-turn conversation test; the determined performance test items can be at least one of the above-mentioned semantic understanding effect test, semantic understanding efficiency test, and system stability test.
[0084] It can be understood that the selected functional test items should match the functional design of the system under test, that is, functional tests that the system under test can handle should be selected. For example, if the system under test is designed as an information retrieval system, then the functional test for the system under test should be an information retrieval test. Additionally, the named entity recognition test and sensitive information discrimination test can be selected. Since the system under test itself is not designed with a dialogue function, the dialogue guidance test, context-related multi-turn conversation test, etc. are not applicable to the system under test. The selection of performance test items is the same.
[0085] Therefore, the embodiments of the present application only exemplarily show the set of test items specified in the present application. In actual applications, appropriate test items should be selected from them according to the functional requirements and test requirements of the system under test. Since it is impossible to enumerate all the functions and forms of the systems under test, the embodiments of the present application will not list them one by one.
[0086] After determining the semantic understanding test items, for each semantic understanding test item, the corresponding test can be realized by performing the following steps S102 to S104 respectively:
[0087] S102. Select test data corresponding to the semantic understanding test item from the pre-constructed test database, and generate a test data set that matches the semantic understanding test item by using the selected test data.
[0088] Specifically, before the test starts, the embodiments of the present application pre-construct a test database composed of text data through manual writing or collection. In this test database, it includes text data of defined scenarios or services, as well as text data of undefined scenarios or services; the text data of defined scenarios or services includes general text data of defined scenarios or services and common text data of defined scenarios or services; the text data of undefined scenarios or services includes general text data of undefined scenarios or services in the same field, common text data of undefined scenarios or services in the same field, chat text data, and text data that is meaningless or illogical.
[0089] During actual testing, text data can be selected from this test database as test data according to needs. The specific content of the test database will be introduced in detail in the following embodiments.
[0090] It should be noted that during the intelligent voice interaction process, for the system to understand the intention of the speaker, it usually needs to first perform speech recognition on the speaker's voice to determine the speech content of the speaker, and then perform semantic understanding on the speech content of the speaker to understand the intention of the speaker. It can be seen that the semantic understanding (sub)-system in the intelligent voice interaction application actually performs semantic understanding processing on the text obtained by the intelligent voice interaction system through speech recognition. That is to say, the input data of the semantic understanding (sub)-system is text data. Therefore, for the test of the semantic understanding (sub)-system, text data should also be used as test data. Therefore, the embodiments of the present application pre-construct a test database composed of text data to provide data support for the semantic understanding function test.
[0091] It can be understood that when the type of input data of the semantic understanding (sub)-system changes, a test database composed of corresponding type of data can be constructed for the semantic understanding function test. For example, assume that the input of the semantic understanding (sub)-system is voice data. Then, when testing this semantic understanding (sub)-system, a test database composed of voice data is pre-constructed.
[0092] Based on the above test database, when performing the test of a certain semantic understanding test item, select the test data corresponding to this semantic understanding test item from the test database. Among them, the number of test data selected from the test database can be flexibly set according to the test requirements. Then, use the selected test data to generate a test data set that matches this semantic understanding test item.
[0093] Specifically, the test data selected from the test database is only general text corpus. When it is applied to the specific test project scenario, in some cases, it needs to be processed to meet the test requirements of the test project.
[0094] For example, assume that the semantic understanding test item is to test the text modification function of the system under test, that is, to modify the obvious error text content in the test data text. Then the test data should contain the obvious error text content. However, the text data selected from the test database may not contain the obvious error text content. Therefore, it is necessary to change some of the text content in the selected text data to obvious error text content, and then use the modified text as the test data to test the text modification function of the system under test.
[0095] Another example, assume that the semantic understanding test item is to test the effect of named entity recognition of the system under test, that is, to test whether the system under test can correctly recognize the named entities in the test data. To verify whether the system under test correctly recognizes the named entities in the test data, it is necessary to set the named entity annotation labels corresponding to the test data for comparison with the results of the system under test recognizing the named entities, so as to determine whether the named entities recognized by the system under test are correct. Therefore, it is necessary to perform named entity annotation on the selected test data, and then use the selected test data and the named entity labels corresponding to the test data together as the test data to test the effect of named entity recognition of the system under test.
[0096] According to the above introduction, for each semantic understanding test item, select the test data corresponding to the semantic understanding test item from the test database respectively, and process each selected test data to generate the test data matching the semantic understanding test item (using it as the test data after processing, or directly using it as the test data), so as to obtain the test data set corresponding to the semantic understanding test item.
[0097] S103. Through a programmable test tool, input the test data in the test data set into the system under test respectively to test the semantic understanding test item of the system under test.
[0098] Specifically, the programmable test tool calls the input interface of the system under test, inputs each test data in the test data set into the system under test respectively, and calls the function corresponding to the semantic understanding test item of the system under test to realize the test of the semantic understanding test item of the system under test.
[0099] For example, assume that the semantic understanding test item is the named entity recognition function test, specifically to test whether the system under test provides the named entity recognition function. Then input each test data in the test data set corresponding to the test item, that is, the text containing the named entity, into the system under test, so that the system under test performs named entity recognition processing on the input text and outputs the running result of named entity recognition, that is, to realize the named entity recognition test of the system under test.
[0100] S104. Use a test statistics tool to obtain the running results of the system under test during the semantic understanding test project, and determine the test results corresponding to the semantic understanding test project based on the obtained running results and the test content of the semantic understanding test project.
[0101] Specifically, the test statistics tool records the running results of the system under test during the semantic understanding test project, and then determines the test results corresponding to the semantic understanding test project based on the obtained running results and the test content of the semantic understanding test project.
[0102] For example, assume that the semantic understanding test project is named entity recognition test, specifically to test whether the system under test provides the named entity recognition function. After the test data in the test dataset are respectively input into the system under test, the test statistics tool records the running results of the system under test, and determines whether the running results of the system under test are the recognition results of the named entities in the test data. If so, it can be determined that the system under test provides the named entity recognition function; if not, it can be determined that the system under test does not provide the named entity recognition function.
[0103] Another example, assume that the semantic understanding test project is to test the semantic understanding test effect of the system under test, specifically to test the named entity recognition effect of the system under test, that is, to test the accuracy of the named entity recognition of the system under test. After the test data in the test dataset are respectively input into the system under test, the test statistics tool records the running results of the system under test, that is, records the recognition results of the named entity recognition of the test data by the system under test. By comparing the named entity recognition results of the system under test for the test data with the named entity labels corresponding to the test data, the accuracy of the named entity recognition of the system under test is calculated and determined.
[0104] Referring to the above introduction and examples, with the help of programmable test tools and test statistics tools, perform online and / or offline functional tests and / or performance tests on the system under test. That is, through the programmable test tool, input the test datasets corresponding to each semantic understanding test project into the system under test respectively to perform the semantic understanding test project on the system under test, and then the test statistics tool obtains the running results of the system under test under the tests of each test project, and determines the test results corresponding to the test project according to the obtained running results and the specific test content of the test project.
[0105] For the specific test content and test methods of the above functional test projects and performance test projects, refer to the specific test content and test methods of each test project in the following text.
[0106] As can be seen from the above introduction, the intelligent voice interaction test method proposed in the embodiments of this application formulates functional test items and performance test items. When conducting semantic understanding functional tests on the system under test, semantic understanding test items can be selected and determined from each functional test item and / or performance test item according to the functional requirements of the system under test and the semantic understanding test requirements. At the same time, the embodiments of this application also pre-construct a test database, from which test data corresponding to the determined semantic understanding test items can be selected for the semantic understanding functional test of the system under test. After determining the semantic understanding test items and test data, by inputting the test data into the system under test, it is possible to conduct tests on the semantic understanding test items of the system under test, and then, according to the operation results of the system under test during the test process and the test content of the semantic understanding test items, the test results corresponding to the semantic understanding test items can be determined.
[0107] By adopting the technical solution of the embodiments of this application, it is possible to implement the semantic understanding functional test of the system under test. Moreover, the embodiments of this application formulate complete functional test items and performance test items, and can conduct semantic understanding functional tests on the system under test from multiple aspects. The tests are more comprehensive, more scientific, and more credible.
[0108] At the same time, the embodiments of this application propose a systematic and improved test scheme, especially a new design scheme for the construction of test data. By constructing a test database according to the embodiments of this application and selecting test data from it for the semantic understanding functional test of the system under test, the test can be made closer to the real scenario, improving the scientificity and accuracy of the test.
[0109] Next, the above-mentioned test database of the embodiments of this application will be introduced.
[0110] The test database pre-constructed in the embodiments of this application consists of test data. The types and requirements of the test data in this test database can be seen in Table 2. All the test data therein are test texts, and the types and requirements of this test text can be seen in Table 1.
[0111] Table 1 Types and Requirements of Test Texts
[0112]
[0113] Table 2 Types and Requirements of Test Data
[0114]
[0115] As can be seen from the above Table 1 and Table 2, the pre-constructed test database in the embodiments of the present application includes text data of defined scenarios or services, as well as text data of undefined scenarios or services. Among them, the text data of defined scenarios or services refers to the text data whose corresponding scenarios or services have been clearly defined. For example, the dialogue text data in the Q&A scenario, the text data in the intelligent medical service, etc.; while the text data of undefined scenarios or services refers to the text data whose corresponding scenarios or services have not been clearly defined. For example, some text data can only be applied in a certain scenario or service, but the scenario or service does not have a clearly defined scenario name or service name, then this type of text data is the text data of undefined scenarios or services.
[0116] As can be seen from Table 2, the text data of defined scenarios or services, as well as the text data of undefined scenarios or services, both include common text, special text, and abnormal text, and the common text, special text, and abnormal text among them all meet the text types and requirements shown in Table 1.
[0117] As shown in Table 1, the above-mentioned common text includes single-character or word text, short text, single-sentence text, dialogue text, paragraph text, and article text with intention representation, and the number of each type of text is not less than 5.
[0118] The above-mentioned special text includes sensitive information text, named entity text (such as person names, place names, etc., which need to cover named entities related to defined services), special format text (such as numbers, date and time, English upper and lower cases, etc.), specific language text (such as Chinese, English, Korean, etc.), special character set encoding text (such as UTF-8, GBK, etc.), and special symbol text (such as text containing symbols such as commas, periods, question marks, etc.). Among them, the number of sensitive information text and named entity text is not less than 1000 respectively, and the number of special format text, specific language text, special character set encoding text, and special symbol text is not less than 5 respectively;
[0119] The above-mentioned abnormal text includes garbled text (such as UTF-8 and GBK mixed encoding text, etc.) and unsupported language text, and the number of each type of text is not less than 5.
[0120] Referring to Table 2, the text data of defined scenarios or services in the test database is divided into general text data of defined scenarios or services and frequently used text data of defined scenarios or services according to the application frequency;
[0121] Among them, the general text data of defined scenarios or services refers to the text data of defined scenarios or services with general application frequency; the frequently used text data of defined scenarios or services refers to the text data of defined scenarios or services with relatively high application frequency.
[0122] In the general text data of defined scenarios or services, the text data corresponding to each service is not less than 200 pieces; in the common text data of defined scenarios or services, there are at least 3 pieces of real text data corresponding to each service, and continuous collection can be carried out.
[0123] Among them, the above-mentioned real text data refers to the text data generated during the historical operation of the system under test. For example, the real text data corresponding to a certain service can be the data generated during the historical operation of the system under test when performing this task.
[0124] The text data of undefined scenarios or services includes general text data of undefined scenarios or services in the same field, common text data of undefined scenarios or services in the same field, chat text data, and meaningless or illogical text data;
[0125] Among them, the general text data of undefined scenarios or services in the same field refers to the text data of undefined scenarios or services that belong to the same field and have a general application frequency; the common text data of undefined scenarios or services in the same field refers to the text data of undefined scenarios or services that belong to the same field and have a relatively high application frequency.
[0126] In the general text data of undefined scenarios or services in the same field and the common text data of undefined scenarios or services in the same field, there are at least 3 pieces of real text data corresponding to each service. At the same time, the common text data of undefined scenarios or services in the same field can be continuously collected. The number of the above-mentioned chat text data is not less than 1000 pieces, and the average number of characters in each chat text is not less than 5 characters; the above-mentioned meaningless or illogical text data is not less than 100 pieces, and each piece of meaningless or illogical text data is not less than 5 characters.
[0127] Furthermore, when a voice interaction system performs semantic understanding, it usually performs speech recognition on the speaker's voice and then performs semantic understanding processing on the speech recognition result. In order to adapt to the working process and data processing characteristics of the voice interaction system, the general text data of defined scenarios or services, the common text data of defined scenarios or services, the general text data of undefined scenarios or services in the same field, and the common text data of undefined scenarios or services in the same field in the test database should be the text data obtained by the system under test through speech recognition, that is, collect the speech recognition results of the voice interaction system to obtain the above-mentioned test data.
[0128] Moreover, the embodiments of the present application require that the length distribution of all text data of the same type among the general text data of a defined scenario or service, the common text data of a defined scenario or service, the general text data of an undefined scenario or service in the same field, and the common text data of an undefined scenario or service in the same field complies with the set data distribution requirements. For example, the length distribution of all text data of the same type is controlled to comply with the normal distribution. It should be noted that when testing the length distribution of text data, the text length distribution should be statistically analyzed under the condition of a large amount of data, and the text length quantity distribution should be controlled based on this. When the amount of data of the same type of text data is small, the distribution of the average length of the common text should be used as the length distribution of this type of text data, and the text length quantity distribution should be controlled based on this.
[0129] In the common semantic understanding function test schemes in this field, a standardized test data set is not constructed as introduced in the embodiments of the present application above. Usually, when semantic understanding function testing is required, data is collected temporarily for testing. Due to the lack of test data planning, the test data used in the existing semantic understanding function test schemes either cannot cover various business scenarios, or the coverage quantity of various application scenarios is insufficient, or does not comprehensively include various types of text, etc. These deficiencies will directly lead to incomplete, unbalanced, and unscientific testing of the semantic understanding function.
[0130] The intelligent voice interaction test method proposed by the embodiments of the present application can solve the above problems. The intelligent voice interaction test method proposed by the embodiments of the present application constructs a test database, which includes text data of defined scenarios or services, and text data of undefined scenarios or services; the text data of defined scenarios or services includes general text data of defined scenarios or services and common text data of defined scenarios or services; the text data of undefined scenarios or services includes general text data of undefined scenarios or services in the same field, common text data of undefined scenarios or services in the same field, chat text data, and meaningless or illogical text data. At the same time, the present application also designs the specific text types, text quantities, length distributions, etc. of various text data in the above test database. The test data in this test database is text data related to various scenarios or services. Based on this database for semantic understanding function testing, test data matching the real working scenario or business scenario of the system under test can be selected from the database, and the semantic understanding function of the system under test can be tested using the selected test data, which can make the testing more scientific, closer to the real scenario, and have a better testing effect.
[0131] Meanwhile, a sufficient amount of text data can ensure that there is sufficient, comprehensive, and balanced test data for the semantic understanding function test, enabling a comprehensive and in-depth test of the semantic understanding function.
[0132] Through Figure 1 the introduction of the embodiments shown, it can be determined that the test process of the semantic understanding test items proposed in the embodiments of the present application includes the processes of selecting test data, inputting test data, obtaining operation results, and generating test results.
[0133] The following will refer to the above processing logic to elaborate on the specific test contents and test methods of each test item in detail.
[0134] First, the specific test contents and test methods of each functional test item will be introduced. The functional test items are mainly used to test whether the system under test has the functions corresponding to the functional test items. The functional test items of the intelligent voice interaction test method proposed in the embodiments of the present application specifically include:
[0135] 1. Intent understanding test
[0136] The test content of the intent understanding test includes checking whether the system under test provides the function of understanding the speaker's intent. The above function of understanding the speaker's intent includes, but is not limited to, the following specific functions:
[0137] a) Fuzzy recognition: It can correctly handle problems such as typos, synonyms, extra words, and missing words. That is, when there are typos, synonyms, extra words, or missing words in the test text, it can still determine the speaker's intent through fuzzy recognition.
[0138] b) Semantic extraction: It can extract semantic elements and the speaker's key intent, including:
[0139] Named entity extraction, the system under test can automatically extract the named entities expressing the key intent in the text;
[0140] Keyword extraction, the system under test can automatically extract the keywords expressing the intent in the text;
[0141] Semantic relationship extraction, the system under test can automatically extract the triples expressing semantic relationships in the text.
[0142] c) Semantic sorting: The system under test can give multiple sorted understanding results in the semantic understanding result for the speaker to select or reconfirm.
[0143] d) Intent classification: The system under test can predict the speaker's key intent, map the input text data to one or more predefined intents, and mark the intent category to which the text data belongs.
[0144] In the process of specific intention understanding test, at least one of the above functions is tested on the system under test to check whether the system under test provides the corresponding function.
[0145] The specific test method is as follows:
[0146] Select text data from the predefined scenario or business text data in the pre-constructed test database as the test data corresponding to the intention understanding test.
[0147] Input the test data into the system under test through a programmable test tool, enabling the system under test to perform intention understanding processing on the input test data. Then, obtain the operation result of the system under test through a test statistics tool. Determine the operation result based on the obtained operation result and the test content of the above-mentioned intention understanding test, that is, determine whether the operation result of the system under test during the test is the operation result of understanding the speaker's intention. If so, it indicates that the system under test provides the function of understanding the speaker's intention; otherwise, it indicates that the system under test does not provide the function of understanding the speaker's intention, thus obtaining the test result.
[0148] Furthermore, it is also possible to examine whether the speaker's intention obtained from the operation of the system under test is correct to determine whether the system under test provides the function of understanding the speaker's intention. For example, when the speaker's intention obtained from the operation of the system under test is correct, it is determined that the system under test provides the function of understanding the speaker's intention; if the speaker's intention obtained from the operation of the system under test is incorrect, it is determined that the system under test does not provide the function of understanding the speaker's intention.
[0149] Among them, whether the speaker's intention obtained from the operation of the system under test is correct can be determined by comparing the result of understanding the speaker's intention obtained from the operation of the system under test with the speaker's intention label corresponding to the test data.
[0150] 2. Named Entity Recognition Test
[0151] The test content of the named entity recognition test includes checking whether the system under test provides the function of finding and accurately annotating named entities in the text, that is, testing whether the system under test can identify named entities from the test text data.
[0152] The specific test method is as follows:
[0153] Select named entity text from the pre-constructed test database as the test data for the named entity recognition test. Among them, the named entity text selected from the test database can be the named entity text included in any type of test data. In theory, as long as the text data contains named entities, it can be used as the test data for the named entity recognition test.
[0154] Test data is input into the system under test through a programmable test tool, enabling the system under test to perform named entity recognition processing on the test data. The operation results of the system under test are obtained through a test statistics tool. Based on the obtained operation results and the test content of the above-mentioned named entity recognition test, the operation results are judged, that is, it is judged whether the operation result of the system under test is the recognition result of the named entities in the test data. If so, it indicates that the system under test provides the named entity recognition function; if not, it indicates that the system under test does not provide the named entity recognition function, thereby obtaining the test result.
[0155] Furthermore, it is also possible to examine whether the named entity recognition result obtained by the system under test is correct to determine whether the system under test provides the named entity recognition function. For example, when the named entity recognition result obtained by the system under test is correct, it is determined that the system under test provides the named entity recognition function; if the named entity recognition result obtained by the system under test is incorrect, it is determined that the system under test does not provide the named entity recognition function.
[0156] Among them, whether the named entity recognition result obtained by the system under test is correct can be determined by comparing the named entity recognition result obtained by the system under test with the named entity label corresponding to the test data.
[0157] 3. Sensitive Information Discrimination Test
[0158] The test content of the sensitive information discrimination test includes checking whether the system under test provides the function of distinguishing sensitive content in the input text according to the context, that is, testing whether the system under test can identify sensitive information content from the test text.
[0159] The specific test method is as follows:
[0160] Select sensitive information texts from a pre-constructed test database as the test data for the sensitive information discrimination test. Among them, the sensitive information texts selected from the test database can be sensitive information texts included in any type of test data. In theory, as long as the text data contains sensitive information, it can be used as the test data for the sensitive information discrimination test. The above-mentioned sensitive information specifically includes content related to porn, violence, terrorism, etc.
[0161] Test data is input into the system under test through a programmable test tool, enabling the system under test to perform sensitive information discrimination processing on the test data. The operation result of the system under test is obtained through a test statistics tool. The operation result is judged based on the obtained operation result and the test content of the above-mentioned sensitive information discrimination test, that is, to determine whether the operation result of the system under test is the recognition result of the sensitive information in the test data. If so, it indicates that the system under test provides a sensitive information discrimination function; if not, it indicates that the system under test does not provide a sensitive information discrimination function, thereby obtaining the test result.
[0162] Furthermore, it is also possible to examine whether the sensitive information recognition result obtained by the system under test is correct to determine whether the system under test provides a sensitive information discrimination function. For example, when the sensitive information recognition result obtained by the system under test is correct, it is determined that the system under test provides a sensitive information discrimination function; if the sensitive information recognition result obtained by the system under test is incorrect, it is determined that the system under test does not provide a sensitive information discrimination function.
[0163] Among them, whether the sensitive information recognition result obtained by the system under test is correct can be determined by comparing the sensitive information recognition result obtained by the system under test with the sensitive information label corresponding to the test data.
[0164] 4. Semantic Rejection Recognition Test
[0165] The test content of the semantic rejection recognition test includes checking whether the system under test provides a function of distinguishing and rejecting invalid text input content that cannot be processed or should not be processed, that is, testing whether the system under test rejects when invalid text that cannot be processed or should not be processed is input into the system under test, that is, whether the operation result of the system under test is a rejection result. Among them, the above-mentioned content that cannot be processed includes content that is not supported by the system under test or is irrelevant to the business; the above-mentioned content that should not be processed includes completely meaningless content.
[0166] The specific test method is as follows:
[0167] Select text from the text data of undefined scenarios or services in the pre-constructed test database, specifically select invalid text that cannot be processed or should not be processed, as the test data for the semantic rejection recognition test.
[0168] Test data is input into the system under test through a programmable test tool, enabling the system under test to perform semantic recognition processing on the test data. The operation result of the system under test is obtained through a test statistics tool. The operation result is judged based on the obtained operation result and the test content of the above-mentioned semantic rejection recognition test, that is, to determine whether the operation result of the system under test is a rejection result. If so, it indicates that the system under test provides a semantic rejection recognition function; if not, it indicates that the system under test does not provide a semantic rejection recognition function, thereby obtaining the test result.
[0169] 5. Information Retrieval Test
[0170] The test content of the information retrieval test includes checking whether the system under test provides an information retrieval function. The information retrieval includes at least personalized dictionary retrieval and / or third-party information source retrieval and / or custom knowledge base retrieval, that is, retrieving information related to the test data from the personalized dictionary and / or third-party information source and / or custom knowledge base.
[0171] Among them, personalized dictionary retrieval can specifically be contact list retrieval, song list retrieval, point of interest retrieval, etc.; third-party information source retrieval can specifically be weather retrieval, flight retrieval, hotel retrieval, stock retrieval, etc.
[0172] The specific test method is as follows:
[0173] Select text data from the text data of the predefined scenarios or services in the pre-constructed test database as the test data for the information retrieval test.
[0174] Input the test data into the system under test through a programmable test tool, enable the system under test to perform information retrieval processing based on the test data, and obtain the operation result of the system under test through a test statistics tool. Determine the operation result according to the obtained operation result and the test content of the above information retrieval test, that is, judge whether the operation result of the system under test is to retrieve information related to the test data from the retrieval database (such as personalized dictionary, third-party information source, custom knowledge base, etc.). If so, it indicates that the system under test provides an information retrieval function; if not, it indicates that the system under test does not provide an information retrieval function, thereby obtaining the test result.
[0175] Furthermore, it is also possible to examine whether the information retrieval result obtained by the system under test is correct to determine whether the system under test provides an information retrieval function. For example, when the information retrieval result obtained by the system under test is information related to the test data, it is determined that the system under test provides an information retrieval function; if the information retrieval result obtained by the system under test is not information related to the test data, it is determined that the system under test does not provide an information retrieval function.
[0176] 6. Text Similarity Calculation Test
[0177] The test content of the text similarity calculation test includes checking whether the system under test provides a function to calculate the degree of semantic information consistency between the input text data and the existing text, that is, testing whether the system under test has the function to determine the type of semantic information consistency between the input text data and the existing text. The types of semantic information consistency include, but are not limited to, the following specific contents:
[0178] a) The words used in the sentence have changed, but the semantic information is similar. That is, the words used in the input text data have changed compared to the existing text, but the semantic information of the two is similar.
[0179] b) The sentence structure has changed, but the semantic information is similar. That is, the sentence structure of the input text data has changed compared to the existing text, but the semantic information of the two is similar.
[0180] c) The words and structure of the sentence are similar, but the semantic information is not similar. That is, the words and structure of the input text data are similar to the existing text, but the semantic information of the two is not similar.
[0181] The specific test method is as follows:
[0182] Select text data from the text data of the predefined scenarios or services in the pre - constructed test database as the test data for the text similarity calculation test.
[0183] Input the test data into the system under test through a programmable test tool, enabling the system under test to perform text similarity calculation processing on the test data and the existing text, and obtain the operation result of the system under test through a test statistics tool. Determine the operation result according to the obtained operation result and the test content of the above - mentioned text similarity calculation test, that is, judge whether the operation result of the system under test is the semantic information consistency type of the input text data and the existing text. If so, it indicates that the system under test provides the text similarity calculation function; if not, it indicates that the system under test does not provide the text similarity calculation function, thereby obtaining the test result.
[0184] Furthermore, it is also possible to examine whether the text similarity calculation result obtained from the operation of the system under test is correct to determine whether the system under test provides the text similarity calculation function. For example, when the semantic information consistency type of the test data obtained from the operation of the system under test and the existing text is correct, it is determined that the system under test provides the text similarity calculation function; if the semantic information consistency type of the test data obtained from the operation of the system under test and the existing text is incorrect, it is determined that the system under test does not provide the text similarity calculation function.
[0185] Among them, whether the semantic information consistency type of the test data obtained from the operation of the system under test and the existing text is correct can be determined by comparing the semantic information consistency type of the test data obtained from the operation of the system under test and the existing text with the pre - labeled semantic information consistency type label of the test data and the existing text.
[0186] 7. Text modification test
[0187] The test content of the text modification test includes checking whether the system under test provides the function of modifying the previous sentence text in the dialogue, that is, testing whether the system under test can combine the obtained subsequent sentence text to modify the previous sentence text in the dialogue scenario. Specifically, it can be to modify the incorrect text content in the text or to modify the whole text.
[0188] The specific test method is as follows:
[0189] Select text data from the text data of the predefined scenarios or services in the pre-constructed test database as the test data for the text modification test.
[0190] Input the test data into the system under test through a programmable test tool, so that the system under test modifies and processes the previous sentence text in the dialogue text, and obtain the operation result of the system under test through a test statistics tool. Determine the operation result according to the obtained operation result and the above test content of the text modification test, that is, judge whether the operation result of the system under test is the result of modifying the previous sentence text in the dialogue. If so, it means that the system under test provides the text modification function; if not, it means that the system under test does not provide the text modification function, thus obtaining the test result.
[0191] Furthermore, it is also possible to examine whether the text modification result obtained by the system under test is correct to determine whether the system under test provides the text modification function. For example, when the text modification result of modifying the previous sentence text in the dialogue obtained by the system under test is correct, it is determined that the system under test provides the text modification function; if the text modification result of modifying the previous sentence text in the dialogue obtained by the system under test is incorrect, it is determined that the system under test does not provide the text modification function.
[0192] Among them, whether the text modification result of modifying the previous sentence text in the dialogue obtained by the system under test is correct can be determined by comparing the text modification result of modifying the previous sentence text in the dialogue obtained by the system under test with the pre-labeled text modification label of modifying the previous sentence text in the dialogue.
[0193] 8. Semantic Correction Test
[0194] The test content of the semantic correction test includes checking whether the system under test provides the function of automatically correcting the results of semantic understanding errors, that is, testing whether the system under test can automatically correct the results of semantic understanding errors corresponding to the test data according to the test data.
[0195] Among them, semantic understanding errors include syntactic errors, Chinese word segmentation errors, coreference resolution errors, etc.
[0196] The specific test method is as follows:
[0197] Select text data from the text data of predefined scenarios or services in the pre-constructed test database as the test data for text modification testing. At the same time, create semantic recognition results with semantic understanding errors corresponding to the test data, which are also used as test data.
[0198] Input the test data and the semantic recognition results with semantic understanding errors corresponding to the test data into the system under test through a programmable test tool, so that the system under test processes and modifies the input semantic recognition results corresponding to the test data according to the input test data, and obtain the operation result of the system under test through a test statistics tool. Determine the operation result according to the obtained operation result and the test content of the above semantic correction test, that is, judge whether the operation result of the system under test is the result of correcting the input semantic recognition result. If so, it indicates that the system under test provides a semantic correction function; if not, it indicates that the system under test does not provide a semantic correction function, thereby obtaining the test result.
[0199] Furthermore, it is also possible to examine whether the semantic correction result obtained from the operation of the system under test is correct to determine whether the system under test provides a semantic correction function. For example, when the semantic correction result of the semantic recognition result corresponding to the test data obtained from the operation of the system under test is correct, it is determined that the system under test provides a semantic correction function; if the semantic correction result of the semantic recognition result corresponding to the test data obtained from the operation of the system under test is incorrect, it is determined that the system under test does not provide a semantic correction function.
[0200] Among them, whether the semantic correction result of the semantic recognition result corresponding to the test data obtained from the operation of the system under test is correct can be determined by comparing the semantic correction result of the semantic recognition result corresponding to the test data obtained from the operation of the system under test with the correct semantic recognition result corresponding to the test data.
[0201] 9. Natural Language Generation Test
[0202] The test content of the natural language generation test includes checking whether the system under test provides the function of generating natural language texts that conform to the speaker's intention and meet the requirements of voice interaction responses based on the semantic understanding results, that is, testing whether the system under test can generate natural language texts that meet the above requirements according to the test data. Among them, the generated natural language texts include but are not limited to simple reply texts, reply texts based on predefined templates, reply texts that understand and conform to the speaker's intention, and reasonable, guiding, or recommended reply texts when the speaker's intention is not clear.
[0203] The specific test method is as follows:
[0204] Select text data from the text data of defined scenarios or businesses in a pre-built test database as test data for natural language generation testing.
[0205] Input the test data into the system under test through a programmable test tool, enabling the system under test to generate natural language text based on the input test data, and obtain the operation result of the system under test through a test statistics tool. Determine the operation result based on the obtained operation result and the test content of the above natural language generation test, that is, judge whether the operation result of the system under test is natural language text. If so, it indicates that the system under test provides the natural language generation function; if not, it indicates that the system under test does not provide the natural language generation function, thereby obtaining the test result.
[0206] Furthermore, it is also possible to examine whether the natural language text obtained from the operation of the system under test conforms to the speaker's intention and meets the requirements of voice interaction response to determine whether the system under test provides the natural language generation function. For example, when the natural language text obtained from the operation of the system under test conforms to the speaker's intention and meets the requirements of voice interaction response, it is determined that the system under test provides the natural language generation function; if the natural language text obtained from the operation of the system under test does not conform to the speaker's intention and / or does not meet the requirements of voice interaction response, it is determined that the system under test does not provide the natural language generation function.
[0207] 10. Logical reasoning test
[0208] The test content of the logical reasoning test includes checking whether the system under test provides the function of logical calculation and derivation for text content, that is, testing whether the system under test can perform logical calculation and derivation based on the test data to obtain a logical reasoning result. For example, assuming the test data is "2020", then through logical reasoning, the logical reasoning result of "2020 is a leap year" can be obtained; assuming the test data is "father's mother", then through logical reasoning, the logical reasoning result of "father's mother is called grandma" can be obtained.
[0209] The specific test method is as follows:
[0210] Select text data from the text data of defined scenarios or businesses in a pre-built test database as test data for the logical reasoning test.
[0211] Test data is input into the system under test through a programmable test tool, enabling the system under test to perform logical reasoning based on the input test data. The operating results of the system under test are obtained through a test statistics tool. Based on the obtained operating results and the test content of the above-mentioned logical reasoning test, the operating results are judged, that is, it is determined whether the operating results of the system under test are the logical reasoning results obtained through logical reasoning based on the test data. If so, it indicates that the system under test provides a logical reasoning function; if not, it indicates that the system under test does not provide a logical reasoning function, thereby obtaining the test results.
[0212] Furthermore, it is also possible to examine whether the logical reasoning results obtained by the system under test are correct to determine whether the system under test provides a logical reasoning function. For example, when the logical reasoning results obtained by the system under test are correct, it is determined that the system under test provides a logical reasoning function; if the logical reasoning results obtained by the system under test are incorrect, it is determined that the system under test does not provide a logical reasoning function.
[0213] Among them, whether the logical reasoning results output by the system under test are correct can be determined by comparing the operating results of the system under test with the correct logical reasoning results obtained based on the test data.
[0214] 11. Dialogue guidance test
[0215] The test content of the dialogue guidance test includes checking whether the system under test provides the function of dynamically generating guiding prompt words according to the speaker's intention and scenario requirements to guide the user to state their ultimate purpose. That is, during the conversation with the user, according to the user's intention and scenario requirements, guiding prompt words for the user are generated, enabling the user to have a conversation according to the guiding prompt words, thereby stating their ultimate purpose. The output of the guiding prompt words can enable the user to state their intention or purpose to the intelligent question-and-answer system more directly and efficiently.
[0216] Among them, the guiding prompt words output by the intelligent voice interaction system include, but are not limited to, the following content:
[0217] a) Personalized dictionary, that is, selecting text from the personalized dictionary as the guiding prompt words.
[0218] b) Information mined and classified according to the user's behavior habits, that is, according to the user's historical conversation content, mining the user's frequently used information, and thus mining guiding prompt words that conform to the current scenario based on the user's frequently used information. For example, assuming that the user often books air tickets from place A to place B through the intelligent voice interaction system, then during the current user conversation, assuming that the user inputs the conversation content of "Help me book a plane ticket" to the intelligent voice interaction system, the system can output the guiding prompt word "Do you want to book a plane ticket from place A to place B?"
[0219] c) The knowledge within the defined knowledge base, that is, retrieving the text content that conforms to the current scenario from the defined knowledge base as the guiding prompt.
[0220] d) Third-party source information, that is, retrieving relevant information from third-party sources to generate guiding prompts. For example, assuming that the user inputs the conversation content of "Help me book a flight from Place A to Place B" into the intelligent voice interaction system, the system can retrieve the flight information from Place A to Place B from the Civil Aviation Administration website. For example, if the available flight information is retrieved, at this time, generate the guiding prompt of "The available flights from Place A to Place B are as follows. Which flight do you want to book?"
[0221] e) The associated information retrieved from the massive data, that is, retrieving the associated information of the user's conversation content from the massive database as the guiding prompt.
[0222] f) Rejection prompt, that is, when the content input by the user is something that the system cannot process or should not process, the system outputs a rejection prompt to the user as a guiding prompt to the user, so as to prompt the user to input the correct information. For example, assuming that the intelligent voice interaction system is a fast food ordering system, if the user inputs the conversation content of "Help me book a flight" to the system, the system can output the guiding prompt of "Sorry, I don't understand your needs".
[0223] The specific test method is as follows:
[0224] Select text data from the text data of the defined scenarios or services and / or undefined scenarios or services in the pre-constructed test database as the test data for the dialogue guidance test.
[0225] Input the test data into the system under test through a programmable test tool, so that the system under test generates guiding prompts according to the input test data, and obtain the operation result of the system under test through a test statistics tool. Determine the operation result according to the obtained operation result and the test content of the above dialogue guidance test, that is, judge whether the operation result of the system under test is the guiding prompt generated based on the test data. If so, it means that the system under test provides the dialogue guidance function. If not, it means that the system under test does not provide the dialogue guidance function, so as to obtain the test result.
[0226] Furthermore, it is also possible to examine whether the guiding prompt words obtained from the operation of the system under test conform to the speaker's intention and the requirements of the scenario, so as to determine whether the system under test provides a dialogue guiding function. For example, when the guiding prompt words obtained from the operation of the system under test conform to the speaker's intention and the requirements of the scenario, that is, when it can achieve the role of guiding the user, it is determined that the system under test provides a dialogue guiding function; if the guiding prompt words obtained from the operation of the system under test do not conform to the speaker's intention and / or do not conform to the requirements of the scenario, it is determined that the system under test does not provide a dialogue guiding function.
[0227] 12. Context-related multi-turn conversation test
[0228] The test content of the context-related multi-turn conversation test includes checking whether the system under test provides the function of context-related multi-turn conversation processing, that is, testing whether the system under test can conduct context-related multi-turn conversations with the user.
[0229] When the intelligent voice interaction system conducts context-related multi-turn conversations with the user, it should at least be able to achieve the following functions during the multi-turn conversation process:
[0230] a) Dialogue state tracking, that is, being able to track the progress of the conversation in real time. For example, when the conversation pauses or the paused conversation restarts, it can respond in real time.
[0231] b) Dialogue strategy management, that is, being able to conduct conversations with the user with dialogue strategies that conform to the current scenario. For example, when the user's input dialogue content has clearly exceeded the capabilities of the intelligent voice interaction system, it should adopt the strategy of interrupting the conversation in a timely manner to improve the conversation efficiency.
[0232] c) Dialogue intention switching and jumping, that is, when the user's dialogue intention switches or jumps, being able to adaptively adjust the dialogue content, or being able to actively switch the dialogue intention to conduct conversations with the user according to the user's current dialogue intention. For example, assuming that the user's previous dialogue intention is to book a flight through the intelligent voice interaction system, after the flight booking is completed, the system can output the dialogue content "Do you need me to book a hotel for you" to achieve the purpose of switching the dialogue intention.
[0233] d) Historical information inheritance, that is, classifying and mining the historical dialogue content with the user, so as to be able to select the content that conforms to the current dialogue intention from the historical dialogue content during the conversation with the user and achieve multi-turn conversations with the user.
[0234] The specific test method is as follows:
[0235] Select text data from the text data of defined scenarios or services in a pre-built test database as the test data for the dialogue guidance test. It should be noted that since it is necessary to test whether the system under test can provide a multi-round conversation function related to the context, the selected test data is presented in the form of test data groups, and each test data group includes multiple pieces of context-related test data.
[0236] Input each group of test data into the system under test through a programmable test tool, so that the system under test generates conversation content according to the input test data. For example, assume that the data in a certain test data group includes three pieces of test data {A1, A2, A3}, and the three pieces of test data have a clear context relationship, such as the context order of A1→A2→A3. Then input A1, A2, and A3 into the system under test in sequence, so that the system under test generates conversation content according to each input test data. For example, when A1 is input, the system under test generates conversation content a1; when A2 is input, the system under test generates conversation content a2; when A3 is input, the system under test generates conversation content a3. The above process is the multi-round conversation between the system under test and the user.
[0237] Input test data into the system under test in the above manner, and obtain the operation result of the system under test through a test statistics tool. Judge the operation result according to the obtained operation result and the test content of the above context-related multi-round conversation test, that is, judge whether the operation result of the system under test is to output a context-related multi-round conversation based on the test data. If so, it means that the system under test provides a context-related multi-round conversation function; if not, it means that the system under test does not provide a context-related multi-round conversation function, so as to obtain the test result.
[0238] Furthermore, it is also possible to examine whether the conversation content obtained by running the system under test meets the requirements related to the context to determine whether the system under test provides a dialogue guidance function. For example, examine whether the conversation content obtained by running the system under test during the test meets some or all of the above requirements for dialogue state tracking, dialogue strategy management, dialogue intention switching or jumping, and historical information inheritance to determine whether the conversation content output by the system under test is context-related conversation content. If so, it is determined that the system under test provides a context-related multi-round conversation function; if not, it is determined that the system under test does not provide a context-related multi-round conversation function.
[0239] Each functional test item in the intelligent voice interaction test method proposed in the embodiments of the present application tests functions such as intention understanding, named entity recognition, sensitive information discrimination, semantic rejection recognition, information retrieval, text similarity calculation, text modification, semantic correction, natural language generation, logical reasoning, dialogue guidance, and context-related multi-turn conversations of the object under test. These functions basically cover various functions that may need to be implemented in the semantic understanding scenario. When actually applying the technical solutions of the embodiments of the present application, testers can flexibly select test items from the above-mentioned functional test items according to the functional requirements or test purposes of the system under test to perform functional tests on the system under test. Therefore, it can meet the semantic understanding tests for various scenarios and various requirements.
[0240] Conventional semantic understanding test schemes usually target a specific object under test and execute functional test items that match the object under test. Moreover, the specific processing content of each functional test item is relatively simple. Therefore, the existing semantic understanding test schemes are not universal. When the object under test changes, they are usually not applicable and need to redesign the test items. Moreover, their test content is not comprehensive enough to accurately test the true state of the object under test.
[0241] Compared with the defects of the existing semantic understanding test schemes, the intelligent voice interaction test method proposed in the embodiments of the present application stipulates more comprehensive functional test items and rich and specific test content for each test item. When facing any object under test, testers can use the semantic understanding test scheme proposed in the embodiments of the present application to test various functions of the object under test, so as to more deeply and comprehensively test all aspects of the system under test in the intention understanding scenario and obtain more rigorous, more accurate, and more scientific test results.
[0242] Next, the test content and test methods of each performance test item will be introduced. Performance testing is mainly used to test the performance of the system under test when executing functions corresponding to the functional test items. The performance test items of the intelligent voice interaction test method proposed in the embodiments of the present application specifically include:
[0243] 1. Semantic understanding effect test
[0244] The test content of the semantic understanding effect test includes detecting the semantic understanding effect of the system under test when performing functions, specifically including detecting at least one of precision, recall rate, rejection rate, accuracy rate, F1 value, mean reciprocal rank, and normalized discounted cumulative gain of the system under test when performing functions. Among them, the system under test performing functions specifically means that the system under test performs the functions it possesses, such as the system under test performing named entity recognition functions, intent understanding functions, sensitive information discrimination functions, etc. In the embodiments of the present application, the system under test performing functions may also refer to the processing performed by the system under test during function testing. For example, assuming that the system under test is performing a named entity recognition test, it can be considered that the system under test is performing a named entity recognition function, that is, performing named entity recognition processing.
[0245] The following introduces each test index in the semantic understanding effect test:
[0246] a) Precision
[0247] Precision is used to detect the semantic understanding ability of the system under test, that is, the ratio of the number of times the system under test actually responds correctly to valid texts to the total number of times all texts are responded correctly. The calculation formula is as follows:
[0248]
[0249] Among them, P SS represents the semantic understanding precision, and N SS represents the number of times the valid text is actually responded correctly; N S represents the total number of times all texts are responded correctly.
[0250] b) Recall rate
[0251] The recall rate is used to detect the semantic understanding ability of the system under test, that is, the ratio of the number of times the system under test actually responds correctly to valid texts to the total number of times that should be responded correctly. The calculation formula is as follows:
[0252]
[0253] Among them, R SS represents the semantic understanding recall rate, and N SS represents the number of times the valid text is actually responded correctly; N SC represents the total number of times the valid text should be responded correctly.
[0254] c) Rejection rate
[0255] The rejection rate is used to detect the semantic rejection ability of the system under test, that is, the ratio of the number of times the system under test correctly responds to invalid text to the total number of times of invalid text input. Among them, invalid text includes text data that the system under test does not support or is irrelevant to the business and completely meaningless noise data. The formula for calculating the rejection rate is as follows:
[0256]
[0257] Among them, SR represents the semantic rejection rate, and N SR represents the number of times the invalid text is actually correctly responded; N R represents the total number of times of invalid text input.
[0258] d) Accuracy rate
[0259] The accuracy rate is used to detect the semantic understanding ability of the system under test, that is, the ratio of the number of times the system under test correctly responds to all text to the total number of times of all text responses. The formula for calculating it is as follows:
[0260]
[0261] Among them, A SS represents the semantic understanding accuracy rate, and N SS represents the number of times the valid text is actually correctly responded, and N SR represents the number of times the invalid text is actually correctly responded; N represents the total number of all text responses.
[0262] e) F1 value
[0263] The F1 value is used to detect the semantic understanding ability of the system under test, that is, the weighted harmonic mean of the semantic understanding precision rate and the semantic understanding recall rate. The formula for calculating it is as follows:
[0264]
[0265] Among them, F1 represents the semantic understanding F1 value, and P SS represents the semantic understanding precision rate, and R SS represents the semantic understanding recall rate.
[0266] f) Mean Reciprocal Rank
[0267] The Mean Reciprocal Rank is used to detect the information retrieval ability of the system under test, that is, the average of the reciprocals of the ranking positions of the correct results in the results given by the system under test. The formula for calculating it is as follows:
[0268]
[0269] Among them, MRR represents the Mean Reciprocal Rank, Q represents the total number of information retrievals, i represents the i-th information retrieval, and rank iIndicates the sorting position where the correct result appears in the i-th information retrieval.
[0270] g) Normalized Discounted Cumulative Gain
[0271] Normalized Discounted Cumulative Gain is used to detect the information retrieval ability of the system under test, that is, the ratio of the sorting relevance score of the results given by the system under test to the sorting relevance score of the ideal results. Its calculation formula is as follows:
[0272]
[0273]
[0274] NDCG = DCG / IDCG
[0275] Among them, DCG represents Discounted Cumulative Gain, that is, the sorting relevance score of the results given by the system under test. K represents the number of information retrieval results. j represents the j-th retrieval result, and rel j represents the relevance score of the j-th retrieval result. IDCG represents the Discounted Cumulative Gain of the ideal results, that is, the sorting relevance score of the ideal results. |REL K | represents that the number of information retrieval results is sorted from large to small according to the relevance score. NDCG represents Normalized Discounted Cumulative Gain.
[0276] It can be understood that the above test indicators are used to detect the performance of the system under test when performing different functions. Therefore, when testing the semantic understanding effect of the system under test, the test indicators applicable to the function should be selected according to the function performed by the system under test for the semantic understanding effect test.
[0277] The effect test indicators applicable to each function are shown in Table 3:
[0278] Table 3 Different functions and their applicable effect test indicators
[0279]
[0280] Referring to Table 3, it can be understood that when the function performed by the system under test is the intent understanding function, to detect the semantic understanding effect of the system under test when performing the function, at least the precision rate and recall rate of the system under test when performing semantic extraction on the test data need to be detected. In addition, one or more of the rejection rate, accuracy rate, and F1 value of the system under test when performing semantic extraction on the test data can also be detected. Among them, for the effect test of the intent understanding function, it is only carried out on the semantic extraction function of the system, that is, only the semantic extraction function is tested. As long as the semantic extraction is correct, it is considered that the intent understanding result of the system under test is correct.
[0281] When the function executed by the system under test is named entity recognition, to detect the semantic understanding effect of the system under test during function execution, at least the precision and recall of the system under test for named entity recognition of test data need to be detected. In addition, one or more of the rejection rate, accuracy, and F1 value of the system under test for named entity recognition of test data can also be detected. Among them, for the effect test of the named entity recognition function, only the named entity recognition effect of the system is tested. As long as the named entity recognition result is correct, it is considered that the intention understanding result of the system under test is correct.
[0282] When the function executed by the system under test is sensitive information discrimination, to detect the semantic understanding effect of the system under test during function execution, at least the precision and recall of the system under test for distinguishing sensitive content in the input text need to be detected. In addition, one or more of the rejection rate, accuracy, and F1 value of the system under test for distinguishing sensitive content in the input text can also be detected.
[0283] When the function executed by the system under test is semantic rejection, to detect the semantic understanding effect of the system under test during function execution, it includes detecting the rejection rate of the system under test for distinguishing and rejecting invalid text input content that cannot be processed or should not be processed.
[0284] When the function executed by the system under test is information retrieval, to detect the semantic understanding effect of the system under test during function execution, at least the normalized discounted cumulative gain of the system under test for retrieving information matching the test data needs to be detected. In addition, one or more of the precision, recall, rejection rate, and reciprocal of the average rank of the system under test for retrieving information matching the test data can also be detected.
[0285] When the function executed by the system under test is text correction, to detect the semantic understanding effect of the system under test during function execution, at least the recall of the system under test for correcting the previous sentence text in the dialogue needs to be detected. Further, one or more of the precision, rejection rate, accuracy, and F1 value can also be detected.
[0286] When the function executed by the system under test is semantic correction, to detect the semantic understanding effect of the system under test during function execution, at least the precision and recall of the system under test for automatically correcting the results of semantic understanding errors need to be detected. One or more of the rejection rate, accuracy, and F1 value can also be detected.
[0287] When the function executed by the system under test is logical reasoning, to detect the semantic understanding effect of the system under test during function execution, at least the precision and recall of the system under test for logical calculation and derivation of text content need to be detected. One or more of the rejection rate, accuracy, and F1 value can also be detected. [[ID=:19]]
[0288] When the function executed by the system under test is a context - related multi - turn dialogue function, to detect the semantic understanding effect of the system under test when executing the function, at least the accuracy rate of the system under test providing context - related multi - turn conversation processing needs to be detected, and one or more of precision rate, recall rate, rejection rate, and F1 value can also be detected. Among them, in the effect test of the context - related multi - turn dialogue function, whether the dialogue finally achieves the speaker's intention should be selected to judge whether the dialogue is correct.
[0289] In addition, it should be noted that for some functions executed by the system under test, such as text similarity calculation, natural language generation, dialogue guidance, etc., there may be no clear standard to define whether the system operation result is correct. Therefore, the above - mentioned effect test indicators are not applicable to these functions, and other test items are needed to test these functions. For example, the semantic understanding efficiency test introduced in the embodiments below of this application can be used for testing.
[0290] Referring to the above introduction, when the system under test executes a function, its semantic understanding effect can be tested, that is, calculate each applicable effect test indicator when the system under test executes the function. The specific test method is as follows:
[0291] First, select test data from a pre - constructed test database (including text data with defined scenarios or services and text data with undefined scenarios or services), and perform manual annotation on each test data to obtain the label corresponding to the test data. This label serves as a standard result comparison file for comparing with the operation result of the system under test to judge whether the operation result of the system under test is correct.
[0292] Then, according to the function and performance requirements of the system under test and the application scenario, configure the corresponding software and hardware environment to obtain an adapted test environment.
[0293] Under the above - mentioned test environment, input the test data into the system under test in online / offline state through a programmable test tool, and obtain the operation result of the system under test through a test statistics tool. And, according to the applicable relationship given in Table 3, obtain the value of the effect evaluation indicator when the system under test executes the function. For example, input the named entity text into the system under test and obtain the operation result of the system under test. By comparing the operation result of the system under test with the label corresponding to the input text data, judge whether the named entity recognition result of the system under test is correct. According to this judgment result, the precision rate and recall rate when the system under test executes the named entity recognition function can be calculated.
[0294] Finally, according to the system operation results and the test results of semantic understanding effects, a test result file is compiled and generated, which includes the test set name, the number of test sets, the results of index items, etc. If the operation results of the system under test meet the technical requirements of the system under test or relevant standards and specifications, the test passes; otherwise, it fails.
[0295] As a preferred test method, an embodiment of the present application proposes that the above-mentioned semantic understanding effect test can be carried out while performing a function test on the system under test.
[0296] When testing a function test item of the system under test, test data is input into the system under test to make the system under test execute the function corresponding to the function test item, and the operation results of the system under test are obtained. On the one hand, the operation results are used to verify whether the system under test provides the function corresponding to the function test item to complete the test of the function test item; on the other hand, the operation results are compared with the labels corresponding to the test data to determine whether the operation results are correct. Furthermore, according to the correctness judgment of the operation results corresponding to each test data, the semantic understanding effect evaluation index matching the function executed by the system under test in the function test item can be calculated, and then the test result of the semantic understanding effect can be obtained. For the selection of test data, the construction of the test environment and the test process in this test process, please refer to the above introduction.
[0297] It can be understood that performing the semantic understanding effect test while performing the function test on the system under test can obtain the test results of two different test items in the same test process, improving the test efficiency.
[0298] In a conventional semantic understanding test scheme, when testing the semantic understanding effect of the object under test, it is usually only simply tested whether the object under test can accurately understand the user's semantics, that is, usually the semantic understanding accuracy rate in the object under test is tested. However, the semantic understanding accuracy rate cannot fully represent the semantic understanding effect of the object under test.
[0299] When the intelligent voice interaction test method proposed by an embodiment of the present application performs a semantic understanding effect test on the system under test, it tests from aspects such as the precision rate, recall rate, rejection rate, accuracy rate, F1 value, average reciprocal rank, and normalized discounted cumulative gain of the semantic understanding of the system under test, and standardizes the applicability of the above-mentioned various test indicators to various functions. Based on the above technical solutions of the embodiments of the present application, when a user performs a semantic understanding effect test on the system under test, the applicable test indicators can be selected according to the functions of the system under test to test the semantic understanding effect of the system under test, so that the semantic understanding effect of the system under test with any function can be tested from multiple dimensions, and objective and comprehensive test results can be obtained.
[0300] 2. Semantic Understanding Efficiency Test
[0301] The test content of semantic understanding efficiency test includes detecting the semantic understanding efficiency of the system under test when executing functions, specifically including one or more of the parameters such as the average response time of semantic understanding, the distribution of semantic understanding response time, and the semantic understanding throughput rate when the system under test is executing functions.
[0302] The specific meanings of the above parameters are as follows:
[0303] a) Average response time of semantic understanding
[0304] The semantic understanding response time refers to the time when the system under test gives the semantic understanding result of a piece of text after inputting a piece of text; the average response time of semantic understanding is the ratio of the semantic understanding response time of all test data to the total number of all input test data on the test data set, and its calculation formula is:
[0305]
[0306] Among them, T avu represents the average response time of semantic understanding, W represents the test data set, T i represents the semantic understanding response time corresponding to test sample i, and N represents the total number of input test data.
[0307] b) Distribution of semantic understanding response time
[0308] The distribution of speech understanding response time refers to the distribution and proportion of the semantic understanding response time of all test data in the test data set in each response time interval. As a preferred method, the embodiments of the present application count the proportion of the semantic understanding response time of all test data in the test data set in the interval below 100ms, the proportion in the interval of 100ms - 200ms, and the proportion in the interval above 200ms to obtain the distribution of semantic understanding response time. When actually applying the technical solutions of the embodiments of the present application, the above response time intervals can be set as needed.
[0309] c) Semantic understanding throughput rate
[0310] The semantic understanding throughput rate refers to the size of the text for which semantic understanding is achieved by the system under test within the unit response time, that is, the efficiency of inputting a large amount of (business-related) test text data set at one time and giving the semantic understanding result at one time, and its calculation formula is as follows:
[0311]
[0312] Among them, TP represents the semantic understanding throughput rate, W represents the test data set, S i represents the size (kb) of the text corresponding to test sample i on the test data set, T iIndicates the semantic understanding response time corresponding to test sample i.
[0313] Referring to the above introduction, when the system under test executes its functions, its semantic understanding efficiency can be tested, that is, one or more of the parameters such as the average semantic understanding response time, the semantic understanding response time distribution, and the semantic understanding throughput rate when the system under test executes its functions are calculated. The specific test method is as follows:
[0314] First, select test data from a pre-constructed test database (including text data of defined scenarios or services and text data of undefined scenarios or services) to obtain a test data set.
[0315] Then, according to the function and performance requirements of the system under test and the application scenario, configure the corresponding software and hardware environment to obtain an adapted test environment.
[0316] In the above test environment, input the test data into the system under test in the online / offline state through a programmable test tool, and obtain the operation results of the system under test through a test statistics tool. Also, calculate the values of various semantic understanding efficiency evaluation indicators when the system under test executes its functions. For example, input the named entity text into the system under test, obtain the operation results of the system under test, and calculate and determine the average semantic understanding response time, the semantic understanding response time distribution, the semantic understanding throughput rate, etc. of the system under test by statistically analyzing the response time, the size of the file that can be processed, etc. during named entity recognition.
[0317] Finally, according to the system operation results and the semantic understanding efficiency test results, organize and generate a test result file, which includes the test set name, the number of test sets, the results of the index items, etc. If the operation results of the system under test meet the technical requirements of the system under test or relevant standard specifications, the test passes; otherwise, it fails.
[0318] As a preferred test method, the embodiment of the present application proposes that the above semantic understanding efficiency test can be carried out simultaneously with the function test of the system under test.
[0319] When conducting tests on the functional test items of the system under test, test data is input into the system under test to enable the system under test to execute the functions corresponding to the functional test items, and the operation results of the system under test are obtained. On the one hand, the operation results are used to verify whether the system under test provides the functions corresponding to the functional test items to complete the tests of the functional test items; on the other hand, based on the obtained operation results of the system under test, the values of various semantic understanding efficiency evaluation indicators when the system under test executes functions are calculated, that is, the average semantic understanding response time, the semantic understanding response time distribution, the semantic understanding throughput rate, etc. when the system under test executes functions are calculated, and then the semantic understanding efficiency test results can be obtained. The selection of test data, the construction of the test environment, and the test process in this test process can refer to the above introduction.
[0320] It can be understood that when conducting semantic understanding efficiency tests while conducting functional tests on the system under test, the test results of two different test items can be obtained in the same test process, improving the test efficiency.
[0321] In the conventional semantic understanding test scheme, when testing the semantic understanding efficiency of the object under test, usually only the response time for the object under test to understand the user's semantics is simply tested. If the time is shorter, the semantic understanding efficiency is higher. However, simply using the semantic understanding response time cannot objectively reflect the semantic understanding efficiency of the object under test.
[0322] When the intelligent voice interaction test method proposed in the embodiments of the present application conducts semantic understanding efficiency tests on the system under test, it tests from aspects such as the average response time, response time distribution, and throughput rate of the semantic understanding of the system under test. The embodiments of the present application can conduct detailed statistical analysis on the response time of the semantic understanding of the system under test, and can analyze the semantic understanding efficiency of the system under test in combination with the semantic understanding throughput rate, so more objective and accurate test results can be obtained.
[0323] 3. System stability test
[0324] The test content of the system stability test includes detecting whether the system operation situation and resource usage situation are stable when the system under test executes functions under the conditions of set software and hardware configurations and system concurrent paths. Among them, the above resource usage situation at least includes the usage rates of system physical memory, virtual memory, CPU, GPU, handles, and network resources.
[0325] Specifically, the stability test of the system under test includes a stable operation test and a resource usage test.
[0326] Among them, the stable operation test refers to detecting the ability of the system under test to run all its functions without crashing, freezing, or malfunctioning and to be able to continue running normally under the conditions of given software and hardware configurations and system concurrent paths.
[0327] Resource usage testing refers to the ability to detect that when the tested system runs all its functions under the given software and hardware configurations and the number of concurrent system paths, the resource usage rates of system physical memory, virtual memory, CPU, GPU, handles, network resources, etc. remain continuously stable. The resource usage situation of the tested system can be monitored through resource monitoring tools.
[0328] It should be noted that the above-mentioned given software and hardware configurations and the number of concurrent system paths need to meet the normal operation ability of the tested system, so as to prevent the tested system from being restricted by the software and hardware configurations and the number of concurrent system paths and being unable to exert its normal ability, affecting the objectivity of the test.
[0329] The specific test method for system stability testing is as follows:
[0330] First, select test data from the test database to form a test data set, and clarify the software and hardware configurations and the number of concurrent system paths. Among them, since the tested system needs to be continuously tested, the quantity of test data should be sufficient to support the requirements of the entire stability test for test data.
[0331] Then, according to the functional and performance requirements of the tested system, as well as the application scenario and the clarified software and hardware configurations and the number of concurrent system paths, configure the corresponding software and hardware environment to obtain an adapted test environment.
[0332] Under the above-mentioned test environment, through a programmable test tool, continuously input test data to the tested system in a loop for 7 days in the online scenario and 3 days in the offline scenario without interruption, so that the tested system executes all its functions that it can execute under the given software and hardware configurations and the number of concurrent system paths, and through a test statistics tool and a resource monitoring tool, monitor the running situation and resource usage situation of the tested system in real time.
[0333] Among them, the reason for choosing to test for 7 days in the online scenario and 3 days in the offline scenario is that usually, any semantic understanding system or device will not run continuously online for more than 7 days, nor will it run continuously offline for more than 3 days. Moreover, the online running situation of the tested system is more important than the offline running situation. Therefore, it is necessary to conduct more continuous online tests on the tested system. Therefore, the embodiments of the present application control the tested system to run continuously online for 7 days and continuously offline for 3 days to fully test its running stability in the online scenario and the offline scenario. When actually applying the technical solutions of the embodiments of the present application, the test duration in the online scenario and the offline scenario can also be flexibly adjusted with reference to the technical concept of the embodiments of the present application.
[0334] Specifically, the operation results of the system under test during the test are obtained through a test statistics tool to determine whether the system under test is operating normally. For example, if the system under test can continuously output operation results, it indicates that it is operating normally; if the system under test stops outputting operation results, it indicates that an abnormality has occurred in its operation. Through a resource monitoring tool, the resource usage of the system under test during the test can be monitored.
[0335] Finally, according to the system stability test results, a test result file is compiled and generated, which includes the test set name, the number of test sets, software and hardware configurations, the number of concurrent system paths, and the results of metric items, etc. If the operation results of the system under test meet the technical requirements of the system under test or relevant standard specifications, the test passes; otherwise, it fails.
[0336] Compared with the conventional semantic understanding test scheme that does not perform stability tests on the object under test or only tests its stability in a single scenario, the intelligent voice interaction test method proposed in the embodiments of the present application can perform stability tests on the system under test separately from the offline scenario and the online scenario, so as to be able to more comprehensively test the stability of the system under test in various working scenarios.
[0337] In summary, the intelligent voice interaction test method proposed in the embodiments of the present application specifies the specific test items for semantic understanding tests, including functional test items and performance test items, and also specifies the specific item contents of each test item and the specific processing procedures for each item content.
[0338] In the existing semantic understanding test schemes, various test items are not clearly specified. Only individual test items are selected for testing within the scope of business, and the test items are not comprehensive, or the test contents of some test items are not comprehensive. When these test schemes are applied to the tests of other semantic understanding systems or devices, they are often unable to be applied due to the mismatch between the test items and the functions of the systems or devices.
[0339] Based on the comprehensive and rich test items specified by the intelligent voice interaction test method proposed in the embodiments of the present application, testers can select test items that are applicable to the system or device under test or match the test purpose according to the test requirements, and perform test processing according to the specific processing contents of each test item recorded in the embodiments of the present application, so as to achieve semantic understanding tests that meet the requirements for any test object.
[0340] The above embodiments respectively introduce the content of the test database, the specific test content and the specific test method of each test item in the intelligent voice interaction test method proposed by this application. Based on the introduction of the above embodiments, by selecting test data from the test database and test items from each test item, and inputting the test data and obtaining the system operation results through programmable test tools and test statistics tools, the test results can be sorted out according to the system operation results and the test content of the test items, so as to realize scientific and effective testing of the system under test.
[0341] The intelligent voice interaction test method introduced above is essentially an objective test method. However, intelligent voice interaction applications need to interact with users. Therefore, on the basis of the above test method, the embodiments of this application also propose to conduct a subjective experience test on the system under test to further improve the comprehensiveness and scientificity of intelligent voice interaction testing.
[0342] See Figure 2 As shown, the intelligent voice interaction test method proposed by the embodiments of this application further includes:
[0343] S205. Determine the target speaker intention according to the functional requirements of the system under test and the semantic understanding test requirements.
[0344] Specifically, according to the scenario or business requirements, the functional requirements of the system under test, and the semantic understanding test requirements, the speaker intention is generated in a random manner as the target speaker intention for the subjective experience test of the system under test.
[0345] S206. Generate a subjective experience test data set according to the functional requirements of the system under test, the semantic understanding test requirements, and the target speaker intention.
[0346] Specifically, according to the functional requirements of the system under test and the semantic understanding test requirements, the test data is obtained by means of manual writing or collection, or by selecting from a pre-constructed test database, so as to obtain the subjective experience test data set.
[0347] S207. Have testers conduct a subjective experience test on the system under test according to the subjective experience test data set to obtain a subjective experience test result.
[0348] Specifically, no less than 20 testers of different genders, different age groups, and different educational backgrounds select test data from the subjective experience test data set, input it into the system under test through a programmable test tool, and have a conversation with the system under test, so as to conduct a subjective experience test on the system under test. Through the log analysis and statistics tool, the conversation situation between the system under test and the user is statistically analyzed to obtain the subjective experience test result.
[0349] Among them, the specific test content and test process for testers to conduct subjective experience tests on the system under test include:
[0350] The tester selects test data from the subjective experience test dataset and inputs it into the system under test, so as to have a conversation with the system under test. During the conversation, the average number of dialogue turns for the system under test to understand the intention of the target speaker specified by the tester, the task completion rate in one conversation, and the satisfaction of the tester with the entire conversation process are recorded through a log analysis and statistics tool.
[0351] Among them, the average number of dialogue turns for the system under test to understand the intention of the target speaker specified by the tester refers to the situation where the tester has a conversation with the system under test regarding the intention of one target speaker. When the tester believes that the system under test has understood the intention of the target speaker expressed by the tester during the conversation, the conversation stops. The average of the number of dialogue turns in multiple such conversation processes is the average number of dialogue turns for the system under test to understand the intention of the target speaker specified by the tester. This average number of dialogue turns can characterize the efficiency of the system under test in understanding the speaker's intention.
[0352] The above-mentioned average number of dialogue turns can be calculated and determined through the following formula:
[0353]
[0354] Among them, R avg represents the average number of dialogue turns, N represents the total number of request conversations, and R n represents the number of interaction turns in each conversation.
[0355] The above-mentioned task completion rate in one conversation refers to the success rate of the system under test in understanding at least one intention of the target speaker specified by the tester during the process of having a conversation with the system under test.
[0356] During one conversation process, the tester can specify multiple intentions of the target speaker, that is, during one conversation process, the tester can state multiple intentions of the target speaker, and each intention of the target speaker is represented as a task. If the system under test successfully understands one intention of the target speaker, it is considered that the system under test has completed this task. Therefore, during one call between the system under test and the tester, by statistically calculating the success rate of the system under test in understanding at least one intention of the target speaker specified by the tester, the task completion rate of the system under test in one conversation can be determined.
[0357] The above-mentioned task completion rate can be calculated and determined through the following formula:
[0358]
[0359] Among them, Rreach It represents the task completion rate in a conversation. M represents the total number of requested tasks, and R m represents the number of tasks achieved in each conversation. There can be multiple tasks in a conversation.
[0360] The satisfaction of the tester with the entire conversation process refers to the satisfaction level evaluation given by the tester to the system under test after completing the subjective experience test of the system under test. Considering aspects such as task completion and response speed, it is generally divided into: very satisfied, satisfied, average, dissatisfied, very poor.
[0361] Finally, based on the average number of conversation turns for the system under test to understand the intended target speaker's intention specified by the tester, the task completion rate in a conversation, and the satisfaction of the tester with the entire conversation process, the subjective experience test result of the system under test is determined.
[0362] Exemplarily, the above information such as the average number of conversation turns, the task completion rate in a conversation, and the satisfaction of the tester with the entire conversation process corresponding to the system under test can be jointly used as the subjective experience test result of the system under test.
[0363] The higher the task completion rate of the system under test in a conversation, the more favorable the subjective experience test result of the system under test; when the task completion rates in a conversation are the same, the fewer the average number of conversation turns for the system under test to understand the intended target speaker's intention specified by the tester, the more favorable the subjective experience test result of the system under test; when the task completion rate in a conversation and the average number of conversation turns for the system under test to understand the intended target speaker's intention specified by the tester are both the same, the higher the satisfaction of the tester with the entire conversation process with the system under test, the more favorable the subjective experience test result of the system under test.
[0364] For the above-mentioned subjective experience test conducted by the tester on the system under test, the software and hardware environment should be configured according to the functional and performance requirements of the system under test and the usage scenario, so as to obtain a suitable test environment.
[0365] S208. Generate the semantic understanding function test result of the system under test according to the test results corresponding to each semantic understanding test item and the subjective experience test result.
[0366] Specifically, as an exemplary implementation, the test results corresponding to each speech understanding test item and the subjective experience test result of the system under test can be integrated as the semantic understanding function test result of the system under test.
[0367] Alternatively, it is also possible to integrate and quantify the test results according to the test results corresponding to each semantic understanding test item and the subjective experience test results according to preset quantification rules, and finally obtain a quantified test result for the semantic understanding function of the system under test, such as obtaining a test score for the semantic understanding function test of the system under test, etc.
[0368] In addition, Figure 2 Steps S201 to S204 in the method embodiment shown respectively correspond to Figure 1 Steps S101 to S104 in the method embodiment shown. For the specific content, please refer to Figure 1 the introduction of the embodiment shown, which will not be elaborated here. [[ID=…]]
[0369] When the embodiment of the present application conducts a semantic understanding test on the system under test, manual test content is also added. Combining this manual test with the above-mentioned functional test and performance test can more comprehensively test the system under test from both subjective and objective aspects, thereby obtaining more accurate test results. The above-mentioned semantic understanding test scheme combining subjective test and objective test is a pioneer in this field.
[0370] Corresponding to the above intelligent voice interaction test method, the embodiment of the present application also proposes an intelligent voice interaction test device. Refer to Figure 3 shown, the device includes:
[0371] A project selection unit 100, configured to determine semantic understanding test items according to the functional requirements of the system under test and the semantic understanding test requirements;
[0372] A test data set generation unit 110, configured to select test data corresponding to the semantic understanding test items from a pre-constructed test database, and generate a test data set matching the semantic understanding test items by using the selected test data;
[0373] Wherein, the test database includes text data of defined scenarios or services, and text data of undefined scenarios or services;
[0374] The text data of the defined scenarios or services includes general text data of the defined scenarios or services and common text data of the defined scenarios or services; among the general text data of the defined scenarios or services, the text data corresponding to each service is not less than 200 pieces; among the common text data of the defined scenarios or services, there are at least 3 pieces of real text data corresponding to each service;
[0375] The text data of the undefined scenarios or services includes general text data of undefined scenarios or services in the same field, common text data of undefined scenarios or services in the same field, chat text data, and meaningless or illogical text data; in the general text data of undefined scenarios or services in the same field and the common text data of undefined scenarios or services in the same field, there are at least 3 pieces of real text data corresponding to each service; the number of the chat text data is not less than 1000 pieces, and the average number of characters in each chat text is not less than 5 characters; the number of the meaningless or illogical text data is not less than 100 pieces, and each piece is not less than 5 characters.
[0376] The general text data of the defined scenarios or services, the common text data of the defined scenarios or services, the general text data of undefined scenarios or services in the same field, and the common text data of undefined scenarios or services in the same field are all composed of common text, special text, and abnormal text obtained by the system under test through speech recognition. Moreover, the length distribution of all text data belonging to the same type conforms to the set quantity distribution requirements.
[0377] The common text includes single-word or word-and-phrase texts, short-text texts, single-sentence texts, dialogue texts, paragraph texts, and article texts with intention expressions, and the number of each type of text is not less than 5 pieces.
[0378] The special text includes sensitive information texts, named entity texts, special format texts, specific language texts, special character set encoding texts, and special symbol texts. Among them, the number of sensitive information texts and named entity texts is not less than 1000 pieces respectively, and the number of special format texts, specific language texts, special character set encoding texts, and special symbol texts is not less than 5 pieces respectively.
[0379] The abnormal text includes garbled texts and unsupported language texts, and the number of each type of text is not less than 5 pieces.
[0380] The test processing unit 120 is used to input the test data in the test dataset into the system under test respectively through a programmable test tool to test the semantic understanding test items of the system under test.
[0381] The test statistics unit 130 is used to obtain the operation result of the system under test when conducting the test of the semantic understanding test item through a test statistics tool, and determine the test result corresponding to the semantic understanding test item based on the obtained operation result and the test content of the semantic understanding test item.
[0382] The intelligent voice interaction test device proposed in the embodiments of the present application formulates functional test items and performance test items. When conducting a semantic understanding function test on the system under test, semantic understanding test items can be selected and determined from each functional test item and / or performance test item according to the functional requirements of the system under test and the semantic understanding test requirements. At the same time, the embodiments of the present application also pre-construct a test database, from which test data corresponding to the determined semantic understanding test items can be selected for the semantic understanding function test of the system under test. After determining the semantic understanding test items and test data, by inputting the test data into the system under test, a test of the semantic understanding test items of the system under test can be realized. Furthermore, according to the operation results of the system under test during the test process and the test content of the semantic understanding test items, the test results corresponding to the semantic understanding test items can be determined.
[0383] By using the device in the embodiments of the present application, the semantic understanding function test of the system under test can be realized. Moreover, the embodiments of the present application formulate complete functional test items and performance test items, and can conduct semantic understanding function tests on the system under test from multiple aspects. The test is more comprehensive, more scientific, and has higher credibility.
[0384] Optionally, the semantic understanding test items include functional test items and / or performance test items. The functional test items are used to test whether the system under test has functions corresponding to the functional test items, and specifically include at least one of intention understanding test, named entity recognition test, sensitive information discrimination test, semantic rejection recognition test, information retrieval test, text similarity calculation test, text modification test, semantic correction test, natural language generation test, logical reasoning test, dialogue guidance test, and context-related multi-round conversation test;
[0385] The performance test items are used to test the performance of the system under test when executing functions corresponding to the functional test items, and specifically include at least one of semantic understanding effect test, semantic understanding efficiency test, and system stability test.
[0386] Optionally, the test content of the intention understanding test includes checking whether the system under test provides a function to understand the intention of the speaker. The function of understanding the intention of the speaker includes at least one of functions such as fuzzy recognition of text, semantic extraction, semantic sorting, and intention classification;
[0387] The test content of the named entity recognition test includes checking whether the system under test provides a function to find and accurately label named entities in the text;
[0388] The test content of the sensitive information discrimination test includes checking whether the system under test provides a function to distinguish sensitive content in the input text according to the context;
[0389] The test content of semantic rejection recognition test includes checking whether the system under test provides the function of distinguishing and rejecting invalid text input content that cannot be processed or should not be processed;
[0390] The test content of information retrieval test includes checking whether the system under test provides an information retrieval function, and the information retrieval includes at least personalized dictionary retrieval and / or third-party information source retrieval and / or custom knowledge base retrieval;
[0391] The test content of text similarity calculation test includes checking whether the system under test provides the function of calculating the degree of semantic information consistency between the input text data and the existing text;
[0392] The test content of text modification test includes checking whether the system under test provides the function of modifying the previous sentence text in the conversation;
[0393] The test content of semantic correction test includes checking whether the system under test provides the function of automatically correcting the results of semantic understanding errors;
[0394] The test content of natural language generation test includes checking whether the system under test provides the function of generating natural language text that conforms to the speaker's intention and meets the requirements of voice interaction response according to the semantic understanding results;
[0395] The test content of logical reasoning test includes checking whether the system under test provides the function of performing logical calculations and derivations on the text content;
[0396] The test content of dialogue guidance test includes checking whether the system under test provides the function of dynamically generating guiding prompt words according to the speaker's intention and scenario requirements to guide the user to state their ultimate purpose;
[0397] The test content of context-related multi-turn conversation test includes checking whether the system under test provides the function of context-related multi-turn conversation processing;
[0398] Correspondingly, selecting the test data corresponding to the semantic understanding test items from the pre-constructed test database includes:
[0399] When the semantic understanding test item is any test item among intention understanding test, information retrieval test, text similarity calculation test, text modification test, semantic correction test, natural language generation test, logical reasoning test, and context-related multi-turn conversation test, select text data from the text data of the defined scenarios or services in the pre-constructed test database;
[0400] When the semantic understanding test item is named entity recognition test, select named entity text from the pre-constructed test database;
[0401] When the semantic understanding test item is a sensitive information discrimination test, select sensitive information text from a pre-constructed test database;
[0402] When the semantic understanding test item is a semantic rejection recognition test, select text data from the text data of undefined scenarios or services in a pre-constructed test database;
[0403] When the semantic understanding test item is a dialogue guidance test, select text data from the text data of defined scenarios or services, and / or the text data of undefined scenarios or services in a pre-constructed test database.
[0404] Optionally, the test content of the semantic understanding effect test includes detecting the semantic understanding effect of the system under test when executing functions, specifically including detecting at least one of the precision rate, recall rate, rejection recognition rate, accuracy rate, F1 value, mean reciprocal rank, and normalized discounted cumulative gain of the system under test when executing functions.
[0405] The calculation formula of the rejection recognition rate is as follows:
[0406]
[0407] Among them, SR represents the semantic rejection recognition rate, N SR represents the number of times the actual response to invalid text is correct; N R represents the total number of invalid text inputs;
[0408] The calculation formula of the F1 value is as follows:
[0409]
[0410] Among them, F1 represents the semantic understanding F1 value, P SS represents the semantic understanding precision rate, R SS represents the semantic understanding recall rate;
[0411] The calculation formula of the mean reciprocal rank is as follows:
[0412]
[0413] Among them, MRR represents the mean reciprocal rank, Q represents the total number of information retrievals, i represents the i-th information retrieval, and rank i represents the sorting position where the correct result appears in the i-th information retrieval;
[0414] The calculation formula of the normalized discounted cumulative gain is as follows:
[0415]
[0416]
[0417] NDCG = DCG / IDCG
[0418] Among them, DCG represents the discounted cumulative gain, which is the ranking correlation score of the results given by the system under test. K represents the number of information retrieval results, j represents the j-th retrieval result, and rel j represents the relevance score of the j-th retrieval result. IDCG represents the discounted cumulative gain of the ideal results, which is the ranking correlation score of the ideal results. |REL K | represents that the number of information retrieval results is sorted in descending order according to the relevance score. NDCG represents the normalized discounted cumulative gain.
[0419] When the function executed by the system under test is the intent understanding function, detect the semantic understanding effect of the system under test when executing the function, including at least detecting the precision rate and recall rate when the system under test performs semantic extraction on the test data;
[0420] When the function executed by the system under test is the named entity recognition function, detect the semantic understanding effect of the system under test when executing the function, including at least detecting the precision rate and recall rate when the system under test performs named entity recognition on the test data;
[0421] When the function executed by the system under test is the sensitive information discrimination function, detect the semantic understanding effect of the system under test when executing the function, including at least detecting the precision rate and recall rate when the system under test distinguishes sensitive content in the input text;
[0422] When the function executed by the system under test is the semantic rejection recognition function, detect the semantic understanding effect of the system under test when executing the function, including detecting the rejection rate when the system under test distinguishes and rejects invalid text input content that cannot be processed or should not be processed;
[0423] When the function executed by the system under test is the information retrieval function, detect the semantic understanding effect of the system under test when executing the function, including at least detecting the normalized discounted cumulative gain when the system under test retrieves information matching the test data;
[0424] When the function executed by the system under test is the text correction function, detect the semantic understanding effect of the system under test when executing the function, including at least detecting the recall rate when the system under test corrects the previous text in the dialogue;
[0425] When the function executed by the system under test is the semantic correction function, detect the semantic understanding effect of the system under test when executing the function, including at least detecting the precision rate and recall rate when the system under test automatically corrects the results with semantic understanding errors;
[0426] When the function executed by the system under test is a logical reasoning function, the semantic understanding effect of the system under test during function execution is detected, including at least detecting the precision rate and recall rate when the system under test performs logical calculations and derivations on text content;
[0427] When the function executed by the system under test is a context-related multi-turn dialogue function, the semantic understanding effect of the system under test during function execution is detected, including at least detecting the accuracy rate when the system under test provides context-related multi-turn conversation processing.
[0428] Optionally, the test content of the semantic understanding efficiency test includes detecting the semantic understanding efficiency of the system under test during function execution, specifically including at least detecting at least one of the average semantic understanding response time, semantic understanding response time distribution, and semantic understanding throughput rate of the system under test during function execution;
[0429] Among them, the average semantic understanding response time refers to the ratio of the semantic understanding response time of all test data in the test data set to the total number of all test data, and the semantic understanding response time refers to the time when the system under test gives the semantic understanding result of a piece of text after inputting a piece of text;
[0430] The calculation formula for the average semantic understanding response time is as follows:
[0431]
[0432] Among them, T avu represents the average semantic understanding response time, W represents the test data set, T i represents the semantic understanding response time corresponding to test sample i, and N represents the total number of input test data;
[0433] The semantic understanding response time distribution refers to the distribution and proportion of the semantic understanding response times of all test data in the test data set in each response time interval;
[0434] The semantic understanding throughput rate refers to the text size for which the system under test achieves semantic understanding within the unit response time;
[0435] The calculation formula for the semantic understanding throughput rate is as follows:
[0436]
[0437] Among them, TP represents the semantic understanding throughput rate, W represents the test data set, S i represents the size (kb) of the text corresponding to test sample i on the test data set, and T i represents the semantic understanding response time corresponding to test sample i.
[0438] Optionally, the test content of the system stability test includes detecting whether the system operation and resource usage are stable when the system under test executes functions under the set software and hardware configurations and the system concurrency; the resource usage at least includes the usage rates of the system's physical memory, virtual memory, CPU, GPU, handles, and network resources, and the resource usage is monitored by a resource monitoring tool;
[0439] Performing a system stability test on the system under test includes:
[0440] Through a programmable test tool, continuously input test data into the system under test in a loop for 7 days in the online scenario and 3 days in the offline scenario without interruption, so that the system under test executes functions under the set software and hardware configurations and the system concurrency, and real-time monitor the operation and resource usage of the system under test through a test statistics tool and a resource monitoring tool.
[0441] Optionally, the device further includes:
[0442] A subjective experience test unit, which is used to determine the target speaker's intention according to the functional requirements of the system under test and the semantic understanding test requirements; generate a subjective experience test data set according to the functional requirements of the system under test, the semantic understanding test requirements, and the target speaker's intention; have a tester perform a subjective experience test on the system under test according to the subjective experience test data set to obtain a subjective experience test result; generate a semantic understanding function test result for the system under test according to the test results corresponding to each semantic understanding test item and the subjective experience test result.
[0443] Optionally, the step of having a tester perform a subjective experience test on the system under test according to the subjective experience test data set to obtain a subjective experience test result includes:
[0444] Have a tester have a conversation with the system under test according to the subjective experience test data set, record the average number of conversation turns for the system under test to understand the target speaker's intention specified by the tester, the task completion rate in one conversation, and the tester's satisfaction with the entire conversation process; where the task completion rate in one conversation refers to the success rate of the system under test in understanding at least one target speaker's intention specified by the tester during a conversation between the tester and the system under test;
[0445] Determine the subjective experience test result for the system under test according to the average number of conversation turns for the system under test to understand the target speaker's intention specified by the tester, the task completion rate in one conversation, and the tester's satisfaction with the entire conversation process.
[0446] Specifically, for the specific content of each embodiment of the above intelligent voice interaction test device and the specific working content of each unit, please refer to the content of the above method embodiments, which will not be elaborated here.
[0447] Another embodiment of the present application further provides an intelligent voice interaction test device. Refer to Figure 4 As shown, the device includes:
[0448] A memory 200 and a processor 210;
[0449] Among them, the memory 200 is connected to the processor 210 and is used to store programs;
[0450] The processor 210 is used to implement the intelligent voice interaction test method disclosed in any of the above embodiments by running the programs stored in the memory 200.
[0451] Specifically, the above intelligent voice interaction test device may further include: a bus, a communication interface 220, an input device 230, and an output device 240.
[0452] The processor 210, the memory 200, the communication interface 220, the input device 230, and the output device 240 are interconnected through the bus. Among them:
[0453] The bus may include a path for transmitting information between various components of the computer system.
[0454] The processor 210 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0455] The processor 210 may include a main processor and may also include a baseband chip, a modem, etc.
[0456] The program for implementing the technical solution of the present invention is stored in the memory 200, and the operating system and other key services can also be stored. Specifically, the program may include program code, and the program code includes computer operation instructions. More specifically, the memory 200 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash memory, and so on.
[0457] The input device 230 may include devices for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor, etc.
[0458] The output device 240 may include devices for allowing information to be output to a user, such as a display screen, a printer, a speaker, etc.
[0459] The communication interface 220 may include any transceiver-like device for communicating with other devices or communication networks, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.
[0460] The processor 2102 executes the program stored in the memory 200 and calls other devices, and can be used to implement the steps of the intelligent voice interaction test method provided by the embodiments of the present application.
[0461] Another embodiment of the present application further provides a storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the intelligent voice interaction test method provided by any of the above embodiments are implemented.
[0462] Specifically, for the specific working content of each part of the above intelligent voice interaction test device, and for the specific processing content when the computer program on the above storage medium is run by a processor, reference can be made to the content of each embodiment of the above intelligent voice interaction test method, which will not be elaborated here.
[0463] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be adopted in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0464] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For device embodiments, since they are basically similar to method embodiments, they are described relatively simply. For related parts, reference can be made to the corresponding descriptions in the method embodiments.
[0465] The steps in the methods of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs. The technical features recorded in each embodiment can be replaced or combined.
[0466] In the embodiments of the present application, the modules and sub-modules in the devices and terminals can be combined, divided, and deleted according to actual needs.
[0467] In several embodiments provided by the present application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are only illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.
[0468] The modules or sub-modules described as separate components may or may not be physically separated. The components serving as modules or sub-modules may or may not be physical modules or sub-modules, that is, they can be located in one place, or distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0469] In addition, in each embodiment of the present application, the functional modules or sub-modules can be integrated in a processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated in one module. The above-mentioned integrated modules or sub-modules can be implemented in the form of hardware, or in the form of software functional modules or sub-modules.
[0470] Those skilled in the art may further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0471] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software units executed by a processor, or a combination of the two. The software units can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0472] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0473] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An intelligent voice interaction test method, characterized in that, Including: Determine semantic understanding test items according to the functional requirements and semantic understanding test requirements of the system under test; For each semantic understanding test item, perform the following test processes respectively: Select test data corresponding to the semantic understanding test item from a pre-constructed test database, and generate a test data set that matches the semantic understanding test item using the selected test data; Among them, the test database includes text data of defined scenarios or services, as well as text data of undefined scenarios or services; The text data of the defined scenarios or services includes general text data of the defined scenarios or services and common text data of the defined scenarios or services; in the general text data of the defined scenarios or services, the text data corresponding to each service is not less than 200 pieces; in the common text data of the defined scenarios or services, there are at least 3 pieces of real text data corresponding to each service; The text data of the undefined scenarios or services includes general text data of undefined scenarios or services in the same field, common text data of undefined scenarios or services in the same field, chat text data, and meaningless or illogical text data; in the general text data of undefined scenarios or services in the same field and the common text data of undefined scenarios or services in the same field, there are at least 3 pieces of real text data corresponding to each service; the number of the chat text data is not less than 1000 pieces, and the average number of characters in each chat text is not less than 5 characters; the number of the meaningless or illogical text data is not less than 100 pieces, and each piece is not less than 5 characters; The general text data of the defined scenarios or services, the common text data of the defined scenarios or services, the general text data of undefined scenarios or services in the same field, and the common text data of undefined scenarios or services in the same field are all composed of common text, special text, and abnormal text obtained by speech recognition of the system under test, and the length distribution of all text data belonging to the same type meets the set quantity distribution requirements; The common text includes single-word or word text, short text, single-sentence text, dialogue text, paragraph text, and article text with intention representation, and the number of each type of text is not less than 5 pieces; The special text includes sensitive information text, named entity text, special format text, specific language text, special character set encoding text, and special symbol text. Among them, the number of the sensitive information text and the named entity text is not less than 1000 pieces respectively, and the number of the special format text, the specific language text, the special character set encoding text, and the special symbol text is not less than 5 pieces respectively; The abnormal text includes garbled text and unsupported language text, and the number of each type of text is not less than 5 pieces; Input the test data in the test data set into the system under test respectively through a programmable test tool to test the semantic understanding test item of the system under test; By using a test statistics tool, obtain the operation result of the system under test during the test of the semantic understanding test item, and determine the test result corresponding to the semantic understanding test item based on the obtained operation result and the test content of the semantic understanding test item.
2. The method according to claim 1, wherein The semantic understanding test item includes a function test item and / or a performance test item. The function test item is used to test whether the system under test has the function corresponding to the function test item, and specifically includes at least one of an intention understanding test, a named entity recognition test, a sensitive information discrimination test, a semantic rejection recognition test, an information retrieval test, a text similarity calculation test, a text modification test, a semantic correction test, a natural language generation test, a logical reasoning test, a dialogue guidance test, and a context-related multi-round conversation test. The performance test item is used to test the performance of the system under test when executing the function corresponding to the function test item, and specifically includes at least one of a semantic understanding effect test, a semantic understanding efficiency test, and a system stability test.
3. The method according to claim 2, wherein The test content of the intention understanding test includes checking whether the system under test provides a function to understand the intention of the speaker. The function to understand the intention of the speaker includes at least one of the functions of fuzzy recognition of text, semantic extraction, semantic sorting, and intention classification. The test content of the named entity recognition test includes checking whether the system under test provides a function to find and accurately label named entities in the text. The test content of the sensitive information discrimination test includes checking whether the system under test provides a function to distinguish sensitive content in the input text according to the context. The test content of the semantic rejection recognition test includes checking whether the system under test provides a function to distinguish and reject invalid text input content that cannot be processed or should not be processed. The test content of the information retrieval test includes checking whether the system under test provides an information retrieval function, and the information retrieval includes at least personalized dictionary retrieval and / or third-party information source retrieval and / or custom knowledge base retrieval. The test content of the text similarity calculation test includes checking whether the system under test provides a function to calculate the degree of semantic information consistency between the input text data and the existing text. The test content of the text modification test includes checking whether the system under test provides a function to modify the previous sentence text in the dialogue. The test content of the semantic correction test includes checking whether the system under test provides a function to automatically correct the result of semantic understanding error. The test content of the natural language generation test includes checking whether the system under test provides a function to generate natural language text that conforms to the intention of the speaker and meets the requirements of voice interaction response according to the semantic understanding result. The test content of the logical reasoning test includes checking whether the system under test provides a function to perform logical calculation and derivation on the text content. The test content of the dialogue guidance test includes checking whether the system under test provides a function to dynamically generate guiding prompt words according to the intention of the speaker and the scene requirements, and guide the user to state their ultimate goal. The test content of the context-related multi-round conversation test includes checking whether the system under test provides a function for context-related multi-round conversation processing. Correspondingly, selecting test data corresponding to the semantic understanding test items from the pre-constructed test database includes: When the semantic understanding test item is any one of the intent understanding test, information retrieval test, text similarity calculation test, text modification test, semantic correction test, natural language generation test, logical reasoning test, and context-related multi-turn conversation test, select text data from the text data of the defined scenarios or services in the pre-constructed test database; When the semantic understanding test item is the named entity recognition test, select named entity text from the pre-constructed test database; When the semantic understanding test item is the sensitive information discrimination test, select sensitive information text from the pre-constructed test database; When the semantic understanding test item is the semantic rejection recognition test, select text data from the text data of the undefined scenarios or services in the pre-constructed test database; When the semantic understanding test item is the dialogue guidance test, select text data from the text data of the defined scenarios or services, and / or the text data of the undefined scenarios or services in the pre-constructed test database.
4. The method according to claim 2, wherein The test content of the semantic understanding effect test includes detecting the semantic understanding effect of the system under test when executing functions, specifically including detecting at least one of the precision rate, recall rate, rejection recognition rate, accuracy rate, F1 value, average reciprocal rank, and normalized discounted cumulative gain of the system under test when executing functions; The calculation formula of the rejection recognition rate is as follows: Among them, SR represents the semantic rejection rate, and N SR represents the number of times the actual response to the invalid text is correct; N R represents the total number of times the invalid text is input; The calculation formula of the F1 value is as follows: Among them, F1 represents the semantic understanding F1 value, P SS represents the semantic understanding precision rate, R SS represents the semantic understanding recall rate; The calculation formula of the average reciprocal rank is as follows: Among them, MRR represents the mean reciprocal rank, Q represents the total number of information retrievals, i represents the i-th information retrieval, and rank i represents the ranking position where the correct result appears in the i-th information retrieval; The calculation formula of the normalized discounted cumulative gain is as follows: NDCG = DCG / IDCG Among them, DCG represents the cumulative gain of losses, that is, the ranking correlation score of the results given by the system under test. K represents the number of information retrieval results, j represents the j-th retrieval result, and rel j represents the relevance score of the j-th retrieval result. IDCG represents the cumulative gain of losses of the ideal results, that is, the ranking correlation score of the ideal results. |REL K | represents that the number of information retrieval results is sorted from largest to smallest according to the relevance score. NDCG represents the normalized cumulative gain of losses; When the function executed by the system under test is the intent understanding function, detecting the semantic understanding effect of the system under test when executing the function includes at least detecting the precision rate and recall rate when the system under test performs semantic extraction on the test data; When the function executed by the system under test is the named entity recognition function, detecting the semantic understanding effect of the system under test when executing the function includes at least detecting the precision rate and recall rate when the system under test performs named entity recognition on the test data; When the function executed by the system under test is the sensitive information discrimination function, detecting the semantic understanding effect of the system under test when executing the function includes at least detecting the precision rate and recall rate when the system under test distinguishes sensitive content in the input text; When the function executed by the system under test is the semantic rejection recognition function, detecting the semantic understanding effect of the system under test when executing the function includes detecting the rejection recognition rate when the system under test distinguishes and rejects invalid text input content that cannot be processed or should not be processed; When the function executed by the system under test is the information retrieval function, detecting the semantic understanding effect of the system under test when executing the function includes at least detecting the normalized discounted cumulative gain when the system under test retrieves information matching the test data; When the function executed by the system under test is the text modification function, detecting the semantic understanding effect of the system under test when executing the function includes at least detecting the recall rate when the system under test modifies the previous sentence text in the dialogue; When the function executed by the system under test is a semantic correction function, detect the semantic understanding effect of the system under test when executing the function, including at least detecting the precision rate and recall rate when the system under test automatically corrects the results of semantic understanding errors; When the function executed by the system under test is a logical reasoning function, detect the semantic understanding effect of the system under test when executing the function, including at least detecting the precision rate and recall rate when the system under test performs logical calculations and derivations on the text content; When the function executed by the system under test is a multi-turn conversation function related to context, detect the semantic understanding effect of the system under test when executing the function, including at least detecting the accuracy rate when the system under test provides multi-turn session processing related to context.
5. The method according to claim 2, wherein The test content of the semantic understanding efficiency test includes detecting the semantic understanding efficiency of the system under test when executing the function, specifically including at least detecting at least one of the average semantic understanding response time, semantic understanding response time distribution, and semantic understanding throughput rate of the system under test when executing the function; Among them, the average semantic understanding response time refers to the ratio of the semantic understanding response time of all test data in the test data set to the total number of all test data, and the semantic understanding response time refers to the time when the system under test gives the semantic understanding result of a piece of text after inputting a piece of text; The calculation formula of the average semantic understanding response time is as follows: Among them, T avu represents the average response time of semantic understanding, W represents the test data set, and T i represents the semantic understanding response time corresponding to test sample i, and N represents the total number of input test data; The speech understanding response time distribution refers to the distribution and proportion of the semantic understanding response time of all test data in the test data set in each response time interval; The semantic understanding throughput rate refers to the text size of the semantic understanding achieved by the system under test within the unit response time; The calculation formula of the semantic understanding throughput rate is as follows: Among them, TP represents the semantic understanding throughput rate, W represents the test data set, S i represents the size of the text corresponding to the test sample i on the test data set, T i represents the semantic understanding response time corresponding to the test sample i.
6. The method according to claim 2, wherein The test content of the system stability test includes detecting whether the system operation situation and resource usage situation of the system under test are stable when executing the function under the conditions of the set software and hardware configuration and system concurrent number; the resource usage situation at least includes the usage rates of the system physical memory, virtual memory, CPU, GPU, handles, and network resources, and this resource usage situation is monitored by a resource monitoring tool; Conduct a system stability test on the system under test, including: Through a programmable test tool, continuously input test data to the system under test in a loop for 7 days in the online scenario and 3 days in the offline scenario without interruption, so that the system under test executes the function under the conditions of the set software and hardware configuration and system concurrent number, and real-time monitor the operation situation and resource usage situation of the system under test through a test statistics tool and a resource monitoring tool.
7. The method according to claim 2, wherein The method further includes: Determine the target speaker intention according to the function requirements of the system under test and the semantic understanding test requirements; Generate a subjective experience test data set according to the function requirements of the system under test, the semantic understanding test requirements, and the target speaker intention. The tester conducts a conversation with the system under test based on the subjective experience test dataset, and records the average number of dialogue turns for the system under test to understand the intention of the target speaker specified by the tester, the task completion rate in a single conversation, and the tester's satisfaction with the entire conversation process. Among them, the task completion rate in a single conversation refers to the success rate of the system under test in understanding the intention of at least one target speaker specified by the tester during a conversation with the tester. Based on the average number of dialogue turns for the system under test to understand the intention of the target speaker specified by the tester, the task completion rate in a single conversation, and the tester's satisfaction with the entire conversation process, determine the subjective experience test result of the system under test. Generate the semantic understanding function test result of the system under test according to the test results corresponding to each semantic understanding test item and the subjective experience test result.
8. An intelligent voice interaction test device, characterized in that, Including: A project selection unit for determining semantic understanding test items according to the functional requirements of the system under test and the semantic understanding test requirements. A test dataset generation unit for selecting test data corresponding to the semantic understanding test items from a pre-constructed test database and generating a test dataset matching the semantic understanding test items using the selected test data. Among them, the test database includes text data of defined scenarios or services, as well as text data of undefined scenarios or services. The text data of defined scenarios or services includes general text data of defined scenarios or services and common text data of defined scenarios or services. Among the general text data of defined scenarios or services, there are no less than 200 pieces of text data corresponding to each service. Among the common text data of defined scenarios or services, there are at least 3 pieces of real text data corresponding to each service. The text data of undefined scenarios or services includes general text data of undefined scenarios or services in the same field, common text data of undefined scenarios or services in the same field, chat text data, and text data that is meaningless or illogical. Among the general text data of undefined scenarios or services in the same field and the common text data of undefined scenarios or services in the same field, there are at least 3 pieces of real text data corresponding to each service. The number of chat text data is no less than 1000, and the average number of characters in each chat text is no less than 5 characters. The number of meaningless or illogical text data is no less than 100, and each piece is no less than 5 characters. The general text data of defined scenarios or services, the common text data of defined scenarios or services, the general text data of undefined scenarios or services in the same field, and the common text data of undefined scenarios or services in the same field are all composed of common text, special text, and abnormal text obtained by the system under test through speech recognition. Moreover, the length distribution of all text data belonging to the same type conforms to the set quantity distribution requirements. The common texts include single-word or word-group texts, short-text texts, single-sentence texts, dialogue texts, paragraph texts, and article texts with intention expressions, and the number of each type of text is not less than 5; The special texts include sensitive information texts, named entity texts, special format texts, specific language texts, special character set encoding texts, and special symbol texts. Among them, the number of sensitive information texts and named entity texts is not less than 1000 respectively, and the number of special format texts, specific language texts, special character set encoding texts, and special symbol texts is not less than 5 respectively; The abnormal texts include garbled texts and unsupported language texts, and the number of each type of text is not less than 5; A test processing unit, configured to input the test data in the test dataset into the system under test respectively through a programmable test tool, so as to perform tests on the semantic understanding test items of the system under test; A test statistics unit, configured to obtain the operation result of the system under test when performing tests on the semantic understanding test items through a test statistics tool, and determine the test result corresponding to the semantic understanding test item based on the obtained operation result and the test content of the semantic understanding test item.
9. An intelligent voice interaction test device, characterized in that, Comprising: A memory and a processor; The memory is connected to the processor and is used for storing computer programs; The processor is configured to implement the intelligent voice interaction test method according to any one of claims 1 to 7 by running the program in the memory.
10. A storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by the processor, the intelligent voice interaction test method according to any one of claims 1 to 7 is implemented.