Testing Method, Electronic Device and Storage Medium of Speech Synthesis System
By automatically obtaining test templates and test plans for speech synthesis systems, and generating target test cases, it solves the problem of impact on the test progress caused by artificially formulating test plans and solutions in the existing technology, and achieves more efficient test progress and efficiency.
Patent Information
- Application Number
- CN202210906435.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-07-29
AI Technical Summary
In the prior art, testing plans and testing plans need to be formulated manually, resulting in the increase in testing functions of the software products to be tested, and the testing progress is affected.
A test method for speech synthesis system is proposed. By obtaining system attribute data from the application side, filtering preliminary test templates and test scheme data, and generating target test cases, we realize automated testing of speech synthesis system.
Automatically obtaining test templates and test plans is realized, reducing human intervention and improving testing efficiency and progress.
Smart Images

Figure CN115237785B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a test method, an electronic device, and a storage medium for a speech synthesis system. Background Art
[0002] Software testing refers to the process of verifying and validating software products (including phased products).
[0003] In the related art, it is necessary to manually formulate a test plan and a test scheme for the software product to be tested. When the test functions of the software product to be tested increase, if the above method is used to test the software product to be tested, the test progress of the software product to be tested will be affected. Summary of the Invention
[0004] The main purpose of the embodiments of the present disclosure is to propose a test method, an electronic device, and a storage medium for a speech synthesis system, which can accelerate the test progress of the speech synthesis system and improve the test efficiency of the speech synthesis system.
[0005] To achieve the above object, a first aspect of the embodiments of the present disclosure proposes a test method for a speech synthesis system, including:
[0006] Obtaining system attribute data of the speech synthesis system to be tested from an application end; wherein, the speech synthesis system runs on the application end, and the system attribute data includes source information of the application end and version type information of the speech synthesis system;
[0007] Screening out a preliminary test template from a preset first test database according to the source information;
[0008] Performing a filling operation on the preliminary test template according to the version type information to obtain target test data;
[0009] Screening out test scheme data from a preset second test database according to the target test data;
[0010] Obtaining an initial test case according to the test scheme data;
[0011] Screening the initial test case according to a preset screening condition to obtain a target test case;
[0012] Testing the speech synthesis system according to the target test case.
[0013] In some embodiments, the test scheme data includes a test type, and the test type includes functional testing;
[0014] The obtaining an initial test case according to the test scheme data includes:
[0015] If the test type is the functional test, obtain the test scenario information of the speech synthesis system;
[0016] Filter out the initial test cases from a preset test case database according to the test scenario information; or, input the test scenario information into a preset text generation model for text generation to obtain the initial test cases.
[0017] In some embodiments, the filtering condition includes a matching pass result, and the test plan data includes test requirement information;
[0018] The filtering of the initial test cases according to preset filtering conditions to obtain target test cases includes:
[0019] Obtain the test title information of the initial test cases;
[0020] Obtain the test requirement information according to the test plan data;
[0021] Compare the test requirement information with the test title information;
[0022] If the test requirement information matches the test title information, obtain the matching pass result;
[0023] Use the initial test cases as the target test cases according to the matching pass result.
[0024] In some embodiments, the filtering condition includes a comparison pass result;
[0025] The filtering of the initial test cases according to preset filtering conditions to obtain target test cases includes:
[0026] Obtain a first test audio generated by the speech synthesis system according to the initial test cases;
[0027] Input the first test audio into a preset target speech recognition model for recognition to obtain a test text;
[0028] Compare the test text with the initial test cases;
[0029] If the test text is consistent with the initial test cases, obtain the comparison pass result;
[0030] Use the initial test cases as the target test cases according to the comparison pass result.
[0031] In some embodiments, the system attribute data further includes acoustic training parameters;
[0032] Before screening the initial test cases according to the preset screening conditions to obtain target test cases, the test method further includes constructing the target speech recognition model, specifically including:
[0033] Obtain language knowledge according to a preset language database; wherein, the language knowledge includes acoustic knowledge, phonetic knowledge, and language prior knowledge;
[0034] Input the acoustic knowledge and the phonetic knowledge into the original acoustic model for training to obtain a preliminary acoustic model;
[0035] Input the language prior knowledge into the original language model for training to obtain a preliminary language model;
[0036] Construct a preliminary speech recognition model according to the preliminary acoustic model and the preliminary language model;
[0037] Obtain speech information, and obtain speech parameters according to the speech information; wherein, the speech parameters match the acoustic training parameters;
[0038] Match the speech parameters with the preliminary speech recognition model to obtain the target speech recognition model.
[0039] In some embodiments, the test method further includes:
[0040] If the version type information indicates an iterative type, obtain the historical test cases of the speech synthesis system;
[0041] Obtain actual test cases according to the historical test cases and the target test cases;
[0042] Test the speech synthesis system according to the actual test cases.
[0043] In some embodiments, the test method further includes:
[0044] Obtain a second test audio generated by the speech synthesis system according to the target test case;
[0045] Analyze the audio content of the second test audio to obtain an analysis result; wherein, the analysis result includes an error result indicating a functional defect of the speech synthesis system;
[0046] Generate a prompt message for prompting function improvement according to the error result.
[0047] In some embodiments, the preliminary test template includes fillable fields;
[0048] The filling operation of the preliminary test template according to the version type information to obtain target test data includes:
[0049] If the version type information indicates an iterative type, obtain the historical test data of the speech synthesis system;
[0050] Obtain the field attribute of the filling field; wherein, the field attribute includes a reusable type;
[0051] Use the filling field with the field attribute of the reusable type as the field to be processed;
[0052] Perform a filling operation on the field to be processed according to the historical test data to obtain the target test data.
[0053] To achieve the above object, a second aspect of the embodiments of the present disclosure provides an electronic device, including:
[0054] At least one memory;
[0055] At least one processor;
[0056] At least one computer program;
[0057] The computer program is stored in the memory, and the processor executes the at least one computer program to implement:
[0058] The test method of the speech synthesis system as described in any item of the first aspect.
[0059] To achieve the above object, a third aspect of the embodiments of the present disclosure provides a computer-readable storage medium, and the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute:
[0060] The test method of the speech synthesis system as described in any item of the first aspect.
[0061] The test method of the speech synthesis system provided by the embodiments of the present application, according to the preset first test database and second test database, enables the preliminary test module and test scheme data to be screened according to the system attribute data of the speech synthesis system to be tested, so that the target test cases for testing can be obtained according to the test scheme data and screening conditions. It can be seen that the test method of the speech synthesis system provided by the embodiments of the present application can automatically obtain the preliminary test template (i.e., the test plan) and test scheme data (i.e., the test scheme), and on this basis, realize the automatic test of the speech synthesis system, avoiding the method of manually formulating the test plan and test scheme in the related art. Therefore, the test method of the speech synthesis system provided by the embodiments of the present application can speed up the test progress and improve the test efficiency. Description of the Drawings
[0062] Figure 1It is a schematic flowchart of a test method for a voice synthesis system according to an embodiment of the present application;
[0063] Figure 2 It is another schematic flowchart of a test method for a voice synthesis system according to an embodiment of the present application;
[0064] Figure 3 It is another schematic flowchart of a test method for a voice synthesis system according to an embodiment of the present application;
[0065] Figure 4 It is another schematic flowchart of a test method for a voice synthesis system according to an embodiment of the present application;
[0066] Figure 5 It is another schematic flowchart of a test method for a voice synthesis system according to an embodiment of the present application;
[0067] Figure 6 It is another schematic flowchart of a test method for a voice synthesis system according to an embodiment of the present application;
[0068] Figure 7 It is another schematic flowchart of a test method for a voice synthesis system according to an embodiment of the present application;
[0069] Figure 8 It is another schematic flowchart of a test method for a voice synthesis system according to an embodiment of the present application;
[0070] Figure 9 It is a block diagram of a module of a test device for a voice synthesis system according to an embodiment of the present application;
[0071] Figure 10 It is a schematic diagram of a hardware structure of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0072] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0073] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.
[0074] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.
[0075] First, some nouns involved in this application are parsed as follows:
[0076] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also refers to the theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0077] Natural Language Processing (NLP): NLP uses a computer to process, understand, and apply human languages (such as Chinese, English, etc.). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, and is often referred to as computational linguistics. Natural language processing includes syntactic analysis, semantic analysis, discourse understanding, etc. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information retrieval, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and linguistic research related to language computing.
[0078] Text to Speech (TTS): also known as text-to-speech conversion, is a technology that converts text information generated by the computer itself or input externally into understandable and fluent speech output. In speech synthesis technology, it is mainly divided into a language analysis part and an acoustic system part, also known as the front-end part and the back-end part. Among them, the language analysis part mainly analyzes the input text information to generate the corresponding linguistic specification; the acoustic system part mainly generates the corresponding audio according to the linguistic specification provided by the language analysis part, thereby realizing the function of voice. Specifically, the language analysis part includes text input, text structure and language judgment, text standardization, text conversion to phonemes, sentence reading and rhythm prediction, etc. Among them, the text structure and language judgment is to judge the language of the input text to be synthesized, for example, to judge whether the text to be synthesized is Chinese, English, Japanese, etc., and then according to the grammatical rules of the corresponding language, the entire text in the text to be synthesized is divided into single sentences, and the segmented sentences are transmitted to the subsequent processing module. Text standardization is to standardize all the contents in the text to be synthesized. For example, when there are Arabic numerals or letters in the text to be synthesized, it is necessary to convert the Arabic numerals or letters into text according to the set rules to facilitate the subsequent text phonetic notation. Text phonetic elements are used to determine the pronunciation of the current synthesized text. For example, in Chinese speech synthesis, the text is mainly annotated with pinyin, so the text phonetic elements need to convert the text into the corresponding pinyin. When some characters are polyphonic, it is also necessary to determine which tone the text is specifically pronounced through word segmentation, part of speech and syntactic analysis, etc., and determine the tone of the text. Sentence and rhythm prediction is to predict the rhythm of the text, that is, to determine where the synthesized speech needs to pause, how long the pause is, which words or phrases need to be stressed, which words need to be read lightly, etc., so that the sound of the synthesized speech is high and low, and the ups and downs, so as to achieve a more realistic imitation of the human voice. There are three main technical implementation methods for the acoustic system, namely waveform splicing speech synthesis technology, parameter speech synthesis technology, and end-to-end speech synthesis technology. Among them, waveform splicing speech synthesis technology splices syllables in an existing library to achieve speech synthesis. Parametric speech synthesis technology mainly uses data methods to model the spectral characteristic parameters of existing recordings, build a mapping relationship between text sequences and speech features, and generate a parametric synthesizer. End-to-end speech synthesis technology uses a neural network learning method to output synthesized audio for directly input text or phonetic characters based on the intermediate "black box part".
[0079] Automatic Speech Recognition (ASR): It is used to convert the lexical content in human speech into computer-readable input, such as keystrokes, binary codes, or character sequences. An automatic speech recognition system includes four parts: signal processing and feature extraction, acoustic model, language model, and decoding search. Among them, signal processing and feature extraction take an audio signal as input, enhance the speech by eliminating noise and channel distortion, transform the signal from the time domain to the frequency domain, and extract appropriate and representative feature vectors for the acoustic model. The acoustic model integrates acoustic and phonetic knowledge, takes the features generated by the feature extraction part as input, and generates an acoustic model score for the variable-length feature sequence. The language model is used to learn the mutual relationship between words through training corpora to estimate the likelihood of a hypothesized word sequence, which is also called the language model score. Therefore, if one understands the prior knowledge in the corresponding field or the prior knowledge related to the task, it will be possible to effectively improve the language model score.
[0080] Prior knowledge: It is knowledge prior to experience, which can be used to adjust and optimize the parameter range of a model or to constrain the model to improve the accuracy of the model output data. For example, when using an image recognition model to recognize license plates, the prior knowledge includes the aspect ratio of the license plate, the background color of the license plate, the character color, the character font, the number of character rows, the distribution characteristics of character intervals, and the character set to which each character belongs.
[0081] Knowledge mining: It refers to obtaining information such as entities, new entity links, and new association rules from data. The main techniques include entity linking and disambiguation, knowledge rule mining, knowledge graph representation learning, etc. Among them, entity linking and disambiguation are content mining of knowledge; knowledge rule mining is structure mining; representation learning is to map the knowledge graph to a vector space and then conduct mining.
[0082] Phonetics: It is used to study the pronunciation mechanism, speech characteristics, and variation rules in speech, etc. The research objects of phonetics include vowels, consonants, tones, stresses, rhythms, sound changes, prosodies, etc.
[0083] Linguistics: It is a discipline that takes human language as the research object. The research objects of linguistics include language structure, word formation, syntax, semantics, etc.
[0084] Software Testing: It is a process of auditing or comparing the actual output with the expected output, used to identify the correctness, integrity, security, and quality of the software to be tested, etc. Divided according to the development stage, software testing can be divided into unit testing, integration testing, and system testing. Among them, unit testing is also called module testing, and unit testing is aimed at the smallest unit of software design. The purpose of unit testing is to check whether each program unit can correctly implement the module functions, performance, interfaces, and design constraints, etc. required in the detailed design specifications, and to discover various errors that may exist within each module. Unit testing needs to design test cases starting from the internal structure of the program, and multiple modules can be independently unit-tested in parallel. Integration testing is also called assembly testing. Integration testing is usually carried out on the basis of unit testing. Integration testing is a process of orderly and incremental testing of all program modules. Integration testing is used to verify the interface relationships of program units or components, in order to gradually integrate them into program components or the entire system that meets the requirements of the general design. System testing is a process of checking whether the complete program system can be correctly configured and connected with the system (including hardware, peripherals, network, system software, support platforms, etc.) and finally meet the design requirements in the real system operating environment.
[0085] Test Case: It refers to the description of the test tasks for a specific software product, reflecting the test plan, methods, techniques, and strategies. A test case is a set of test inputs, execution conditions, and expected results compiled for a specific goal, used to verify whether the software product to be tested meets a specific software requirement. A test case mainly includes four contents: use case title, precondition, test steps, and expected result. Among them, the use case title mainly describes the test of a certain function; the precondition refers to the conditions that the use case title needs to meet; the test steps are used to describe the operation steps of the test case; the expected result refers to meeting the expected requirements.
[0086] Test Plan: It is a document data at the organizational management level, used to stipulate and restrict the test scope, organization, resources, principles, etc. of the entire software testing process, and to formulate the task allocation and time schedule for each stage of the entire testing process, and to put forward the evaluation of each task, risk analysis, and management requirements.
[0087] Test Scheme: It is a further refinement and clarification of the test plan, and it is a document data at the technical level. The test scheme is used to describe the characteristics of the software product to be tested, the test methods, the planning of the test environment, the design and selection of test tools, the design method of test cases, the design scheme of test code, etc.
[0088] Functional testing: Also known as behavioral testing, it tests the characteristics and operable behaviors of the software product to be tested according to the characteristics of the software product to be tested, operation descriptions, and user scenarios, in order to determine whether it meets the design requirements.
[0089] Ad-hoc testing: It is used to retest the important functions of the software product to be tested and conduct spot checks on the functions and performance of the software product to be tested. It is an effective way and process to ensure the integrity of test coverage.
[0090] Interface testing: It is a type of testing for testing the interfaces between system components. Interface testing is mainly used to detect the interaction points between external systems and the software product to be tested, as well as between the subsystems within the software product to be tested. The focus of testing is to check data exchange, transmission, control management, and the mutual logical dependencies between systems, etc.
[0091] Performance testing: It refers to the process of using automated testing tools or code means to simulate normal and peak load access to the software product to be tested, in order to observe whether the performance indicators of the software product to be tested are qualified. Among them, performance indicators include response time, throughput, resource utilization rate, error rate, etc.
[0092] Security testing: It is the process of verifying the security services of the software product to be tested and identifying potential security defects, including user management and access control testing, communication and data encryption testing, data backup and recovery testing, etc.
[0093] Regression testing: It refers to a testing method of retesting the software product to be tested by modifying the old version to obtain the new version, in order to confirm that the modification does not introduce new errors or cause errors in other unmodified functions.
[0094] In the related art, it is necessary to manually formulate a test plan and test scheme for the software product to be tested. When the test functions of the software product to be tested increase, if the above method is used to test the software product to be tested, it will affect the test progress of the software product to be tested.
[0095] Based on this, the embodiments of the present application provide a testing method, an electronic device, and a storage medium for a speech synthesis system, which can perform automated testing on the speech synthesis system, thereby accelerating the test progress to a certain extent and improving the test efficiency of the speech synthesis system.
[0096] The embodiments of the present application provide a testing method, an electronic device, and a storage medium for a speech synthesis system, which are specifically described through the following embodiments. First, the testing method for the speech synthesis system in the embodiments of the present application is described.
[0097] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0098] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, mechatronics, etc. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0099] The test method for the speech synthesis system provided by the embodiments of the present application relates to the field of artificial intelligence technology, and particularly relates to the field of software testing technology. The test method for the speech synthesis system provided by the embodiments of the present application can be applied to a terminal, can also be applied to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, or a smart watch, etc.; the server can be an independent server, or can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms, etc.; the software can be an application implementing the test method for the speech synthesis system, etc., but is not limited to the above forms.
[0100] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0101] In a first aspect, with reference to Figure 1, embodiments of the present application provide a method for testing a speech synthesis system, and the testing method includes but is not limited to steps S110 to S170.
[0102] S110. Obtain system attribute data of the speech synthesis system to be tested from the application side; wherein, the speech synthesis system runs on the application side, and the system attribute data includes source information of the application side and version type information of the speech synthesis system.
[0103] It can be understood that the speech synthesis system to be tested is the software product to be tested in the embodiments of the present application, and it can be loaded on application sides such as the APP side of a terminal device or the WEB side of a web page. Obtain the system attribute data of the speech synthesis system to determine the basic attribute information of the speech synthesis system. Among them, the system attribute data includes source information for characterizing the application side loaded by the speech synthesis system and version information for characterizing the speech synthesis system. Specifically, when the version information is expressed as version 1.0, it indicates that the speech synthesis system is developed for the first time; when the version type information is expressed as version 2.0, it indicates that the speech synthesis system is iteratively developed. It can be understood that version 1.0 and version 2.0 are only exemplary, that is, according to actual needs, an iterative development version of 1.1 can also be set, and the embodiments of the present application do not make specific limitations on this.
[0104] It can be understood that the terminal device can be a mobile terminal device or a non-mobile terminal device. Among them, the mobile terminal device can be a mobile phone, a tablet computer, a notebook computer, a handheld computer, a vehicle-mounted terminal device, a wearable device, a super mobile personal computer, a netbook, a personal digital assistant, a CPE, a UFI (wireless hotspot device), etc. The non-mobile terminal device can be a personal computer, a television, a teller machine or a self-service machine, etc. The embodiments of the present application do not make specific limitations on this.
[0105] S120. Screen out a preliminary test template from a preset first test database according to the source information.
[0106] It can be understood that when the software product to be tested is loaded on different application sides, the focus of software testing will be different. For example, when the application side is the APP side of a terminal device, the test focuses include terminal device model matching test, system compatibility test, random test, performance test, etc.; when the application side is the WEB side of a web page, the test focuses include browser compatibility test, website usability test, concurrent performance test, etc.
[0107] Specifically, according to historical software test data, a first test database is pre-constructed. The first test database includes multiple test templates and the source information corresponding to each test template. For example, when the source information includes the APP side of the terminal device and the WEB side of the web page, the first test database includes test templates corresponding to the APP side of the terminal device and test templates corresponding to the WEB side of the web page. Obtain the test template corresponding to the application side loaded by the speech synthesis system from the first test database according to the obtained source information, and use this test template as the preliminary test template. It can be understood that the test template is used to represent the test plan in software testing.
[0108] S130. Perform a filling operation on the preliminary test template according to the version type information to obtain target test data;
[0109] It can be understood that the preliminary test template obtained according to step S120 is only a general template corresponding to the source information. Therefore, in order to match the current speech synthesis system to be tested, it is also necessary to fill the content of the preliminary test template according to the specific information of the speech synthesis system.
[0110] Specifically, determine whether the speech synthesis system to be tested is for first development or iterative development according to the version type information. When it is determined that the speech synthesis system is for first development, filling information can be obtained by means such as user input and automatic recognition of development project files; when it is determined that the speech synthesis system is for iterative development, filling information can be obtained by means such as user input, automatic recognition of development project files, and information of the previous iteration version. Perform a filling operation on the preliminary test template according to the filling information to obtain target test data.
[0111] S140. Screen out test plan data from a preset second test database according to the target test data;
[0112] It can be understood that the test plan data is used to characterize the test plan in software testing. The test plan includes the test types of this test, and the test types include functional testing, random testing, interface testing, performance testing, security testing, regression testing, etc. From the above description, it can be seen that the test plan is the technical-level document data corresponding to the test plan. Therefore, according to the target test data, the test plan data corresponding to the target test data can be screened out from the pre-constructed second test database. For example, if the target test data indicates that the voice synthesis system to be tested only involves the expansion or contraction of the cluster, only performance testing of the voice synthesis system needs to be performed at this time to complete the corresponding requirement verification. Therefore, according to the target test data, a test plan (i.e., test plan data) including performance testing is screened out from the second test database. Or, when the target test data indicates that the voice synthesis system to be tested is for iterative development, according to the target test data, a test plan data including security testing, random testing, and regression testing is screened out from the second test database.
[0113] S150. Obtain initial test cases according to the test plan data;
[0114] It can be understood that the test plan is used to describe the design method of test cases. Therefore, according to the test plan data, initial test cases matching the voice synthesis system to be tested can be obtained. For example, when the screened test plan data only includes functional testing, only the initial test cases for performing functional testing are obtained or constructed to speed up the testing progress of the voice synthesis system.
[0115] S160. Screen the initial test cases according to the preset screening conditions to obtain target test cases;
[0116] It can be understood that according to step S150, multiple initial test cases will be obtained or constructed. In order to ensure that the test cases loaded on the voice synthesis system to be tested meet the test requirements and match the functions of the voice synthesis system, etc., the multiple initial test cases also need to be screened. Specifically, according to the preset screening conditions such as the test requirements and functions of the voice synthesis system to be tested, the initial test cases that meet the screening conditions are used as target test cases.
[0117] S170. Test the voice synthesis system according to the target test cases.
[0118] It can be understood that the obtained target test cases are used as the input data set of the speech synthesis system to implement the software test of the speech synthesis system. Specifically, taking the function test as an example, assume that the function of the speech synthesis system to be tested is to convert Chinese into Chinese. At this time, the target test cases are in the text format including Chinese content. The target test cases are input into the application end loaded with the speech synthesis system, and the output data of the application end is compared with the target test cases, so as to implement the function test of the speech synthesis system.
[0119] It can be understood that the test method of the speech synthesis system provided by the embodiments of the present application can be any one of unit tests, integration tests, and system tests of the speech synthesis system, and the embodiments of the present application do not make specific limitations on this. Secondly, during the process of testing the speech synthesis system by the test method of the speech synthesis system provided by the embodiments of the present application, test report data can also be automatically synthesized according to the generated target test data, test plan data, target test cases, and test results when testing according to the target test cases. For example, a test report module is preset, and the test report template is filled according to the target test data, test plan data, target test cases, test results, etc. to obtain the test report data.
[0120] The test method of the speech synthesis system provided by the embodiments of the present application enables the preliminary test module and test plan data to be screened according to the system attribute data of the speech synthesis system to be tested based on the preset first test database and second test database, so as to obtain the target test cases for testing according to the test plan data and screening conditions. It can be seen that the test method of the speech synthesis system provided by the embodiments of the present application can realize the automatic acquisition of the preliminary test template (i.e., the test plan) and test plan data (i.e., the test plan), and on this basis, realize the automatic test of the speech synthesis system, avoiding the method of manually formulating the test plan and test plan in the related technology. Therefore, the test method of the speech synthesis system provided by the embodiments of the present application can speed up the test progress and improve the test efficiency.
[0121] Refer to Figure 2 , in some embodiments, the test plan data includes a test type, and the test type includes a function test. Step S150 includes but is not limited to sub-steps S210 to S220.
[0122] S210. If the test type is a function test, obtain the test scenario information of the speech synthesis system;
[0123] It can be understood that the test scenario data is used to characterize the test scenario in software testing. The test scenario includes the test type of this test, and the test type includes functional testing, random testing, interface testing, performance testing, security testing, regression testing, etc. When it is determined according to the test scenario data that the test type of the speech synthesis system to be tested is functional testing, the test scenario information of the speech synthesis system is obtained by means of user input, automatic recognition of development project files, etc. When the speech synthesis system to be tested is applied to a specific scenario, the test scenario information corresponds to the specific scenario. For example, when the speech synthesis system to be tested is applied to the insurance scenario, the test scenario information includes keyword information related to insurance such as insurance recommendation, insurance introduction, and purchasing insurance. When the speech synthesis system to be tested is applied to a non-specific scenario, the test scenario information includes keyword information of multiple different scenarios.
[0124] It can be understood that the test scenario information can be in the format of an excel file, that is, multiple keyword information is written into the excel file, and the multiple keyword information is set in different rows in the excel file for subsequent calling of the keyword information.
[0125] S220. Screen out the initial test cases from the preset test case database according to the test scenario information; or, input the test scenario information into the preset text generation model for text generation to obtain the initial test cases.
[0126] It can be understood that the embodiments of the present application provide two methods for obtaining the initial test cases according to the test scenario information. Among them, the first one is to pre-construct a test case database, which includes keyword information corresponding to different scenarios and the text corresponding to each keyword information. For example, for the keyword information of insurance recommendation, the corresponding text is "Hello, which aspect of insurance do you want to buy? Do you have a need for accidental value preservation or disease protection?" Screen out multiple texts from the test case database according to the keyword information in the excel file, and use the multiple texts as the initial test cases. It can be understood that one keyword information can correspond to one text or multiple texts, and the embodiments of the present application do not make specific limitations on this.
[0127] The second one is to pre-construct a text generation model that can generate random text according to keyword information, use the keyword information in the excel file as the input data of the text generation model, and use the output data of the text generation model as the initial test cases. It can be understood that the text generation model can be a GAN network model or other network models, and the embodiments of the present application do not make specific limitations on this.
[0128] Refer to Figure 3, in some embodiments, the screening condition includes a matching pass result, and the test scenario data includes test requirement information. Step S160 includes, but is not limited to, sub-steps S310 to S350.
[0129] S310. Obtain the test title information of the initial test case;
[0130] It can be understood that the initial test case includes a use case title (i.e., test title information), and the use case title is used to describe the function to be tested. Specifically, when obtaining the initial test case according to the first method above, the test case database includes the text corresponding to each keyword information and the test title information corresponding to the text. For example, it includes test title information such as correct text synthesis of audio, synthesis of specified audio by adding spaces to the text, etc. Among them, the text corresponding to correct text synthesis of audio is used to test the text synthesis function of the speech synthesis system; the text corresponding to synthesis of specified audio by adding spaces to the text includes special characters such as spaces, and this text is used to test the function of the speech synthesis system to recognize special characters. It can be understood that one test title information can correspond to multiple texts. For example, multiple texts can be set to test the special character recognition function of the speech synthesis system.
[0131] When obtaining the initial test case according to the second method above, technologies such as OCR (Optical Character Recognition) can be used to determine whether the initial test case includes special characters, etc., and then determine the test title information corresponding to the initial test case.
[0132] S320. Obtain the test requirement information according to the test scenario data;
[0133] It can be understood that the test scenario data includes the test requirement information of the speech synthesis system to be tested, and the test requirement information is used to describe the function points to be tested of the speech synthesis system. For example, the test requirement information includes correct synthesis of audio, recognition of special characters, etc.
[0134] S330. Compare the test requirement information with the test title information;
[0135] It can be understood that the test requirement information is traversed and compared with multiple test title information to determine whether there is test title information corresponding to the test requirement information, and then determine whether all the multiple initial test cases cover the function points to be tested.
[0136] S340. If the test requirement information matches the test title information, obtain a matching pass result;
[0137] It can be understood that the test requirement information is respectively matched and compared with multiple test title information, and matching results are obtained. Among them, the matching results include passing match results and failing match results. It can be understood that one test requirement information can be matched with multiple test title information. For example, when the test requirement information indicates identifying special characters, it can be matched with the test title information of adding spaces to the text to synthesize a specified audio, and the test title information of adding a hash character (#) to the text to synthesize a specified audio.
[0138] S350. Use the initial test case as the target test case according to the passing match result.
[0139] It can be understood that the initial test case with a passing match result is used as the target test case, and the speech synthesis system to be tested is software-tested according to the target test case.
[0140] It can be understood that in order to ensure that all test requirement information is covered, when none of the multiple test title information matches the test requirement information, new test cases are obtained according to the test requirement information, and the new test cases are used as the target test cases to ensure the comprehensiveness of the target test cases.
[0141] Refer to Figure 4 , in some embodiments, the screening condition includes the passing comparison result. Step S160 includes but is not limited to sub-steps S410 to S450.
[0142] S410. Obtain the first test audio generated by the speech synthesis system according to the initial test case;
[0143] S420. Input the first test audio into a preset target speech recognition model for recognition to obtain the test text;
[0144] S430. Compare the test text with the initial test case;
[0145] S440. If the test text is consistent with the initial test case, obtain the passing comparison result;
[0146] S450. Use the initial test case as the target test case according to the passing comparison result.
[0147] It can be understood that in some embodiments, it is also necessary to judge the correctness of the initial test case, that is, use the initial test case as the input data of the speech synthesis system to be tested, and determine whether the initial test case can be correctly recognized and processed by the speech synthesis system according to the output data and standard data of the speech synthesis system.
[0148] Specifically, in step S410 of some embodiments, the initial test case is used as the input data of the speech synthesis system to be tested, and the first test audio generated by the speech synthesis system according to the initial test case is obtained.
[0149] In step S420 of some embodiments, in order to avoid affecting the test progress and test accuracy of the speech synthesis system when comparing the output data with the preset data manually, the test method of the speech synthesis system provided by the embodiments of the present application pre-constructs a target speech recognition model. The output data of the speech synthesis system to be tested is used as the input data of the target speech recognition model, and the output data of the target speech recognition model is used as the data to be compared (i.e., the test text).
[0150] In step S430 of some embodiments, the output data of the target speech recognition model (i.e., the test text) is compared with the standard data (i.e., the initial test case) to determine whether the initial test case can be correctly recognized and processed by the speech synthesis system.
[0151] In step S440 of some embodiments, according to the comparison process in step S430, corresponding comparison results will be generated. The comparison results include a comparison pass result and a comparison fail result. If the output data of the target speech recognition model is consistent with the standard data, a comparison pass result indicating that the corresponding initial test case can be recognized and processed by the speech synthesis system is generated, that is, the initial test case is correct. At this time, the initial test case is used as the target test case to perform real software testing on the speech synthesis system in subsequent operations. It can be understood that "consistent" means that the error between the output data of the target speech recognition model and the standard data is within a preset range, and the specific value of the preset range can be adaptively adjusted according to actual needs, and the embodiments of the present application do not make specific limitations on this.
[0152] Refer to Figure 5 , in some embodiments, the system attribute data further includes acoustic training parameters. Before step S420, the test method of the language synthesis system provided by the embodiments of the present application further includes: constructing a target speech recognition model, specifically including but not limited to steps S510 to S570.
[0153] S510. Obtain language knowledge according to a preset language database; wherein, the language knowledge includes acoustic knowledge, phonetic knowledge, and language prior knowledge;
[0154] It can be understood that voice information, language information, etc. of each scenario and application field are obtained in advance, and a language database is constructed based on the voice information and the language information. Signal processing operations and knowledge mining operations are performed on the information in the language database to obtain language knowledge such as acoustic knowledge, phonetic knowledge, and language prior knowledge. Among them, the voice information includes information related to phonetics, and the language information includes information related to linguistics. The acoustic knowledge includes knowledge such as pause distribution and language emotion, and the phonetic knowledge includes knowledge such as phonemes and syllables. The language prior knowledge includes dictionary knowledge, grammar knowledge, syntactic knowledge, etc. related to the field (or scenario).
[0155] S520. Input the acoustic knowledge and the phonetic knowledge into the original acoustic model for training to obtain a preliminary acoustic model;
[0156] It can be understood that multiple pieces of acoustic knowledge and multiple pieces of phonetic knowledge obtained according to the above steps are used as training parameters of the original acoustic model to train multiple preliminary acoustic models corresponding to different fields and different application scenarios; and / or, multiple preliminary acoustic models corresponding to the same scenario but using different training parameters are trained.
[0157] S530. Input the language prior knowledge into the original language model for training to obtain a preliminary language model;
[0158] It can be understood that multiple pieces of language prior knowledge obtained according to the above steps are used as training parameters of the original language model to train multiple preliminary language models corresponding to different fields and different application scenarios; and / or, multiple preliminary language models corresponding to the same scenario but using different training parameters are trained.
[0159] S540. Construct a preliminary speech recognition model based on the preliminary acoustic model and the preliminary language model;
[0160] It can be understood that multiple preliminary acoustic models are respectively matched with the preliminary language models applied to the relevant fields (or scenarios) to construct multiple preliminary speech recognition models.
[0161] S550. Obtain the voice information from the target object;
[0162] It can be understood that in actual deployment, the speech synthesis system to be tested will be applied to specific fields or specific scenarios. Therefore, the acoustic training parameters of the speech synthesis system are parameters related to the field (or scenario). For example, when the speech synthesis system is deployed on the shopping platform APP of a terminal device, the acoustic training parameters include speech characteristic parameters such as the linear prediction coefficients (LPC) and mel-frequency cepstral coefficients (MFCC) of the target object (i.e., the speaker, such as a customer service representative) of the shopping platform, so that the output data of the speech synthesis system matches the timbre, pitch, etc. of the target object of the shopping platform.
[0163] It can be understood that in order to ensure the accuracy of comparing the test text with the initial test case, the target speech recognition model should be a preliminary speech recognition model with high recognition ability for the voice of the target object. Therefore, obtain the speech information of the target object, for example: obtain the audio of the target object, and the audio is used to represent any content.
[0164] S560. Extract feature parameters from the speech information according to the acoustic training parameters to obtain speech parameters;
[0165] It can be understood that determine the parameter type of the acoustic training parameters, and extract feature parameters from the speech information according to the parameter type to obtain speech parameters of the same type as the parameter type. For example, when the acoustic training parameters include linear prediction coefficients (LPC), mel-frequency cepstral coefficients (MFCC), etc., linear prediction coefficients (LPC), mel-frequency cepstral coefficients (MFCC), etc. will be extracted from the speech information as speech parameters.
[0166] S570. Select the target speech recognition model from the preliminary speech recognition models according to the speech parameters.
[0167] It can be understood that the speech parameters obtained according to the above steps will be respectively matched and compared with the corresponding parameters of multiple preliminary speech recognition models to obtain the preliminary speech recognition model with the highest matching probability. The preliminary speech recognition model with the highest matching probability is used as the target speech recognition model.
[0168] Refer to Figure 6 , in some embodiments, the test method for the speech synthesis system provided by the embodiments of the present application further includes but is not limited to steps S610 to S630.
[0169] S610. If the version type information indicates an iterative type, obtain the historical test cases of the speech synthesis system;
[0170] It can be understood that when it is determined that the speech synthesis system is iteratively developed according to the system attribute data of the speech synthesis system, obtain the historical test cases corresponding to the historical development versions of the speech synthesis system.
[0171] S620. Obtain the actual test cases based on the historical test cases and the target test cases;
[0172] It can be understood that when the speech synthesis system is developed iteratively, it should be determined whether there is a phenomenon of repeated requirement use cases, that is, it should be determined whether the target test cases are repeated with the historical test cases. When it is determined that the target test cases are repeated with the historical test cases, regression testing should be performed on the speech synthesis system according to the target test cases, that is, the target test cases are used as the actual test cases to retest the speech synthesis system. When it is determined that the target test cases are not repeated with the historical test cases, it indicates that the target test cases correspond to the new requirements of the current development version of the speech synthesis system. Therefore, it is necessary to judge the correctness of the target test cases, use the target test cases that pass the correctness judgment as the actual test cases, and archive the target test cases to be used as the historical test cases of the speech synthesis system in the next iterative version. It can be understood that the method for judging the correctness of the target test cases is the same as the method for judging the correctness of the initial test cases described above, and the embodiments of the present application will not elaborate on this again.
[0173] S630. Test the speech synthesis system according to the actual test cases.
[0174] It can be understood that the actual test cases obtained according to step S620 are used as the input data of the speech synthesis system to implement the software test of the speech synthesis system.
[0175] Refer to Figure 7 , in some embodiments, the test method of the speech synthesis system provided by the embodiments of the present application further includes but is not limited to steps S710 to S730.
[0176] S710. Obtain the second test audio generated by the speech synthesis system according to the target test cases;
[0177] It can be understood that the target test cases are used as the input data of the speech synthesis system to be tested, so as to determine whether the target test cases are contradictory use cases according to the output data (i.e., the second test audio) of the speech synthesis system, that is, to determine whether the target test cases are contradictory to the existing requirement functions of the speech synthesis system. When the audio content of the second test audio corresponds to the target test cases, it indicates that the target test cases are not contradictory use cases; when the audio content of the second test audio is content such as warnings or error reports, it indicates that the target test cases are contradictory use cases.
[0178] For example, the existing required function of the speech synthesis system for the pound sign (#) is set to splicing. That is, when the text corresponding to the target test case is "Script A#Script B", if the second test audio output by the speech synthesis system is the spliced audio of Script A and Script B (i.e., outputting "Script A Script B"), it is determined that the target test case is not a contradictory use case. When the text corresponding to the target test case is "17#302" used to represent a house number, since it is contradictory to the current splicing function of the speech synthesis system, the speech synthesis system will output a second test audio such as "Splicing error" or "Splicing failed". At this time, it is determined that the target test case is a contradictory use case.
[0179] S720. Analyze the audio content of the second test audio to obtain an analysis result; wherein, the analysis result includes an error result indicating a functional defect in the speech synthesis system.
[0180] It can be understood that by means of speech recognition and the like, the audio content of the second test audio is analyzed to obtain an analysis result including a correct result and an error result. When the audio content of the second test audio corresponds to the target test case, a correct result is obtained; when the audio content of the second test audio is "Splicing error", "Splicing failed", etc., an error result is obtained.
[0181] S730. Generate a prompt message for prompting function improvement according to the error result.
[0182] It can be understood that when an error result is obtained, it indicates that the target test case currently input to the speech synthesis system is contradictory to the functional requirements of the speech synthesis system. Therefore, in order to improve the functional requirements of the speech synthesis system, a corresponding prompt message will be generated to prompt the user to improve the function.
[0183] Refer to Figure 8 , in some embodiments, the preliminary test template includes filled fields, and step S130 includes but is not limited to sub-steps S810 to S840.
[0184] S810. If the version type information indicates an iterative type, obtain the historical test data of the speech synthesis system.
[0185] It can be understood that when it is determined according to the system attribute data of the speech synthesis system that the speech synthesis system is iteratively developed, the historical test data of the historical development version of the speech synthesis system is obtained. It can be understood that the historical test data includes the historical test data (i.e., the historical test plan) for which filling operations have been performed.
[0186] S820. Obtain the field attribute of the filled field; wherein, the field attribute includes a reusable type.
[0187] It can be understood that the preliminary test template includes multiple filling fields, and each filling field has a field attribute of reusable or non-reusable type. Among them, the reusable type indicates that the content corresponding to the filling field is the general content of the speech synthesis system (such as the content corresponding to filling fields like background, purpose, etc.). Therefore, the content corresponding to the filling field can be reused in the preliminary test templates of different iterative versions of the speech synthesis system to obtain the corresponding target test data. The non-reusable type indicates that the content corresponding to the filling field is only applicable to the speech synthesis system of the corresponding iterative version, such as the content corresponding to filling fields like version number, newly added required functions, etc.
[0188] S830. Use the filling fields with the field attribute of reusable type as the fields to be processed;
[0189] Specifically, different types of identifiers can be set for the filling fields in the preliminary test template, and the field attribute of the corresponding filling field can be determined by identifying the type of the identifier. Or, through technologies such as OCR (Optical Character Recognition), the text content of the filling field is recognized, and the recognized text content is compared with the text content in the preset reusable database, so as to determine the field attribute of the corresponding filling field. It can be understood that the above methods for determining the field attribute of the filling field are only exemplary, and the embodiments of the present application do not make specific limitations in this regard. When it is determined that the field attribute of the filling field is of the reusable type, the filling field is used as the field to be processed.
[0190] S840. Perform a filling operation on the fields to be processed according to the historical test data to obtain the target test data.
[0191] It can be understood that when the field attribute of the filling field is of the reusable type, it indicates that the preliminary test template can be filled according to the content of the corresponding filling field in the historical test data. Therefore, find the filling field corresponding to the field to be processed from the historical test data, associate the content corresponding to the filling field with the field to be processed, that is, fill the preliminary test template according to the content corresponding to the filling field, and use the preliminary test template after all filling operations are completed as the target test data.
[0192] It can be understood that when the field attribute of the filling field is of the non-reusable type, the content corresponding to the filling field can be obtained by means such as user input, automatic recognition of development project files, etc., and the embodiments of the present application do not make specific limitations in this regard.
[0193] It can be understood that completing all filling operations includes filling the content corresponding to all reusable filling fields and filling the content corresponding to all non-reusable filling fields.
[0194] The test method for the speech synthesis system provided by the embodiments of the present application first realizes the automatic generation of target test data through a preset first test database and a filling operation; and realizes the automatic search for test plan data through a preset second test database, thus avoiding the manual method in the related art for formulating target test data and test plan data, and to a certain extent, accelerating the test progress of the speech synthesis system and improving the test efficiency of the speech synthesis system. Secondly, by judging the full coverage, correctness, repeatability, and contradictoriness of test cases, the accuracy of the speech synthesis system test is ensured.
[0195] It can be understood that the test method for the speech synthesis system provided by the embodiments of the present application can also judge the test risks according to methods such as a decision tree classification model. Specifically, the test risks are classified into three categories: high, medium, and low. When it is determined according to the decision tree classification model that the current test operation is a high-risk operation, a corresponding warning prompt signal can be generated. Among them, the decision factors of the decision tree classification model include the judgment results of the full coverage, correctness, repeatability, and contradictoriness of test cases, as well as the random selection results in random testing, the pressure, load, etc. in performance testing. The embodiments of the present application do not make specific limitations on this.
[0196] Refer to Figure 9 , in some embodiments, the embodiments of the present application also provide a test device for a speech synthesis system. The test device includes:
[0197] A system attribute data acquisition module 910, configured to acquire system attribute data of the speech synthesis system to be tested from the application end; wherein, the speech synthesis system runs on the application end, and the system attribute data includes the source information of the application end and the version type information of the speech synthesis system;
[0198] A target test data acquisition module 920, configured to screen out a preliminary test template from a preset first test database according to the source information; and perform a filling operation on the preliminary test template according to the version type information to obtain target test data;
[0199] A test plan data acquisition module 930, configured to screen out test plan data from a preset second test database according to the target test data;
[0200] A test case generation module 940, configured to obtain an initial test case according to the test plan data; and screen the initial test case according to a preset screening condition to obtain a target test case;
[0201] A test module 950, configured to test the speech synthesis system according to the target test case.
[0202] It can be seen that the content in the embodiments of the test method of the above voice synthesis system is applicable to the embodiments of the test device of the present voice synthesis system. The functions specifically implemented by the embodiments of the test device of the present voice synthesis system are the same as those of the embodiments of the test method of the above voice synthesis system, and the beneficial effects achieved are also the same as those of the embodiments of the test method of the above voice synthesis system.
[0203] An embodiment of the present application also provides an electronic device, including:
[0204] At least one memory;
[0205] At least one processor;
[0206] At least one program;
[0207] The program is stored in the memory, and the processor executes at least one program to implement the test method of the voice synthesis system described above in the present application. The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a vehicle-mounted computer, etc.
[0208] Referring to Figure 10 , Figure 10 FIG. shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0209] A processor 1010, which can be implemented in a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., for executing relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0210] A memory 1020, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called by the processor 1010 to execute the test method of the voice synthesis system of the embodiments of the present application;
[0211] An input / output interface 1030, for implementing information input and output;
[0212] A communication interface 1040 for implementing communication interaction between this device and other devices, which can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0213] A bus 1050 for transmitting information between various components of the device (such as a processor 1010, a memory 1020, an input / output interface 1030, and a communication interface 1040);
[0214] Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 achieve communication connections with each other inside the device through the bus 1050.
[0215] The embodiment of the present application also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the test method of the above voice synthesis system.
[0216] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0217] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0218] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.
[0219] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0220] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.
[0221] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0222] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0223] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0224] The unit described as a separate component may or may not be physically separated, and the component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of these units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0225] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0226] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc., various media that can store programs.
[0227] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.
Claims
1. A testing method for a speech synthesis system, characterized in that, it includes: Obtain the system attribute data of the speech synthesis system to be tested from the application end; wherein, the speech synthesis system runs on the application end, and the system attribute data includes the source information of the application end and the version type information of the speech synthesis system; Filter out a preliminary test template from a preset first test database according to the source information; Perform a filling operation on the preliminary test template according to the version type information to obtain target test data; Filter out test plan data from a preset second test database according to the target test data; Obtain an initial test case according to the test plan data; Filter the initial test case according to a preset filtering condition to obtain a target test case; Test the speech synthesis system according to the target test case; Wherein, the system attribute data further includes acoustic training parameters, and the method further includes constructing a target speech recognition model, specifically including: Obtain language knowledge according to a preset language database; wherein, the language knowledge includes acoustic knowledge, phonetic knowledge, and language prior knowledge; Input the acoustic knowledge and the phonetic knowledge into an original acoustic model for training to obtain a preliminary acoustic model; Input the language prior knowledge into an original language model for training to obtain a preliminary language model; Construct a preliminary speech recognition model according to the preliminary acoustic model and the preliminary language model; Obtain speech information from a target object; Extract feature parameters from the speech information according to the acoustic training parameters to obtain speech parameters; Select the target speech recognition model from the preliminary speech recognition model according to the speech parameters, and the target speech recognition model is used to filter the initial test case.
2. The testing method for a speech synthesis system according to claim 1, characterized in that, The test plan data includes a test type, and the test type includes a functional test; The obtaining an initial test case according to the test plan data includes: If the test type is the functional test, obtain the test scenario information of the speech synthesis system; Filter out the initial test case from a preset test case database according to the test scenario information; or, input the test scenario information into a preset text generation model for text generation to obtain the initial test case.
3. The testing method for a speech synthesis system according to claim 1, characterized in that, The filtering condition includes a matching pass result, and the test plan data includes test requirement information; The filtering the initial test case according to a preset filtering condition to obtain a target test case includes: Obtain the test title information of the initial test case; Obtain the test requirement information according to the test plan data; Compare the test requirement information with the test title information; If the test requirement information matches the test title information, obtain the matching pass result; Use the initial test case as the target test case according to the matching pass result.
4. The test method for a speech synthesis system according to claim 1, wherein, the screening conditions include comparison passing results; the screening of the initial test cases according to the preset screening conditions to obtain target test cases includes: obtaining a first test audio generated by the speech synthesis system according to the initial test cases; inputting the first test audio into the target speech recognition model for recognition to obtain a test text; comparing the test text with the initial test cases; if the test text is consistent with the initial test cases, obtaining the comparison passing results; taking the initial test cases as the target test cases according to the comparison passing results.
5. The test method for a speech synthesis system according to any one of claims 1 to 4, wherein, the test method further includes: if the version type information indicates iterative, obtaining the historical test cases of the speech synthesis system; obtaining actual test cases according to the historical test cases and the target test cases; testing the speech synthesis system according to the actual test cases.
6. The test method for a speech synthesis system according to any one of claims 1 to 4, wherein, the test method further includes: obtaining a second test audio generated by the speech synthesis system according to the target test cases; analyzing the audio content of the second test audio to obtain an analysis result; wherein, the analysis result includes an error result indicating a functional defect of the speech synthesis system; generating a prompt message for prompting function improvement according to the error result.
7. The test method for a speech synthesis system according to any one of claims 1 to 4, wherein, the preliminary test template includes filling fields; the filling operation on the preliminary test template according to the version type information to obtain target test data includes: if the version type information indicates iterative, obtaining the historical test data of the speech synthesis system; obtaining the field attributes of the filling fields; wherein, the field attributes include reusable type; taking the filling fields with the field attribute of the reusable type as the fields to be processed; performing a filling operation on the fields to be processed according to the historical test data to obtain the target test data.
8. An electronic device, wherein, it includes: at least one memory; at least one processor; at least one computer program; the computer program is stored in the memory, and the processor executes the at least one computer program to implement the test method for a speech synthesis system according to any one of claims 1 to 7.
9. A computer-readable storage medium, wherein, the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the test method for a speech synthesis system according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for testing speech synthesis quality
CN109473121A
Voice intelligent customer service automatic testing method and system and storage medium
CN113836010A