A semantic testing method, device and computer-readable storage medium
By automatically generating simulated speech and setting preset intervals between them, the problems of mutual interference and low efficiency in speech function testing in the prior art are solved, and efficient and accurate semantic testing is achieved.
Patent Information
- Application Number
- CN202311101274.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-08-30
AI Technical Summary
In the existing voice function test, multiple testers are prone to interfering with each other when inputting voice through the microphone, and the efficiency of manually inputting voice is low, making it difficult to achieve efficient semantic testing.
By automatically generating multiple analog voices, converting them into analog voices using a preset text list, and setting a preset interval between analog voices to generate analog audio to simulate tone pauses between real words, thereby improving the accuracy of semantic recognition.
It realizes that simulated speech can be automatically generated without the participation of multiple testers, avoids interference in semantic testing, and improves testing efficiency and accuracy.
Smart Images

Figure CN117037768B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of testing, and in particular to a semantic testing method, apparatus and computer-readable storage medium. Background Art
[0002] Existing voice function testers usually speak voice commands into a microphone and then test the functions they developed. When conducting batch testing on thousands of voices, testers need to use devices such as artificial mouths to complete automated testing. During the testing process, multiple testers input voices into their respective microphones. When adjacent testers output their respective voices, interference will occur, and the testing efficiency of semantic testing by manually inputting voices is relatively low. Summary of the Invention
[0003] In view of the above problems, the embodiments of the present application provide a semantic testing method, apparatus and computer-readable storage medium to automatically generate multiple simulated voices and improve the semantic testing efficiency.
[0004] According to one aspect of the embodiments of the present application, a semantic testing method is provided. The testing method includes: determining whether the acquisition mode of voice data is a target acquisition mode according to the read application configuration information; if it is the target acquisition mode, converting multiple preset texts in a preset text list into multiple simulated voices, and generating a simulated audio according to the multiple simulated voices; wherein the simulated audio includes multiple simulated voices with a preset interval duration between the simulated voices; identifying semantic information according to the simulated audio, and determining a test result according to the semantic information.
[0005] In an optional manner, the determining whether the acquisition mode of voice data is a target acquisition mode according to the read application configuration information further includes: responding to a user-triggered test instruction, reading the application configuration information specified by the test instruction; detecting whether the read application configuration information is target application configuration information, and determining whether the acquisition mode of voice data is a target acquisition mode according to the detection result; if the detection result indicates that the read application configuration information is the target application configuration information, determining that the acquisition mode of the voice data is the target acquisition mode.
[0006] In an optional manner, before converting multiple preset texts in the preset text list into multiple simulated voices, the testing method further includes: writing the multiple preset texts into an initial text list, and setting blank texts with a preset length between each preset text to obtain the preset text list.
[0007] In an alternative manner, generating the simulated audio according to the multiple simulated voices further includes: obtaining the blank text of the preset length between each preset text, and converting the blank text of the preset length into blank voices of the preset interval duration; setting the blank voices of the preset interval duration between every two simulated voices to generate the simulated audio.
[0008] In an alternative manner, recognizing semantic information according to the simulated audio and determining a test result according to the semantic information further includes: traversing the simulated voices in the simulated audio, and using the traversed simulated voice as a target simulated voice; recognizing target semantic information according to the target simulated voice, and determining a test result of the target simulated voice according to the target semantic information, so as to obtain test results corresponding to all simulated voices.
[0009] In an alternative manner, determining the test result of the target simulated voice according to the target semantic information further includes: executing a target action corresponding to the target semantic information, and determining the test result of the target simulated voice according to a target execution result corresponding to the target action; if the target execution result indicates that the target action is successfully executed, determining a test result indicating that the target simulated voice is successfully tested; if the target execution result indicates that the target action fails to be executed, determining a test result indicating that the target simulated voice fails the test.
[0010] In an alternative manner, the test method further includes: obtaining a start time when the target semantic information is executed, and obtaining an end time corresponding to when the test result of the target simulated voice is determined; storing the start time, the end time, and the test result of the target simulated voice into a system log.
[0011] According to another aspect of the embodiments of the present application, there is provided a semantic test device, where the test device includes: a determination module, configured to determine whether a collection method of voice data is a target collection method according to read application configuration information; a conversion module, configured to, if it is the target collection method, convert multiple preset texts in a preset text list into multiple simulated voices, and generate simulated audio according to the multiple simulated voices; where the simulated audio includes multiple simulated voices with a preset interval duration between the simulated voices; a test module, configured to recognize semantic information according to the simulated audio, and determine a test result according to the semantic information.
[0012] According to one aspect of the embodiments of the present application, there is provided an electronic device, including: a controller; a memory, configured to store one or more programs, and when the one or more programs are executed by the controller, execute the above test method.
[0013] According to one aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of a computer, the computer is caused to execute the above-mentioned testing method.
[0014] According to one aspect of the embodiments of the present application, there is also provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned testing method.
[0015] The embodiments of the present application determine whether the acquisition mode of voice data is a target acquisition mode according to the read application configuration information, so as to flexibly select the acquisition mode of voice data; if it is the target acquisition mode, multiple preset texts in a preset text list are converted into multiple simulated voices, and a simulated audio is generated according to the multiple simulated voices; wherein, the simulated audio includes multiple simulated voices with a preset interval duration between the simulated voices; the embodiments of the present application preset an interval duration between the simulated voices to simulate the tone pause duration between real utterances, which is convenient for sentence segmentation processing, thereby improving the accuracy of semantic recognition; semantic information is recognized according to the simulated audio, and a test result is determined according to the semantic information. The embodiments of the present application can automatically generate simulated voices to improve the efficiency of semantic testing.
[0016] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to be able to understand the technical means of the embodiments of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the embodiments of the present application more obvious and understandable, the following specifically describes the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is a schematic flowchart of a method for testing semantics shown in an exemplary embodiment of the present application.
[0019] Figure 2 is a schematic structural diagram of a voice data acquisition interface shown in an exemplary embodiment of the present application.
[0020] Figure 3is based on Figure 1 The flowchart diagram shows another semantic test method illustrated by the exemplary embodiment shown.
[0021] Figure 4 The diagram shows the application scenario of the semantic test method of this application.
[0022] Figure 5 The flowchart diagram shows the process of determining the semantic test scenario illustrated by an exemplary embodiment of this application.
[0023] Figure 6 The structural diagram shows the semantic test device illustrated by an exemplary embodiment of this application.
[0024] Figure 7 The structural diagram shows the computer system of the electronic device illustrated by an exemplary embodiment of this application. Detailed Description of the Embodiment
[0025] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.
[0026] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0027] The flowcharts shown in the drawings are only exemplary descriptions and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.
[0028] In this application, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0029] Semantic testing for tens of thousands of voices usually requires joint development by multiple testers. Each tester inputs voice into his or her own microphone, which, on the one hand, affects the work efficiency of other testers at the workstation. On the other hand, adjacent testers interfere with each other when outputting their own voices. In addition, the efficiency of semantic testing through manual voice input is low.
[0030] To this end, one aspect of the present application provides a semantic testing method that can automatically generate multiple simulated voices without the need for the tester to input voice through a microphone, thereby avoiding interference in the semantic testing process and improving the efficiency of the semantic testing. Figure 1 , Figure 1 1 is a flowchart of a semantic testing method shown in an exemplary embodiment of the present application. The testing method at least includes S110 to S130, which are described in detail as follows:
[0031] S110: Determine whether the voice data collection method is the target collection method according to the read application configuration information.
[0032] The voice data collection method of this embodiment includes two methods, one is to collect the voice data input by the test personnel in real time, and the other is to generate simulated audio in this embodiment. Figure 2 , Figure 2 It is a structural diagram of a voice data acquisition interface shown in an exemplary embodiment of the present application. Among them, the voice data acquisition interface is a Record interface, and the Record interface connects a physical device recorder and a TTS recorder; the physical device recorder includes an Android recording system; the TTS recorder includes a TTS instance and a text reading module; the physical device recorder is used to record sounds in an actual environment; and the TTS recorder is used in semantic testing scenarios. Voice data collected by different collection methods can all be assigned to the Record object, that is, voice data collected by different collection methods share the same interface, and follow the Liskov Substitution Principle, so that the voice data collection method can be seamlessly switched. Especially when the interface is used frequently such as pausing and resuming, there is no need to implant code, and the same parameter list is used during initialization, which greatly reduces the probability of problems during the switching of collection methods.
[0033] S120: If it is a target collection method, multiple preset texts in the preset text list are converted into multiple simulated voices, and simulated audio is generated according to the multiple simulated voices; wherein the simulated audio includes multiple simulated voices with preset intervals between the simulated voices.
[0034] The target acquisition method in this embodiment is the second acquisition method mentioned above, by setting Figure 2The shown TTS (Text To Speech) recorder Figure 2 is a schematic structural diagram of the TTS recorder shown in an exemplary embodiment of the present application. Among them, the TTS recorder includes a TTS instance and a text reading module; among them, the TTS instance converts multiple preset texts in a preset text list into multiple simulated voices to generate a simulated audio; the text reading module can read texts in multiple formats to generate multiple preset texts in the preset text list, for example, read multiple preset texts from a txt text or an excel text.
[0035] Among them, each line of text in the preset text list corresponds to generating a simulated voice, and there is a preset interval duration between each simulated voice, for example, 3 - 5 seconds, that is, after the current simulated voice finishes playing, it is necessary to pause for 3 - 5 seconds to play the next simulated voice, so as to simulate the intonation pause between real words, which is convenient for sentence segmentation and improves the accuracy of recognizing the semantics of each simulated voice at the same time.
[0036] The simulated voice and the simulated audio in the present application can both be in a mute form, that is, they can be played and tested in a mute form, etc., so as to avoid sound interference during the test process.
[0037] S130: Recognize semantic information based on the simulated audio, and determine the test result according to the semantic information.
[0038] Exemplarily, recognize the semantic information of each simulated voice in the simulated audio one by one, and execute the action instruction corresponding to each semantic information one by one to obtain the test result of each semantic information, and summarize to obtain a result test table.
[0039] Another exemplarily, execute the action instruction corresponding to each semantic information simultaneously to determine the corresponding test result according to the execution result corresponding to each semantic information, and summarize to obtain a result test table.
[0040] In this embodiment, it is determined whether the acquisition method of the voice data is the target acquisition method according to the read application configuration information, so as to flexibly select the acquisition method of the voice data; if it is the target acquisition method, multiple preset texts in the preset text list are converted into multiple simulated voices, and a simulated audio is generated according to the multiple simulated voices; among them, the simulated audio includes multiple simulated voices with a preset interval duration between the simulated voices; in this embodiment, by presetting the interval duration between the simulated voices, the intonation pause duration between real words is simulated, which is convenient for sentence segmentation processing, thereby improving the accuracy of semantic recognition; recognize semantic information based on the simulated audio, and determine the test result according to the semantic information. This embodiment can automatically generate simulated voices, without the need for multiple testers to participate in the test, saving labor costs while improving the semantic test efficiency.
[0041] In another exemplary embodiment of the present application, details are introduced on how to determine whether the acquisition method of voice data is the target acquisition method according to the read application configuration information. For details, please refer to Figure 3 , Figure 3 is based on Figure 1 shown in the exemplary embodiment is a schematic flowchart of another semantic test method. This test method further includes S310 to S330 in S110 as shown in Figure 1 shown below in detail:
[0042] S310: In response to the received user-triggered test instruction, read the application configuration information specified by the test instruction.
[0043] The user can select the corresponding voice data acquisition method according to actual needs. The user triggers the test instruction to the execution end of this embodiment. The execution end, in response to this test instruction, reads the application configuration information specified in this test instruction, so as to determine the corresponding voice data acquisition method according to the read application configuration information.
[0044] Exemplarily, the execution end, in response to the received user-triggered test instruction, reads the number of the application configuration information specified in the test instruction, and determines the corresponding application configuration information according to this number. For example, number 1 corresponds to the first application configuration information, and number 2 corresponds to the second application configuration information.
[0045] S320: Detect whether the read application configuration information is the target application configuration information, and determine whether the acquisition method of voice data is the target acquisition method according to the detection result.
[0046] Exemplarily, number 1 corresponds to the first application configuration information, and number 2 corresponds to the second application configuration information (i.e., the target application configuration information of this embodiment). If the number of the application configuration information specified in the test instruction is 2, it is determined that the currently read application configuration information is the second application configuration information (target application configuration information), and the acquisition method of voice data is the target acquisition method.
[0047] S330: If the detection result indicates that the read application configuration information is the target application configuration information, determine that the acquisition method of voice data is the target acquisition method.
[0048] If the detection result indicates that the read application configuration information is not the target application configuration information, determine that the acquisition method of voice data is not the target acquisition method.
[0049] This embodiment provides a method for determining a target acquisition method for voice data. By reading the application configuration information specified by a test instruction triggered by a user and matching the application configuration information with the target application configuration information corresponding to the target acquisition method, if the match is successful, the acquisition method for voice data can be quickly determined as the target acquisition method.
[0050] In another exemplary embodiment of the present application, how to construct a preset text list is introduced in detail. The specific steps are as follows: Write multiple preset texts into an initial text list, and set blank texts with a preset length between each preset text to obtain a preset text list.
[0051] The initial text list can be understood as a blank text list, that is, there are no relevant preset texts in the initial text list. The preset text list is a list including at least one preset text, and there are blank texts with a preset length between each preset text.
[0052] Exemplarily, the specific content of the preset text list is "A preset text 00000B preset text 00000C preset text"; among them, "00000" represents the blank text with a preset length.
[0053] Another exemplarily, construct the preset text list shown in Table 1. Table 1 is an exemplary preset text list. Among them, "00000" represents the blank text with a preset length.
[0054] The first preset text Turn on the Bluetooth player 00000 The second preset text Open the window 00000 The third preset text Turn on the radio 00000 …… ……
[0055] Table 1
[0056] In this embodiment, there are blank texts with a preset length between each preset text, so that after generating a simulated audio including multiple simulated voices according to the preset texts, there are blank voices with a preset interval duration between each simulated voice to simulate the tone pause duration between real words, which is convenient for sentence segmentation processing, thereby improving the accuracy of semantic recognition.
[0057] Furthermore, how to generate a simulated audio according to multiple simulated voices is introduced, that is, the above S120 further includes S1201 to S1202, which are introduced in detail as follows:
[0058] S1201: Obtain the blank texts with a preset length between each preset text, and convert the blank texts with a preset length into blank voices with a preset interval duration.
[0059] The blank text can be converted into a blank voice, and the duration of the blank voice is positively correlated with the length of the blank text.
[0060] S1202: Set blank voices with a preset interval duration between every two simulated voices to generate a simulated audio.
[0061] Exemplarily, obtain the first preset text to the third preset text in Table 1, as well as the blank text of "00000", convert the blank text into a 3-second blank voice, and arrange them in the original order of the preset text to generate a simulated audio. If the simulated audio is played, then "open the window" will be played 3 seconds after "turn on the Bluetooth player", and then "turn on the radio" will be played after an interval of 3 seconds. During the process of generating the simulated audio, the preset texts can also be sorted randomly or in other orders, and this embodiment does not limit this.
[0062] In some embodiments, a simulated audio can be generated with a preset text and a blank text, so as to perform semantic tests based on multiple simulated audios, that is, a semantic test is performed every time a simulated audio is generated.
[0063] This embodiment provides a method for generating a simulated audio. By setting a blank voice with a preset interval duration between every two simulated voices, the generated simulated audio is made to be more in line with the human voice audio in the real scenario, better simulating the pause duration of the tone between real words, facilitating sentence segmentation processing, and thus improving the accuracy of semantic recognition.
[0064] In another exemplary embodiment of the present application, how to obtain semantic information based on the recognition of the simulated audio and determine the test result based on the semantic information is introduced in detail, that is, the above S130 further includes S1301 to S1302, which are introduced in detail as follows:
[0065] S1301: Traverse the simulated voices in the simulated audio, and use the traversed simulated voice as the target simulated voice.
[0066] In this embodiment, by means of traversal, it is ensured that each simulated voice in the simulated audio is semantically tested without omission, so as to obtain the test results corresponding to all the simulated voices in the simulated audio.
[0067] S1302: Obtain the target semantic information based on the target simulated voice, and determine the test result of the target simulated voice based on the target semantic information, so as to obtain the test results corresponding to all the simulated voices.
[0068] Exemplarily, when the second simulated voice is traversed, an identification operation is performed on the second simulated voice to obtain the semantic information corresponding to the second simulated voice, that is, the current target semantic information. Based on the target semantic information, relevant comparative analysis is performed to determine the test result of the second simulated voice, and so on, to obtain the test results corresponding to all the simulated voices in the simulated audio.
[0069] This embodiment provides a method for determining the test results corresponding to all analog voices in an analog audio. By traversing each analog voice in the analog audio and identifying the semantic information of the traversed analog voice, the test results of the traversed analog voice are determined, so as to perform semantic tests on each analog voice in the analog audio without omission, and accurately obtain the test results corresponding to all analog voices in the analog audio.
[0070] In another exemplary embodiment of the present application, how to determine the test results of the target analog voice according to the target semantic information is introduced in detail, that is, the above S1302 further includes S13021 to S13023, which are introduced in detail as follows:
[0071] S13021: Execute the target action corresponding to the target semantic information, and determine the test results of the target analog voice according to the target execution result corresponding to the target action.
[0072] S13022: If the target execution result indicates that the target action is successfully executed, determine the test result indicating that the target analog voice test is successful.
[0073] S13023: If the target execution result indicates that the target action fails, determine the test result indicating that the target analog voice test fails.
[0074] Exemplarily, the target action corresponding to the target semantic information is "play the songs of singer A". If the executed action is "played the XX song of singer A", it indicates that the target action has been successfully executed. Furthermore, the test results of the target analog voice indicate that the target analog voice test is successful; if the executed action is "played the songs of singer B" or "turned on the radio", it indicates that the target action fails, and furthermore, the test results of the target analog voice indicate that the target analog voice test fails.
[0075] In some embodiments, all test results can be summarized in a test result list. The test result list not only includes each analog voice and its corresponding test results, but also includes the original preset text corresponding to each analog voice, screenshots of the test pages before and after performing the corresponding actions, etc., and can also include the timestamps before and after performing the corresponding actions. Save the above test result list to the result file path configured in the gradle script file. If Excel is defined, save it in an excel file, and other formats with pictures and texts such as doc are also supported. Testers can perform test analysis by viewing the table results and the corresponding logs after the test is completed.
[0076] This embodiment provides a method for determining the test result of a target simulated voice. By performing a target action corresponding to the target semantic information, the test result of the target simulated voice is indirectly determined based on the execution result, and a positive relationship is established between the execution result of the target action and the test result of the target simulated voice, so as to facilitate the quick determination of the test result of the target simulated voice.
[0077] In some embodiments, the start and end timestamps of performing the target action corresponding to the target semantic information are also recorded to facilitate testers to obtain the corresponding timestamps through the system log to optimize the test process. Specifically, the test method further includes: obtaining the start time of performing the target semantic information and the end time corresponding to determining the test result of the target simulated voice; storing the start time, the end time, and the test result of the target simulated voice in the system log.
[0078] Exemplarily, the start time of performing the target semantic information is 10:00, and the end time corresponding to determining the test result of the target simulated voice is 10:01, that is, the end time of performing the target semantic information is 10:01. Then, the above two timestamps and the test result of the target voice can be recorded in the test result list, saved to the corresponding file, and then stored in the system log to facilitate testers to obtain relevant information to analyze and optimize the test process. In addition, the system log can be compressed and saved to prevent relevant data from being overwritten.
[0079] In some embodiments, the execution duration of the target action can also be calculated based on the start time and the end time and stored in the system log to facilitate testers to query the execution duration of the target action.
[0080] In another exemplary embodiment of the present application, the application scenarios of the above multiple test methods are exemplarily described. For details, please refer to Figure 4 , Figure 4 is a schematic diagram of the application scenario of the semantic test method of the present application. It includes a server 100, and the server 100 can control a TTS 101, a recording module 102, a speech recognition engine module 103, a semantic execution module 104, a test result module 105, and an instruction set 106. They can be connected by wireless communication or wired means, and the present application does not limit the connection means between them.
[0081] Server 100, as an execution end, controls the corresponding modules to execute the test method shown in any of the above exemplary embodiments. Exemplarily, Server 100 determines whether the acquisition mode of the voice data is the target acquisition mode according to the read application configuration information; if it is the target acquisition mode, it converts multiple preset texts in the preset text list into multiple simulated voices, and generates a simulated audio according to the multiple simulated voices; wherein, the simulated audio includes multiple simulated voices with a preset interval duration between the simulated voices; Server 100 identifies semantic information according to the simulated audio, and determines the test result according to the semantic information.
[0082] To more clearly elaborate on the functions of each module and the logical relationships between each module, the following provides a detailed description of each module:
[0083] A Record object is defined in the recording module 102. As Figure 2 shown, the Record interface is connected to the TTS recorder (TTS101) and the physical device recorder. If the recording module 102 determines that the current scenario is in the automated semantic test scenario, it will initialize TTS101. The text reading module in TTS101 will read the corresponding preset instruction text from the instruction set 106, and assign the object after instantiation to the Record object. Since both implementation classes use the same interface and follow the Liskov substitution principle, the recording module 102 can seamlessly switch between the physical device recorder and the TTS recorder, and there is no need to implant any code when using interfaces such as pause and resume frequently, and the parameter list is the same during initialization, greatly reducing the probability of software problems. Among them, the recording data is all in the pcm format, that is, the recording module 102 can convert the data format passed in by the physical device recorder into the pcm format, and the data format generated by TTS101 and passed into the recording module 102 is also the pcm format. In addition, the data format output by the test result module 105 is also the pcm format, and the output data is used to broadcast the test result.
[0084] The recognition speed of the speech recognition engine module 103 is limited. For example, with a sampling rate of 16Khz and a 16Bit mono format, it means that the recording module 102 has to generate 16000 * 2Byte of pcm data per second. If its generation speed is not restricted, it will cause the speech recognition engine module 103 to become stuck. However, if the recording module 102 generates pcm data too slowly, the audio data recognized by the speech recognition engine module 103 will be stuck, and it cannot simulate the scenario of semantic testing through a real physical recording device.
[0085] The semantic execution module 104 receives the semantic information sent by the speech recognition engine module 103, executes the actions corresponding to the semantic information, and sends the execution result to the test result module 105. The test result module 105 determines the corresponding test result based on the received execution result, generates PCM format voice data carrying the test result, and sends it to the playback terminal to play the test result.
[0086] There are multiple preset texts corresponding to the test instructions in the instruction set 106, that is, each test instruction corresponds to a corresponding preset text, so that the text reading module in the TTS 101 can read the preset text corresponding to the corresponding test instruction from it.
[0087] The server 100 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. Among them, multiple servers can form a blockchain, and the server is a node on the blockchain. The server 100 can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. There is no limitation here either.
[0088] Combined with the above application scenarios, it further explains how to determine that the current semantic test scenario is an automated semantic test scenario. The automated semantic test scenario is a scenario that uses the TTS 101 to generate simulated audio for semantic testing. The details are introduced as follows:
[0089] Two paths are configured in the script file of the automated semantic test scenario, one is the test instruction set file path, and the other is the test result file path. These two file paths can be read through BuildConfig in the recording module 102, and then combined with the interface of the file system to determine whether the test instruction set file exists in this file. If it exists, it will be determined that the current test scenario is an automated semantic test scenario. For specific details, please refer to Figure 5 , Figure 5 which is a schematic flowchart of the process for determining the semantic test scenario shown in an exemplary embodiment of the present application.
[0090] First, define two variables of string type in the gradle file: the test instruction set file path and the test result file path, and configure and define the variables so that the defined variable values can be read through BuildConfig at the code level;
[0091] Then, according to the read defined variable values, detect whether the address defined in the test instruction set file is empty and whether it is a valid address; if the address is non-empty and is a valid address, determine that the current semantic test scenario is an automated semantic test scenario.
[0092] Finally, the text reading module in TTS101 reads the preset text corresponding to the test instruction in the excel file in instruction set 106, and the TTS instance in TTS101 generates a pcm format simulated audio according to the read preset text and inputs it into the recording module 102.
[0093] Another aspect of the present application also provides a semantic test device, as Figure 6 shown Figure 6 is a schematic structural diagram of the semantic test device shown in an exemplary embodiment of the present application. The test device 600 includes:
[0094] A determination module 610, configured to determine whether the acquisition mode of the voice data is a target acquisition mode according to the read application configuration information.
[0095] A conversion module 630, configured to, if it is the target acquisition mode, convert multiple preset texts in the preset text list into multiple simulated voices, and generate a simulated audio according to the multiple simulated voices; wherein, the simulated audio includes multiple simulated voices with a preset interval duration between the simulated voices.
[0096] A test module 650, configured to identify semantic information according to the simulated audio, and determine a test result according to the semantic information.
[0097] In an optional manner, the determination module 610 further includes:
[0098] A response unit, configured to respond to a user-triggered test instruction and read the application configuration information specified by the test instruction.
[0099] A detection unit, configured to detect whether the read application configuration information is target application configuration information, and determine whether the acquisition mode of the voice data is a target acquisition mode according to the detection result.
[0100] An acquisition mode determination unit, configured to, if the detection result indicates that the read application configuration information is target application configuration information, determine that the acquisition mode of the voice data is a target acquisition mode.
[0101] In an optional manner, the test device 600 further includes:
[0102] A preset text list generation module, configured to write multiple preset texts into an initial text list, and set blank texts with a preset length between each preset text to obtain a preset text list.
[0103] In an alternative manner, the conversion module 630 further includes:
[0104] An acquisition unit, configured to acquire blank texts of a preset length between each preset text, and convert the blank texts of the preset length into blank voices of a preset interval duration.
[0105] A setting unit, configured to set blank voices of a preset interval duration between every two simulated voices to generate a simulated audio.
[0106] In an alternative manner, the testing module 650 further includes:
[0107] A traversal unit, configured to traverse the simulated voices in the simulated audio, and use the traversed simulated voices as target simulated voices.
[0108] A testing unit, configured to identify target semantic information according to the target simulated voices, and determine test results of the target simulated voices according to the target semantic information, so as to obtain test results corresponding to all the simulated voices.
[0109] In an alternative manner, the testing module 650 further includes:
[0110] An execution unit, configured to execute a target action corresponding to the target semantic information, and determine a test result of the target simulated voice according to a target execution result corresponding to the target action.
[0111] A successful execution unit, configured to determine a test result indicating that the test of the target simulated voice is successful if the target execution result indicates that the target action is successfully executed.
[0112] A failed execution unit, configured to determine a test result indicating that the test of the target simulated voice fails if the target execution result indicates that the target action fails to be executed.
[0113] In an alternative manner, the testing device 600 further includes:
[0114] A time acquisition module, configured to acquire a start time for executing the target semantic information, and acquire an end time corresponding to when the test result of the target simulated voice is determined.
[0115] A storage module, configured to store the start time, the end time, and the test result of the target simulated voice into a system log.
[0116] The test device of the present application determines whether the acquisition mode of voice data is the target acquisition mode according to the read application configuration information, so as to flexibly select the acquisition mode of voice data; if it is the target acquisition mode, multiple preset texts in the preset text list are converted into multiple simulated voices, and a simulated audio is generated according to the multiple simulated voices; wherein, the simulated audio includes multiple simulated voices with a preset interval duration between the simulated voices; the test device of the present application simulates the tone pause duration between real words by presetting an interval duration between the simulated voices, which is convenient for sentence segmentation processing, thereby improving the accuracy of semantic recognition; the semantic information is recognized according to the simulated audio, and the test result is determined according to the semantic information. The test device of the present application can automatically generate simulated voices to improve the efficiency of semantic testing.
[0117] It should be noted that the test device provided in the above embodiment and the test method provided in the foregoing embodiment belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiment, and will not be elaborated here.
[0118] On the other hand, the present application also provides an electronic device, including: a controller; a memory for storing one or more programs, which, when executed by the controller, execute the above-mentioned test method.
[0119] Please refer to Figure 7 , Figure 7 FIG. is a schematic structural diagram of a computer system of an electronic device shown in an exemplary embodiment of the present application, which shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application.
[0120] It should be noted that Figure 7 the computer system 700 of the electronic device shown is only an example, and should not bring any limitation to the functions and usage scope of the embodiments of the present application.
[0121] As Figure 7 shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage part 708 into the random access memory (RAM) 703, such as executing the method in the above embodiment. In the RAM 703, various programs and data required for system operation are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.
[0122] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as required. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as required so that a computer program read from it can be installed into the storage section 708 as required.
[0123] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by a central processing unit (CPU) 701, various functions defined in the system of the present application are executed.
[0124] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0125] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in an order different from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0126] The units involved in the embodiments of the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the unit itself in some cases.
[0127] On the other hand, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned test method is implemented. The computer-readable storage medium can be included in the electronic device described in the above embodiments, or can exist alone without being assembled into the electronic device.
[0128] On the other hand, the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the test methods provided in the above various embodiments.
[0129] According to one aspect of the embodiments of the present application, a computer system is also provided, including a central processing unit (CPU). It can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage section into a random access memory (RAM), such as executing the method in the above embodiments. In the RAM, various programs and data required for system operation are also stored. The CPU, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.
[0130] The following components are connected to the I / O interface: an input section including a keyboard, a mouse, etc.; an output section including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section including a hard disk, etc.; and a communication section including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section performs communication processing via a network such as the Internet. A drive is also connected to the I / O interface as required. A removable medium, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive as required, so that a computer program read from it can be installed into the storage section as required.
[0131] The above content is only a preferred exemplary embodiment of the present application and is not intended to limit the implementation of the present application. Those of ordinary skill in the art can easily make corresponding adaptations or modifications according to the main concept and spirit of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope required by the claims.
Claims
1. A semantic testing method, characterized in that, the testing method includes: detecting whether the address defined in the test instruction set file is empty and whether it is a valid address; if the address is non-empty and is the valid address, determining the acquisition method as the target acquisition method, where the target acquisition method is a method of acquiring the simulated audio generated by the TTS module in an automated voice testing scenario; the test instruction set file includes preset texts corresponding to multiple test instructions; if it is the target acquisition method, obtaining a preset text list to the TTS module based on the address, so that the TTS module converts multiple preset texts in the preset text list into multiple simulated voices, converts the blank texts between each preset text into blank voices with a preset interval duration, and generates the simulated audio according to the multiple simulated voices and the blank voices; wherein, the simulated audio includes multiple simulated voices with blank voices having a preset interval duration between the simulated voices; identifying semantic information according to the simulated audio, and determining a test result according to the semantic information.
2. The testing method according to claim 1, characterized in that, the testing method further includes: writing the multiple preset texts into an initial text list, and setting blank texts with a preset length between each preset text to obtain the preset text list.
3. The testing method according to claim 2, characterized in that, generating the simulated audio according to the multiple simulated voices and the blank voices further includes: obtaining the blank texts with the preset length between each preset text, and converting the blank texts with the preset length into blank voices with the preset interval duration; setting blank voices with the preset interval duration between every two simulated voices to generate the simulated audio.
4. The testing method according to claim 1, characterized in that, identifying semantic information according to the simulated audio, and determining a test result according to the semantic information further includes: traversing the simulated voices in the simulated audio, and taking the traversed simulated voice as the target simulated voice; identifying target semantic information according to the target simulated voice, and determining the test result of the target simulated voice according to the target semantic information to obtain the test results corresponding to all simulated voices.
5. The testing method according to claim 4, characterized in that, determining the test result of the target simulated voice according to the target semantic information further includes: executing the target action corresponding to the target semantic information, and determining the test result of the target simulated voice according to the target execution result corresponding to the target action; if the target execution result indicates that the target action is successfully executed, determining a test result indicating that the target simulated voice test is successful; if the target execution result indicates that the target action fails to be executed, determining a test result indicating that the target simulated voice test fails.
6. The testing method according to claim 5, characterized in that, the testing method further includes: Obtain the start time of executing the target semantic information, and obtain the end time corresponding to the test result of determining the target simulated speech; Store the start time, end time and the test result of the target simulated speech in the system log.
7. A semantic test device, Characterized in that, The test device includes: A determination module, configured to detect whether the address defined in the test instruction set file is empty and whether it is a valid address; if the address is non-empty and is the valid address, determine that the acquisition method is the target acquisition method, and the target acquisition method is a method for acquiring the simulated audio generated by the TTS module in an automated speech test scenario; the test instruction set file includes preset texts corresponding to multiple test instructions; A conversion module, configured to, if it is the target acquisition method, obtain a preset text list to the TTS module based on the address, so that the TTS module converts multiple preset texts in the preset text list into multiple simulated speeches, converts the blank texts between the preset texts into blank speeches with a preset interval duration, and generates the simulated audio according to the multiple simulated speeches; wherein, the simulated audio includes multiple simulated speeches with blank speeches having a preset interval duration between the simulated speeches; A test module, configured to identify semantic information according to the simulated audio and determine a test result according to the semantic information.
8. An electronic device, Characterized in that, It includes: A controller; A memory, configured to store one or more programs, and when the one or more programs are executed by the controller, enable the controller to implement the test method according to any one of claims 1 to 6.
9. A computer-readable storage medium, Characterized in that, Computer-readable instructions are stored thereon, and when the computer-readable instructions are executed by a processor of a computer, cause the computer to execute the test method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Unattended cloud voice library collection and intelligent product testing system and method
CN108109633A
Voice product test method, device and equipment and computer readable medium
CN109003602A