Voice assistant verification method and device, computer equipment, storage medium and program product

Through the automated voice assistant verification method, the test client and log information comparison technology is used to solve the inaccuracy and high cost problems caused by manual detection in voice assistant verification, and an efficient and reliable voice assistant performance evaluation is achieved.

CN120276996APending Publication Date: 2025-07-08CHENGDU CELIS TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510751186.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing voice assistant verification methods rely on manual detection, resulting in inaccurate verification results, inefficient, high cost and poor versatility.

Method used

Through the test client, the synthetic voice with different tones is obtained, and the voice assistant to be tested is played, and its response process information is recorded. The log information is automatically compared with the expected response process information to generate verification results.

Benefits of technology

It improves the reliability and efficiency of verification results, reduces human subjective bias and cost, and can automatically conduct multi-tone recognition and long-term dialogue ability testing, which is suitable for various devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276996A_ABST
    Figure CN120276996A_ABST
Patent Text Reader

Abstract

The invention relates to a voice assistant verification method and device, computer equipment, a storage medium and a program product. The method comprises the following steps: a test client obtains at least one piece of synthetic voice with different timbres and plays the synthetic voice to a tested voice assistant; the test client obtains the response process information of the voice assistant to the synthesized voice by obtaining the log information associated with the voice assistant; and the test client further obtains a comparison result of the response process information of the voice assistant and the expected response process information corresponding to the synthesized voice, and generates a verification result of the tested voice assistant based on the comparison result. The whole verification process can be automatically executed according to a set program, the tested voice assistant is tested through synthesized voices with different timbres, whether all processing links of the voice assistant meet expectations or not is determined through information comparison, corresponding verification results are automatically generated accordingly, the objectivity of the verification results is ensured, and the verification efficiency is improved. And meanwhile, the verification efficiency is also improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice technologies, and in particular, to a method, apparatus, computer device, storage medium, and computer program product for verifying a voice assistant. Background Art

[0002] The process verification of existing voice assistants mainly relies on manual detection (relying on the mouth, eyes, and ears). To comprehensively verify the performance of a voice assistant, a large number of test tasks need to be executed; during the manual detection process, the results are mainly obtained through the subjective observation of testers, and different testers may obtain different verification results, and it is difficult to effectively observe the speech recognition performance or speech synthesis performance of the voice assistant.

[0003] It can be seen that the verification method of traditional technologies for voice assistants has at least the problem of inaccurate verification results. Summary of the Invention

[0004] Based on this, to address the above technical problems, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for verifying a voice assistant.

[0005] In a first aspect, this application provides a method for verifying a voice assistant, which is applied to a test client. The method includes:

[0006] Obtain at least one synthesized speech, each synthesized speech having a different timbre;

[0007] Play the synthesized speech to the voice assistant to be tested;

[0008] Obtain the log information of the voice assistant, where the log information includes the response process information of the voice assistant to the synthesized speech;

[0009] Obtain the comparison result between the response process information of the voice assistant and the expected response process information corresponding to the synthesized speech;

[0010] Generate the verification result of the voice assistant based on the comparison result.

[0011] In combination with the first aspect, in one of the embodiments, the response process information of the voice assistant includes at least one of: the recognized corpus obtained by recognizing the synthesized speech, the target response corpus matched based on the recognized corpus, and the response time for the synthesized speech;

[0012] The expected response process information corresponding to the synthesized speech includes at least one of: the actual corpus corresponding to the synthesized speech, the expected response corpus corresponding to the synthesized speech, and the expected response time corresponding to the synthesized speech.

[0013] In combination with the first aspect, in one embodiment, the method further includes:

[0014] Based on the log information recorded during the speech synthesis process, obtain the actual corpus corresponding to the synthesized speech;

[0015] Obtain the target text information corresponding to the actual corpus from a preset information library, use the target text information as the expected response corpus, and obtain the expected response time of the synthesized speech according to the target text information and a preset speech rate; wherein, the preset information library is the information library corresponding to the speech assistant under test.

[0016] In combination with the first aspect, in one embodiment, the comparison result of obtaining the response process information of the speech assistant and the expected response process information corresponding to the synthesized speech includes at least one of the following:

[0017] Obtain the first comparison result between the recognized corpus of the speech assistant and the actual corpus corresponding to the synthesized speech;

[0018] Obtain the second comparison result between the target response corpus of the speech assistant and the expected corpus corresponding to the synthesized speech;

[0019] Obtain the third comparison result between the actual response time of the speech assistant and the expected response time.

[0020] In combination with the first aspect, in one embodiment, the method further includes:

[0021] Collect the audio stream of the response speech of the speech assistant in reply to the synthesized speech;

[0022] In response to the end of this round of voice interaction, obtain the actual response corpus of the speech assistant based on the speech recognition result of the audio stream of the response speech;

[0023] Obtain the fourth comparison result between the actual response corpus of the speech assistant and the target response corpus matched by the speech assistant;

[0024] The generating the verification result of the speech assistant based on the comparison result includes:

[0025] Generate the verification result of the speech assistant based on at least one of the first comparison result, the second comparison result, the third comparison result, and the fourth comparison result.

[0026] In combination with the first aspect, in one embodiment, the generating the verification result of the speech assistant based on at least one of the first comparison result, the second comparison result, the third comparison result, and the fourth comparison result includes:

[0027] In response to any of the comparison results being failed, generate a verification failure report for the voice assistant;

[0028] In response to all comparison results being passed, generate a verification pass report for the voice assistant.

[0029] Combined with the first aspect, in one embodiment, the obtaining methods of the first comparison result, the second comparison result, and the fourth comparison result include:

[0030] For two corpora to be compared, determine the edit distance between the two corpora;

[0031] In response to the edit distance being greater than a preset similarity threshold, determine that the comparison result of the two corpora is passed; otherwise, determine that the comparison result of the two corpora is failed;

[0032] Wherein, the similarity threshold is determined based on a preset error propagation probability and the length of the propagated string.

[0033] Combined with the first aspect, in one embodiment, the obtaining method of the third comparison result includes:

[0034] Obtain the time difference between the actual response time and the expected response time of the voice assistant;

[0035] In response to the time difference being less than or equal to a preset timeout threshold, determine that the third comparison result is passed; otherwise, determine that the third comparison result is failed.

[0036] Combined with the first aspect, in one embodiment, the obtaining of at least one synthesized voice includes:

[0037] Obtain a corpus list to be synthesized and a timbre list to be synthesized; the corpus list contains at least one text corpus, and the timbre list contains at least one timbre information;

[0038] Send the corpus list and the timbre list to the server, and generate synthesized voices with different timbres corresponding to each text corpus in the corpus list based on a voice synthesis model preset on the server side.

[0039] In a second aspect, there is also provided a voice assistant verification device, which is applied to a test client. The device includes:

[0040] A synthesized voice acquisition module, configured to acquire at least one synthesized voice, and each synthesized voice has a different timbre;

[0041] A broadcast control module, configured to play the synthesized voice to the voice assistant to be tested;

[0042] A log information acquisition module, configured to acquire log information of the voice assistant, where the log information includes information on the response process of the voice assistant to the synthesized speech;

[0043] A comparison module, configured to obtain a comparison result between the response process information of the voice assistant and expected response process information corresponding to the synthesized speech;

[0044] A result generation module, configured to generate a verification result of the voice assistant based on the comparison result.

[0045] In a third aspect, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method according to any one of the above first aspects are implemented.

[0046] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method according to any one of the above first aspects are implemented.

[0047] In a fifth aspect, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the method according to any one of the above first aspects are implemented.

[0048] For the above voice assistant verification method, device, computer device, storage medium, and computer program product, the test client acquires at least one synthesized speech with different timbres and plays the synthesized speech to the voice assistant to be tested; after receiving the synthesized speech, the voice assistant to be tested performs relevant processing, and the processing process information is recorded in the log associated with the voice assistant; the test client obtains the response process information of the voice assistant to the synthesized speech by acquiring the log information associated with the voice assistant; the test client can further obtain a comparison result between the response process information of the voice assistant and the expected response process information corresponding to the synthesized speech, and generate a verification result of the voice assistant to be tested based on these comparison results. The entire verification process can be automatically executed according to a set program, and at the same time, the voice assistant to be tested is tested with synthesized speeches of different timbres, which can improve the reliability of the verification result; in addition, the test device can also determine whether each processing link of the voice assistant meets the expectations based on the log information of the voice assistant during the test process and through information comparison, and automatically generate a corresponding verification result, avoiding subjective biases existing in manual verification and evaluation, ensuring the objectivity of the verification result, and also improving the verification efficiency and reducing the verification cost compared with the manual verification method. Description of the Drawings

[0049] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the accompanying drawings required for the description of the embodiments or related technologies. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0050] Figure 1 It is an application environment diagram of the voice assistant verification method in an embodiment;

[0051] Figure 2 It is a schematic flowchart of the process for a test client to implement the voice assistant verification method in an embodiment;

[0052] Figure 3 It is a schematic overall flowchart of the process for a test system to implement the voice assistant verification method in an embodiment;

[0053] Figure 4 It is a schematic specific flowchart of the process for a test system to implement the voice assistant verification method in an embodiment;

[0054] Figure 5 It is a schematic timing diagram of the process for a test client and a server to implement the voice assistant verification method in an embodiment;

[0055] Figure 6 It is a schematic flowchart of the execution process of relevant threads during the process of a test client executing the single - time voice assistant verification method in an embodiment;

[0056] Figure 7 It is a structural block diagram of a voice assistant verification device in an embodiment;

[0057] Figure 8 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0058] In order to make the purpose, technical solutions and advantages of the present application clearer and more understandable, the following further details the present application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0059] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data need to comply with the relevant regulations.

[0060] Before elaborating on the embodiments of the present application in detail, some terms related to the embodiments of the present application are described as follows:

[0061] Speech: Refers to the sound signals generated by the human vocal organs, which is the oral manifestation of language. Its main feature is that it exists in the form of audio (such as recording files, real-time conversations), and contains paralinguistic information such as intonation, rhythm, and stress. Usually, speech can be converted into text through speech recognition technology (ASR, Automatic Speech Recognition).

[0062] Corpus: Refers to a collection of language data collected according to specific rules, usually in text form. Its main feature is that it exists in the form of text, and has been annotated or structured (such as part-of-speech tagging, syntactic analysis). In most cases, it is used to train language models or study language rules.

[0063] ASR, Automatic Speech Recognition, also known as speech-to-text technology. ASR is a technology that converts human speech signals into text or machine-readable instructions. Its basic principle is to first preprocess the speech signal, including noise reduction, endpoint detection, etc., to improve the quality of the speech signal and determine the start and end positions of the speech; then, extract features from the preprocessed speech signal to convert the speech signal into a sequence of feature vectors; next, use the acoustic model and language model to process and recognize the feature vectors. The acoustic model is used to match the speech features with acoustic units such as phonemes or syllables, and the language model adjusts and optimizes the recognition results according to the knowledge of language grammar, vocabulary, and semantics, and finally outputs the recognized text content.

[0064] TTS, the abbreviation of Text-to-Speech, that is, text-to-speech technology, also known as speech synthesis technology in some cases.

[0065] With the development of speech technology, speech assistants in intelligent devices such as cars and robots provide users with richer interaction methods and information transmission technologies. Taking the speech assistant in a car as an example, it can help users timely understand the parameters of the device and ensure the normal operation and driving safety of the vehicle.

[0066] Regarding the related functions of the voice assistant in the device, it is usually necessary to verify through process phenomena to ensure that the voice assistant is in the best working condition. In the traditional solution, the process verification of the voice assistant is mainly manual detection (relying on the mouth, eyes, and ears). Generally, a large number of test tasks need to be executed. During the manual detection process, the ASR (Automatic Speech Recognition / Speech-to-Text Technology) task and TTS (Text-to-Speech) task of the voice assistant are also the results obtained through the subjective observation of the testers. It is very easy to overlook many aspects such as the quantification of the response time of the voice assistant, false wake-up and wake-up, continuous conversation ability, performance under noise, accent adaptability, edge processing, and stability.

[0067] It can be seen that the traditional voice assistant verification method has at least the following deficiencies: First, the verification efficiency is low. The process of verifying the voice assistant involves manual participation, which is likely to result in errors. Second, the verification cost is high. A large number of test tasks need to be executed, requiring a large amount of human resources. Third, the accuracy is difficult to guarantee. Different testers have their own subjective judgment factors when verifying the voice assistant, resulting in different verification results. Fourth, the generality is poor. When verifying the voice assistants of different devices, testers with different experiences need to participate.

[0068] In view of this, the present application provides a verification method for a voice assistant. The method includes: obtaining at least one synthesized voice with different timbres through a test client and playing the synthesized voice to the voice assistant to be tested; after receiving the synthesized voice, the voice assistant to be tested performs relevant processing, and the processing process information is recorded in the log associated with the voice assistant; the test client obtains the response process information of the voice assistant by obtaining the log information associated with the voice assistant; the test client can further compare the obtained response process information of the voice assistant with the corresponding expected response process information, and automatically generate the verification result of the voice assistant to be tested based on the comparison result. Compared with the traditional technology, the verification method of the present application: (1) solves the mismeasurement caused by human subjective factors in the verification process; (2) facilitates the stable test of the long-term conversation ability of the voice assistant and solves the problem that human resources cannot have a long-term conversation during the verification process; (3) can perform performance tests on the multi-timbre recognition of the voice assistant and solves the problem of insufficient human resources during the verification process; (4) can also verify the text-to-speech ability of the voice assistant and solves the problem that it is difficult for humans to verify the recognition process; (5) at the same time, the verification process of the voice assistant is fully automated, solving the problems of low efficiency and high labor cost; (6) moreover, the verification method of the voice assistant can be reused in various vehicle models, robots, and any project with voice tests, solving the problem of poor generality of the traditional technology.

[0069] The verification method for the voice assistant provided by the embodiments of the present application can be applied to a test system, which includes a test device and a device under test. The device under test is configured with a voice assistant, and voice interaction can be performed between the test device and the device under test.

[0070] Among them, as Figure 1 shown, the test device may include a test client 101 and a backend model service (also called the server) 102. The test client 101 communicates with the server 102 through the TCP / IP protocol. Among them, the client 101 mainly includes a sound collection device, a host computer, and a sound playback device. The host computer can be signal-connected to the sound playback device through a serial port, and the host computer can also be signal-connected to the sound collection device through a message middleware. Its main functions are voice cloning, multi-timbre synthesis, sound processing, device under test detection, text comparison, verification report output, intermediate file temporary storage, speech recognition, audio stream transmission, etc. The server 102 deploys a speech synthesis model, a text processing model, and a speech recognition model, and can provide speech synthesis and speech recognition during the test process, and is also used to provide functions such as text similarity comparison and message middleware.

[0071] Among them, the device under test is configured with a voice assistant. During the test, the voice assistant is in an on state. When it hears the voice played by the test device, the voice assistant of the device under test can perform processes such as speech recognition, text matching, and text-to-speech on the heard voice signal, and finally output a reply voice or a response voice. Specifically, the device under test can be a car or a robot, and the present application does not limit this.

[0072] In an exemplary embodiment, as Figure 2 shown, a verification method for the voice assistant function is provided. Taking the test client in Figure 1 as an example, it includes the following steps S201 to S205. Among them:

[0073] Step S201, obtain at least one synthesized voice, and each synthesized voice has a different timbre.

[0074] Among them, the synthesized voice can be understood as the voice synthesized based on the corpus and the selected timbre, which is different from the real voice when a person speaks.

[0075] In some embodiments, the test client can synthesize voices with multiple different timbres based on the speech synthesis model of the server. Among them, multiple different timbres can be extracted from the real voices of different people through voice cloning technology in advance.

[0076] In the embodiments of the present application, multiple synthesized voices with different timbres can be obtained to participate in the verification test of the voice assistant, so as to verify the recognition ability of the voice assistant under test for voices with different timbres and ensure the accuracy and reliability of the verification results.

[0077] Step S202: Play the synthesized voice to the voice assistant under test.

[0078] In the embodiments of the present application, the voice assistant under test is deployed in the corresponding device under test, and the device under test can be a new energy vehicle, a robot or other intelligent terminals. Based on this, playing the synthesized voice from the test client to the voice assistant under test can be understood as playing the synthesized voice from the test client to the device under test, and the voice assistant in the device under test is in an on state and can listen to external voices in real time and make responses.

[0079] In some embodiments, if there are multiple synthesized voices, the test client can sort the multiple synthesized voices and play them in the order of sorting. It can be understood that the verification method of the present application is implemented based on the voice interaction between the test client and the voice assistant. Therefore, the multiple synthesized voices can be played discontinuously. For example, after receiving the response voice from the voice assistant, the next synthesized voice can be played continuously.

[0080] Step S203: Obtain the log information of the voice assistant, and the log information includes the response process information of the voice assistant to the synthesized voice.

[0081] Among them, the voice assistant can generate log information in real time during operation. Through the log information, the ASR recognition process information of the voice assistant for the received voice is recorded, the process information of matching the corresponding corpus based on TTS for response is recorded, and the process information of generating a response voice based on the matched corpus and playing the response voice is also recorded. Therefore, the response process information of the voice assistant to the synthesized voice includes but is not limited to voice recognition results, intermediate processing data, and processing time information, etc. The processing time information can be the processing time of each process or the processing time of the entire response process, which can be selected based on the actual situation.

[0082] In some embodiments, the device where the test client is located can be connected to the device under test through a data cable to obtain the log information associated with the voice assistant in the device under test in real time; it can be understood that in other embodiments, the test device can also obtain the log information of the voice assistant when the current round of voice interaction ends. The present application does not limit the acquisition timing and acquisition method of the log information.

[0083] In addition, after the test client obtains the log information, it can be cached first, and after the current round of voice interaction ends, the log information can be analyzed to obtain the response process information of the voice assistant during the current round of testing.

[0084] Step S204, obtain the comparison result between the response process information of the voice assistant and the expected response process information corresponding to the synthesized speech.

[0085] On the one hand, the test client obtains the log information of the voice assistant and obtains the response process information in the current round of voice interaction of the voice assistant based on the log information. On the other hand, the test client can obtain the actual corpus corresponding to the synthesized speech based on the historical information corresponding to the synthesized speech, and can also obtain the expected response corpus and the expected response time corresponding to the actual corpus of each synthesized speech by querying the pre-configured information library. The actual corpus, the expected response corpus, and the expected response time corresponding to the synthesized speech, etc. are used as the expected response process information of the synthesized speech.

[0086] On this basis, the test client can compare the response process information of the voice assistant with the expected response process information of the synthesized speech, and then verify whether the processing process of the voice assistant at multiple process nodes meets the expectations. Correspondingly, the comparison result between the response process information of the voice assistant and the expected response process information corresponding to the synthesized speech includes: (1) The response processes of multiple process nodes of the voice assistant all meet the expectations; (2) The response processes of multiple process nodes of the voice assistant do not meet the expectations; (3) Among the response processes of multiple processing nodes of the voice assistant, some processes meet the expectations and some processes do not meet the expectations.

[0087] Step S205, generate the verification result of the voice assistant based on the comparison result.

[0088] Based on the comparison result between the response process information of the voice assistant and the expected response process information corresponding to the synthesized speech, the test client can obtain the final verification result according to the pre-configured verification rules. For example, when the response processes of multiple process nodes of the voice assistant all meet the expectations, it is determined that the verification result of the voice assistant passes. As long as the response process of any one process node does not meet the expectations, it is determined that the verification result of the voice assistant fails.

[0089] In some embodiments, after the test client obtains the verification result of the voice assistant, it can also generate a verification report for the tested voice assistant according to the preset report template, which is convenient for relevant personnel to understand the relevant information of the verification.

[0090] Based on the verification method of the above embodiments, the test client can obtain at least one synthesized speech with different timbres and play the synthesized speech to the voice assistant to be tested; after receiving the synthesized speech, the voice assistant to be tested performs relevant processing, and the processing process information is recorded in the log associated with the voice assistant; the test client obtains the response process information of the voice assistant to the synthesized speech by obtaining the log information associated with the voice assistant; the test client can further obtain the comparison result between the response process information of the voice assistant and the expected response process information corresponding to the synthesized speech, and generate the verification result of the voice assistant to be tested based on these comparison results. The entire verification process can be automatically executed according to the set program. At the same time, by testing the voice assistant to be tested with synthesized speeches of different timbres, the reliability of the verification result can be improved; in addition, the test device can also determine whether each processing link of the voice assistant meets the expectations according to the log information during the test process of the voice assistant and through information comparison, and automatically generate the corresponding verification result accordingly, avoiding the subjective deviation existing in the manual verification evaluation, ensuring the objectivity of the verification result, and also improving the verification efficiency and reducing the verification cost compared with the manual verification method.

[0091] In an exemplary embodiment, the response process information of the voice assistant to be tested pointed out in the foregoing embodiments may specifically include at least one of: the recognized corpus obtained by recognizing the synthesized speech, the target response corpus matched based on the recognized corpus, and the response time for the synthesized speech.

[0092] Among them, the recognized corpus obtained by the voice assistant to be tested by recognizing the synthesized speech can be understood as the corpus information corresponding to the synthesized speech obtained by the voice assistant to be tested after receiving the synthesized speech and performing ASR recognition on the synthesized speech.

[0093] Among them, the target response corpus matched for the recognized corpus of the synthesized speech can be understood as the response corpus corresponding to the text corpus information obtained by the voice assistant to be tested after performing ASR recognition on the synthesized speech and based on the pre-set information library, and after performing text-to-speech processing based on this response corpus, it can be used to reply to the received synthesized speech.

[0094] Among them, the response time of the voice assistant for the synthesized speech can be understood as the processing time information of each processing link involved in the process from the voice assistant receiving the synthesized speech to replying with the response speech or the processing time information of the entire process. For example, it can be the time information of the recognition process of the synthesized speech by the voice assistant, or the time information of the voice assistant matching the corresponding target response corpus after recognizing the corpus of the synthesized speech, or the time information from matching the corresponding target response corpus to playing the response speech, and it can also be the time information of the entire process from the voice assistant receiving the synthesized speech to playing the response speech.

[0095] Correspondingly, the expected response process information corresponding to the synthesized speech played by the test client may include at least one of: the actual corpus corresponding to the synthesized speech, the expected response corpus corresponding to the synthesized speech, and the expected response time corresponding to the synthesized speech.

[0096] Among them, for the actual corpus corresponding to the synthesized speech, the test client can obtain it based on the corpus information during the synthesis process of the synthesized speech without performing ASR recognition on the synthesized speech; the expected response corpus corresponding to the synthesized speech can be obtained based on the pre-configured information library corresponding to the voice assistant. This information library can store the common interaction corpora of this voice assistant. Based on the actual corpus corresponding to the synthesized speech, the response corpus of the voice assistant under normal circumstances can be matched therefrom as the expected response corpus corresponding to the synthesized speech.

[0097] Among them, the expected response time corresponding to the synthesized speech can be the expected processing time information of each processing link involved in the process from when the voice assistant hears the synthesized speech to when it replies with the response speech or the expected processing time information of the entire process. In some exemplary embodiments, the test client can estimate the speech recognition time of the voice assistant based on the number of characters of the actual corpus corresponding to the synthesized speech, and estimate the time taken for the voice assistant to output the response speech based on the number of characters of the expected response corpus and the preset speech rate.

[0098] Based on the above embodiments, the test client can obtain the information of multiple processing processes of the voice assistant to be tested. Correspondingly, the test client can also obtain the expected response information for multiple processing timeouts of the voice assistant. Based on the acquisition and comparison of the information of multiple processing processes, the performance of the voice assistant can be verified more comprehensively to ensure the accuracy of the finally generated verification result.

[0099] In some embodiments, the verification method of the present application may further include the steps of: obtaining the actual corpus corresponding to the synthesized speech based on the log information recorded during the speech synthesis process; matching the target text information corresponding to the actual corpus from the preset information library, using the target text information as the expected response corpus, and obtaining the expected response time of the synthesized speech according to the target text information and the preset speech rate.

[0100] Based on the method of this embodiment, the test client can obtain the actual corpus corresponding to the synthesized speech without performing ASR recognition on the synthesized speech, which is beneficial to saving system resources and improving verification efficiency; at the same time, it can also reduce the error in the ASR recognition process, thereby making the final verification result more accurate.

[0101] In some exemplary embodiments, the comparison result of obtaining the response process information of the voice assistant and the expected response process information corresponding to the synthesized speech as pointed out in the foregoing embodiments may include at least one of the following: obtaining a first comparison result between the recognized corpus of the voice assistant and the actual corpus corresponding to the synthesized speech; obtaining a second comparison result between the target response corpus of the voice assistant and the expected corpus corresponding to the synthesized speech; obtaining a third comparison result between the actual response time of the voice assistant and the expected response time.

[0102] Based on the first comparison result, the speech recognition performance of the voice assistant can be verified. Based on the second comparison result, the matching accuracy of the voice assistant to the text information library can be verified. Based on the third comparison result, the processing efficiency or response timeliness of the voice assistant can be verified. Thus, the performance of different processing processes of the voice assistant can be realized, or the voice assistant can be verified from different dimensions, making the verification result more comprehensive.

[0103] Combined with any of the above embodiments, in other embodiments, the verification method of the present application may further include the following steps: collecting an audio stream of the response speech played by the voice assistant for the synthesized speech; in response to the end of this round of voice interaction, obtaining the actual response corpus of the voice assistant based on the speech recognition result of the audio stream of the response speech; and obtaining a fourth comparison result between the actual response corpus of the voice assistant and the target response corpus; wherein, the target response corpus is the response corpus matched by the voice assistant from the corresponding text information library for the recognized corpus of the synthesized speech.

[0104] In some embodiments, the test client can obtain the speech recognition result of the audio stream of the response speech of the voice assistant from the server. With the help of the speech recognition module of the server, while ensuring the accuracy of the recognition result, the burden on the test client can be reduced. And in this way, since the server and the test site can be physically isolated, the server is less affected by external environmental sounds when performing speech recognition on the response speech of the voice assistant, which is beneficial to improving the accuracy of speech recognition.

[0105] Thus, by further collecting the audio stream of the response speech actually replied by the voice device to be tested, performing speech recognition on the audio stream to obtain the actual corpus (i.e., the actual response corpus) actually replied by the voice assistant, and then the test client comparing the actual response corpus with the target response corpus matched by the voice assistant from the information library, it can be verified whether there is an abnormality in the process of the voice assistant synthesizing the response speech based on the target response corpus, further improving the comprehensiveness of the verification of the voice assistant.

[0106] Combined with the above embodiments, in some embodiments, the specific manner in which the test client in this application generates the verification result of the voice assistant based on the comparison result may include: The test client generates the verification result of the voice assistant to be tested based on at least one of the foregoing first comparison result, second comparison result, third comparison result, and fourth comparison result.

[0107] As a specific example, if any one of the foregoing first comparison result, second comparison result, third comparison result, and fourth comparison result is not passed, the test client generates a verification failure report for the voice assistant; if all comparison results are passed, the test client generates a verification pass report for the voice assistant.

[0108] It can be understood that the test client can also generate the verification result of the voice assistant based on other rules according to the actual situation. For example, if the performance impact of a certain link on the voice assistant is small, then when the comparison result of the information corresponding to this link is not passed and the comparison results of the information corresponding to other links are passed, a verification pass report for the voice assistant can still be generated at this time.

[0109] Based on this, the test client automatically generates the verification result of the voice assistant based on the comparison results of information in different links. This process is not affected by human subjective biases, and the generation efficiency of the verification result is not affected by human factors, taking into account both the objectivity and generation efficiency of the verification result.

[0110] In some embodiments, the acquisition methods of the foregoing first comparison result, second comparison result, and fourth comparison result may include:

[0111] For two corpora to be compared, obtain the edit distance between the two corpora; wherein, the two corpora to be compared can be the recognition corpus of the voice assistant and the actual corpus corresponding to the synthesized speech, or the target response corpus of the voice assistant and the expected corpus corresponding to the synthesized speech, or the actual response corpus of the voice assistant and the target response corpus. Among them, the edit distance between two corpora can be understood as a measurement method for measuring the similarity between two strings, representing the minimum number of deletion, insertion, and replacement operations required to transform one corpus string into another corpus string. For example, the edit distance between "test" and "tent" is 1, because only "s" needs to be replaced with "n"; when the two corpora are the same, the edit distance between them is 0.

[0112] In response to the edit distance between the two corpora being greater than a preset similarity threshold, determine that the comparison result of the two corpora to be compared is passed; otherwise, determine that the comparison result of the two corpora is not passed. Among them, the preset similarity threshold can be determined based on the preset error propagation probability and string length.

[0113] As a specific example, the Levenshtein distance core algorithm can be used to calculate the edit distance between two corpora to be compared. For example, use the two-dimensional matrix D[i][j] to represent the distance between the first i characters of corpus A and the first j characters of corpus B, where D[i][j] represents the minimum number of insertion, deletion, and replacement steps required to convert corpus A into the string of corpus B. If D[i][j] is 0, the strings of the two corpora are the same.

[0114] In another specific example, the similarity calculation method for two corpora to be compared can be as follows:

[0115]

[0116] Where max(m, n) represents the larger value between the string length m of corpus A and the string length n of corpus B. The closer the similarity_ratio of the two corpora is to 1, the higher the similarity between the two corpora.

[0117] In yet another specific example, the error propagation probability model can also be referred to, and the acceptable difference range can be adjusted according to the string lengths of the two corpora to be compared. Assume that the recognition error probability of each character is p (allowing k errors), and it is estimated based on the ASR model test data:

[0118]

[0119] In the above formula, the cumulative probability of <=k error characters appearing in n characters is calculated based on the binomial distribution. Its main function is to solve the maximum number of allowable error characters k when the text length is n with a confidence level of 99%. For example: when the number of characters in "Please open the window" is 5, the preset confidence level is 99%, and correspondingly p = 0.01. Based on the above formula, k can be solved to be 0 errors, and the specific constraint is the above formula P(k) >= confidence level 99%.

[0120] Based on the k solved from the above formula, the final similarity threshold can be set as:

[0121]

[0122] Where n represents the length of the sentence / string, and k + 1 can be understood as reserving a safety margin of 1 character. For different corpus lengths, different k values may be solved, and then different similarity thresholds can be determined. Finally, the comparison result between the two corpora to be compared is obtained by comparing the Levenshtein distance with the similarity threshold. As an example, the similarity threshold can be as shown in Table 1 below.

[0123] Table 1

[0124]

[0125] Based on this, according to the string lengths and edit distances of the corpora to be compared, the comparison results of the two corpora can be obtained. The greater the edit distance between the corpora and the smaller the string length of the corpora, the greater the probability that the comparison of the two corpora fails. The smaller the edit distance between the corpora and the greater the string length of the corpora, the greater the probability that the comparison between the two corpora passes. Based on the corpus comparison results, the test client can more accurately verify the speech recognition performance or text-to-speech performance of the voice assistant, avoid being too strict or too lenient in verification, and ensure the referenceability of the verification results.

[0126] Combined with the foregoing embodiments, the method for obtaining the third comparison result may include: obtaining the time difference between the actual response time and the expected response time of the voice assistant; in response to the time difference being less than or equal to a preset timeout threshold, determining that the third comparison result is passed; otherwise, determining that the third comparison result is not passed.

[0127] Wherein, the timeout threshold can be set in advance and represents the tolerance for the response delay of the voice assistant. For a voice assistant with high requirements for response real-time performance, the timeout threshold can be set smaller, and for a voice assistant with low requirements for response real-time performance, the timeout threshold can be set relatively larger.

[0128] In some embodiments, as described in the foregoing embodiments, the actual response time and the expected response time of the voice assistant can both be the times of multiple processing procedures. Therefore, in the embodiments of the present application, the test client can compare the actual response time and the expected response time corresponding to different processing procedures in the manner of the foregoing embodiments to obtain multiple third comparison results. On this basis, as an optional implementation manner, the test client can finally determine that the third comparison result is passed when all the multiple third comparison results are passed.

[0129] Based on this, determining whether the response timeliness of the voice assistant passes the verification through the timeout threshold has higher accuracy than human perception.

[0130] In some embodiments, the method for the test client to obtain at least one synthesized voice may include: obtaining a corpus list to be synthesized and a timbre list to be synthesized; the corpus list contains at least one text corpus, and the timbre list contains at least one timbre information; inputting the corpus list and the timbre list into a preset speech synthesis model, and generating synthesized voices with different timbres corresponding to each text corpus in the corpus list based on the speech synthesis model.

[0131] For the voice assistant to be tested, a corpus list can be configured in advance on the test client, and a voice color list to be synthesized can be prepared in advance. Through the interaction between the test client and the backend server, synthesized voices of multiple voice colors can be synthesized, which occupies less resources of the test client. At the same time, the voice assistant to be tested can be more comprehensively verified to verify its response ability to different voices.

[0132] Combining some or all of the above embodiments, continue to refer to Figure 3 , the overall process of the verification method of the present application includes:

[0133] S1, construct a voice verification system, and this system can refer to the system shown in Figure 1 .

[0134] S2, turn on the voice broadcast and reception, including turning on devices such as an external artificial mouth, a sound card, and a microphone.

[0135] S3, start the detection of the test device, including whether the voice broadcast and reception devices are turned on. The test host computer screens and outputs the log system of the device to be tested. For example, through a regular expression, for example, match the text content in the middle that starts with "target=" and ends with ".", and output this text content.

[0136] S4, remotely generate voices. The test client remotely requests the server to synthesize voices, and the test client plays the voices.

[0137] Through interaction with the server, synthesized voices of multiple voice colors are synthesized, and the test client plays the synthesized voices through an artificial mouth.

[0138] During the process of the test client playing the voices, the test client can collect the currently played audio stream based on the external microphone device. After that, the test client can also collect the audio stream of the response voice of the device to be tested based on the microphone device and send it to the server, and obtain the ASR result of the response voice of the device to be tested by the server. Subsequently, the ASR result can be analyzed and compared to verify the relevant performance of the voice assistant of the device to be tested.

[0139] S5, the test client conducts ASR verification and TTS verification.

[0140] The test client obtains the processing process information of the voice assistant based on the log information associated with the voice assistant of the device to be tested. Based on this processing process information, verify whether the voice recognition ASR processing process of the voice assistant is normal (that is, compare the recognized corpus of the voice assistant and the actual corpus of the synthesized voice), verify whether the corpus matching process of the voice assistant is normal (that is, compare the actual response corpus of the voice assistant and the expected response corpus of the synthesized voice), and verify whether the TTS processing process of the voice assistant is normal (that is, compare the target response corpus of the voice assistant and the actual response corpus).

[0141] S6, Report result: Based on the above verification results, generate a test report for the entire verification process.

[0142] Based on the above overall process, the following will refer to Figure 4 , and give an exemplary description of the specific process of the voice assistant verification method of this application.

[0143] S41, Start the service of the server.

[0144] The server is deployed with a trained text-to-speech model, a trained voice model, the Levenshtein core algorithm, and an asynchronous message library (such as lightweight and high-performance ZeroMQ). The service of the server is used to provide functions such as speech synthesis, text similarity comparison, and message middleware during the test process.

[0145] S42, Start the test client and associated hardware, and start the device under test and the voice assistant, that is, ensure that the test device, the device under test, and the voice-related modules function properly.

[0146] S43, Input corpus.

[0147] Start the test execution program of the test client and drive the data according to the corpus.

[0148] S44, Speech synthesis with voice.

[0149] Synthesize voices of various tones through the speech synthesis model of the server.

[0150] S45, Device broadcasting.

[0151] The test client calls an external device to conduct a conversation with the device under test. Optionally, the external receiving device of the client can collect the audio stream of the synthesized voice played to verify the synthesis effect and ensure the quality of the synthesized voice used for testing.

[0152] S46, Match ASR logs and TTS logs with regular expressions.

[0153] During the process of the device under test recognizing the synthesized voice, it can match the ASR logs based on regular expressions, and during the process of generating the response voice, it can match the TTS logs based on regular expressions.

[0154] Therefore, by screening the log information of the device under test to obtain the ASR logs and TTS logs, the actual ASR information and TTS information can be obtained, such as the recognized corpus of the synthesized voice, the target corpus corresponding to the recognized corpus, the actual response time, etc. The obtained log information is stored in a queue for final analysis, and the relevant processing process information of the voice assistant of the device under test can be obtained by analyzing the logs.

[0155] S47, Audio ASR Verification.

[0156] Obtain the audio stream of the response speech replied by the device under test through a radio, process it into data with a specific sampling rate through streaming audio, and then generate the text of this speech through the asr_by_bytes method.

[0157] S48, Result Upload Message Middleware ZeroMQ (Asynchronous Message Library).

[0158] ZeroMQ deployed on the server is responsible for saving the audio ASR verification results of the server. The separation of the server and the client avoids the problem of mutual interference between receiving and playing sounds.

[0159] S49, Verification Results.

[0160] The expected value corresponding to the synthesized speech played can be stored by the test client. After the test execution is completed, collect the log information of the device under test, and obtain the actual ASR values of each stage of the device under test from ZeroMQ on the server. The test client compares the actual returned information of the speech assistant with the corresponding expected value through time difference, using direct character comparison or the Levenshtein core algorithm to obtain the comparison result. If all comparison results are True, this test passes, and the test client generates a detailed process report. If any one of the comparison results is False, the test client generates the reason for the error and provides information such as the corresponding log location and time for developers to troubleshoot.

[0161] Based on the verification method of the above embodiments, compared with the traditional technology, the technical problems that can be solved / technical effects achieved by the solution of the present disclosure include: (1) accurately detecting the response time of the speech assistant, solving the problem of false measurement caused by human subjective factors during the verification process; (2) stably testing the long-term conversation ability of the speech assistant, solving the problem that human resources cannot have long-term conversations during the verification process; (3) testing the performance of the speech assistant's recognition of multiple voices, solving the problem of insufficient human resources during the verification process; (4) verifying the model ability of the speech assistant's text-to-speech conversion, solving the problem that cannot be verified during the verification process; (5) fully automating the speech assistant verification process, solving the problems of serious low efficiency and high labor costs; (6) the speech assistant verification method can be reused in various vehicle models, robots, and any project with speech tests, solving the problem of poor generality.

[0162] Based on the above overall process and specific process, the following refers to Figure 5 , and further exemplarily illustrates the specific process of the speech assistant verification method of the present application.

[0163] S501. On the one hand, the test client starts the voice recording service thread and registers the audio stream processing service. On the other hand, it starts the voice service and the message middleware service of the server.

[0164] In addition, ensure that the voice assistant of the device under test is in the on state.

[0165] S502. The test client starts the log monitoring thread and registers the log stream processing service.

[0166] S503. The test client inputs a corpus list to the voice service of the server and obtains the synthesized voice synthesized by the voice service. The voice service of the server can synthesize synthesized voices of various timbres.

[0167] After that, the test execution thread of the test client can be started, and data driving is performed based on the synthesized voice.

[0168] S504. The test client triggers the log monitoring thread and starts collecting the log information of the device under test.

[0169] S505. The test client plays the synthesized voice through the artificial mouth.

[0170] S506. The audio stream of the played synthesized voice is synchronously collected through the audio collection thread and the voice recording device.

[0171] In some embodiments, the collected audio stream can be synchronously sent to the server for speech recognition processing in real time or at regular intervals. In other embodiments, after the audio stream is collected, it is first cached locally on the client, and after the current round of voice interaction ends, the cached audio stream is sent to the server for speech recognition.

[0172] S507. In the test scenario of continuously responding to the voice assistant of the device under test, the test client continues to play other synthesized voices during the test through the artificial mouth.

[0173] During this process, the audio stream of other synthesized voices played during the test and the audio stream of the response voice replied by the voice assistant of the device under test are collected through the audio collection thread and the voice recording device.

[0174] S508. After this round of voice interaction is completed, the audio collection thread and the voice recording device stop collecting the audio stream to avoid collecting the audio stream of other ambient voices and affecting the subsequent verification results.

[0175] S509. After this round of voice interaction is completed, the log monitoring thread stops collecting the log information of the device under test.

[0176] S510. After this round of voice interaction is completed, the test client sends the audio stream collected during this round of test to the server.

[0177] S511. The voice service of the server performs speech recognition on the audio stream.

[0178] Optionally, the test client may send the audio stream of the synthesized speech it plays and the audio stream of the response speech of the device under test to the speech model of the server. In this way, the speech model can use the synthesized speech as context to recognize the response speech of the device under test, improving the accuracy of speech recognition.

[0179] In some embodiments, the voice service of the server may also upload the recognition results in segments to the message middleware service for storage, to be retrieved by the test client for analysis later.

[0180] S512. In some embodiments, the test client may also save the audio stream collected during this round of testing as an intermediate test file and store it in the storage service of the server as the source data for the verification process.

[0181] S513. After this round of voice interaction is completed, the log monitoring thread starts to analyze the log information of the device under test that has been monitored.

[0182] During the test, the device under test may write the processing process information of the voice assistant into the ASR log or TTS log. The log monitoring thread collects the full amount of log information of the device under test and filters out the ASR log or TTS log among them, then the relevant processing process information of the voice assistant of the device under test can be obtained. The specific content of the processing process information can refer to the description of the foregoing embodiments and will not be elaborated here.

[0183] It can be understood that the execution timing of steps S510 - S513 is not limited to the illustration shown. They can be executed synchronously or in other orders.

[0184] S514. In some embodiments, the log analysis results obtained by the log monitoring thread analysis can be saved locally on the test client. The original log content monitored by the log monitoring thread can also be saved to the intermediate test file of the server as the source data for the test process.

[0185] S515. The test client requests the message middleware service of the server to obtain the speech recognition results during the test process.

[0186] S516. The test client combines the log analysis results of the device under test and the speech recognition results during the test process, analyzes the test results of the voice assistant of the device under test for this round, generates a corresponding verification report, and displays the report.

[0187] Continue to refer to the following Figure 6, introduce the log monitoring thread of the test client involved in the verification process.

[0188] (1) Initialize the log detection object as the device under test.

[0189] The adb -s {{device under test}} shell "logcat |grep{{filter}}" subprocess can be started through a python thread to keep the process active, block and obtain logs in the Iobuffer, and distribute the logs.

[0190] (2) Collect filtered logs and distribute them to the test client

[0191] Redirect all the logs of the device under test into the file content in real time, and add a notifier to distribute them to the receiving queue of the test client. If the test is not being executed at this time, no distribution will be made.

[0192] Continue to refer to Figure 6 , describe the processing flow of the test execution thread (VoiceServerThread) of the test client involved in the verification process:

[0193] (1) Synthesize the corpus.

[0194] The server can be requested to synthesize voice by inputting the corpus list. Before starting to play the synthesized voice, control the server to turn on the sound collection. After playing the voice, control to turn off the sound collection. In the next round of voice broadcast and sound collection interaction, this thread controls the server to turn on and off the sound collection again, and also obtains the log information of the voice assistant of the device under test into the cache during the process.

[0195] (2) Judge the response time.

[0196] Judge the response time of the voice assistant of the device under test to each synthesized voice through the expected response corpus and the expected voice playback speed, and wait for the device under test to fully reply.

[0197] Continue to refer to Figure 6 , describe the audio collection program involved in the verification process.

[0198] Deploy the sound collection control program through fastapi for the client to start and stop sound collection.

[0199] (3) Continuously collect audio and video data streams.

[0200] Obtain the data stream of the sound collection hardware, write get_voice_tobytes into the memory. When the sound collection ends, the data stream stops being collected. Send the data stream in the memory to the voice recognition model of the server through the fastapi interface asr_by_bytes to obtain the recognized corpus.

[0201] (4) Organize the data structure and serialize the processing flow.

[0202] Each collected audio stream has a unique identifier id. The corpus returned by the speech recognition service on the server side forms a dictionary, such as: {{id: recognized content}, {id2: recognized content}}. Finally, it is serialized into the json format and uploaded to the ZeroMQ service on the server side, waiting for the test client to obtain.

[0203] The analysis of the test client includes:

[0204] (1) When analyzing the process log time, first, the time to play the synthesized voice should be earlier than the ASR log time filtered, that is, the filtered ASR log time should be later than the time to play the synthesized voice.

[0205] (2) Each voice ASR is correct.

[0206] Obtain the ASR text of each id audio through the ZeroMQ service, compare it with the expected value. If they are the same or the similarity meets the conditions, then this test passes.

[0207] (3) Report generation.

[0208] The test client can fill in the test case results, the time to play the synthesized voice, the log filtering result, the actual test process log, the ASR recognition result, and the Levenshtein distance value in an excel table to provide a voucher for the test process information.

[0209] Through the verification method disclosed in this application, the verification items and test cases for the voice assistant may include the content shown in Table 2 below. For different verification items, synthetic voices can be generated for prediction specifically, including generating synthetic voices based on different corpora, generating different synthetic voices based on different voices, dialects or accents, etc. The test processes for different verification items can all refer to the description of the foregoing embodiments.

[0210] Table 2

[0211]

[0212] Compared with the existing technology, the technical effects of the technical solution of the above-mentioned embodiment of the present application include at least: (1) Improving the accuracy of voice verification, using a model combined with a program to control the reception and broadcasting and audio synthesis, reducing the accuracy problem of manual verification. (2) Reducing the detection cost, forming a set of program automation control combined with AI technology voice process verification method, which is simple to operate. You only need to operate the verification system / program to complete the voice verification process and generate a report, which greatly reduces the human resource cost. (3) Improving versatility, through program control of voice, for the test equipment, only log collection is required, and AI technology is run to train various tones in advance, thereby improving verification efficiency and coverage, and the entire process can be used in various test equipment, including but not limited to various vehicle computer projects or robot projects.

[0213] It should be understood that, although the steps in the flowcharts involved in the above embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0214] Based on the same inventive concept, the embodiment of the present application also provides a voice assistant function verification device for implementing the voice assistant function verification method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in the embodiments of the verification device for one or more voice assistant functions provided below can be referred to the limitations of the voice assistant function verification method above, and will not be repeated here.

[0215] In an exemplary embodiment, Figure 7 As shown, a voice-assisted verification device is provided, which is applied to a test client, and the device includes:

[0216] The synthesized speech acquisition module 701 is used to acquire at least one synthesized speech, each synthesized speech has a different timbre;

[0217] A broadcast control module 702, used to play the synthesized voice to the voice assistant under test;

[0218] The log information acquisition module 703 is used to acquire the log information of the voice assistant, wherein the log information includes the response process information of the voice assistant to the synthesized speech;

[0219] A comparison module 704, configured to obtain a comparison result between the response process information of the voice assistant and the expected response process information corresponding to the synthesized speech;

[0220] A result generation module 705, configured to generate a verification result of the voice assistant based on the comparison result.

[0221] Based on the above voice assistant verification device, the test client obtains at least one synthesized speech with different timbres through the synthesized speech acquisition module, and plays the synthesized speech to the voice assistant to be tested through the broadcast control module; after receiving the synthesized speech, the voice assistant to be tested performs relevant processing, and the processing process information is recorded in the log associated with the voice assistant; the test client obtains the log information associated with the voice assistant through the log information acquisition module to obtain the response process information of the voice assistant to the synthesized speech; the test client can further obtain the comparison result between the response process information of the voice assistant and the expected response process information corresponding to the synthesized speech based on the comparison module, and finally the result generation module can generate the verification result of the voice assistant to be tested based on these comparison results. In the whole verification process, each functional module can be automatically executed according to the set program without human participation. At the same time, the voice assistant to be tested is tested with synthesized speeches of different timbres to improve the reliability of the verification result. In addition, the verification process is objective and can automatically generate the corresponding verification result, avoiding the subjective deviation existing in the artificial verification evaluation. At the same time, compared with the artificial verification method, the verification efficiency is improved and the verification cost is reduced.

[0222] Each module in the above voice assistant verification device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.

[0223] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 8As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it realizes a method for verifying a voice assistant. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0224] Those skilled in the art can understand that Figure 8 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0225] Those skilled in the art can understand that Figure 8 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0226] In an exemplary embodiment, a computer device is provided, and its structural diagram can be as Figure 8As shown, the computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it is used to: obtain at least one synthesized voice, each synthesized voice having a different timbre; play the synthesized voice to the voice assistant to be tested; obtain the log information of the voice assistant, where the log information includes the response process information of the voice assistant to the synthesized voice; obtain the comparison result between the response process information of the voice assistant and the expected response process information corresponding to the synthesized voice; generate a verification result of the voice assistant based on the comparison result.

[0227] In addition, when the processor executes the computer program, it can also be used to implement the steps in the above-mentioned other method embodiments.

[0228] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it realizes: obtaining at least one synthesized voice, each synthesized voice having a different timbre; playing the synthesized voice to the voice assistant to be tested; obtaining the log information of the voice assistant, where the log information includes the response process information of the voice assistant to the synthesized voice; obtaining the comparison result between the response process information of the voice assistant and the expected response process information corresponding to the synthesized voice; generating a verification result of the voice assistant based on the comparison result;

[0229] In addition, when the computer program is executed by a processor, it can also implement the steps in the above-mentioned other method embodiments.

[0230] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, it realizes: obtaining at least one synthesized voice, each synthesized voice having a different timbre; playing the synthesized voice to the voice assistant to be tested; obtaining the log information of the voice assistant, where the log information includes the response process information of the voice assistant to the synthesized voice; obtaining the comparison result between the response process information of the voice assistant and the expected response process information corresponding to the synthesized voice; generating a verification result of the voice assistant based on the comparison result;

[0231] In addition, when the computer program is executed by a processor, it can also implement the steps in the above-mentioned other method embodiments.

[0232] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0233] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0234] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0235] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method for verifying a voice assistant, characterized in that Applied to a test client, the method includes: Obtain at least one synthesized speech, each synthesized speech having a different timbre; Play the synthesized speech to the speech assistant to be tested; Obtain the log information of the speech assistant, where the log information includes the response process information of the speech assistant to the synthesized speech; Obtain the comparison result between the response process information of the speech assistant and the expected response process information corresponding to the synthesized speech; Generate a verification result of the speech assistant based on the comparison result.

2. The method according to claim 1, wherein The response process information of the speech assistant includes at least one of: the recognized corpus obtained by recognizing the synthesized speech, the target response corpus matched based on the recognized corpus, and the actual response time for the synthesized speech; The expected response process information corresponding to the synthesized speech includes at least one of: the actual corpus corresponding to the synthesized speech, the expected response corpus corresponding to the synthesized speech, and the expected response time corresponding to the synthesized speech.

3. The method according to claim 2, wherein The method further includes: Based on the log information recorded during the speech synthesis process, obtain the actual corpus corresponding to the synthesized speech; Obtain the target text information corresponding to the actual corpus from a preset information library, use the target text information as the expected response corpus, and based on the target text information and a preset speech rate, obtain the expected response time of the synthesized speech; wherein, the preset information library is the information library corresponding to the speech assistant to be tested.

4. The method according to claim 2, wherein Obtaining the comparison result between the response process information of the speech assistant and the expected response process information corresponding to the synthesized speech includes at least one of the following: Obtain the first comparison result between the recognized corpus of the speech assistant and the actual corpus corresponding to the synthesized speech; Obtain the second comparison result between the target response corpus of the speech assistant and the expected corpus corresponding to the synthesized speech; Obtain the third comparison result between the actual response time of the speech assistant and the expected response time.

5. The method according to claim 4, characterized in that, The method further includes: Collect the audio stream of the response speech replied by the speech assistant to the synthesized speech; In response to the end of this round of voice interaction, based on the speech recognition result of the audio stream of the response speech, obtain the actual response corpus of the speech assistant; Obtain the fourth comparison result between the actual response corpus of the speech assistant and the target response corpus matched by the speech assistant; Generating the verification result of the speech assistant based on the comparison result includes: Generate the verification result of the speech assistant based on at least one of the first comparison result, the second comparison result, the third comparison result, and the fourth comparison result.

6. The method according to claim 5, wherein Generating the verification result of the speech assistant based on at least one of the first comparison result, the second comparison result, the third comparison result, and the fourth comparison result includes: In response to any one of the comparison results being not passed, generate a verification failure report of the speech assistant; In response to all comparison results being passed, generate a verification passed report of the speech assistant.

7. The method according to claim 5, wherein The obtaining methods of the first comparison result, the second comparison result, and the fourth comparison result include: For two corpora to be compared, determine the edit distance between the two corpora; In response to the edit distance between the two corpora being greater than a preset similarity threshold, determine that the comparison result of the two corpora is passed; otherwise, determine that the comparison result of the two corpora is not passed; Wherein, the similarity threshold is determined based on a preset error propagation probability and the length of the propagated string.

8. The method according to claim 5, wherein The obtaining method of the third comparison result includes: Obtain the time difference between the actual response time and the expected response time of the voice assistant; In response to the time difference being less than or equal to a preset timeout threshold, determine that the third comparison result is passed; otherwise, determine that the third comparison result is not passed.

9. The method according to any one of claims 1 to 8, characterized in that, The obtaining of at least one synthesized voice includes: Obtain a corpus list to be synthesized and a timbre list to be synthesized; the corpus list contains at least one text corpus, and the timbre list contains at least one timbre information; Send the corpus list and the timbre list to the server, and generate synthesized voices with different timbres corresponding to each text corpus in the corpus list based on a voice synthesis model preset on the server.

10. A voice assistant verification device, characterized in that, Applied to a test client, the device includes: A synthesized voice acquisition module, configured to acquire at least one synthesized voice, and each synthesized voice has a different timbre; A broadcast control module, configured to play the synthesized voice to the voice assistant to be tested; A log information acquisition module, configured to acquire the log information of the voice assistant, and the log information contains the response process information of the voice assistant to the synthesized voice; A comparison module, configured to obtain the comparison result between the response process information of the voice assistant and the expected response process information corresponding to the synthesized voice; A result generation module, configured to generate a verification result of the voice assistant based on the comparison result.

11. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Service testing method and device based on artificial intelligence

    CN112037763A

  • Performance test method, device and equipment of voice interaction equipment and readable medium

    CN114999454A

  • Response capability testing method and device, electronic equipment and medium

    CN116705068A

  • Test system

    CN118038854A

  • Speech recognition test method and device, computer equipment and speech recognition test system

    CN119170016A