Voiceprint recognition-based testing methods, systems, storage media, and electronic devices

CN116312624BActive Publication Date: 2025-12-02HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD +2

Patent Information

Application Number
CN202211097967.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-08
Publication Date
2025-12-02
Estimated Expiration
2042-09-08

AI Technical Summary

Technical Problem

[0005]本申请提供一种基于声纹识别的测试方法、系统、存储介质及电子装置,用以解决现有技术中人工测试效率低下,错误率较高的缺陷,通过本申请不仅能够大大的提升声纹识别应用效果的测试效率,而且能够自动生成测试报告,便于直观地分析出应用过程中的错误问题,从而大大地提高识别应用的准确性

Benefits of technology

[0020]本申请提供的一种基于声纹识别的测试方法、系统、存储介质及电子装置,通过获取预设的音频文件,其中,所述音频文件包括目标音频数据及对应的执行预期值;将所述目标音频数据发送至声纹识别接口,并接收经由所述声纹识别接口返回的声纹特征和用户意图,其中,所述声纹特征和所述用户意图为所述声纹识别接口调用对应的识别算法对所述目标音频数据解析生成;将所述声纹特征和所述用户意图发送至设备执行接口,生成设备执行命令;其中,所述设备执行命令为所述设备执行接口根据所述声纹特征和所述用户意图调用命令生成服务进行处理生成;将所述设备执行命令发送至对应的虚拟设备,接收所述虚拟设备根据所述设备执行命令返回的执行结果;根据所述声纹特征和所述执行结果分别与所述执行预期值进行对比,生成测试报告。不仅能够大大的提升声纹识别应用效果的测试效率,而且能够自动生成测试报告,便于直观地分析出应用过程中的错误问题,从而大大地提高识别应用的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312624B_ABST
    Figure CN116312624B_ABST
Patent Text Reader

Abstract

This application discloses a testing method, system, storage medium, and electronic device based on voiceprint recognition, relating to the field of voiceprint recognition technology. The testing method includes: acquiring a preset audio file; sending target audio data from the audio file to a voiceprint recognition interface and receiving voiceprint features and user intent returned via the voiceprint recognition interface; sending the voiceprint features and user intent to a device execution interface to generate a device execution command; sending the device execution command to a corresponding virtual device and receiving the execution result returned by the virtual device based on the device execution command; and comparing the voiceprint features and execution result with the expected execution value to generate a test report. The embodiments provided by this invention not only greatly improve the testing efficiency of voiceprint recognition applications but also automatically generate test reports, facilitating intuitive analysis of errors during the application process, thereby significantly improving the accuracy of the recognition application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voiceprint recognition technology, and in particular to a test method, system, storage medium and electronic device based on voiceprint recognition. Background Technology

[0002] With the development of voiceprint recognition technology, its application scenarios are increasing and the requirements for related technical indicators are becoming more and more stringent.

[0003] Voiceprint recognition has a wide range of applications, mainly because each person's voice characteristics play a role. When we apply voiceprint recognition technology to smart home appliances, we can correctly identify the voiceprint characteristics of the current user, such as the voiceprint characteristics of the elderly, children, adults, men, and women. Improving the voiceprint recognition experience has always been a key goal for the implementation of voiceprint applications.

[0004] Currently, the application effect of voiceprint recognition technology is usually tested manually. This means verifying the implementation of the voiceprint function through the actual interaction process of the voiceprint recognition application device. Although this verification and testing method is feasible, it can only verify a limited amount of data. If it is necessary to verify a large amount of different voiceprint data, or to reproduce incorrect voiceprint judgments, the manual testing method is unstable and is limited by factors such as environment, distance, equipment, and the current speaker. It not only has a relatively high error rate, but is also inefficient. Summary of the Invention

[0005] This application provides a testing method, system, storage medium, and electronic device based on voiceprint recognition to address the shortcomings of low efficiency and high error rate in manual testing in the prior art. This application can not only greatly improve the testing efficiency of voiceprint recognition application effects, but also automatically generate test reports, which facilitates intuitive analysis of errors in the application process, thereby greatly improving the accuracy of recognition applications.

[0006] In a first aspect, this application provides a testing method based on voiceprint recognition, comprising:

[0007] Obtain a preset audio file, wherein the audio file includes target audio data and corresponding expected execution values;

[0008] The target audio data is sent to the voiceprint recognition interface, and the voiceprint features and user intent returned by the voiceprint recognition interface are received. The voiceprint features and user intent are generated by the voiceprint recognition interface calling the corresponding recognition algorithm to parse the target audio data.

[0009] The voiceprint features and the user intent are sent to the device execution interface to generate a device execution command; wherein, the device execution command is generated by the device execution interface calling a command generation service based on the voiceprint features and the user intent.

[0010] Send the device execution command to the corresponding virtual device, and receive the execution result returned by the virtual device based on the device execution command;

[0011] A test report is generated by comparing the voiceprint features and the execution results with the expected execution values.

[0012] Secondly, this application also provides a testing system based on voiceprint recognition, comprising:

[0013] The acquisition module is used to acquire a preset audio file, wherein the audio file includes target audio data and corresponding expected execution values;

[0014] The recognition module is used to send the target audio data to the voiceprint recognition interface and receive the voiceprint features and user intent returned by the voiceprint recognition interface, wherein the voiceprint features and user intent are generated by the voiceprint recognition interface calling the corresponding recognition algorithm to parse the target audio data;

[0015] A device execution command generation module is used to send the voiceprint features and the user intent to the device execution interface to generate a device execution command; wherein, the device execution command is generated by the device execution interface calling a command generation service based on the voiceprint features and the user intent;

[0016] The execution result receiving module is used to send the device execution command to the corresponding virtual device and receive the execution result returned by the virtual device based on the device execution command;

[0017] The test report generation module is used to compare the voiceprint features and the execution results with the expected execution values ​​to generate a test report.

[0018] Thirdly, this application also provides a computer-readable storage medium comprising a stored program, wherein the program, when executed, implements a voiceprint recognition-based testing method as described in any of the above embodiments.

[0019] Fourthly, this application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute, through the computer program, a test method based on voiceprint recognition as described above.

[0020] This application provides a testing method, system, storage medium, and electronic device based on voiceprint recognition. The method involves acquiring a preset audio file, including target audio data and a corresponding expected execution value; sending the target audio data to a voiceprint recognition interface and receiving voiceprint features and user intent returned by the interface, wherein the voiceprint features and user intent are generated by the voiceprint recognition interface calling a corresponding recognition algorithm to parse the target audio data; sending the voiceprint features and user intent to a device execution interface to generate a device execution command; wherein the device execution command is generated by the device execution interface calling a command generation service based on the voiceprint features and user intent; sending the device execution command to a corresponding virtual device and receiving the execution result returned by the virtual device; and comparing the voiceprint features and execution result with the expected execution value to generate a test report. This method not only significantly improves the testing efficiency of voiceprint recognition applications but also automatically generates test reports, facilitating intuitive analysis of errors during the application process and thus greatly improving the accuracy of the recognition application. Attached Figure Description

[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of the hardware environment for a test method based on voiceprint recognition provided in this application;

[0024] Figure 2 This is a flowchart illustrating a testing method based on voiceprint recognition provided in this application;

[0025] Figure 3 This application provides Figure 2 A flowchart illustrating step S220;

[0026] Figure 4 This application provides Figure 2 A flowchart illustrating step S230;

[0027] Figure 5 This is a schematic diagram of the structure of a test system based on voiceprint recognition provided in this application;

[0028] Figure 6 This is a schematic diagram of the electronic device provided in this application. Detailed Implementation

[0029] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0031] The following is combined with Figures 1-6 This application describes the test method, system, storage medium, and electronic device based on voiceprint recognition.

[0032] According to one aspect of the embodiments of this application, a testing method based on voiceprint recognition is provided. This voiceprint recognition-based testing method is widely applicable to whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligence house ecosystems. Optionally, in this embodiment, the above-mentioned voiceprint recognition-based testing method can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.

[0033] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.

[0034] like Figure 2 As shown, it is a schematic diagram of the implementation process of the voiceprint recognition-based testing method provided in the embodiments of this application. The voiceprint recognition-based testing method may include, but is not limited to, steps S210 to S250.

[0035] S210, Obtain a preset audio file, wherein the audio file includes target audio data and corresponding expected execution values;

[0036] S220, the target audio data is sent to the voiceprint recognition interface, and the voiceprint features and user intent returned by the voiceprint recognition interface are received, wherein the voiceprint features and user intent are generated by the voiceprint recognition interface calling the corresponding recognition algorithm to parse the target audio data;

[0037] S230, the voiceprint feature and the user intent are sent to the device execution interface to generate a device execution command; wherein, the device execution command is generated by the device execution interface calling the command generation service based on the voiceprint feature and the user intent;

[0038] S240, the device execute command is sent to the corresponding virtual device, and the execution result returned by the virtual device according to the device execute command is received;

[0039] S250, a test report is generated by comparing the voiceprint features and the execution result with the expected execution value.

[0040] In step S210 of some embodiments, a preset audio file is obtained.

[0041] Understandably, the specific execution steps can be as follows: perform cluster analysis based on the attributes of the acquired multiple voice data to determine multiple categories of sub-voice datasets; and obtain the corresponding target audio data from the multiple categories of sub-voice datasets according to a preset ratio to obtain the audio file.

[0042] It should be further noted that a pre-set audio acquisition device can be used to collect and process audio data from multiple different categories of people, thereby obtaining multiple voice data sets to determine multiple categories of sub-voice datasets.

[0043] It should be noted that the audio file includes the target audio data and the corresponding expected execution value.

[0044] In step S220 of some embodiments, the target audio data is sent to the voiceprint recognition interface, and the voiceprint features and user intent returned by the voiceprint recognition interface are received, wherein the voiceprint features and user intent are generated by the voiceprint recognition interface calling the corresponding recognition algorithm to parse the target audio data.

[0045] It is understandable that after completing step S210 to obtain the preset audio file, the specific execution steps can be as follows: First, if the measurement value of each target audio data exceeds a preset measurement value threshold, each target audio data is segmented according to a preset order to obtain multiple audio data packets corresponding to each target audio data; the segmented multiple audio data packets are sent to the voiceprint recognition interface, and the multiple audio data packets received by the voiceprint recognition interface are processed according to the preset order to obtain the target audio data; if the voiceprint recognition interface detects the target audio data, the voiceprint recognition algorithm and semantic parsing algorithm corresponding to the voiceprint recognition interface are invoked; the voiceprint recognition algorithm and semantic parsing algorithm are used to parse the target audio data to obtain the voiceprint features and user intent corresponding to the target audio data.

[0046] It should be noted that the voiceprint features and the user intent are generated by the voiceprint recognition interface calling the corresponding recognition algorithm to parse the audio data.

[0047] More specifically, when the target audio data is detected by the voiceprint recognition interface, the voiceprint recognition algorithm corresponding to the voiceprint recognition interface is invoked, and the voiceprint recognition algorithm is used to parse and process the target audio data to obtain the voiceprint features corresponding to the target audio data.

[0048] When the voiceprint recognition interface detects the target audio data, the semantic parsing algorithm corresponding to the voiceprint recognition interface is invoked to parse and process the target audio data to obtain the user intent corresponding to the target audio data.

[0049] It should be noted that voiceprint features are used to distinguish which user is making the voice. For example, they can be used to distinguish between male and female voiceprint features, or to distinguish between elderly women and female children.

[0050] It should be further explained that user intent is the actual intention expressed in their voice data, such as, "I want to turn on the air conditioner and let the cold air blow on me." "Turn on the air conditioner and let the cold air blow on me."

[0051] In step S230 of some embodiments, the voiceprint feature and the user intent are sent to the device execution interface to generate a device execution command; wherein, the device execution command is generated by the device execution interface calling a command generation service based on the voiceprint feature and the user intent.

[0052] It is understandable that after completing step S220, which involves sending the audio data to the voiceprint recognition interface and receiving the voiceprint features and user intent returned by the voiceprint recognition interface, the specific execution steps can be as follows: first, encapsulate the voiceprint features and user intent to generate a request file; send the request file to the device execution interface; and if the device execution interface finds the request file, generate the device execution command, wherein the device execution command includes a device execution instruction and a device scene execution instruction corresponding to the voiceprint features and user intent.

[0053] In step S240 of some embodiments, the device execution command is sent to the corresponding virtual device, and the execution result returned by the virtual device according to the device execution command is received.

[0054] It is understandable that after completing step S230, which sends the voiceprint feature and the user intent to the device execution interface to generate a device execution command, the specific execution steps can be: sending the device execution command obtained in step S230 to the corresponding virtual device, and receiving the execution result returned by the virtual device based on the device execution command.

[0055] Optionally, virtual devices may include, but are not limited to, virtual home appliances, virtual smart locks, etc.

[0056] It is understandable that virtual devices can be class functions, which means that the corresponding physical device is encapsulated into a corresponding class function on a computer cloud server through code. This is used to execute commands based on the device and return the corresponding execution results.

[0057] For example, taking a smart air conditioner as a virtual device, the audio data is an audio sample of a child. This audio data is sent to a voiceprint recognition interface, and the voiceprint features and user intent returned by the interface are received. These are then sent to a device execution interface to generate a device execution command. This command is sent to the corresponding virtual device (the class function corresponding to the smart air conditioner). The execution result returned by the virtual device (the class function corresponding to the smart air conditioner) based on the command is received. The audio sample is "I'm too hot." After the above steps, the resulting device execution command is: turn on the smart air conditioner to child mode, set the cooling mode temperature to 26 degrees Celsius, and the fan speed to low. The server receives the execution result returned by the virtual device (the class function corresponding to the smart air conditioner) based on the command and uses it to generate a test report.

[0058] In step S250 of some embodiments, a test report is generated by comparing the voiceprint features and the execution result with the expected execution value.

[0059] In some embodiments, the audio file may further include at least: the file path of the audio file, the file name of the audio file, and the voiceprint features of the audio data.

[0060] It should be noted that, in addition to the target audio data and the corresponding expected execution value, the audio file may also include, but is not limited to, the file path of the audio file, the file name of the audio file, and the voiceprint features of the audio data.

[0061] Furthermore, the file path of the audio file can be quickly retrieved from the query. The file name of the audio file can quickly identify the target audio file to be queried. Based on the voiceprint characteristics of the target audio data, the user intent corresponding to the audio data can be quickly analyzed and parsed.

[0062] In some embodiments, obtaining the preset audio file includes: performing cluster analysis based on the attributes of multiple acquired speech data to determine multiple categories of sub-speech datasets; and obtaining corresponding target audio data from the multiple categories of sub-speech datasets according to a preset ratio to obtain the audio file.

[0063] Understandably, the process involves first acquiring multiple voice data sets, then using a pre-defined clustering analysis algorithm to perform clustering analysis on the attributes of the multiple voice data sets, thereby determining multiple categories of sub-voice datasets. Then, according to a pre-defined ratio, the corresponding target audio data is obtained from each of the multiple categories of sub-voice datasets, thus obtaining the audio file.

[0064] It should be noted that a pre-set voice acquisition device can be used to collect and process the voice data of multiple different categories of people, and then send the collected voice data to the server so that the server can obtain multiple voice data.

[0065] It is understandable that voice acquisition devices can include, but are not limited to, home appliances with voice modules, smart speakers, voice recorders, and other electronic devices capable of acquiring and saving voice data. It is also understandable that different categories of people can be categorized by gender or age group, and their voice data can be collected and processed separately to obtain multiple sets of voice data. For example, voice data for male children, female children, male adults, female adults, elderly men, and elderly women.

[0066] Furthermore, audio files can be stored in a database or maintained and managed through a list, which may include, but is not limited to: target audio data and corresponding expected execution values, the file path of the audio file, the file name of the audio file, and the voiceprint features of the audio data.

[0067] In some embodiments, reference Figure 3 As shown, step S220 may also include, but is not limited to, steps S310 to S340.

[0068] S310, when the measurement value of each target audio data exceeds the preset measurement value threshold, each target audio data is segmented in a preset order to obtain multiple audio data packets corresponding to each target audio data.

[0069] S320, the segmented multiple audio data packets are sent to the voiceprint recognition interface, and the multiple audio data packets received by the voiceprint recognition interface are processed according to the preset order to obtain the target audio data;

[0070] S330, when the target audio data is detected by the voiceprint recognition interface, the voiceprint recognition algorithm and semantic parsing algorithm corresponding to the voiceprint recognition interface are invoked;

[0071] S340, using the voiceprint recognition algorithm and semantic parsing algorithm, the target audio data is parsed and processed to obtain the voiceprint features and user intent corresponding to the target audio data.

[0072] In step S310 of some embodiments, if the measurement value of each target audio data exceeds a preset measurement value threshold, each target audio data is segmented in a preset order to obtain multiple audio data packets corresponding to each target audio data.

[0073] It is understandable that when the measured value of each target audio data exceeds the preset measured value threshold, each target audio data is segmented in a preset order to obtain multiple audio data packets corresponding to each target audio data.

[0074] For example, if the target audio data has a measured value of 100kb, and the preset measured value threshold is 20kb, then the target audio data needs to be segmented according to a preset order. An example of the segmentation process is as follows:

[0075] The target audio data sample 1 for children is divided into audio data packets A, B, and C.

[0076] The audio data packets A, B, and C are arranged in a preset order. In the embodiments of this application, the preset order can be sequential, reverse, or arranged according to a preset rule. Therefore, no further limitation is made on the preset order.

[0077] The target audio data is segmented into multiple audio data packets to facilitate their transmission to the voiceprint recognition interface. This interface then uses the corresponding voiceprint recognition and semantic parsing algorithms to analyze the target audio data, obtaining the corresponding voiceprint features and the user's intent. This significantly improves the transmission efficiency of the target audio data, especially when the audio data sample is large.

[0078] In step S320 of some embodiments, the segmented multiple audio data packets are sent to the voiceprint recognition interface, and the multiple audio data packets received by the voiceprint recognition interface are processed according to the preset order to obtain the target audio data.

[0079] It is understandable that after performing step S310, in which the measurement value of each target audio data exceeds the preset measurement value threshold, the specific execution steps can be to send the segmented multiple audio data packets to the voiceprint recognition interface, and process the multiple audio data packets received by the voiceprint recognition interface according to the preset order to obtain the target audio data.

[0080] For example, the server first sends a message to the voiceprint recognition interface indicating that it needs to transmit three audio data packets containing a target audio data. When the voiceprint recognition interface receives these multiple audio data packets from the server, it checks the number of received audio data packets containing the same target audio data to determine if all the audio data packets containing the same target audio data have been received. This avoids the problem of audio data packets being lost during transmission.

[0081] For example, the A, B, and C audio data packets obtained in step S310 are assembled in a preset order to obtain the target audio data. Furthermore, assembling the target audio data is for invoking the corresponding recognition algorithm of the voiceprint recognition service based on the target audio data. It should be noted that the preset order is the same as the order in which the target audio data was segmented in step S310.

[0082] In step S330 of some embodiments, when the voiceprint recognition interface detects the target audio data, the voiceprint recognition algorithm and semantic parsing algorithm corresponding to the voiceprint recognition interface are invoked.

[0083] It is understandable that after step S320, which involves sending the segmented audio data packets to the voiceprint recognition interface and processing the audio data packets received by the voiceprint recognition interface according to the preset order to obtain the target audio data, the specific execution steps can be as follows: the voiceprint recognition interface will periodically detect whether there is target audio data within a preset time period according to a preset detection strategy. When target audio data is detected, the voiceprint recognition algorithm and semantic parsing algorithm corresponding to the voiceprint recognition interface will be called to parse and process the target audio data using the voiceprint recognition algorithm and semantic parsing algorithm to obtain the voiceprint features and user intent corresponding to the target audio data.

[0084] In step S340 of some embodiments, the voiceprint recognition algorithm and semantic parsing algorithm are used to parse the target audio data to obtain the voiceprint features and user intent corresponding to the target audio data.

[0085] It is understandable that, after step S330, where the target audio data is detected by the voiceprint recognition interface, and the corresponding voiceprint recognition algorithm and semantic parsing algorithm of the voiceprint recognition interface are invoked, the specific execution steps can be as follows:

[0086] When the voiceprint recognition interface detects the target audio data, the voiceprint recognition algorithm corresponding to the voiceprint recognition interface is invoked, and the voiceprint recognition algorithm is used to parse and process the target audio data to obtain the voiceprint feature corresponding to the target audio data.

[0087] When the voiceprint recognition interface detects the target audio data, the semantic parsing algorithm corresponding to the voiceprint recognition interface is invoked to parse and process the target audio data to obtain the user intent corresponding to the target audio data.

[0088] In some embodiments, reference Figure 4 As shown, step S230 may also include, but is not limited to, steps S410 to S430.

[0089] S410, encapsulate the voiceprint features and the user intent to generate a request file;

[0090] S420, the request file is sent to the device execution interface;

[0091] S430, if the request file is found in the device execution interface, the device execution command is generated, wherein the device execution command includes a device execution instruction and a device scene execution instruction corresponding to the voiceprint feature and the user intent.

[0092] In step S410 of some embodiments, the voiceprint features and the user intent are encapsulated to generate a request file.

[0093] It is understood that in step S340, the voiceprint recognition algorithm and semantic parsing algorithm are used to parse the target audio data to obtain the voiceprint features and user intent corresponding to the target audio data. The obtained voiceprint features and user intent are then encapsulated to obtain a request file, which is used to send the entire request file to the device execution interface.

[0094] In step S420 of some embodiments, the request file is sent to the device execution interface.

[0095] It is understandable that after performing step S410 to encapsulate the voiceprint features and the user intent and generate a request file, the specific execution steps can be as follows: the server will send the request file obtained in step S410 to the device execution interface, so that the device execution interface can call the command generation service to parse and process the request file.

[0096] In step S430 of some embodiments, if the request file is found by the device execution interface, the device execution command is generated, wherein the device execution command includes a device execution instruction and a device scene execution instruction corresponding to the voiceprint feature and the user intent.

[0097] It is understandable that after the step S420 of sending the request file to the device execution interface is completed, the specific execution steps can be as follows: the device execution interface periodically queries the database for a request file within a preset time period to determine whether the request file sent by the server has been received. If the request file is found, the command generation service is triggered to parse and process the request file according to the command generation service.

[0098] The command generation service is used to analyze and process the voiceprint features and user intent in the request file to generate the device execution command. The command generation service is invoked to analyze and process the voiceprint features and user intent in the request file to generate the device execution command. The device execution command includes a device execution instruction and a device scene execution instruction corresponding to the voiceprint features and user intent. A corresponding device execution result is generated based on the device execution instruction corresponding to the voiceprint features and user intent, and a corresponding scene execution result is generated based on the device scene execution instruction corresponding to the voiceprint features and user intent.

[0099] It should be noted that the scene execution result can include the execution results of multiple devices. For example, if the scene execution command "Good morning" is sent to the execution device, then some scene devices will execute this scene command, causing multiple execution devices in multiple scenes to execute. For each execution device in a scene, there will be a corresponding execution device result, and the multiple execution results of multiple execution devices in a scene constitute the scene execution result.

[0100] In some embodiments, the step of generating a test report by comparing the voiceprint features and the execution result with the expected execution value includes:

[0101] The voiceprint features and the expected voiceprint recognition value are compared in a first comparison process to obtain a first comparison result; the device execution result and the expected device execution value are compared in a second comparison process to obtain a second comparison result; and the scene execution result and the expected scene execution value are compared in a third comparison process to obtain a third comparison result.

[0102] Based on the first comparison result, the second comparison result, and the third comparison result, a test report corresponding to the target audio data is generated.

[0103] The above embodiments can quickly test the application effect of large amounts of audio data, which can not only greatly improve the testing efficiency of voiceprint recognition application effect, but also automatically generate test reports, making it easy to intuitively analyze the errors in the application process, thereby greatly improving the accuracy of recognition application.

[0104] It is understood that the execution result includes at least: device execution result and scene execution result. The expected execution value includes at least: voiceprint recognition expected value, device execution expected value, and scene execution expected value.

[0105] A second comparison process is performed between the device execution result obtained from the device execution instruction and the expected device execution value to obtain a second comparison result. Similarly, a third comparison process is performed between the scene execution result obtained from the scene execution instruction and the expected scene execution value to obtain a third comparison result.

[0106] Further, if the first comparison result indicates that the voiceprint feature and the expected value of voiceprint recognition are consistent, and the second comparison result and the third comparison result indicate that the execution result is consistent with the expected value of execution, a first test result corresponding to the target audio data is generated;

[0107] If the first comparison result indicates that the voiceprint feature and the expected value of voiceprint recognition are inconsistent, or if the second comparison result and / or the third comparison result indicate that the execution result is inconsistent with the expected value of execution, a second test result corresponding to the target audio data is generated; the test report includes the first test result and the second test result.

[0108] Understandably, if the device execution result is the same as the expected value, the scene execution result is the same as the expected value, and the identified voiceprint features are the same as the expected value, then a first test result corresponding to the target audio data is generated. Based on this first test result, it can be proven that the voiceprint features identified using the voiceprint recognition algorithm, the user intent obtained using the semantic parsing algorithm, and the scene execution result obtained from the preset scene execution command are all correct.

[0109] It should be noted that if at least any one set of comparisons shows a discrepancy between the device execution result and the expected value, the scene execution result and the expected value, and the identified voiceprint features and the expected value, a second test result is obtained. Based on the second test result, it can be determined that an error occurred during the recognition algorithm's identification of the target audio data and the testing process. The specific error step can be identified, and the cause of the error can be analyzed, allowing for rapid iterative optimization and improving the accuracy of the voiceprint recognition test.

[0110] In some embodiments of this application, after the step of generating a second test result corresponding to the target audio data when the first comparison result indicates that the voiceprint feature and the expected value of voiceprint recognition are inconsistent, or the second comparison result and / or the third comparison result indicate that the execution result is inconsistent with the expected execution value, the method further includes:

[0111] The error types of the error test cases corresponding to the second test result are analyzed and processed to determine the target error type from the multiple error types;

[0112] The neural network model to be optimized is determined based on the target error type, and the error test cases corresponding to the target error type are input into the neural network model for optimization training.

[0113] Understandably, the error types of the test cases corresponding to the second test results are analyzed and processed to determine the target error type from multiple error types.

[0114] It should be noted that the multiple error types include at least: voiceprint recognition error type, device execution error type, and scene execution error type. For example, a target error type is determined from the multiple error types, where the target error type is a voiceprint recognition error type. Based on the target error type, a neural network model to be optimized is determined, and the error test cases corresponding to the target error type are input into the neural network model for optimization training.

[0115] It should be noted that the neural network model can be a convolutional neural network model.

[0116] Furthermore, the convolutional neural network model consists of an input layer, multiple convolutional layers, multiple fully connected layers, and an output layer. The convolutional neural network model is trained using gradient descent and backpropagation algorithms. The activation function of the convolutional layers is preferably the ReLU function, and the stride, convolutional size, and number of convolutions of each convolutional layer can be freely set.

[0117] In some embodiments of this application, after the voiceprint recognition algorithm is optimized, based on previous test data, and assuming that the test environment excludes the influence of environmental noise, the speaker's volume, distance, and differences in the microphone of the receiving device, the test cases can be re-executed to achieve the comparison of the algorithm optimization test task. Furthermore, by comparing the voiceprint recognition test results before and after the algorithm optimization, a comparison report of the algorithm optimization effect can be quickly provided, thereby rapidly improving test efficiency.

[0118] This application provides a testing method based on voiceprint recognition, which can not only greatly improve the testing efficiency of voiceprint recognition application effects, but also automatically generate test reports, making it easy to intuitively analyze errors in the application process, thereby greatly improving the accuracy of recognition applications.

[0119] The following describes a voiceprint recognition-based testing system provided in this application. The voiceprint recognition-based testing system described below and the voiceprint recognition-based testing method described above can be referred to and correspond to each other.

[0120] refer to Figure 5 The diagram shows a structural schematic of a voiceprint recognition-based testing system provided in this application. The system includes: an acquisition module 510 for acquiring a preset audio file, wherein the audio file includes target audio data and a corresponding expected execution value; an identification module 520 for sending the target audio data to a voiceprint recognition interface and receiving voiceprint features and user intent returned by the voiceprint recognition interface, wherein the voiceprint features and user intent are generated by the voiceprint recognition interface calling a corresponding identification algorithm to parse the target audio data; a device execution command generation module 530 for sending the voiceprint features and user intent to a device execution interface to generate a device execution command; wherein the device execution command is generated by the device execution interface calling a command generation service based on the voiceprint features and user intent; an execution result receiving module 540 for sending the device execution command to a corresponding virtual device and receiving the execution result returned by the virtual device based on the device execution command; and a test report generation module 550 for comparing the voiceprint features and the execution result with the expected execution value to generate a test report.

[0121] In some embodiments, the acquisition module 510 is configured to perform cluster analysis based on the attributes of the acquired multiple voice data to determine multiple categories of sub-voice datasets; and to acquire corresponding target audio data from the multiple categories of sub-voice datasets according to a preset ratio to obtain the audio file.

[0122] In some embodiments, the identification module 520 is configured to: segment each target audio data in a preset order when the measurement value of each target audio data exceeds a preset measurement value threshold, to obtain multiple audio data packets corresponding to each target audio data; send the segmented multiple audio data packets to the voiceprint recognition interface, and process the multiple audio data packets received by the voiceprint recognition interface based on the preset order to obtain the target audio data; when the voiceprint recognition interface detects the target audio data, invoke the voiceprint recognition algorithm and semantic parsing algorithm corresponding to the voiceprint recognition interface; and use the voiceprint recognition algorithm and semantic parsing algorithm to parse the target audio data to obtain the voiceprint features and user intent corresponding to the target audio data.

[0123] In some embodiments, a device execution command generation module 530 is used to encapsulate the voiceprint features and the user intent to generate a request file; send the request file to the device execution interface; and, if the device execution interface finds the request file, generate the device execution command, wherein the device execution command includes a device execution instruction and a device scene execution instruction corresponding to the voiceprint features and the user intent; and use the command generation service to analyze and process the voiceprint features and the user intent in the request file to generate the device execution command.

[0124] The execution results include at least: device execution results and scene execution results;

[0125] The expected execution values ​​include at least: expected voiceprint recognition values, expected device execution values, and expected scene execution values;

[0126] In some embodiments, the test report generation module 550 is used to perform a first comparison process on the voiceprint features and the expected voiceprint recognition value to obtain a first comparison result; perform a second comparison process on the device execution result and the expected device execution value to obtain a second comparison result; and perform a third comparison process on the scene execution result and the expected scene execution value to obtain a third comparison result; and generate the test report corresponding to the target audio data based on the first comparison result, the second comparison result, and the third comparison result.

[0127] In some embodiments, the test report generation module 550 is configured to generate a first test result corresponding to the target audio data when the first comparison result indicates that the voiceprint feature and the expected value of voiceprint recognition are consistent, and the second comparison result and the third comparison result indicate that the execution result is consistent with the expected execution value; and to generate a second test result corresponding to the target audio data when the first comparison result indicates that the voiceprint feature and the expected value of voiceprint recognition are inconsistent, or the second comparison result and / or the third comparison result indicate that the execution result is inconsistent with the expected execution value; the test report includes the first test result and the second test result.

[0128] In some embodiments, the testing system is further configured to: analyze and process the error types of the error test cases corresponding to the second test result, determine a target error type from a plurality of error types; determine a neural network model to be optimized based on the target error type, and input the error test cases corresponding to the target error type into the neural network model for optimization training.

[0129] This application provides a voiceprint recognition-based testing system that not only greatly improves the testing efficiency of voiceprint recognition applications, but also automatically generates test reports, facilitating intuitive analysis of errors during the application process and thus significantly improving the accuracy of the recognition application.

[0130] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6 As shown, the electronic device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640. The processor 610, communications interface 620, and memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions from the memory 630 to execute the voiceprint recognition-based testing method described in the above embodiment.

[0131] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0132] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute a test method based on voiceprint recognition provided in the above embodiments.

[0133] In another aspect, this application also provides a computer-readable storage medium comprising a stored program, wherein the program executes a voiceprint recognition-based testing method provided in the above embodiments when it is run.

[0134] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0135] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A testing method based on voiceprint recognition, characterized in that, Applied to servers, including: Obtain a preset audio file, wherein the audio file includes target audio data and corresponding expected execution values; If the measurement value of each target audio data exceeds the preset measurement value threshold, each target audio data is segmented in a preset order to obtain multiple audio data packets corresponding to each target audio data. The segmented audio data packets are sent to the voiceprint recognition interface, and the audio data packets received by the voiceprint recognition interface are processed according to the preset order to obtain the target audio data. When the target audio data is detected by the voiceprint recognition interface, the corresponding voiceprint recognition algorithm and semantic parsing algorithm of the voiceprint recognition interface are invoked. The target audio data is parsed using the aforementioned voiceprint recognition algorithm and semantic parsing algorithm to obtain the voiceprint features and user intent corresponding to the target audio data. The voiceprint features and the user intent are encapsulated to generate a request file; The request file is sent to the device execution interface; If the request file is found in the device execution interface, a device execution command is generated, wherein the device execution command includes a device execution instruction and a device scene execution instruction corresponding to the voiceprint feature and the user intent; The target audio data is sent to the voiceprint recognition interface, and the voiceprint features and user intent returned by the voiceprint recognition interface are received. The voiceprint features and user intent are generated by the voiceprint recognition interface calling the corresponding recognition algorithm to parse the target audio data. The voiceprint features and the user intent are sent to the device execution interface to generate a device execution command; wherein, the device execution command is generated by the device execution interface calling a command generation service based on the voiceprint features and the user intent. Send the device execution command to the corresponding virtual device, and receive the execution result returned by the virtual device based on the device execution command; A test report is generated by comparing the voiceprint features and the execution results with the expected execution values.

2. The testing method based on voiceprint recognition according to claim 1, characterized in that, The process of obtaining the preset audio file includes: Cluster analysis is performed on the attributes of multiple acquired speech data to determine multiple categories of sub-speech datasets; The target audio data is obtained from the sub-speech datasets of multiple categories according to a preset ratio to obtain the audio file.

3. The testing method based on voiceprint recognition according to claim 1, characterized in that, The execution results include at least: device execution results and scene execution results; The expected execution values ​​include at least: expected voiceprint recognition values, expected device execution values, and expected scene execution values; The step of comparing the voiceprint features and the execution result with the expected execution value to generate a test report includes: The voiceprint features and the expected voiceprint recognition value are compared in a first comparison process to obtain a first comparison result; the device execution result and the expected device execution value are compared in a second comparison process to obtain a second comparison result; and the scene execution result and the expected scene execution value are compared in a third comparison process to obtain a third comparison result. Based on the first comparison result, the second comparison result, and the third comparison result, a test report corresponding to the target audio data is generated.

4. The testing method based on voiceprint recognition according to claim 3, characterized in that, The step of generating the test report corresponding to the target audio data based on the first comparison result, the second comparison result, and the third comparison result includes: If the first comparison result indicates that the voiceprint feature and the expected value of voiceprint recognition are consistent, and the second comparison result and the third comparison result indicate that the execution result is consistent with the expected value of execution, a first test result corresponding to the target audio data is generated. If the first comparison result indicates that the voiceprint feature and the expected value of voiceprint recognition are inconsistent, or if the second comparison result and / or the third comparison result indicate that the execution result is inconsistent with the expected value of execution, a second test result corresponding to the target audio data is generated; the test report includes the first test result and the second test result.

5. The testing method based on voiceprint recognition according to claim 4, characterized in that, After the step of generating a second test result corresponding to the target audio data when the first comparison result indicates that the voiceprint feature and the expected value of voiceprint recognition are inconsistent, or the second comparison result and / or the third comparison result indicate that the execution result is inconsistent with the expected value of execution, the method further includes: The error types of the error test cases corresponding to the second test result are analyzed and processed to determine the target error type from the multiple error types; The neural network model to be optimized is determined based on the target error type, and the error test cases corresponding to the target error type are input into the neural network model for optimization training.

6. A testing system based on voiceprint recognition, characterized in that, Applied to servers, including: The acquisition module is used to acquire a preset audio file, wherein the audio file includes target audio data and corresponding expected execution values; The recognition module is used to segment each target audio data in a preset order when the measurement value of each target audio data exceeds a preset measurement value threshold, obtaining multiple audio data packets corresponding to each target audio data; sending the segmented multiple audio data packets to a voiceprint recognition interface, and processing the multiple audio data packets received by the voiceprint recognition interface based on the preset order to obtain the target audio data; when the voiceprint recognition interface detects the target audio data, calling the voiceprint recognition algorithm and semantic parsing algorithm corresponding to the voiceprint recognition interface; using the voiceprint recognition algorithm and semantic parsing algorithm to parse the target audio data, obtaining the voiceprint features and user intent corresponding to the target audio data; encapsulating the voiceprint features and user intent to generate a request file; sending the request file to a device execution interface; and when the device execution interface finds the request file, generating a device execution command, wherein the device execution command includes a device execution instruction and a device scene execution instruction corresponding to the voiceprint features and user intent. The target audio data is sent to the voiceprint recognition interface, and the voiceprint features and user intent returned by the voiceprint recognition interface are received. The voiceprint features and user intent are generated by the voiceprint recognition interface calling the corresponding recognition algorithm to parse the target audio data. A device execution command generation module is used to send the voiceprint features and the user intent to the device execution interface to generate a device execution command; wherein, the device execution command is generated by the device execution interface calling a command generation service based on the voiceprint features and the user intent; The execution result receiving module is used to send the device execution command to the corresponding virtual device and receive the execution result returned by the virtual device based on the device execution command; The test report generation module is used to compare the voiceprint features and the execution results with the expected execution values ​​to generate a test report.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 5.

8. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 5 through the computer program.

Citation Information

Patent Citations

  • Voice skill testing method and device, and equipment

    CN113223496A

Cited By

  • Real-time dialogue analysis method and system based on voiceprint recognition

    CN121565180A