Intelligent television voice and voiceprint recognition three-element test method and automatic test system
By combining repeated voiceprint recognition testing, environmental adaptability testing, and core scenario application methods, the instability and accuracy issues of voiceprint recognition testing in smart TVs have been resolved. This has enabled more efficient testing that is closer to user scenarios, and improved recognition accuracy and stability.
Patent Information
- Application Number
- CN202511505737.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2026-02-24
AI Technical Summary
Existing smart TV voiceprint recognition testing methods lack stability checks, the test results differ greatly from the accuracy in actual use, and the recognition accuracy is low in complex environments, making it impossible to effectively check the accuracy of voiceprint recognition.
A testing method combining repeated voiceprint recognition testing, environmental adaptability testing, and voiceprint core scenario application testing was adopted. Multiple tests were conducted in different environments and core application scenarios, and voiceprint recognition testing was carried out in conjunction with an automated testing system.
It improves the accuracy and stability of voiceprint recognition testing, enhances recognition capabilities in complex environments, increases testing efficiency and accuracy, and meets the needs of users in diverse usage scenarios.
Smart Images

Figure CN121565145A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent speech recognition testing technology, specifically, to a method and automated testing system for testing the three elements of voiceprint recognition in intelligent televisions. Background Technology
[0002] The voiceprint function of smart TVs is an advanced application of voice recognition technology. By analyzing the unique characteristics of a user's voice (such as pitch, timbre, and speech rate), it can accurately identify family members and provide personalized services and security management. Current voiceprint recognition testing methods typically involve testing multiple users' audio in a single, isolated environment. The drawbacks are the limited testing methods, the lack of stability checks, and the high accuracy rates shown in test data, while resulting in low accuracy in actual use. The main issues include: voiceprints not being recognized or misrecognized in complex, noisy environments; low recognition rates at off-angle angles; probabilistic failure to execute core application functions integrated with voiceprint recognition; and probabilistic failure to recognize voiceprints due to timeouts in the voiceprint recognition process. Summary of the Invention
[0003] The purpose of this invention is to provide a three-element testing method and automated testing system for voiceprint recognition in smart TVs, which solves the problems of inaccurate and unstable results of existing voiceprint recognition testing methods, the large difference between the accuracy data of voiceprint recognition testing and the accuracy found in actual use, and the inability of existing methods to effectively check the accuracy of voiceprint recognition.
[0004] The present invention solves the above problems through the following technical solution:
[0005] A three-element test method for voiceprint recognition in smart TVs is proposed, which combines a voiceprint recognition repetition test method, an environmental adaptability test method, and a voiceprint core scene application method.
[0006] Furthermore, the environmental adaptability testing method includes:
[0007] Set up a test environment, which includes a home scenario, a speech scenario at an off-angle, and a noisy scenario;
[0008] Speech recognition tests were conducted in each test environment.
[0009] Furthermore, the voiceprint recognition repetition test method is as follows:
[0010] Record the voiceprint of the person being tested and register a voiceprint account on a smart TV;
[0011] Record the voiceprint audio of the test subjects. For each test subject's voiceprint audio, use the same person's single audio to repeatedly and continuously test multiple times to check the stability of voiceprint recognition. The test subjects' voiceprint audio contains voice commands in the core application scenarios of voiceprint.
[0012] Furthermore, the test subjects are representative individuals sampled according to different ages and genders, and the method for recording the voiceprint audio of the test subjects is as follows:
[0013] Representatives recorded their speech data in a quiet, soundproof environment using high-fidelity speakers, according to the test cases, to form the voiceprint audio of the tested individuals.
[0014] Furthermore, the voiceprint core scenario application method involves conducting voiceprint recognition tests in each core application scenario based on the core application scenarios of the smart TV.
[0015] Furthermore, the core application scenarios include:
[0016] Personalized content recommendation scenario: Used to automatically push content based on family members' viewing history and preferences;
[0017] Multi-user permission management: used to assign different operation permissions through voiceprint in parental control mode;
[0018] Family interaction optimization scenario: Used to set different desktop content based on family members with registered voiceprints, realizing personalized customization based on voiceprints;
[0019] In a security scenario, it is used to enter visitor mode when an unfamiliar voiceprint is detected, restricting access to viewing history and smart home control.
[0020] Furthermore, the speech recognition testing method that combines the voiceprint recognition repetitive testing method, the environmental adaptability testing method, and the voiceprint core scene application method is as follows:
[0021] Using the voiceprint audio of each test subject, voiceprint recognition tests were conducted on various combinations of test environments and core application scenarios.
[0022] An automated testing system for intelligent TV voiceprint recognition, which implements the aforementioned three-element testing method for intelligent TV voiceprint recognition, includes an intelligent TV, a test control terminal, and test fixtures connected via communication, and also includes an artificial mouth device, wherein:
[0023] The testing fixture is used to test the voice recognition function of the smart TV according to the control instructions of the test control terminal, obtain the voice recognition results of the smart TV, and feed them back to the test control terminal; it is also used to play the voiceprint audio of the person being tested through the audio equipment, or to send the voiceprint audio of the person being tested to the artificial mouth device.
[0024] Artificial mouth device is used to acquire the voiceprint audio of the person being tested from audio equipment or testing fixtures and play it as a voice command to a smart TV.
[0025] The smart TV is used to perform voiceprint recognition on the acquired audio of the test subject and send the voice recognition test results to the test fixture.
[0026] The test control terminal is used to issue control commands to the test fixture, control the start and stop of the test, set the number of tests and record the results, and generate test reports from the speech recognition test results.
[0027] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0028] (1) This invention adopts a three-element test method for voiceprint recognition, which solves the problems of inaccurate and unstable voiceprint recognition test results. It solves the problem that the voiceprint recognition test results deviate greatly from the actual recognition accuracy in existing methods, and cannot effectively check the voiceprint recognition accuracy.
[0029] (2) In addition to manual testing, the present invention can also be imported into automated testing to conduct long-term stability testing.
[0030] (3) Based on the principle of voiceprint recognition technology, the three-element test method of voiceprint recognition is closer to the user's usage scenario. It adds multiple environments, core application scenarios of voiceprint, and long-term stability test of voiceprint, which can effectively check the accuracy of voiceprint recognition and improve the accuracy and efficiency of software test results.
[0031] (4) By adding environmental adaptability testing, this invention can check the accuracy of voiceprint recognition under different environments. Attached Figure Description
[0032] Figure 1 This is a flowchart of the present invention;
[0033] Figure 2 A schematic diagram illustrating the setup of a home environment;
[0034] Figure 3 This is a schematic diagram of the construction of a noise scene;
[0035] Figure 4 A schematic diagram illustrating the setup of a scene where speech is delivered from a slightly off-center angle;
[0036] Figure 5 This is a schematic diagram of the automated testing system of the present invention. Detailed Implementation
[0037] The present invention will be further described in detail below with reference to embodiments, but the implementation of the present invention is not limited thereto.
[0038] Example 1:
[0039] Combined with appendix Figure 1 As shown, a three-element test method for voiceprint recognition in smart TVs is proposed, which combines a voiceprint recognition repetition test method, an environmental adaptability test method, and a voiceprint core scene application method. The method includes:
[0040] (a) Repeated test method for voiceprint recognition
[0041] Voiceprint audio recording: The voiceprint audio of the test subjects (i.e., real human voiceprint audio) is recorded in advance. The specific steps are as follows: Representatives of men, women, the elderly, and children are sampled and their speech is recorded in a quiet, professionally soundproofed environment using high-fidelity speakers to create a .wav file of their voices, according to the test case data. (Techniques to maintain audio fidelity: Recording is done using professional high-fidelity speakers, and in a quiet environment to avoid other sources of interference).
[0042] The representative samples need to consider voiceprint recognition and regional accents, selecting the elderly, middle-aged, young adults, children under 16 years old, and both male and female, in order to broaden the representative sample range and better reflect user characteristics. The test case corpus contains voice commands from core application scenarios, with each voice command consisting of a short phrase of a few words, and the corpus size is 10-30 samples.
[0043] (ii) Environmental adaptability testing method
[0044] Set up a testing environment that includes a home setting, a speech scene at an off-center angle, and a noisy environment, for example:
[0045] Home environment: such as Figure 2 As shown, the ambient sound around the TV is less than 50dB (the sound pressure level of a user's home environment is generally between 40 and 50dB when tested with a sound pressure meter, simulating a user's home usage scenario), while the sound level of the TV playing video is about 60dB. When a voice command is broadcast 90 degrees to the TV, the sound pressure level of the broadcast voice command is about 70dB when tested at the microphone of the entire unit.
[0046] Noise environment: such as Figure 3 As shown, the ambient sound around the TV is 65dB (the sound pressure level of a user's home environment is generally around 65dB when tested with a sound pressure meter to simulate a noisy environment). When a voice command is broadcast 90 degrees to the TV, the sound pressure level of the broadcast voice command is approximately 70dB at the microphone of the TV.
[0047] Situations where the speaker is speaking from a different angle: such as Figure 4 As shown, the ambient sound around the TV in a home environment is less than 50dB. When the announcer stands to the left or right of the TV at an angle (30 degrees, 150 degrees) and faces the TV to issue instructions, the sound pressure level of the voice instructions is about 70dB when tested at the microphone of the whole machine.
[0048] (III) Voiceprint Core Scene Application Method:
[0049] The core application scenarios of voiceprint recognition mainly include:
[0050] 1) Personalized content recommendations: Automatically push movies, music, and other content based on family members' viewing history and preferences. For example, if Dad says "I want to watch a movie," the TV will recommend science fiction films; when the child gives the command, the system switches to the cartoon list.
[0051] 2) Multi-user permission management: In parental control mode, the TV assigns different operation permissions through voiceprint. For example, children cannot purchase paid content by voice, and only parents' voiceprints can unlock advanced settings.
[0052] 3) Optimized Family Interaction: Different desktop content can be set for family members with registered voiceprints, enabling personalized customization based on voiceprint. For example, seniors can set a senior mode, children can set a children's mode, and parents can set a standard mode. By recognizing different voiceprints registered in the family, the TV can switch to the corresponding member's desktop content. For example, if a child says "Open my desktop," the TV will automatically switch to the children's mode desktop.
[0053] 4) Security Protection: When an unfamiliar voiceprint is detected, the system enters visitor mode, restricting access to viewing history and smart home control to protect privacy. For example, if a guest visits an unregistered member and says "I want to watch a movie," the TV will not recommend the family member's viewing history or other private information.
[0054] Before testing, the test subjects need to register their voiceprints. Turn on the TV and connect to the internet. Install the Family Member Center app, open the Family Member Center app, add members, record the test subject's voiceprint, and after successful registration, find the registered member's voiceprint ID number "voiceid:registers-*****". Continue registering multiple voiceprint accounts. For example, you can obtain the registered member's voiceprint ID number using logcat|grep voiceid, such as: name:User 1 voiceid: registers-D7B90420560005W50601000J-00014. Different people register different voiceprint ID numbers, and the voiceprint ID number of the same person after registration is unique.
[0055] In various testing environments, recorded audio from test subjects was played, tailored to different representatives and application scenarios. Using each subject's voiceprint audio, voiceprint recognition tests were conducted across various combinations of the testing environment and core application scenario. The voiceprint recognition results were then used for voiceprint identification and judgment. For example, in home environments, noisy environments, and scenarios with speech at an off-center angle, test subjects or human voices played real audio, and a smart TV received voice commands, such as "I want to watch a movie," enabling far-field and near-field voice activation and recognition.
[0056] In the core voiceprint scenario, voice commands are used. The person being tested or the artificial mouth device issues a command, which is received by the television. Based on the voiceprint recognition logic (feature extraction, model training, real-time comparison), when the user issues a command, the system compares the real-time voiceprint with a pre-stored model, completing identity verification within 0.5 seconds. Based on the voiceprint recognition model, a voiceprint ID for the command is generated. According to the principles of family member center voiceprint registration and voiceprint model technology, different people register different voiceprint IDs, while the voiceprint ID of the same person is unique after registration.
[0057] The voiceprint recognition result obtained from the user's voice feedback result is "action=speechFeedback;text=I want to watch a movie". The recognition result of "text=" is the keyword of the voice execution operation in this round, which can control the TV to execute feedback. This feedback result can be used to check whether the core scene function of the voiceprint is executed correctly.
[0058] Voiceprint recognition information in the speech recognition feedback results:
[0059] vprUserId=registers-D7B90420560005W50601000J-00014, this voiceprint ID number accurately represents the voiceprint recognition keyword of this round of voice commands.
[0060] Voiceprint recognition result judgment: Check whether the registered voiceprint ID and the voiceprint recognition ID are completely consistent.
[0061] Based on the technical principle of the uniqueness of the voiceprint ID number after the same person registers, the voiceprint of the registrant "voiceid:registers-D7B90420560005W50601000J-00014" is compared with the voiceprint recognition result "vprUserId=registers-D7B90420560005W50601000J-00014" to determine whether the strings are completely consistent. If they are completely consistent, it means that the voiceprint recognition is correct; otherwise, the voiceprint recognition is incorrect.
[0062] Upon completion of this test, if it was conducted manually, the results will be recorded automatically; if it was conducted automatically, the automated testing system can automatically record and generate a test report file. This report file can meet the development needs for problem analysis and test documentation. (Note: An automated testing tool is an open-source testing tool based on a Unix system, containing the commands used for testing.)
[0063] The voiceprint recognition result is judged based on the speech recognition principle and voiceprint model (the speech principle and voiceprint model adopt existing methods, which are not within the scope of this invention and will not be described here).
[0064] This invention utilizes real-person audio samples from different age groups, genders, speaking styles, and accents to effectively identify issues with the voiceprint model, such as timeouts in voiceprint recognition, pre-recording muting, and recognition errors. The personalized voice characteristics of each individual better reflect user experiences. By adding environmental adaptability testing, it examines voiceprint recognition accuracy under various environments. Unlike single-environment testing methods, the added scenarios, such as noise and angle deviation, more closely resemble typical user scenarios. Furthermore, by incorporating application testing of core application scenarios for voiceprint recognition—scenarios that users can perceive—this method effectively enhances the user experience.
[0065] Example 2:
[0066] Based on Example 1, such as Figure 5 As shown, the automated testing system for voiceprint recognition in smart TVs includes a smart TV with communication connectivity, a test control terminal, and test fixtures, as well as an artificial mouth device, wherein:
[0067] The testing fixture is used to test the voice recognition function of the smart TV according to the control instructions of the test control terminal, obtain the voice recognition results of the smart TV, and feed them back to the test control terminal; it is also used to play the voiceprint audio of the person being tested through the audio equipment, or to send the voiceprint audio of the person being tested to the artificial mouth device.
[0068] Artificial mouth device is used to acquire the voiceprint audio of the person being tested from audio equipment or testing fixtures and play it as a voice command to a smart TV.
[0069] The smart TV is used to perform voiceprint recognition on the acquired audio of the test subject and send the voice recognition test results to the test fixture.
[0070] The test control terminal is used to issue control commands to the test fixture, control the start and stop of the test, set the number of tests and record the results, and generate test reports from the speech recognition test results.
[0071] The testing fixture includes a fixture remote control, fixture relays, and automated testing tools. These tools control the fixture relays, remote control, and router to test voice recognition functionality. The voice recognition test results are automatically sent to the testing computer via a data cable, enabling remote, unattended monitoring. The artificial mouth device includes a computer, an external multi-channel sound card, and an equalizer. The automated testing system can continuously cycle through tests and record results. Commands from the automated testing tools can be set to perform hundreds or thousands of stress tests.
[0072] It's worth noting that manual testing can be used in the early unit and integration testing phases of development, while automated testing can be used in the later system testing and acceptance testing phases. This is beneficial for discovering low-probability and long-term stability issues. Automated testing utilizes voiceprint recognition information to write automated test scripts for automatic judgment. Through the establishment and cooperation of a human-machine environment, automated testing is achieved, and test results are automatically statistically analyzed, providing a basis for development and test analysis. Furthermore, it can perform hundreds or thousands of automated tests, replacing repetitive manual testing operations. This method has been implemented, and automated stress testing is beneficial for checking the accuracy and stability of voiceprint recognition.
[0073] Based on the principles of voiceprint recognition technology, the three-element test method for voiceprint recognition is closer to user scenarios. It adds multi-environment testing, core application scenarios of voiceprint, and long-term stability testing of voiceprint, which can effectively check the accuracy of voiceprint recognition and improve the accuracy and efficiency of software testing results.
[0074] Although the present invention has been described herein with reference to illustrative embodiments, the above embodiments are merely preferred embodiments of the present invention, and the implementation of the present invention is not limited to the above embodiments. It should be understood that those skilled in the art can devise many other modifications and implementations, which will fall within the scope and spirit of the principles disclosed in this application.
Claims
1. A method for testing the three elements of voiceprint recognition in smart TVs, characterized in that, A voiceprint recognition testing method combining repeated testing, environmental adaptability testing, and core scenario application of voiceprint recognition is proposed.
2. The method for testing the three elements of voiceprint recognition in smart TVs according to claim 1, characterized in that, The environmental adaptability testing method includes: Set up a test environment, which includes a home scenario, a speech scenario at an off-angle, and a noisy scenario; Speech recognition tests were conducted in each test environment.
3. The method for testing the three elements of voiceprint recognition in smart TVs according to claim 1, characterized in that, The voiceprint recognition repetition test method is as follows: Record the voiceprint of the person being tested and register a voiceprint account on a smart TV; Record the voiceprint audio of the test subjects. For each test subject's voiceprint audio, use the same person's single audio to repeatedly and continuously test multiple times to check the stability of voiceprint recognition. The audio of the tested person's voiceprint contains voice commands from the core application scenarios of the voiceprint.
4. The method for testing the three elements of voiceprint recognition in smart TVs according to claim 3, characterized in that, The test subjects were representative individuals sampled according to different ages and genders, and the method for recording the voiceprint audio of the test subjects was as follows: Representatives recorded their speech data in a quiet, soundproof environment using high-fidelity speakers, according to the test cases, to form the voiceprint audio of the tested individuals.
5. The method for testing the three elements of voiceprint recognition in smart TVs according to claim 3, characterized in that, The core scenario application method of voiceprint is based on the core application scenarios of smart TVs, and voiceprint recognition tests are conducted in each core application scenario.
6. The method for testing the three elements of voiceprint recognition in smart TVs according to claim 5, characterized in that, The core application scenarios include: Personalized content recommendation scenario: Used to automatically push content based on family members' viewing history and preferences; Multi-user permission management: used to assign different operation permissions through voiceprint in parental control mode; Family interaction optimization scenario: Used to set different desktop content based on family members with registered voiceprints, realizing personalized customization based on voiceprints; In a security scenario, it is used to enter visitor mode when an unfamiliar voiceprint is detected, restricting access to viewing history and smart home control.
7. The method for testing the three elements of voiceprint recognition in smart TVs according to claim 6, characterized in that, The voiceprint recognition testing method, which combines the voiceprint recognition repetition test method, the environmental adaptability test method, and the voiceprint core scene application method, is as follows: Using the voiceprint audio of each test subject, voiceprint recognition tests were conducted on various combinations of test environments and core application scenarios.
8. An automated testing system for intelligent television voiceprint recognition that implements the three-element testing method for intelligent television voiceprint recognition as described in any one of claims 1-7, characterized in that, It includes a smart TV with communication connectivity, a test control terminal and test fixtures, and an artificial mouth device, among which: The testing fixture is used to test the voice recognition function of the smart TV according to the control instructions of the test control terminal, obtain the voice recognition results of the smart TV, and feed them back to the test control terminal; it is also used to play the voiceprint audio of the person being tested through the audio equipment, or to send the voiceprint audio of the person being tested to the artificial mouth device. Artificial mouth device is used to acquire the voiceprint audio of the person being tested from audio equipment or testing fixtures and play it as a voice command to a smart TV. The smart TV is used to perform voiceprint recognition on the acquired audio of the test subject and send the voice recognition test results to the test fixture. The test control terminal is used to issue control commands to the test fixture, control the start and stop of the test, set the number of tests and record the results, and generate test reports from the speech recognition test results.