Smart television voice recognition test method and test system based on real person broadcast

By recording live speaking audio files and broadcasting with artificial mouth equipment, the recognition errors and inefficiency problems caused by sound differences in the existing smart TV voice recognition testing methods are solved, and higher test accuracy and efficiency are achieved.

CN120220650APending Publication Date: 2025-06-27SICHUAN CHANGHONG ELECTRIC CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510402444.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing smart TV voice recognition testing method uses electronic component sound generators to broadcast sound, which is a big gap with real-person voice, resulting in low recognition errors and accuracy, low manual re-reading test efficiency and large workload.

Method used

The test method based on live broadcast is adopted to test the voice recognition system of the smart TV by recording live speech audio files and using artificial mouth equipment to broadcast these audio files as voice commands.

Benefits of technology

This method can be closer to user usage scenarios, improve the accuracy and efficiency of speech recognition testing, and has high stability in the test process and can run for a long time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220650A_ABST
    Figure CN120220650A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of smart televisions, and particularly discloses a test method and a test system for voice recognition of a smart television based on real person broadcasting, which are based on the principle of voice technology, test voice recognition by using real person audio broadcasting, acquire key fields, realize automatic test and automatically generate a test report file. The test result is closer to the use scene of the user, and the test accuracy and the test efficiency are improved. The method is closer to the real human voice and can be closer to the actual use scene of the smart television, and the accuracy of the smart television voice recognition test is improved. The mode that a real person speaks an audio file and is matched with manual mouth equipment broadcasting is adopted, and compared with manual on-site broadcasting, the testing efficiency is greatly improved; by controlling a real person to speak the audio playing box to broadcast with a manual mouth, the testing process can be stable for a long time and is high in stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart TVs, in particular to a test method and a test system for voice recognition of smart TVs based on real-person broadcast. Background Art

[0002] The voice recognition index of smart TVs is one of the important indicators of TV products and also one of the important functions of TV products. There are still many aspects that need to be improved in the voice recognition technology of smart TVs. For example, in a noisy living room or when multiple people are speaking simultaneously, the voice recognition system of smart TVs is easily affected by background noise, reverberation, and interference signals, resulting in recognition errors and low voice recognition accuracy; or when the volume of the TV program is relatively large, voice commands may not be accurately recognized. If the voice recognition system of a smart TV is difficult to accurately recognize and understand the differences in pitch, timbre, etc. in the actual environment, it is very easy to have recognition errors, and sometimes it cannot accurately understand the user's intention, requiring the user to repeat or rephrase the command multiple times, reducing the naturalness and fluency of the interaction and resulting in insufficient personalized experience for users. Therefore, reliable testing of the voice recognition system of smart TVs is an important technical means to discover voice recognition problems of smart TVs and improve the accuracy of voice recognition of smart TVs.

[0003] Currently, in the smart TV testing industry for voice recognition, an electronic component sound generator built into the computer system is usually used to broadcast sounds for voice recognition testing. Since the sounds broadcast by the electronic component sound generator built into the computer system are quite different from real human voices, for example, the broadcast sounds have emotional loss, the sound effects are relatively pale, the speed of words and sentences is too fast or too slow, resulting in phenomena such as stuttering, missing words, and discontinuous sentences. In addition, the sounds broadcast by the electronic component sound generator are relatively thinner than real human voices, and low-frequency sounds are easily missing in recognition. There are problems that the voice recognition test results are inaccurate and do not conform to the user's usage scenarios; another test method is to use artificial repeated reading of test terms for voice recognition testing. This type of method not only has low efficiency but also requires a large amount of workload, and cannot reach the test quantity of repeating tens of millions of times, unable to meet the needs of smart TV testing. Summary of the Invention

[0004] The purpose of the present invention is to: address the problem that in the existing voice recognition test methods for smart TVs, due to the significant difference between the sounds broadcast by the electronic component sound generator and real human voices, the voice recognition test results are inaccurate and do not conform to the user's usage scenarios; and provide a test method and a test system for voice recognition of smart TVs based on real-person broadcast.

[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0006] A test method for voice recognition of smart TVs based on real-person broadcast includes the following steps:

[0007] S1. Recording audio files of real people speaking: Select representative people from men, women, old people and children, and record audio files of real people speaking according to the corpus of the test case; the corpus includes at least one phrase;

[0008] S2, broadcasting the audio file of real person speaking as voice command: activating the smart TV to receive the voice command; controlling the artificial mouth device to broadcast the audio file of real person speaking, so that the artificial mouth device issues the corresponding voice command;

[0009] S3. Obtain and determine the voice recognition result: after the smart TV receives the voice command, obtain the voice recognition result of the smart TV; and determine whether the voice recognition of the smart TV is correct based on the real person speaking audio file.

[0010] The test method for smart TV voice recognition based on real person broadcast of the present invention selects a person to record a real person speech audio file, and the obtained real person speech audio file is closer to the user usage scenario, and the voice characteristics of men, women, the elderly and children are all recorded representatively. In addition, the real person speech audio file is controlled by an artificial mouth device to broadcast the real person speech audio file, which is closer to the real human voice, avoids the situation that the sound effect of the electronic components of the computer system is relatively pale, and the speed of words and sentences is too fast or too slow, can be closer to the actual usage scenario of the smart TV, and improve the accuracy of the smart TV voice recognition test; the real person speech audio file is used in conjunction with the artificial mouth device for broadcasting, which greatly improves the test efficiency compared with the manual on-site broadcast; by controlling the real person speech audio player and the artificial mouth broadcast, the test process can be stable for a long time, and the test process stability is high.

[0011] Preferably, the test method for smart TV voice recognition based on real-person broadcast described in the present invention, the method for obtaining the voice recognition result of the smart TV, specifically includes the following steps: collecting and identifying the execution operation information of the smart TV through the serial port command between the smart TV and the test control terminal; the execution operation information includes: activation status and recognition status.

[0012] As a preferred solution of the present invention, by respectively obtaining the activation state and recognition state in the execution operation information of the smart TV, it is not only conducive to quickly and accurately obtaining the voice recognition results, but also can keep the collected and recognized operation information clear and not confusing, which provides convenience for the subsequent judgment of the recognition results.

[0013] Preferably, the test method for smart TV voice recognition based on real person broadcasting of the present invention determines whether the voice recognition of the smart TV is correct, and specifically comprises the following steps:

[0014] Determine whether the execution operation information contains the first keyword field; if the execution operation information contains the first keyword field, determine that the activation status is activated; if the execution operation information does not contain the first keyword field, determine that the activation status is not activated;

[0015] Identify the second keyword field of the execution operation information; compare the second keyword field with the text of the corresponding live speech audio file, and determine whether the second keyword field is exactly the same as the text;

[0016] If the second keyword field is exactly the same as the text, determine that the speech recognition is correct; if there are differences between the second keyword field and the text, determine that the speech recognition is incorrect.

[0017] As a preferred solution of the present invention, by determining whether the execution operation information contains the first keyword field and identifying the second keyword field, and then respectively determining the activation status and whether the speech recognition is correct, the error of result determination caused by data collection or comparison confusion is reduced, and the accuracy of the test result is further improved, thereby improving the reliability of the intelligent TV speech recognition test.

[0018] Preferably, for the test method of intelligent TV speech recognition based on live broadcast of the present invention, to determine whether the speech recognition of the intelligent TV is correct, the following steps are further included: screen and classify the activation status and recognition status, and then obtain multiple groups of test results; the test results include: not activated and not recognized, activated and not recognized, activated and recognized correctly, activated and recognized incorrectly. Record multiple groups of the test results in rows and columns, and generate a test report.

[0019] As a preferred solution of the present invention, by screening and classifying the activation status and recognition status, obtaining multiple groups of test results, recording according to different test results, and generating a test report, it is beneficial for testers to conveniently and quickly evaluate the intelligent TV speech recognition function; it promotes R & D personnel to timely discover the reasons for inaccurate intelligent TV speech recognition, and facilitates the analysis and research of problems.

[0020] Preferably, for the test method of intelligent TV speech recognition based on live broadcast of the present invention, to activate the intelligent TV to receive voice commands, the following steps are specifically included: for far-field speech, wake up the intelligent TV through the intelligent voice assistant, and then activate the intelligent TV to receive voice commands; for near-field speech, activate the intelligent TV to receive voice commands through the control remote control.

[0021] As a preferred embodiment of the present invention, the method for activating a smart TV to receive voice commands can adapt to different activation methods for far-field voice and near-field voice respectively, improve the activation efficiency and accuracy, reduce the probability of test failure caused by non-activation, and further enhance the applicability and reliability of the test.

[0022] Preferably, in the test method for smart TV voice recognition based on real-person broadcast of the present invention, the number of language materials is 30 to 500.

[0023] As a preferred embodiment of the present invention, by setting the number of language materials to 30 to 500, it is possible to provide as many tests as possible that conform to the actual application scenarios, and further improve the reliability of the test.

[0024] Preferably, in the test method for smart TV voice recognition based on real-person broadcast of the present invention, during the period from when the artificial mouth device issues the corresponding voice command to when the smart TV receives the voice command, the sound intensity in the test room is maintained at <50 dB.

[0025] As a preferred embodiment of the present invention, by building a test environment, using professional high-fidelity audio equipment to record the real-person speech audio file of the representative person, and cooperating with controlling the sound intensity in the voice test room to be <50 dB, it is more in line with the home environment noise sound pressure range, simulates the smart TV used in the user's home use scenario, enhances the environmental matching degree of the smart TV voice recognition, and further improves the reliability of the smart TV voice recognition test.

[0026] To achieve the purpose of the present invention, the present invention provides another technical solution:

[0027] A test system for smart TV voice recognition includes: an audio device, a smart TV, a test control terminal, a test tooling, and an artificial mouth device; the audio device is used to record the real-person speech audio file; the smart TV is used to recognize the voice command; the test control terminal is used to control the artificial mouth device to broadcast and control the test tooling to obtain the voice recognition result; the test tooling is used to cooperate with the smart TV to recognize the voice command and to collect the execution operation information; the artificial mouth device is used to broadcast the real-person speech audio file as the corresponding voice command.

[0028] The test system for voice recognition of smart TVs described in the present invention can realize automated testing of voice recognition of smart TVs by recording audio files of real storytelling and then controlling the broadcast of an artificial mouth device through a test control terminal, in conjunction with the control and activation functions of the smart TV itself; correspondingly, by controlling the test tooling to obtain voice recognition results, a test report file can be automatically generated, so that the test results are closer to the user's usage scenarios, and the accuracy and efficiency of the test are improved; through the mutual cooperation of the various modules of the above-mentioned system, long-term and stable voice recognition testing can be achieved with high efficiency and good stability.

[0029] Preferably, in the test system for smart TV voice recognition described in the present invention, the test tool comprises: a tool remote control, a relay, a routing device and a control module; the tool remote control is used to cooperate with the control module to control the test tool; the relay is used to respond to the operation instructions of the control module; the routing device is used to provide network connection; the control module is used to control the remote control, relay and routing device to recognize voice instructions and collect execution operation information, and send the voice recognition results to the test control terminal through a data line.

[0030] As a preferred solution of the present invention, the control module controls the relay, tooling remote control and routing equipment to realize the test of the voice recognition function, and automatically sends the voice recognition test results to the test control terminal through the data line to realize remote unmanned monitoring, which further facilitates the test personnel to operate and obtain the voice recognition results.

[0031] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0032] 1. The test method is closer to real human voices, avoids the situation where the sound effects of the electronic components of the computer system are relatively pale, and the speed of words and sentences is too fast or too slow. It can be closer to the actual use scenarios of smart TVs and improve the accuracy of smart TV voice recognition tests. Compared with manual on-site broadcasting, the test efficiency is greatly improved. By controlling the real person speaking audio player and the artificial mouth broadcast, the test process can be stable for a long time and the test process is highly stable.

[0033] 2. The test system can realize the automated test of smart TV voice recognition by recording the audio file of real storytelling and controlling the broadcast of the artificial mouth device through the test control terminal, in combination with the control and activation functions of the smart TV. By controlling the test tooling to obtain the voice recognition results, the test report file can be automatically generated, so that the test results are closer to the user's usage scenarios, and the test accuracy and efficiency are improved. Through the cooperation of various modules, long-term and stable voice recognition testing can be achieved with high efficiency and good stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is a schematic flowchart of a test method for speech recognition of an intelligent TV based on real-person broadcasting in the present invention;

[0035] Figure 2 is a schematic diagram of the module structure of a test system for speech recognition of an intelligent TV in the present invention. Specific embodiments

[0036] The present invention will be described in detail below with reference to the accompanying drawings.

[0037] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0038] Embodiment 1:

[0039] Referring to Figure 1 and Figure 2 as shown, this embodiment provides a test method for speech recognition of an intelligent TV based on real-person broadcasting, including the following steps:

[0040] S1. Record the real-person speech audio file: Sample representative persons from men, women, the elderly, and children, and record the real-person speech audio file of the representative persons according to the corpus of the test cases; the corpus includes at least one phrase;

[0041] S2. Broadcast the real-person speech audio file as a voice command: Activate the intelligent TV to receive voice commands; control the artificial mouth device to broadcast the real-person speech audio file, so that the artificial mouth device issues the corresponding voice command;

[0042] S3. Obtain and judge the speech recognition result: After the intelligent TV receives the voice command, obtain the speech recognition result of the intelligent TV; judge whether the speech recognition of the intelligent TV is correct according to the real-person speech audio file.

[0043] It should be noted that the corpus is understood as the speech and text discourse set according to the semantics of the speech that the intelligent TV to be tested usually needs to recognize, such as the text or speech of "I want to watch Legend of Chu Qiao". Correspondingly, it is understood that there is a one-to-one correspondence between the real-person speech audio file and the voice command. For example, if the content recorded in the real-person speech audio file is "I want to watch Legend of Chu Qiao", then the content and text semantics of the voice command are also "I want to watch Legend of Chu Qiao".

[0044] The artificial mouth device of the present invention is understood as a sound source device that can simulate the sound production of a human mouth. It is usually composed of a small speaker installed in a baffle or closed cavity with a specific shape, capable of simulating the sound field near the human mouth, having a sound field directivity and radiation pattern similar to that of the human mouth, and is mostly used for acoustic testing, providing a stable sound source; the frequency response range of the artificial mouth device is usually between 100Hz and 8000Hz, but some high-end products can reach 31.5kHz. Taking the AM022 model artificial mouth device as an example, it has the continuous sweep frequency capabilities of 94dB, 110dB, and 80dB, and the total harmonic distortion is less than 3%; specifically, the artificial mouth device may include: a microprocessor, an external multi-channel sound card, and an equalizer.

[0045] The real-person speech audio file of the present invention, for example, the audio file obtained by the sampled representative person speaking according to the content of the corpus, such as a wav audio file. It should be noted that in order to maintain the sound effect of the present invention, it can be recorded by using professional high-fidelity speaker equipment, and at the same time, it is recorded in a quiet environment to avoid other interference sources. Moreover, during the test process, the sound in the speech test room is less than 50dB, and the sound pressure range of the user's home environment noise is generally between 40 and 50dB when measured by a sound pressure meter. Simulating the user's home use scenario can improve the reliability of the test environment. For example, during the period from when the artificial mouth device issues a corresponding voice command to when the smart TV receives the voice command, the sound intensity in the test room is maintained <50dB.

[0046] Specifically, obtaining the speech recognition result of the smart TV specifically includes the following steps: collecting and recognizing the execution operation information of the smart TV through the serial port command between the smart TV and the test control terminal; the execution operation information includes: activation status and recognition status. The smart TV is connected to the test control terminal through a serial cable, and the test control terminal is connected to the test fixture. According to the serial port command, such as logcat|grepCH_ER_COLLECT, the execution operation information of the TV is collected.

[0047] Specifically, judging whether the speech recognition of the smart TV is correct specifically includes the following steps:

[0048] Judging whether the execution operation information contains the first keyword field; if the execution operation information contains the first keyword field, it is determined that the activation status is activated; if the execution operation information does not contain the first keyword field, it is determined that the activation status is not activated;

[0049] Identifying the second keyword field of the execution operation information; comparing the second keyword field with the text of the corresponding real-person speech audio file to judge whether the second keyword field and the text are exactly the same;

[0050] If the second keyword field is exactly the same as the text, the speech recognition is determined to be correct; if there are differences between the second keyword field and the text, the speech recognition is determined to be incorrect.

[0051] Specifically, the test control terminal collects the execution operation information of the TV according to the serial port command logcat|grepCH_ER_COLLECT. The judgment method is as follows: One is to determine whether the first keyword field is collected in the print information of the execution operation. For example, if there is a field with the keyword action=wakeUp, it means it has been activated; if not, it means it has not been activated. The second is to compare the second keyword field in the print information of the execution operation, such as the keyword field action=speechText; text=, with the text of the human speech audio file of the voice command to determine whether the speech recognition is correct. If the text of the second keyword field is exactly the same as the text of the human speech audio file, it means the speech recognition is correct; otherwise, the speech recognition is incorrect. For example, taking the human speech audio file recorded with the corpus "I want to watch Legend of Chu Qiao" as an example, the test results in Table 1 are obtained.

[0052] Table 1: Test results of the human speech audio file "I want to watch Legend of Chu Qiao"

[0053] Category Activation collection result Recognition collection result NO1 Voice activation recognition correct There is action=wakeUp action=speechText; text=I want to watch The Princess Weiyoung NO2 Voice not activated and not recognized There is no action=wakeUp There is no action=speechText NO3 Voice activated but not recognized There is action=wakeUp action=speechText; text= NO4 Voice activated and recognition incorrect There is action=wakeUp action=speechText; text=Watch The Princess Weiyoung

[0054] Specifically, to determine whether the speech recognition of the smart TV is correct, the following steps are also included: Screen and classify the activation state and recognition state, and then obtain multiple groups of test results. The test results include: not activated and not recognized, activated and not recognized, activated and recognized correctly, activated and recognized incorrectly. Record the multiple groups of test results in rows and columns and generate a test report. For example, record the test results of the keywords action=speechText; text= of NO1, NO3, and NO4 in Table 1 into the test report; mark NO2 without collecting the action=wakeUp as "not activated" and record it in the test report. After the test is completed, a test report is automatically generated. Through the MESSAGE instruction of the automated test monkey tool, record the test results of each round row by row and column, and automatically generate a test report file. The test report file can meet the needs of development for problem analysis and test recording.

[0055] In the present invention, the collection and judgment of the speech recognition results are completed, and this round of test can be ended, and the initial state can be restored for the next round of test. Different test corpora can be designed according to the use cases for continuous loop testing and result recording. Through the automated test tool instructions of the test tooling, hundreds or thousands of loop stress tests can be set.

[0056] Specifically, activating the smart TV to receive voice commands specifically includes the following steps: For far-field voice, wake up the smart TV through the intelligent voice assistant, and then activate the smart TV to receive voice commands; for near-field voice, activate the smart TV to receive voice commands through the control remote control. For example, for far-field voice, activate the voice recognition function of the smart TV by shouting the wake-up word of Changhong Xiaobai. For near-field voice, after pairing the smart TV with the Bluetooth remote control, control the test tooling through the test control terminal. It should be noted that the test control terminal is understood as a computer terminal with control and execution functions. For example, for a test computer, activate the voice by pressing the voice key on the remote control and controlling the smart TV through Bluetooth. Specifically, within six seconds after the voice is activated, the test tooling controls the artificial mouth through the data cable to play the wav file of the real person's speech audio, and then issues the voice command "I want to watch ***", and the smart TV under test receives the voice command.

[0057] Embodiment 2:

[0058] Reference Figure 2 As shown, on the basis of Embodiment 1, this embodiment provides a test system for smart TV voice recognition, which is used to implement the test method for smart TV voice recognition based on real person's broadcast in Embodiment 1, including: audio equipment, smart TV, test control terminal, test tooling, and artificial mouth equipment; the audio equipment is used to record the real person's speech audio file; the smart TV is used to recognize voice commands; the test control terminal is used to control the artificial mouth equipment to broadcast and control the test tooling to obtain the voice recognition result; the test tooling is used to cooperate with the smart TV to recognize voice commands and to collect execution operation information; the artificial mouth equipment is used to broadcast the real person's speech audio file as the corresponding voice command.

[0059] It should be noted that the test tooling is understood as a device that communicates and is electrically connected to the smart TV and the test control terminal respectively; specifically, the test tooling includes: a tooling remote control, a relay, a routing device, and a control module; the tooling remote control is used to cooperate with the control module to control the test tooling; the relay is used to respond to the operation instructions of the control module; the routing device is used to provide a network connection; the control module is used to control the remote control, relay, and routing device to recognize voice commands and collect execution operation information, and send the voice recognition result to the test control terminal through the data cable.

[0060] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A test method for smart TV voice recognition based on real person broadcast, characterized in that: The following steps are involved: S1. Recording audio files of real people speaking: Select representative people from men, women, old people and children, and record audio files of real people speaking according to the corpus of the test case; the corpus includes at least one phrase; S2, broadcasting the audio file of real people speaking as voice commands: activating the smart TV to receive the voice commands; Controlling the artificial mouth device to broadcast the real person speech audio file, so that the artificial mouth device issues a corresponding voice command; S3. Obtain and determine the voice recognition result: after the smart TV receives the voice command, obtain the voice recognition result of the smart TV; and determine whether the voice recognition of the smart TV is correct based on the real person speaking audio file.

2. The test method for intelligent TV speech recognition based on real person broadcast according to claim 1 is characterized in that: Obtaining the voice recognition result of the smart TV specifically includes the following steps: collecting and identifying the execution operation information of the smart TV through the serial port command between the smart TV and the test control terminal; the execution operation information includes: activation state and recognition state.

3. The test method for intelligent TV voice recognition based on real person broadcast according to claim 2 is characterized in that: Determining whether the voice recognition of the smart TV is correct specifically includes the following steps: Determine whether the execution operation information includes a first key field; if the execution operation information includes the first key field, determine that the activation state is activated; if the execution operation information does not include the first key field, determine that the activation state is inactivated; Identify the second key field of the execution operation information; compare the second key field with the text of the corresponding real person speaking audio file to determine whether the second key field and the text are completely consistent; If the second key field and the text are completely consistent, it is determined that the speech recognition is correct; if the second key field and the text are different, it is determined that the speech recognition is wrong.

4. The test method for intelligent TV voice recognition based on real person broadcast according to claim 2, characterized in that: Judging whether the voice recognition of the smart TV is correct also includes the following steps: screening and classifying the activation status and recognition status to obtain multiple groups of test results; the test results include: not activated and not recognized, activated and not recognized, activated and recognized correctly, activated and recognized incorrectly, recording the multiple groups of test results in rows and columns, and generating a test report.

5. The method for testing smart TV speech recognition based on real person broadcast according to any one of claims 1 to 4, characterized in that: Activating the smart TV to receive voice commands specifically includes the following steps: for far-field voice, the smart TV is awakened by the smart voice assistant, and then the smart TV is activated to receive voice commands; for near-field voice, the smart TV is activated to receive voice commands by controlling the remote control.

6. The method for testing smart TV speech recognition based on real person broadcast according to any one of claims 1 to 4, characterized in that: The number of the corpus is 30 to 500.

7. The method for testing smart TV speech recognition based on real person broadcast according to any one of claims 1 to 4, characterized in that: During the period from when the artificial mouth device issues a corresponding voice command to when the smart TV receives the voice command, the sound intensity in the test room is maintained at <50dB.

8. A test system for smart TV speech recognition, characterized in that: A test method for implementing smart TV voice recognition based on real-person broadcast as described in any one of claims 1 to 7, comprising: an audio device, a smart TV, a test control terminal, a test tool and an artificial mouth device; the audio device is used to record the real-person speech audio file; the smart TV is used to recognize the voice command; the test control terminal is used to control the artificial mouth device to broadcast and control the test tool to obtain the voice recognition result; the test tool is used to cooperate with the smart TV to recognize voice commands and to collect execution operation information; the artificial mouth device is used to broadcast the real-person speech audio file as a corresponding voice command.

9. The test system for intelligent TV speech recognition according to claim 8, characterized in that: The test tool includes: a tool remote control, a relay, a routing device and a control module; the tool remote control is used to cooperate with the control module to control the test tool; the relay is used to respond to the operation instructions of the control module; the routing device is used to provide network connection; the control module is used to control the remote control, relay and routing device to recognize voice instructions and collect execution operation information, and send the voice recognition results to the test control terminal through a data line.