A test system and method for voice wake-up and recognition
The voice wake-up and recognition testing system, which utilizes the image and audio recognition of a rotating lifting platform and camera equipment, solves the problem of inaccurate judgment of false wake-up of the device under test in the existing technology, realizes non-destructive testing and reduces false recognition, and provides a tool for adjusting device performance.
Patent Information
- Application Number
- CN202311044598.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-18
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2043-08-18
AI Technical Summary
Existing speech recognition testing methods cannot accurately determine whether the device under test has been falsely woken up, and require damaging the product under test for testing. They cannot confirm false wake-ups caused by misrecognition, and the test data cannot reflect the performance differences of the device at different positions and distances.
A voice wake-up and recognition test system was designed, including a playback device, an execution device, a status detection device, and a rotating lifting platform. The angle and height of the device are adjusted by rotating the lifting platform, and the wake-up status and recognition result of the device under test are determined by combining the image recognition and audio recognition of the camera device.
Testing can be performed without damaging the device under test, reducing the probability of false identification, saving labor costs, and generating 2D recognition maps to help engineers adjust product performance and identify problems.
Smart Images

Figure CN117292676B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice wake-up recognition testing technology, and in particular to a test system and method for voice wake-up and recognition. Background Technology
[0002] Smart home products are rapidly evolving, and voice recognition is an important component of human-computer interaction in smart home products. The quality of voice recognition directly affects the success rate, latency, and user experience of the product.
[0003] Smart home products with voice recognition capabilities require testing before mass production or shipment. Existing methods typically confirm wake-up status by printing logs from the device under test (DUT). This approach cannot definitively rule out false wake-ups due to misidentification and often requires software modifications or additional wiring from the DUT. Another existing testing method utilizes the DUT's voice module to recognize wake words. Once recognized, it sends a command to the control device, instructing it to play the command. The drawbacks of this method are that false wake-ups are recorded, and if the DUT fails to activate the wake-up function, it's impossible to determine whether the wake-up was recognized but not activated, or whether the failure to activate the wake-up function was due to the absence of the wake-up word.
[0004] Since the voice recognition of the device under test is not limited to one direction, but varies depending on the placement of the device and the distance and orientation of the user, the test data is especially important for engineers to understand product performance and locate product problems. Summary of the Invention
[0005] This invention addresses the technical problems of current testing methods that require damaging the product under test and cannot confirm false wake-ups during testing by designing a voice wake-up and recognition testing system and method.
[0006] This invention provides a voice wake-up and recognition testing system for testing devices under test. The system includes: a playback device, a server, an execution device, a status detection device, and a rotating lifting platform. The playback device is communicatively connected to the server and is used to play a wake-up voice message sent by the server. The wake-up voice message includes a wake-up command and a voice command. The execution device is connected to the device under test (DUT). The DUT recognizes the wake-up voice and obtains the instruction corresponding to the wake-up voice. Then, the DUT controls the execution device to execute the instruction corresponding to the wake-up voice. The server is communicatively connected to the execution device, and the execution device executes the instruction corresponding to the wake-up voice and then uploads the execution result to the server; The server is used to determine whether the execution result matches the sent wake-up voice, and to record the determination result; The status detection device is communicatively connected to the server and is used to detect the working status of the device under test and the execution device; The rotating lifting platform is used to place the device under test, and the rotating lifting platform is used to control the device under test to be at different angles and / or different heights.
[0007] Furthermore, the status detection device is a camera device; The camera device is used to record video of the device under test and the execution device; The camera device is used to record the voice generated by the device under test and the surrounding environment.
[0008] Furthermore, the rotary lifting platform includes: a microcontroller, a rotary motor, a transmission device, a loading platform, and a lifting motor; The microcontroller is connected to the rotary motor and is used to control the rotary motor to rotate. The microcontroller is connected to the lifting motor and is used to control the lifting motor to move up and down. The rotary motor and the lifting motor are connected to the platform via a transmission device, which are used to drive the platform to rotate at a certain angle and to lift at a certain height. Furthermore, the rotating lifting platform includes: a wireless communication module; The wireless communication module is connected to the microcontroller and communicates with the server to upload the lifting height and rotation angle of the rotating platform to the server.
[0009] This invention provides a voice wake-up recognition testing method, comprising: The server sends a rotation signal to the rotating platform, the rotating platform rotates at a fixed angle, and uploads the rotation angle to the server; The server controls the playback device to play the wake-up voice. The device under test identifies the wake-up voice, controls the execution device to execute the instruction corresponding to the wake-up voice, and obtains the execution result. The execution result is uploaded to the server via the execution device; The server determines whether the execution result matches the voice command; if they do, the recognition is considered successful.
[0010] Furthermore, the step of rotating the lifting platform by a fixed angle and uploading the rotation angle to the server includes: After rotating a full circle, the height is raised or lowered by a fixed amount via the rotating lifting platform, and the height is then uploaded to the server.
[0011] Furthermore, controlling the playback device to play the wake-up voice via the server includes: The wake-up command is played through the playback device, and the voice command is played after a first time interval; The playback device plays the wake-up command and voice command while simultaneously playing ambient noise.
[0012] Furthermore, the step of receiving and recognizing the wake-up voice through the device under test, controlling the execution device to execute the instruction corresponding to the wake-up voice, and obtaining the execution result includes: If the device under test does not receive the wake-up command, the execution device cannot execute; If the device under test receives the wake-up command but does not recognize the voice command, the execution device cannot execute it; If the device under test receives the wake-up command or recognizes the voice command incorrectly, the execution device will be unable to execute. If the device under test receives the wake-up command and recognizes the voice command, it controls the execution device to execute the corresponding voice command.
[0013] Furthermore, the step of uploading the execution result to the server via the execution device and receiving the execution result via the server includes: If the server does not receive the execution result within the second time period, it retrieves the video captured by the camera device and extracts the device status image at the corresponding time. Perform image recognition on the device status image to determine whether the device under test has been started; if so, confirm that the device under test has been woken up. If the server does not receive the execution result within the second time period, it retrieves the audio recorded by the camera device and extracts the audio segment of the corresponding time period. The audio segment is subjected to audio recognition to determine whether the device under test has been started; if so, it is confirmed that the device under test has been woken up.
[0014] Furthermore, the step of determining whether the execution result and the voice command match through the server, and determining successful recognition if they match, includes: Record the successful recognition results from different angles and heights; A coordinate system is established based on different angles and heights. The successful recognition results are used to generate recognition two-dimensional maps corresponding to different heights of the device under test. The recognition two-dimensional maps include the test recognition rate of the device under test in different directions.
[0015] The present invention has the following advantages: First, an execution device is added to the backend of the device under test (DUT). This eliminates the need to print logs from the DUT before identification and judgment, allowing the test environment to be set up without damaging the DUT. Any software or hardware modifications can be made only in the execution device without affecting the accuracy of the DUT. The execution device will only be controlled when the DUT is woken up and correctly identified. The execution device can be used to determine whether the recognized voice commands are correct, reducing the probability of misrecognition.
[0016] Secondly, by adding video equipment to record the testing process, when the server does not receive the execution result, the status of the device under test and the execution device can be further judged through image recognition and audio recognition. This can help testers determine whether the device has not been woken up, has an identification error, or is a malfunction of other machines, thus saving labor costs.
[0017] Third, by adding a rotating lifting platform to place the device under test, the device under test can be placed at different heights and angles for wake-up voice testing. The angle and height feedback from the rotating lifting platform can be used to generate a recognition two-dimensional map, which makes it easier for engineers to adjust product performance, locate product problems, and can also be used to indicate the best installation position of the device in the instruction manual. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the overall system structure of the present invention.
[0020] Figure 2 This is a flowchart of the method of the present invention.
[0021] Figure 3 This is a two-dimensional recognition image of the device under test in this invention at a height of 1 meter.
[0022] Figure 4 This is a two-dimensional recognition image of the device under test in this invention at a height of 1.5 meters. Detailed Implementation
[0023] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.
[0025] like Figure 1 As shown, the present invention provides a voice wake-up and recognition testing system for testing devices under test. The system includes: a playback device, a server, an execution device, a status detection device, and a rotating lifting platform. The playback device is communicatively connected to the server and is used to play a wake-up voice message sent by the server. The wake-up voice message includes a wake-up command and a voice command. In this embodiment, a high-fidelity speaker can be used as the standard player to play wake-up commands and voice commands (smart home products can be set with one wake-up word and multiple command words; the wake-up word is used to wake up the device under test, and the command words are used to make the device under test perform corresponding actions, such as "Xiao Li Butler, turn on the light," where "Xiao Li Butler" is the wake-up word and "turn on the light" is the command word). The playback device is connected to the server, and the wake-up commands and voice commands are stored on the server. Different home products have different wake-up voices and voice commands, and a corresponding command database can be established. According to the actual product requirements, the corresponding wake-up voices and voice commands can be retrieved. The server can also adjust the volume and words of the wake-up voices and voice commands played by the playback device. The execution device is connected to the device under test (DUT). The DUT recognizes the wake-up voice and obtains the corresponding instruction. Then, the DUT controls the execution device to execute the instruction corresponding to the wake-up voice. The server is communicatively connected to the execution device. After the execution device executes the instruction corresponding to the wake-up voice, it uploads the execution result to the server. The server is used to determine whether the execution result matches the sent wake-up voice and records the determination result. In this embodiment, the execution device can be a switch panel. The purpose of the execution device is to separate it from the device under test (DUT), making the test environment easier to set up. If data were directly transmitted from the DUT to the server, it would often require damaging the DUT, affecting the accuracy of the test. The execution device can also determine whether the identified command words are correct, reducing the probability of misidentification. The server then judges and records the execution results uploaded by the execution device.
[0026] The status detection device is communicatively connected to the server and is used to detect the working status of the device under test and the execution device.
[0027] In this embodiment, detecting the operating status of the device under test (DUT) and the execution device may specifically include: recording video of the DUT and the execution device, and recording audio generated by the DUT and the surrounding environment. In other embodiments, other status detection methods may also be used, such as detecting the operating status of the DUT and / or the execution device, such as displacement and shape changes, using sensors or other electronic devices.
[0028] In this embodiment, the status detection device is a camera device. In other embodiments, the status detection device may also be other electronic devices with video recording and audio recording functions.
[0029] In a preferred embodiment, the camera device records and monitors the entire process of the device under test (DUT) and the executing device in real time. It connects to the server and takes screenshots for image recognition when no voice command is recognized. The state of the DUT before and after wake-up is different, facilitating subsequent data analysis. Some home appliances will generate light prompts after being woken up. The state graph of the DUT can be used to determine whether it has been woken up. Furthermore, the state graph of the executing device (switch panel) can be used to confirm whether the DUT has correctly recognized the voice command through image recognition of the switch panel's state. For example, when the playback device plays "Xiao Li Butler, turn on the light," and the executing device does not upload an execution result, the server controls the camera device to take a screenshot. Based on the states of the DUT and the executing device, it can be determined whether the device was woken up but not recognized, or neither was woken up nor recognized.
[0030] For some home appliances that do not have light indicators but can interact with humans and respond to voice commands, the audio generated by the device under test (DUT) can be recorded by a camera to confirm whether the DUT has been woken up. Furthermore, by analyzing the audio segments recorded by the camera after the DUT receives the voice command, it can be confirmed whether the DUT has made a correct voice response, thus confirming whether the DUT has correctly recognized the voice command.
[0031] In summary, image recognition and audio recognition, or other recognition methods, can be combined. For example, when audio recognition cannot confirm whether the device under test has correctly recognized the voice command, the state diagram of the device can be used to confirm whether the device has performed an action, such as a switch change.
[0032] The rotating lifting platform is used to place the device under test, and the rotating lifting platform is used to control the device under test to be at different angles and / or different heights.
[0033] Furthermore, the rotary lifting platform includes: a microcontroller, a rotary motor, a transmission device, a loading platform, a lifting motor, and a wireless communication module; The microcontroller is connected to the rotary motor and is used to control the rotary motor to rotate. The microcontroller is connected to the lifting motor and is used to control the lifting motor to move up and down. The rotary motor and the lifting motor are connected to the platform via a transmission device, which are used to drive the platform to rotate at a certain angle and to lift at a certain height. The wireless communication module is connected to the microcontroller and communicates with the server to upload the lifting height and rotation angle of the rotating platform to the server.
[0034] In this embodiment, the operator can manually adjust the angle and height of the rotating lifting platform with each rotation. The rotating lifting platform automatically records these adjustments and sends the data to the server during each rotation and lift. The device is placed in an 86-box container, then inside a cement shell, to simulate a real-world user scenario. The device under test is then placed on the rotating lifting platform.
[0035] like Figure 2 As shown, the present invention provides a testing method for voice wake-up and recognition, comprising: S1 The server sends a rotation signal to the rotating lifting platform, the rotating lifting platform rotates at a fixed angle, and uploads the rotation angle to the server; After S101 rotates back to the starting point, it raises or lowers the fixed height via the rotating lifting platform and uploads the raised or lowered height to the server.
[0036] In this step, at the starting point, after one round of testing, a rotation signal is sent by the server, and the rotating platform rotates by a fixed angle (e.g., 1°). After the next round of testing is fully executed, the server sends another rotation signal, and the rotating platform rotates another 1° to begin the next round of testing, repeating this cycle. Since the testing time for different devices is not fixed, the server needs to send rotation signals. This facilitates the collection of test data from different angles of the device under test, and the testing system can be matched to different devices under test without manually setting the rotation time interval of the rotating platform. When the platform returns to the starting point, it rises or falls by a fixed height (this height can be set according to requirements), and the rotation test continues as described above.
[0037] S2 controls the playback device to play the wake-up voice through the server; S201 plays the wake-up command through the playback device, and plays the voice command after a first time interval; S202 plays the wake-up command and voice command through the playback device while simultaneously playing ambient noise.
[0038] In this step, in order to reproduce the actual wake-up recognition process of the device under test, a wake-up command needs to be issued first, followed by a voice command. According to the test requirements, noise is added to the environment. The noise addition method can be a fixed audio or news broadcast, or a silent environment, which can further test the recognition ability of the device under test in a noisy environment.
[0039] S3 receives and recognizes the wake-up voice through the device under test, controls the execution device to execute the instruction corresponding to the wake-up voice, and obtains the execution result; S301 If the device under test does not receive the wake-up command, the execution device cannot execute; S302 If the device under test receives the wake-up command but does not recognize the voice command, the execution device cannot execute; S303 If the device under test receives the wake-up command or recognizes the voice command incorrectly, the execution device cannot execute; S304 If the device under test receives the wake-up command and recognizes the voice command, it controls the execution device to execute the corresponding voice command.
[0040] In this step, we'll use "Xiao Li Butler, turn on the light" as an example. The server controls the playback device to play the wake-up voice. When the device under test receives the wake-up voice, it will transition from a waiting-to-wake state to a wake-up state. If the device under test receives both the wake-up command and the voice command, it will control the execution device (switch panel) to execute the "turn on the light" command. The switch panel will open, and the execution device will synchronously upload the opening signal to the server for judgment and recording. If the device under test only receives the wake-up command but not the command word, or if the command word is misrecognized, the execution device will not be able to execute the relevant command. The execution device will only execute the command if the device under test receives both the wake-up command and the voice command is correctly recognized. The execution device can determine whether the recognized voice command is correct, reducing the probability of misrecognition.
[0041] S4 uploads the execution result to the server via the execution device; S401 If the server does not receive the execution result within the second time period, it retrieves the video captured by the camera device and extracts the device status image at the corresponding time. S402 performs image recognition on the device status image to determine whether the device under test has been started; if so, it confirms that the device under test has been woken up. S403 If the server does not receive the execution result within the second time period, then retrieve the video captured by the camera device and extract the video and audio at the corresponding time. S404 performs audio recognition on the video and audio to determine whether the device under test has been started. If so, it confirms that the device under test has been woken up.
[0042] In this step, if the device under test (DUT) only receives the wake-up command but not the command word, or if the command word is misrecognized, the executing device will be unable to execute the relevant commands, and no execution result will be uploaded to the server. Taking "Xiao Li Butler, turn on the light" as an example, the server will typically receive the execution result within 5-8 seconds during testing. If the wake-up word and command word are longer, the waiting time will be even longer. If no execution result is received within this timeframe, it is considered that there is a problem with the testing process. In this case, video footage captured by the camera during this timeframe can be retrieved, and the status diagrams of the DUT and executing devices or the audio from this timeframe can be extracted to confirm the situation.
[0043] Some home appliances will generate light prompts after being woken up. At this time, the status graph of the device under test can be used to determine whether the device under test has been woken up. Furthermore, the status graph of the execution device (switch panel) can be used to confirm whether the device under test has correctly recognized the voice command by recognizing the status of the switch panel through image recognition.
[0044] Some home appliances do not have light indicators, but they can interact with humans and respond to voice commands. Therefore, the audio recorded by the camera can be used to confirm whether the device under test has been woken up. Furthermore, by analyzing the audio segments recorded by the camera after the device under test receives the voice command, it can be confirmed whether the device under test has made the correct voice response, thus confirming whether the device under test has correctly recognized the voice command.
[0045] Image recognition and audio recognition can be combined. For example, when audio recognition cannot confirm whether the device under test has correctly recognized the voice command, the state diagram of the device can be used to confirm whether the device has performed an action, such as a switch change.
[0046] S5 determines whether the execution result and the voice command match through the server; if they match, the recognition is considered successful.
[0047] S501 records the execution results of successful recognition at different angles and heights; S502 establishes a coordinate system based on different angles and heights, and uses the successful recognition results to generate recognition two-dimensional maps corresponding to different heights of the device under test.
[0048] In this step, after all angles at different heights of the device under test have been tested, a two-dimensional recognition map for each height is generated based on the successful recognition records at different heights and angles. This two-dimensional recognition map provides a visual representation of the recognition rate in different directions at the same height. Figure 3 and Figure 4 As shown, due to the intuitiveness of two-dimensional drawings, recognition of two-dimensional drawings at different heights is achieved by generating them. Figure 3 It is a two-dimensional image recognition at a height of 1 meter. Figure 4 This is a 2D recognition diagram at a height of 1.5 meters. As can be seen from the diagram, recognition is better at a frontal angle (0 degrees) and worse at a back-facing angle (180 degrees). This is mainly due to factors such as the device's structure, microphone position, microphone selection, and algorithm. The better performance at 1.5 meters compared to 1 meter may be because the device under test and the playback device are at a relatively horizontal position, making it easier for the device under test to pick up audio. This makes it easier for engineers to adjust product performance, locate product problems, and can also be used in the instruction manual to indicate the optimal installation position for the device.
[0049] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0050] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0051] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0052] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0053] It should be noted that any reference signs placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0054] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention.
[0055] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
[0056] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0057] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms should not be construed as necessarily referring to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
Claims
1. A testing method for voice wake-up and recognition, characterized in that, According to a voice wake-up and recognition test system, the system is used to test the device under test, including: a playback device, a server, an execution device, a status detection device, and a rotating lifting platform, wherein the status detection device is a camera device; The method includes: The server sends a rotation signal to the rotating platform, the rotating platform rotates at a fixed angle, and uploads the rotation angle to the server; The server controls the playback device to play the wake-up voice message. The device under test identifies the wake-up voice, controls the execution device to execute the instruction corresponding to the wake-up voice, and obtains the execution result. Uploading the execution result to the server via the execution device includes: If the server does not receive the execution result within the second time period, it retrieves the video captured by the camera device and extracts the device status image at the corresponding time. Perform image recognition on the device status image to determine whether the device under test has been started; if so, confirm that the device under test has been woken up. If the server does not receive the execution result within the second time period, it retrieves the audio recorded by the camera device and extracts the audio segment of the corresponding time period. The audio segment is subjected to audio recognition to determine whether the device under test has been started; if so, it is confirmed that the device under test has been woken up. The server determines whether the execution result matches the voice command; if they do, the recognition is considered successful.
2. The test method for voice wake-up and recognition according to claim 1, characterized in that, The step of rotating the platform by a fixed angle and uploading the rotation angle to the server includes: After rotating a full circle, the height is raised or lowered by a fixed amount via the rotating lifting platform, and the height is then uploaded to the server.
3. The test method for voice wake-up and recognition according to claim 1, characterized in that, The step of controlling the playback device to play the wake-up voice through the server includes: The wake-up command is played through the playback device, and the voice command is played after a first time interval. The playback device plays wake-up commands and voice commands while simultaneously playing ambient noise.
4. The test method for voice wake-up and recognition according to claim 3, characterized in that, The step of receiving and recognizing the wake-up voice through the device under test, controlling the execution device to execute the instruction corresponding to the wake-up voice, and obtaining the execution result includes: If the device under test does not receive the wake-up command, the execution device cannot execute; If the device under test receives the wake-up command but does not recognize the voice command, the execution device cannot execute it; If the device under test receives the wake-up command or recognizes the voice command incorrectly, the execution device will be unable to execute. If the device under test receives the wake-up command and recognizes the voice command, it controls the execution device to execute the corresponding voice command.
5. The test method for voice wake-up and recognition according to claim 1, characterized in that, The step of determining whether the execution result and the voice command match through the server, and determining successful recognition if they match, also includes: Record the successful recognition results from different angles and heights; A coordinate system is established based on different angles and heights. The successful recognition results are used to generate recognition two-dimensional maps corresponding to different heights of the device under test. The recognition two-dimensional maps include the test recognition rate of the device under test in different directions.
6. A test system for voice wake-up and recognition, characterized in that, A test method for implementing a voice wake-up and recognition method according to any one of claims 1-5 includes: The playback device is communicatively connected to the server and is used to play a wake-up voice message sent by the server. The wake-up voice message includes a wake-up command and a voice command. The execution device is connected to the device under test (DUT), the DUT recognizes the wake-up voice to obtain the instruction corresponding to the wake-up voice, and the DUT controls the execution device to execute the instruction corresponding to the wake-up voice. The server is communicatively connected to the execution device, and the execution device executes the instruction corresponding to the wake-up voice and then uploads the execution result to the server; The server is used to determine whether the execution result matches the sent wake-up voice, and to record the determination result; The status detection device is communicatively connected to the server and is used to detect the working status of the device under test and the execution device; The rotating lifting platform is used to place the device under test, and the rotating lifting platform is used to control the device under test to be at different angles and / or different heights.
7. The test system for voice wake-up and recognition according to claim 6, characterized in that, The camera device is used to record video of the device under test and the execution device; The camera device is used to record the voice generated by the device under test and the surrounding environment.
8. The test system for voice wake-up and recognition according to claim 6, characterized in that, The rotary lifting platform includes: a microcontroller, a rotary motor, a transmission device, a loading platform, and a lifting motor; The microcontroller is connected to the rotary motor and is used to control the rotation of the rotary motor; The microcontroller is connected to the lifting motor and is used to control the lifting motor to lift and lower. The rotary motor and the lifting motor are connected to the platform via a transmission device, which drives the platform to rotate at a certain angle and to lift at a certain height.
9. The test system for voice wake-up and recognition according to claim 8, characterized in that, The rotating lifting platform includes: a wireless communication module; The wireless communication module is connected to the microcontroller and communicates with the server to upload the lifting height and rotation angle of the rotating platform to the server.
Citation Information
Patent Citations
Speech recognition terminal evaluation system and method
CN110211567A
Voice wake-up and recognition automatic test method, storage medium and test terminal
CN112151029A