Voice equipment testing device, voice equipment testing method and storage medium
By designing an automated voice equipment test device, the problems of inefficiency of traditional test methods and inaccurate results are solved, and efficient and accurate evaluation is achieved in different environments, which is suitable for testing of various voice equipment.
Patent Information
- Application Number
- CN202510565247.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional speech recognition testing methods rely on manual or semi-automated operations, resulting in cumbersome testing processes and inefficient efficiency, making it difficult to simulate the real environment, and the accuracy and consistency of results are difficult to ensure.
A voice equipment testing device is designed, including a mobile module, a pronunciation module and a radio module. The control module automatically controls the pronunciation module to move in the test space, plays voice commands that simulate human voice and background noise, and receives feedback voice through the radio module. The integrated module realizes automated evaluation, which is suitable for various environments and equipment.
It improves testing efficiency, enhances testing flexibility and adaptability, can more accurately evaluate the device's voice processing capabilities in different environments, reduces manual intervention, and the test results are closer to actual use.
Smart Images

Figure CN120356459A_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the field of speech testing technologies, and particularly relates to a speech device testing apparatus, a speech device testing method, and a storage medium. Background Art
[0002] Speech interaction and recognition technologies are developing rapidly and have become an important part of fields such as smart homes, virtual assistants, and in-vehicle systems. With the expansion of these application scenarios, speech recognition systems not only need to provide efficient and accurate speech input processing capabilities in diverse environments, but also need to maintain good recognition performance in the presence of noise interference, at different orientations, and at different distances. However, existing speech recognition system testing methods still have some limitations.
[0003] Traditional speech recognition testing methods usually rely on manual or semi-automated operations, which means that the testing process is cumbersome and error-prone. For example, testers need to manually set up the test environment, play test corpora one by one, and record and analyze the test results. This approach not only takes time, but also makes it difficult to ensure the accuracy and consistency of test results due to human factors. In addition, traditional testing methods are difficult to simulate various complex scenarios in a real environment, thus limiting the comprehensive evaluation of the performance of speech recognition systems. Moreover, each functional module operates independently and cannot form a unified automated testing process. This scattered testing method leads to low testing efficiency, fragmented data processing, and difficulties in ensuring the accuracy and consistency of test results. Summary of the Invention
[0004] The following is an overview of the subject matter described in detail in this document. This overview is not intended to limit the scope of protection of the claims.
[0005] Embodiments of this application provide a speech device testing apparatus, a speech device testing method, and a storage medium, which are applicable to various different types of devices under test and test scenarios, can simulate the influence of various environments on the devices under test, and help evaluate the performance of the devices in different environments.
[0006] To achieve the above object, a first aspect of the embodiments of the present application provides a voice device testing apparatus, which is characterized by including: a moving module for driving the voice device testing apparatus to move within a preset testing space; a pronunciation module disposed on the moving module, the pronunciation module is used to play a voice command simulating a human voice, so that the device under test emits a feedback voice according to the voice command, wherein the voice command includes a human voice recording and background noise; a sound receiving module disposed on the moving module, the sound receiving module is used to receive the feedback voice; a control module, the moving module, the pronunciation module and the sound receiving module are respectively electrically connected to the control module. During the process that the control module controls the moving module to drive the voice device testing apparatus to move, the control module controls the pronunciation module to play the voice command, and receives the feedback voice of the device under test through the sound receiving module, and evaluates the voice processing ability of the device under test according to the feedback voice.
[0007] In some embodiments, an artificial head structure is disposed on the moving module. The pronunciation module includes a mouth simulator disposed on the mouth of the artificial head structure, and the sound receiving module includes an ear simulator disposed on the ear of the artificial head structure.
[0008] In some embodiments, an attitude adjustment component is disposed at the bottom of the artificial head structure. The attitude adjustment component is used to adjust the attitude of the artificial head structure to control the height and orientation of the mouth simulator and the ear simulator.
[0009] In some embodiments, a sound pressure calibration component is further disposed on the pronunciation module. Before the pronunciation module plays the voice command, the control module calibrates the sound pressure of the pronunciation module through the sound pressure calibration component.
[0010] In some embodiments, the voice device testing apparatus further includes a device camera module for acquiring an operation image of the device under test after receiving the voice command. The control module identifies the mechanical state of the device under test according to the operation image of the device under test to evaluate the voice processing ability of the device under test.
[0011] In some embodiments, the voice device testing apparatus further includes a mechanical fixing component and a mechanical sensing component. Both the mechanical fixing component and the mechanical sensing component are disposed on the moving module. The mechanical fixing component is used to fix the device under test, and the mechanical sensing component is used to detect the mechanical state of the device under test.
[0012] In some embodiments, a capability assertion model is provided in the control module. After the control module obtains the feedback voice of the device under test through the sound collection module, or obtains the mechanical state of the device under test when receiving the voice command through the mechanical sensing component, the capability assertion model discriminates the voice processing capability of the device under test based on the feedback voice and the mechanical state, and generates a voice capability report of the device under test.
[0013] In some embodiments, a radar component and an environmental camera component are provided on the mobile module. The mobile module scans the test space where the voice device testing apparatus is located through the radar component, and the mobile module also takes pictures of the test space where the voice device testing apparatus is located through the environmental camera component. The mobile module performs autonomous navigation and obstacle avoidance in the test space based on the environmental information obtained by the radar component and the environmental camera component.
[0014] To achieve the above object, a second aspect of the present application proposes a voice device testing method. The voice device testing method is applied to the voice device testing apparatus described in the first aspect. The voice device testing method includes: controlling the voice device testing apparatus to move between preset test points in the test space; modulating the human voice recording and background noise in the voice command, and playing the modulated voice command through the pronunciation module; obtaining the feedback voice of the device under test after receiving the voice command through the sound collection module, and identifying and processing the feedback voice to obtain the voice processing capability of the device under test.
[0015] To achieve the above object, a third aspect of the present application proposes a computer-readable storage medium. The computer-readable storage medium includes a stored computer program; wherein, the computer program controls the device where the computer-readable storage medium is located to execute the voice device testing method described in the second aspect when running.
[0016] According to the solution provided by the embodiments of the present application, by setting a moving module and arranging the pronunciation module and the sound collection module on the moving module, the moving module drives the pronunciation module and the sound collection module to move to preset positions in the test space, so as to be able to simulate the voice processing process of the device under test when the controller issues voice commands to the device under test at different positions in the actual application scenario where the device under test is fixed. By issuing voice commands to the device under test at different positions, the voice processing capabilities of the device under test for voice commands in different directions, at different distances, and with different received volumes can be tested, which helps to simulate various sound environments that the device under test may encounter in the real world and makes the test results closer to the actual usage situation. The voice commands played by the pronunciation module include human voice recordings and background noise. For different test requirements, the control module can adjust the human voice recordings and background noise, so that the sound signals received by the device under test are more in line with the sounds issued by the controller in the actual scenario, which helps to simulate various sound environments that the device under test may encounter in the real world and makes the test results closer to the actual usage situation. The control module can automatically control the pronunciation module to play sound commands and receive feedback voices through the sound collection module, realizing an automated evaluation of the voice processing capabilities of the device under test. The integrated moving, pronunciation, and sound collection modules make the test process without manual intervention, reducing the workload of testers, while accelerating the test speed and improving the efficiency of the test process. Moreover, the control module can adjust the moving path of the moving module and the sound commands played by the pronunciation module according to different test requirements, making it applicable to various different types of devices under test and test scenarios, enhancing the flexibility and adaptability of the test. The voice device test device proposed in the present application can move in the test space to simulate the influence of various environments on the device under test, which helps to evaluate the performance of the device in different environments.
[0017] Other features and advantages of the present application will be described in the subsequent description, and, in part, will become apparent from the description, or will be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the description, the claims, and the drawings. Brief Description of the Drawings
[0018] The drawings are used to provide a further understanding of the technical solutions of the present application, and constitute a part of the description. Together with the embodiments of the present application, they are used to explain the technical solutions of the present application, and do not constitute a limitation to the technical solutions of the present application.
[0019] Figure 1 It is an optional system block diagram of the voice device test device provided by the embodiments of the present application; Figure 2 It is an optional structural schematic diagram of the voice device test device provided by the embodiments of the present application; Figure 3An optional structural schematic diagram of the artificial head structure provided by the embodiment of the present application; Figure 4 Another optional structural schematic diagram of the voice device testing device provided by the embodiment of the present application; Figure 5 Provided by the embodiment of the present application Figure 4 The specific structural schematic diagram of position A in Figure 6 An optional structural schematic diagram of the mechanical fixing component and the mechanical sensing component provided by the embodiment of the present application; Figure 7 An optional structural schematic diagram of the voice device testing system provided by the embodiment of the present application; Figure 8 An optional structural schematic diagram of the autonomous navigation system provided by the embodiment of the present application; Figure 9 An optional flow schematic diagram of the voice device testing method provided by the embodiment of the present application; Figure 10 An optional structural schematic diagram of the testing device provided by the embodiment of this application; Figure 11 An optional hardware structural schematic diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0020] This part will describe the specific embodiments of the present application in detail. The preferred embodiments of the present application are shown in the drawings. The role of the drawings is to supplement the description of the text part of the specification, enabling people to intuitively and vividly understand each technical feature and the overall technical solution of the present application, but it cannot be understood as a limitation on the protection scope of the present application.
[0021] In the description of the present application, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc. is based on the orientation or positional relationship shown in the drawings, and it is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present application.
[0022] In the description of the present application, the meaning of several is one or more, the meaning of multiple is more than two, greater than, less than, exceeding, etc. are understood as not including the original number, and above, below, within, etc. are understood as including the original number. If there is a description of first and second, it is only for the purpose of distinguishing technical features and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or the sequence relationship of the indicated technical features.
[0023] In the description of this application, unless otherwise clearly defined, terms such as setting, installation, electrical connection, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above terms in this application in combination with the specific content of the technical solution.
[0024] Currently, traditional speech recognition testing methods usually rely on manual or semi-automated operations, which means that the testing process is cumbersome and error-prone. For example, testers need to manually set up the test environment, play test corpora one by one, and record and analyze the test results. This method not only takes time but also makes it difficult to ensure the accuracy and consistency of test results due to human factors. In addition, traditional testing methods are difficult to simulate various complex scenarios in a real environment, thus limiting the comprehensive evaluation of the performance of speech recognition systems. And each functional module operates independently, unable to form a unified automated testing process. This scattered testing method leads to low testing efficiency, fragmented data processing, and difficulty in ensuring the accuracy and consistency of test results.
[0025] To address the problems of low testing efficiency, fragmented data processing, and difficulty in ensuring the accuracy and consistency of test results, this application provides a speech device testing apparatus, a speech device testing method, and a storage medium. According to the solution provided by the embodiments of this application, it is possible to simulate the influence of various environments on the device under test, which helps to evaluate the performance of the device in different environments.
[0026] The speech device testing apparatus, speech device testing method, and storage medium provided by the embodiments of this application are specifically described through the following embodiments. First, the speech device testing apparatus in the embodiments of this application is described.
[0027] The following further elaborates on the embodiments of this application in conjunction with the accompanying drawings.
[0028] Referring to Figures 1 to 5 , an embodiment of this application provides a speech device testing apparatus, including: A moving module 100, configured to drive the speech device testing apparatus to move within a preset test space; A pronunciation module 200, disposed on the moving module 100, and the pronunciation module 200 is configured to play a voice command simulating a human voice, so that the device under test emits a feedback voice according to the voice command, where the voice command includes a human voice recording and background noise; A sound collection module 300, disposed on the moving module 100, and the sound collection module 300 is configured to receive the feedback voice; The control module 400, the movement module 100, the pronunciation module 200, and the sound receiving module 300 are electrically connected to the control module 400 respectively. During the process that the control module 400 controls the movement module 100 to drive the voice device testing apparatus to move, the control module 400 controls the pronunciation module 200 to play a sound command, and receives the feedback voice of the device under test through the sound receiving module 300, and evaluates the voice processing ability of the device under test according to the feedback voice.
[0029] It can be understood that by setting the movement module 100, and setting the pronunciation module 200 and the sound receiving module 300 on the movement module 100, the movement module 100 drives the pronunciation module 200 and the sound receiving module 300 to move at preset positions in the test space, so as to be able to simulate the voice processing process of the device under test when the controller issues a voice command to the device under test at different positions in the actual application scenario where the device under test is fixed. By issuing voice commands to the device under test at different positions, it is possible to test the voice processing ability of the device under test for voice commands in different directions, at different distances, and with different received volumes, which helps to simulate various sound environments that the device under test may encounter in the real world and makes the test results closer to the actual usage situation; the voice commands played by the pronunciation module 200 include human voice recordings and background noise. For different test requirements, the control module 400 can adjust the human voice recordings and background noise, so that the sound signal received by the device under test is more in line with the sound issued by the controller in the actual scenario, which helps to simulate various sound environments that the device under test may encounter in the real world and makes the test results closer to the actual usage situation; the control module 400 can automatically control the pronunciation module 200 to play a sound command and receive the feedback voice through the sound receiving module 300, realizing the automatic evaluation of the voice processing ability of the device under test. The integrated movement, pronunciation, and sound receiving module 300 makes the test process without manual intervention, reduces the workload of the test personnel, speeds up the test speed, and improves the efficiency of the test process. Moreover, the control module 400 can adjust the movement path of the movement module 100 and the sound command played by the pronunciation module 200 according to different test requirements, making it applicable to various different types of devices under test and test scenarios, enhancing the flexibility and adaptability of the test; the voice device testing apparatus proposed in this application can move in the test space, simulating the influence of various environments on the device under test, which helps to evaluate the performance of the device in different environments.
[0030] Refer to Figure 7As shown, the voice device test apparatus proposed in the embodiments of the present application and the device under test to be tested are co-deployed in a preset test space, which is used to simulate the actual application scenario of the device under test. The test apparatus is connected to the main server through a network. The main server issues test tasks to the test apparatus through a local area network and controls the start and stop of the test tasks. The test apparatus uploads the test data generated during the execution of the test tasks to the main server, which performs statistics, analysis, and storage.
[0031] The voice device test apparatus proposed in the embodiments of the present application can test facilities integrated with intelligent voice response modules on the market, and is not limited to intelligent home appliances such as PCs, mobile phones, tablet computers, smart air conditioners, smart range hoods, smart refrigerators, smart ovens, smart stoves, smart washing machines, smart water heaters, smart washing equipment, smart dishwashers, smart projection devices, smart TVs, smart drying racks, smart curtains, smart audio and video, smart sockets, smart speakers, smart soundboxes, smart fresh air devices, smart kitchen and bathroom equipment, smart bathroom equipment, smart floor cleaning robots, smart window cleaning robots, smart mopping robots, smart air purification devices, smart steam ovens, smart microwaves, smart kitchen water heaters, smart purifiers, smart water dispensers, smart door locks, etc.
[0032] In addition, referring to Figure 3 and Figure 5 As shown, in an embodiment of the present application, a mannequin head structure 500 is provided on the mobile module 100, and the pronunciation module 200 includes a mouth simulator 201 provided on the mouth of the mannequin head structure 500, and the sound collection module 300 includes an ear simulator 301 provided on the ear of the mannequin head structure 500.
[0033] The mannequin head structure 500 used in the embodiments of the present application is a high-precision human head model designed specifically for electroacoustic testing. It integrates a mouth simulator 201 and an ear simulator 301 with high-standard requirements, and can truly reproduce the head acoustic characteristics of ordinary adults. The mannequin head structure 500 is mainly applied to the acoustic testing of products such as telephone handsets (including mobile phones and cordless phones), headphones, audio conferencing devices, microphones, hearing aids, and hearing protectors.
[0034] In a specific embodiment, the mouth simulator 201 complies with the ITU-T Rec. P.58 standard and is a sound source with high and low frequency response capabilities. The mouth simulator 201 is built into a specially designed closed acoustic cavity and generates standardized voice signals through a high-performance speaker drive unit. This acoustic cavity design enables it to simulate and mimic the average acoustic characteristics of the human mouth, including acoustic damping, radiation, and transmission characteristics; the frequency response range of the mouth simulator 201 is from 100 Hz to 10 kHz, and it can continuously output a sound pressure of ≥110 dBSPL (in the range of 200 Hz to 2 kHz), with a distortion rate of less than 1% (in the range of 200 Hz to 10 kHz). These characteristics ensure the high fidelity and consistency of the mouth simulator 201 in voice tests.
[0035] Before playing the voice command through the mouth simulator 201, the sound pressure calibration component 700 is fixed 25 millimeters directly in front of the mouth simulator 201. The sound pressure calibration component 700 is electrically connected to the control module 400. The sound pressure calibration component 700 receives the test voice of the mouth simulator 201 and calibrates the sound pressure of the mouth simulator 201 through the control module 400.
[0036] Two ear simulators 301 that comply with the IEC 60318-4 / ITU-T Rec. P.57 Type 3.3 standard are also provided on the artificial head structure 500. They are equipped with silicone soft earlobes for truly simulating the acoustic characteristics of the human ear. The ear simulator 301 has a 1 / 2-inch microphone built-in and is connected to the microphone preamplifier through a corner adapter to ensure high-precision acoustic signal acquisition. The frequency response range of the simulated ear is from 100 Hz to 10 kHz, and there is high-frequency response consistency between the two ears (±1 dB when ≤5 kHz, ±3 dB when ≤10 kHz). This consistency is crucial for the accurate measurement of binaural acoustic signals.
[0037] The control module 400 obtains the corpus data, processes and amplifies the signal through the internal audio processing circuit, and finally outputs it through the mouth simulator 201. To ensure that the played voice command has high fidelity, the system has a multi-level sound pressure calibration and distortion control mechanism built-in. During the continuous playback of the voice command by the mouth simulator 201, the control module 400 receives the voice command through the first microphone and continuously monitors and adjusts the authenticity of the voice command.
[0038] At the bottom of the artificial head structure 500, there is an attitude adjustment component 600, which is used to adjust the attitude of the artificial head structure 500 to control the height and orientation of the mouth simulator 201 and the ear simulator 301. In a specific embodiment, the attitude adjustment component 600 includes a motor, a universal coupling, and a platform fixedly connected to the movable end of the universal coupling. The artificial head structure 500 is arranged on the platform. The motor controls the telescopic, rotational, or tilting movement of the movable rod in the universal shaft, thereby controlling the different orientations and heights of the artificial head structure 500. When the distance and orientation between the voice device testing apparatus and the device under test are fixed, by adjusting the attitude of the artificial head structure 500 through the attitude adjustment component 600, it is possible to simulate the scenario where the controller speaks voice commands in different directions, and the left and right ears receive the feedback voice from the device under test, thus improving the sensitivity of the voice device testing apparatus and expanding the testing content and testing scenarios of the voice device testing apparatus.
[0039] In addition, referring to Figure 6 As shown, in an embodiment of the present application, the voice device testing apparatus further includes a mechanical fixing component 910 and a mechanical sensing component 920. Both the mechanical fixing component 910 and the mechanical sensing component 920 are arranged on the moving module 100. The mechanical fixing component 910 is used to fix the device under test, and the mechanical sensing component 920 is used to detect the mechanical state of the device under test.
[0040] During the testing process, the device under test needs to be maintained at a specific position and angle. The mechanical fixing component 910 can accurately fix the device under test to ensure that the device under test does not displace during the testing process. The mechanical fixing component 910 can fix the device under test on one side of the pronunciation module 200 and the sound collection module 300, thereby simulating the receiving ability of the device under test for voice commands when the controller faces, is side-facing, or has their back to the device under test. The mechanical sensing module can also simulate the arm movements of the controller to perform various complex operation behaviors, such as single-clicking, double-clicking, long-pressing, sliding, zooming, rotating, pinching, etc., so that the actual situation of the interaction between the controller and the device under test can be reproduced during the testing. Moreover, the mechanical fixing component 910 can also fix the device under test at any preset angle and contact the mechanical moving parts of the device under test through the mechanical sensing component 920, thereby detecting and analyzing the mechanical reaction of the device under test based on voice commands.
[0041] When it is necessary to simulate the operation of a device under test by a human in a complex environment, if the device under test can simultaneously emit feedback voice and generate mechanical motion after receiving a voice command, the mechanical fixing module and the mechanical sensing module can cooperate with the pronunciation module 200 and the sound receiving module 300. The mechanical fixing module fixes the device under test at a preset position. While the pronunciation module 200 plays the voice command, the mechanical sensing module operates on the mechanical components of the device under test. The device under test reacts according to the voice command and the operation of the mechanical sensing module. The sound receiving module 300 receives the feedback voice played by the device under test, and the mechanical sensing module in contact with the mechanical components of the device under test analyzes the mechanical reaction of the device under test, so as to comprehensively analyze the composite reaction of the device under test to the voice command in a complex environment.
[0042] For example, when it is necessary to test the artificial intelligence platform integrated on a portable computer, the mechanical fixing module fixes the portable computer directly in front of the artificial head structure 500, so that the mouth simulator 201 faces the screen of the portable computer, thus simulating the scenario of a controller using the artificial intelligence platform. The mouth simulator 201 plays the voice command for the artificial intelligence platform, and at the same time, the mechanical sensing module operates on the keyboard, mouse, touchpad or touch screen of the portable computer according to preset requirements. During the interaction with the portable computer, the feedback voice of the artificial intelligence platform is received through the ear simulator 301, so that the voice processing ability of the artificial intelligence platform integrated on the portable computer can be analyzed and evaluated.
[0043] In addition, the mechanical fixing component 910 can also fix multiple devices under test at the same time, so as to test the voice processing abilities of multiple devices under test at the same time.
[0044] In a specific embodiment, when the mass of the device under test is greater than the bearing capacity of the mechanical fixing component 910, the voice device testing device can move to one side of the device under test. The device under test generates mechanical motion according to the voice command. The mechanical sensing component 920 contacts the component of the device under test that generates mechanical motion. By obtaining the force, stroke, motion time and speed curve of the mechanical sensing component 920, the voice processing ability and response ability of the device under test are judged; for example, assuming that the device under test is an intelligent refrigerator integrated with an intelligent voice module, before the pronunciation module 200 plays the voice command, the mechanical sensing component 920 contacts the refrigerator door of the intelligent refrigerator. The pronunciation module 200 plays the voice command of "open the refrigerator door". During the opening process of the refrigerator door, the mechanical sensing component 920 moves along with the refrigerator door. According to the movement situation of the mechanical sensing component 920 itself, the outward thrust, outward stroke, motion time and speed curve during the opening process of the refrigerator door can be obtained.
[0045] Elastic pads are provided on the surfaces of the mechanical fixing component 910 and the mechanical sensing component 920 to prevent damage to the surfaces of the mechanical fixing component 910, the mechanical sensing component 920, or the device under test.
[0046] In a specific embodiment, the first microphone is mounted 25 millimeters directly in front of the mouth simulator 201 through the mechanical fixing component 910. The sound of the mouth simulator 201 is received by the first microphone, and the sound pressure of the mouth simulator 201 is calibrated by the control module 400.
[0047] In addition, in an embodiment of the present application, a capability assertion model is provided in the control module 400. After the control module 400 obtains the feedback voice of the device under test through the sound collection module 300, or obtains the mechanical state of the device under test when receiving a voice command through the mechanical sensing component 920, the capability assertion model discriminates the voice processing ability of the device under test based on the feedback voice and the mechanical state, and generates a voice ability report of the device under test.
[0048] The capability assertion model is used to determine whether the feedback voice and the mechanical state of the device under test are correct or incorrect. After the control module 400 recognizes and accepts the feedback voice of the device under test through the sound collection module 300, the control module 400 recognizes the feedback voice through the built-in voice recognition model, extracts the text, response time, volume, and clarity of the feedback voice. Through the text of the feedback voice, the capability assertion model can determine whether the device under test plays the correct voice. By analyzing the response time of the feedback voice, the capability assertion model can evaluate the performance of the voice model built into the device under test. In an environment with noise, by analyzing the volume and clarity of the feedback voice, the quality of the speaker in the device under test can be determined, so as to obtain a voice ability report of the device under test in terms of voice feedback.
[0049] In addition, the control module 400 also analyzes the mechanical state of the device under test through the mechanical sensing module, so as to obtain motion information such as the force, stroke, speed, response time, and motion time of the device under test during mechanical motion. By comparing the above motion information with the preset device information through the capability assertion model, a voice ability report of the device under test in terms of mechanical response can be obtained.
[0050] By comprehensively analyzing the voice ability reports of the device under test in terms of voice feedback and mechanical response respectively, the voice processing ability of the device under test based on voice commands can be comprehensively evaluated.
[0051] In addition, in an embodiment of the present application, the voice device testing apparatus further includes a device camera module 800. The device camera module 800 is configured to obtain an operation image of the device under test after receiving a voice command. The control module 400 identifies the mechanical state of the device under test based on the operation image of the device under test, so as to evaluate the voice processing ability of the device under test.
[0052] When the device under test cannot be fixed by the mechanical fixing module or the mechanical movement of the device under test cannot be detected by the mechanical sensing module, during the process of the control module 400 playing a voice command through the pronunciation module 200, the device camera module 800 obtains the operation image of the device under test, and the control module 400 analyzes the operation images within a certain period of time to obtain the mechanical movement state of the device under test within a certain period of time.
[0053] Alternatively, if it is necessary to test the voice processing ability of the device under test when the distance between the controller and the device under test is relatively far, the mechanical sensing module cannot directly detect the mechanical movement of the device under test after receiving a voice command. At this time, by obtaining the operation image of the device under test through the device camera module 800, it is also possible to analyze the feedback voice and mechanical movement state of the device under test at a relatively long distance, so as to be able to evaluate the voice processing ability more comprehensively.
[0054] For example, when it is necessary to perform a voice test on a wall-mounted air conditioner integrated with an intelligent voice module, since the height of the wall-mounted air conditioner is generally relatively high, the mechanical sensing module cannot directly contact the main unit, fan blade or air deflector of the wall-mounted air conditioner. At this time, a voice command of "turn on the cooling mode" is played through the pronunciation module 200, and the wall-mounted air conditioner is continuously photographed through the device camera module 800, and the opening and closing states of the fan blade and the air deflector in the photographed air conditioner operation image are analyzed, so as to be able to analyze the response speed of the wall-mounted air conditioner to the voice command and the corresponding execution duration.
[0055] During the process of the voice device testing apparatus moving and playing a voice command at the same time, the device camera module 800 locks the device under test, and the control module 400 continuously obtains the operation image of the device under test through the device camera module 800. While analyzing the mechanical movement state of the device under test through the operation image, the distance between the voice device testing apparatus and the device under test is calculated through the operation image. By analyzing the distance between the voice device testing apparatus and the device under test and the feedback voice received by the sound collection module 300, the maximum voice recognition distance and the maximum volume propagation distance of the device under test are respectively determined.
[0056] For example, the voice module 200 alternately plays voice commands of "turn on the cooling mode" and "turn off the cooling mode" for the wall-mounted air conditioner, and the mobile module 100 drives the voice device test device to slowly move away from directly below the wall-mounted air conditioner. The device camera module 800 continuously obtains operation images of the wall-mounted air conditioner. When it is recognized that the fan blades and air deflector of the wall-mounted air conditioner stop rotating, it indicates that the position where the voice command of "turn on the cooling mode" or "turn off the cooling mode" corresponding to the current opening and closing state of the fan blades and air deflector was most recently played is the maximum voice recognition distance of the wall-mounted air conditioner. Then, the mobile module 100 can also drive the voice device test device to move closer to the wall-mounted air conditioner again. When it is determined that the fan blades and air deflector start rotating again, the maximum voice recognition distance of the wall-mounted air conditioner can be verified.
[0057] In addition, during the process of driving the voice device test device to move within the space to be measured by the mobile module 100, the operation images are analyzed to select test points where there are obstacles between the voice device test device and the device under test, so as to be able to test the reception ability of voice commands in the case of obstacles blocking.
[0058] In addition, as shown in Figure 8 In an embodiment of the present application, a radar component 110 and an environmental camera component 120 are provided on the mobile module 100. The mobile module 100 scans the test space where the voice device test device is located through the radar component 110. The mobile module 100 also takes pictures of the test space where the voice device test device is located through the environmental camera component 120. The mobile module 100 performs autonomous navigation and obstacle avoidance in the test space based on the environmental information obtained by the radar component 110 and the environmental camera component 120.
[0059] Specifically, the mobile module 100 integrates a self-developed and optimized SLAM model (simultaneous localization and mapping). The overall framework of the test space is pre-entered into the mobile module 100. Obstacles of different sizes, positions, and types are randomly placed in the test space to simulate a real usage scenario. During the process of the mobile module 100 moving in the test space, the mobile module 100 obtains the placement of obstacles near the voice device test device through the three-dimensional radar component 110, and real-time draws a map of the drivable area in the test space based on the data detected by the radar component 110, so that the voice device test device can autonomously perform voice tests on the device under test according to a preset test procedure without the interference of staff.
[0060] In addition, the mobile module 100 is also provided with an environmental camera component 120. There are obstacles in the test space that cannot be clearly detected by the radar component 110. Therefore, during the movement, the environment in which the voice equipment test device is located is photographed by the environmental camera component 120, the photographed environmental image is analyzed, and the obstacles in the environmental image are identified, so as to adjust the map available for driving in the test space, thereby ensuring the safety of the voice equipment test device during the movement.
[0061] In addition, refer to Figure 9 As shown, the embodiment of the present application further proposes a voice device testing method, which is applied to the voice device testing device in the above embodiment. The voice device testing method can be executed by a server, or can also be executed by a terminal, or can also be executed by a server in conjunction with a terminal. The voice device testing method includes but is not limited to the following steps S910 to S930: Step S910, controlling the voice equipment testing device to move between preset test points in the test space; Step S920, modulating the human voice recording and background noise in the voice command, and playing the modulated voice command through the pronunciation module; Step S930, obtaining feedback voice of the device under test after receiving the voice command through the sound receiving module, recognizing and processing the feedback voice, and obtaining the voice processing capability of the device under test.
[0062] It can be understood that the specific implementation of the voice device testing method is basically the same as the specific embodiment of the above-mentioned voice device testing device, and will not be repeated here.
[0063] In addition, refer to Figure 10 The present application also provides a testing device 1000, comprising: The driving module 1001 is used to control the voice equipment testing device to move between preset test points in the test space; The testing module 1002 is used to modulate the human voice recording and background noise in the voice command, and play the modulated voice command through the pronunciation module; The processing module 1003 is used to obtain the feedback voice of the device under test after receiving the voice command through the sound receiving module, recognize and process the feedback voice, and obtain the voice processing capability of the device under test.
[0064] The above-mentioned testing device 1000 and the voice device testing method are based on the same inventive concept, which will not be described in detail here.
[0065] In addition, refer to Figure 11 , Figure 11 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes: The processor 1101 can be implemented in the form of a general - purpose CPU (Central Processing Unit), a microprocessor, an application - specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application; The memory 1102 can be implemented in the form of a read - only memory (ROM), a static storage device, a dynamic storage device, or a random - access memory (RAM), etc. The memory 1102 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1102 and are called by the processor 1101 to execute the voice device testing method of the embodiments of the present application; The input / output interface 1103 is used to implement information input and output; The communication interface 1104 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.); The bus 1105 transmits information between the various components of the device (such as the processor 1101, the memory 1102, the input / output interface 1103, and the communication interface 1104); Among them, the processor 1101, the memory 1102, the input / output interface 1103, and the communication interface 1104 are communicatively connected to each other inside the device through the bus 1105.
[0066] The embodiments of the present application also provide a storage medium. The storage medium is a computer - readable storage medium for computer - readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above - mentioned voice device testing method.
[0067] As a non - transient computer - readable storage medium, the memory can be used to store non - transient software programs and non - transient computer - executable programs. In addition, the memory can include high - speed random - access memory, and can also include non - transient memory, such as at least one magnetic disk storage device, a flash memory device, or other non - transient solid - state storage devices. In some embodiments, the memory optionally includes a memory remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above - mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0068] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0069] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0070] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0071] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0072] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0073] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0074] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0075] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0076] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.
Claims
1. A voice device testing apparatus, characterized in that Comprising: A moving module for driving the voice device testing apparatus to move within a preset testing space; A pronunciation module disposed on the moving module, the pronunciation module for playing voice instructions simulating human voices so that the device under test emits feedback voice according to the voice instructions, wherein the voice instructions include human voice recordings and background noise; A sound collection module disposed on the moving module, the sound collection module for receiving the feedback voice; A control module, the moving module, the pronunciation module and the sound collection module are respectively electrically connected to the control module. During the process that the control module controls the moving module to drive the voice device testing apparatus to move, the control module controls the pronunciation module to play the voice instructions, and receives the feedback voice of the device under test through the sound collection module, and evaluates the voice processing ability of the device under test according to the feedback voice.
2. The voice device testing apparatus according to claim 1, wherein An artificial head structure is disposed on the moving module, the pronunciation module includes a mouth simulator disposed on the mouth of the artificial head structure, and the sound collection module includes an ear simulator disposed on the ear of the artificial head structure.
3. The voice device testing apparatus according to claim 2, wherein An attitude adjustment component is disposed at the bottom of the artificial head structure, and the attitude adjustment component is used to adjust the attitude of the artificial head structure to control the height and orientation of the mouth simulator and the ear simulator.
4. The voice device testing apparatus according to claim 1, wherein A sound pressure calibration component is further disposed on the pronunciation module. Before the pronunciation module plays the voice instructions, the control module calibrates the sound pressure of the pronunciation module through the sound pressure calibration component.
5. The voice device testing apparatus according to claim 1, wherein The voice device testing apparatus further includes a device camera module for acquiring an operation image of the device under test after receiving the voice instructions, and the control module identifies the mechanical state of the device under test according to the operation image of the device under test to evaluate the voice processing ability of the device under test.
6. The voice device testing apparatus according to claim 1, wherein The voice device testing apparatus further includes a mechanical fixing component and a mechanical sensing component. Both the mechanical fixing component and the mechanical sensing component are disposed on the moving module. The mechanical fixing component is used to fix the device under test, and the mechanical sensing component is used to detect the mechanical state of the device under test.
7. The voice device testing apparatus according to claim 6, wherein An ability assertion model is disposed in the control module. After the control module acquires the feedback voice of the device under test through the sound collection module, or acquires the mechanical state of the device under test when receiving the voice instructions through the mechanical sensing component, the ability assertion model discriminates the voice processing ability of the device under test based on the feedback voice and the mechanical state, and generates a voice ability report of the device under test.
8. The voice device testing apparatus according to claim 1, wherein A radar component and an environmental camera component are provided on the mobile module. The mobile module scans the test space where the voice device testing apparatus is located through the radar component, and the mobile module also takes pictures of the test space where the voice device testing apparatus is located through the environmental camera component. The mobile module performs autonomous navigation and obstacle avoidance in the test space based on the environmental information obtained by the radar component and the environmental camera component.
9. A method for testing a voice device, characterized in that, The voice device testing method is applied to the voice device testing apparatus according to any one of claims 1 to 8. The voice device testing method includes: Controlling the voice device testing apparatus to move between preset test points in the test space; Modulating the human voice recording and background noise in the voice command, and playing the modulated voice command through the pronunciation module; Obtaining the feedback voice of the device under test after receiving the voice command through the receiving module, and identifying and processing the feedback voice to obtain the voice processing ability of the device under test.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program; wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the voice device testing method according to claim 9.
Citation Information
Patent Citations
Test method and system
CN113257247A
Equipment testing method and device, electronic equipment and computer storage medium
CN117409763A
Test method, system and device of intelligent voice interaction system and medium
CN118135998A
Intelligent audio equipment test box
CN209676482U