Voice wake-up test method and apparatus, electronic device, and computer-readable storage medium

CN122738451APending Publication Date: 2026-09-11FALCON INNOVATIONS TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610971134.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

然而,智能眼镜的使用场景具有高度动态性,因此这种静态环境下的测试,无法真实反映智能眼镜在复杂实际工况下的语音唤醒鲁棒性

Benefits of technology

测试参数维度包括佩戴姿态参数、运动状态参数、噪声参数和声场距离参数,可以全面覆盖待测设备在真实使用场景中可能遇到的各类动态工况;基于多个测试参数维度的测试参数生成测试用例的组合,实现了对复杂场景的系统性遍历与自动化编排,避免了人工测试的随机性与片面性,确保了测试覆盖的完整性与可复现性;控制模拟设备执行测试用例的组合,可以精准还原用户在实际佩戴过程中因姿态变化、身体运动、环境噪声干扰以及声源远近差异所形成的综合声学环境,这种多维联动的模拟方式,使得唤醒词采集条件与真实用户体验高度一致,从而能够客观评估语音唤醒算法在真实工况下的实际表现,显著提升了测试结果的真实性与可信度;通过采集唤醒结果数据并自动生成包含唤醒率、拒识率、误唤醒率、响应时延和鲁棒性评分等指标的测试报告,实现了测试数据的量化输出与标准化评价。如此,不仅为唤醒算法迭代优化提供了明确的数据支撑,还大幅降低了人工统计与分析的成本,支持全流程自动化的高效测试,可以准确测试出复杂工况的语音唤醒鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122738451A_ABST
    Figure CN122738451A_ABST
Patent Text Reader

Abstract

This application discloses a voice wake-up testing method, apparatus, electronic device, and computer-readable storage medium, relating to the field of voice wake-up technology. The method includes: acquiring test parameters across multiple test parameter dimensions, including: wearing posture parameters, motion state parameters, noise parameters, and sound field distance parameters; generating a combination of test cases based on the test parameters across these dimensions; controlling a simulation device to execute the combination of test cases to simulate corresponding wearing postures and motion states, playing corresponding environmental noise, and playing a wake-up word within the environmental noise according to the sound field distance parameters; collecting wake-up result data; and generating a voice wake-up test report based on the wake-up result data, the test report including one or more of the following: wake-up rate, rejection rate, false wake-up rate, response latency, and robustness score. Thus, this solution can accurately test the robustness of voice wake-up under complex working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice wake-up technology, specifically to a voice wake-up testing method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] The reliability of voice wake-up functionality in smart glasses directly impacts user experience. Industry testing protocols for voice wake-up performance are primarily based on those used for fixed or handheld devices such as smart speakers, smartphones, and smart home devices. These traditional tests are typically conducted in controlled, static environments, placing the device under test in a fixed position and angle, and evaluating wake-up success rates by playing a standard wake-up phrase. However, smart glasses are used in highly dynamic scenarios, so these static environment tests cannot accurately reflect the robustness of voice wake-up in complex real-world conditions. Summary of the Invention

[0003] This application provides a voice wake-up testing method, apparatus, electronic device, and computer-readable storage medium, which can accurately test the voice wake-up robustness under complex working conditions.

[0004] In a first aspect, embodiments of this application provide a voice wake-up testing method, including: The test parameters are obtained from multiple test parameter dimensions, including: wearing posture parameters, motion state parameters, noise parameters, and sound field distance parameters. Based on the test parameters of the multiple test parameter dimensions, a combination of test cases is generated; The control simulation device executes a combination of the test cases to simulate the corresponding wearing posture and movement state, plays the corresponding environmental noise, and plays the wake-up word in the environmental noise according to the sound field distance parameter; Collect wake-up result data, and generate a voice wake-up test report based on the wake-up result data. The test report includes one or more of the following: wake-up rate, rejection rate, false wake-up rate, response latency, and robustness score.

[0005] Secondly, embodiments of this application provide a voice wake-up testing device, including: The parameter acquisition module is used to acquire test parameters of multiple test parameter dimensions, including: wearing posture parameters, motion state parameters, noise parameters, and sound field distance parameters. The test case generation module is used to generate combinations of test cases based on the test parameters of the multiple test parameter dimensions. The test execution module is used to control the simulation device to execute the combination of the test cases to simulate the corresponding wearing posture and movement state, play the corresponding environmental noise, and play the wake-up word in the environmental noise according to the sound field distance parameter; The report generation module is used to collect wake-up result data and generate a voice wake-up test report based on the wake-up result data. The test report includes one or more of the following: wake-up rate, rejection rate, false wake-up rate, response latency, and robustness score.

[0006] Thirdly, embodiments of this application also provide an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the above-described voice wake-up test method.

[0007] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the above-described voice wake-up test method.

[0008] Fifthly, embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described in embodiments of this application.

[0009] The embodiments of this application have the following beneficial effects: The test parameters include wearing posture parameters, motion state parameters, noise parameters, and sound field distance parameters, comprehensively covering various dynamic working conditions that the device under test may encounter in real-world usage scenarios. The combination of test cases generated based on multiple test parameter dimensions enables systematic traversal and automated orchestration of complex scenarios, avoiding the randomness and bias of manual testing and ensuring the completeness and reproducibility of test coverage. Controlling the simulated device to execute test cases accurately recreates the comprehensive acoustic environment formed by user posture changes, body movements, environmental noise interference, and differences in sound source distance during actual wear. This multi-dimensional simulation method ensures that the wake-up word acquisition conditions are highly consistent with real user experience, thus enabling objective evaluation of the actual performance of the voice wake-up algorithm under real-world conditions and significantly improving the authenticity and credibility of the test results. By collecting wake-up result data and automatically generating test reports containing indicators such as wake-up rate, rejection rate, false wake-up rate, response latency, and robustness score, the quantitative output and standardized evaluation of test data are achieved. This not only provides clear data support for the iterative optimization of the wake-up algorithm, but also significantly reduces the cost of manual statistics and analysis, supports efficient testing with full-process automation, and can accurately test the robustness of voice wake-up under complex working conditions. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram of the steps of a voice wake-up testing method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the modules of a voice wake-up testing system provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a voice wake-up testing device provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0012] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0013] In one embodiment, such as Figure 1 As shown, a voice wake-up testing method is provided. Although the logical order is illustrated in the step diagram, in some cases, the steps shown or described can be performed in a different order than that shown in the diagram. Specifically, this voice wake-up testing method can be applied to a voice wake-up testing system. Figure 2 This is a schematic diagram of the structure of a voice wake-up testing system provided in an embodiment of this application; as shown below. Figure 2 As shown, this voice wake-up testing system can include a user configuration layer, a control and scheduling layer, a hardware simulation layer, a device under test (DUT) layer, and a data acquisition and evaluation layer. The user configuration layer supports unified configuration of test parameters such as posture, motion, noise, near and far fields, and continuous duration. The control and scheduling layer can include a combined traversal engine, a timing control module, and an automated execution engine to achieve fully automated scheduling. The hardware simulation layer can include posture / motion simulation devices, sound source devices (such as an artificial mouth), and noise playback devices. The data acquisition and evaluation layer can automatically collect wake-up results, record false wake-ups, statistically analyze indicators, and generate test reports. The DUT layer can include the DUT, which may include an audio acquisition module and a voice wake-up module. The audio acquisition module can be a microphone, used to acquire the wake-up word audio signal from the external environment and the background noise signal superimposed on the wake-up word audio signal during the test, and convert the acquired acoustic signals into audio data. The voice wake-up module communicates with the audio acquisition module to receive audio data. It sequentially performs voice endpoint detection, voice feature extraction, and wake-up word template matching on the audio data. When the matching result meets the preset wake-up conditions, the voice wake-up module outputs a wake-up success event and corresponding wake-up delay information. When the matching result does not meet the preset wake-up conditions, the voice wake-up module outputs a wake-up failure event or a rejection event. The wake-up success event, wake-up failure event, rejection event, and wake-up delay information constitute the wake-up result data, which is collected by the test system through the device log interface for subsequent voice wake-up robustness evaluation.

[0014] In one embodiment, the device under test can be smart glasses, which can be wearable optical see-through smart glasses. Specifically, the wearable optical see-through smart glasses may further include at least a glasses frame, optical display components, electronic circuit components, sensors, etc., and the sensors built into the glasses include a heart rate monitor, a blood glucose meter, a microphone, and / or an eye tracker.

[0015] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the priority of the embodiments.

[0016] according to Figure 1 The voice wake-up test method shown includes at least steps S110 to S140, which are described in detail below: In step S110, test parameters of multiple test parameter dimensions are obtained.

[0017] In step S120, a combination of test cases is generated based on the test parameters of the multiple test parameter dimensions.

[0018] In step S130, the simulation device is controlled to execute the combination of the test cases to simulate the corresponding wearing posture and movement state, play the corresponding environmental noise, and play the wake-up word in the environmental noise according to the sound field distance parameter.

[0019] In step S140, wake-up result data is collected, and a voice wake-up test report is generated based on the wake-up result data. The test report includes one or more of the following: wake-up rate, rejection rate, false wake-up rate, response latency, and robustness score.

[0020] Test parameter dimensions refer to categorical variables used to characterize the various external usage conditions and acoustic environment conditions faced by the device under test during actual use. Each test parameter dimension corresponds to an independent category of factors affecting voice wake-up performance. Test parameters may include, but are not limited to: wearing posture parameters, motion state parameters, noise parameters, and sound field distance parameters. Wearing posture parameters refer to physical quantities reflecting the spatial angle and relative displacement of the device under test (DUT) relative to a reference plane of the user's body parts (such as the head or face). Motion state parameters describe the type and intensity of limb movements performed by the wearer of the DUT during the test. Noise parameters refer to the set of sound types and their energy intensities used to simulate background interference in a real environment. Sound field distance parameters refer to the spatial relationship between the wake-up word's sound source and the audio acquisition module of the DUT.

[0021] In one embodiment, each test parameter dimension may include multiple test parameters. For example, wearing posture parameters may specifically include pitch angle (angle of head down or up), yaw angle (angle of left and right side turn), roll angle (angle of frame tilt), and the tightness or displacement of the temples in contact with the head. Motion state parameters refer to parameters used to simulate the dynamic physical behavior of a user while wearing smart glasses, including but not limited to states such as stillness, constant speed walking, rapid arm swinging, head turning, and external turbulence. Noise parameters refer to background interference parameters used to reproduce a realistic acoustic environment, covering noise scene types (such as quiet indoors, noisy shopping malls, street wind noise, etc.) and noise intensity (sound pressure level, for example, 30dB to 85dB). Sound field distance parameters refer to parameters used to define the spatial positional relationship of the wake word sound source relative to the smart glasses microphone, including the distance classification of the sound source (near field, mid field, far field) and the incident angle of the sound wave (horizontal azimuth and vertical pitch angle). Table 1 is an example of multiple test parameters included in multiple test parameter dimensions.

[0022] Table 1

[0023] As shown in Table 1, each test parameter dimension may include one or more test parameter sub-dimensions, and the test parameters of each test parameter sub-dimension are independent of each other and can be freely combined. As an example, generating a combination of test cases based on the test parameters of the multiple test parameter dimensions may include: traversing and combining the test parameters of each dimension in the multiple test parameter dimensions to generate a combination of test cases covering various scenario conditions. By cross-traversing the test parameters of each test parameter sub-dimension, a complete set of test cases covering real-world usage scenarios can be constructed. For example, selecting a test parameter from each test parameter sub-dimension can yield test parameters such as head-down, looseness, walking, office noise scenario, 40dB, near-field, and horizontal, and a combination of test cases can be generated based on this test parameter.

[0024] A test case combination refers to a complete set of test instructions for a single automated execution, composed of the intersection of specific values ​​from multiple test parameter dimensions. Generating test case combinations based on multiple test parameters involves using permutation and combination algorithms to cross-match parameter values ​​from different dimensions, thereby constructing a test matrix covering the entire scenario. This combination method can systematically traverse all possible operating conditions, avoiding omissions by manual testing.

[0025] It can control one or more simulation devices to execute combinations of test cases, thereby physically reproducing a preset test environment to simulate corresponding wearing postures and movement states, play corresponding environmental noise, and play a wake-up word within that environmental noise according to selected sound field distance parameters. For example, a multi-degree-of-freedom robotic arm or posture simulation platform can be used to adjust the angle of the device under test to simulate wearing posture, a vibration platform or displacement device can be used to simulate movement states, a high-fidelity speaker can be used to play pre-recorded environmental noise, and a standard wake-up word audio can be played through a standard sound source (such as an artificial mouth) at a specified distance and angle. This process ensures a high degree of consistency and repeatability of test conditions.

[0026] When executing test case combinations, the system can detect whether voice wake-up is achieved based on the voice wake-up module of the device under test (DUT), collect wake-up result data, and generate a test report to quantitatively evaluate the test results. Specifically, the system can capture the internal logs of the DUT in real time via wired or wireless interfaces, recording the success or failure of each wake-up attempt, response time, and whether false triggers occur. Finally, through statistical analysis of multiple data points, key indicators such as wake-up rate, rejection rate, and false wake-up rate are calculated, and a robustness score is derived. The robustness score is a quantitative value calculated based on a weighted average of multiple wake-up performance indicators, used to comprehensively evaluate the stability of the algorithm in complex environments.

[0027] The technical solution adopted in this application includes test parameters such as wearing posture parameters, motion state parameters, noise parameters, and sound field distance parameters, which can comprehensively cover various dynamic working conditions that the device under test may encounter in real-world usage scenarios. The combination of test cases generated based on multiple test parameter dimensions enables systematic traversal and automated orchestration of complex scenarios, avoiding the randomness and bias of manual testing and ensuring the completeness and reproducibility of test coverage. Controlling the simulated device to execute the combination of test cases can accurately reproduce the comprehensive acoustic environment formed by changes in posture, body movement, environmental noise interference, and differences in the distance of sound sources during actual wear. This multi-dimensional linkage simulation method makes the wake-up word acquisition conditions highly consistent with real user experience, thereby objectively evaluating the actual performance of the voice wake-up algorithm under real-world conditions and significantly improving the authenticity and credibility of the test results. By collecting wake-up result data and automatically generating test reports containing indicators such as wake-up rate, rejection rate, false wake-up rate, response latency, and robustness score, quantitative output and standardized evaluation of test data are achieved. This not only provides clear data support for the iterative optimization of the wake-up algorithm, but also significantly reduces the cost of manual statistics and analysis, supports efficient testing with full-process automation, and can accurately test the robustness of voice wake-up under complex working conditions.

[0028] In one embodiment, the user configuration layer of the voice wake-up testing system can receive test requirements configured by the tester. For example, in terms of wearing posture parameters, the tester can select three states: head down 15 degrees, head up 20 degrees, and tilted to the left 10 degrees; in terms of motion parameters, they can select two states: stationary and simulated walking; in terms of noise parameters, they can select 45dB office environment noise and 70dB street noise; and in terms of sound field distance parameters, they can select 0.5 meters in the near field directly in front and 1.5 meters in the far field at a 45-degree angle to the right front. After receiving these parameters, the control scheduling layer automatically generates 3×2×2×2=24 test cases using the combination traversal engine. Subsequently, the timing control module and the automated execution engine begin to control the hardware simulation layer and the device under test layer to perform simulated testing: first, the posture simulation platform is controlled to adjust the device under test to a 15-degree head down position, the vibration platform is started to simulate walking rhythm, and 45dB of office background noise is played through surround sound. On this basis, a standard artificial mouth located 0.5 meters in front clearly pronounces the wake-up word. The audio acquisition module of the device under test (DUT) can acquire audio signals, and the voice wake-up module processes the received audio signals using its internal algorithm. The data acquisition and evaluation layer reads the log file output by the DUT in real time via a USB cable. If the DUT returns a wake-up success flag within one second, it is recorded as a successful wake-up, with a latency of 800 milliseconds; otherwise, it is recorded as a wake-up failure. After completing a cyclical test of all 24 test cases, the system also underwent a 4-hour silent monitoring to count the number of false wake-ups.

[0029] Finally, the data acquisition and evaluation layer generated a test report: 22 out of 24 active wake-up tests were successful, with a wake-up rate of 91.7%; one false wake-up occurred during 4 hours of monitoring; the average response latency was 750 milliseconds. The system calculated a robustness score of 88 points based on a preset weighting formula and automatically generated a test report containing trend charts for various indicators.

[0030] Based on the above technical solution, as an example, the data collected for wake-up results may include one or more of the following: data on whether wake-up was successful or failed, response latency data, and rejection data.

[0031] Wake-up success or failure data refers to the record of whether the device under test (DUT) correctly recognizes the target wake-up word and enters the wake-up state in a single wake-up test. Response latency data refers to the time interval from the start of wake-up word playback to the output of a wake-up success event by the DUT, used to characterize the response speed of the voice wake-up algorithm. Rejection data refers to records where the DUT fails to trigger a wake-up response after receiving a voice signal that is not the target wake-up word, used to characterize the algorithm's ability to reject irrelevant voice.

[0032] After the test system controls the simulated device to execute scenario simulations according to the combination of test cases and play the wake-up words, the voice wake-up module of the device under test (DUT) processes the received audio data and generates corresponding response events. The test system captures the event logs output by the DUT in real time through the device log interface, parses and extracts the result information corresponding to each wake-up attempt. The device log interface refers to the communication channel between the test system and the DUT used to transmit device operation logs and wake-up event information, including but not limited to the ADB debug bridging protocol.

[0033] When the device under test (DUT) correctly recognizes the wake word and successfully enters the wake-up state, the test system collects the wake-up success result data and records the time difference between the start of the wake word playback and the output of the wake-up success event by the DUT, i.e., the response latency data. This response latency data reflects the processing speed of the voice wake-up algorithm and is one of the important indicators for evaluating user experience. When the DUT fails to recognize the wake word and does not output any wake-up response, the test system collects the wake-up failure result data. When the audio data received by the DUT contains a speech signal but the speech signal is not the target wake word (e.g., ordinary conversation or irrelevant speech fragments in the environment), and the DUT does not trigger a wake-up response, the test system records this situation as rejection data, which is used to measure the voice wake-up algorithm's ability to reject non-target speech.

[0034] The above three types of data can be collected individually or in combination, depending on the configuration requirements of the test scenario. For example, in a test scenario that only evaluates the wake-up success rate, only the result data of successful or failed wake-up can be collected; in a test scenario that requires a comprehensive evaluation of wake-up performance, the above three types of data can be collected simultaneously to comprehensively depict the wake-up performance of the device under test under different operating conditions.

[0035] By adopting the technical solution of this application embodiment, the voice wake-up performance of the device under test can be comprehensively characterized from three perspectives: wake-up success rate, response speed, and anti-interference capability by classifying and collecting wake-up success or failure result data, response latency data, and rejection data. This provides a complete data foundation for the subsequent generation of test reports and robustness scores containing multi-dimensional indicators, avoiding the one-sidedness of evaluation by a single indicator.

[0036] In one embodiment, the wearing posture parameters include, but are not limited to, one or more of the following: pitch angle parameters, yaw angle parameters, temple spread angle parameters, and displacement parameters; the motion state parameters include motion type parameters, which include, but are not limited to, one or more of the following: stationary, walking, arm swinging, head turning, and bumping; the noise parameters include, but are not limited to, noise scene identification and noise intensity parameters; the sound field distance parameters include, but are not limited to, sound source distance classification and sound source incident angle parameters, which include, but are not limited to, one or more of the following: horizontal angle parameters and pitch angle parameters.

[0037] The pitch angle parameter refers to the tilt angle of the device under test (DUT) along the vertical direction of the user's face or head, used to characterize the impact of head-down or head-up posture on the pickup orientation of the audio acquisition device. The yaw angle parameter refers to the horizontal rotation angle of the DUT along the horizontal direction of the user's face or head, used to characterize the impact of head-turning or head-tilting posture on the orientation of the audio acquisition device relative to the sound source. The temple spread angle parameter refers to the spread of the temples of the DUT relative to the frame plane, used to characterize the impact of temple tightness on the overall stability of the glasses. The displacement parameter refers to the offset of the DUT relative to its normal wearing reference position, used to characterize the impact of looseness, slippage, or offset on the microphone pickup area.

[0038] The pitch angle parameter characterizes the degree of tilt of the device under test (DUT) relative to the user's face in the vertical direction. For example, when a user looks down at their phone, their glasses tilt downwards with their head, or when they look up into the distance, their glasses tilt upwards with their head. Changes in this pitch angle parameter directly alter the vertical relative position between the microphone and the sound source, thus affecting the sound pickup effect. The yaw angle parameter characterizes the degree of horizontal rotation of the DUT in the left-right direction. For example, when a user tilts their head to talk to someone or turns their head to observe their surroundings, their glasses yaw accordingly. Changes in this yaw angle parameter change the horizontal angle between the microphone and the sound source. The temple opening angle parameter characterizes the opening range of the temples of the DUT relative to the frame plane. When the temple opening angle is too large, the overall fit between the glasses and the head decreases, which may cause the microphone position to deviate from the designed pickup area for normal wear. The displacement parameter characterizes the amount of offset of the DUT relative to its normal wearing position, including frame slippage due to looseness, lateral frame shift due to movement, or forward tilting of the frame due to nose pad deformation. The above four types of parameters can be configured individually or in combination to fully simulate various wearing posture changes that users may experience in real use.

[0039] Motion type parameters are parameters used to identify the categories of dynamic behaviors applied to the device under test (DUT) during testing, including static and various motion states. The static state corresponds to the baseline test condition where the DUT is fixed in a standard wearing position and does not generate any external movement. The walking state simulates the natural rise and fall and sway of the head during normal walking, reflecting the interference of everyday walking scenarios on microphone pickup. The arm swing state simulates the micro-vibration of the frame and dynamic shift of the microphone pickup area caused by the large arm swing of the user during walking. The head turning state simulates the changes in angular velocity of posture and airflow disturbance caused by the user's rapid or slow head turning during wear. The bumpy state simulates the continuous interference of high-frequency vibrations caused by uneven road surfaces when the user is traveling, affecting microphone pickup. These motion types can be applied individually or in combination; for example, a head turning motion can be superimposed on the walking state to construct a more complex composite motion scenario.

[0040] Noise scene identifiers are information used to distinguish different types of background noise environments, with each scene identifier corresponding to different noise spectrum characteristics. Noise intensity parameters are parameters used to set the sound pressure level of the background noise during testing. Noise scene identifiers specify the type of background noise simulated in the test, such as a quiet indoor environment, an open-plan office environment, a noisy shopping mall environment, a city street environment, or an outdoor wind noise environment. Different scenes correspond to different noise spectrum characteristics and temporal distribution features. Noise intensity parameters are used to set the sound pressure level of the background noise, with an adjustable range covering 30dB to 85dB, quantifying the interference level of different noise environments from an energy perspective. By jointly configuring noise scene identifiers and noise intensity parameters, a continuous noise gradient from quiet to noisy can be accurately reproduced, making the test conditions highly consistent with the real acoustic environment.

[0041] Sound source distance classification refers to the categorization of the distance between the wake-up word sound source and the device under test (DUT) based on their proximity. The sound source incident angle parameter defines the spatial direction of the wake-up word sound wave as it reaches the microphone of the DUT. The horizontal angle parameter refers to the azimuth angle of the sound source relative to the front of the DUT on the horizontal plane. The pitch angle parameter refers to the incident tilt angle of the sound source relative to the DUT on the vertical plane.

[0042] Sound source distance classification categorizes the spatial distance between the wake-up word sound source and the microphone of the device under test into near-field, mid-field, and far-field categories. Near-field corresponds to the typical distance at which the user speaks (approximately 0.3 to 0.5 meters), mid-field corresponds to a slightly longer interaction distance (approximately 1 to 1.5 meters), and far-field corresponds to sound sources at even greater distances (approximately 2 meters and above). The horizontal angle parameter characterizes the azimuth angle of the sound source relative to the front of the device under test on the horizontal plane, such as 0 degrees directly in front, 45 degrees to the left front, or 90 degrees to the right front. The pitch angle parameter characterizes the incident angle of the sound source relative to the device under test on the vertical plane, such as being at the same level as the glasses, or higher or lower than the glasses position. By jointly configuring the sound source distance classification and the sound source incident angle parameter, the spatial location of the wake-up word sound source can be accurately defined in three-dimensional space, comprehensively covering real-world scenarios where users initiate voice wake-up at different distances and orientations.

[0043] By employing the technical solution of this application embodiment, and through the refined definition of four test parameter dimensions, the testing system can accurately model and flexibly configure real-world usage scenarios for smart glasses from four aspects: wearing posture, motion state, noise environment, and sound field position. The composability of each dimension parameter ensures that test cases can cover various working conditions from simple to complex, supporting both targeted investigation of single factors and comprehensive evaluation of multiple factors, significantly improving the scenario coverage and result granularity of the test.

[0044] Based on the above technical solution, as an embodiment, the combination of control simulation device executing the test cases may include: controlling the posture / motion simulation device to wear the device under test at a corresponding wearing angle according to the wearing posture parameters; controlling the posture / motion simulation device to perform a corresponding movement according to the movement state parameters; controlling the noise playback device to play environmental noise at a corresponding sound pressure level according to the noise parameters; and controlling the sound source device to play the wake-up word at a corresponding position and angle according to the sound field distance parameters.

[0045] Combinations of control simulation devices to execute test cases can include controlling the posture / motion simulation device to wear the device under test at a corresponding wearing angle. The posture / motion simulation device is a mechanical device with multi-degree-of-freedom adjustment capabilities, comprising an adjustable wearing bracket and a contouring module that simulates head shape. Before the test begins, the test system first extracts the specific values ​​of the wearing posture parameters from the current test case, and then sends corresponding motion control commands to the posture / motion simulation device. Upon receiving the commands, the posture / motion simulation device drives the wearing bracket to adjust to the target angle in the pitch, yaw, and roll directions using its built-in servo motors or stepper motors. Simultaneously, it adjusts the relative position between the contouring module and the wearing bracket according to the displacement parameters to simulate different wearing displacement states such as normal wearing, loosening, or slipping. The device under test is fixed to the contouring module, thereby positioning its microphone and frame in a spatial orientation corresponding to the target wearing posture. Through the above control process, the spatial changes in the microphone pickup area can be accurately reproduced under real wearing conditions such as head tilting, head raising, head tilting, frame tilting, or frame slipping.

[0046] Combinations of control simulation devices to execute test cases can include controlling the posture / motion simulation device to perform corresponding movements. After completing the initial setting of the wearing posture, the test system continues to extract specific values ​​of motion state parameters from the current test case and sends corresponding motion control commands to the posture / motion simulation device. The posture / motion simulation device executes the corresponding motion mode according to the command. For example, when the motion state parameter is stationary, the wearing bracket maintains the current posture; when the motion state parameter is walking, the wearing bracket generates periodic up-and-down undulations and back-and-forth swaying according to a preset step frequency and amplitude; when the motion state parameter is arm swinging, micro-vibrations of the torso caused by arm swinging are superimposed on the walking motion; when the motion state parameter is head turning, the wearing bracket performs rotational motion in the horizontal plane according to a preset angular velocity; when the motion state parameter is bumping, the wearing bracket generates random vibrations according to a preset high-frequency vibration spectrum. The above motion modes can be executed individually or in combination, for example, a head turning action can be superimposed on the walking motion to simulate the real behavior of a user looking back while walking. Through the above control process, the combined effects of body swaying, airflow disturbance, and micro-vibration of the glasses frame on microphone pickup under different motion states can be accurately reproduced.

[0047] Combinations of control over the simulation equipment to execute test cases may include controlling a noise playback device to play ambient noise at a corresponding sound pressure level. The noise playback device is a controllable sound source device placed in the environment surrounding the device under test, and may include a high-fidelity speaker array and a power amplifier. The test system extracts specific values ​​of noise parameters from the current test case, including noise scene identifiers and noise intensity parameters. The noise audio library refers to a collection of various background noise audio files categorized by scene and pre-stored in the test system. Different scene identifiers correspond to different noise spectral characteristics and temporal distribution characteristics. Based on the noise scene identifiers, the test system selects noise audio files corresponding to the target scene from the pre-stored noise audio library, such as low-energy steady noise for a quiet indoor environment, mixed human voices and equipment operating noise for an open office environment, high-energy broadband noise for a noisy shopping mall environment, mixed traffic and pedestrian noise for an urban street environment, or low-frequency airflow noise for an outdoor wind noise environment.

[0048] Based on the noise intensity parameters, the test system adjusts the sound pressure level of the selected noise audio to the target value via a power amplifier, with an adjustable range covering 30dB to 85dB. The noise playback device continuously plays ambient noise throughout the test to create a stable background noise field at the microphone of the device under test. Through this control process, a continuous noise gradient from quiet to noisy can be accurately simulated, ensuring that the test conditions closely match the user's actual acoustic environment.

[0049] Combinations of control over the simulation device to execute test cases may include controlling the sound source device to play a wake-up word at corresponding positions and angles. The sound source device is a standard sound source device, preferably an artificial mouth, which can output a pre-recorded wake-up word audio signal with standardized acoustic characteristics. The test system extracts specific values ​​of sound field distance parameters from the current test case, including sound source distance classification and sound source incident angle parameters. According to the sound source distance classification, the test system controls the sound source device to move along a preset guide rail or robotic arm to a position corresponding to the target distance category. The near-field position corresponds to a distance of approximately 0.3 meters to 0.5 meters, simulating a typical scenario of the user speaking; the mid-field position corresponds to a distance of approximately 1 meter to 1.5 meters, simulating a slightly farther interaction scenario; and the far-field position corresponds to a distance of approximately 2 meters or more, simulating a distant sound source scenario.

[0050] Based on the horizontal angle parameter in the sound source incident angle parameters, the test system controls the sound source device to adjust to the target azimuth angle on the horizontal plane, such as 0 degrees directly in front, 45 degrees to the left front, or 90 degrees to the right front. Based on the pitch angle parameter in the sound source incident angle parameters, the test system controls the sound source device to adjust to the target incident angle on the vertical plane to simulate situations where the sound source is higher or lower than the eyeglass position. After completing the position and angle settings, the sound source device outputs the standard wake-up word audio signal according to the preset playback sequence, against a background of continuous ambient noise playback by a noise playback device. Through the above control process, the spatial position and incident direction of the wake-up word sound source can be accurately defined in three-dimensional space, comprehensively covering real-world scenarios where users initiate voice wake-up at different distances and orientations.

[0051] The control of multiple devices involves a timing coordination relationship. Specifically, the posture / motion simulation device is controlled to wear the device under test at the corresponding wearing angle, and the posture / motion simulation device is controlled to execute the corresponding movement priority to complete the initial setting of the wearing posture and the initiation of the movement state; the noise playback device is controlled to play environmental noise at the corresponding sound pressure level before playing the wake-up word to ensure that the environmental noise field has been stabilized; the sound source device is controlled to play the wake-up word at the corresponding position and angle after the above three steps are completed, so as to play the wake-up word in the target acoustic environment. The test system uses a timing control module to uniformly schedule the execution order and timing of the above four sub-steps to ensure that the hardware devices work together and achieve accurate reproduction of the test scenario.

[0052] The technical solution of this application decomposes the execution of test cases into four independent hardware control sub-steps: wearing posture simulation, motion state simulation, environmental noise simulation, and wake-up word playback. This achieves precise decoupling and independent control of test conditions in each dimension. Each sub-step drives the corresponding simulation device through parameterized instructions, ensuring the configurability, repeatability, and reproducibility of the test scenario. This avoids random errors caused by manual operation and provides reliable hardware execution assurance for the quantitative evaluation of wake-up performance.

[0053] Based on the above technical solution, as an embodiment, the voice wake-up test method may further include: recording false wake-up event information within a set continuous monitoring period, wherein the false wake-up event information includes timestamps and scene attribution information; and establishing a continuous false wake-up statistical model based on the false wake-up event information.

[0054] False wake-up events can be recorded within a set continuous monitoring period. A false wake-up event refers to an abnormal event in which the voice wake-up module of the device under test (DUT) triggers a wake-up response on its own without receiving a target wake-up word. The continuous monitoring period is a pre-configured monitoring timeframe for false wake-up statistics, with selectable values ​​including but not limited to 2 hours, 4 hours, 12 hours, and 24 hours. Within this continuous monitoring period, the test system continuously monitors the event logs output by the voice wake-up module of the DUT through the device log interface. When the voice wake-up module of the DUT triggers a wake-up response on its own without receiving a target wake-up word, the test system classifies this event as a false wake-up event and records it.

[0055] Each false wake-up event record includes at least a timestamp and scene attribution information. The timestamp precisely identifies the moment the false wake-up event occurred, enabling subsequent analysis of its temporal distribution patterns, such as whether frequent false wake-ups occur within a specific timeframe. Scene attribution information identifies the corresponding test scene conditions, including the specific values ​​of wearing posture parameters, motion state parameters, noise parameters, and sound field distance parameters. By recording scene attribution information, it is possible to trace the combination of operating conditions under which each false wake-up occurred, thus providing data to pinpoint the cause of the false wake-up.

[0056] During the continuous monitoring period, the test system maintains the scenario conditions set by the current test case, continuously collecting acoustic signals from the environment and inputting them to the device under test (DUT). During this period, ambient noise is continuously played, but no wake-up word is actively played. The DUT's voice wake-up module remains in standby listening mode. If it outputs a wake-up success event without any target wake-up word input, the test system captures and records it as a false wake-up event. In some test scenarios, scenario conditions can be dynamically switched according to a preset strategy within the continuous monitoring period to examine the probability of false wake-ups during different scenario transitions.

[0057] A continuous false wake-up statistical model can be established based on false wake-up event information. This model quantifies false wake-up behavior from multiple dimensions, including time distribution, scene distribution, and frequency distribution, based on all false wake-up event information recorded during continuous monitoring. After the continuous monitoring period ends, the test system summarizes and analyzes all recorded false wake-up event information to construct the continuous false wake-up statistical model. This model quantifies false wake-up behavior from multiple dimensions: in the time dimension, it statistically analyzes the time interval distribution of false wake-up events and the trend of cumulative frequency over time; in the scene dimension, it counts the number of false wake-ups according to different wearing postures, movement states, noise scenes, and sound field distance conditions, identifying high-incidence combinations of working conditions; and in the frequency dimension, it calculates the false wake-up incidence rate per unit time and generates a histogram of false wake-up frequency distribution.

[0058] The continuous false wake-up statistical model can further output quantitative indicators of the false wake-up rate, such as the average number of false wake-ups per hour and the cumulative number of false wake-ups in 24 hours, providing an objective evaluation benchmark for horizontal comparison between different algorithm versions or different hardware solutions.

[0059] The technical solution adopted in this application, by introducing a mechanism for recording and statistically modeling false wake-up events within a continuous monitoring period, overcomes the limitation of traditional wake-up testing that only focuses on the success rate of active wake-up, and achieves long-term, multi-dimensional, and quantifiable evaluation of false wake-up behavior. This mechanism can accurately locate the combination of working conditions and time patterns that lead to frequent false wake-ups, providing clear data guidance for the iterative optimization of voice wake-up algorithms, and significantly improving the comprehensiveness and scientific nature of the test evaluation.

[0060] To facilitate better implementation of the voice wake-up testing method of this application, this application also provides a voice wake-up testing device based on the above-described voice wake-up testing method. The meanings of the terms used are the same as in the above-described voice wake-up testing method, and specific implementation details can be found in the description of the method embodiments.

[0061] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of the voice wake-up testing device provided in the embodiments of this application, wherein the voice wake-up testing device includes: The parameter acquisition module 301 is used to acquire test parameters of multiple test parameter dimensions, including: wearing posture parameters, motion state parameters, noise parameters and sound field distance parameters; The test case generation module 302 is used to generate a combination of test cases based on the test parameters of the multiple test parameter dimensions; The test execution module 303 is used to control the simulation device to execute the combination of the test cases to simulate the corresponding wearing posture and movement state, play the corresponding environmental noise, and play the wake-up word in the environmental noise according to the sound field distance parameter; The report generation module 304 is used to collect wake-up result data and generate a voice wake-up test report based on the wake-up result data. The test report includes one or more of the following: wake-up rate, rejection rate, false wake-up rate, response latency, and robustness score.

[0062] In one embodiment, the use case generation module 302 is specifically used to perform: The test parameters of each dimension in the multiple test parameter dimensions are traversed and combined to generate a combination of test cases that cover the conditions of each scenario.

[0063] In one embodiment, the test execution module 303 is specifically used to execute: Based on the wearing posture parameters, control the posture / motion simulation device to wear the device under test at the corresponding wearing angle; Based on the motion state parameters, the posture / motion simulation device is controlled to perform the corresponding motion; Based on the noise parameters, the noise playback device is controlled to play ambient noise at the corresponding sound pressure level; Based on the sound field distance parameter, the sound source device is controlled to play the wake-up word at the corresponding position and angle.

[0064] In one embodiment, the device under test includes an audio acquisition module and a voice wake-up module.

[0065] In one embodiment, the report generation module 304 is specifically configured to perform: Collect one or more of the following: wake-up success or failure result data, response latency data, and rejection data.

[0066] In one embodiment, the device further includes: The recording module is used to record false wake-up event information within a set continuous monitoring period. The false wake-up event information includes timestamps and scene attribution information. A module is established to build a statistical model for continuous false wake-up based on the false wake-up event information.

[0067] In one embodiment, the wearing posture parameters include one or more of the following: pitch angle parameters, yaw angle parameters, temple spread angle parameters, and displacement parameters; The motion state parameters include motion type parameters, which include one or more of the following: stationary, walking, arm swinging, head turning, and bumping. The noise parameters include noise scene identifiers and noise intensity parameters; The sound field distance parameters include sound source distance classification and sound source incident angle parameters, wherein the sound source incident angle parameters include one or more of horizontal angle parameters and pitch angle parameters.

[0068] The technical solution adopted in this application includes test parameters such as wearing posture parameters, motion state parameters, noise parameters, and sound field distance parameters, which can comprehensively cover various dynamic working conditions that the device under test may encounter in real-world usage scenarios. The combination of test cases generated based on multiple test parameter dimensions enables systematic traversal and automated orchestration of complex scenarios, avoiding the randomness and bias of manual testing and ensuring the completeness and reproducibility of test coverage. Controlling the simulated device to execute the combination of test cases can accurately reproduce the comprehensive acoustic environment formed by changes in posture, body movement, environmental noise interference, and differences in the distance of sound sources during actual wear. This multi-dimensional linkage simulation method makes the wake-up word acquisition conditions highly consistent with real user experience, thereby objectively evaluating the actual performance of the voice wake-up algorithm under real-world conditions and significantly improving the authenticity and credibility of the test results. By collecting wake-up result data and automatically generating test reports containing indicators such as wake-up rate, rejection rate, false wake-up rate, response latency, and robustness score, quantitative output and standardized evaluation of test data are achieved. This not only provides clear data support for the iterative optimization of the wake-up algorithm, but also significantly reduces the cost of manual statistics and analysis, supports efficient testing with full-process automation, and can accurately test the robustness of voice wake-up under complex working conditions.

[0069] Specific limitations regarding the voice wake-up testing device can be found in the limitations of the voice wake-up testing method described above, and will not be repeated here. Each module in the aforementioned voice wake-up testing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0070] In addition, this application also provides an electronic device, such as Figure 4 As shown, the electronic device can be smart glasses, and the diagram illustrates the structure of the electronic device involved in this application. Specifically: The electronic device may include components such as a processor 401 with one or more processing cores and a memory 402 with one or more computer-readable storage media. Those skilled in the art will understand that... Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 401 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402, it performs various functions and processes data, thereby providing overall monitoring of the electronic device. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.

[0071] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0072] In one embodiment, the electronic device further includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power equipment debugging circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0073] In one embodiment, the electronic device may further include an input unit 404, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0074] When the specific electronic device is smart glasses, in addition to the above structure, it also includes at least the glasses frame, optical display components, electronic circuit components, sensors, etc. The sensors built into the glasses include heart rate monitors, blood glucose meters, microphones, cameras and / or eye trackers.

[0075] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402, thereby implementing the steps in any of the voice wake-up test methods provided in the embodiments of this application.

[0076] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0077] In one embodiment, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the methods described in any embodiment of this application.

[0078] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in any embodiment of this application.

[0079] In some embodiments, a computer program product is also provided, including a computer program or instructions that, when executed by a processor, implement the methods described in any embodiment of this application.

[0080] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0081] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0082] To this end, this application provides a computer-readable storage medium storing a computer program that can be loaded by a processor to execute the steps in any of the voice wake-up testing methods provided in this application.

[0083] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0084] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0085] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the voice wake-up testing methods provided in this application, the beneficial effects that any of the voice wake-up testing methods provided in this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0086] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0087] The above provides a detailed description of a voice wake-up testing method, apparatus, electronic device, and computer-readable storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A voice wake-up testing method, characterized in that, include: The test parameters are obtained from multiple test parameter dimensions, including: wearing posture parameters, motion state parameters, noise parameters, and sound field distance parameters. Based on the test parameters of the multiple test parameter dimensions, a combination of test cases is generated; The control simulation device executes a combination of the test cases to simulate the corresponding wearing posture and movement state, plays the corresponding environmental noise, and plays the wake-up word in the environmental noise according to the sound field distance parameter; Collect wake-up result data, and generate a voice wake-up test report based on the wake-up result data. The test report includes one or more of the following: wake-up rate, rejection rate, false wake-up rate, response latency, and robustness score.

2. The method according to claim 1, characterized in that, The combination of test cases generated based on the test parameters of the multiple test parameter dimensions includes: The test parameters of each dimension in the multiple test parameter dimensions are traversed and combined to generate a combination of test cases that cover the conditions of each scenario.

3. The method according to claim 1, characterized in that, The combination of control simulation devices executing the test cases includes: Based on the wearing posture parameters, control the posture / motion simulation device to wear the device under test at the corresponding wearing angle; Based on the motion state parameters, the posture / motion simulation device is controlled to perform the corresponding motion; Based on the noise parameters, the noise playback device is controlled to play ambient noise at the corresponding sound pressure level; Based on the sound field distance parameter, the sound source device is controlled to play the wake-up word at the corresponding position and angle.

4. The method according to claim 3, characterized in that, The device under test includes an audio acquisition module and a voice wake-up module.

5. The method according to claim 1, characterized in that, The data collected for the wake-up results includes: Collect one or more of the following: wake-up success or failure result data, response latency data, and rejection data.

6. The method according to claim 1, characterized in that, The method further includes: Within the set continuous monitoring period, record false wake-up event information, which includes timestamps and scene attribution information; Based on the false wake-up event information, a statistical model for continuous false wake-ups is established.

7. The method according to claim 1, characterized in that, The wearing posture parameters include one or more of the following: pitch angle parameters, yaw angle parameters, temple opening angle parameters, and displacement parameters; The motion state parameters include motion type parameters, which include one or more of the following: stationary, walking, arm swinging, head turning, and bumping. The noise parameters include noise scene identifiers and noise intensity parameters; The sound field distance parameters include sound source distance classification and sound source incident angle parameters, wherein the sound source incident angle parameters include one or more of horizontal angle parameters and pitch angle parameters.

8. A voice wake-up testing device, characterized in that, include: The parameter acquisition module is used to acquire test parameters of multiple test parameter dimensions, including: wearing posture parameters, motion state parameters, noise parameters, and sound field distance parameters. The test case generation module is used to generate combinations of test cases based on the test parameters of the multiple test parameter dimensions. The test execution module is used to control the simulation device to execute the combination of the test cases to simulate the corresponding wearing posture and movement state, play the corresponding environmental noise, and play the wake-up word in the environmental noise according to the sound field distance parameter; The report generation module is used to collect wake-up result data and generate a voice wake-up test report based on the wake-up result data. The test report includes one or more of the following: wake-up rate, rejection rate, false wake-up rate, response latency, and robustness score.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the voice wake-up test method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the voice wake-up test method as described in any one of claims 1 to 7.