Detection method for intelligent voice interaction recognition
By using robots and large model recognition technology in the testing environment and automatically adjusting the signal-to-noise ratio, the problems of cluttered intelligent voice interaction testing equipment and difficulty in data acquisition are solved, and an efficient and flexible testing solution is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DONGGUAN RISEN XINPU ACOUSTIC TECH CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-17
AI Technical Summary
In existing technologies, intelligent voice interaction testing equipment is disorganized, has a poor acoustic environment, cannot cover all products under test, and the testing process is time-consuming and labor-intensive, requiring communication with multiple manufacturers to obtain test data.
The robot moves within the testing environment, using microphones and cameras to collect responses. It then uses a large model to identify keywords, generate a test report, automatically adjust the signal-to-noise ratio, and directly acquire test data without the need for laying tracks or communicating with manufacturers.
It simplifies the testing environment, improves the quality of the acoustic environment, saves testing time and processes, is applicable to various voice interaction products, and reduces testing complexity.
Smart Images

Figure CN121884804A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice device testing technology, and in particular to a detection method for intelligent voice interaction recognition. Background Technology
[0002] In modern electronic products such as smart homes, smart cockpits, mobile phones, and computers, voice wake-up and interactive control are indispensable components of a good human-computer experience.
[0003] Intelligent voice interaction testing involves playing back a corpus of speech signals while controlling the signal-to-noise ratio at the device under test (DUT) based on noise playback, and then capturing and returning the results from the DUT. Typically, this requires playing the speech from different positions and angles to comprehensively simulate real-world usage. Intelligent voice interaction recognition testing includes both intelligence testing and sound quality testing. Intelligence testing includes metrics such as wake-up rate, false wake-up rate, speech recognition capability, voiceprint recognition capability, speech interruption, and response time. Sound quality testing includes metrics such as sound pressure level, frequency response, idle noise, and distortion.
[0004] In existing technologies, intelligent voice interaction testing often employs a four-axis (xyz + horizontal rotation) operating platform constructed using linear motors, a lifting table, and a turntable, upon which a sound source is placed for testing. This method requires laying multiple moving tracks in the laboratory, leading to cluttered equipment and a poor acoustic environment.
[0005] In addition, existing technologies that analyze product response by entering the product developer mode and capturing logs place high demands on the product. Generally, only products in the development stage or with factory firmware can achieve this. In reality, most test products are commercially available finished products, and the factory firmware has long blocked the log capture channel. Existing testing methods are insufficient to cover all products under test, which leads to the need to communicate with multiple component manufacturers to provide background data during the testing process before testing, rather than obtaining it directly, which is time-consuming and labor-intensive.
[0006] The above problems urgently need to be addressed. Summary of the Invention
[0007] This invention discloses a detection method for intelligent voice interaction recognition, which aims to solve the technical problems existing in the prior art.
[0008] The present invention employs the following technical solution, comprising: calibrating the background noise equalization in the detection environment based on a microphone within the detection environment; setting a robot to move according to predetermined coordinate positions and playing speech data at any point within the predetermined coordinate positions, wherein the predetermined coordinate positions include multiple point positions; collecting response content transmitted from an intelligent voice interaction device based on the robot's built-in microphone and camera; transmitting the response content to a large model, identifying keywords in the response content based on the large model, and recording the response content and the keywords; determining whether there are any unreached point positions within the predetermined coordinate positions; if there are unreached point positions within the predetermined coordinate positions, moving the robot to the unreached point position and playing speech data; if there are no unreached point positions within the predetermined coordinate positions, outputting the response content and keywords collected at all point positions to form a detection report.
[0009] Optionally, the method further includes: after the robot moves to a position point in the predetermined coordinates and plays the speech data, if the factory firmware interface of the intelligent voice interaction device is not closed, connecting the factory firmware interface with the terminal, obtaining the log of the intelligent voice interaction device, and recording the log, wherein the log includes the response content of the intelligent voice interaction device.
[0010] Optionally, the microphones in the detection environment include four active full-range speakers and one active subwoofer. The four active full-range speakers are placed in the four corners of the detection environment, and the active subwoofer is placed between any two adjacent active full-range speakers in the detection environment.
[0011] Optionally, calibrating the background noise equalization in the detection environment based on the microphones within the detection environment includes: calibrating the sensitivity of the microphones; calibrating the active subwoofer individually, maintaining the bandwidth of the active subwoofer between 50Hz and 125Hz; sequentially calibrating four active full-range speakers, maintaining the bandwidth of the active full-range speakers between 125Hz and 10kHz; selecting any wideband noise to verify the calibration result; ending the calibration if the calibration result meets a preset standard; and recalibrating if the calibration result does not meet the preset standard.
[0012] Optionally, the step of sequentially calibrating the four active full-range speakers to maintain their bandwidth between 125Hz and 10kHz includes: performing single calibration on each of the four active full-range speakers sequentially to maintain their bandwidth between 125Hz and 10kHz; performing double calibration on any two adjacent active full-range speakers and then performing double calibration on the remaining two active full-range speakers to maintain their bandwidth between 125Hz and 10kHz; and performing quadruple calibration on all four active full-range speakers simultaneously to maintain their bandwidth between 125Hz and 10kHz.
[0013] Optionally, after the robot moves to a predetermined coordinate position and before playing the speech corpus at any point within the predetermined coordinate position, the method further includes: measuring the background noise at the intelligent voice interaction device with a first sound pressure level; stopping the background noise, the robot playing the speech corpus, and measuring the second sound pressure level at the intelligent voice interaction device; calculating the actual signal-to-noise ratio based on the first sound pressure level and the second sound pressure level; determining the difference between the actual signal-to-noise ratio and the target signal-to-noise ratio range, and adjusting the channel volume of the robot playing the speech corpus based on the difference.
[0014] Optionally, determining the difference between the actual signal-to-noise ratio and the target signal-to-noise ratio range, and adjusting the channel volume of the robot playing the corpus based on the difference, includes: if the actual signal-to-noise ratio is within the target signal-to-noise ratio range, not adjusting the channel volume of the robot playing the corpus; if the actual signal-to-noise ratio is less than the target signal-to-noise ratio range, lowering the channel volume of the robot playing the corpus; and if the actual signal-to-noise ratio is greater than the target signal-to-noise ratio range, increasing the channel volume of the robot playing the corpus.
[0015] Optionally, before setting the robot to move according to a predetermined coordinate position, the method further includes: calibrating the intrinsic parameters of the robot's built-in camera; recording the detection environment based on the camera and the robot's built-in inertial measurement unit to complete the extrinsic parameter calibration; controlling the robot to move autonomously within the detection environment, with the camera continuously acquiring images within the detection environment and the inertial measurement unit recording the robot's posture changes; using the SLAM algorithm to extract and match feature points from the acquired images, calculating the robot's own pose in conjunction with the robot's posture changes, and constructing an initial two-dimensional grid map of the detection environment with the initial position point as the origin; eliminating the accumulated error of the map through a loop closure detection algorithm, correcting the coordinate deviation in the initial two-dimensional grid map, and generating a two-dimensional grid map.
[0016] Optionally, based on the robot's built-in microphone and camera, the system collects the response content transmitted by the intelligent voice interaction device, including: the robot's built-in microphone collects the response audio emitted by the intelligent voice interaction device at a sampling rate of 48kHz and a precision of 24bit, and performs real-time noise reduction and echo cancellation processing to filter the background noise; the continuous response audio is segmented into multiple independent audio segments according to voice pauses and stored as response content, wherein the voice pauses are used to indicate audio segments with a silence duration of ≥200ms; the behavioral characteristics of the intelligent voice interaction device are pre-recorded, and the coordinate position of the intelligent voice interaction device is collected through the camera; the robot's built-in camera collects the images displayed on the screen of the intelligent voice interaction device and the behavioral actions of the intelligent voice interaction device in real time; the text on the images displayed on the screen of the intelligent voice interaction device is extracted using OCR text recognition technology, and the behavioral actions and behavioral characteristics of the intelligent voice interaction device are matched using dynamic recognition technology and stored as response content.
[0017] Optionally, the response content is transmitted to a large model, and keywords in the response content are identified based on the large model. The response content and the keywords are then recorded. This includes: establishing a communication link between the robot acquisition module and the large model, wherein the robot acquisition module includes a built-in microphone and camera; encapsulating the acquired response content into a standard JSON format and transmitting it to the large model; the large model removing redundant characters from the response content and correcting typos generated during identification; using built-in prompt words in the large model, and extracting initial keywords from the response content based on the prompt words; performing confidence filtering on the initial keywords, retaining keywords with a confidence score greater than 0.8 output by the large model; and storing and recording the response content and the keywords.
[0018] The technical solution adopted in this invention can achieve at least one of the following beneficial effects: In this embodiment of the invention, by setting up a microphone and a camera, the response content in the intelligent voice interaction device is automatically identified. At the same time, the response content is input into a large model for processing. Test data can be obtained directly after the test, without having to contact multiple manufacturers, saving test time and test process. In addition, setting up a mobile robot to walk in the testing environment eliminates the need to lay tracks, effectively reducing the complexity of the testing environment and improving the quality of the acoustic environment. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below, forming part of the present invention. The illustrative embodiments of the present invention and their descriptions explain the present invention and do not constitute an improper limitation of the present invention. In the accompanying drawings: Figure 1 This is a flowchart of a detection method for intelligent voice interaction recognition in Embodiment 1 of the present invention; Figure 2 This is a diagram showing the layout of the background noise detection equipment in an intelligent voice interaction recognition detection method according to Embodiment 1 of the present invention. Figure 3 This is a flowchart of noise equalization in a detection method for intelligent voice interaction recognition according to Embodiment 1 of the present invention; Figure 4 This is a flowchart of the signal-to-noise ratio adjustment at the intelligent voice interaction device in the detection method for intelligent voice interaction recognition according to Embodiment 1 of the present invention; Figure 5 This is a flowchart of an optional intelligent voice interaction recognition detection method in Embodiment 2 of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. In the description of this invention, it should be noted that the term "or" is generally used to include the meaning of "and / or," unless otherwise expressly indicated.
[0021] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or a magnetic connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances. Furthermore, in the description of this application, the terms "first," "second," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance. In the description of this invention, "a plurality of" means at least two, such as two, three, or more, unless otherwise explicitly specified.
[0022] Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0023] First, to facilitate understanding of the embodiments of the present invention, some terms or nouns involved in the present invention will be explained below: An active full-range speaker is an integrated speaker device with a built-in power amplifier that can cover the entire audio range (20Hz-20kHz, the complete frequency range that the human ear can hear) and can be used directly without the need for an additional power amplifier.
[0024] An active subwoofer is an audio device specifically designed to reproduce low-frequency sounds (typically 20Hz-200Hz). It has a built-in power amplifier and is mainly used to supplement the low-frequency extension and impact of the audio system.
[0025] To address the problems existing in related technologies, this application provides a detection method for intelligent voice interaction recognition.
[0026] Example 1 This embodiment provides a detection method for intelligent voice interaction recognition, such as... Figure 1 As shown, Figure 1 This is a flowchart of a detection method for intelligent voice interaction recognition according to Embodiment 1 of the present invention. The method includes: Step S102: Based on the microphones in the detection environment, calibrate the background noise equalization in the detection environment; Optionally, background noise uniformity can be detected by collecting noise data from multiple microphones at multiple measurement points, analyzing the spatial distribution differences of sound pressure levels, and judging and correcting the uniformity deviation of environmental noise.
[0027] Optionally, when testing the background noise uniformity in the testing environment, all test equipment and ventilation systems within the testing environment must be shut down (maintaining a minimum airflow if necessary) to avoid momentary interference such as personnel movement and equipment start-up and shutdown. The static time should be no less than 10 minutes to ensure noise stability. All microphones used for testing must be calibrated for sensitivity beforehand using a sound level calibrator (such as a piston generator). The calibration frequency should be selected from the core frequency band of the test (e.g., 1kHz), and the calibration coefficients should be recorded to ensure microphone consistency. Connect the microphones to a sound level meter or data acquisition instrument, and set frequency weighting (A-weighting, conforming to noise testing standards) and time weighting (slow setting S, reducing the impact of momentary fluctuations). The sampling frequency should be no less than 20kHz, and the sampling duration should be 10 to 30 seconds for each measurement point to ensure data representativeness.
[0028] In some preferred embodiments, the microphones in the detection environment include four active full-range speakers and one active subwoofer. The four active full-range speakers are placed in the four corners of the detection environment, and the active subwoofer is placed between any two adjacent active full-range speakers in the detection environment. Figure 2 As shown, Figure 2 This is a diagram showing the layout of the background noise detection equipment in an intelligent voice interaction recognition detection method according to Embodiment 1 of the present invention.
[0029] Optionally, background noise uniformity mainly depends on the spatial distribution uniformity. Measurement points should cover key locations within the testing area. Using the geometric center of the testing environment as the origin, points should be placed using a grid method, such as a 5×5 or 3×3 grid. The spacing between measurement points is determined by the size of the space (0.5 to 1m for small laboratories, 1 to 2m for large spaces). In small testing environments, measurement points can be placed in the corners of the area, i.e., active full-range speakers can be placed in the four corners of the room.
[0030] Specifically, the operating frequency band of active full-range speakers is 200Hz to 20kHz, and some can extend down to 100Hz. They can simulate most of the clearly identifiable sounds in daily life, such as interpersonal conversations, radio broadcasts, and children's laughter. They can also simulate the ticking of clocks, keyboard typing, and car horns. By using active full-range speakers to simulate these everyday sounds, noise interference can be created in daily life.
[0031] In addition, active subwoofers simulate low-frequency sound waves, such as the rumble of thunder, the sound of distant cannons, the idling sound of a car engine, and the low-frequency roar of an airplane taking off. By simulating relatively low-frequency sounds in life, active subwoofers can create noise interference in the testing environment, allowing intelligent voice interaction devices to recognize the speech data produced by robots in various noise environments.
[0032] In some preferred embodiments, the background noise equalization in the detection environment is calibrated based on the microphones within the detection environment, including: calibrating the microphone sensitivity; calibrating a single active subwoofer, maintaining its bandwidth between 50Hz and 125Hz; sequentially calibrating four active full-range speakers, maintaining their bandwidth between 125Hz and 10kHz; selecting any broadband noise level to verify the calibration results; ending the calibration if the results meet a preset standard; and recalibrating if the results do not meet the preset standard. Figure 3 As shown, Figure 3 This is a flowchart of noise equalization in a detection method for intelligent voice interaction recognition according to Embodiment 1 of the present invention.
[0033] In some preferred embodiments, four active full-range speakers are calibrated sequentially to maintain their bandwidth between 125Hz and 10kHz. This includes: performing single calibration on each of the four active full-range speakers sequentially to maintain their bandwidth between 125Hz and 10kHz; performing double calibration on any two adjacent active full-range speakers and then performing double calibration on the remaining two active full-range speakers to maintain their bandwidth between 125Hz and 10kHz; and performing quadruple calibration on all four active full-range speakers simultaneously to maintain their bandwidth between 125Hz and 10kHz.
[0034] Step S104: Set the robot to move according to the predetermined coordinate position and play the corpus at any position point in the predetermined coordinate position, wherein the predetermined coordinate position includes multiple position points; Optionally, the robot mainly consists of three parts: a mouth simulator, a processing core, and an intelligent chassis. The robot has a built-in AM3100 active mouth simulator, primarily used for playing speech data. The processing core includes a high-performance industrial computer, an RS1224 analog signal transceiver unit, an external communication plug-in, and a battery pack. The robot also features two RST4000 standard microphones and dual cameras (top and bottom) for receiving external sound and image signals. The processing core is the robot's brain, running the Sega SMS (MasterSystem) self-control version, receiving commands and controlling the robot's operation, movement, and speech. The intelligent chassis is equipped with powered wheels and features automatic mapping, laser obstacle avoidance, and automatic recharging when the battery is low, responsible for the robot's movement.
[0035] Optionally, the robot models the test room using a smart chassis, and can then be directed to any coordinates and orientation via a terminal. The robot is equipped with laser obstacle avoidance and low battery warnings, and returns to its charging dock to reduce unexpected malfunctions during testing.
[0036] In some preferred embodiments, after the robot moves to a predetermined coordinate position and before playing the speech data at any point within the predetermined coordinate position, the method further includes: measuring a first sound pressure level of background noise at the intelligent voice interaction device; stopping the background noise, the robot playing the speech data, and measuring a second sound pressure level at the intelligent voice interaction device; calculating the actual signal-to-noise ratio based on the first and second sound pressure levels; determining the difference between the actual signal-to-noise ratio and a target signal-to-noise ratio range, and adjusting the channel volume of the robot playing the speech data based on the difference. Figure 4 As shown, Figure 4 This is a flowchart of the signal-to-noise ratio adjustment at the intelligent voice interaction device in the detection method for intelligent voice interaction recognition according to Embodiment 1 of the present invention.
[0037] Optionally, the signal-to-noise ratio can be calculated by measuring the background noise and the sound pressure of the speech corpus, and then the volume can be dynamically adjusted to match the target signal-to-noise ratio range to ensure that the intelligent voice interaction device can clearly recognize the speech corpus.
[0038] Optional, signal-to-noise ratio is the corpus sound pressure level. sound pressure level with background noise The difference is commonly calculated in decibels in acoustics, and the formula is: by =40dB(A), Taking 60dB(A) as an example, at this time... =20dB.
[0039] This indicates that the higher the signal-to-noise ratio, the higher the accuracy of speech recognition; the target signal-to-noise ratio range for intelligent voice interaction is 15 to 30 dB.
[0040] In some preferred embodiments, the difference between the actual signal-to-noise ratio and the target signal-to-noise ratio range is determined, and the channel volume of the robot playing the corpus is adjusted based on the difference, including: if the actual signal-to-noise ratio is within the target signal-to-noise ratio range, the channel volume of the robot playing the corpus does not need to be adjusted; if the actual signal-to-noise ratio is less than the target signal-to-noise ratio range, the channel volume of the robot playing the corpus is reduced; if the actual signal-to-noise ratio is greater than the target signal-to-noise ratio range, the channel volume of the robot playing the corpus is increased.
[0041] Step S106: Based on the robot's built-in microphone and camera, collect the response content transmitted by the intelligent voice interaction device. Optionally, to accurately obtain the response results of voice interaction, the intelligent voice robot is equipped with two low-noise microphones and two binocular cameras. When the intelligent voice interaction device cannot obtain the response status by directly capturing logs, it can simulate a real person's response by recording audio or taking pictures. This result is then converted into a recognition result by connecting to a local large model and sent back to the MasterSystem self-control board, and finally transmitted back to the client wirelessly.
[0042] In some preferred embodiments, after the robot moves to a location point at a predetermined coordinate position and plays the speech data, if the factory firmware interface of the intelligent voice interaction device is not closed, the robot connects the factory firmware interface to the terminal, obtains the logs of the intelligent voice interaction device, and records the logs, which include the response content of the intelligent voice interaction device.
[0043] It should be noted that the response format of the intelligent voice interaction device (DUT) is determined manually in advance, and the system is currently unable to automatically recognize the response format.
[0044] Optionally, the robot transmits the speech data, and then the DUT responds (taking voice response as an example). The microphone on the robot will continuously record the DUT's response after the speech data is played, and send the recorded speech data to the large model to be converted into text. Finally, the large model searches for keywords in the text result and determines whether the response was successful.
[0045] Optionally, for devices with displays but no voice feedback, such as smart screens, if the robot's camera cannot directly capture the speech recognition results on the display during the testing process, a fixed camera can be installed in front of the display to capture and recognize the speech response. This camera is directly connected to the client computer.
[0046] In some preferred embodiments, the intrinsic parameters of the robot's built-in camera are calibrated, and the detection environment is recorded based on the camera and the robot's built-in inertial measurement unit to complete the extrinsic parameter calibration. The robot is controlled to move autonomously within the detection environment, the camera continuously acquires images within the detection environment, and the inertial measurement unit records the robot's posture changes. The SLAM algorithm is used to extract and match feature points from the acquired images, and the robot's own pose is calculated in combination with the robot's posture changes. An initial two-dimensional grid map of the detection environment is constructed with the initial position point as the origin. The cumulative error of the map is eliminated by the loop closure detection algorithm, the coordinate deviation in the initial two-dimensional grid map is corrected, and a two-dimensional grid map is generated.
[0047] In some preferred embodiments, the robot's built-in microphone and camera collect the response content transmitted by the intelligent voice interaction device, including: the robot's built-in microphone collects the response audio emitted by the intelligent voice interaction device at a sampling rate of 48kHz and a precision of 24bit, and performs noise reduction and echo cancellation processing in real time to filter background noise; the continuous response audio is segmented into multiple independent audio segments according to the voice pauses and stored as response content, wherein the voice pauses are used to indicate audio segments with a silence duration of ≥200ms; the behavioral characteristics of the intelligent voice interaction device are pre-recorded, and the coordinate position of the intelligent voice interaction device is collected through the camera; the robot's built-in camera collects the images displayed on the screen of the intelligent voice interaction device and the behavior of the intelligent voice interaction device in real time; the text on the images displayed on the screen of the intelligent voice interaction device is extracted through OCR text recognition technology, and the behavior of the intelligent voice interaction device is matched with the behavioral characteristics through dynamic recognition technology and stored as response content.
[0048] Step S108: Transmit the response content to the large model, identify keywords in the response content based on the large model, and record the response content and keywords. In some preferred embodiments, the response content is transmitted to a large model, and keywords in the response content are identified based on the large model. The response content and keywords are then recorded. This includes: establishing a communication link between the robot acquisition module and the large model, wherein the robot acquisition module includes a built-in microphone and camera; encapsulating the acquired response content into a standard JSON format and transmitting it to the large model; removing redundant characters from the response content and correcting typos generated during identification in the response content; using built-in prompt words in the large model, the large model extracts initial keywords from the response content based on the prompt words; performing confidence screening on the initial keywords and retaining keywords with a confidence score greater than 0.8 output by the large model; and storing and recording the response content and keywords.
[0049] Optionally, the large model can be directly deployed inside the robot. In this case, the TCP / IP protocol can be adopted to ensure low latency and high stability. It can also be used as a cloud-based large model, and in this case, the HTTP / HTTPS protocol is adopted, and the API interface is used to achieve data transmission. An authentication key (Token) needs to be configured to prevent the leakage of transmitted data.
[0050] Optionally, the radio collects voice responses, which are converted into text data through the built-in automatic speech recognition (ASR) module of the robot. The camera collects visual information (such as user gestures and expressions), which is converted into text descriptions (such as "the user waves their hand to indicate refusal" and "the user nods to indicate confirmation") through an image recognition model. The collected response content is encapsulated in JSON format. The encapsulation needs to include metadata and the response content to ensure that the large model can parse it.
[0051] Optionally, after receiving the JSON data, the large model performs content preprocessing: First, redundant characters are removed, specifically by filtering out meaningless symbols in the text (such as filler words like "um" and "ah", repeated punctuation, and random whitespace). In addition, spelling mistakes are corrected, that is, recognition errors are corrected based on the context semantics (such as changing "tiaojie" to "adjust" and "diantou" to "nod"). Prompt words are built into the large model, and when configuring the prompt words, the extraction rules need to be clearly defined to adapt to the voice interaction scenario. The large model extracts initial keywords based on the above prompt words. The large model extracts initial keywords from the processed content. For example, for the processed content: "Hello, the volume adjustment is okay, and I nod to confirm", the initial keywords are: volume adjustment, confirmation, nod, volume.
[0052] Optionally, the large model outputs a confidence level for each initial keyword (reflecting the degree of association between the keyword and the core semantics of the text, with a value ranging from 0 to 1); the screening rule is to retain keywords with a confidence level > 0.8 and eliminate redundant words with low confidence levels. Specifically as follows: The large model outputs the initial keyword + confidence level as: volume adjustment(0.95), confirmation(0.90), nod(0.75), volume(0.85). At this time, the screened keywords are volume adjustment, confirmation, volume. And the above-screened keywords are stored.
[0053] Step S110, determine whether there are unarrived position points among the predetermined coordinate positions. If there are unarrived position points among the predetermined coordinate positions, the robot moves to the unarrived position point to play the corpus. If there are no unarrived position points among the predetermined coordinate positions, the response content and keywords collected at all position points are output to form a detection report.
[0054] Through steps S102 to S110 above, based on this robot and combined with a background noise playback system, a complete and easy-to-use intelligent automated voice testing solution is generated. It directly utilizes robot mobility, offering advantages such as fast installation and construction, low cost, precise control, and ease of later adjustment. Furthermore, the overall environment setup is flexible, debugging is fast, and efficiency is high. Simultaneously, the testing environment also features automatic signal-to-noise ratio adjustment to meet the testing requirements of products under test with different signal-to-noise ratios. In addition, for acquiring interaction results, a combination of recording, screen capture, and AI large-scale model recognition is used to seamlessly acquire interaction results from a human-like perspective. No restrictions are placed on its firmware or open interface format, making it suitable for all voice interaction products.
[0055] Example 2 Based on the above embodiments, the present invention also proposes an optional implementation method. Figure 5 This is a flowchart of an optional intelligent voice interaction recognition detection method according to Embodiment 2 of the present invention, as shown below. Figure 5 As shown, the method includes: Step S1: In the testing environment (listening room), four active full-range speakers and one active subwoofer are arranged in accordance with the requirements of the ETSI ES 202 396-1 standard. Step S2: Place a standard microphone at the location of the intelligent voice interaction device (product under test, DUT); Step S3: Connect all standard microphone hardware and terminal devices; Step S4: Noise equalization in the room is achieved through a standard microphone running on the terminal device. Step S5: The terminal and the robot connect to each other via a wireless network; Step S6: Adjust the signal-to-noise ratio at the intelligent voice interaction device; Step S7: Control the robot to move to the test point at the preset coordinates; Step S8: Control the robot to speak, playing the wake word and corpus; Step S9 involves capturing the response status of the intelligent voice interaction device through methods such as ADB log capture, microphone recording, and camera capture. ADB (Android Debug Bridge) allows developers or testers to perform various operations on the Android device via USB or wireless connection, such as installing / uninstalling applications, transferring files, executing shell commands, and capturing system logs. Logs are Android system or application runtime logs, automatically generated text records during device operation, containing key information such as system status, application behavior, error messages, and event triggers.
[0056] Step S10: Complete a single test.
[0057] By following steps S1 to S10 above, the entire corpus test can be completed. It supports multi-corpus playback and multi-point playback. The test results will be automatically statistically analyzed, and a complete corpus test report will be output. It also comes with a built-in noise library (simulating typical noises in different occasions such as human voices, traffic, offices, and cafes), and can also provide wake words and commonly used corpora. The corpus playback supports calling the TTS speech engine and using text-to-speech conversion to play multiple types of corpora (different accents for men, women, and children).
[0058] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A detection method for intelligent voice interaction recognition, characterized in that, include: Based on the microphones in the detection environment, the background noise equalization in the detection environment is calibrated. The robot is programmed to move according to predetermined coordinate positions and play speech data at any point within the predetermined coordinate positions, wherein the predetermined coordinate positions include multiple points. Based on the built-in microphone and camera of the robot, the response content transmitted by the intelligent voice interaction device is collected. The response content is transmitted to a large model, and keywords in the response content are identified based on the large model. The response content and the keywords are then recorded. If there are any unreached locations within the predetermined coordinates, the robot moves to the unreached location and plays the audio corpus. If there are no unreached locations within the predetermined coordinates, the robot outputs the collected responses and keywords from all locations to form a detection report.
2. The detection method for intelligent voice interaction recognition according to claim 1, characterized in that, The method also includes: After the robot moves to the predetermined coordinate position and plays the speech data, if the factory firmware interface of the intelligent voice interaction device is not closed, the robot connects the factory firmware interface to the terminal, obtains the log of the intelligent voice interaction device, and records the log, wherein the log includes the response content of the intelligent voice interaction device.
3. The detection method for intelligent voice interaction recognition according to claim 1, characterized in that, The microphones in the detection environment include four active full-range speakers and one active subwoofer. The four active full-range speakers are placed in the four corners of the detection environment, and the active subwoofer is placed between any two adjacent active full-range speakers in the detection environment.
4. The detection method for intelligent voice interaction recognition according to claim 3, characterized in that, The calibration of background noise equalization in the detection environment based on the microphone within the detection environment includes: Calibrate the sensitivity of the microphone; The active subwoofer is calibrated individually to maintain its bandwidth between 50Hz and 125Hz. The four active full-range speakers were calibrated sequentially to maintain their bandwidth between 125Hz and 10kHz. Select any broadband noise verification calibration result. If the calibration result meets the preset standard, end the calibration. If the calibration result does not meet the preset standard, recalibrate.
5. The detection method for intelligent voice interaction recognition according to claim 4, characterized in that, The sequential calibration of the four active full-range speakers, maintaining their bandwidth between 125Hz and 10kHz, includes: Each of the four active full-range speakers was individually calibrated in turn, maintaining a bandwidth between 125Hz and 10kHz. Double calibration is performed on any two adjacent active full-range speakers among the four active full-range speakers, and then double calibration is performed on the remaining two active full-range speakers, keeping the bandwidth between 125Hz and 10kHz. The four active full-range speakers were simultaneously quad-calibrated to maintain a bandwidth between 125Hz and 10kHz.
6. The detection method for intelligent voice interaction recognition according to claim 1, characterized in that, After the robot moves to a predetermined coordinate position, and before playing the corpus at any point within the predetermined coordinate position, the method further includes: The background noise is measured at the first sound pressure level of the intelligent voice interaction device. The background noise is stopped, the robot plays the speech corpus, and the second sound pressure level at the intelligent voice interaction device is measured. The actual signal-to-noise ratio is calculated based on the first sound pressure level and the second sound pressure level. Determine the difference between the actual signal-to-noise ratio and the target signal-to-noise ratio range, and adjust the channel volume of the robot playing the corpus based on the difference.
7. The detection method for intelligent voice interaction recognition according to claim 6, characterized in that, Determine the difference between the actual signal-to-noise ratio and the target signal-to-noise ratio range, and adjust the channel volume of the robot playing the corpus based on the difference, including: When the actual signal-to-noise ratio is within the target signal-to-noise ratio range, it is not necessary to adjust the channel volume of the robot playing the corpus; If the actual signal-to-noise ratio is less than the target signal-to-noise ratio range, the volume of the channel through which the robot plays the corpus is reduced. If the actual signal-to-noise ratio is greater than the target signal-to-noise ratio range, increase the channel volume of the robot playing the corpus.
8. The detection method for intelligent voice interaction recognition according to claim 1, characterized in that, Before setting the robot to move to a predetermined coordinate position, the method also includes: The internal parameters of the robot's built-in camera are calibrated, and the detection environment is recorded based on the camera and the robot's built-in inertial measurement unit to complete the external parameter calibration. The robot is controlled to move autonomously within the detection environment, the camera continuously acquires images of the detection environment, and the inertial measurement unit records the robot's posture changes; The SLAM algorithm is used to extract and match feature points in the acquired images. Combined with the robot's posture changes, the robot's own pose is calculated. Using the initial position point as the origin, an initial two-dimensional grid map of the detection environment is constructed. The cumulative error of the map is eliminated by the loop closure detection algorithm, the coordinate deviation in the initial two-dimensional raster map is corrected, and a two-dimensional raster map is generated.
9. The detection method for intelligent voice interaction recognition according to claim 1, characterized in that, Based on the robot's built-in microphone and camera, the system collects responses from intelligent voice interaction devices, including: The robot's built-in microphone collects the response audio emitted by the intelligent voice interaction device at a sampling rate of 48kHz and a precision of 24bit, and performs noise reduction and echo cancellation processing in real time to filter the background noise. The continuous response audio is segmented into multiple independent audio segments based on the voice pauses and stored as response content. The voice pauses are used to indicate audio segments with a silence duration of ≥200ms. The behavioral characteristics of the intelligent voice interaction device are pre-recorded, and the coordinate position of the intelligent voice interaction device is collected through the camera. The robot's built-in camera captures images displayed on the screen of the intelligent voice interaction device and the device's actions in real time. The text on the screen displayed in the intelligent voice interaction device is extracted using OCR text recognition technology, and the behavior and characteristics of the intelligent voice interaction device are matched using dynamic recognition technology and stored as response content.
10. The detection method for intelligent voice interaction recognition according to claim 1, characterized in that, The response content is transmitted to a large model, and keywords in the response content are identified based on the large model. The response content and the keywords are then recorded, including: Establish a communication link between the robot data acquisition module and the large model, wherein the robot data acquisition module includes a built-in microphone and a camera; The collected response content is encapsulated in standard JSON format and transmitted to the large model; The large model removes redundant characters from the response content and corrects typos generated during the identification process. The large model has built-in prompt words, and the large model extracts initial keywords from the response content based on the prompt words; The initial keywords are filtered by confidence, and keywords with a confidence score greater than 0.8 from the output of the large model are retained; The response content and the keywords are stored and recorded.