Voice environment detection method and computer readable storage medium

By using a voice environment detection method, the process of debugging voice-activated song selection is simplified, the user threshold is lowered and the troubleshooting cost for professionals is reduced, the user experience is improved, and autonomous debugging and troubleshooting of the voice environment are achieved.

CN117496949BActive Publication Date: 2025-12-05FUJIAN STAR NET EVIDEO INFORMATION SYST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311481405.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-08
Publication Date
2025-12-05
Estimated Expiration
2043-11-08

AI Technical Summary

Technical Problem

The current voice-activated song selection process lacks intuitive debugging tools, resulting in a poor user experience, a high barrier to entry, high costs for professional troubleshooting, and the voice recognition effect is greatly affected by environmental factors.

Method used

A method for detecting a speech environment is provided. By receiving test audio, volume detection and speech recognition are performed to determine whether the speech environment is acceptable. This includes the detection of volume range and speech recognition results, and a visual interface is used to guide the user in debugging.

Benefits of technology

It simplifies the voice detection process, lowers the barrier to entry for users and reduces the cost of troubleshooting, improves the user experience, and helps ordinary users solve voice environment problems independently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117496949B_ABST
    Figure CN117496949B_ABST
Patent Text Reader

Abstract

The application discloses a voice environment detection method and a computer readable storage medium. The method comprises the following steps: receiving test audio; performing volume detection on the test audio to determine the maximum decibel value of the test audio; performing voice recognition on the test audio to obtain a voice recognition result, and performing voice recognition detection according to the voice recognition result and the content of the test audio; if a preset condition is met, determining that the voice environment detection is passed, wherein the preset condition comprises that the voice recognition detection is passed and the maximum decibel value of the test audio is within a preset volume range. The application can simplify the voice detection process, reduce the user threshold and problem troubleshooting cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech recognition testing technology, and in particular to a speech environment detection method and a computer-readable storage medium. Background Technology

[0002] In the KTV industry, the effectiveness of voice-activated song selection is greatly affected by environmental factors. While the test results are good in the testing environment, after actual use, there are frequent reports of insensitive voice recognition, poor recognition rate, and even no response to voice input. The market response is not in line with expectations.

[0003] Voice-activated karaoke setup and debugging is highly specialized and requires professional personnel to operate. From a hardware perspective, there are numerous recording wiring solutions involved in voice-activated karaoke setup and debugging, including common methods such as voice cable wiring, voice box wiring, effects processor wiring, and center speaker wiring. Recording devices and interfaces are also diverse, with common devices including BBS microphones, wired microphones, and USB microphones, and common recording input ports such as RCA connectors, 3.5mm, and 6.0 jacks. From a software perspective, voice-activated karaoke setup and debugging involves network-related issues such as network fluctuations and timeouts, as well as software authentication, initialization, and operational problems. Therefore, all existing voice-activated karaoke setup and debugging methods require prior setup and testing by professional personnel.

[0004] Voice-activated karaoke debugging presents numerous and complex voice environment issues, making troubleshooting difficult. Taking voice cable wiring as an example, this method is widely used due to its simplicity and low cost. The specific wiring is as follows: one microphone output goes to the effects unit and then to the amplifier output, while the other is directly input to the karaoke machine via the AV port. Because the microphone output is split into two paths, there is a discrepancy between the software recording effect and what the human ear hears. The effect heard by the human ear is not the actual recording effect, interfering with the user's judgment of whether the recording environment is normal. Furthermore, the microphone input volume to the effects unit cannot be too high, otherwise it can easily cause feedback from the amplifier. For example, if the maximum microphone volume is 31, it can generally only be set to around 20, with slight variations depending on the effects unit and amplifier. In the link from microphone to effects unit / amplifier, sound quality is affected by gain and effects processing; even if the microphone output volume is low, a good listening experience can be achieved by adjusting the effects unit and amplifier. However, from microphone to karaoke machine, the input volume has a significant impact on voice recognition performance. For example, using the same sound source, connecting it to different machines via a microphone splitter, recording PCM, and comparing the waveforms... Figure 1-2 As shown, where, Figure 1 This is a normal waveform. Figure 2The graph shows the waveform of the volume loss, with the x-axis representing time (s) and the y-axis representing decibels (dDFS). From the recording input to the PCB circuit environment and then to the software recording, there is an unavoidable volume loss in the link itself, which varies depending on the hardware environment. In some of the tested machines, the volume loss was greater than 30%. Due to amplifier feedback and volume loss, the problem of low input volume is even more pronounced.

[0005] Since there are currently no intuitive debugging tools for voice-activated song selection, users often have no way of knowing where the problem lies when voice recognition fails to respond. This results in a high barrier to entry for the voice-activated song selection function and a poor user experience. For technicians, the cost of implementation and troubleshooting is high. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a voice environment detection method and a computer-readable storage medium, which can simplify the voice detection process, reduce the user threshold and the cost of troubleshooting.

[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a voice environment detection method, comprising:

[0008] Receive test audio;

[0009] The volume of the test audio is detected to determine the maximum decibel value of the test audio.

[0010] The test audio is subjected to speech recognition to obtain a speech recognition result, and speech recognition detection is performed based on the speech recognition result and the content of the test audio.

[0011] If the preset conditions are met, the voice environment detection is deemed to have passed. The preset conditions include passing the voice recognition detection and the maximum decibel value of the test audio being within a preset volume range.

[0012] The present invention also proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described above.

[0013] The beneficial effects of this invention are as follows: by guiding users step by step through an automated detection process to troubleshoot common problems in voice song selection, the user threshold is lowered, allowing ordinary users to detect and handle problems themselves, thereby improving the user experience. At the same time, it reduces the cost of construction and troubleshooting for technicians. Attached Figure Description

[0014] Figure 1 This is a schematic diagram of a normal waveform;

[0015] Figure 2 This is a schematic diagram of the loss waveform;

[0016] Figure 3 This is a flowchart of a speech environment detection method according to the present invention;

[0017] Figure 4 This is a flowchart of the method according to Embodiment 1 of the present invention;

[0018] Figure 5 This is a schematic diagram of the voice environment detection interface according to Embodiment 1 of the present invention. Detailed Implementation

[0019] To explain the technical content, objectives, and effects of the present invention in detail, the following description is provided in conjunction with the embodiments and accompanying drawings.

[0020] Please see Figure 3 A speech environment detection method, comprising:

[0021] Receive test audio;

[0022] The volume of the test audio is detected to determine the maximum decibel value of the test audio.

[0023] The test audio is subjected to speech recognition to obtain a speech recognition result, and speech recognition detection is performed based on the speech recognition result and the content of the test audio.

[0024] If the preset conditions are met, the voice environment detection is deemed to have passed. The preset conditions include passing the voice recognition detection and the maximum decibel value of the test audio being within a preset volume range.

[0025] As can be seen from the above description, the beneficial effects of the present invention are: it can simplify the voice detection process, reduce the user's usage threshold and the cost of troubleshooting.

[0026] Further, the step of performing volume detection on the test audio to determine the maximum decibel value of the test audio includes:

[0027] Based on a preset period, the test audio within the current period is acquired in real time to obtain the audio for the current period;

[0028] The decibel value of the current period audio is calculated based on the dynamic range of the current period audio and the preset sampling precision.

[0029] If the decibel value of the current period audio is greater than the maximum decibel value, then the maximum decibel value is updated based on the current period audio's decibel value, where the initial value of the maximum decibel value is the decibel value of the first period audio.

[0030] Furthermore, the step of performing volume detection on the test audio to determine the maximum decibel value of the test audio also includes:

[0031] The volume dashboard displays the decibel value of the current audio cycle in real time.

[0032] If the maximum decibel value is updated, the updated maximum decibel value will be displayed in the volume dashboard according to the preset display duration.

[0033] As described above, the system can alert the user when the maximum decibel value has changed.

[0034] Further, the step of calculating the decibel value of the current period audio based on the dynamic range of the current period audio and the preset sampling precision specifically involves:

[0035] The decibel value of the current period's audio is calculated using the decibel calculation formula, which is as follows:

[0036] I = 10 × lg(D / 2) x ) 2 ,

[0037] Where I is the decibel value of the current period audio, D is the dynamic range of the current period audio, and x is the preset sampling precision.

[0038] Furthermore, after performing volume detection on the test audio and determining the maximum decibel value of the test audio, the method further includes:

[0039] If the maximum decibel value of the test audio is not within the preset volume range, a prompt will be made indicating that the volume is too high or too low, and the audio signal level will be adjusted or the volume will be adjusted according to the preset gain value.

[0040] Furthermore, adjusting the volume according to the preset gain value specifically involves:

[0041] Based on the preset sampling precision, the test audio is sampled to obtain each audio sampling point;

[0042] Obtain the gain value and simulate the gain value using the tan function to obtain the scaling factor;

[0043] The amplitude of each audio sampling point is scaled according to the scaling factor.

[0044] As described above, by providing an adjustment method for the audio signal scaling factor, users can easily adjust the volume, allowing them to independently resolve issues of the volume being too high or too low.

[0045] Furthermore, the test audio content includes preset voice wake-up words and voice commands;

[0046] The step of performing speech recognition on the test audio to obtain a speech recognition result, and performing speech recognition detection based on the speech recognition result and the content of the test audio, includes:

[0047] The test audio is subjected to speech recognition. If a preset voice wake-up word is recognized, voice command detection is performed to obtain the voice command recognition result.

[0048] If the voice command recognition result matches the preset voice command, the voice recognition detection is deemed successful.

[0049] As can be seen from the above description, the voice wake-up word is recognized first, and then the voice command is recognized after the wake-up word is recognized, which is in line with the actual application scenario.

[0050] Furthermore, the step of performing speech recognition on the test audio to obtain a speech recognition result, and performing speech recognition detection based on the speech recognition result and the content of the test audio, further includes:

[0051] If no voice command matching the preset voice command is detected within a preset time after the preset voice wake-up word is recognized, the voice recognition detection is deemed to have failed.

[0052] As described above, if no voice command is detected within a certain period of time, the recognition is considered to have failed.

[0053] Furthermore, the preset conditions also include passing the runtime environment detection, which includes network fluctuation detection, network timeout detection, authentication detection, initialization detection, and runtime problem detection.

[0054] As described above, by performing runtime environment testing, operational problems of voice software can be identified, thus improving the comprehensiveness of the testing.

[0055] The present invention also proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described above.

[0056] Example 1

[0057] Please refer to Figure 4-5 Embodiment 1 of the present invention is: a voice environment detection method, which can be applied to a KTV song selection system. In practical application scenarios, users can perform voice environment detection based on a voice environment detection interface. This voice environment detection interface may include a service switch area, a volume dashboard area, a voice recognition detection interface area, a microphone volume adjustment area, a voice wake-up word setting area, and a voice broadcast volume setting area. For example... Figure 5 As shown, the volume gauge in the volume gauge area is displayed as a horizontally placed volume bar, divided into 60 equal segments, each representing 1 dBFS. It is displayed from left to right, with the leftmost segment representing -60 dBFS and the rightmost segment representing 0 dBFS.

[0058] like Figure 4 As shown, the method in this embodiment includes the following steps:

[0059] S1: Perform runtime environment checks. These checks include network fluctuation detection, network timeout detection, authentication detection, initialization detection, and runtime problem detection.

[0060] In this embodiment, network fluctuation detection involves pinging a fixed domain name at regular intervals (e.g., every 10 minutes) to determine network connectivity. Network timeout detection involves setting a preset timeout (e.g., 10 seconds) for all existing network requests and then checking whether connection and data transmission / reception operations are completed within that time. Authentication detection involves calling the data center interface during initialization to determine authorization; if authorization fails, the user is notified. Initialization detection checks whether system submodules (such as voice broadcasting and voice recognition engines) initialize correctly. Runtime problem detection detects other software anomalies during initialization.

[0061] The runtime environment detection automatically monitors throughout the entire detection process lifecycle and provides real-time feedback to the user through the interface.

[0062] S2: Receive test audio, which includes preset voice wake-up words and voice commands.

[0063] During testing, the voice environment detection interface will prompt the user with the content of the test audio, such as "Hey Hey Hey, I want to change the song," where "Hey Hey Hey" is the voice wake-up phrase and "change the song" is the voice command. The user speaks the corresponding content into the microphone according to the prompts. In real-world scenarios, the user can repeat the content of the test audio.

[0064] S3: Perform volume detection on the test audio to determine the maximum decibel value of the test audio.

[0065] Specifically, according to a preset period, the test audio within the current period is acquired in real time to obtain the audio for the current period. In this embodiment, the period is 70ms, meaning that the latest test audio within the last 70ms is acquired every 70ms. Then, according to the decibel calculation formula I = 10 × lg(D / 2) x ) 2The algorithm calculates the decibel value of the current audio cycle, where I is the decibel value of the current audio cycle, D is the dynamic range of the current audio cycle, and x is the preset sampling precision. The decibel value of the first audio cycle is used as the initial maximum decibel value. Then, after calculating the decibel value for each subsequent cycle, it is compared with the current maximum decibel value. If the current decibel value is greater than the maximum decibel value, the maximum decibel value is updated, becoming the new maximum decibel value. Finally, after calculating the decibel values ​​for all audio cycles of the test audio, the latest maximum decibel value is taken as the maximum decibel value for the test audio.

[0066] In this embodiment, the sampling precision is 16-bit. The dynamic range of sound that can be recorded in the 16-bit format is 2 to the power of 16, which is 65,536 units. Therefore, the decibel calculation formula can be expressed as:

[0067] I = 10 × lg(D / 65536) 2

[0068] It can be seen that when D < 65536, I is always negative; when D = 65536, I is 0; and when D > 65536, I is greater than 0, which is an abnormal case. In addition, when D equals 0, that is, when there is no voltage level, I equals negative infinity, and the voltage level will also become the notation "-∞".

[0069] Furthermore, the volume dashboard displays the decibel value of the current audio cycle in real time; that is, after calculating the decibel value of the current audio cycle, the display is refreshed and shown on the volume dashboard. For example, assuming the calculated decibel value is -40dBFS, the volume bars corresponding to -60 to -40dBFS in the volume dashboard will be displayed in green, and the volume bars corresponding to -40 to 0dBFS will be displayed in gray.

[0070] Furthermore, if the maximum decibel value is updated, the updated maximum decibel value will be displayed on the volume dashboard according to the preset display duration. That is, when the maximum decibel value is updated, the updated maximum decibel value will be displayed on the volume dashboard in a specific color (such as orange), and then hidden after a certain delay (such as 3 seconds), thereby notifying the user that the maximum decibel value has changed.

[0071] During the voice environment detection interface display, the volume dashboard will display the volume in real time based on the input audio. After one round of volume detection, the maximum decibel value detected during the period will be judged. In this embodiment, if the maximum decibel value is within the range of -40dBFS to -20dBFS, the volume is considered moderate; if the maximum decibel value is within the range of -60dBFS to -40dBFS, the volume is considered low; and if the maximum decibel value is within the range of -20dBFS to 0dBFS, the volume is considered high.

[0072] S4: Perform speech recognition on the test audio to obtain the speech recognition result, and perform speech recognition detection based on the speech recognition result and the content of the test audio.

[0073] Specifically, speech recognition is performed on the test audio. If a preset voice wake-up word is recognized, voice command detection is performed to obtain a voice command recognition result. If the voice command recognition result matches the preset voice command, the speech recognition detection is deemed to have passed.

[0074] In other words, speech recognition first identifies the wake word. When the recognition result completely matches the wake word (such as "Xiao Hai Xiao Hai"), it performs voice command detection. As long as the voice command (such as "Qie Ge") is recognized, the speech recognition detection passes.

[0075] Furthermore, users can repeat the wake word until it is successfully recognized; there is no timeout period before that. Users can also repeat the voice command; if the voice command is not recognized within a certain time (e.g., 8 seconds) after the wake word is recognized, the recognition is considered to have failed, and the voice recognition test is deemed unsuccessful.

[0076] In real-world scenarios, test audio can be transmitted to the recognition engine in real time. The recognition results are then displayed as characters on the speech environment detection interface. The recognition results are matched with the content of the test audio, and the corresponding characters are displayed in different colors based on the matching results. For example, matched characters are displayed in green, and unmatched characters are displayed in red, allowing users to quickly identify the accuracy of the speech recognition results based on the character colors.

[0077] S5: Determine whether the preset conditions are met. If so, proceed to step S6. The preset conditions include passing the operating environment test, passing the voice recognition test, and the maximum decibel value of the test audio being within a preset volume range.

[0078] S6: Voice environment detection passed.

[0079] To ensure a good voice environment for users when using the voice-activated song selection function, users must pass the voice environment detection test on the interface before they can start the voice service. The current voice environment and any detected issues are displayed to the user in a visual way.

[0080] Furthermore, if the speech recognition test fails, a new round of testing will begin, repeating the above steps. If the speech recognition test continues to fail, the audio can be played or exported via the interface buttons for further troubleshooting.

[0081] If the volume detection fails, the user will be prompted that the volume is too high or too low. Furthermore, when the volume detection fails, the voice environment detection interface can automatically pop up a volume adjustment dialog box for visual adjustment. The user can input voice commands according to the interface prompts, adjusting the volume while observing the volume gauge until the maximum decibel value is within a suitable range, at which point the adjustment is complete. Providing a visual method for adjusting volume facilitates user adjustments to the external recording environment.

[0082] Specifically, when the volume is too high or too low, the volume is adjusted. The volume adjustment mainly includes the following two methods: amplifying or reducing the audio signal level by using a hardware level knob, or adjusting the volume according to a preset gain value.

[0083] First: A common hardware solution is to amplify or reduce the audio signal level using a hardware level knob to increase or decrease the volume. In practical applications, this hardware solution is preferred when the volume is too high or too low.

[0084] Second: Adjust the volume according to the preset gain value, that is, use software algorithms to scale the test audio data (PCM data) received by the system.

[0085] The second method is as follows: the audio is sampled according to the preset sampling precision to obtain each audio sampling point; then, the gain value input by the user is obtained, and the gain value is simulated by a non-linear tan function to obtain the scaling factor; finally, the amplitude of each audio sampling point is scaled according to the scaling factor.

[0086] For 16-bit mono audio, i.e., when the sampling precision is 16-bit, the amplitude range of the sampling points is [-2]. 15 ,2 15 The scaling factor [-1], or [-32768, 32767], multiplies the amplitude of each sampling point by the scaling factor to amplify or reduce the volume (a scaling factor greater than 1 amplifies the volume, while a scaling factor less than 1 reduces the volume). The scaling factor should be chosen according to the logarithmic principle, meaning that human auditory response is based on relative rather than absolute changes in sound, and the logarithmic scale perfectly mimics the human ear's response to sound. Once a gain value, such as `level`, is determined, the actual scaling factor needs to be simulated using the tan function, i.e., `float multiplier = tan(level / 100.0)`.

[0087] Furthermore, overflow processing is needed for data exceeding the amplitude range (i.e., [-32768, 32767]), and the excess amplitude is clipped (i.e., the portion exceeding the range is directly discarded) to ensure the values ​​are within the correct range. However, if too much overflow data occurs in an audio frame, it will cause audio distortion. Therefore, it is necessary to dynamically select the scaling factor appropriately. The minimum and maximum values ​​within 70ms of sampled data can be pre-calculated, and then the minimum and maximum gain values ​​corresponding to the distortion can be calculated respectively, and then displayed on the interface for the user to adjust.

[0088] Furthermore, the amplification or reduction of the audio sampling point amplitude is displayed in real time on the volume dashboard. Specifically, when the user adjusts the gain value, the volume detection dashboard can also display the amplification or reduction effect in real time. The user can complete the volume adjustment operation by adjusting the gain value to bring the volume detection value into a moderate range.

[0089] This embodiment provides an integrated solution for troubleshooting and debugging issues in the voice-activated karaoke environment. It can help and guide users to independently troubleshoot complex and diverse problems, reduce user costs, simplify troubleshooting, and thus improve user experience.

[0090] Example 2

[0091] This embodiment is a computer-readable storage medium corresponding to the above embodiments, on which a computer program is stored. When the program is executed by a processor, it implements the various steps of the speech environment detection method in the above embodiments and can achieve the same technical effect, which will not be repeated here.

[0092] In summary, the speech environment detection method and computer-readable storage medium provided by this invention can offer users a convenient and efficient speech environment detection method and visual adjustment means, reducing the user's usage threshold and thus improving the user experience; at the same time, it can reduce the cost of construction and troubleshooting for technicians.

[0093] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent modifications made based on the content of the present invention specification and drawings, or direct or indirect applications in related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A voice environment detection method, characterized by, The method comprises the following steps: receiving test audio; detecting the volume of the test audio to determine the maximum decibel value of the test audio; performing speech recognition on the test audio to obtain a speech recognition result, and performing speech recognition detection according to the speech recognition result and the content of the test audio; if a preset condition is met, determining that the speech environment detection is passed, wherein the preset condition comprises that the speech recognition detection is passed and the maximum decibel value of the test audio is within a preset volume range; after the volume of the test audio is detected to determine the maximum decibel value of the test audio, the method further comprises the following steps: if the maximum decibel value of the test audio is not within the preset volume range, prompting that the volume is too high or too low, and adjusting the audio signal level or adjusting the volume according to a preset gain value.

2. The voice environment detection method of claim 1, wherein, The method for detecting the volume of the test audio to determine the maximum decibel value of the test audio comprises the following steps: obtaining the test audio in the current period in real time according to a preset period to obtain current period audio; calculating the decibel value of the current period audio according to the dynamic range of the current period audio and a preset sampling accuracy; if the decibel value of the current period audio is greater than the maximum decibel value, updating the maximum decibel value according to the decibel value of the current period audio, and the initial value of the maximum decibel value is the decibel value of the first period audio.

3. The voice environment detection method of claim 2, wherein, The method for detecting the volume of the test audio to determine the maximum decibel value of the test audio further comprises the following steps: displaying the decibel value of the current period audio in real time through a volume dashboard; if the maximum decibel value is updated, displaying the updated maximum decibel value in the volume dashboard for a preset display time.

4. The voice environment detection method of claim 2, wherein, The method for calculating the decibel value of the current period audio according to the dynamic range of the current period audio and a preset sampling accuracy comprises the following steps: calculating the decibel value of the current period audio according to a decibel calculation formula, wherein the decibel calculation formula is I = 10 x lg (D / 2 x ) 2 , wherein I is the decibel value of the current period audio, D is the dynamic range of the current period audio, and x is a preset sampling accuracy.

5. The voice environment detection method of claim 1, wherein, The method for adjusting the volume according to a preset gain value comprises the following steps: sampling the test audio according to a preset sampling accuracy to obtain audio sampling points; obtaining a gain value and simulating the gain value through a tan function to obtain a scaling factor; scaling the amplitude of the audio sampling points according to the scaling factor.

6. The voice environment detection method of claim 1, wherein, The content of the test audio comprises a preset voice wake-up word and a voice instruction. The method for performing speech recognition on the test audio to obtain a speech recognition result, and performing speech recognition detection according to the speech recognition result and the content of the test audio comprises the following steps: performing speech recognition on the test audio, and if the preset voice wake-up word is recognized, performing voice command detection to obtain a voice instruction recognition result; if the voice instruction recognition result matches the preset voice instruction, determining that the speech recognition detection is passed.

7. The voice environment detection method of claim 6, wherein, The method for performing speech recognition on the test audio to obtain a speech recognition result, and performing speech recognition detection according to the speech recognition result and the content of the test audio further comprises the following steps: If no voice instruction recognition result matching the preset voice instruction is detected within a preset time after the preset voice wake-up word is recognized, it is determined that the voice recognition detection is not passed.

8. The voice environment detection method of claim 1, wherein, The preset condition further includes that a running environment detection is passed, and the running environment detection includes network fluctuation detection, network timeout detection, authentication detection, initialization detection, and runtime problem detection.

9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program, when executed by a processor, implements the method of any one of claims 1-8.

Citation Information

Patent Citations

  • Method and system for test of intelligent sound wakeup word identification rate

    CN108511000A

  • Speech interaction scene recognition method, device and intelligent sound box

    CN109218899A