Voice interaction method and system of intelligent electronic equipment

By employing adaptive noise reduction processing and a target wake word mechanism, user identity verification is achieved, ensuring that only authorized personnel can execute commands. This solves the problem of unauthorized users being unable to interact in existing technologies, and improves the security and flexibility of voice interaction.

CN120853601AActive Publication Date: 2025-10-28SHENZHEN ZECHIN ELECTRONICS
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511177854.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-10-28
Estimated Expiration
2045-08-21

AI Technical Summary

Technical Problem

In the voice interaction process using existing technologies, unauthorized users cannot be effectively verified, resulting in insufficient security and flexibility of the voice interaction system.

Method used

Voiceprint features are extracted through adaptive noise reduction processing to determine the user's identity. If the identity verification fails, a target wake word is prompted. The user can then enter the wake word within a preset time to execute the command.

Benefits of technology

It improves the security and flexibility of voice interaction, ensures that instructions are only executed for authorized personnel, solves the problem of identity verification failure caused by voiceprint recognition errors, and enhances the fault tolerance and effectiveness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853601A_ABST
    Figure CN120853601A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice interaction, and provides a voice interaction method and system of intelligent electronic equipment. The method comprises the following steps: receiving an interactive voice input by a user, and carrying out adaptive noise reduction processing on the interactive voice to obtain a noise-reduced interactive voice; extracting voiceprint characteristics of the noise reduction interaction voice, and judging whether the user is an authorized person of the intelligent electronic equipment or not based on the voiceprint characteristics; if not, prompting the user to input a target wake-up word; wherein the target wake-up word is a newly generated wake-up word sent by the authorized officer to the user; judging whether the user inputs the target wake-up word in a preset time period or not; and if yes, executing an instruction corresponding to the noise reduction interaction voice. According to the method, the situation that qualified but unauthorized users cannot perform voice interaction is avoided, and the flexibility and fault tolerance of the system are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice interaction technology, and in particular to a voice interaction method and system for an intelligent electronic device. Background Technology

[0002] With the rapid development of speech recognition, voice interaction has become one of the important human-computer interaction methods in smart electronic devices. Users can control devices and query information through voice, improving the convenience and intelligence of device use. To enhance the security of voice interaction, more and more voice interaction systems are introducing voiceprint recognition mechanisms to achieve personalized services and access control. However, in the process of voiceprint recognition-based voice interaction, existing technologies usually directly refuse to execute voice commands after the user's identity verification fails, lacking a flexible supplementary verification mechanism. This method can easily prevent qualified but unauthorized users from achieving voice interaction with smart electronic devices. Summary of the Invention

[0003] This application provides a voice interaction method and system for intelligent electronic devices to solve the problems mentioned in the background.

[0004] In a first aspect, this application provides a voice interaction method for a smart electronic device, comprising: Receive interactive voice input from the user and perform adaptive noise reduction processing on the interactive voice to obtain noise-reduced interactive voice; Extract the voiceprint features of the noise-reduced interactive voice, and determine whether the user is an authorized person of the smart electronic device based on the voiceprint features; If the user is not an authorized person, the system prompts the user to enter a target wake-up word; wherein, the target wake-up word is the latest generated wake-up word sent by the authorized person to the user. Determine whether the user has entered the target wake word within a preset time period; If so, execute the command corresponding to the noise-reduced interactive voice.

[0005] In one possible implementation, after determining whether the user is an authorized user of the smart electronic device based on the voiceprint features, the method further includes: If the user is an authorized person, they will execute the commands corresponding to the noise-reduced interactive voice.

[0006] In one possible implementation, the step of denoising the interactive voice to obtain denoised interactive voice includes: The interactive voice is parsed to obtain the background sound information carried by the interactive voice, and the background sound information is input into a background classification model for classification to obtain classification labels; wherein, the background classification model is a pre-trained neural network model; Check if a target noise reduction model corresponding to the classification label exists in the database; If it exists, the interactive speech is denoised based on the target denoising model to obtain denoised interactive speech; If it does not exist, the interactive voice is denoised based on the general denoising model in the database to obtain denoised interactive voice.

[0007] In one possible implementation, extracting the voiceprint features of the noise-reduced interactive speech includes: The noise-reduced interactive speech is sequentially pre-emphasized, framed, and windowed to obtain multiple short-time speech frames of fixed length; wherein, there is partial overlap between two adjacent short-time speech frames. For each of the short-time speech frames, the feature sequence corresponding to the short-time speech frame is analyzed; The aforementioned feature sequences are fused to obtain a fused feature sequence; the fused feature sequence is the voiceprint feature.

[0008] In one possible implementation, analyzing the feature sequence corresponding to the short-time speech frame includes: Power spectrum analysis is performed on the short-time speech frame to obtain the power spectrum corresponding to the short-time speech frame; The power spectrum is input into a preset Mel filter bank to obtain the Mel spectrum; Perform a logarithmic operation on the Mel spectrum to obtain the logarithmic Mel spectrum; Perform a discrete cosine transform on the logarithmic Mel frequency spectrum to obtain an initial Mel frequency cepstral coefficient sequence, and extract the first preset number of Mel frequency cepstral coefficients from the initial Mel frequency cepstral coefficient sequence to obtain the target Mel frequency cepstral coefficient sequence. The target Mel frequency cepstral coefficient sequence is normalized to obtain the normalized target Mel frequency cepstral coefficient sequence. The normalized target Mel frequency cepstral coefficient sequence is subjected to a first-order difference operation to obtain a first-order difference coefficient sequence, and the first-order difference coefficient sequence is normalized to obtain a normalized first-order difference coefficient sequence. The normalized first-order difference coefficient sequence is subjected to a second-order difference operation to obtain a second-order difference coefficient sequence, and the second-order difference coefficient sequence is then normalized to obtain a normalized second-order difference coefficient sequence. The normalized target Mel frequency cepstral coefficient sequence, the first-order difference coefficient sequence, and the second-order difference coefficient sequence are arranged in sequence to obtain the feature sequence.

[0009] In one possible implementation, the voiceprint feature is a sequence of numbers consisting of multiple digits, where the absolute value of each digit is no greater than 1. The step of determining whether the user is an authorized user of the smart electronic device based on the voiceprint feature includes: Acquire multiple standard voiceprint features of intelligent electronic devices; the standard voiceprint features are a sequence of numbers consisting of multiple digits, and the absolute value of each digit is no greater than 1; For each of the aforementioned standard voiceprint features, the similarity between the standard voiceprint feature and the voiceprint feature is calculated; The standard voiceprint feature corresponding to the maximum similarity is determined as the target standard voiceprint feature; Calculate the feature difference sequence between the target standard voiceprint feature and the voiceprint feature; the feature difference sequence consists of multiple differences; Each absolute value corresponding to a difference in the feature difference sequence is compared with a preset absolute value. If any of the absolute values ​​is greater than the preset absolute value, it is determined that the user is not an authorized person of the smart electronic device; If none of the absolute values ​​are greater than the preset absolute value, the user is determined to be an authorized person of the smart electronic device.

[0010] In one possible implementation, the authorized personnel include multiple individuals, and the method further includes: before prompting the user to input a target wake word. The system prompts the user to input the first terminal identifier of the first mobile terminal they are currently carrying, and obtains the user's image information based on the first terminal identifier, and generates a target wake-up word based on the image information. The first terminal identifier, the image information, the target wake-up word, and the preset voice command are sent to the second mobile terminals of each authorized person, so that when the second mobile terminal of the authorized person receives the preset voice command, it prompts the authorized person to determine whether to send the target wake-up word to the first mobile terminal based on the image information. When the authorized person determines to send the target wake-up word to the first mobile terminal, the authorized person sends the target wake-up word to the first mobile terminal based on the first terminal identifier.

[0011] In one possible implementation, generating the target wake word based on the image information includes: The image information is processed to obtain grayscale image information; Extract the grayscale value of each pixel from the grayscale image information; Determine the average gray value, maximum gray value, and minimum gray value corresponding to each of the aforementioned gray values; Obtain a preset grayscale value-string matching table, wherein the grayscale value-string matching table includes a grayscale value column and a character column; Extract the target string corresponding to the maximum gray value from the gray value-string matching table to obtain a string empty space, and shift the strings corresponding to each gray value between the minimum gray value and the maximum gray value down by one string empty space to obtain a target string empty space, and insert the target string into the target string empty space to obtain the target gray value-string matching table; In the target grayscale value-string matching table, determine the first string corresponding to the average grayscale value, the second string corresponding to the maximum grayscale value, and the third string corresponding to the minimum grayscale value; The first string, the second string, and the third string are arranged in sequence to obtain the target wake-up word.

[0012] Secondly, this application provides a voice interaction system for an intelligent electronic device, comprising: The receiving module is used to receive interactive voice input by the user and perform adaptive noise reduction processing on the interactive voice to obtain noise-reduced interactive voice. The first judgment module is used to extract the voiceprint features of the noise-reduced interactive voice and determine whether the user is an authorized person of the smart electronic device based on the voiceprint features. The prompting module is used to prompt the user to enter a target wake-up word if the user is not an authorized person; wherein, the target wake-up word is the latest generated wake-up word sent by the authorized person to the user; The second judgment module is used to determine whether the user has entered the target wake-up word within a preset time period; The execution module is used to execute the instruction corresponding to the noise-reduced interactive voice when the user inputs the target wake-up word within a preset time period.

[0013] This application provides a voice interaction method and system for intelligent electronic devices. The method includes: receiving interactive voice input from a user, and performing adaptive noise reduction processing on the interactive voice to obtain noise-reduced interactive voice; extracting the voiceprint features of the noise-reduced interactive voice, and determining whether the user is an authorized personnel of the intelligent electronic device based on the voiceprint features; if not an authorized personnel, prompting the user to input a target wake-up word; wherein the target wake-up word is a newly generated wake-up word sent by the authorized personnel to the user; determining whether the user has input the target wake-up word within a preset time period; if so, executing the instruction corresponding to the noise-reduced interactive voice. This method, on the one hand, by introducing voiceprint feature recognition technology during voice interaction, verifies the user's identity, ensuring that only authorized personnel are directly executed with corresponding instructions, thereby improving the security of voice interaction; on the other hand, by setting a target wake-up word as a supplementary verification method, it solves the problem of identity verification failure due to voiceprint recognition errors or other factors. Users can complete identity verification by inputting the target wake-up word, avoiding the situation where qualified but unauthorized users cannot perform voice interaction, enhancing the system's flexibility and fault tolerance; and furthermore, by limiting the user to input the target wake-up word within a preset time period, it improves the effectiveness of the verification process. Attached Figure Description

[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0015] Figure 1 A flowchart illustrating the voice interaction method for an intelligent electronic device provided in this application embodiment; Figure 2 A schematic block diagram of the structure of a voice interaction system for an intelligent electronic device provided in the embodiments of this application; Figure 3 A schematic block diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it require execution in the described order. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0018] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0019] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the relevant listed items and all possible combinations, and includes such combinations.

[0020] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0021] Please see Figure 1 , Figure 1 This is a flowchart illustrating the voice interaction method for an intelligent electronic device provided in an embodiment of this application, as shown below. Figure 1 As shown, the voice interaction method for intelligent electronic devices provided in this application includes steps S1 to S5.

[0022] Step S1: Receive the interactive voice input by the user and perform adaptive noise reduction processing on the interactive voice to obtain noise-reduced interactive voice.

[0023] Step S2: Extract the voiceprint features of the noise-reduced interactive voice, and determine whether the user is an authorized person of the smart electronic device based on the voiceprint features.

[0024] Step S3: If the user is not an authorized person, prompt the user to enter a target wake-up word; wherein, the target wake-up word is the latest generated wake-up word sent by the authorized person to the user.

[0025] Step S4: Determine whether the user has entered the target wake word within a preset time period.

[0026] Step S5: If yes, execute the command corresponding to the noise reduction interactive voice.

[0027] It should be noted that the smart electronic devices in this embodiment include, but are not limited to, smartphones, smart speakers, smart computers, smart air conditioners, and smart in-vehicle devices.

[0028] This embodiment specifically includes: as described in step S1 above, receiving interactive voice input from the user and performing adaptive noise reduction processing on the interactive voice to obtain noise-reduced interactive voice. Specifically, when the user's interactive voice is received, the interactive voice is parsed to obtain the background sound information carried by the interactive voice, and the background sound information is input into a background classification model for classification to obtain a classification label; it is checked whether a target noise reduction model corresponding to the classification label exists in the database; if it exists, the interactive voice is noise-reduced based on the target noise reduction model to obtain noise-reduced interactive voice; if it does not exist, the interactive voice is noise-reduced based on a general noise reduction model in the database to obtain noise-reduced interactive voice.

[0029] As described in step S2 above, the voiceprint features of the noise-reduced interactive speech are extracted, and the user is determined to be an authorized person of the smart electronic device based on the voiceprint features. Specifically, the noise-reduced interactive speech is sequentially pre-emphasized, framed, and windowed to obtain multiple short-time speech frames of fixed length; wherein, there is partial overlap between two adjacent short-time speech frames; for each short-time speech frame, the feature sequence corresponding to the short-time speech frame is analyzed; the feature sequences are fused to obtain a fused feature sequence; the fused feature sequence is the voiceprint feature.

[0030] As described in step S3 above, if the user is not an authorized person, the user is prompted to input a target wake-up word. The target wake-up word is the latest generated wake-up word sent by the authorized person to the user. Specifically, when it is detected that the user is not an authorized person, the user is prompted to input the first terminal identifier of their currently carried first mobile terminal. Based on the first terminal identifier, the user's image information is obtained, and a target wake-up word is generated based on the image information. The first terminal identifier, the image information, the target wake-up word, and a preset voice command are sent to the second mobile terminals of each authorized person. This allows the second mobile terminal of each authorized person to remind the authorized person, upon receiving the preset voice command, to determine whether to send the target wake-up word to the first mobile terminal based on the image information. When the authorized person determines to send the target wake-up word to the first mobile terminal, the authorized person sends the target wake-up word to the first mobile terminal based on the first terminal identifier. When the first mobile terminal receives the target wake-up word, it reminds the user to input the target wake-up word. Based on the reminder from the first mobile terminal, the user inputs the target wake-up word to the smart electronic device via voice interaction.

[0031] As described in step S4 above, determine whether the user has entered the target wake-up word within a preset time period.

[0032] As described in step S5 above, if yes, the instruction corresponding to the noise-reducing interactive voice is executed. Specifically, if the user inputs the target wake-up word within a preset time period, the noise-reducing interactive voice is semantically recognized based on a preset semantic recognition model to obtain the text information corresponding to the noise-reducing interactive voice, and the instruction corresponding to the text information is executed.

[0033] The method provided in this embodiment, on the one hand, verifies the user's identity by introducing voiceprint feature recognition technology during voice interaction, ensuring that only authorized personnel can directly execute corresponding commands, thereby improving the security of voice interaction. On the other hand, by setting a target wake-up word as a supplementary verification method, it solves the problem of identity verification failure due to voiceprint recognition errors or other factors. Users can complete identity verification by inputting the target wake-up word, avoiding the situation where qualified but unauthorized users cannot perform voice interaction, thus enhancing the system's flexibility and fault tolerance. Furthermore, by limiting the user to input the target wake-up word within a preset time period, the effectiveness of the verification process is improved.

[0034] In some embodiments, after determining whether the user is an authorized user of the smart electronic device based on the voiceprint features, the method further includes the following steps: If the user is an authorized person, they will execute the commands corresponding to the noise-reduced interactive voice.

[0035] In some embodiments, the noise reduction process for the interactive voice to obtain noise-reduced interactive voice includes the following steps: The interactive voice is parsed to obtain the background sound information carried by the interactive voice, and the background sound information is input into a background classification model for classification to obtain classification labels; wherein, the background classification model is a pre-trained neural network model; The system checks whether a target denoising model corresponding to the classification label exists in the database; wherein, the database contains multiple correspondences between classification labels and their corresponding denoising models. If it exists, the interactive speech is denoised based on the target denoising model to obtain denoised interactive speech; If it does not exist, the interactive voice is denoised based on the general denoising model in the database to obtain denoised interactive voice.

[0036] In this embodiment, after receiving interactive voice, the intelligent electronic device parses the interactive voice, extracts the background sound information in the interactive voice, and passes the extracted background sound information to the background classification model. The background classification model outputs a label corresponding to the background sound information based on the spectral characteristics in the audio, such as "traffic noise", "environmental noise", or "mechanical noise". According to the generated classification label, the intelligent electronic device searches the database to see if there is a target noise reduction model that matches the classification label. For example, if the classification label is "traffic noise", the database will be queried to see if there is a noise reduction model that matches traffic noise. If a matching noise reduction model is found, the model is loaded for subsequent processing. If no matching model is found, a general noise reduction model is used.

[0037] The method provided in this embodiment optimizes the adaptability of noise reduction processing by extracting background sound information from interactive speech and inputting it into a background classification model. The classification labels of the background classification model provide precise guidance for different noise sources, enabling the noise reduction process to remove irrelevant noise more efficiently. On the other hand, by finding the corresponding target noise reduction model based on the classification label, the targeting of noise reduction processing is improved, allowing each type of noise to be processed using the most suitable noise reduction algorithm, avoiding the insufficient noise reduction effect that may result from general processing methods. Furthermore, by switching to a general noise reduction model when no target noise reduction model is found, the noise reduction needs when there is a lack of specialized models are addressed, ensuring the clarity of the speech signal, reducing the impact of noise on speech recognition, and guaranteeing the effectiveness of interactive speech.

[0038] In some embodiments, extracting the voiceprint features of the noise-reduced interactive speech includes the following steps: The noise-reduced interactive speech is sequentially pre-emphasized, framed, and windowed to obtain multiple short-time speech frames of fixed length; wherein, there is partial overlap between two adjacent short-time speech frames. For each of the short-time speech frames, the feature sequence corresponding to the short-time speech frame is analyzed; The aforementioned feature sequences are fused to obtain a fused feature sequence; the fused feature sequence is the voiceprint feature.

[0039] In this embodiment, a pre-emphasis operation is performed on the noise-reduced interactive speech to enhance high-frequency components and reduce interference caused by low-frequency components in the speech signal. The pre-emphasis operation performs linear filtering on the original speech signal to generate a speech signal with enhanced frequency characteristics. The pre-emphasis processed speech signal is further divided into multiple fixed-length short-time speech frames by a frame segmentation module. Each short-time speech frame represents a signal segment within a certain time range. The length of the short-time speech frame is based on a preset frame length and frame overlap rate to ensure partial overlap between adjacent frames, thereby preserving the speech continuity feature. For each short-time speech frame, a Hamming window is used to weight the short-time speech frame to reduce the spectral effect at the signal edges. The feature sequence of the short-time speech frame after windowing is analyzed, and the average value of each value corresponding to the same sequence number of each feature sequence is calculated sequentially to obtain the voiceprint feature.

[0040] The method provided in this embodiment, on the one hand, enhances high-frequency components and reduces low-frequency interference by performing pre-emphasis processing on the denoised interactive speech, making subsequent speech feature extraction clearer and improving the recognizability of the speech signal. On the other hand, by performing frame segmentation and windowing processing on the denoised interactive speech to form multiple short-time speech frames and weighting them with Hamming windows, the signal edge effect is reduced and the spectral characteristics of each frame signal are improved, making the feature extraction of each frame more accurate. Furthermore, by fusing the feature sequences of each frame and calculating and obtaining the fused feature sequence, the generated voiceprint features provide more stable and representative voice identity features.

[0041] In some embodiments, analyzing the feature sequence corresponding to the short-time speech frame includes the following steps: Power spectrum analysis is performed on the short-time speech frame to obtain the power spectrum corresponding to the short-time speech frame; specifically, a short-time Fourier transform is performed on the short-time speech frame to obtain the amplitude spectrum, and the amplitude spectrum is converted into a power spectrum. The power spectrum is input into a preset Mel filter bank to obtain the Mel spectrum; Perform a logarithmic operation on the Mel spectrum to obtain the logarithmic Mel spectrum; Perform a discrete cosine transform on the logarithmic Mel frequency spectrum to obtain an initial Mel frequency cepstral coefficient sequence, and extract the first preset number of Mel frequency cepstral coefficients from the initial Mel frequency cepstral coefficient sequence to obtain the target Mel frequency cepstral coefficient sequence. The target Mel frequency cepstral coefficient sequence is normalized to obtain the normalized target Mel frequency cepstral coefficient sequence. The normalized target Mel-frequency cepstral coefficient sequence is subjected to a first-order difference operation to obtain a first-order difference coefficient sequence, and the first-order difference coefficient sequence is then normalized to obtain a normalized first-order difference coefficient sequence. For example, if the normalized target Mel-frequency cepstral coefficient sequence is 0.3, 0.2, -0.1, 0.5, -0.3, 0.2, then the first-order difference coefficient sequence is -0.4, 0.3, -0.2, -0.3. The normalized first-order difference coefficient sequence is subjected to a second-order difference operation to obtain a second-order difference coefficient sequence, and then the second-order difference coefficient sequence is normalized to obtain a normalized second-order difference coefficient sequence; for example, if the normalized first-order difference coefficient sequence is -0.4, 0.3, -0.2, -0.3, then the second-order difference coefficient sequence is 0.2, -0.6; The normalized target Mel frequency cepstral coefficient sequence, the first-order difference coefficient sequence, and the second-order difference coefficient sequence are arranged sequentially to obtain the feature sequence. For example, if the normalized target Mel frequency cepstral coefficient sequence is 0.3, 0.2, -0.1, 0.5, -0.3, 0.2, the normalized first-order difference coefficient sequence is -0.4, 0.3, -0.2, -0.3, and the normalized second-order difference coefficient sequence is 0.2, -0.6, then the feature sequence is 0.3, 0.2, -0.1, 0.5, -0.3, 0.2, -0.4, 0.3, -0.2, -0.3, 0.2, -0.6.

[0042] The method provided in this embodiment, on the one hand, extracts the frequency domain features of short-time speech frames by performing power spectrum analysis, enhancing the ability to capture the details of the speech signal in short-time speech frames and effectively reflecting the changes of the speech signal in the time and frequency domain. On the other hand, by inputting the power spectrum into a Mel filter bank and performing logarithmic operations to convert it into Mel spectrum and log-Mel spectrum, it simulates the perceptual characteristics of the human ear for different frequencies, improving the perceptual accuracy of speech features. Furthermore, by performing discrete cosine transform to extract Mel frequency cepstral coefficients and performing first-order and second-order difference processing, it captures the dynamic change features of the speech signal, making the extracted voiceprint features more stable.

[0043] In some embodiments, the voiceprint feature is a sequence of numbers consisting of multiple digits, where the absolute value of each digit is no greater than 1. The step of determining whether the user is an authorized user of the smart electronic device based on the voiceprint feature includes the following steps: Acquire multiple standard voiceprint features of intelligent electronic devices; the standard voiceprint features are a sequence of numbers consisting of multiple digits, and the absolute value of each digit is no greater than 1; For each of the standard voiceprint features, the similarity between the standard voiceprint feature and the voiceprint feature is calculated; specifically, the cosine value between the vector corresponding to the standard voiceprint feature and the vector corresponding to the voiceprint feature is determined as the similarity. The standard voiceprint feature corresponding to the maximum similarity is determined as the target standard voiceprint feature; Calculate the feature difference sequence between the target standard voiceprint feature and the voiceprint feature; the feature difference sequence consists of multiple differences; for example, if the target standard voiceprint feature is 0.3, 0.2, -0.1, 0.5, -0.3, 0.2, -0.4, 0.3, -0.2, -0.3, 0.2, -0.6, and the voiceprint feature is 0.3, 0.3, -0.1, 0.4, -0.3, 0.2, -0.3, 0.3, -0.1, -0.3, 0.2, -0.6, then the feature difference sequence is 0, 0.1, 0, -0.1, 0, 0, 0.1, 0, 0, 0.1, 0, 0, 0. Each absolute value corresponding to a difference in the feature difference sequence is compared with a preset absolute value. If any of the absolute values ​​is greater than the preset absolute value, it is determined that the user is not an authorized person of the smart electronic device; If none of the absolute values ​​are greater than the preset absolute value, the user is determined to be an authorized person of the smart electronic device.

[0044] The method provided in this embodiment, on the one hand, refines the comparison process of voiceprint features by determining the standard voiceprint feature corresponding to the maximum similarity and calculating the difference sequence between it and the voiceprint feature, thereby enhancing the sensitivity to subtle differences and improving the accuracy of recognition. On the other hand, by comparing the absolute value corresponding to each difference in the feature difference sequence with a preset absolute value, the user's identity is reliably verified.

[0045] In some embodiments, the authorized personnel include multiple individuals, and the method further includes the following steps before prompting the user to input a target wake word: The system prompts the user to input the first terminal identifier of the first mobile terminal they are currently carrying, and obtains the user's image information based on the first terminal identifier, and generates a target wake-up word based on the image information. The first terminal identifier, the image information, the target wake-up word, and the preset voice command are sent to the second mobile terminals of each authorized person, so that when the second mobile terminal of the authorized person receives the preset voice command, it prompts the authorized person to determine whether to send the target wake-up word to the first mobile terminal based on the image information. When the authorized person determines to send the target wake-up word to the first mobile terminal, the authorized person sends the target wake-up word to the first mobile terminal based on the first terminal identifier.

[0046] The method provided in this embodiment enhances the timeliness of wake word generation, which helps improve the reliability of user authentication. On the other hand, by sending the first terminal identifier, the image information, the target wake word, and the preset voice command to the second mobile terminal of each authorized person, it helps prevent the problem of not being able to send the target wake word to the user in a timely manner when the above information is sent to a single authorized person. Furthermore, by having the authorized person determine whether to send the target wake word based on the image information, the security of the wake word sending process is improved.

[0047] In some embodiments, generating a target wake-up word based on the image information includes the following steps: The image information is processed to obtain grayscale image information; Extract the grayscale value of each pixel from the grayscale image information; Determine the average gray value, maximum gray value, and minimum gray value corresponding to each of the aforementioned gray values; Obtain a preset grayscale value-string matching table, wherein the grayscale value-string matching table includes a grayscale value column and a character column; Extract the target string corresponding to the maximum gray value from the gray value-string matching table to obtain a string empty space. Then, shift the strings corresponding to each gray value between the minimum gray value and the maximum gray value down by one string empty space to obtain a target string empty space. Finally, insert the target string into the target string empty space to obtain the target gray value-string matching table. Herein, each gray value between the minimum gray value and the maximum gray value includes the minimum gray value but does not include the maximum gray value. In the target grayscale value-string matching table, determine the first string corresponding to the average grayscale value, the second string corresponding to the maximum grayscale value, and the third string corresponding to the minimum grayscale value; The first string, the second string, and the third string are arranged in sequence to obtain the target wake-up word.

[0048] The method provided in this embodiment, on the one hand, optimizes the generation method of wake words by utilizing the matching relationship between grayscale values ​​and strings to flexibly adjust the position in the mapping process between image information and characters, thereby improving the personalization and security of wake words. On the other hand, by extracting the target string corresponding to the maximum grayscale value from the grayscale value-string matching table to obtain a string empty space, and shifting the strings corresponding to each grayscale value between the minimum and maximum grayscale values ​​down by one string empty space to obtain a target string empty space, and inserting the target string into the target string empty space to obtain a target grayscale value-string matching table, the personalization and security of wake words are further improved.

[0049] Please see Figure 2 , Figure 2 A schematic block diagram of the structure of the voice interaction system 100 of the intelligent electronic device provided in the embodiments of this application, as shown below. Figure 2 As shown in the embodiment of this application, the voice interaction system 100 for an intelligent electronic device includes: The receiving module 110 is used to receive interactive voice input by the user and perform adaptive noise reduction processing on the interactive voice to obtain noise-reduced interactive voice.

[0050] The first judgment module 120 is used to extract the voiceprint features of the noise-reduced interactive voice and determine whether the user is an authorized person of the smart electronic device based on the voiceprint features.

[0051] The prompting module 130 is used to prompt the user to input a target wake-up word if the user is not an authorized person; wherein the target wake-up word is the latest generated wake-up word sent by the authorized person to the user.

[0052] The second judgment module 140 is used to determine whether the user has entered the target wake-up word within a preset time period.

[0053] The execution module 150 is used to execute the instruction corresponding to the noise-reduced interactive voice when the user inputs the target wake-up word within a preset time period.

[0054] The determination module 160 is used to determine the target second PTZ camera based on the first global monitoring area and the target monitoring area when the target monitoring area is not included in the first global monitoring area, and to start the target second PTZ camera within the preset time period.

[0055] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system and its modules described above can be referred to the process in the aforementioned embodiment of the voice interaction method for intelligent electronic devices, and will not be repeated here.

[0056] The voice interaction system 100 of the intelligent electronic device provided in the above embodiments can be implemented in the form of a computer program, which can be used in, for example... Figure 3 The terminal device 200 shown is running on it.

[0057] Please see Figure 3 , Figure 3 The following is a schematic block diagram of the structure of a terminal device 200 provided in an embodiment of this application. The terminal device 200 includes a processor 201 and a memory 202, which are connected through a system bus 203. The memory 202 may include a non-volatile storage medium and internal memory.

[0058] The non-volatile storage medium can store a computer program. The computer program includes program instructions that, when executed by the processor 201, cause the processor 201 to perform the voice interaction method of any of the aforementioned intelligent electronic devices.

[0059] The processor 201 provides computing and control capabilities to support the operation of the entire terminal device 200.

[0060] The internal memory provides an environment for the execution of computer programs in non-volatile storage media. When the computer program is executed by the processor 201, the processor 201 can execute the voice interaction method of any of the above-mentioned intelligent electronic devices.

[0061] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the terminal device 200 involved in the present application. The specific terminal device 200 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0062] It should be understood that processor 201 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, the general-purpose processor can be a microprocessor or any conventional processor.

[0063] In some embodiments, the processor 201 is configured to run a computer program stored in memory to perform the following steps: Receive interactive voice input from the user and perform adaptive noise reduction processing on the interactive voice to obtain noise-reduced interactive voice; Extract the voiceprint features of the noise-reduced interactive voice, and determine whether the user is an authorized person of the smart electronic device based on the voiceprint features; If the user is not an authorized person, the system prompts the user to enter a target wake-up word; wherein, the target wake-up word is the latest generated wake-up word sent by the authorized person to the user. Determine whether the user has entered the target wake word within a preset time period; If so, execute the command corresponding to the noise-reduced interactive voice.

[0064] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the terminal device 200 described above can be referred to the corresponding process of the voice interaction method of the aforementioned intelligent electronic device, and will not be repeated here.

[0065] This application also provides a computer-readable storage medium storing a computer program that, when executed by one or more processors, causes the one or more processors to implement the voice interaction method of the intelligent electronic device provided in this application.

[0066] The computer-readable storage medium can be an internal storage unit of the terminal device 200 in the aforementioned embodiments, such as a hard disk or memory of the terminal device 200. The computer-readable storage medium can also be an external storage device of the terminal device 200, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped with the terminal device 200.

[0067] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A voice interaction method for an intelligent electronic device, characterized in that, include: Receive interactive voice input from the user and perform adaptive noise reduction processing on the interactive voice to obtain noise-reduced interactive voice; Extract the voiceprint features of the noise-reduced interactive voice, and determine whether the user is an authorized person of the smart electronic device based on the voiceprint features; If the user is not an authorized person, the system prompts the user to enter a target wake-up word; wherein, the target wake-up word is the latest generated wake-up word sent by the authorized person to the user. Determine whether the user has entered the target wake word within a preset time period; If so, execute the command corresponding to the noise-reduced interactive voice.

2. The voice interaction method for intelligent electronic devices according to claim 1, characterized in that, After determining whether the user is an authorized user of the smart electronic device based on the voiceprint features, the method further includes: If the user is an authorized person, they will execute the commands corresponding to the noise-reduced interactive voice.

3. The voice interaction method for intelligent electronic devices according to claim 1, characterized in that, The step of denoising the interactive voice to obtain denoised interactive voice includes: The interactive voice is parsed to obtain the background sound information carried by the interactive voice, and the background sound information is input into a background classification model for classification to obtain classification labels; wherein, the background classification model is a pre-trained neural network model; Check if a target noise reduction model corresponding to the classification label exists in the database; If it exists, the interactive speech is denoised based on the target denoising model to obtain denoised interactive speech; If it does not exist, the interactive voice is denoised based on the general denoising model in the database to obtain denoised interactive voice.

4. The voice interaction method for intelligent electronic devices according to claim 1, characterized in that, The extraction of voiceprint features from the noise-reduced interactive speech includes: The noise-reduced interactive speech is sequentially pre-emphasized, framed, and windowed to obtain multiple short-time speech frames of fixed length; wherein, there is partial overlap between two adjacent short-time speech frames. For each of the short-time speech frames, the feature sequence corresponding to the short-time speech frame is analyzed; The aforementioned feature sequences are fused to obtain a fused feature sequence; the fused feature sequence is the voiceprint feature.

5. The voice interaction method for intelligent electronic devices according to claim 4, characterized in that, The analysis of the feature sequence corresponding to the short-time speech frame includes: Power spectrum analysis is performed on the short-time speech frame to obtain the power spectrum corresponding to the short-time speech frame; The power spectrum is input into a preset Mel filter bank to obtain the Mel spectrum; Perform a logarithmic operation on the Mel spectrum to obtain the logarithmic Mel spectrum; Perform a discrete cosine transform on the logarithmic Mel frequency spectrum to obtain an initial Mel frequency cepstral coefficient sequence, and extract the first preset number of Mel frequency cepstral coefficients from the initial Mel frequency cepstral coefficient sequence to obtain the target Mel frequency cepstral coefficient sequence. The target Mel frequency cepstral coefficient sequence is normalized to obtain the normalized target Mel frequency cepstral coefficient sequence. The normalized target Mel frequency cepstral coefficient sequence is subjected to a first-order difference operation to obtain a first-order difference coefficient sequence, and the first-order difference coefficient sequence is normalized to obtain a normalized first-order difference coefficient sequence. The normalized first-order difference coefficient sequence is subjected to a second-order difference operation to obtain a second-order difference coefficient sequence, and the second-order difference coefficient sequence is then normalized to obtain a normalized second-order difference coefficient sequence. The normalized target Mel frequency cepstral coefficient sequence, the first-order difference coefficient sequence, and the second-order difference coefficient sequence are arranged in sequence to obtain the feature sequence.

6. The voice interaction method for intelligent electronic devices according to claim 4, characterized in that, The voiceprint feature is a sequence of numbers consisting of multiple digits, where the absolute value of each digit is no greater than 1. The step of determining whether the user is an authorized user of the smart electronic device based on the voiceprint feature includes: Acquire multiple standard voiceprint features of intelligent electronic devices; the standard voiceprint features are a sequence of numbers consisting of multiple digits, and the absolute value of each digit is no greater than 1; For each of the aforementioned standard voiceprint features, the similarity between the standard voiceprint feature and the voiceprint feature is calculated; The standard voiceprint feature corresponding to the maximum similarity is determined as the target standard voiceprint feature; Calculate the feature difference sequence between the target standard voiceprint feature and the voiceprint feature; the feature difference sequence consists of multiple differences; Each absolute value corresponding to a difference in the feature difference sequence is compared with a preset absolute value. If any of the absolute values ​​is greater than the preset absolute value, it is determined that the user is not an authorized person of the smart electronic device; If none of the absolute values ​​are greater than the preset absolute value, the user is determined to be an authorized person of the smart electronic device.

7. The voice interaction method for intelligent electronic devices according to claim 1, characterized in that, The authorized personnel include multiple individuals, and the method further includes, before prompting the user to input the target wake word: The system prompts the user to input the first terminal identifier of the first mobile terminal they are currently carrying, and obtains the user's image information based on the first terminal identifier, and generates a target wake-up word based on the image information. The first terminal identifier, the image information, the target wake-up word, and the preset voice command are sent to the second mobile terminals of each authorized person, so that when the second mobile terminal of the authorized person receives the preset voice command, it prompts the authorized person to determine whether to send the target wake-up word to the first mobile terminal based on the image information. When the authorized person determines to send the target wake-up word to the first mobile terminal, the authorized person sends the target wake-up word to the first mobile terminal based on the first terminal identifier.

8. The voice interaction method for intelligent electronic devices according to claim 7, characterized in that, The generation of the target wake-up word based on the image information includes: The image information is processed to obtain grayscale image information; Extract the grayscale value of each pixel from the grayscale image information; Determine the average gray value, maximum gray value, and minimum gray value corresponding to each of the aforementioned gray values; Obtain a preset grayscale value-string matching table, wherein the grayscale value-string matching table includes a grayscale value column and a character column; Extract the target string corresponding to the maximum gray value from the gray value-string matching table to obtain a string empty space, and shift the strings corresponding to each gray value between the minimum gray value and the maximum gray value down by one string empty space to obtain a target string empty space, and insert the target string into the target string empty space to obtain the target gray value-string matching table; In the target grayscale value-string matching table, determine the first string corresponding to the average grayscale value, the second string corresponding to the maximum grayscale value, and the third string corresponding to the minimum grayscale value; The first string, the second string, and the third string are arranged in sequence to obtain the target wake-up word.

9. A voice interaction system for an intelligent electronic device, characterized in that, include: The receiving module is used to receive interactive voice input by the user and perform adaptive noise reduction processing on the interactive voice to obtain noise-reduced interactive voice. The first judgment module is used to extract the voiceprint features of the noise-reduced interactive voice and determine whether the user is an authorized person of the smart electronic device based on the voiceprint features. The prompting module is used to prompt the user to enter a target wake-up word if the user is not an authorized person; wherein, the target wake-up word is the latest generated wake-up word sent by the authorized person to the user; The second judgment module is used to determine whether the user has entered the target wake-up word within a preset time period; The execution module is used to execute the instruction corresponding to the noise-reduced interactive voice when the user inputs the target wake-up word within a preset time period.

Citation Information

Patent Citations

  • Voice control method and device, and storage medium

    CN115312068A

  • Control method and control system of intelligent sound equipment

    CN120452435A

  • Systems and methods for authentication based on human teeth pattern

    IN201621010871A

  • Individual authentication system by voice

    JP2003302999A

  • Authentication with dynamic user identification

    US12361109B1