A voice information processing method and apparatus

By acquiring and analyzing the feature information of voice information, the problem of inaccurate identification of target users in existing technologies has been solved, achieving precise user permission matching and improved interaction capabilities.

CN115719592BActive Publication Date: 2026-05-26ZTE CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZTE CORP
Filing Date
2016-08-15
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing speech recognition technology cannot accurately identify target users, resulting in a poor user experience and an inability to accurately match user operation permissions.

Method used

By acquiring the first and second feature information of the voice information, analyzing and judging its relationship with the preset voice information, it is determined whether to perform the corresponding operation.

Benefits of technology

It enables accurate identification of target users and precise matching of user operation permissions, improving the interaction between users and devices and reducing the risk of erroneous operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115719592B_ABST
    Figure CN115719592B_ABST
Patent Text Reader

Abstract

This invention discloses a voice information processing method, the method comprising: acquiring first voice information; analyzing and processing the first voice information to obtain first feature information and second feature information of the first voice information; determining the relationship between the first voice information and preset voice information based on the first feature information and the second feature information of the first voice information, and determining whether to perform an operation corresponding to the first voice information according to the determination result. This invention also discloses a voice information processing device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to voice processing technology in the field of communications, and more particularly to a voice information processing method and apparatus. Background Technology

[0002] With the rapid development of internet technology, wireless networks have been widely used across various industries, and voice services have emerged accordingly. Based on the convenience of voice interaction, voice technology has been widely applied to people's work, entertainment, sports, and other lifestyles. For example, Google integrated a new voice dictation tool into its Google Docs application, enabling users to interact with computers without traditional keyboard input; Microsoft and Apple have also integrated their respective voice products, Cortana and Siri, from their handheld devices into their computer systems; even some smartphones or wearable devices can interact with terminal devices via voice technology. Existing voice recognition technology mainly involves converting the language information contained in the user's voice into text, either locally or in the cloud, and comparing it with the corresponding text in the sampled data. Simultaneously, it compares the frequency resonance of the two voice segments to distinguish between different users.

[0003] However, existing speech recognition technology only performs some experiential identification on certain "physical features" in the user's voice, without considering the changes in voice frequency caused by factors such as the user's speaking habits and mood. This results in a large error between the sampled user's voice frequency and the actual user's voice frequency, making it impossible to accurately identify the target user. This can lead to situations where some users' operations are outside their actual operating permissions, resulting in a poor user experience. Summary of the Invention

[0004] To address the aforementioned technical problems, embodiments of the present invention provide a voice information processing method and apparatus, which at least partially solves the problem in the prior art that the target user cannot be accurately identified based on voice information.

[0005] To achieve the above objectives, the technical solution of this invention is implemented as follows:

[0006] A voice information processing method, the method comprising:

[0007] Obtain the first voice information;

[0008] The first voice information is analyzed and processed to obtain the first feature information and the second feature information of the first voice information;

[0009] Based on the first feature information and the second feature information of the first voice information, the relationship between the first voice information and the preset voice information is determined, and the operation corresponding to the first voice information is determined according to the determination result.

[0010] Optionally, before analyzing the first speech information to obtain the first feature information and the second feature information of the first speech information, the method further includes:

[0011] Obtain the first time-domain waveform corresponding to the first voice information;

[0012] Determine whether the first time-domain waveform of the first voice information is continuous;

[0013] If the first time-domain waveform of the first speech information is continuous, then the analysis and processing of the first speech information is performed to obtain the first feature information and the second feature information of the first speech information.

[0014] If the first time-domain waveform of the first voice information is discontinuous, then the first voice information is reacquired.

[0015] Optionally, the step of analyzing and processing the first speech information to obtain the first feature information and the second feature information of the first speech information includes:

[0016] Perform spectral analysis on the first time-domain waveform of the first speech information to obtain the frequency-domain waveform of the first speech information;

[0017] Based on the frequency domain waveform of the first voice information, the first feature information of the first voice information is obtained;

[0018] The first time-domain waveform of the first voice information is filtered and processed using a delay compensation mechanism to obtain the second time-domain waveform of the first voice information;

[0019] The second feature information of the first voice information is obtained based on the second time-domain waveform of the first voice information.

[0020] Optionally, the step of determining the relationship between the first voice information and preset voice information based on the first feature information and the second feature information of the first voice information, and determining whether to perform the operation corresponding to the first voice information according to the determination result, includes:

[0021] Analyze the relationship between the first feature information of the first speech information and the first feature information of the preset speech information to obtain the first feature coefficient of the first speech information;

[0022] Analyze the relationship between the second feature information of the first speech information and the second feature information of the preset speech information to obtain the second feature coefficient of the first speech information;

[0023] Determine whether the first feature coefficient is less than a first threshold and whether the second feature coefficient is less than a second threshold;

[0024] If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the operation corresponding to the first voice information is executed.

[0025] Optionally, the step of determining that the first voice information matches the preset voice information and performing the operation corresponding to the first voice information if the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold includes:

[0026] If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the preset operation permission of the preset voice information is obtained;

[0027] Identify the first voice information and obtain the first operation corresponding to the first voice information;

[0028] Determine whether the first operation is within the preset operation permissions. If the first operation is within the preset operation permissions, then execute the first operation.

[0029] Optionally, the method further includes:

[0030] Perform spectral analysis on the first time-domain waveform of the first speech information to obtain the frequency of the first speech information;

[0031] Determine whether the frequency of the first voice information is within a preset frequency range;

[0032] If the frequency of the first voice information is within the preset frequency range, then the operation permissions of the user corresponding to the first voice information are set.

[0033] A voice information processing device, the device comprising: a first acquisition unit, a second acquisition unit, and a first processing unit; wherein:

[0034] The first acquisition unit is used to acquire first voice information;

[0035] The second acquisition unit is used to analyze and process the first voice information to obtain the first feature information and the second feature information of the first voice information;

[0036] The first processing unit is configured to determine the relationship between the first voice information and preset voice information based on the first feature information and the second feature information of the first voice information, and determine whether to perform an operation corresponding to the first voice information based on the determination result.

[0037] Optionally, the device further includes: a third acquisition unit, a first judgment unit, and a second processing unit; wherein:

[0038] The third acquisition unit is used to acquire the first time-domain waveform corresponding to the first voice information;

[0039] The first determining unit is used to determine whether the first time-domain waveform of the first voice information is continuous;

[0040] The second processing unit is configured to perform the analysis and processing of the first speech information to obtain the first feature information and the second feature information of the first speech information if the first time-domain waveform of the first speech information is continuous.

[0041] The second processing unit is further configured to reacquire the first voice information if the first time-domain waveform of the first voice information is discontinuous.

[0042] Optionally, the second acquisition unit includes: a first acquisition module and a second acquisition module; wherein:

[0043] The first acquisition module is used to perform spectrum analysis on the first time-domain waveform of the first speech information to obtain the frequency-domain waveform of the first speech information.

[0044] The first acquisition module is further configured to acquire the first feature information of the first voice information based on the frequency domain waveform of the first voice information;

[0045] The second acquisition module is used to filter the first time-domain waveform of the first voice information and process it using a delay compensation mechanism to obtain the second time-domain waveform of the first voice information.

[0046] The second acquisition module is further configured to acquire the second feature information of the first voice information based on the second time-domain waveform of the first voice information.

[0047] Optionally, the first processing unit includes: a third acquisition module, a judgment module, and a processing module; wherein:

[0048] The third acquisition module is used to analyze the relationship between the first feature information of the first voice information and the first feature information of the preset voice information to obtain the first feature coefficient of the first voice information.

[0049] The third acquisition module is further configured to analyze the relationship between the second feature information of the first voice information and the second feature information of the preset voice information, and obtain the second feature coefficient of the first voice information;

[0050] The judgment module is used to determine whether the first feature coefficient is less than the first threshold and whether the second feature coefficient is less than the second threshold.

[0051] The processing module is configured to determine that the first voice information matches the preset voice information and perform an operation corresponding to the first voice information if the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold.

[0052] Optionally, the processing module is further configured to:

[0053] If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the preset operation permission of the preset voice information is obtained;

[0054] Identify the first voice information and obtain the first operation corresponding to the first voice information;

[0055] Determine whether the first operation is within the preset operation permissions. If the first operation is within the preset operation permissions, then execute the first operation.

[0056] Optionally, the device further includes: a fourth acquisition unit, a second judgment unit, and a setting unit; wherein,

[0057] The fourth acquisition unit is used to perform spectrum analysis on the first time-domain waveform of the first speech information to acquire the frequency of the first speech information.

[0058] The second judgment unit is used to determine whether the frequency of the first voice information is within a preset frequency range;

[0059] The setting unit is used to set the operation permissions of the user corresponding to the first voice information if the frequency of the first voice information is within the preset frequency range.

[0060] The voice information processing method and apparatus provided in this invention can acquire first voice information, analyze and process the first voice information to obtain first feature information and second feature information of the first voice information, and determine the relationship between the first voice information and preset voice information based on the first feature information and second feature information of the first voice information. Finally, it determines whether to execute the operation corresponding to the first voice information based on the determination result. In this way, when performing user voice information recognition, the first feature information and second feature information of the user voice information can be considered simultaneously to identify the relationship between the user voice information and preset voice information. This solves the problem in the prior art that it is impossible to accurately identify the target user based on voice information, can accurately identify the target user and accurately match the user's operation permissions, avoid the situation where some users' operations are not within the actual operation permissions, and improve the interaction capability between the user and the device. Attached Figure Description

[0061] Figure 1 A flowchart illustrating a voice information processing method provided in an embodiment of the present invention;

[0062] Figure 2 A flowchart illustrating another voice information processing method provided in an embodiment of the present invention;

[0063] Figure 3 A flowchart illustrating another voice information processing method provided in an embodiment of the present invention;

[0064] Figure 4 A flowchart illustrating another voice information processing method provided in an embodiment of the present invention;

[0065] Figure 5 This is a schematic diagram of the structure of a voice information processing device provided in an embodiment of the present invention;

[0066] Figure 6 This is a schematic diagram of another voice information processing device provided in an embodiment of the present invention;

[0067] Figure 7 This is a schematic diagram of the structure of another voice information processing device provided in an embodiment of the present invention;

[0068] Figure 8 This is a schematic diagram of the structure of a voice information processing device according to another embodiment of the present invention;

[0069] Figure 9 This is a schematic diagram of another voice information processing device provided in another embodiment of the present invention. Detailed Implementation

[0070] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0071] This invention provides a voice information processing method, referring to... Figure 1 As shown, the method includes the following steps:

[0072] Step 101: Obtain the first voice information.

[0073] Specifically, step 101, acquiring the first voice information, can be implemented by a voice information processing device. The voice information processing device can be a smart device capable of voice recognition and executing corresponding operations, such as a smartphone, navigator, tablet, smart TV, smart refrigerator, smart relay, or air conditioner. The first voice information can be real-time voice information sent by the user that controls the smart device to perform related operations. The first voice information can be collected from the moment the user begins speaking and stops collecting after the user stops speaking for a specified period of time, which can be set by the user according to their preference, for example, 5 seconds.

[0074] Step 102: Analyze and process the first speech information to obtain the first feature information and the second feature information of the first speech information.

[0075] Specifically, step 102 involves analyzing and processing the first speech information to obtain first and second feature information, which can be achieved by a speech information processing device. The first feature information may include physical features of the sound such as timbre, resonance, and mode of resonance, while the second feature information may include behavioral features of the sound such as volume and speech rate.

[0076] Step 103: Based on the first feature information and the second feature information of the first voice information, determine the relationship between the first voice information and the preset voice information, and determine whether to perform the operation corresponding to the first voice information according to the determination result.

[0077] Specifically, step 103, based on the first and second feature information of the first voice information, determines the relationship between the first voice information and preset voice information, and determines whether to execute the operation corresponding to the first voice information according to the determination result. This can be implemented by a voice information processing device. The device compares the relationship between the first feature information of the first voice information and the first feature information of the preset voice information, and simultaneously compares the relationship between the second feature information of the first voice information and the second feature information of the preset voice information to determine whether the first voice information and the preset voice information match. If the first voice information matches the preset voice information, the operation corresponding to the first voice information is executed. If the first voice information does not match the preset voice information, the operation corresponding to the first voice information is not executed, and a corresponding prompt voice can be issued, such as "You do not have permission to operate."

[0078] The voice information processing method provided in this invention can acquire first voice information, analyze and process the first voice information to obtain first feature information and second feature information of the first voice information, and determine the relationship between the first voice information and preset voice information based on the first feature information and second feature information of the first voice information. Finally, it determines whether to execute the operation corresponding to the first voice information based on the judgment result. In this way, when performing user voice information recognition, the first feature information and second feature information of the user voice information can be considered simultaneously to identify the relationship between the user voice information and preset voice information. This solves the problem in the prior art that it is impossible to accurately identify the target user based on voice information. It can accurately identify the target user and accurately match the user's operation permissions, avoid the situation where some users' operations are not within the actual operation permissions, and improve the interaction capability between the user and the device.

[0079] This invention provides a voice information processing method, referring to... Figure 2 As shown, the method includes the following steps:

[0080] Step 201: The voice information processing device acquires the first voice information.

[0081] Step 202: The voice information processing device acquires the first time-domain waveform corresponding to the first voice information.

[0082] Specifically, the first time-domain waveform corresponding to the first speech information is the raw waveform of the acquired first speech information before processing.

[0083] Step 203: The voice information processing device determines whether the first time-domain waveform of the first voice information is continuous.

[0084] Specifically, it is determined whether the time-domain waveform of the first voice information, i.e., the real-time voice information sent by the user, is continuous within the reception time. The real-time voice information sent by the user has a continuous time-domain waveform throughout the reception time, while the recorded voice information has a time-domain waveform obtained through sampling by a digital device, and its corresponding time-domain waveform is generally discontinuous within the reception time.

[0085] In step 203, it is determined whether the first time-domain waveform of the first voice information is continuous. Step 204 or steps 205-206 can be executed. If the first time-domain waveform of the first voice information is discontinuous, step 204 is executed. If the first time-domain waveform of the first voice information is continuous, steps 205-206 are executed.

[0086] Step 204: If the first time-domain waveform of the first voice information is discontinuous, the voice information processing device reacquires the first voice information.

[0087] Specifically, if the first time-domain waveform of the first voice information is discontinuous, it indicates that the first voice information currently acquired is recorded voice information. In this case, the first voice information can be directly deleted and the voice information can be acquired again to avoid accidental operation.

[0088] Step 205: If the first time-domain waveform of the first speech information is continuous, the speech information processing device analyzes and processes the first speech information to obtain the first feature information and the second feature information of the first speech information.

[0089] Specifically, if the first time-domain waveform of the first voice information is continuous, it indicates that the first voice information is real-time voice information sent by the user. At this time, spectrum analysis can be performed on the first time-domain waveform of the first voice information to obtain the frequency-domain waveform of the first voice information. The first characteristic information such as timbre, resonance, and resonance mode of the first voice information can be obtained from the frequency-domain waveform of the first voice information. At the same time, the second characteristic information such as volume and speaking speed of the first voice information can be obtained from the first time-domain waveform of the first voice information.

[0090] Step 206: The voice information processing device determines the relationship between the first voice information and the preset voice information based on the first feature information and the second feature information of the first voice information, and determines whether to perform the operation corresponding to the first voice information based on the determination result.

[0091] Specifically, the preset voice information is at least one user voice information that is pre-recorded, sampled, and stored in the local system of the smart device or the corresponding cloud. When recording and sampling the preset voice information in the local system, a high compression rate of the sampled file can be used to reduce the user's network usage costs while ensuring the audio quality of the preset voice information as much as possible. The preset voice information can be stored in the cloud by the smart device communicating with the cloud via a wireless network to store the preset voice information obtained in the local system in the cloud.

[0092] It should be noted that the explanation of the same steps or concepts in this embodiment as in other embodiments can be found in the descriptions in other embodiments, and will not be repeated here.

[0093] The voice information processing method provided in this invention can acquire first voice information, analyze and process the first voice information to obtain first feature information and second feature information of the first voice information, and determine the relationship between the first voice information and preset voice information based on the first and second feature information. Finally, it determines whether to execute the operation corresponding to the first voice information based on the determination result. In this way, when performing user voice information recognition, the first and second feature information of the user voice information can be considered simultaneously to identify the relationship between the user voice information and the preset voice information. This solves the problem in the prior art that it is impossible to accurately identify the target user based on voice information, and can accurately identify the target user and accurately match the user's operation permissions, avoiding situations where some users' operations are not within their actual operation permissions, thus improving the interaction capability between the user and the device. Furthermore, it reduces the risk of unnecessary loss of life and property safety caused by erroneous operation due to the recognition of recorded voice information during the voice recognition process.

[0094] This invention provides a voice information processing method, referring to... Figure 3 As shown, the method includes the following steps:

[0095] Step 301: The voice information processing device acquires the first voice information.

[0096] Step 302: The voice information processing device acquires the first time-domain waveform corresponding to the first voice information.

[0097] Step 303: The voice information processing device determines whether the first time-domain waveform of the first voice information is continuous.

[0098] In step 303, the voice information processing device determines whether the first time-domain waveform of the first voice information is continuous. It can choose to execute step 304 or steps 305 to 312. If the first time-domain waveform of the first voice information is discontinuous, step 304 is executed. If the first time-domain waveform of the first voice information is continuous, steps 305 to 312 are executed.

[0099] Step 304: If the first time-domain waveform of the first voice information is discontinuous, the voice information processing device reacquires the first voice information.

[0100] Step 305: If the first time-domain waveform of the first speech information is continuous, the speech information processing device performs spectral analysis on the first time-domain waveform of the first speech information to obtain the frequency-domain waveform of the first speech information.

[0101] Specifically, the Fourier transform method can be used to perform spectral analysis on the first time-domain waveform of the first speech information to obtain the frequency-domain waveform of the first speech information.

[0102] Step 306: The voice information processing device obtains the first feature information of the first voice information based on the frequency domain waveform of the first voice information.

[0103] Specifically, a first feature analysis can be performed on the frequency domain waveform of the first speech information to obtain the first feature information of the first speech information. The specific implementation method of analyzing the frequency domain waveform to obtain the first feature information can refer to the implementation method of the existing technology, and will not be elaborated here.

[0104] Step 307: The voice information processing device filters the first time-domain waveform of the first voice information and processes it using a delay compensation mechanism to obtain the second time-domain waveform of the first voice information.

[0105] Specifically, the blank signals at the beginning and end of the first time-domain waveform of the first voice information can be filtered, and then a delay compensation mechanism can be used to process the first time-domain waveform of the filtered first voice information to obtain the second time-domain waveform of the second voice information; wherein, the second time-domain waveform and the time-domain waveform of the preset voice information can achieve dynamic consistency in terms of waveform distribution, peak-to-trough spacing, timestamp, etc.

[0106] Step 308: The voice information processing device obtains the second feature information of the first voice information based on the second time-domain waveform of the first voice information.

[0107] Specifically, a second feature analysis is performed on the second time-domain waveform of the first speech information to obtain the second feature information of the first speech information. The method for analyzing the time-domain waveform to obtain the second feature information can refer to the implementation method of the existing technology, which will not be elaborated here.

[0108] Step 309: The voice information processing device analyzes the relationship between the first feature information of the first voice information and the first feature information of the preset voice information to obtain the first feature coefficient of the first voice information.

[0109] Specifically, the first feature information of the first speech information can be obtained by subtracting the first feature information of the preset speech information and taking the absolute value. Of course, other methods used in the prior art can also be used to analyze the relationship between the first feature information of the first speech information and the first feature information of the preset speech information, and it is not limited to the implementation proposed in this invention. The method for obtaining the first feature information of the preset speech information can be the same as the method for obtaining the first feature information of the first speech information.

[0110] Step 310: The voice information processing device analyzes the relationship between the second feature information of the first voice information and the second feature information of the preset voice information to obtain the second feature coefficient of the first voice information.

[0111] Specifically, the second feature information of the first speech information can be subtracted from the second feature information of the preset speech information, and the absolute value can be taken to obtain the second feature coefficient of the first speech information. Of course, other methods used in the prior art can also be used to analyze the relationship between the second feature information of the first speech information and the second feature information of the preset speech information, and it is not limited to the implementation proposed in this invention. The method for obtaining the second feature information of the preset speech information can be the same as the method for obtaining the second feature information of the first speech information.

[0112] Step 311: The voice information processing device determines whether the first feature coefficient is less than the first threshold and whether the second feature coefficient is less than the second threshold.

[0113] Specifically, the first threshold can be a single value set for all first feature coefficients, or different values ​​can be set for different first feature coefficients. For example, the first thresholds of the three first feature coefficients obtained from the first feature information (timbre, resonance, and resonance mode) of the first speech information can be set to the same value. Alternatively, the first threshold corresponding to the first feature information (timbre) of the first speech information can be set as a first value, the first threshold corresponding to the first feature information (resonance) can be set as a second value, and the first feature coefficient corresponding to the first feature information (resonance mode) can be set as a third value. The second threshold can be a single value set for all second feature coefficients, or different values ​​can be set for different second feature coefficients. For example, the second thresholds of the two second feature coefficients obtained from the second feature information (volume level and speech rate) of the first speech information can be set to the same value. Alternatively, the second threshold corresponding to the second feature information (volume level) of the first speech information can be set as a fourth value, and the second threshold corresponding to the second feature information (speech rate) can be set as a fifth value. Users can set the first and second thresholds according to the actual application scenario and desired effect.

[0114] Step 312: If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, the voice information processing device determines that the first voice information matches the preset voice information and obtains the preset operation permission of the preset voice information.

[0115] Specifically, if the first feature coefficient of the first speech information is greater than or equal to the first threshold, and the second feature coefficient is greater than or equal to the second threshold, or if both the first and second feature coefficients of the first speech information are greater than or equal to the first threshold, then the first speech information is considered to not match the preset speech information, and the operation corresponding to the first speech information is not executed. During use, if, after judgment, it is determined that the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, the semantics of the first speech information can be matched with the semantics of the preset speech information to strengthen the verification process and ensure security.

[0116] It should be noted that step 312 can be implemented in the following specific way:

[0117] Step 312a: If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, the voice information processing device determines that the first voice information matches the preset voice information and obtains the preset operation permission of the preset voice information.

[0118] Specifically, preset operation permissions can be the range of operations that different users can perform on the smart device based on the device's function settings, thus improving the security of the operation; these preset operation permissions can be pre-set by the user and stored in the smart device.

[0119] Step 312b: The voice information processing device identifies the first voice information and obtains the first operation corresponding to the first voice information.

[0120] Specifically, semantic recognition can be performed on the first voice information to obtain the first operation corresponding to the first voice information; wherein, the first operation can be the operation that the user wants the smart device to perform, and the implementation method of semantic recognition can refer to the implementation method of existing technology, which will not be elaborated here.

[0121] Step 312c: The voice information processing device determines whether the first operation is within the preset operation permissions. If the first operation is within the preset operation permissions, the first operation is executed.

[0122] Specifically, determining whether the first operation is within the preset operation permissions can be achieved by determining whether the operation corresponding to the first voice information can be found in the preset operation range with the same operation. If the operation corresponding to the first voice information is the same as at least one operation in the preset operation range, the smart device responds and executes the first operation.

[0123] Based on the above embodiments, referring to Figure 4 As shown, in other embodiments of the present invention, the voice information processing method further includes:

[0124] Step 313: The voice information processing device performs spectrum analysis on the first time-domain waveform of the first voice information to obtain the frequency of the first voice information.

[0125] Step 314: The voice information processing device determines whether the frequency of the first voice information is within the preset frequency range.

[0126] Specifically, the preset frequency can be set according to the different voice frequencies corresponding to different age groups of users. In this embodiment, the preset frequency range can be set to the voice frequency range corresponding to minors; for example, taking the male voice frequency as an example: the voice frequency before puberty (minors) is 174.614Hz~184.997Hz, and the voice frequency after puberty (adults) is 87.307Hz~92.499Hz.

[0127] Step 315: If the frequency of the first voice information is within the preset frequency range, the voice information processing device sets the operation permissions for the user corresponding to the first voice information.

[0128] Specifically, if the frequency range of the first voice message is within a preset frequency range, it indicates that the user sending the voice message is a minor, and it is necessary to restrict the functions that minors can use on smart devices. For example, it is possible to disable smart devices or restrict certain functions of smart devices from being used, such as the smart relay stopping power supply to the power socket, preventing the use of pay channels on smart TVs, or preventing the use of game functions on smart devices; when the frequency of the first voice message is within the preset frequency range, the operation corresponding to the first voice message within the restricted function range is not executed; when the frequency of the first voice message is outside the preset range, the relationship between the first voice message and the preset voice message is determined, and subsequent processing procedures are executed according to the relationship between the first voice message and the preset voice message.

[0129] It should be noted that, in other embodiments of the present invention, the preset frequency range can also be the sound frequency range corresponding to an adult. If the frequency of the first voice information is outside the preset frequency range, the voice information processing device sets the operation permissions for the user corresponding to the first voice information. The setting of the preset frequency range can be based on the user's specific needs and wishes, or it can be set at the factory of the smart device. In all embodiments of the present invention, the first feature information can be the physical feature information of the sound, and the first feature coefficient can be the physical feature coefficient corresponding to the physical feature information; the second feature information can be the behavioral feature information of the sound, and the second feature coefficient can be the behavioral feature coefficient corresponding to the behavioral feature information.

[0130] It should be noted that the explanation of the same steps or concepts in this embodiment as in other embodiments can be found in the descriptions in other embodiments, and will not be repeated here.

[0131] The voice information processing method provided in this invention can acquire first voice information, analyze and process the first voice information to obtain first feature information and second feature information of the first voice information, and determine the relationship between the first voice information and preset voice information based on the first and second feature information. Finally, it determines whether to execute the operation corresponding to the first voice information based on the determination result. In this way, when performing user voice information recognition, the first and second feature information of the user voice information can be considered simultaneously to identify the relationship between the user voice information and the preset voice information. This solves the problem in the prior art that it is impossible to accurately identify the target user based on voice information, and can accurately identify the target user and accurately match the user's operation permissions, avoiding situations where some users' operations are not within their actual operation permissions, thus improving the interaction capability between the user and the device. Furthermore, it reduces the risk of unnecessary loss of life and property safety caused by erroneous operation due to the recognition of recorded voice information during the voice recognition process.

[0132] This invention provides a voice information processing device 4, which can be applied to... Figures 1-4 In a corresponding embodiment of a voice information processing method, referring to Figure 5 As shown, the device includes: a first acquisition unit 41, a second acquisition unit 42, and a first processing unit 43, wherein:

[0133] The first acquisition unit 41 is used to acquire the first voice information.

[0134] The second acquisition unit 42 is used to analyze and process the first speech information to obtain the first feature information and the second feature information of the first speech information.

[0135] The first processing unit 43 is used to determine the relationship between the first voice information and preset voice information based on the first feature information and the second feature information of the first voice information, and to determine whether to perform the operation corresponding to the first voice information according to the determination result.

[0136] The voice information processing device provided in this embodiment of the invention can acquire first voice information, analyze and process the first voice information to obtain first feature information and second feature information of the first voice information, and determine the relationship between the first voice information and preset voice information based on the first feature information and second feature information of the first voice information. Finally, it determines whether to execute the operation corresponding to the first voice information based on the judgment result. In this way, when performing user voice information recognition, the first feature information and second feature information of the user voice information can be considered simultaneously to identify the relationship between the user voice information and the preset voice information. This solves the problem in the prior art that it is impossible to accurately identify the target user based on voice information. It can accurately identify the target user and accurately match the user's operation permissions, avoid the situation where some users' operations are not within the actual operation permissions, and improve the interaction capability between the user and the device.

[0137] Specifically, refer to Figure 6 As shown, the device further includes: a third acquisition unit 44, a first judgment unit 45, and a second processing unit 46, wherein:

[0138] The third acquisition unit 44 is used to acquire the first time-domain waveform corresponding to the first voice information.

[0139] The first judgment unit 45 is used to determine whether the first time-domain waveform of the first voice information is continuous.

[0140] The second processing unit 46 is used to perform analysis and processing on the first speech information if the first time-domain waveform of the first speech information is continuous, so as to obtain the first feature information and the second feature information of the first speech information.

[0141] The second processing unit 46 is further configured to reacquire the first speech information if the first time-domain waveform of the first speech information is discontinuous.

[0142] Specifically, refer to Figure 7 As shown, the second acquisition unit 42 includes: a first acquisition module 421 and a second acquisition module 422, wherein:

[0143] The first acquisition module 421 is used to perform spectrum analysis on the first time-domain waveform of the first speech information to obtain the frequency-domain waveform of the first speech information.

[0144] The first acquisition module 421 is further configured to acquire the first feature information of the first speech information based on the frequency domain waveform of the first speech information.

[0145] The second acquisition module 422 is used to filter the first time-domain waveform of the first voice information and process it using a delay compensation mechanism to obtain the second time-domain waveform of the first voice information.

[0146] The second acquisition module 422 is further configured to acquire the second feature information of the first speech information based on the second time-domain waveform of the first speech information.

[0147] Specifically, refer to Figure 8 As shown, the first processing unit 43 includes: a third acquisition module 431, a judgment module 432, and a processing module 433, wherein:

[0148] The third acquisition module 431 is used to analyze the relationship between the first feature information of the first speech information and the first feature information of the preset speech information to obtain the first feature coefficient of the first speech information.

[0149] The third acquisition module 431 is also used to analyze the relationship between the second feature information of the first speech information and the second feature information of the preset speech information to obtain the second feature coefficient of the first speech information.

[0150] The judgment module 432 is used to determine whether the first feature coefficient is less than the first threshold and whether the second feature coefficient is less than the second threshold.

[0151] The processing module 433 is used to match the first speech information with preset speech information and perform the operation corresponding to the first speech information if the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold.

[0152] Specifically, optionally, processing module 433 is used to perform the following steps:

[0153] If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information matches the preset voice information and obtains the preset operation permission of the preset voice information.

[0154] Identify the first voice information and obtain the first operation corresponding to the first voice information.

[0155] Determine whether the first operation is within the preset operation permissions. If the first operation is within the preset operation permissions, then execute the first operation.

[0156] Specifically, refer to Figure 9 As shown, the device further includes: a fourth acquisition unit 47, a second judgment unit 48, and a setting unit 49, wherein:

[0157] The fourth acquisition unit 47 is used to perform spectrum analysis on the first time-domain waveform of the first speech information to acquire the frequency of the first speech information.

[0158] The second judgment unit 48 is used to determine whether the frequency of the first voice information is within a preset frequency range.

[0159] Setting unit 49 is used to set the operation permissions of the user corresponding to the first voice information if the frequency of the first voice information is within a preset frequency range.

[0160] It should be noted that the interaction process between the various units and modules in this embodiment can be referred to Figures 1-4 The corresponding embodiment provides an interactive process in a voice information processing method, which will not be described in detail here.

[0161] The voice information processing device provided in this invention can acquire first voice information, analyze and process the first voice information to obtain first feature information and second feature information of the first voice information, and determine the relationship between the first voice information and preset voice information based on the first and second feature information. Finally, it determines whether to execute the operation corresponding to the first voice information based on the determination result. In this way, when performing user voice information recognition, the first and second feature information of the user voice information can be considered simultaneously to identify the relationship between the user voice information and the preset voice information. This solves the problem in the prior art that it is impossible to accurately identify the target user based on voice information, and can accurately identify the target user and accurately match the user's operation permissions, avoiding situations where some users' operations are not within their actual operation permissions, thus improving the interaction capability between the user and the device. Furthermore, it reduces the risk of unnecessary loss to the user's life and property safety due to incorrect operation caused by recognizing recorded voice information during the voice recognition process.

[0162] In practical applications, the first acquisition unit 41, the second acquisition unit 42, the first processing unit 43, the third acquisition unit 44, the first judgment unit 45, the second processing unit 46, the fourth acquisition unit 47, the second judgment unit 48, the setting unit 49, the first acquisition module 421, the second acquisition module 422, the third acquisition module 431, the judgment module 432, and the processing module 433 can all be implemented by a central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP), or field programmable gate array (FPGA) located in the wireless data transmission device.

[0163] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0164] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0165] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0166] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0167] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention.

Claims

1. A voice information processing method characterized by comprising: The method includes: Obtain the first voice information; The first voice information is analyzed and processed to obtain first feature information and second feature information of the first voice information; wherein, the first feature information includes physical feature information of the sound, and the second feature information includes behavioral feature information of the sound. Analyze the relationship between the first feature information of the first voice information and the first feature information of the preset voice information, subtract the first feature information of the first voice information from the first feature information of the preset voice information and take the absolute value to obtain the first feature coefficient of the first voice information; Analyze the relationship between the second feature information of the first voice information and the second feature information of the preset voice information, subtract the second feature information of the first voice information from the second feature information of the preset voice information and take the absolute value to obtain the second feature coefficient of the first voice information; Determine whether the first feature coefficient is less than a first threshold and whether the second feature coefficient is less than a second threshold; If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the operation corresponding to the first voice information is executed. Before analyzing and processing the first speech information to obtain the first feature information and the second feature information of the first speech information, the method further includes: Obtain the first time-domain waveform corresponding to the first speech information, where the first time-domain waveform corresponding to the first speech information is the unprocessed original waveform of the first speech information collected. Determine whether the first time-domain waveform of the first voice information is continuous within the reception time. If the first time-domain waveform of the first voice information is continuous within the reception time, it indicates that the first voice information is real-time voice information sent by the user. If the first time-domain waveform of the first voice information is discontinuous within the reception time, it indicates that the first voice information is recorded voice information. If the first time-domain waveform of the first speech information is continuous, then the analysis and processing of the first speech information is performed to obtain the first feature information and the second feature information of the first speech information. If the first time-domain waveform of the first voice information is discontinuous, then the first voice information is reacquired.

2. The method according to claim 1, characterized in that, The step of analyzing and processing the first speech information to obtain first feature information and second feature information of the first speech information includes: Perform spectral analysis on the first time-domain waveform of the first speech information to obtain the frequency-domain waveform of the first speech information; Based on the frequency domain waveform of the first voice information, the first feature information of the first voice information is obtained; The first time-domain waveform of the first voice information is filtered and processed using a delay compensation mechanism to obtain the second time-domain waveform of the first voice information; The second feature information of the first voice information is obtained based on the second time-domain waveform of the first voice information.

3. The method according to claim 1, characterized in that, The step of determining that the first voice information matches the preset voice information and performing the operation corresponding to the first voice information if the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold includes: If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the preset operation permission of the preset voice information is obtained; Identify the first voice information and obtain the first operation corresponding to the first voice information; Determine whether the first operation is within the preset operation permissions. If the first operation is within the preset operation permissions, then execute the first operation.

4. A method for processing voice information, characterized in that, The method includes: Acquire first voice information, which is real-time voice information sent by the user that can control the smart device to perform related operations, and the first time-domain waveform corresponding to the first voice information is continuous within the receiving time. The first voice information is analyzed and processed to obtain first feature information and second feature information of the first voice information; wherein, the first feature information includes physical feature information of the sound, and the second feature information includes behavioral feature information of the sound. Analyze the relationship between the first feature information of the first voice information and the first feature information of the preset voice information, subtract the first feature information of the first voice information from the first feature information of the preset voice information and take the absolute value to obtain the first feature coefficient of the first voice information; Analyze the relationship between the second feature information of the first voice information and the second feature information of the preset voice information, subtract the second feature information of the first voice information from the second feature information of the preset voice information and take the absolute value to obtain the second feature coefficient of the first voice information; Determine whether the first feature coefficient is less than a first threshold and whether the second feature coefficient is less than a second threshold; If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the operation corresponding to the first voice information is executed. The step of analyzing and processing the first speech information to obtain the first feature information and the second feature information of the first speech information includes: Perform spectral analysis on the first time-domain waveform of the first speech information to obtain the frequency-domain waveform of the first speech information; Based on the frequency domain waveform of the first voice information, the first feature information of the first voice information is obtained; The first time-domain waveform of the first voice information is filtered and processed using a delay compensation mechanism to obtain the second time-domain waveform of the first voice information; The second feature information of the first voice information is obtained based on the second time-domain waveform of the first voice information.

5. The method according to claim 4, characterized in that, The step of determining that the first voice information matches the preset voice information and performing the operation corresponding to the first voice information if the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold includes: If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the preset operation permission of the preset voice information is obtained; Identify the first voice information and obtain the first operation corresponding to the first voice information; Determine whether the first operation is within the preset operation permissions. If the first operation is within the preset operation permissions, then execute the first operation.

6. A method for processing voice information, characterized in that, The method includes: Acquire first voice information, which is real-time voice information sent by the user that can control the smart device to perform related operations, and the first time-domain waveform corresponding to the first voice information is continuous within the receiving time. The first voice information is analyzed and processed to obtain first feature information and second feature information of the first voice information; wherein, the first feature information includes physical feature information of the sound, and the second feature information includes behavioral feature information of the sound. Analyze the relationship between the first feature information of the first voice information and the first feature information of the preset voice information, subtract the first feature information of the first voice information from the first feature information of the preset voice information and take the absolute value to obtain the first feature coefficient of the first voice information; Analyze the relationship between the second feature information of the first voice information and the second feature information of the preset voice information, subtract the second feature information of the first voice information from the second feature information of the preset voice information and take the absolute value to obtain the second feature coefficient of the first voice information; Determine whether the first feature coefficient is less than a first threshold and whether the second feature coefficient is less than a second threshold; If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the operation corresponding to the first voice information is executed. Wherein, the step of determining that the first voice information matches the preset voice information and performing the operation corresponding to the first voice information if the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold includes: If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the preset operation permission of the preset voice information is obtained; Identify the first voice information and obtain the first operation corresponding to the first voice information; Determine whether the first operation is within the preset operation permissions. If the first operation is within the preset operation permissions, then execute the first operation.

7. The method according to claim 6, characterized in that, The method further includes: Perform spectral analysis on the first time-domain waveform of the first speech information to obtain the frequency of the first speech information; Determine whether the frequency of the first voice information is within a preset frequency range; If the frequency of the first voice information is within the preset frequency range, then the operation permissions of the user corresponding to the first voice information are set.

8. A method for processing voice information, characterized in that, The method includes: Acquire first voice information, which is real-time voice information sent by the user that can control the smart device to perform related operations, and the first time-domain waveform corresponding to the first voice information is continuous within the receiving time. The first voice information is analyzed and processed to obtain first feature information and second feature information of the first voice information; wherein, the first feature information includes physical feature information of the sound, and the second feature information includes behavioral feature information of the sound. Analyze the relationship between the first feature information of the first voice information and the first feature information of the preset voice information, subtract the first feature information of the first voice information from the first feature information of the preset voice information and take the absolute value to obtain the first feature coefficient of the first voice information; Analyze the relationship between the second feature information of the first voice information and the second feature information of the preset voice information, subtract the second feature information of the first voice information from the second feature information of the preset voice information and take the absolute value to obtain the second feature coefficient of the first voice information; Determine whether the first feature coefficient is less than a first threshold and whether the second feature coefficient is less than a second threshold; If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the operation corresponding to the first voice information is executed. The method further includes: Perform spectral analysis on the first time-domain waveform of the first speech information to obtain the frequency of the first speech information; Determine whether the frequency of the first voice information is within a preset frequency range; If the frequency of the first voice information is within the preset frequency range, then the operation permissions of the user corresponding to the first voice information are set.

9. A method for processing voice information, characterized in that, The method includes: Obtain the first voice information; The first voice information is analyzed and processed to obtain first feature information and second feature information of the first voice information; wherein, the first feature information includes physical feature information of the sound, and the second feature information includes behavioral feature information of the sound. Based on the first feature information and the second feature information of the first voice information, determine the relationship between the first voice information and the preset voice information, and determine whether to perform the operation corresponding to the first voice information according to the determination result; Perform spectral analysis on the first time-domain waveform of the first speech information to obtain the frequency of the first speech information; Determine whether the frequency of the first voice information is within a preset frequency range; If the frequency of the first voice information is within the preset frequency range, then set the operation permissions for the user corresponding to the first voice information. Before analyzing and processing the first speech information to obtain the first feature information and the second feature information of the first speech information, the method further includes: Obtain the first time-domain waveform corresponding to the first speech information, where the first time-domain waveform corresponding to the first speech information is the unprocessed original waveform of the first speech information collected. Determine whether the first time-domain waveform of the first voice information is continuous within the reception time. If the first time-domain waveform of the first voice information is continuous within the reception time, it indicates that the first voice information is real-time voice information sent by the user. If the first time-domain waveform of the first voice information is discontinuous within the reception time, it indicates that the first voice information is recorded voice information. If the first time-domain waveform of the first speech information is continuous, then the analysis and processing of the first speech information is performed to obtain the first feature information and the second feature information of the first speech information. If the first time-domain waveform of the first voice information is discontinuous, then the first voice information is reacquired. The step of determining the relationship between the first voice information and preset voice information based on the first and second feature information of the first voice information, and determining whether to execute the operation corresponding to the first voice information based on the determination result, includes: Analyze the relationship between the first feature information of the first voice information and the first feature information of the preset voice information, subtract the first feature information of the first voice information from the first feature information of the preset voice information and take the absolute value to obtain the first feature coefficient of the first voice information; Analyze the relationship between the second feature information of the first voice information and the second feature information of the preset voice information, subtract the second feature information of the first voice information from the second feature information of the preset voice information and take the absolute value to obtain the second feature coefficient of the first voice information; Determine whether the first feature coefficient is less than a first threshold and whether the second feature coefficient is less than a second threshold; If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the operation corresponding to the first voice information is executed.

10. The method according to claim 9, characterized in that, The step of analyzing and processing the first speech information to obtain first feature information and second feature information of the first speech information includes: Perform spectral analysis on the first time-domain waveform of the first speech information to obtain the frequency-domain waveform of the first speech information; Based on the frequency domain waveform of the first voice information, the first feature information of the first voice information is obtained; The first time-domain waveform of the first voice information is filtered and processed using a delay compensation mechanism to obtain the second time-domain waveform of the first voice information; The second feature information of the first voice information is obtained based on the second time-domain waveform of the first voice information.

11. The method according to claim 9, characterized in that, The step of determining that the first voice information matches the preset voice information and performing the operation corresponding to the first voice information if the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold includes: If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the preset operation permission of the preset voice information is obtained; Identify the first voice information and obtain the first operation corresponding to the first voice information; Determine whether the first operation is within the preset operation permissions. If the first operation is within the preset operation permissions, then execute the first operation.

12. A method for processing voice information, characterized in that, The method includes: Acquire first voice information, which is real-time voice information sent by the user that can control the smart device to perform related operations, and the first time-domain waveform corresponding to the first voice information is continuous within the receiving time. The first voice information is analyzed and processed to obtain first feature information and second feature information of the first voice information; wherein, the first feature information includes physical feature information of the sound, and the second feature information includes behavioral feature information of the sound. Based on the first feature information and the second feature information of the first voice information, determine the relationship between the first voice information and the preset voice information, and determine whether to perform the operation corresponding to the first voice information according to the determination result; Perform spectral analysis on the first time-domain waveform of the first speech information to obtain the frequency of the first speech information; Determine whether the frequency of the first voice information is within a preset frequency range; If the frequency of the first voice information is within the preset frequency range, then set the operation permissions for the user corresponding to the first voice information. The step of analyzing and processing the first speech information to obtain the first feature information and the second feature information of the first speech information includes: Perform spectral analysis on the first time-domain waveform of the first speech information to obtain the frequency-domain waveform of the first speech information; Based on the frequency domain waveform of the first voice information, the first feature information of the first voice information is obtained; The first time-domain waveform of the first voice information is filtered and processed using a delay compensation mechanism to obtain the second time-domain waveform of the first voice information; Based on the second time-domain waveform of the first voice information, the second feature information of the first voice information is obtained; The step of determining the relationship between the first voice information and preset voice information based on the first and second feature information of the first voice information, and determining whether to execute the operation corresponding to the first voice information based on the determination result, includes: Analyze the relationship between the first feature information of the first voice information and the first feature information of the preset voice information, subtract the first feature information of the first voice information from the first feature information of the preset voice information and take the absolute value to obtain the first feature coefficient of the first voice information; Analyze the relationship between the second feature information of the first voice information and the second feature information of the preset voice information, subtract the second feature information of the first voice information from the second feature information of the preset voice information and take the absolute value to obtain the second feature coefficient of the first voice information; Determine whether the first feature coefficient is less than a first threshold and whether the second feature coefficient is less than a second threshold; If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the operation corresponding to the first voice information is executed.

13. The method according to claim 12, characterized in that, The step of determining that the first voice information matches the preset voice information and performing the operation corresponding to the first voice information if the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold includes: If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the preset operation permission of the preset voice information is obtained; Identify the first voice information and obtain the first operation corresponding to the first voice information; Determine whether the first operation is within the preset operation permissions. If the first operation is within the preset operation permissions, then execute the first operation.

14. A voice information processing device, characterized in that, The device includes: a first acquisition unit, a second acquisition unit, and a first processing unit; wherein: The first acquisition unit is used to acquire first voice information; The second acquisition unit is used to analyze and process the first voice information to obtain first feature information and second feature information of the first voice information; wherein, the first feature information includes physical feature information of the sound, and the second feature information includes behavioral feature information of the sound. The first processing unit is configured to determine the relationship between the first voice information and preset voice information based on the first feature information and the second feature information of the first voice information, and determine whether to perform an operation corresponding to the first voice information based on the determination result. The first processing unit includes: a third acquisition module, a judgment module, and a processing module; wherein: The third acquisition module is used to analyze the relationship between the first feature information of the first voice information and the first feature information of the preset voice information, and to subtract the first feature information of the first voice information from the first feature information of the preset voice information and take the absolute value to obtain the first feature coefficient of the first voice information. The third acquisition module is further configured to analyze the relationship between the second feature information of the first voice information and the second feature information of the preset voice information, subtract the second feature information of the first voice information from the second feature information of the preset voice information and take the absolute value to obtain the second feature coefficient of the first voice information; The judgment module is used to determine whether the first feature coefficient is less than the first threshold and whether the second feature coefficient is less than the second threshold. The processing module is configured to determine that the first voice information matches the preset voice information and perform an operation corresponding to the first voice information if the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold. The device further includes: a third acquisition unit, a first judgment unit, and a second processing unit; wherein: The third acquisition unit is used to acquire the first time-domain waveform corresponding to the first voice information, wherein the first time-domain waveform corresponding to the first voice information is the unprocessed original waveform of the acquired first voice information. The first judgment unit is used to judge whether the first time-domain waveform of the first voice information is continuous. If the first time-domain waveform of the first voice information is continuous, it indicates that the first voice information is real-time voice information sent by the user. If the first time-domain waveform of the first voice information is discontinuous, it indicates that the first voice information is recorded voice information. The second processing unit is configured to perform the analysis and processing of the first speech information to obtain the first feature information and the second feature information of the first speech information if the first time-domain waveform of the first speech information is continuous. The second processing unit is further configured to reacquire the first voice information if the first time-domain waveform of the first voice information is discontinuous.

15. The apparatus according to claim 14, characterized in that, The second acquisition unit includes: a first acquisition module and a second acquisition module; wherein: The first acquisition module is used to perform spectrum analysis on the first time-domain waveform of the first speech information to obtain the frequency-domain waveform of the first speech information. The first acquisition module is further configured to acquire the first feature information of the first voice information based on the frequency domain waveform of the first voice information; The second acquisition module is used to filter the first time-domain waveform of the first voice information and process it using a delay compensation mechanism to obtain the second time-domain waveform of the first voice information. The second acquisition module is further configured to acquire the second feature information of the first voice information based on the second time-domain waveform of the first voice information.

16. The apparatus according to claim 14, characterized in that, The processing module is further specifically used for: If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the preset operation permission of the preset voice information is obtained; Identify the first voice information and obtain the first operation corresponding to the first voice information; Determine whether the first operation is within the preset operation permissions. If the first operation is within the preset operation permissions, then execute the first operation.

17. A voice information processing device, characterized in that, The device includes: a first acquisition unit, a second acquisition unit, and a first processing unit; wherein: The first acquisition unit is used to acquire first voice information, which is real-time voice information sent by the user that can control the smart device to perform related operations; The second acquisition unit is used to analyze and process the first voice information to obtain first feature information and second feature information of the first voice information; wherein, the first feature information includes physical feature information of the sound, and the second feature information includes behavioral feature information of the sound. The first processing unit is configured to determine the relationship between the first voice information and preset voice information based on the first feature information and the second feature information of the first voice information, and determine whether to perform an operation corresponding to the first voice information based on the determination result. The first processing unit includes: a third acquisition module, a judgment module, and a processing module; wherein: The third acquisition module is used to analyze the relationship between the first feature information of the first voice information and the first feature information of the preset voice information, and to subtract the first feature information of the first voice information from the first feature information of the preset voice information and take the absolute value to obtain the first feature coefficient of the first voice information. The third acquisition module is further configured to analyze the relationship between the second feature information of the first voice information and the second feature information of the preset voice information, subtract the second feature information of the first voice information from the second feature information of the preset voice information and take the absolute value to obtain the second feature coefficient of the first voice information; The judgment module is used to determine whether the first feature coefficient is less than the first threshold and whether the second feature coefficient is less than the second threshold. The processing module is configured to determine that the first voice information matches the preset voice information and perform an operation corresponding to the first voice information if the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold. The second acquisition unit includes: a first acquisition module and a second acquisition module; wherein: The first acquisition module is used to perform spectrum analysis on the first time-domain waveform of the first speech information to obtain the frequency-domain waveform of the first speech information. The first acquisition module is further configured to acquire the first feature information of the first voice information based on the frequency domain waveform of the first voice information; The second acquisition module is used to filter the first time-domain waveform of the first voice information and process it using a delay compensation mechanism to obtain the second time-domain waveform of the first voice information. The second acquisition module is further configured to acquire the second feature information of the first voice information based on the second time-domain waveform of the first voice information.

18. The apparatus according to claim 17, characterized in that, The processing module is further specifically used for: If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the preset operation permission of the preset voice information is obtained; Identify the first voice information and obtain the first operation corresponding to the first voice information; Determine whether the first operation is within the preset operation permissions. If the first operation is within the preset operation permissions, then execute the first operation.

19. A voice information processing device, characterized in that, The device includes: a first acquisition unit, a second acquisition unit, and a first processing unit; wherein: The first acquisition unit is used to acquire first voice information, which is real-time voice information sent by the user that can control the smart device to perform related operations; The second acquisition unit is used to analyze and process the first voice information to obtain first feature information and second feature information of the first voice information; wherein, the first feature information includes physical feature information of the sound, and the second feature information includes behavioral feature information of the sound. The first processing unit is configured to determine the relationship between the first voice information and preset voice information based on the first feature information and the second feature information of the first voice information, and determine whether to perform an operation corresponding to the first voice information based on the determination result. The first processing unit includes: a third acquisition module, a judgment module, and a processing module; wherein: The third acquisition module is used to analyze the relationship between the first feature information of the first voice information and the first feature information of the preset voice information, and to subtract the first feature information of the first voice information from the first feature information of the preset voice information and take the absolute value to obtain the first feature coefficient of the first voice information. The third acquisition module is further configured to analyze the relationship between the second feature information of the first voice information and the second feature information of the preset voice information, subtract the second feature information of the first voice information from the second feature information of the preset voice information and take the absolute value to obtain the second feature coefficient of the first voice information; The judgment module is used to determine whether the first feature coefficient is less than the first threshold and whether the second feature coefficient is less than the second threshold. The processing module is configured to determine that the first voice information matches the preset voice information and perform an operation corresponding to the first voice information if the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold. The device further includes: a fourth acquisition unit, a second judgment unit, and a setting unit; wherein... The fourth acquisition unit is used to perform spectrum analysis on the first time-domain waveform of the first speech information to acquire the frequency of the first speech information. The second judgment unit is used to determine whether the frequency of the first voice information is within a preset frequency range; The setting unit is used to set the operation permissions of the user corresponding to the first voice information if the frequency of the first voice information is within the preset frequency range.

20. A voice information processing device, characterized in that, The device includes: a first acquisition unit, a second acquisition unit, and a first processing unit; wherein: The first acquisition unit is used to acquire first voice information; The second acquisition unit is used to analyze and process the first voice information to obtain first feature information and second feature information of the first voice information; wherein, the first feature information includes physical feature information of the sound, and the second feature information includes behavioral feature information of the sound. The first processing unit is configured to determine the relationship between the first voice information and preset voice information based on the first feature information and the second feature information of the first voice information, and determine whether to perform an operation corresponding to the first voice information based on the determination result. The device further includes: a fourth acquisition unit, a second judgment unit, and a setting unit; wherein... The fourth acquisition unit is used to perform spectrum analysis on the first time-domain waveform of the first speech information to acquire the frequency of the first speech information. The second judgment unit is used to determine whether the frequency of the first voice information is within a preset frequency range; The setting unit is used to set the operation permissions of the user corresponding to the first voice information if the frequency of the first voice information is within the preset frequency range. The device further includes: a third acquisition unit, a first judgment unit, and a second processing unit; wherein: The third acquisition unit is used to acquire the first time-domain waveform corresponding to the first voice information, wherein the first time-domain waveform corresponding to the first voice information is the unprocessed original waveform of the acquired first voice information. The first judgment unit is used to judge whether the first time-domain waveform of the first voice information is continuous. If the first time-domain waveform of the first voice information is continuous, it indicates that the first voice information is real-time voice information sent by the user. If the first time-domain waveform of the first voice information is discontinuous, it indicates that the first voice information is recorded voice information. The second processing unit is configured to perform the analysis and processing of the first speech information to obtain the first feature information and the second feature information of the first speech information if the first time-domain waveform of the first speech information is continuous. The second processing unit is further configured to reacquire the first voice information if the first time-domain waveform of the first voice information is discontinuous. The first processing unit includes: a third acquisition module, a judgment module, and a processing module; wherein: The third acquisition module is used to analyze the relationship between the first feature information of the first voice information and the first feature information of the preset voice information, and to subtract the first feature information of the first voice information from the first feature information of the preset voice information and take the absolute value to obtain the first feature coefficient of the first voice information. The third acquisition module is further configured to analyze the relationship between the second feature information of the first voice information and the second feature information of the preset voice information, subtract the second feature information of the first voice information from the second feature information of the preset voice information and take the absolute value to obtain the second feature coefficient of the first voice information; The judgment module is used to determine whether the first feature coefficient is less than the first threshold and whether the second feature coefficient is less than the second threshold. The processing module is configured to determine that the first voice information matches the preset voice information and perform an operation corresponding to the first voice information if the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold.

21. The apparatus according to claim 20, characterized in that, The second acquisition unit includes: a first acquisition module and a second acquisition module; wherein: The first acquisition module is used to perform spectrum analysis on the first time-domain waveform of the first speech information to obtain the frequency-domain waveform of the first speech information. The first acquisition module is further configured to acquire the first feature information of the first voice information based on the frequency domain waveform of the first voice information; The second acquisition module is used to filter the first time-domain waveform of the first voice information and process it using a delay compensation mechanism to obtain the second time-domain waveform of the first voice information. The second acquisition module is further configured to acquire the second feature information of the first voice information based on the second time-domain waveform of the first voice information.

22. The apparatus according to claim 20, characterized in that, The processing module is further specifically used for: If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the preset operation permission of the preset voice information is obtained; Identify the first voice information and obtain the first operation corresponding to the first voice information; Determine whether the first operation is within the preset operation permissions. If the first operation is within the preset operation permissions, then execute the first operation.

23. A voice information processing device, characterized in that, The device includes: a first acquisition unit, a second acquisition unit, and a first processing unit; wherein: The first acquisition unit is used to acquire first voice information, which is real-time voice information sent by the user that can control the smart device to perform related operations; The second acquisition unit is used to analyze and process the first voice information to obtain first feature information and second feature information of the first voice information; wherein, the first feature information includes physical feature information of the sound, and the second feature information includes behavioral feature information of the sound. The first processing unit is configured to determine the relationship between the first voice information and preset voice information based on the first feature information and the second feature information of the first voice information, and determine whether to perform an operation corresponding to the first voice information based on the determination result. The device further includes: a fourth acquisition unit, a second judgment unit, and a setting unit; wherein... The fourth acquisition unit is used to perform spectrum analysis on the first time-domain waveform of the first speech information to acquire the frequency of the first speech information. The second judgment unit is used to determine whether the frequency of the first voice information is within a preset frequency range; The setting unit is used to set the operation permissions of the user corresponding to the first voice information if the frequency of the first voice information is within the preset frequency range. The second acquisition unit includes: a first acquisition module and a second acquisition module; wherein: The first acquisition module is used to perform spectrum analysis on the first time-domain waveform of the first speech information to obtain the frequency-domain waveform of the first speech information. The first acquisition module is further configured to acquire the first feature information of the first voice information based on the frequency domain waveform of the first voice information; The second acquisition module is used to filter the first time-domain waveform of the first voice information and process it using a delay compensation mechanism to obtain the second time-domain waveform of the first voice information. The second acquisition module is further configured to acquire the second feature information of the first voice information based on the second time-domain waveform of the first voice information; The first processing unit includes: a third acquisition module, a judgment module, and a processing module; wherein: The third acquisition module is used to analyze the relationship between the first feature information of the first voice information and the first feature information of the preset voice information, and to subtract the first feature information of the first voice information from the first feature information of the preset voice information and take the absolute value to obtain the first feature coefficient of the first voice information. The third acquisition module is further configured to analyze the relationship between the second feature information of the first voice information and the second feature information of the preset voice information, subtract the second feature information of the first voice information from the second feature information of the preset voice information and take the absolute value to obtain the second feature coefficient of the first voice information; The judgment module is used to determine whether the first feature coefficient is less than the first threshold and whether the second feature coefficient is less than the second threshold. The processing module is configured to determine that the first voice information matches the preset voice information and perform an operation corresponding to the first voice information if the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold.

24. The apparatus according to claim 23, characterized in that, The processing module is further specifically used for: If the first feature coefficient is less than the first threshold and the second feature coefficient is less than the second threshold, then the first voice information is determined to match the preset voice information and the preset operation permission of the preset voice information is obtained; Identify the first voice information and obtain the first operation corresponding to the first voice information; Determine whether the first operation is within the preset operation permissions. If the first operation is within the preset operation permissions, then execute the first operation.