Voice security interaction method, terminal device and storage medium

Through dynamic encryption of sensitive vocabulary and multi-level voice command processing, the problem of low data security in traditional AI voice interaction is solved, and effective protection of user privacy and data security improvement of voice interaction is achieved.

CN120164463AInactive Publication Date: 2025-06-17SHENZHEN KAICHUANG FUTURE TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510316882.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The lack of processing of private information during traditional AI voice interaction leads to low data security.

Method used

By dynamically encrypting sensitive words, users' privacy is protected in remote control instructions, remote control instructions are processed differently, excessive encryption of local control instructions, and multi-level voice command processing is implemented to form an end-to-end security link.

Benefits of technology

It effectively protects user privacy, reduces the risk of privacy leakage, reduces dependence on cloud security measures, and improves the data security of voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164463A_ABST
    Figure CN120164463A_ABST
Patent Text Reader

Abstract

The invention is suitable for the field of data processing, and discloses a voice secure interaction method, terminal equipment and a storage medium. The voice security interaction method comprises the following steps: collecting voice information of a user; if the voice information carries the voice instruction, judging whether the voice instruction is a remote control instruction; if yes, whether the voice instruction carries sensitive vocabularies or not is judged; and if the voice instruction carries the sensitive vocabulary, encrypting the sensitive vocabulary in the voice instruction to obtain a target voice instruction. According to the invention, the dependence on cloud security measures is reduced, and the data security of voice interaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of audio processing, and particularly relates to a secure voice interaction method, a terminal device, and a storage medium. Background Art

[0002] With the popularization of intelligent devices, voice interaction has become an important way of human-computer interaction. As a portable intelligent device, the AI headset has become an ideal carrier for voice interaction due to its portability and real-time performance.

[0003] In the traditional AI voice interaction process, the processing of privacy information is often lacking. For example, when a user uses a voice assistant to input a voice command and needs to call a cloud service, the voice assistant often directly converts the user's voice command into text and sends the privacy information to the cloud server for processing. This method has the technical problem of low data security. A new technology is needed to solve the above technical problem. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a secure voice interaction method, a terminal device, and a storage medium, which can solve the problem of low data security in AI voice interaction in related technologies.

[0005] The first aspect of the present invention provides a secure voice interaction method, including: Collecting the voice information of the user; If the voice information carries a voice command, determining whether the voice command is a remote control command; If so, determining whether the voice command carries sensitive words; If the voice command carries the sensitive words, encrypting the sensitive words in the voice command to obtain a target voice command.

[0006] Optionally, in the first implementation manner of the first aspect of the present invention, the step of collecting the voice information of the user includes: Controlling a microphone array to collect the initial voice information of the user; Performing audio processing on the initial voice information to obtain the voice information.

[0007] Optionally, in the second implementation manner of the first aspect of the present invention, before the step of if the voice information carries a voice command, determining whether the voice command is a remote control command, the method further includes: Invoking a pre-trained voice recognition model to process the voice information; If the voice recognition model determines that the voice information carries a preset audio feature, determining that the voice information carries the voice command.

[0008] Optionally, in the third implementation manner of the first aspect of the present invention, the step of controlling the microphone array to collect the initial voice information of the user further includes Invoking a beamforming algorithm to control the microphone array to collect the initial voice information of the user.

[0009] Optionally, in the fourth implementation manner of the first aspect of the present invention, the step of performing audio processing on the initial voice information to obtain the voice information includes: Performing noise reduction and echo cancellation processing on the initial voice information to obtain the voice information.

[0010] Optionally, in the fifth implementation manner of the first aspect of the present invention, after the step of encrypting the sensitive word in the voice command to obtain a target voice command if the voice command carries the sensitive word, the method further includes Sending the target voice command to the connected target device via Bluetooth.

[0011] Optionally, in the sixth implementation manner of the first aspect of the present invention, the step of determining whether the voice command is a remote control command includes: Determining whether it is necessary to forward the voice command from the connected target device to a remote server; If so, determining that the voice command is a remote control command.

[0012] Optionally, in the seventh implementation manner of the first aspect of the present invention, after the step of determining whether the voice command is a remote control command if the voice information carries a voice command, the method further includes: If not, determining that the voice command is a local control command and sending the voice command to the target device via Bluetooth.

[0013] In a second aspect, an embodiment of the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned secure voice interaction method are implemented.

[0014] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned secure voice interaction method are implemented.

[0015] In a fourth aspect, an embodiment of the present invention provides a computer program product, which when running on a terminal device causes the terminal device to execute the above-mentioned secure voice interaction method.

[0016] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows: By dynamically encrypting sensitive words, user privacy is protected in the remote control instructions, reducing the risk of privacy leakage. The remote control instructions are processed separately, effectively avoiding excessive encryption of local control instructions and reducing resource waste. Multilevel voice instruction processing is implemented, and sensitive word detection and encryption are embedded throughout the process from voice collection to instruction execution, forming an end-to-end secure link, enabling sensitive information to be encrypted at the generation stage and effectively avoiding the leakage of plaintext data. The dependence on cloud security measures is significantly reduced, and the data security of voice interaction is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.

[0018] Figure 1 It is a schematic diagram of an embodiment of the secure voice interaction method in the embodiments of the present invention; Figure 2 It is a schematic diagram of a specific embodiment of step S101 of the secure voice interaction method in the embodiments of the present invention; Figure 3 It is a schematic diagram of a specific embodiment of step S101 of the secure voice interaction method in the embodiments of the present invention; Figure 4 It is a schematic diagram of a specific embodiment of step S102 of the secure voice interaction method in the embodiments of the present invention; Figure 5 It is a schematic diagram of an embodiment of the terminal device in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further elaborates on the present invention in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0020] It should be noted that the terms "including", "comprising", "having" and any variations thereof in the specification, claims and above-mentioned drawings of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, terminal, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices. In the claims, specification and drawings of the present invention, relational terms such as "first" and "second" are only used to distinguish one entity / operation / object from another entity / operation / object, and do not necessarily require or imply any such actual relationship or order between these entities / operations / objects.

[0021] Reference to "embodiment" herein means that a particular feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0022] With the popularization of intelligent devices, voice interaction has become an important way of human-computer interaction. As a portable intelligent device, the AI headset has become an ideal carrier for voice interaction due to its portability and real-time nature.

[0023] In the traditional AI voice interaction process, the processing of privacy information is often lacking. For example, when a user uses a voice assistant to input a voice command and needs to call a cloud service, the voice assistant often directly converts the user's voice command into text and sends the privacy information to the cloud server for processing. This method has the technical problem of low data security. A new technology is needed to solve the above technical problems.

[0024] In view of this, the embodiments of the present invention provide a secure voice interaction method, terminal device and storage medium, which protect user privacy in remote control instructions by dynamically encrypting sensitive words, reducing the risk of privacy leakage. Differentiating the processing of remote control instructions effectively avoids over-encryption of local control instructions and reduces resource waste. Implementing multi-level voice instruction processing, embedding sensitive word detection and encryption throughout the process from voice collection to instruction execution, forming an end-to-end secure link, enables sensitive information to be encrypted at the generation stage, effectively avoiding the leakage of plaintext data. Significantly reducing the dependence on cloud security measures and enhancing the data security of voice interaction.

[0025] The professional terms involved in the embodiments of the present invention include but are not limited to: Microphone Array: A hardware system composed of multiple (at least two) microphones that enhances the voice acquisition ability through spatial distribution and collaborative work.

[0026] Beamforming Technology: A signal processing technology that forms a directional receiving beam by adjusting the phase and gain of each microphone to focus on sound sources in a specific direction.

[0027] Acoustic Echo Cancellation (AEC): Eliminates the interference of the sound played by the device's own speaker on the microphone input through an algorithm.

[0028] Definition of Gain Control: Dynamically adjusts the volume amplitude of the voice signal to ensure stable output.

[0029] Speech Recognition Model: An algorithm model trained based on artificial intelligence (such as deep learning) used to convert voice signals into text instructions.

[0030] Preset Audio Features: Predetermined voice features, including wake-up words (such as "Xiaoi Assistant"), command keywords (such as "play"), or voiceprint features.

[0031] Mel Frequency Cepstral Coefficients (MFCC): An algorithm for speech feature extraction that simulates the auditory characteristics of the human ear and converts voice signals into frequency-domain feature vectors.

[0032] Sensitive Word Encryption: Encrypts privacy information (such as passwords and ID numbers) in voice commands to generate unreadable ciphertext.

[0033] End-to-End Encryption (E2EE): An encryption method where data only exists in plaintext at the sending and receiving ends, and the transmission link and intermediate nodes (such as the cloud) cannot decrypt it.

[0034] Key Management: Securely controls the generation, storage, distribution, and destruction of encryption keys.

[0035] BLE Packet (Bluetooth Low Energy Packet): A data unit transmitted in the Bluetooth low energy mode, containing command content, check information, etc.

[0036] Remote Control Command: A command that needs to be executed through a cloud server (such as sending a message, querying the weather).

[0037] Local Control Command: A command that only needs to be executed locally on the target device (such as adjusting the volume, playing music).

[0038] Natural Language Processing (NLP) model: An AI model for understanding the semantics of speech instructions (such as intent recognition, entity extraction).

[0039] To illustrate the technical solution of the present invention, the following will be described through specific embodiments.

[0040] Figure 1 Fig. 1 shows a schematic flowchart of the implementation of a secure voice interaction method provided by an embodiment of the present invention. This method can be applied to a terminal device. The terminal device can be a mobile phone, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, etc.

[0041] Specifically, the above-mentioned secure voice interaction method may include the following steps S101 to S103.

[0042] Step S101: Collect the user's voice information.

[0043] Among them, the AI earphone collects the user's voice signal in real time through a built-in microphone array (at least two microphones). The beamforming technology is used to enhance the user's voice directionally and suppress environmental noise interference.

[0044] Perform noise reduction processing on the collected original voice signal to filter out background noise (such as wind noise, environmental noise). Perform echo cancellation to eliminate the interference of the earphone's own playback sound on the microphone input. Adjust the gain of the voice signal to ensure stable volume and provide clear input for subsequent recognition.

[0045] Input the preprocessed voice signal into a locally deployed speech recognition model (such as an edge-side AI model). The model analyzes the voice features and extracts the voice instructions in text form. Determine whether it carries a voice instruction.

[0046] Step S102: If the voice information carries a voice instruction, determine whether the voice instruction is a remote control instruction.

[0047] Among them, if the model detects that the voice contains preset audio features (such as a wake-up word or an instruction keyword), it is determined as a valid voice instruction. If there is no valid instruction, the process terminates or returns to step S101.

[0048] Step S103: If so, determine whether the voice instruction carries sensitive words.

[0049] Among them, analyze the content of the voice instruction to determine whether it needs to be forwarded to a remote server through a mobile device (such as sending a message, cloud search, etc.).

[0050] If the instruction involves an external service (such as sending a message, accessing cloud data), it is marked as a remote control instruction.

[0051] If the instruction only needs to be executed locally (such as playing music, adjusting the volume), it is marked as a local control instruction.

[0052] Perform keyword matching on the text content of the remote control instruction (such as preset sensitive words like "password", "ID number", "bank account number", etc.).

[0053] Step S104, if the voice instruction carries sensitive words, encrypt the sensitive words in the voice instruction to obtain the target voice instruction.

[0054] Among them, if sensitive words are detected, use a local encryption algorithm (such as AES encryption) to encrypt the sensitive part to generate the target voice instruction.

[0055] For example, the original instruction: "Send a message: My password is 123456." The encrypted instruction: "Send a message: My password is [encrypted data]." Transmit the encrypted target voice instruction through a wireless communication module (such as Bluetooth) to the connected mobile device. The mobile device performs subsequent operations according to the instruction type (remote or local) (such as forwarding to the cloud and / or local processing).

[0056] If the cloud needs to perform semantic understanding or execute operations on the instruction content (such as sending a message, accessing a database), the encrypted sensitive words must be decrypted.

[0057] The cloud service needs to have decryption permissions (such as holding a key), or the decryption is completed by a trusted third-party service.

[0058] Key management must strictly follow security protocols (such as the key is only stored in the cloud security module) to ensure that the decryption process is not maliciously intercepted.

[0059] If the instruction only needs to be transparently transmitted or stored (such as directly forwarding the encrypted instruction to the recipient), the cloud does not need to decrypt.

[0060] Optionally, the user sends the encrypted message "The password is [encrypted data]" to a friend, and the friend's device decrypts it locally. In this scenario, the cloud only serves as a transmission channel and does not participate in data processing.

[0061] If the purpose of encryption is to prevent privacy leakage during transmission, the cloud needs to decrypt to execute the instruction (such as sending a message to a third-party platform).

[0062] The encryption operation is completed locally on the AI headset, and the key is managed by the headset or the user device. The cloud needs to apply for a temporary key through a secure interface for decryption. Even if the cloud is attacked, the attacker cannot directly obtain the key, and sensitive information is still protected.

[0063] If end-to-end encryption technology is adopted, the cloud cannot decrypt the data, and only the instruction receiver (such as the message receiver) can decrypt it.

[0064] Only encrypt the sensitive part, and the non-sensitive content can still be processed in plain text, reducing the data range for cloud decryption. The beneficial effects of the embodiments of the present invention compared with the prior art are as follows: By dynamically encrypting sensitive words, user privacy is protected in the remote control instruction, and the risk of privacy leakage is reduced. Distinguish and process remote control instructions, effectively avoiding over-encryption of local control instructions and reducing resource waste. Implement multi-level voice instruction processing, embed sensitive word detection and encryption throughout the process from voice collection to instruction execution, forming an end-to-end secure link, enabling sensitive information to be encrypted at the generation stage, and effectively avoiding the leakage of plaintext data. Significantly reduce the dependence on cloud security measures and improve the data security of voice interaction.

[0065] Traditional single microphone devices have poor voice collection quality in noisy environments, resulting in a high recognition error rate. Based on this, an optional embodiment of the present invention is proposed.

[0066] Refer to Figure 2 , Figure 2 It is a schematic diagram of a specific embodiment of step S102 of the secure voice interaction method in the embodiments of the present invention. Step S101 also includes the following specific implementation manners.

[0067] Step S1011, control the microphone array to collect the initial voice information of the user.

[0068] In the embodiment of the present invention, the AI headset is built-in with an array composed of at least two microphones, and the voice collection ability is enhanced through the collaborative work of multiple microphones.

[0069] Optionally, call the beamforming algorithm to control the microphone array to collect the initial voice information of the user. Specifically, call the beamforming algorithm to adjust the phase and gain of each microphone to form a directional beam and focus on the user's voice direction. Suppress environmental noise (such as wind noise and background human voices) and lateral interference, and improve the signal-to-noise ratio of the target voice.

[0070] Each microphone synchronously collects the original sound signal (i.e., the initial voice information), and retains the multi-channel audio data for subsequent processing.

[0071] Step S1012, perform audio processing on the initial voice information to obtain voice information.

[0072] In the embodiment of the present invention, digital signal processing algorithms (such as spectral subtraction and adaptive filtering) are used to filter out environmental noise. In a noisy street, vehicle noise can be filtered out and the user's voice can be retained.

[0073] Optionally, the initial sound information is denoised and echo-canceled to obtain sound information. Specifically, through the adaptive echo cancellation algorithm (AEC), the interference of the sound played by the headphone's own speaker to the microphone input is eliminated. This can effectively avoid the delayed echo of the user's own voice (such as the "self-excitation" phenomenon during a call).

[0074] Dynamically adjust the gain of the voice signal according to the ambient volume to ensure stable output volume. Automatically increase the volume when the user speaks in a low voice to avoid voice recognition failure. The signal after denoising, echo cancellation, and gain control is used as the sound information and input into the subsequent voice recognition module.

[0075] In the embodiment of the present invention, the user's voice is directionally enhanced and ambient noise is suppressed through a microphone array and beamforming technology. The voice signal is further purified through processing such as denoising and echo cancellation. This can improve the accuracy of voice recognition, especially in noisy scenarios.

[0076] Traditional technologies still have deficiencies in aspects such as the accuracy of voice recognition in noisy environments, user privacy protection, and the cooperation efficiency with mobile devices. Based on this, an alternative embodiment of the present invention is proposed.

[0077] Refer to Figure 3 , Figure 3 For a schematic diagram of a specific embodiment of step S102 of the secure interaction method for voice in the embodiment of the present invention, the following specific implementation manners are also included before step S102.

[0078] Step S201, call a pre-trained voice recognition model to process the sound information.

[0079] In the embodiment of the present invention, the AI headphone calls a pre-trained voice recognition model (such as an end-side neural network model), and this model is deployed in a local device (such as a headphone chip or a connected mobile device) without relying on the cloud.

[0080] Analyze the preprocessed sound information (the audio processing result from claim 2), extract voice features (such as Mel Frequency Cepstral Coefficients MFCC). Identify the voice content and determine whether it contains valid instructions.

[0081] The model performs real-time analysis on the input voice to detect whether there are preset audio features, such as: wake-up words. Instruction keywords (such as "send", "query", "play"). Voiceprint features (such as the voice pattern of a specific user). Determine whether the voice contains preset features through pattern matching or probability classification (such as support vector machines, deep learning classifiers).

[0082] Each microphone synchronously collects the original sound signal (i.e., the initial sound information), and retains the multi-channel audio data for subsequent processing.

[0083] Step S202: If the voice recognition model determines that the voice information carries a preset audio feature, it is determined that the voice information carries a voice command.

[0084] In an embodiment of the present invention, if the voice recognition model detects a preset audio feature (such as a wake-up word + command keyword), it is determined that the current voice information carries a valid voice command. For example, when the user says "Xiaoming, send a message to Zhang San", the model detects that "Xiaoming" is the wake-up word and "send a message" is the command keyword, and then triggers the subsequent process.

[0085] If no preset feature is detected (such as environmental noise or irrelevant conversation), the process terminates or returns to the voice collection step. After confirming the existence of the voice command, the subsequent steps are executed.

[0086] In an embodiment of the present invention, the voice features are accurately extracted through a local voice recognition model (such as a deep learning model), which can effectively distinguish valid commands from noise. The preset audio feature (such as a wake-up word) is used as a double verification, which can effectively improve the detection robustness.

[0087] In the traditional AI voice interaction process, the processing of privacy information is often lacking. For example, when the user uses a voice assistant to input a voice command and needs to call the cloud service, the voice assistant often directly converts the user's voice command into text and sends the privacy information to the cloud server for processing. This method has the technical problem of low data security. Based on this, the present invention proposes an alternative embodiment.

[0088] Step S101 further includes the following specific embodiments.

[0089] Step S1011: Send the target voice command to the connected target device through Bluetooth.

[0090] In an embodiment of the present invention, after encrypting the sensitive words locally, the encrypted sensitive part and the non-sensitive content are integrated into the target voice command.

[0091] For example, the original command: "Send a message: The password is 123456." The encrypted command: "Send a message: The password is [encrypted data]." The AI headset and the target device (such as a smart phone) establish a secure connection through the Bluetooth protocol.

[0092] Use a Bluetooth pairing key (such as a PIN code or a key exchange protocol) to ensure the encryption of the communication link.

[0093] Verify the legitimacy of the target device (such as filtering through a device whitelist).

[0094] Start the Bluetooth data transmission module and configure the transmission parameters (such as the packet size, transmission rate).

[0095] Encapsulate the target voice command into a data format supported by the Bluetooth protocol (such as a BLE data packet).

[0096] Send the encrypted command to the target device via the Bluetooth link.

[0097] After receiving the data, the target device returns an acknowledgment signal (ACK) to ensure that the command is delivered intact.

[0098] If the transmission fails, trigger a retransmission mechanism (such as up to 3 retries).

[0099] After receiving the encrypted command, the target device forwards the encrypted command to a remote server (such as a cloud service).

[0100] In the embodiment of the present invention, even if the Bluetooth link is monitored, the attacker can only obtain the encrypted command and cannot crack the sensitive information.

[0101] In the traditional AI voice interaction process, the processing of privacy information is often lacking. For example, when a user uses a voice assistant to input a voice command and needs to call a cloud service, the voice assistant often directly converts the user's voice command into text and sends the privacy information to the cloud server for processing. This method has the technical problem of low data security. Based on this, the present invention proposes an alternative embodiment.

[0102] Refer to Figure 4 , Figure 4 For a specific embodiment diagram of step S102 of the voice secure interaction method in the embodiment of the present invention, step S102 further includes the following specific implementation manners.

[0103] Step S1021, determine whether it is necessary to forward the voice command from the connected target device to a remote server.

[0104] In the embodiment of the present invention, after local voice recognition, parse the text content of the voice command to determine its intention and operation type. For example, "Send a message to Zhang San: The meeting time is 2 pm tomorrow." "Query the weather tomorrow." According to the command content, determine whether the operation depends on a remote server (such as a cloud AI service, a third-party API).

[0105] The determination logic, operations that require cloud services include but are not limited to sending messages (depending on a message server), querying the weather (depending on a weather API), and online searching (depending on a search engine). Operations that do not require cloud services include but are not limited to playing local music, adjusting the headphone volume, etc.

[0106] Step S1022, if so, determine that the voice command is a remote control command.

[0107] In an embodiment of the present invention, if an instruction requires cloud services, the target device (such as a smart phone) forwards the instruction to a remote server for execution.

[0108] Optionally, a list of locally preset service types is used to classify the instruction type through keyword matching or a natural language processing (NLP) model. If it needs to be forwarded to a remote server, the instruction is marked as a remote control instruction, triggering subsequent sensitive word detection and encryption processes.

[0109] Optionally, if not, the voice instruction is determined to be a local control instruction and is sent to the target device via Bluetooth.

[0110] In an embodiment of the present invention, by distinguishing between remote and local instructions and triggering sensitive word encryption only in remote instructions, privacy data can be protected before leaving the device.

[0111] As Figure 5 shown, it is a schematic diagram of a terminal device provided by an embodiment of the present invention. The terminal device 5 may include: a processor 501, a memory 502, and a computer program 503 stored in the memory 502 and executable on the processor 501, such as a secure interaction program for voice. When the processor 501 executes the computer program 503, the steps in the above-mentioned secure interaction embodiments for each voice are implemented.

[0112] The computer program may be divided into one or more modules / units. One or more modules / units are stored in the memory 502 and executed by the processor 501 to complete the present invention. One or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the terminal device.

[0113] The terminal device may include, but is not limited to, a processor 501 and a memory 502. Those skilled in the art can understand that Figure 5 merely examples of the terminal device do not constitute a limitation to the terminal device. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the terminal device may also include input / output devices, network access devices, buses, etc.

[0114] The so-called processor 501 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0115] The memory 502 may be an internal storage unit of the terminal device, such as the hard disk or memory of the terminal device. The memory 502 may also be an external storage device of the terminal device, such as a plug-in hard disk equipped on the terminal device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 502 may also include both the internal storage unit and the external storage device of the terminal device. The memory 502 is used to store computer programs and other programs and data required by the terminal device. The memory 502 may also be used to temporarily store data that has been output or is to be output.

[0116] It should be noted that for the convenience and brevity of description, the structure of the above terminal device may also refer to the specific description of the structure in the method embodiments, which will not be elaborated here.

[0117] The embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned secure interaction method of voice can be implemented.

[0118] The embodiment of the present invention provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can be made to execute the steps in the above-mentioned secure interaction method of voice.

[0119] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not elaborated or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0120] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0121] In the embodiments provided by the present invention, it should be understood that the disclosed terminal devices and methods can be implemented in other ways. For example, the terminal device embodiments described above are merely illustrative. Another point is that the couplings or direct couplings or communication connections shown or discussed among each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0122] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0123] In addition, the functional units in each embodiment of the present invention can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0124] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0125] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A secure voice interaction method, characterized in that: include: Collect user's voice information; If the sound information carries a voice command, determining whether the voice command is a remote control command; If yes, determining whether the voice command contains sensitive words; If the voice instruction carries the sensitive word, the sensitive word in the voice instruction is encrypted to obtain a target voice instruction.

2. The secure voice interaction method according to claim 1, characterized in that: The step of collecting the user's voice information includes: Control the microphone array to collect the user's initial voice information; Audio processing is performed on the initial sound information to obtain the sound information.

3. The secure voice interaction method according to claim 1, characterized in that: If the sound information carries a voice command, before the step of determining whether the voice command is a remote control command, the method further includes: Calling a pre-trained speech recognition model to process the sound information; If the speech recognition model determines that the sound information carries the preset audio feature, then it is determined that the sound information carries the voice command.

4. The secure voice interaction method according to claim 1, characterized in that: The step of controlling the microphone array to collect the user's initial voice information also includes A beamforming algorithm is called to control the microphone array to collect the initial sound information of the user.

5. The secure voice interaction method according to claim 1, characterized in that: The step of performing audio processing on the initial sound information to obtain the sound information comprises: The initial sound information is subjected to noise reduction and echo cancellation processing to obtain the sound information.

6. The secure voice interaction method according to claim 1, characterized in that: If the voice instruction carries the sensitive words, the sensitive words in the voice instruction are encrypted to obtain the target voice instruction. The method further includes: The target voice command is sent to the connected target device via Bluetooth.

7. The secure voice interaction method according to claim 1, characterized in that: The step of determining whether the voice command is a remote control command comprises: Determining whether the voice command needs to be forwarded from the connected target device to the remote server; If so, it is determined that the voice command is a remote control command.

8. The secure voice interaction method according to claim 1, characterized in that: After the step of determining whether the voice command is a remote control command if the sound information carries a voice command, the method further includes: If not, it is determined that the voice command is a local control command, and the voice command is sent to the target device via Bluetooth.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the secure voice interaction method as claimed in any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the secure voice interaction method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Test data preservation method and device based on large model, storage medium and equipment

    CN120743796A

  • Test data preservation method and device based on large model, storage medium and equipment

    CN120743796B