Voice interaction method and corresponding device, server, storage medium

By encrypting the voice output device and forwarding it through the server, the problem of user privacy leakage in VoIP software is solved, and confidential calls and conferences for multiple people are achieved.

CN116192423BActive Publication Date: 2026-02-10ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211518331.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2026-02-10
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

The encryption features of existing VoIP software cannot effectively prevent operating systems and malware from accessing voice call content, resulting in a high risk of user privacy leaks and the possibility of servers eavesdropping on call content.

Method used

A voice interaction device and method are provided, in which a voice output device encrypts a voice signal in the audio domain, and a server forwards it to a voice receiving device for decryption, ensuring that only users with specified permission levels can decrypt and obtain the original voice content.

Benefits of technology

It effectively protects user call privacy, prevents operating systems and malware from obtaining raw voice content, and improves the security and flexibility of voice interaction, making it suitable for multi-person calls and conference scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116192423B_ABST
    Figure CN116192423B_ABST
Patent Text Reader

Abstract

The application provides a voice interaction method and corresponding device, server and storage medium. The voice interaction method comprises: entering a voice encryption state in response to an encryption starting instruction before or during multi-person voice interaction; in the voice encryption state, encrypting an obtained voice signal based on a key to generate an encrypted voice signal; and transmitting the encrypted voice signal to a terminal device or a server connected with a voice output device. According to the technical solution of the application, the voice content during multi-person voice interaction can be effectively prevented from being leaked, and the security of multi-person voice interaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of voice interaction, and in particular to a voice interaction method and a corresponding device, server and storage medium. BACKGROUND

[0002] With the popularity of Dingding and other cloud video conference software, voice calls, conferences and the like among multiple users are more convenient, and at the same time, the possibility of leakage of user voice privacy is greatly increased, for example, the operating system and some malicious software can obtain the content of voice calls. The related technology provides an encryption function in the VoIP (Voice over Internet Protocol, a voice call technology) software, but the encryption function is data level, transmission layer encryption, which cannot resist the operating system layer and malicious software to obtain the microphone input. SUMMARY

[0003] The embodiments of the present application provide a voice interaction method and a corresponding device, server and storage medium to solve the technical problems existing in the prior art.

[0004] In a first aspect, the embodiments of the present application provide a voice output device, comprising:

[0005] A first control switch is located on the outside of the voice output device and is used to generate an encryption start instruction when triggered;

[0006] An encryption control module is electrically connected with the first control switch and is used to enter a voice encryption state in response to the encryption start instruction before or during multi-person voice interaction, and encrypt the obtained voice signal based on a key in the voice encryption state to generate an encrypted voice signal;

[0007] A first communication module is electrically connected with the encryption control module and is used to transmit the encrypted voice signal to a terminal device or a server connected with the voice output device; the terminal device is used to forward the encrypted voice signal to the server, and the server is used to forward the encrypted voice signal to a voice receiving device of an interaction object.

[0008] In a second aspect, the embodiments of the present application provide a server, comprising:

[0009] A second communication module is used to receive the encrypted voice signal transmitted by the voice output device or the terminal device provided in the first aspect of the embodiments of the present application, and forward the encrypted voice signal to a voice receiving device of an interaction object.

[0010] In a third aspect, the embodiments of the present application provide a voice receiving device, comprising:

[0011] The third communication module is configured to receive the encrypted voice signal sent by the server according to the second aspect of the present application.

[0012] The decryption control module is configured to decrypt the encrypted voice signal based on the key.

[0013] In the fourth aspect, the present application provides a voice interaction method, which can be applied to a voice output device, and the method comprises the following steps:

[0014] Before or during the multi-person voice interaction, the voice encryption state is entered in response to an encryption starting instruction.

[0015] In the voice encryption state, the obtained voice signal is encrypted based on the key to generate an encrypted voice signal.

[0016] The encrypted voice signal is transmitted to a terminal device or a server connected to the voice output device, the terminal device is configured to forward the encrypted voice signal to the server, and the server is configured to forward the encrypted voice signal to a voice receiving device of an interaction object.

[0017] In the fifth aspect, the present application provides a voice interaction method, which can be applied to a server, and the method comprises the following steps:

[0018] The encrypted voice signal transmitted by the voice output device or the terminal device according to the first aspect of the present application is received.

[0019] The encrypted voice signal is forwarded to a voice receiving device of an interaction object.

[0020] In the sixth aspect, the present application provides a voice interaction method, which can be applied to a voice receiving device, and the method comprises the following steps:

[0021] The encrypted voice signal sent by the server according to the second aspect of the present application is received.

[0022] The encrypted voice signal is decrypted based on the key, and the key is the key used in the voice interaction method according to the fourth aspect of the present application.

[0023] In the seventh aspect, the present application provides a computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the voice interaction method provided in any of the embodiments of the present application.

[0024] Compared with the prior art, the present application has the following advantages:

[0025] According to the technical solutions provided in the embodiments of the present application, voice interaction can be realized through a voice output device, a server and a voice receiving device, voice content in multi-person voice interaction can be encrypted in an audio domain, and an operating system, malicious software and the like can only obtain encrypted voice data and cannot decrypt and thus cannot obtain original voice content, thereby effectively protecting user call privacy and improving the security of voice interaction in a secret call, a secret meeting and the like; the control operation of whether to enter a voice encryption state can be performed before multi-person voice interaction or in the process of multi-person voice interaction, and the control is relatively flexible; and the encryption operation can be realized in the form of a combination of software and hardware.

[0026] The above description is only a summary of the technical solutions of the present application, in order to more clearly understand the technical means of the present application, the technical solutions can be implemented according to the content of the description, and in order to make the above and other purposes, characteristics and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application are described below. BRIEF DESCRIPTION OF DRAWINGS

[0027] In the drawings, like reference numerals refer to same or similar functionalities throughout the several views. The drawings are not necessarily to scale. It is to be understood that the drawings only depict several embodiments in accordance with the present application and should not be considered as limiting the scope of the present application.

[0028] Figure 1 An application scenario schematic diagram of the voice interaction scheme provided in the embodiments of the present application;

[0029] Figure 2 Another application scenario schematic diagram of the voice interaction scheme provided in the embodiments of the present application;

[0030] Figure 3 A structural framework schematic diagram of a voice output device provided in the embodiments of the present application;

[0031] Figure 4 A structural framework schematic diagram of a server provided in the embodiments of the present application;

[0032] Figure 5 A structural framework schematic diagram of a voice receiving device provided in the embodiments of the present application;

[0033] Figure 6 An interaction schematic diagram between devices in the embodiments of the present application;

[0034] Figure 7 A flow schematic diagram of a voice interaction method provided in the embodiments of the present application;

[0035] Figure 8 A principle schematic diagram of spectrum inversion in the embodiments of the present application;

[0036] Figure 9 Another principle diagram for spectrum inversion in the embodiment of the present application;

[0037] Figure 10 A principle diagram for updating the spectrum cut-off point in the embodiment of the present application;

[0038] Figure 11 A flow diagram of another voice interaction method provided by the embodiment of the present application;

[0039] Figure 12 A flow diagram of another voice interaction method provided by the embodiment of the present application; and

[0040] Figure 13 A flow diagram of another voice interaction method provided by the embodiment of the present application. DETAILED DESCRIPTION

[0041] In the following, only certain exemplary embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the concept or scope of the present application. Therefore, the drawings and the description are considered to be exemplary in nature, rather than limiting.

[0042] With the popularity of VoIP software such as DingDing, voice calls, conferences and other interactions among multiple users are more convenient, and on this basis, more interaction requirements are derived, such as encrypting call content to prevent eavesdropping, some users need to occasionally communicate alone during the process of multiple user calls to communicate more confidential content, and need to be able to realize both separate communication and not affect the overall communication function to improve call efficiency and conference efficiency. Based on the above requirements, the related technology proposes some solutions, but these solutions usually encrypt the call content at the data level and the transmission layer, which cannot resist the acquisition of the microphone input at the operating system level and malicious software, and users cannot perceive the improvement of privacy protection, and the server side of the VoIP software may still have a decryption key, so that it can eavesdrop in transit. In addition, the call content may also be recognized as text by the ASR (Automatic Speech Recognition, automatic speech recognition) system on the server side, so that the keyword is eavesdropped and then the advertisement is pushed.

[0043] Based on the above status quo, the embodiments of the present application provide a voice interaction scheme, which includes a voice interaction device, a server and a voice interaction method executable on the voice interaction device. The voice interaction device can be a voice output device for outputting voice or a voice receiving device for receiving voice, the voice output device can encrypt an input voice signal and output an encrypted voice signal, the server can forward the encrypted voice signal output by the voice output device to the voice receiving device, and the voice receiving device can receive the encrypted voice signal and decrypt it to restore the original voice signal. The voice output device and the voice receiving device are used in cooperation, and encrypted communication can be realized to protect the privacy of users in the communication process.

[0044] Figure 1 An application scenario of the voice interaction scheme is shown. Referring to Figure 1 The voice interaction scheme provided by the embodiments of the present application can be applied to a scenario of multi-person communication interaction, such as online conference of multiple persons, live broadcast, etc. In the scenario, each user can use a voice interaction device (a voice output device or a voice receiving device) to connect with a terminal device, and install VoIP software on the terminal device, so that voice connection between multiple users can be realized through VoIP. In the communication process, a voice output side, a voice transfer side and a voice receiving side are involved. On the voice output side, a speaker can speak into the voice output device, so as to input a voice signal to the voice output device. The voice signal is encrypted by the voice output device and can be sent to the server of the voice transfer side through the VoIP software installed on the terminal device A. The server can forward the encrypted voice signal to the terminal device B of the voice receiving side. On the voice receiving side, the VoIP software installed on the terminal device B can send the received encrypted voice signal to the voice receiving device through the terminal device B. The voice receiving device can decrypt the encrypted voice signal and play the decrypted voice signal to a listener on the voice receiving side.

[0045] In Figure 1 In the application scenario shown, the voice output device can be a microphone (also referred to as a microphone or a microphone) or a headset, a headset conversion head, a smart speaker, etc. The voice output device can also be a mobile phone, a computer, a smart watch, etc. The voice receiving device can be a headset without a sound transmission function or a headset with a sound transmission function, or can be a headset conversion head, a smart speaker, etc. The terminal device can be a mobile phone, a computer, a smart watch, etc.

[0046] Figure 2 Another application scenario for realizing the voice interaction scheme is shown. Referring to Figure 2The voice interaction scheme provided in the embodiments of the present application can be applied to a scenario of multi-person call interaction, for example, an online conference of multiple persons, in which part of the users can connect with a terminal device using a voice output device and install VoIP software on the terminal device, and another part of the users can install VoIP software on a voice receiving device, and voice connection between multiple users can be achieved through VoIP. In the process of the call, a voice output side, a voice transfer side and a voice receiving side are involved. On the voice output side, a speaker (for example, a user using a voice output device) can speak into the voice output device, so as to input a voice signal to the voice output device. The voice signal can be sent to a server of the voice transfer side through the VoIP software installed on the terminal device after being encrypted by the voice output device. The server can forward the encrypted voice signal to a voice receiving device of the voice receiving side. On the voice receiving side, the voice receiving device can decrypt the encrypted voice signal received by the VoIP software and play the decrypted voice signal to a listener of the voice receiving side.

[0047] In Figure 2 The voice output device can be any one of a microphone, a headset with a microphone function, a headset adapter, a smart speaker, or any one of a mobile phone, a computer, a smart watch, and the like. The voice receiving device can be any one of a mobile phone, a computer, a smart watch, and the like.

[0048] In Figure 1 In the application scenario shown in FIG. 1, if the voice interaction device used by each user simultaneously integrates voice output and voice receiving functions, for each user, the user can act as a speaker and a listener, and the voice interaction device of the user can implement encryption and decryption functions. Figure 2 In the scenario shown in FIG. 1, for a user using a voice output device, if the voice output device simultaneously integrates voice receiving functions, the user can act as a speaker and a listener, and the voice output device can implement encryption and decryption functions.

[0049] In Figure 1 and Figure 2 In the application scenarios shown in FIG. 1 and FIG. 2, the speaker can select the listener, and at least part of the listeners are given the permission to listen to the content of the speech of the speaker, so that the encrypted voice signal can be output only to the voice receiving device used by the selected listener.

[0050] The voice interaction scheme provided in the embodiments of the present application can be applied to various actual scenarios, for example, online chat of multiple persons in a life scenario, live broadcast in a small range, online conference, online teaching, internal live broadcast in an office scenario, and the like. The voice interaction scheme provided in the embodiments of the present application can be used when it is necessary to protect voice content from being leaked and improve call efficiency and conference efficiency.

[0051] For the convenience of understanding the technical solutions of the embodiments of the present application, the related technologies of the embodiments of the present application are described as follows. The following related technologies can be combined with the technical solutions of the embodiments of the present application in any way, and all of them belong to the protection scope of the embodiments of the present application.

[0052] The embodiments of the present application provide a voice output device, as shown in the following figure. Figure 3 The device can include a first control switch 301, an encryption control module 302, and a first communication module 303.

[0053] The first control switch 301 can be located on the outside of the voice output device and can be used to generate an encryption start instruction when triggered. The encryption control module 302 can be electrically connected with the first control switch 301 and used to enter a voice encryption state in response to the encryption start instruction before or during multi-person voice interaction, and to encrypt the obtained voice signal based on a key in the voice encryption state to generate an encrypted voice signal. The first communication module 303 can be electrically connected with the encryption control module 302 and can be used to transmit the encrypted voice signal to a terminal device or a server connected with the voice output device; the terminal device is used to forward the encrypted voice signal to the server, and the server is used to forward the encrypted voice signal to a voice receiving device of an interaction object.

[0054] In one example, when the voice output device is a microphone, a headset, a headset adapter, a smart speaker, or the like, it needs to be connected with a mobile phone, a tablet computer, or the like for use during voice interaction. At this time, the voice output device can send the encrypted voice signal to the connected terminal device, the terminal device forwards the encrypted voice signal to the server, and then the server forwards the encrypted voice signal to the voice receiving device. In another example, when the voice output device is a mobile phone, a computer, a smart watch, or the like, it can directly send the encrypted voice signal to the server.

[0055] The first control switch 301 can be any one of a key switch, a toggle switch, a twist switch, a touch switch, etc. When the first control switch 301 is a key switch, an encryption start instruction can be generated by pressing the key switch, so that the voice output device enters a voice encryption state, and an encryption end instruction can be generated by pressing the key switch again, so that the voice output device ends the voice encryption state. When the first control switch 302 is a toggle switch, an encryption start instruction can be generated by toggling the toggle switch, so that the voice output device enters a voice encryption state, and an encryption end instruction can be generated by toggling the toggle switch again or toggling the toggle switch in another direction, so that the voice output device ends the voice encryption state. When the first control switch 302 is a twist switch, an encryption start instruction can be generated by twisting the twist switch, so that the voice output device enters a voice encryption state, and an encryption end instruction can be generated by twisting the twist switch again or twisting the twist switch in another direction, so that the voice output device ends the voice encryption state. When the first control switch 302 is a touch switch, an encryption start instruction can be generated by touching the touch switch, so that the voice output device enters a voice encryption state, and an encryption end instruction can be generated by touching the touch switch again or toggling the toggle switch in another direction, so that the voice output device ends the voice encryption state.

[0056] The type of the first control switch 302 can not be limited to the above-mentioned types, and can also be other types. For the above-mentioned types of switches, the specific operation for generating the corresponding instruction can not be limited to the above-mentioned several ways, and can also be other ways. The encryption control module 302 can be a processor.

[0057] The voice output device provided by the embodiment of the present application can encrypt the voice content in the multi-person voice interaction in the audio domain. The operating system, malicious software, and ASR system can only obtain the encrypted voice signal and cannot decrypt and obtain the original voice content, thereby effectively protecting the user's call privacy, preventing the operating system and malicious software from obtaining the call content, and realizing secure calls, secure meetings, etc. The control operation of whether to enter the voice encryption state can be performed before the multi-person voice interaction or during the multi-person voice interaction, and the control flexibility is strong. The encryption operation can be realized in the form of software and hardware combination.

[0058] In an implementation, the first control switch 301 can have multiple gears, each of which can generate a level of encryption start instruction when triggered, and each level of encryption start instruction is associated with a level of voice encryption state. Correspondingly, the encryption control module 302 can also be configured to enter a level of voice encryption state associated with a level of encryption start instruction received when the level of encryption start instruction is received, and generate voice transmission rule information in the voice encryption state. The first communication module 303 can also be configured to transmit the voice transmission rule information to the terminal device or the server, and the voice transmission rule information can include a specified permission level, which can be a permission level associated with the current voice encryption state.

[0059] The user can switch between the gears in the first control switch 301 according to needs, so as to switch between different levels of voice encryption state. On the outside of the voice output device 300, the gears can be arranged in order of the levels of encryption start instructions that can be generated (from low to high or from high to low), so as to facilitate the user to switch the gears in turn, and also facilitate the user to quickly locate a certain gear.

[0060] The multiple gears in the first control switch 301 can be associated with multiple permission levels by generating different levels of encryption start instructions and entering different voice encryption states. By selecting and triggering the gears, the user can select the interactive object of the corresponding permission level as the object of receiving the encrypted voice signal, which can meet the user's different encryption needs.

[0061] In an implementation, the voice output device described above can further include an output module, which can be configured to output a list of interactive objects of multiple voice interactions for the current voice encryption state, and generate a selection instruction when a selection operation is performed. Correspondingly, the encryption control module 302 can also be configured to determine the selected interactive object as the interactive object of the specified permission level in response to the selection instruction, and generate the voice transmission rule information based on the interactive object of the specified permission level.

[0062] The output module can be a display screen, through which the list of interactive objects can be output. The user can select the interactive object by touching the specified area of the display screen to generate the selection instruction, so that the user can select the interactive object of the specified permission level by himself / herself.

[0063] In an implementation, the voice output device described above can further include a second control switch, which can be configured to generate a key update instruction when triggered. Correspondingly, the encryption control module 302 can be configured to update the key in response to the key update instruction.

[0064] The second control switch can be any type of switch, such as a push-button switch, toggle switch, rotary switch, or touch switch. When the second control switch is a push-button switch, a key update command can be generated by pressing the push-button switch. When the first control switch is a toggle switch, a key update command can be generated by toggling the toggle switch. When the first control switch is a rotary switch, a key update command can be generated by touching the rotary switch.

[0065] In one example, the second control switch can be integrated with the first control switch 301 as a single switch. For instance, the first control switch 301 may include multiple gear positions, one of which, when triggered, can generate an update command.

[0066] The aforementioned voice output device can be any device such as a microphone, earphone, headphone adapter, or smart speaker. When used with terminal devices such as mobile phones, computers, and smartwatches, the voice output device may also include an audio transmission interface for transmitting voice signals to the terminal device, such as a TRS (Tailor Resonance) interface, an XLR (XLR) interface, or an HDMI (High Definition Multimedia Interface) interface. It can be plugged into the corresponding interface of the terminal device for plug-and-play use, or it can be wirelessly connected to the terminal device. This wired or wireless connection method makes it easier for users to perceive privacy protection. For example, users can clearly perceive that when the voice output device is plugged into the terminal device or wirelessly connected to the terminal device, the subsequent speech content is protected by the voice output device. When the voice output device is unplugged from the terminal device or disconnected from the terminal device, the subsequent speech content is no longer protected by the voice output device.

[0067] The aforementioned voice output device can independently encrypt the voice signal, separate from the terminal device. The voice signal is encrypted when it is transmitted from the voice output device to the terminal device, and the operating system and malicious software on the terminal device cannot decrypt it, thus making it impossible to obtain the voice content.

[0068] Based on the same technical concept, embodiments of this application also provide a server, such as... Figure 4 As shown, the server may include a second communication module 401, which can be used to receive encrypted voice signals transmitted by a voice output device or a terminal device, and forward the encrypted voice signals to the voice receiving device of the interactive object. The voice output device can be any type of voice output device provided in this application embodiment, and the encrypted voice signal transmitted by the terminal device can be provided by any type of voice output device provided in this application embodiment.

[0069] In an embodiment, the second communication module 401 can also be configured to receive voice transmission rule information transmitted by the voice output device or the terminal device. Correspondingly, when forwarding the encrypted voice signal to the voice receiving device of the interaction object, the second communication module 401 can be configured to transmit the encrypted voice signal to the voice receiving device of the interaction object corresponding to the specified permission level in the voice transmission rule information, or transmit the encrypted voice signal to the voice receiving device of each interaction object corresponding to the permission level and send the key to the voice receiving device of the interaction object corresponding to the specified permission level.

[0070] Interaction objects of different permission levels can have a containing relationship. For example, interaction objects of a second permission level can contain interaction objects of a first permission level, and have more interaction objects. The interaction objects of the first permission level can be interaction objects of a higher permission, which can listen to the content that the interaction objects of the second permission level have permission to listen to, and also can listen to the content that the interaction objects of the second permission level do not have permission to listen to. For example, the interaction objects of a low permission level include interaction objects of a high permission level. In an example, for an online meeting of an enterprise, all employees of the enterprise have permission to listen to part of the meeting content, which are interaction objects of a second permission level, and the senior management of the enterprise have permission to listen to all the meeting content (including highly confidential content), which are interaction objects of a first permission level.

[0071] Based on the above manner, the server can make the specified interaction object obtain the voice content in two ways. One way is to transmit the encrypted voice signal only to the voice receiving device of the interaction object corresponding to the specified permission level (for example, a high permission level) in the voice transmission rule information, and not to the voice receiving device of the interaction object of other permission levels (for example, a low permission level). That is, the interaction objects of other permission levels cannot obtain the encrypted voice signal, and thus cannot decrypt it, so that it can be prevented that the voice content delivered to part of the interaction objects is leaked to other interaction objects. Another way is to transmit the encrypted voice signal to the voice receiving device of all interaction objects without distinction, but send the key only to the voice receiving device of the interaction object corresponding to the specified permission level (for example, a high permission level). That is, only the voice receiving device of the interaction object corresponding to the specified permission level can decrypt the received encrypted voice signal to obtain the correct voice content. The voice receiving device of the interaction object corresponding to other permission levels cannot decrypt the received encrypted voice signal to obtain the correct voice content.

[0072] In an embodiment, the second communication module 401 is further configured to acquire an interactive object list uploaded by the voice output device or the terminal device, the interactive object list including a plurality of interactive objects and organizational structure information of the plurality of interactive objects, such as employee information of an enterprise and department information to which the employees belong.

[0073] Correspondingly, the server can further include a processing module electrically connected to the second communication module 401, the processing module being configured to determine a permission level of each interactive object in the interactive object list according to the organizational structure information of each interactive object in the interactive object list. In an example, the permission level of each employee can be determined according to the employee information of the enterprise and the department information to which the employees belong. For example, if employee A belongs to the finance department, employee A has the permission to listen to the financial data; if employee A belongs to the board of directors or other senior management department, employee A has the permission to listen to the account information and other information, and employees in the same department have the same permission level.

[0074] Based on the organizational structure information of the interactive objects, the permission level can be quickly determined, and when the encrypted voice signal and the voice transmission rule information are received, the correct voice content can be delivered to the interactive objects with the specified permission level.

[0075] The server provided by the embodiments of the present application can be a server of a voice output device or a voice receiving device, a server of a VoIP software for implementing multi-person voice interaction, or other servers with certain security, and the processing module can be any one of a plurality of types of processors.

[0076] Based on the same technical concept, the embodiments of the present application further provide a voice receiving device, as shown in the accompanying drawings, which can include a third communication module 501 and a decryption control module 502. The third communication module 501 is configured to receive an encrypted voice signal sent by a server, and the server can be any one of the servers provided by the embodiments of the present application. The decryption control module 502 is configured to decrypt the encrypted voice signal based on a key, and the correct voice content can be obtained after decryption, thereby completing effective voice interaction. Figure 5

[0077] In an embodiment, the voice receiving device can further include a third control switch arranged on the outside of the voice receiving device, the third control switch being configured to generate a decryption start instruction when a start operation is performed, and the received encrypted voice signal can be decrypted. The third control switch can be any one of a key switch, a toggle switch, a twist switch, a touch switch, and the like.

[0078] The voice receiving device can be any one of a microphone, a headset, a headset adapter, a smart speaker, a mobile phone, a computer, a smart watch, and the like.​

[0079] In the case that the voice receiving device is a microphone, a headset, a headset adapter or a smart speaker used in cooperation with a terminal device, the decryption can be directly based on the key, which is a hardware decryption method. The voice receiving device can further include an audio transmission interface, such as a TRS interface, an XLR interface, an HDMI interface, etc., for transmitting voice signals with the terminal device. The audio transmission interface can be inserted into the corresponding interface of the terminal device for plug-and-play with the terminal device, or can be wirelessly connected to the terminal device. This wired or wireless connection method is more convenient for users to perceive privacy protection. For example, the user can obviously perceive that the encrypted voice signal received can be decrypted for normal listening when the voice receiving device is inserted into or wirelessly connected to the terminal device, and the encrypted voice signal received cannot be decrypted for normal listening when the voice receiving device is pulled out of or disconnected from the terminal device.

[0080] In the case that the voice receiving device is a mobile phone, a computer, a smart watch or any other device, the voice receiving and voice decryption functions can be integrated into the terminal device. The voice receiving device can pre-receive a plug-in for decryption and install it on the device, so that the encrypted voice signal received by the VoIP software can be decrypted in subsequent voice interaction, which is a software decryption method.

[0081] In an exemplary scenario, the voice output device in the embodiments of the present application can integrate the functions of voice receiving and voice decryption. For example, the voice output device can integrate the components and functions of the voice receiving device. The voice receiving device can integrate the functions of voice output and voice encryption. For example, the voice receiving device can integrate the components and functions of the voice output device. That is, the same device can integrate the functions of voice output, voice encryption, voice receiving and voice decryption. In this way, through the same device, the user can output encrypted voice and decrypt received encrypted voice, and can freely switch between the roles of speaker and listener.

[0082] The devices (voice output device, server and voice receiving device) in the embodiments of the present application can further include a memory for storing a computer program that can be executed to implement voice interaction.

[0083] In the embodiments of the present application, the components in each device (the voice output device, the server and the voice receiving device) can be connected to each other through a bus and complete data interaction between each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 and Figure 5 Only one thick line is used to represent the bus in the figure, but it does not mean that there is only one bus or only one type of bus.

[0084] Optionally, in a specific implementation, if the components in each device are integrated on a chip, the components can complete communication with each other through an internal interface.

[0085] Figure 6 The interaction schematic diagram between the devices is shown, and the details are described in the device embodiments. Figure 6 The mobile phone (terminal device) can be provided with a key, the mobile phone can upload the key to the VoIP server, the VoIP server can obtain a list of authorized listeners (i.e. listeners) (for example, selected by a speaker), and distribute the key to the devices of the authorized listeners, for example, distributing the key at t1 to the key at t1, and distributing the key at t2 to the key at t2. The voice receiving device of the listener receiving the key can listen to the voice content through hardware decryption or software decryption. The specific working principles of the devices provided in the embodiments of the present application will be further introduced in the subsequent method embodiments.

[0086] Based on the same technical concept, referring to Figure 7 The embodiments of the present application also provide a voice interaction method 700, which can be applied to a voice output device. The voice interaction method 700 can include the following steps S701-S703:

[0087] S701, before or during the multi-person voice interaction, in response to an encryption start instruction, entering a voice encryption state.

[0088] The multi-person voice interaction can be an online conference, live broadcast, etc. of multiple people. Before the multi-person voice interaction, it can be a time before starting the multi-person online conference, live broadcast, etc. For example, 1 minute, 5 minutes or other times before the starting time of the multi-person online conference, live broadcast, etc. The encryption start instruction can be generated based on the opening operation of the first control switch of the voice output device by the user. The first control switch of the voice receiving device and the opening operation can refer to the description in the previous device embodiments, which will not be repeated here.

[0089] S702, in the voice encryption state, the obtained voice signal is encrypted based on the key to generate an encrypted voice signal.

[0090] S703, transmitting the encrypted voice signal to a terminal device or a server connected with the voice output device.

[0091] The terminal device can be configured to forward the encrypted voice signal to the server, and the server can be configured to forward the encrypted voice signal to a voice receiving device of the interactive object. The interactive object can be a user participating in the voice receiving side of the multi-person voice interaction, such as the listener shown in FIGS. 1 and 2. Figure 1 and Figure 2 When outputting the encrypted voice signal to the voice receiving device of the interactive object of the multi-person voice interaction, the encrypted voice signal can be output to all or part of the interactive objects participating in the multi-person voice interaction.

[0092] The voice interaction method provided by the embodiments of the present application can encrypt the voice content in the multi-person voice interaction in the audio domain. The operating system, malicious software, and ASR system can only obtain the encrypted voice signal and cannot decrypt it to obtain the original voice content, thereby effectively protecting the user's call privacy and preventing the operating system and malicious software from obtaining the call content, and realizing secure calls and secure meetings. The control operation of whether to enter the voice encryption state can be performed before the multi-person voice interaction or during the multi-person voice interaction, and the control flexibility is strong.

[0093] In an implementation, the encryption start instruction includes a plurality of levels of encryption start instructions. Correspondingly, in the step S701, in response to the encryption start instruction entering the voice encryption state, it can include: in response to receiving one level of encryption start instruction, entering one level of voice encryption state associated with the encryption start instruction of the level;

[0094] Generating voice transmission rule information in the voice encryption state; the voice transmission rule information includes a specified permission level, and the specified permission level is a permission level associated with the current voice encryption state.

[0095] In an implementation, the voice interaction method 700 can further include: outputting a list of interactive objects corresponding to the multi-person voice interaction for the current voice encryption state; and in response to a selection instruction for an interactive object in the list of interactive objects, determining the selected interactive object as an interactive object of the specified permission level. Correspondingly, the voice transmission rule information can be generated in the voice encryption state based on the interactive object of the specified permission level. This way, the user can specify the interactive object currently having the listening permission to meet the actual interaction needs of the user.

[0096] The hardware information of all the interactive objects corresponding to the multi-person voice interaction can be included in the interactive object list, for example, at least one of the device identification of the voice receiving device of the interactive object, the information of the hardware configuration, and the like, and the user information of the interactive object, for example, at least one of the account, the nickname, the avatar, and the like of the user.

[0097] The selection instruction can be generated after the user selects the interactive object in the interactive object list. According to the actual selection of the user, the selected interactive object can include all or part of the interactive objects in the interactive object list.

[0098] According to the actual needs of the user, the operation of outputting the interactive object list corresponding to the multi-person voice interaction for the user to select can be executed once or multiple times to meet the encryption needs for different objects. Each time the operation is executed, the interactive object list corresponding to the multi-person voice interaction can be output in response to the triggering operation of the user on the operation list, so as to assign the listening permission to different interactive objects. For example, in an online conference, the speaker can select which users need to listen to the voice content. By selecting to assign a higher permission to the part of the users as the specified permission level, the encrypted voice signal is sent to the part of the users.

[0099] In an embodiment, the voice interaction method 700 can further include: in the current voice encryption state, switching to another level of voice encryption state in response to a state switching instruction. The state switching instruction can be generated when the gear of the first control switch of the voice output device is switched.

[0100] In different levels of voice encryption states, voice encryption can be performed for interactive objects of different permission levels. When different voice contents need to be sent to interactive objects of different permission levels, there is no need to separately initiate a conference. The voice encryption state can be switched in the same conference. In an example, if only a small part of the content in the same conference needs to be sent to part of the participants, the voice encryption state can be switched to achieve the purpose. This can avoid the situation that part of the participants cannot timely switch back to the original conference after separately initiating a conference. The whole conference process can be smoother, and the conference efficiency is higher.

[0101] The state switching instruction can be generated in any of the following manners: manner one, the state switching instruction is generated after the user switches the gear of the first control switch on the voice output device; manner two, the state switching instruction is generated after the voice content is recognized, for example, the state switching instruction can be generated after the keyword of the voice content that can lead to a higher security level is recognized, for example, “service data” and “only” in “the following service data is only for the colleagues of this project” can be used as the keyword to trigger the generation of the state switching instruction; manner three, the state switching instruction is generated at a specified time point. For some multi-person online conferences whose agenda is clear, that is, it is clear what content is spoken at which time, the starting time point at which the voice encryption state needs to be switched can be set in advance.

[0102] In an implementation manner, in the voice encryption state in the step S702, the encrypted voice signal is generated based on the key. In different voice encryption states, the encrypted voice signal can be generated based on different keys. A key can be assigned to each voice encryption state in advance. The assignment operation can be performed before the multi-person voice interaction to avoid affecting the normal progress of the multi-person voice interaction and improve the overall efficiency of the multi-person voice interaction. The assignment operation can also be performed during the multi-person voice interaction to obtain the real-time demand of the user and perform the association accordingly.

[0103] In different voice encryption states, different keys are used for encryption, which can improve the security of encryption and the confidentiality of voice interaction.

[0104] In an implementation manner, the voice interaction method 700 can further include displaying a voice encryption option on a starting interface or a real-time interaction interface of the multi-person voice interaction, and generating an encryption starting instruction in response to a triggering operation on the voice encryption option.

[0105] The starting interface can be an interface to be entered into the voice interaction. The voice encryption option can be displayed on the starting interface for the user to select whether to enter the voice encryption state. The voice encryption option can be displayed in the form of text or in the form of text combined with an image. The text form can be, for example, “whether to enter the voice encryption state” or “whether to start the voice encryption function”. The real-time interaction interface can be an interface mainly displaying interaction information. The voice encryption option can be displayed in the form of an image in the edge area of the real-time interaction interface. The voice encryption option displayed on the starting interface or the real-time interaction interface can include multiple sub-options, for example, “whether to enter the first-level voice encryption state” or “whether to enter the second-level voice encryption state”.

[0106] In an embodiment, the voice interaction method 700 can further include updating the key in response to a key update instruction, and then encrypting the real-time acquired voice signal based on the updated key. Updating the key can improve the encryption strength to prevent invalid encryption caused by key theft.

[0107] In one example, the specific way of updating the key can be to update the key at preset time points. The time interval between the preset time points can be fixed, for example, the key can be updated every 5 minutes after entering the voice encryption state. The time interval between the preset time points can be unfixed, for example, the key can be updated every 5 minutes after entering the voice encryption state, and then the key can be updated every 10 minutes. The time interval for updating the key can increase or decrease in turn, which is not limited in the embodiments of the present application. The way of updating the key at a fixed time can improve the real-time performance of the key and improve the encryption strength to better prevent the leakage of interaction information.

[0108] In another example, the specific way of updating the key can be to update the key in response to an update instruction. The update instruction can be generated when the user triggers the second control switch in the voice output device. Based on this way, the user can judge whether the key needs to be updated, and trigger the update of the key at the time point when the key is needed. In this way, the user's autonomy and controllability of encryption can be enhanced.

[0109] In an embodiment, the key includes at least one of a frequency spectrum boundary point and a frequency spectrum cutoff point of the acquired voice signal. Correspondingly, in the step S702, encrypting the acquired voice signal based on the key can include: determining at least one frequency band in the acquired voice signal based on at least one of the frequency spectrum boundary point and the frequency spectrum cutoff point, and performing frequency spectrum inversion on the at least one frequency band, for example, performing frequency spectrum inversion on each frequency band in the at least one frequency band. Frequency spectrum inversion is an encryption algorithm in the analog domain.

[0110] In one example, the key can include a frequency spectrum boundary point Lo and a frequency spectrum cutoff point Hi of the voice signal as shown in Figure 8 Based on the frequency spectrum lower limit Lo and the frequency spectrum upper limit Hi shown in (a) of Figure 8 , one unique frequency band can be determined, that is, the frequency band between the frequency spectrum lower limit Lo and the frequency spectrum upper limit Hi. The entire frequency band can be inverted, the original high-frequency signal in the frequency band is placed at a low-frequency position, and the original low-frequency signal in the frequency band is placed at a high-frequency position to obtain the inverted frequency spectrum as shown in (b).

[0111] In another example, the key can include a frequency spectrum boundary point Lo and a frequency spectrum cutoff point Hi of the voice signal as shown in Figure 9The spectrum boundary points Lo and Hi of the speech signal, and the spectrum cut-off point a, wherein Lo is the lower limit of the spectrum, Hi is the upper limit of the spectrum, and the spectrum cut-off point a is a point selected between the lower limit of the spectrum Lo and the upper limit of the spectrum Hi, are shown based on Figure 9 The spectrum lower limit Lo, the spectrum upper limit Hi, and the spectrum cut-off point a shown in FIG. (a) can divide two frequency bands, the first frequency band between the spectrum lower limit Lo and the spectrum cut-off point a can be inverted as a whole, and the second frequency band between the spectrum cut-off point a and the spectrum upper limit Hi can be inverted as a whole, so that the high-frequency signal in each frequency band is placed at a low-frequency position, and the low-frequency signal in each frequency band is placed at a high-frequency position.

[0112] In the embodiments of the present application, the number of spectrum cut-off points can be one or more, Figure 9 The spectrum cut-off point a shown in FIG. (a) is only an example and does not limit the number of spectrum cut-off points. The scheme of setting the spectrum cut-off point between the spectrum boundary points can increase the number of inverted frequency bands, thereby increasing the encryption strength.

[0113] In an embodiment, based on at least one of the spectrum boundary points and the spectrum cut-off points, at least one frequency band in the obtained speech signal can be determined by inputting the obtained speech signal into a low-pass filter to attenuate signals greater than a cutoff frequency, the cutoff frequency of the low-pass filter being the upper limit of the spectrum, and the signals below the cutoff frequency being the determined frequency band, and the cutoff frequency can be a carrier frequency, for example, 4 kHz; or inputting the obtained speech signal into a low-pass filter and at least one band-pass filter, the low-pass filter can attenuate signals greater than a cutoff frequency, and the band-pass filter can attenuate other signals outside the specified frequency band, and based on the low-pass filter and the at least one band-pass filter, at least one spectrum cut-off point and at least two frequency bands can be determined.

[0114] In an example, the obtained speech signal can be input into a low-pass filter with a cutoff frequency of 4 kHz, the low-pass filter can attenuate signals greater than 4 kHz, and the signals below 4 kHz, i.e., 0-4 kHz, are retained, 0 kHz can be used as the lower limit of the spectrum, 4 kHz can be used as the upper limit of the spectrum, and 0-4 kHz is the determined frequency band.

[0115] In another example, the acquired voice signal can be input into a low-pass filter with a cutoff frequency of 1.6 kHz, which can attenuate signals greater than 1.6 kHz, and a band-pass filter with a cutoff frequency of 1.6 kHz and 4 kHz, which can attenuate signals in other frequency ranges outside the range of 1.6-4 kHz, 0 kHz can be the lower limit of the frequency spectrum, 4 kHz can be the upper limit of the frequency spectrum, 1.6 kHz can be the frequency spectrum cut-off point between the lower limit and the upper limit of the frequency spectrum, and 0-1.6 kHz and 1.6-4 kHz are the two frequency bands determined.

[0116] In an embodiment, the spectrum inversion of at least one frequency band can include mixing the voice signal of each frequency band with a carrier wave (sine wave) to obtain two sidebands, i.e., an upper sideband and a lower sideband, and attenuating the upper sideband through a low-pass filter, and the lower sideband is the frequency band after spectrum inversion. In an example, if the frequency band retained by the low-pass filter before inversion is 0-4 kHz, the signal in this frequency band can be mixed with a carrier wave (e.g., 4 kHz), which is equivalent to AM (amplitude modulation) modulation. If the frequency in this frequency band is represented as f, after mixing with the 4 kHz carrier wave, two sidebands of 4k-f (lower sideband) and 4k+f (upper sideband) can be obtained. By using a 4 kHz low-pass filter, 4k+f can be attenuated, and the retained 4k-f is the frequency band after spectrum inversion, i.e., the encrypted frequency band. All encrypted frequency bands can collectively form an encrypted voice signal.

[0117] In an embodiment, after mixing and low-pass filtering, the power of the voice signal will be lost. In order to make up for this loss, the signal obtained after mixing and low-pass filtering can be subjected to loudness normalization processing. The specific processing method can be to multiply the signal obtained after mixing and low-pass filtering by a number greater than 1 to increase the loudness. In the above step S703, the encrypted voice signal output to the terminal device or the server can be the signal after loudness normalization.

[0118] In an embodiment, when updating the key, the information of at least one of the lower limit of the frequency spectrum, the upper limit of the frequency spectrum, and the spectrum cut-off point can be updated. In an example, referring to Figure 10 , the spectrum cut-off point at t1 can be point a, i.e., spectrum cut-off point a, and the frequency band is divided based on the spectrum cut-off point a and inverted, and at t2, the spectrum cut-off point can be updated to point b, i.e., spectrum cut-off point b, and the frequency band is divided based on the spectrum cut-off point b and inverted. The ASR system can include a model retrained for the encryption method of spectrum inversion, and updating the information of at least one of the lower limit of the frequency spectrum, the upper limit of the frequency spectrum, and the spectrum cut-off point can make the retrained model in the ASR system invalid, and better prevent the voice information from being leaked to the ASR system.

[0119] In an embodiment, in the voice interaction method described above, after the spectrum inversion of the at least one frequency band, the voice signal after the spectrum inversion can be input into a low-pass filter, so that the part greater than the cutoff frequency is attenuated, and the encrypted voice signal output can be the signal output after the low-pass filter processing. The cutoff frequency of the low-pass filter can be set according to actual needs, for example, it can be set to 4 kHz (kilohertz), and the upper limit of the VoIP channel is usually 4 kHz. Human ears are usually more sensitive to signals below 4 kHz.

[0120] In an embodiment, the voice interaction method 700 described above can further include: outputting prompt information for prompting the user to turn off the voice enhancement function in the current interaction software (i.e., the VoIP software for implementing voice interaction), and turning off the voice enhancement function in the current interaction software in response to a confirmation operation for the prompt information; or performing reverse compensation on the voice signal before the spectrum inversion or the voice signal after the spectrum inversion. The reverse compensation can be used to compensate for the signal attenuation caused by the spectrum inversion.

[0121] In actual application, some interaction software can have a voice enhancement function, which can implement voice enhancement, noise reduction, etc., and can be equivalent to an equalizer (EQ). The voice enhancement function is usually used to adjust the gain of voice in different frequency bands. The voice enhancement function usually attenuates the part that is not sensitive to human ears (i.e., usually the high-frequency part) and enhances the part that is sensitive to human ears (usually the low-frequency part). The spectrum inversion will place the low-frequency part in the high-frequency position. After the spectrum inversion, using the voice enhancement function to attenuate the high-frequency part and enhance the low-frequency part is equivalent to attenuating the part that is sensitive to human ears and enhancing the part that is not sensitive to human ears, which will make the voice interaction effect worse.

[0122] To solve this problem, the embodiments of the present application provide two solutions: one solution is to prompt the user to turn off the original voice enhancement function, so as to avoid the adverse effect of the voice enhancement function after the spectrum inversion; the other solution is to perform reverse compensation on the voice signal before the spectrum inversion to compensate for the signal attenuation expected to be caused by the spectrum inversion, or to perform reverse compensation on the voice signal after the spectrum inversion to compensate for the signal attenuation that has already been caused by the spectrum inversion.

[0123] In one example, the reverse compensation on the speech signal before the spectrum inversion or the speech signal after the spectrum inversion can include: based on the frequency response of the VoIP channel obtained by testing the VoIP channel in advance, performing enhancement processing on the signal in a specified frequency range of the speech signal before the spectrum inversion or the speech signal after the spectrum inversion, to compensate for the signal attenuation in the specified frequency range after the spectrum inversion and the low-pass filtering processing. The specified frequency range can be a frequency range matching the frequency response of the VoIP channel, and the specified frequency range can be a lower frequency range, for example, 0-400 Hz (Hertz), or a higher frequency range, for example, 3600 Hz-4000 Hz, or other ranges, and the specific range can be determined according to the frequency response of the VoIP channel obtained by testing. The specific way of enhancement processing can be to construct an inverse compensation EQ model based on the frequency response of the VoIP channel obtained by testing the VoIP channel in advance, and to enhance the specified frequency range of the speech signal based on the EQ model.

[0124] In one example, if the specified frequency range is 0-400 Hz, the specific way of reverse compensation on the speech signal before the spectrum inversion can be: to perform enhancement processing on the signal in the range of 0-400 Hz in the speech signal before the spectrum inversion, to compensate for the attenuation caused by inverting 0-400 Hz to the position of 3600-4000 Hz and processing by the low-pass filter; and the specific way of reverse compensation on the speech signal after the spectrum inversion can be: the signal in the original 0-400 Hz is inverted to the position of 3600-4000 Hz after the spectrum inversion, at this time, the signal in the range of 3600-4000 Hz in the inverted speech signal can be enhanced to compensate for the attenuation caused by inverting the signal in the original 0-400 Hz to the position of 3600-4000 Hz and processing by the low-pass filter.

[0125] The specific way of enhancement processing can be to construct an inverse compensation EQ model based on the frequency response of the VoIP channel obtained by testing the VoIP channel in advance, and to enhance the specified frequency range of the speech signal based on the EQ model.

[0126] In one embodiment, the voice interaction method 700 described above can further include: before the spectrum inversion of the at least one frequency band, removing the direct current signal in the obtained speech signal, preventing the generation of direct current bias, and preventing the direct current signal from being inverted to high frequency signal to cause large noise interference in the subsequent spectrum inversion process.

[0127] Figure 11 A specific example of the voice interaction method applied to the voice output device provided by the embodiments of the present application is shown, and with reference to Figure 11 , the voice interaction method 1100 can include the following steps:

[0128] S1101, obtain an audio sampling point, i.e., a sampling point of a speech signal input by a speaker; S1102, remove a direct current bias in the audio sampling point; S1103, perform inverse compensation on the audio sampling point after the direct current bias is removed; S1104, input the audio sampling point after the inverse compensation into a 4 kHz low-pass filter to perform low-pass filtering; S1105, perform mixing processing on the filtered audio sampling point and a 4 kHz carrier (sine wave); S1106, input the mixed signal into a 4 kHz low-pass filter to perform low-pass filtering; and S1107, perform loudness normalization processing on the filtered audio sampling point and then output.

[0129] Based on the same technical concept, referring to Figure 12 The embodiments of the present application also provide a voice interaction method 1200, which can be applied to a server. The voice interaction method 1200 can include the following steps S1201-S1202.

[0130] S1201, receive an encrypted voice signal transmitted by a voice output device or a terminal device.

[0131] The voice output device can be any voice output device provided by the embodiments of the present application, and the encrypted voice signal transmitted by the terminal device can be provided by any voice output device provided by the embodiments of the present application.

[0132] S1202, forward the encrypted voice signal to a voice receiving device of an interaction object.

[0133] In an implementation manner, the voice interaction method 1200 can further include receiving voice transmission rule information transmitted by the voice output device or the terminal device. Correspondingly, the step of forwarding the encrypted voice signal to the voice receiving device of the interaction object can include: transmitting the encrypted voice signal to the voice receiving device of the interaction object corresponding to a specified permission level in the voice transmission rule information; or, transmitting the encrypted voice signal to the voice receiving device of the interaction object corresponding to each permission level, and sending a key to the interaction object corresponding to the specified permission level.

[0134] Based on the above manner, the server can make the specified interactive object obtain the voice content in two ways. One way is to only transmit the encrypted voice signal to the voice receiving device of the interactive object corresponding to the specified permission level (e.g., high permission level) in the voice transmission rule information, and not to the voice receiving device of the interactive object of other permission levels (e.g., low permission level), that is, the interactive object of other permission levels cannot obtain the encrypted voice signal, and thus cannot decrypt, thereby preventing the voice content delivered to part of the interactive objects from being leaked to other interactive objects; another way is to transmit the encrypted voice signal to the voice receiving device of all interactive objects without distinction, but only send the key to the voice receiving device of the interactive object corresponding to the specified permission level (e.g., high permission level), that is, only the voice receiving device of the interactive object corresponding to the specified permission level can decrypt the received encrypted voice signal to obtain the correct voice content, and the voice receiving device of the interactive object corresponding to other permission levels cannot decrypt even if it receives the encrypted voice signal, and cannot obtain the correct voice content.

[0135] In an implementation, the voice interaction method 1200 can further include: obtaining device information of each interactive object corresponding to the multi-person voice interaction; determining whether the voice receiving device of the interactive object corresponding to the specified permission level is a voice receiving device with decryption function according to the device information, if yes, sending the key to the voice receiving device of the interactive object corresponding to the specified permission level, so that the voice receiving device can decrypt the encrypted voice signal according to the key, and if not, maintaining the current state. The key can be obtained from the voice output device.

[0136] In an example, the device information of the interactive object can include a device identifier of the voice receiving device of the interactive object, such as a model or other special identifier, which includes information of whether the voice receiving device has decryption function or is a decryption device, and according to the device identifier, it can be determined whether the voice receiving device of each interactive object is a device with decryption function, that is, a decryption device. In the embodiment of the application, the device with decryption function can be a device that has stored a key when it leaves the factory.

[0137] For the voice receiving device without decryption function, the key can be provided to it so that it can decrypt the encrypted voice signal, so that it can normally listen to the voice information of the speaker. For example, in a multi-person voice interaction scenario, the settings used by each user participating in the voice interaction can not be uniform, and only part of the users use voice receiving devices with decryption function and can directly decrypt the voice signal to normally listen to the voice information, and the remaining users cannot directly decrypt. For the voice receiving device of the user who cannot decrypt, the key can be provided to it so that the part of the user can normally listen to the voice information.

[0138] In an implementation, the voice interaction method 1200 can further include: obtaining an interaction object list uploaded by a voice output device or a terminal device, the interaction object list including a plurality of interaction objects and organizational structure information of the plurality of interaction objects, and determining a permission level of each interaction object in the interaction object list according to the organizational structure information of the interaction object.

[0139] In an example, for a meeting within an enterprise, the organizational structure information can include employee information within the enterprise and department information to which the employee belongs, so as to determine the permission level of each employee. For example, if employee A belongs to the finance department, employee A has the permission to listen to the financial data, if employee A belongs to the board of directors or other senior management department, employee A has the permission to listen to the account information and other information, and employees in the same department have the same permission level.

[0140] In another example, for a meeting between enterprises, the organizational structure information can include user information and enterprise information to which the user belongs. For example, if user B belongs to enterprise C, i.e., user B is an employee of enterprise C, user B has the permission to listen to information related to the internal affairs of enterprise C, if user B belongs to other enterprises, user B does not have the permission to listen to the information related to the internal affairs of enterprise C, and users in the same enterprise can have the same permission level.

[0141] Based on the organizational structure information of the interaction object, the permission level can be quickly divided, and when the encrypted voice signal and the voice transmission rule information are received, the correct voice content can be delivered to the interaction object with the specified permission level.

[0142] In an implementation, the voice interaction method 1200 can further include: obtaining interaction process information corresponding to the multi-person voice interaction, and determining an association relationship between each voice interaction period and each permission level according to the interaction process information. Correspondingly, in the step S1202, forwarding the encrypted voice signal to the voice receiving device of the interaction object can include: during the multi-person voice interaction, forwarding the encrypted voice signal to the first interaction object, and the first interaction object can be an interaction object corresponding to the permission level associated with the voice interaction period to which the current time belongs.

[0143] In an example, for an online meeting, the interaction process information can be a meeting agenda. Based on the meeting agenda, it can be determined which content is shared with which participant at which time period, the participant corresponds to a determined permission level, and then the association relationship between each time period and each permission level can be determined. When the server forwards the encrypted voice signal, the time period to which the current time belongs can be identified, and the encrypted voice signal can be sent to the participant corresponding to the time period.

[0144] In an implementation, the voice interaction method 1200 can further include: sending the association between the voice interaction time periods and the permission levels to a second interaction object during the multi-person voice interaction. The second interaction object is an interaction object other than the first interaction object, i.e., an interaction object of a permission level not associated with the voice interaction time period to which the current time belongs. For example, during a time period T of an online conference, the server can send the association between the conference time periods and the permission levels to a participant of a permission level not associated with the time period T, and the participant can determine whether the conference content of the current time period T is related to the participant and whether the participant has the permission to listen according to the association, and when it is determined that the participant does not have the permission to listen, the participant can choose to continue participating in the conference or temporarily leave the conference.

[0145] Based on the same technical concept, referring to Figure 13 The embodiments of the present application also provide a voice interaction method 1300, which can be applied to a voice receiving device. The method can include the following steps S1301-S1302.

[0146] S1301, receiving an encrypted voice signal sent by a server.

[0147] The server can be any server provided by the embodiments of the present application, and the server can send the encrypted voice signal by the voice interaction method 1200 provided by the embodiments of the present application.

[0148] S1302, decrypting the encrypted voice signal based on a key.

[0149] The key can be the key used in the voice interaction method applicable to the voice output device provided by any of the embodiments of the present application.

[0150] The key in the voice receiving device can be the same key as the voice output device, for example, a voice receiving device with decryption function usually pre-stores the same key as the voice output device; the key in the voice receiving device can also be a real-time acquired key, for example, a voice receiving device without decryption function usually does not pre-store the same key as the voice output device, and can acquire the key or the updated key provided by the voice output device and forwarded by the server in real time.

[0151] In an example scenario, in a multi-person voice interaction process, there can be multiple voice output devices outputting voice at the same time, for example, in a meeting, multiple people speak at the same time, and each voice receiving device can obtain the key provided by each voice output device at the same time, and decrypt the encrypted voice signal output by the corresponding voice output device based on each key. The keys provided by each voice output device received by the same voice receiving device can be the same, for example, uniformly as a pre-agreed key, or different, for example, the update frequency and update rule of the key of different voice output devices can be different, for example, the encryption range selected by the user of different voice output devices and the permission of the same listener set are different, so that the keys output by the voice receiving device of the same listener are different.

[0152] In an embodiment, in the step S1302, decrypting the encrypted voice signal based on the key can include: detecting whether there is suspicious software on the current voice receiving device; in a case where it is determined that there is no suspicious software on the current voice receiving device, decrypting the encrypted voice signal based on the key; and in a case where it is determined that there is suspicious software on the current voice receiving device, not decrypting the encrypted voice signal. The suspicious software can be software that can listen to the voice interaction content, and detecting the suspicious software can determine whether the software environment of the current voice receiving device is safe, and the encrypted voice signal obtained is decrypted in a safe environment, which can further enhance the privacy protection function.

[0153] In the embodiments of the present application, the specific manner of decrypting the encrypted voice signal can be: inverting the spectrum of at least one frequency band in the encrypted voice signal according to the information of at least one of the frequency spectrum lower limit, the frequency spectrum upper limit and the frequency spectrum cut-off point in the key, which is equivalent to performing the same spectrum inversion operation as the encryption process again, that is, the decryption is a symmetric operation of the encryption, and the execution of the encryption operation twice is equivalent to the decryption.

[0154] Based on the same technical concept, the embodiments of the present application provide a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the method provided in the embodiments of the present application.

[0155] The embodiments of the present application also provide a chip, which includes a processor, and is used to call and run instructions stored in a memory, so that a communication device installed with the chip executes the method provided in the embodiments of the present application.

[0156] The embodiments of the present application also provide a chip, which includes: an input interface, an output interface, a processor and a memory, the input interface, the output interface, the processor and the memory are connected through internal connection paths, and the processor is used to execute the code in the memory, and when the code is executed, the processor is used to execute the method provided in the embodiments of the present application.

[0157] It should be understood that the processor in the embodiments of the present application can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. It is worth noting that the processor can be a processor supporting an advanced RISC machine (ARM) architecture.

[0158] Further, the memory in the embodiments of the present application can include a read-only memory and a random access memory. The memory can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM) or a flash memory. The volatile memory can include a random access memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM are available. For example, a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate SDRAM (DDR SDRAM), an enhanced SDRAM (ESDRAM), a sync link DRAM (SLDRAM) and a direct memory bus random access memory (Direct Rambus RAM, DR RAM).

[0159] In the above-described embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded on a computer, all or part of the processes or functions according to the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium.

[0160] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Also, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, a person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.

[0161] In addition, the terms "first", "second", etc. are used only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise explicitly specified.

[0162] Any process or method described in the flowchart or otherwise described herein can be understood as a representation of code including one or more executable instructions for performing a specific logical function or process. Also, the scope of the preferred embodiments of the present application includes additional implementations that can not be shown or discussed explicitly, including implementations in which functions are performed in different orders, in substantially simultaneous fashion, or in reverse order according to the functions involved.

[0163] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as a list of executable instructions for implementing the logic function, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus, such as a computer-based system, a system including a processor, or other system that can fetch and execute instructions from the instruction execution system, device or apparatus, or in conjunction with these instructions execution system, device or apparatus.

[0164] It should be understood that each part of the present application can be realized by hardware, software, firmware or a combination thereof. In the above embodiments, a plurality of steps or methods can be realized by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the above-mentioned embodiment methods can be completed by a program instructing the relevant hardware, which can be stored in a computer readable storage medium and includes one or a combination of the steps of the embodiment methods when executed.

[0165] In addition, each functional unit in each embodiment of the present application can be integrated in one processing module, or each unit can be physically present separately, or two or more units can be integrated in one module. The above-mentioned integrated module can be realized in the form of hardware or in the form of a software functional module. The above-mentioned integrated module, if realized in the form of a software functional module and sold or used as an independent product, can also be stored in a computer readable storage medium. The storage medium can be a read-only memory, a magnetic disk or an optical disk, etc.

[0166] The above is only an exemplary embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A voice output device, characterized in that, include: A first control switch, located on the outside of the voice output device, is used to generate an encryption enable command when triggered; wherein, the first control switch has multiple positions, each position generates an encryption enable command of a level when triggered, and each level of encryption enable command is associated with a level of voice encryption state; An encryption control module, electrically connected to the first control switch, is used to enter a voice encryption state in response to the encryption enable command before or during multi-person voice interaction, and to encrypt the voice signal obtained based on the key pair in the voice encryption state to generate an encrypted voice signal; the encryption control module is also used to enter a level of voice encryption state associated with a level of encryption enable command when a level of encryption enable command is received, and to generate voice transmission rule information in the voice encryption state; The first communication module is electrically connected to the encryption control module and is used to transmit the encrypted voice signal to a terminal device or server connected to the voice output device; the terminal device is used to forward the encrypted voice signal to the server, and the server is used to forward the encrypted voice signal to the voice receiving device of the interactive object based on the specified permission level in the voice transmission rule information. The first communication module is further configured to transmit the voice transmission rule information to the terminal device or the server; the voice transmission rule information includes a specified permission level, which is a permission level associated with the current voice encryption state; the voice transmission rule information is used by the server to send the encrypted voice signal to the voice receiving device of the interactive object corresponding to the specified permission level.

2. The voice output device according to claim 1, characterized in that, Also includes: The output module is used to output a list of multiple voice interaction objects for the current voice encryption state, and to generate selection instructions when a selection operation is performed. The encryption control module is further configured to respond to the selection instruction to determine the selected interaction object as the interaction object of the specified permission level, and generate the voice transmission rule information based on the interaction object of the specified permission level.

3. The voice output device according to any one of claims 1-2, characterized in that, Also includes: The second control switch is used to generate a key update instruction when triggered. The encryption control module is used to update the key in response to the key update command.

4. A server, characterized in that, include: The second communication module is used to receive encrypted voice signals transmitted by the voice output device or terminal device according to any one of claims 1-3, and forward the encrypted voice signals to the voice receiving device of the interactive object.

5. The server according to claim 4, characterized in that, The second communication module is also used to receive voice transmission rule information transmitted by the voice output device or the terminal device; When forwarding the encrypted voice signal to the voice receiving device of the interactive object, the second communication module is used to transmit the encrypted voice signal to the voice receiving device of the interactive object corresponding to the specified permission level in the voice transmission rule information, or to transmit the encrypted voice signal to the voice receiving device of the interactive object corresponding to each permission level and send a key to the voice receiving device of the interactive object corresponding to the specified permission level.

6. The server according to claim 4 or 5, characterized in that, The second communication module is further configured to: obtain a list of interactive objects uploaded by the voice output device or the terminal device; the list of interactive objects includes multiple interactive objects and organizational structure information of the multiple interactive objects; The server further includes a processing module electrically connected to the second communication module, used to determine the corresponding permission level of each interactive object in the interactive object list based on the organizational structure information of each interactive object in the interactive object list.

7. A voice receiving device, characterized in that, include: The third communication module is used to receive encrypted voice signals sent by the server according to any one of claims 4-6; The decryption control module is used to decrypt the encrypted voice signal based on the key.

8. A voice interaction method, characterized in that, Applied to a voice output device, the method includes: When the first control switch located on the outside of the voice output device is triggered, an encrypted activation command is generated. The first control switch has multiple positions, and each position generates an encrypted activation command of a certain level when triggered. Before or during multi-person voice interaction, in response to an encryption enable command, the system enters a voice encryption state. The encryption enable command includes multiple levels of encryption enable commands, and each level of encryption enable command is associated with a level of voice encryption state. Specifically, when an encryption enable command of a certain level is received, the system enters a level of voice encryption state associated with that level of encryption enable command, and voice transmission rule information is generated in that voice encryption state. In the voice encryption state, the acquired voice signal is encrypted based on the key to generate an encrypted voice signal; The encrypted voice signal is transmitted to a terminal device or server connected to the voice output device; the terminal device is used to forward the encrypted voice signal to the server; and the server is used to forward the encrypted voice signal to the voice receiving device of the interactive object based on the specified permission level in the voice transmission rule information. The voice transmission rule information is transmitted to the terminal device or the server; the voice transmission rule information includes a specified permission level, which is a permission level associated with the current voice encryption state; the voice transmission rule information is used by the server to send the encrypted voice signal to the voice receiving device of the interactive object corresponding to the specified permission level.

9. The voice interaction method according to claim 8, characterized in that, Also includes: Output a list of interaction objects corresponding to the current voice encryption state for multi-person voice interaction; In response to a selection instruction for an interaction object in the list of interaction objects, the selected interaction object is determined as the interaction object for the specified permission level; The generation of voice transmission rule information in this voice encryption state includes: Voice transmission rule information is generated based on the interactive object with the specified permission level.

10. The voice interaction method according to claim 8, characterized in that, Also includes: In the current voice encryption state, in response to the state switching command, switch to another level of voice encryption state; The state switching command is generated when the position of the first control switch of the voice output device is switched.

11. The voice interaction method according to claim 10, characterized in that, In the voice encryption state, encrypting the acquired voice signal based on the key includes: Under different levels of voice encryption, the acquired voice signals are encrypted using different keys.

12. The voice interaction method according to any one of claims 8-11, characterized in that, Also includes: The voice encryption option is displayed on the startup interface or real-time interaction interface of the multi-person voice interaction. In response to a trigger operation on the voice encryption option, an encryption enable command is generated.

13. The voice interaction method according to any one of claims 8-11, characterized in that, The key includes at least one of the following: the spectral boundary points and the spectral cutoff points of the acquired speech signal. The encryption of the acquired voice signal based on the key includes: Based on at least one of the spectral boundary points and the spectral cutoff points, at least one frequency band in the acquired speech signal is determined; The spectrum of at least one frequency band is inverted.

14. The voice interaction method according to claim 13, characterized in that, Also includes: Output a prompt message, and in response to a confirmation operation on the prompt message, disable the voice enhancement function in the current interactive software; Alternatively, reverse compensation may be performed on the speech signal before or after the spectrum inversion; the reverse compensation is used to compensate for the signal attenuation caused by the spectrum inversion.

15. A voice interaction method, characterized in that, Applied to a server, the method includes: Receive encrypted voice signals transmitted by the voice output device or terminal device according to any one of claims 1-3; The encrypted voice signal is forwarded to the voice receiving device of the interactive object.

16. The voice interaction method according to claim 15, characterized in that, Also includes: Receive voice transmission rule information transmitted by the voice output device or the terminal device; The forwarding of the encrypted voice signal to the voice receiving device of the interactive object includes: The encrypted voice signal is transmitted to the voice receiving device of the interactive object corresponding to the permission level specified in the voice transmission rule information; Alternatively, the encrypted voice signal may be transmitted to the voice receiving device of the interactive object corresponding to each permission level, and a key may be sent to the interactive object corresponding to the specified permission level.

17. The voice interaction method according to claim 15 or 16, characterized in that, Also includes: Obtain the list of interactive objects uploaded by the voice output device or the terminal device; The list of interactive objects includes multiple interactive objects and organizational structure information of the multiple interactive objects; Assign appropriate permission levels to each interactive object in the interactive object list based on the organizational structure information of each interactive object in the list.

18. The voice interaction method according to claim 15, characterized in that, Also includes: Obtain the interaction flow information corresponding to the multi-person voice interaction; The relationship between each voice interaction period and each permission level is determined based on the interaction process information. The forwarding of the encrypted voice signal to the voice receiving device of the interactive object includes: During the multi-person voice interaction, the encrypted voice signal is forwarded to the first interaction object; The first interaction object is the interaction object corresponding to the permission level associated with the voice interaction time period to which the current moment belongs.

19. The voice interaction method according to claim 18, characterized in that, Also includes: During the multi-person voice interaction, the association between each voice interaction period and each permission level is sent to the second interaction object; The second interaction object is an interaction object other than the first interaction object.

20. A voice interaction method, characterized in that, Applied to a voice receiving device, the method includes: Receive encrypted voice signals sent by the server according to any one of claims 4-6; The encrypted voice signal is decrypted based on a key; the key is the key used in the voice interaction method according to any one of claims 8-14.

21. The voice interaction method according to claim 20, characterized in that, The decryption of the encrypted voice signal based on the key includes: Detect whether there is suspicious software on the current voice receiving device; If it is determined that there is no suspicious software on the current voice receiving device, the encrypted voice signal is decrypted based on the key.

22. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the voice interaction method according to any one of claims 8-21.

Citation Information

Patent Citations

  • Voice encryption method, device and system, electronic equipment and storage medium

    CN113225310A

  • Bluetooth earphone with voice message encryption, sending and receiving functions

    CN115361616A