Earphone function control method and device, earphone and storage medium

By performing component analysis and privacy processing of headphone audio signals, the problem of privacy information leakage in headphone function control is solved, achieving higher security and user experience.

CN120277402APending Publication Date: 2025-07-08GEER TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510399375.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art does not consider personal privacy protection in the headphone function control, resulting in a high risk of privacy information leakage.

Method used

The target multi-layer audio signal processing model is used to analyze the components of the original audio signal, identify sensitive information and perform privacy processing, eliminate privacy audio signals, and generate target voiceprint feature information to control the headphone function.

Benefits of technology

It effectively reduces the possibility of user privacy leakage and improves the security and user experience of headphone function control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277402A_ABST
    Figure CN120277402A_ABST
Patent Text Reader

Abstract

The invention discloses an earphone function control method and device, an earphone and a storage medium, and relates to the technical field of data processing, and the method comprises the steps: carrying out the component analysis of an original audio signal based on a target multilayer audio signal processing model; privacy processing is carried out on the audio signal corresponding to the sensitive information according to the target information privacy strategy; removing the privacy audio signal from the original audio signal, and generating target voiceprint feature information according to the original audio signal after removal; and controlling the earphone function according to the target control instruction. Through the above mode, when the sensitive information exists in the audio component information, the privacy audio signal is removed from the original audio signal, the possibility of leaking user privacy can be reduced to the greatest extent, and after the target control instruction is generated according to the target voiceprint feature information, the earphone function is controlled according to the target control instruction, so that the user experience is improved. Therefore, the safety of controlling the earphone function can be effectively improved, and the user experience is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a method and device for controlling headphone functions, headphones, and a storage medium. Background Art

[0002] With the continuous development of voiceprint recognition technology, more and more devices have begun to use voiceprint recognition technology for function control. For example, in the case of Bluetooth headphones, voiceprint recognition technology is inseparable from the acquisition of audio signals, and the acquired audio signals often involve the personal privacy of users. For example, ID numbers, bank card numbers, and passwords. Currently, in the process of voiceprint recognition, the protection of personal privacy is not considered. If the headphone functions are directly controlled, the above personal privacy will be improperly stored or transmitted, resulting in a risk of leakage. Therefore, the security of controlling headphone functions in the above manner is relatively low.

[0003] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of this application is to provide a method and device for controlling headphone functions, headphones, and a storage medium, aiming to solve the technical problem of relatively low security in controlling headphone functions in the prior art.

[0005] To achieve the above purpose, this application proposes a method for controlling headphone functions, and the method includes:

[0006] Performing component analysis on the original audio signal based on a target multi-layer audio signal processing model to obtain audio component information;

[0007] When there is sensitive information in the audio component information, performing privacy processing on the audio signal corresponding to the sensitive information according to a target information privacy strategy to obtain a privacy-processed audio signal;

[0008] Removing the privacy-processed audio signal from the original audio signal, and generating target voiceprint feature information according to the original audio signal after removal;

[0009] Generating a target control instruction according to the target voiceprint feature information, and controlling the headphone functions according to the target control instruction.

[0010] In an embodiment, the step of performing component analysis on the original audio signal based on a target multi-layer audio signal processing model to obtain audio component information includes:

[0011] Collecting the original audio signal based on a sound collection unit composed of a multi-microphone array;

[0012] Transmitting the original audio signal to a signal processing unit through a target data transmission bus;

[0013] De-noising the original audio signal based on the signal processing unit, and performing frame processing on the de-noised original audio signal to obtain a number of short-frame audio signals;

[0014] Based on the target multi-layer audio signal processing model, component analysis is performed on the several numbers of short-frame audio signals to obtain audio component information.

[0015] In one embodiment, when sensitive information exists in the audio component information, the step of performing privacy processing on the audio signal corresponding to the sensitive information according to the target information privacy policy to obtain the privacy-protected audio signal includes:

[0016] When sensitive information exists in the audio component information, obtaining an audio signal corresponding to the sensitive information;

[0017] Determining the source of the audio signal corresponding to the sensitive information, and determining a target information privacy policy based on the source;

[0018] The audio signal corresponding to the sensitive information is privacy-processed according to the target information privacy policy to obtain a privacy-protected audio signal.

[0019] In one embodiment, the step of determining the target information privacy policy according to the source includes:

[0020] When the source is user conversation content, determining the target information privacy strategy to be a time domain masking privacy strategy;

[0021] When the source is the environmental background, it is determined that the target information privacy strategy is a frequency domain filtering privacy strategy.

[0022] In one embodiment, the step of generating a target control instruction according to the target voiceprint feature information, and controlling the earphone function according to the target control instruction includes:

[0023] Acquiring the user's standard voiceprint feature information based on a pre-stored user voiceprint template;

[0024] Calculating the similarity between the target voiceprint feature information and the standard voiceprint feature information;

[0025] When the similarity is greater than a preset similarity threshold, determining the speech content corresponding to the original audio signal after being eliminated;

[0026] A target control instruction is generated according to the voice content, and the earphone function is controlled according to the target control instruction.

[0027] In one embodiment, after the step of determining the speech content corresponding to the original audio signal after elimination when the similarity is greater than a preset similarity threshold, the method further includes:

[0028] Generating information to be encrypted according to the target voiceprint feature information, the original audio signal after elimination, and the operation record of the original audio signal;

[0029] Encrypting the information to be encrypted according to the target encryption policy to obtain target encrypted information;

[0030] Directly clearing the target encrypted information.

[0031] In one embodiment, before the step of performing privacy processing on the audio signal corresponding to the sensitive information according to the target information privacy policy when there is sensitive information in the audio component information to obtain a privacy-processed audio signal, the method further includes:

[0032] When detecting the invocation of the privacy protection setting interface, displaying a sensitive information setting interface;

[0033] Obtaining user input information and / or user selected information based on the sensitive information setting interface;

[0034] Performing sensitive detection on the user input information and / or user selected information, and generating sensitive information according to the sensitive detection result.

[0035] In addition, to achieve the above object, the present application further provides a headphone function control device, the device includes:

[0036] An analysis module, configured to perform component analysis on the original audio signal based on a target multi-layer audio signal processing model to obtain audio component information;

[0037] A privacy processing module, configured to perform privacy processing on the audio signal corresponding to the sensitive information according to the target information privacy policy to obtain a privacy-processed audio signal when there is sensitive information in the audio component information;

[0038] An elimination module, configured to eliminate the privacy-processed audio signal from the original audio signal, and generate target voiceprint feature information according to the original audio signal after elimination;

[0039] A control module, configured to generate a target control instruction according to the target voiceprint feature information, and control the headphone function according to the target control instruction.

[0040] In addition, to achieve the above object, the present application further provides a headset, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the headset function control method as described above.

[0041] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the headset function control method as described above are implemented.

[0042] One or more technical solutions proposed by the present application have at least the following technical effects: performing component analysis on the original audio signal based on the target multi-layer audio signal processing model to obtain audio component information; when there is sensitive information in the audio component information, performing privacy processing on the audio signal corresponding to the sensitive information according to the target information privacy strategy to obtain a privacy-processed audio signal; removing the privacy-processed audio signal from the original audio signal, and generating target voiceprint feature information according to the original audio signal after removal; generating a target control instruction according to the target voiceprint feature information, and controlling the headset function according to the target control instruction. In the above manner, when it is determined that there is sensitive information in the audio component information, the privacy-processed audio signal is removed from the original audio signal, which can minimize the possibility of leaking user privacy, and after the target voiceprint feature information generates the target control instruction, the headset function is controlled according to the target control instruction, thereby effectively improving the security of controlling the headset function and further improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0044] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the accompanying drawings required for describing the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 It is a schematic flowchart provided for the first embodiment of the headset function control method of the present application;

[0046] Figure 2 It is a schematic flowchart provided for the second embodiment of the headset function control method of the present application;

[0047] Figure 3Schematic diagram of the module structure of the headphone function control device according to an embodiment of the present application;

[0048] Figure 4 Schematic diagram of the device structure of the hardware operating environment involved in the headphone function control method according to an embodiment of the present application.

[0049] The realization of the purpose, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0050] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a microprocessor, etc. that can implement the above functions. Hereinafter, taking the microprocessor as an example, this embodiment and the following embodiments will be described.

[0051] Based on this, an embodiment of the present application provides a headphone function control method, referring to Figure 1 , Figure 1 Schematic diagram of the process of the first embodiment of the headphone function control method of the present application.

[0052] In this embodiment, the headphone function control method includes steps S10 to S40:

[0053] Step S10, perform component analysis on the original audio signal based on the target multi-layer audio signal processing model to obtain audio component information.

[0054] It should be noted that the headphones for function control in this embodiment have a privacy protection function, that is, the headphone function is controlled only after privacy processing and voiceprint verification are passed. The headphones can be ordinary Bluetooth headphones or smart Bluetooth headphones. In addition, the headphones include but are not limited to a microprocessor, a wireless communication module, a power module, and other auxiliary function modules. Among them, the microprocessor is arranged in the headphones. As the core processing chip of the headphones, the microprocessor has the characteristics of high performance and low power consumption, and a signal processing unit is newly added. The signal processing unit uses a dedicated chip with powerful digital signal processing capabilities and can be used for denoising, frame segmentation, and component analysis, etc. The signal processing unit can be a certain high-performance audio processing chip of Texas Instruments.

[0055] It can be understood that the target multi-layer audio signal processing model can be an audio signal processing model constructed by an advanced audio processing algorithm based on deep learning. The target multi-layer audio signal processing model includes a multi-layer convolutional neural network. Among them, the convolutional neural network is used to extract the time-frequency features of the original audio signal, the recurrent neural network is used to model the time series information, extract features and classify the original audio signal, and the analysis neural network is used to analyze the components of the original audio signal in real time and accurately. The component information refers to the information that characterizes the components in the original audio signal.

[0056] Further, step S10 includes: collecting the original audio signal by a sound collection unit composed of a multi-microphone array; transmitting the original audio signal to the signal processing unit through the target data transmission bus; denoising the original audio signal based on the signal processing unit, and performing frame splitting on the denoised original audio signal to obtain a plurality of short-frame audio signals; performing component analysis on the plurality of short-frame audio signals respectively based on the target multi-layer audio signal processing model to obtain audio component information.

[0057] It should be understood that in order to effectively improve the accuracy and directivity of the collected audio signal and suppress environmental noise, in this embodiment, a sound collection unit composed of a multi-microphone array is used to collect the original audio signal, which includes but is not limited to the user's speech sound, the surrounding environmental sound, etc. To improve the quality of the audio signal, the original audio signal is denoised based on the signal processing unit to remove background noise. The algorithms used for denoising can be adaptive filtering algorithms, wavelet denoising algorithms, etc. To facilitate the analysis of audio component information, frame splitting needs to be performed, that is, the continuous denoised original audio signal is segmented into short-frame audio signals with each frame being 20 - 30 milliseconds, and the above-mentioned plurality of short-frame audio signals are successively input into the target multi-layer audio signal processing model, and the target multi-layer audio signal processing model analyzes the audio component information.

[0058] It should be noted that the target data transmission bus has the characteristic of high speed. Through the target data transmission bus, the signal processing unit, the sound collection unit, and the voiceprint recognition unit are closely connected to ensure the efficiency and stability of data internal transmission.

[0059] Step S20, when there is sensitive information in the audio component information, perform privacy processing on the audio signal corresponding to the sensitive information according to the target information privacy policy to obtain a privacy-processed audio signal.

[0060] It can be understood that the target information privacy protection strategy refers to the strategy of privacy-protecting the audio signal corresponding to sensitive information. The target information privacy protection strategy includes, but is not limited to, time-domain masking privacy protection strategy, frequency-domain filtering privacy protection strategy, etc. When it is determined that sensitive information exists in the audio component information, it indicates that there is information related to user privacy in the original audio signal. At this time, according to the target information privacy protection strategy, the audio signal corresponding to the sensitive information is processed into a privacy-protected audio signal.

[0061] Further, before step S20, it further includes: when detecting the invocation of the privacy protection setting interface, displaying a sensitive information setting interface; obtaining user input information and / or user selected information based on the sensitive information setting interface; performing sensitive detection on the user input information and / or user selected information, and generating sensitive information according to the sensitive detection result.

[0062] It should be understood that the privacy protection setting interface is an interface for setting sensitive information. This privacy protection setting interface can be provided by the earphone. The way to invoke the privacy protection setting interface can be the terminal APP, the earphone operation button, etc. After displaying the sensitive information setting interface, the user can set sensitive information by checking, inputting, etc. Among them, the user selected information is obtained by the checking method, and the user input information is obtained by the input method. The user input information and user selected information include, but are not limited to, vocabulary, frequency band, etc.

[0063] It can be understood that in order to effectively improve the accuracy of setting sensitive information, it is also necessary to perform sensitive detection on the user input information and / or user selected information to screen out non-sensitive information, and then generate sensitive information according to the sensitive detection result. In addition, the user can also adjust the deletion rule of the privacy-protected audio signal in the sensitive information setting interface according to their own needs. The deletion rule includes, but is not limited to, deletion intensity, deletion method, etc., to meet the privacy protection needs of different users.

[0064] Step S30, removing the privacy-protected audio signal from the original audio signal, and generating target voiceprint feature information according to the original audio signal after removal.

[0065] It should be understood that after obtaining the privacy-protected audio signal, the privacy-protected audio signal will be immediately removed from the original audio signal to ensure that the privacy-protected audio signal will not be saved or transmitted, thereby avoiding the leakage of user privacy information.

[0066] It should be noted that after obtaining the original audio signal after removal, in order to effectively improve the accuracy and robustness of generating target voiceprint feature information, the Mel Frequency Cepstral Coefficient algorithm can be used, combined with a deep neural network classifier to perform voiceprint recognition on the original audio signal after removal, and extract features from the recognized voiceprint information to obtain target voiceprint feature information.

[0067] Step S40: Generate a target control instruction according to the target voiceprint feature information, and control the headphone functions according to the target control instruction.

[0068] It can be understood that the target control instruction refers to an instruction for controlling the headphone functions. For example, playing / pausing music, answering / hanging up a call, turning up / turning down the volume, etc. In addition, the technical solution of this embodiment can also be applied to other intelligent devices that require voiceprint recognition technology, such as smart home, intelligent vehicle system, etc., so as to fully protect the security of the user's personal privacy while enjoying the convenience brought by the voiceprint recognition technology.

[0069] Further, step S40 includes: obtaining the standard voiceprint feature information of the user based on the pre-stored user voiceprint template; calculating the similarity between the target voiceprint feature information and the standard voiceprint feature information; when the similarity is greater than the preset similarity threshold, determining the speech content corresponding to the original audio signal after removal; generating a target control instruction according to the speech content, and controlling the headphone functions according to the target control instruction.

[0070] It should be understood that the similarity between the target voiceprint feature information and the standard voiceprint feature information can be calculated using algorithms such as dynamic time warping algorithm, hidden Markov model, etc. When it is determined that the similarity is greater than the preset similarity threshold, it indicates successful recognition. At this time, a target control instruction is generated according to the speech content corresponding to the original audio signal after removal. On the contrary, when it is determined that the similarity is less than or equal to the preset similarity threshold, it indicates failed recognition. At this time, a prompt tone or feedback information for the user to re-enter the control instruction is issued.

[0071] Further, after the step of determining the speech content corresponding to the original audio signal after removal when the similarity is greater than the preset similarity threshold, it further includes: generating information to be encrypted according to the target voiceprint feature information, the original audio signal after removal, and the operation record of the original audio signal; encrypting the information to be encrypted according to the target encryption strategy to obtain target encrypted information; directly clearing the target encrypted information.

[0072] It can be understood that, in order to avoid disclosing other personal information of the user, after determining the speech content corresponding to the original audio signal after removal, information to be encrypted is generated according to the target voiceprint feature information, the original audio signal after removal, and the operation record of the original audio signal. First, the information to be encrypted is encrypted using the target encryption strategy to ensure the security of the information to be encrypted. Then, the target encrypted information is directly cleared, that is, the target encrypted information does not exist from beginning to end and the above relevant information cannot be obtained from the Bluetooth headset. Compared with the timed clearing after storage, it can effectively reduce the leakage risk caused by storing information.

[0073] In this embodiment, the original audio signal is subjected to component analysis based on the target multi-layer audio signal processing model to obtain audio component information; when there is sensitive information in the audio component information, the audio signal corresponding to the sensitive information is subjected to privacy processing according to the target information privacy strategy to obtain a privacy-processed audio signal; the privacy-processed audio signal is removed from the original audio signal, and target voiceprint feature information is generated according to the original audio signal after removal; a target control instruction is generated according to the target voiceprint feature information, and the headphone function is controlled according to the target control instruction. In the above manner, when it is determined that there is sensitive information in the audio component information, the privacy-processed audio signal is removed from the original audio signal, which can minimize the possibility of leaking user privacy, and after the target voiceprint feature information generates the target control instruction, the headphone function is controlled according to the target control instruction, thereby effectively improving the security of controlling the headphone function and further improving the user experience.

[0074] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2 , step S20 includes steps S201 to S203:

[0075] Step S201, when there is sensitive information in the audio component information, obtain the audio signal corresponding to the sensitive information.

[0076] It should be noted that when it is determined that there is sensitive information in the audio component information, it indicates that there is information related to user privacy in the original audio signal, and at this time, the audio signal corresponding to the sensitive information is obtained.

[0077] Step S202, determine the source of the audio signal corresponding to the sensitive information, and determine the target information privacy strategy according to the source.

[0078] It can be understood that the target information privacy strategy refers to the strategy for privacy processing of the audio signal corresponding to the sensitive information. Different sources of audio signals result in different determined target information privacy strategies.

[0079] Further, the step of determining the target information privacy strategy according to the source includes: when the source is user conversation content, determining the target information privacy strategy as a time-domain masking privacy strategy; when the source is an environmental background, determining the target information privacy strategy as a frequency-domain filtering privacy strategy.

[0080] It should be understood that when it is determined that the source is the user's conversation content, it indicates that muting processing is required. At this time, the target information privacy protection strategy used is the time-domain masking privacy protection strategy. When it is determined that the source is the environmental background, it indicates that filtering processing is required. At this time, the target information privacy protection strategy used is the frequency-domain filtering privacy protection strategy. When it is determined that the source includes the user's conversation content and the environmental background, the target information privacy protection strategy is determined to be the target information privacy protection strategy and the frequency-domain filtering privacy protection strategy.

[0081] Step S203: Perform privacy protection processing on the audio signal corresponding to the sensitive information according to the target information privacy protection strategy to obtain a privacy-protected audio signal.

[0082] It should be understood that when the target information privacy protection strategy is the time-domain masking privacy protection strategy, mute the audio signal corresponding to the sensitive information according to the time-domain masking privacy protection strategy. When the target information privacy protection strategy is the frequency-domain filtering privacy protection strategy, filter according to the frequency-domain filtering privacy protection strategy to achieve the purpose of privacy protection processing.

[0083] In this embodiment, when there is sensitive information in the audio component information, the audio signal corresponding to the sensitive information is obtained; the source of the audio signal corresponding to the sensitive information is determined, and the target information privacy protection strategy is determined according to the source; the audio signal corresponding to the sensitive information is subjected to privacy protection processing according to the target information privacy protection strategy to obtain a privacy-protected audio signal. In the above manner, when there is sensitive information in the audio component information, the source of the audio signal corresponding to the sensitive information is determined, and the target information privacy protection strategy for privacy protection processing is determined according to the source, and then the audio signal corresponding to the sensitive information is processed into a privacy-protected audio signal according to the target information privacy protection strategy, so as to effectively improve the accuracy of obtaining the privacy-protected audio signal, and further improve the efficiency of removing the privacy-protected audio signal from the original audio signal.

[0084] This application also provides a headphone function control device. Please refer to Figure 3 , the device includes:

[0085] Analysis module 10, configured to perform component analysis on the original audio signal based on the target multi-layer audio signal processing model to obtain audio component information.

[0086] Privacy protection processing module 20, configured to perform privacy protection processing on the audio signal corresponding to the sensitive information according to the target information privacy protection strategy to obtain a privacy-protected audio signal when there is sensitive information in the audio component information.

[0087] The elimination module 30 is used to eliminate the privatized audio signal from the original audio signal and generate target voiceprint feature information according to the original audio signal after elimination.

[0088] The control module 40 is used to generate a target control instruction according to the target voiceprint feature information and control the headphone function according to the target control instruction.

[0089] In this embodiment, the original audio signal is analyzed for its components based on the target multi-layer audio signal processing model to obtain audio component information. When there is sensitive information in the audio component information, the audio signal corresponding to the sensitive information is privatized according to the target information privatization strategy to obtain a privatized audio signal. The privatized audio signal is eliminated from the original audio signal, and target voiceprint feature information is generated according to the original audio signal after elimination. A target control instruction is generated according to the target voiceprint feature information, and the headphone function is controlled according to the target control instruction. In the above manner, when it is determined that there is sensitive information in the audio component information, the privatized audio signal is eliminated from the original audio signal, which can minimize the possibility of leaking user privacy. After the target control instruction is generated from the target voiceprint feature information, the headphone function is controlled according to the target control instruction, thereby effectively improving the security of controlling the headphone function and further enhancing the user experience.

[0090] The headphone function control device provided in this application adopts the headphone function control method in the above embodiment and can solve the technical problem of the low security of controlling the headphone function in the prior art. Compared with the prior art, the beneficial effects of the headphone function control device provided in this application are the same as those of the headphone function control method provided in the above embodiment, and the other technical features in the headphone function control device are the same as the features disclosed in the above embodiment method, which will not be elaborated here.

[0091] In one embodiment, the analysis module 10 is further used to collect the original audio signal based on the sound collection unit composed of a multi-microphone array; transmit the original audio signal to the signal processing unit through the target data transmission bus; denoise the original audio signal based on the signal processing unit and perform frame splitting on the denoised original audio signal to obtain a plurality of short-frame audio signals; perform component analysis on the plurality of short-frame audio signals respectively based on the target multi-layer audio signal processing model to obtain audio component information.

[0092] In one embodiment, the privacy processing module 20 is further configured to display a sensitive information setting interface when detecting a call to the privacy protection setting interface; obtain user input information and / or user selected information based on the sensitive information setting interface; perform sensitive detection on the user input information and / or user selected information, and generate sensitive information according to the sensitive detection result.

[0093] In one embodiment, when there is sensitive information in the audio component information, the privacy processing module 20 is further configured to obtain an audio signal corresponding to the sensitive information; determine the source of the audio signal corresponding to the sensitive information, and determine a target information privacy policy according to the source; perform privacy processing on the audio signal corresponding to the sensitive information according to the target information privacy policy to obtain a privacy processed audio signal.

[0094] In one embodiment, when the source is user conversation content, the privacy processing module 20 is further configured to determine that the target information privacy policy is a time domain masking privacy policy; when the source is an environmental background, determine that the target information privacy policy is a frequency domain filtering privacy policy.

[0095] In one embodiment, the control module 40 is further configured to obtain the standard voiceprint feature information of the user based on a pre-stored user voiceprint template; calculate the similarity between the target voiceprint feature information and the standard voiceprint feature information; when the similarity is greater than a preset similarity threshold, determine the speech content corresponding to the original audio signal after removal; generate a target control instruction according to the speech content, and control the headphone function according to the target control instruction.

[0096] In one embodiment, the control module 40 is further configured to generate information to be encrypted according to the target voiceprint feature information, the original audio signal after removal, and the operation record of the original audio signal; encrypt the information to be encrypted according to a target encryption policy to obtain target encrypted information; directly clear the target encrypted information.

[0097] This application provides a headphone, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the headphone function control method in the first embodiment above.

[0098] Next, refer to Figure 4, which shows a schematic structural diagram of a headset suitable for implementing the embodiments of the present application. The headset in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The shown headset is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0099] As Figure 4 shown, the headset may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the ROM (Read Only Memory) 1002 or the program loaded from the storage device 1003 into the RAM (Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the headset are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the headset to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a headset with various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.

[0100] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. The computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by a processing device 1001, the above functions defined in the methods of the disclosed embodiments of the present application are executed.

[0101] The earphone provided by the present application adopts the earphone function control method in the above embodiment, which can solve the technical problem of low security in controlling the functions of the earphone in the prior art. Compared with the prior art, the beneficial effects of the earphone provided by the present application are the same as those of the earphone function control method provided in the above embodiment, and other technical features in the earphone are the same as those disclosed in the method of the previous embodiment, which will not be elaborated here.

[0102] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0103] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0104] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the earphone function control method in the above embodiment.

[0105] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0106] The above computer-readable storage medium can be included in the earphone; or can exist separately without being assembled into the earphone.

[0107] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0108] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems and methods according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0109] The modules described in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0110] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned headphone function control method, which can solve the technical problem of low security in controlling the functions of headphones in the prior art. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the headphone function control method provided by the above embodiments, and will not be elaborated here.

[0111] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made by using the description of the present application and the content of the accompanying drawings under the technical concept of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A method for controlling the functions of an earphone, characterized in that, The method includes: Performing component analysis on the original audio signal based on a target multi-layer audio signal processing model to obtain audio component information; When there is sensitive information in the audio component information, performing privacy processing on the audio signal corresponding to the sensitive information according to a target information privacy policy to obtain a privacy-processed audio signal; Removing the privacy-processed audio signal from the original audio signal and generating target voiceprint feature information based on the original audio signal after removal; Generating a target control instruction according to the target voiceprint feature information and controlling the headphone function according to the target control instruction.

2. The method according to claim 1, characterized in that, The step of performing component analysis on the original audio signal based on a target multi-layer audio signal processing model to obtain audio component information includes: Collecting the original audio signal based on a sound collection unit composed of a multi-microphone array; Transmitting the original audio signal to a signal processing unit through a target data transmission bus; Performing denoising on the original audio signal based on the signal processing unit and performing frame splitting on the denoised original audio signal to obtain a number of short-frame audio signals; Performing component analysis on the number of short-frame audio signals respectively based on a target multi-layer audio signal processing model to obtain audio component information.

3. The method according to claim 1, characterized in that, The step of, when there is sensitive information in the audio component information, performing privacy processing on the audio signal corresponding to the sensitive information according to a target information privacy policy to obtain a privacy-processed audio signal includes: When there is sensitive information in the audio component information, obtaining the audio signal corresponding to the sensitive information; Determining the source of the audio signal corresponding to the sensitive information and determining a target information privacy policy according to the source; Performing privacy processing on the audio signal corresponding to the sensitive information according to the target information privacy policy to obtain a privacy-processed audio signal.

4. The method according to claim 3, wherein The step of determining a target information privacy policy according to the source includes: When the source is user conversation content, determining that the target information privacy policy is a time-domain masking privacy policy; When the source is an environmental background, determining that the target information privacy policy is a frequency-domain filtering privacy policy.

5. The method according to any one of claims 1 to 4, characterized in that, The step of generating a target control instruction according to the target voiceprint feature information and controlling the headphone function according to the target control instruction includes: Obtaining the standard voiceprint feature information of the user based on a pre-stored user voiceprint template; Calculating the similarity between the target voiceprint feature information and the standard voiceprint feature information; When the similarity is greater than a preset similarity threshold, determining the speech content corresponding to the original audio signal after removal; Generating a target control instruction according to the speech content and controlling the headphone function according to the target control instruction.

6. The method according to claim 5, characterized in that After the step of, when the similarity is greater than a preset similarity threshold, determining the speech content corresponding to the original audio signal after removal, it further includes: Generating information to be encrypted according to the target voiceprint feature information, the original audio signal after removal, and the operation record of the original audio signal; Encrypting the information to be encrypted according to a target encryption policy to obtain target encrypted information; Directly clearing the target encrypted information.

7. The method according to claim 1, wherein Before the step of, when there is sensitive information in the audio component information, performing privacy processing on the audio signal corresponding to the sensitive information according to the target information privacy policy to obtain a privacy-processed audio signal, further includes: When detecting a call to the privacy protection setting interface, display a sensitive information setting interface; Obtain user input information and / or user selected information based on the sensitive information setting interface; Perform sensitive detection on the user input information and / or user selected information, and generate sensitive information according to the sensitive detection result.

8. An earphone function control device, characterized in that, The device includes: An analysis module, configured to perform component analysis on the original audio signal based on a target multi-layer audio signal processing model to obtain audio component information; A privacy processing module, configured to, when there is sensitive information in the audio component information, perform privacy processing on the audio signal corresponding to the sensitive information according to the target information privacy policy to obtain a privacy-processed audio signal; An elimination module, configured to eliminate the privacy-processed audio signal from the original audio signal, and generate target voiceprint feature information according to the original audio signal after elimination; A control module, configured to generate a target control instruction according to the target voiceprint feature information, and control the headphone function according to the target control instruction.

9. A headset, characterized in that, The headphone includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the headphone function control method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the headphone function control method according to any one of claims 1 to 7 are implemented.