Voice processing method and device, earphone and program product

By using scene recognition and tone conversion technologies in VR devices to process voice signals collected by headphones, the shortcomings of existing VR devices in terms of voice privacy, security and personalized experience are solved, and a more advanced immersive experience is achieved.

CN119943069APending Publication Date: 2025-05-06SHENZHEN GRANDSUN ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411976646.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing VR devices have shortcomings in voice privacy, security and personalized experiences to meet users' advanced immersive experience needs.

Method used

By obtaining the original voice signal collected by the headphone microphone, performing scene recognition, determining the target voice sample, and performing tone conversion processing based on the target voice sample, outputting the processed voice signal.

Benefits of technology

It realizes the tone conversion of voice signals, enriches the processing methods of voice signals, and meets users' needs in voice privacy, security or personalized experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943069A_ABST
    Figure CN119943069A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of voice processing, and provides a voice processing method and device, an earphone and a program product, and the method comprises the steps: obtaining an original voice signal collected by a microphone of the earphone; scene recognition is carried out on the original voice signal, a target voice sample corresponding to the original voice signal is determined, and the target voice sample is a candidate voice sample corresponding to the scene of the original voice signal in a plurality of candidate voice samples; according to the target voice sample, tone conversion processing is carried out on the original voice signal to obtain a processed voice signal; and controlling a loudspeaker of the earphone to output the processed voice signal. According to the method, the tone conversion processing is performed on the original voice signal according to the target voice sample corresponding to the original voice signal, so that the processing modes of the voice signal are enriched, the tone conversion of the voice signal is realized, and the requirements of a user on the aspects of voice privacy, safety or personalized experience are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of speech processing technology, and in particular, relates to a speech processing method, device, earphone and program product. Background Art

[0002] With the rapid development of virtual reality (VR) technology, users' requirements for immersive experience are getting higher and higher.

[0003] Current VR devices mainly focus on improving visual effects, with relatively simple processing methods, and are unable to meet users' needs in terms of voice privacy, security, or personalized experience. Summary of the invention

[0004] The embodiments of the present application provide a voice processing method, device, headset and program product, which can enrich the processing method of voice signals and meet the user's needs in voice privacy, security or personalized experience.

[0005] In a first aspect, an embodiment of the present application provides a speech processing method, comprising:

[0006] Obtain the original voice signal collected by the microphone of the headset;

[0007] Performing scene recognition on the original speech signal to determine a target speech sample corresponding to the original speech signal, where the target speech sample is a candidate speech sample corresponding to the scene of the original speech signal among multiple candidate speech samples;

[0008] According to the target speech sample, the original speech signal is processed by timbre conversion to obtain a processed speech signal;

[0009] Control the speaker of the headset to output the processed voice signal.

[0010] In some embodiments, according to the target speech sample, the original speech signal is subjected to timbre conversion processing to obtain a processed speech signal, including:

[0011] Performing audio analysis on the original speech signal to obtain audio feature information of the original speech signal;

[0012] Perform feature recognition on audio feature information to obtain the emotional features and semantic features of the original speech signal;

[0013] The timbre features of the target speech sample are used to perform timbre conversion processing on the emotional features and semantic features to obtain a processed speech signal. The timbre features are used to characterize the timbre of the target speech sample.

[0014] In some embodiments, performing audio analysis on the original speech signal to obtain audio feature information of the original speech signal includes:

[0015] Extract features from the original speech signal to obtain an N-dimensional standard logarithmic Mel spectrum;

[0016] The N-dimensional standard logarithmic Mel spectrum is downsampled to obtain the audio feature information of the original speech signal.

[0017] In some embodiments, feature recognition is performed on the audio feature information to obtain the emotional features and semantic features of the original speech signal, including:

[0018] Perform semantic feature recognition on the audio feature information to obtain the semantic features of the original speech signal;

[0019] The self-attention network is used to perform emotional feature recognition on audio feature information to obtain the emotional characteristics of the original speech signal.

[0020] In some embodiments, performing scene recognition on the original speech signal to determine the target speech sample corresponding to the original speech signal includes:

[0021] Preprocessing the original speech signal to obtain a preprocessed speech signal;

[0022] Performing scene recognition on the preprocessed speech signal to obtain a speech scene of the preprocessed speech signal;

[0023] A candidate speech sample corresponding to the speech scene is selected from multiple candidate speech samples as a target speech sample corresponding to the original speech signal.

[0024] In some embodiments, preprocessing the original speech signal to obtain a preprocessed speech signal includes:

[0025] Performing signal enhancement on the original speech signal to obtain an enhanced speech signal;

[0026] The enhanced speech signal is denoised using a long short-term memory network to obtain a preprocessed speech signal.

[0027] In some embodiments, controlling a speaker of an earphone to output a processed voice signal includes:

[0028] Translating the processed speech signal based on a specified language to obtain a translated speech signal;

[0029] The microphone of the headset is controlled to output the translated voice signal.

[0030] In a second aspect, an embodiment of the present application provides a speech processing device, including:

[0031] An acquisition module, used to acquire the original voice signal collected by the microphone of the headset;

[0032] A scene recognition module is used to perform scene recognition on the original speech signal and determine a target speech sample corresponding to the original speech signal, where the target speech sample is a candidate speech sample corresponding to the scene of the original speech signal among multiple candidate speech samples;

[0033] The timbre conversion module is used to perform timbre conversion processing on the original voice signal according to the target voice sample to obtain a processed voice signal;

[0034] The output module is used to control the speaker of the headset to output the processed voice signal.

[0035] In a third aspect, an embodiment of the present application provides a headset, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any method of the first aspect when executing the computer program.

[0036] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method of any one of the first aspects is implemented.

[0037] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a headset, enables the headset to execute any one of the methods in the first aspect.

[0038] The embodiments of the present application provide a speech processing method, device, headset and program product, the method comprising: obtaining an original speech signal collected by a microphone of the headset; performing scene recognition on the original speech signal to determine a target speech sample corresponding to the original speech signal, the target speech sample being a candidate speech sample corresponding to the scene of the original speech signal among multiple candidate speech samples; performing timbre conversion processing on the original speech signal according to the target speech sample to obtain a processed speech signal; and controlling the speaker of the headset to output the processed speech signal. Utilizing the above technical solution, by performing timbre conversion processing on the original speech signal according to the target speech sample corresponding to the original speech signal, the processing method of the speech signal is enriched, and the timbre conversion of the speech signal is realized, thereby meeting the user's needs in terms of speech privacy, security or personalized experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0040] Figure 1 It is a flowchart of a speech processing method provided by an embodiment of the present application;

[0041] Figure 2 is a flow chart of a speech processing method provided by another embodiment of the present application;

[0042] Figure 3 is a structural block diagram of a speech processing device provided by an embodiment of the present application;

[0043] Figure 4 It is a structural schematic diagram of an earphone provided in one embodiment of the present application. DETAILED DESCRIPTION

[0044] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0045] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.

[0046] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0047] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.

[0048] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0049] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0050] It should be noted that the information collection process (such as face image collection process, voice signal collection process, etc.) / feature extraction process involved in this application is performed with the user's knowledge and permission, that is, the information collection process / feature extraction process complies with the requirements of laws and regulations and does not constitute an act that harms the public interest.

[0051] Figure 1 It is a flowchart of a voice processing method provided in an embodiment of the present application. As an example but not a limitation, the method can be applied to headphones.

[0052] S101. Acquire an original voice signal collected by a microphone of an earphone.

[0053] The original voice signal can be a voice signal originally collected by the microphone of the headset and has not been processed in any way. The source of the original voice signal is not limited. For example, it can be sound input by the headset Bluetooth, including sound input by a phone or sound input by the Bluetooth Audio Distribution Profile (A2DP), etc. It can also be sound input by a line in, or sound from the external environment, etc.

[0054] S102: Perform scene recognition on the original speech signal to determine a target speech sample corresponding to the original speech signal, where the target speech sample is a candidate speech sample corresponding to the scene of the original speech signal among multiple candidate speech samples.

[0055] The target voice sample can be a candidate voice sample corresponding to the scene of the original voice signal among multiple candidate voice samples; the multiple candidate voice samples can be voice samples pre-stored by the user, which can be used as a benchmark for timbre conversion and stored in the memory of the headset; the candidate voice sample can be a voice sample recorded by the voice acquisition module when the user uses the headset for the first time or during the use of the headset; the specific content of the candidate voice sample can be customized by the user according to actual needs to establish a personal timbre database, and each candidate voice sample has different timbre characteristics. Furthermore, different candidate voice samples can correspond to different scenes, or to different characters based on the scenes, so that users can choose different voice samples for conversion according to different scenes or preferences, providing a highly personalized sound experience.

[0056] After obtaining the original voice signal collected by the microphone, this step can perform scene recognition on the original voice signal to determine the target voice sample corresponding to the original voice signal. The specific scene recognition process can be directly performing scene recognition on the collected original voice signal, or the original voice signal can be first processed to perform certain operations, and then the pre-processed voice signal can be subjected to specific scene recognition. This embodiment does not limit this.

[0057] In some embodiments, performing scene recognition on the original speech signal to determine the target speech sample corresponding to the original speech signal includes:

[0058] Preprocessing the original speech signal to obtain a preprocessed speech signal;

[0059] Performing scene recognition on the preprocessed speech signal to obtain a speech scene of the preprocessed speech signal;

[0060] A candidate speech sample corresponding to the speech scene is selected from multiple candidate speech samples as a target speech sample corresponding to the original speech signal.

[0061] In a specific implementation, the original voice signal can be preprocessed first to enable more accurate subsequent scene recognition and timbre processing. The preprocessing means can be determined based on the actual situation of the original voice signal. Different voice sources can correspond to the same or different preprocessing means. For example, the preprocessing means may include howling suppression or echo elimination.

[0062] Subsequently, the preprocessed voice signal can be subjected to scene recognition, so that a candidate voice sample corresponding to the voice scene can be selected from a plurality of candidate voice samples as a target voice sample corresponding to the original voice signal. For example, a plurality of candidate voice samples can be stored according to different scenes, including candidate voice samples for game scenes, candidate voice samples for voice call scenes, and so on. In this embodiment, if the voice scene of the collected original voice signal is identified as a game scene, a candidate voice sample corresponding to the game scene can be selected from a plurality of candidate voice samples as a target voice sample corresponding to the original voice signal. Specifically, scene recognition can, for example, obtain the voice scene of the preprocessed voice signal with the aid of a preset recognition model, such as inputting the preprocessed voice signal into a preset recognition model to directly output the voice scene, and the preset recognition model can be a pre-trained neural network model; or the voice scene can be recognized by analyzing and calculating the preprocessed voice signal.

[0063] In some embodiments, preprocessing the original speech signal to obtain a preprocessed speech signal includes:

[0064] Performing signal enhancement on the original speech signal to obtain an enhanced speech signal;

[0065] The enhanced speech signal is denoised using a long short-term memory network to obtain a preprocessed speech signal.

[0066] In a specific implementation, the preprocessing process may include signal enhancement processing and noise reduction processing, wherein the signal enhancement process may include, for example, performing gain processing on the original speech signal to increase the signal amplitude of the original speech signal, thereby avoiding the situation where certain speech in the original speech signal is lost due to low gain, ensuring the integrity of the speech signal, and facilitating subsequent speech processing.

[0067] Furthermore, the environmental noise elimination technology can be combined to intelligently filter the background noise in the enhanced speech signal. Specifically, a deep learning end-to-end noise suppression network, such as a long short-term memory (LSTM) denoising network, can be used to perform noise reduction processing on the enhanced speech signal to obtain a noise-reduced speech signal. The specific noise reduction process can include: by identifying noise features, reducing the noise signal in the low signal-to-noise ratio area to leave a clear speech signal, and at the same time combining spectral subtraction to model the noise spectrum and subtract the noise component from the audio spectrum to retain the main speech frequency band.

[0068] Furthermore, this embodiment can also use adaptive noise cancellation (Adaptive Noise Control, ANC) at any stage of speech processing to integrate the characteristics of external sounds to make speech processing more natural and realistic, such as by collecting environmental noise data and generating a signal with opposite phase to achieve a noise cancellation effect.

[0069] S103: Perform timbre conversion processing on the original speech signal according to the target speech sample to obtain a processed speech signal.

[0070] After determining the target speech sample through the above steps, the target speech sample can be used to perform timbre conversion processing on the original speech signal. The specific method of timbre conversion is not limited. The original speech signal can be directly converted into the timbre of the target speech sample based on the timbre conversion model, and the timbre conversion model is a pre-trained neural network model; the original speech signal can also be feature analyzed, and the processed speech signal can be obtained by performing timbre processing on the features of the original speech signal, so that the timbre of the processed speech signal is consistent with the timbre of the target speech sample. The specific feature analysis content will not be further expanded here, as long as the timbre conversion processing of the original speech signal can be achieved.

[0071] S104: Control the speaker of the headset to output the processed voice signal.

[0072] This step can directly control the speaker of the headset to output the processed voice signal to the user, or can further combine other processing methods to enrich the processing of voice signals, such as supporting the conversion of external sounds in different languages ​​into specified timbres, achieving cross-language timbre unification, and facilitating users to learn languages ​​or communicate across cultures; or this embodiment can also automatically translate while converting the timbre, for example, translating the processed voice signal based on the specified language to obtain the translated voice signal; and controlling the microphone of the headset to output the translated voice signal. The specified language can be configured according to the actual needs of the user, and this embodiment does not limit this.

[0073] The present embodiment provides a speech processing method, which obtains an original speech signal collected by a microphone of an earphone; performs scene recognition on the original speech signal to determine a target speech sample corresponding to the original speech signal, wherein the target speech sample is a candidate speech sample corresponding to the scene of the original speech signal among multiple candidate speech samples; performs timbre conversion processing on the original speech signal according to the target speech sample to obtain a processed speech signal; and controls the speaker of the earphone to output the processed speech signal. By using this method, by performing timbre conversion processing on the original speech signal according to the target speech sample corresponding to the original speech signal, the processing method of the speech signal is enriched, and the timbre conversion of the speech signal is realized, thereby meeting the user's needs in terms of speech privacy, security or personalized experience.

[0074] Figure 2 This is a flow chart of a speech processing method provided by another embodiment of the present application. This embodiment will perform timbre conversion processing on the original speech signal according to the target speech sample, and the processed speech signal is further optimized as follows: perform audio analysis on the original speech signal to obtain the audio feature information of the original speech signal; perform feature recognition on the audio feature information to obtain the emotional features and semantic features of the original speech signal; use the timbre features of the target speech sample to perform timbre conversion processing on the emotional features and semantic features to obtain the processed speech signal, and the timbre features are used to characterize the timbre of the target speech sample. Figure 2 As shown, the method includes:

[0075] S201. Acquire an original voice signal collected by a microphone of an earphone.

[0076] S202: Perform scene recognition on the original speech signal to determine a target speech sample corresponding to the original speech signal, where the target speech sample is a candidate speech sample corresponding to the scene of the original speech signal among multiple candidate speech samples.

[0077] S203: Perform audio analysis on the original speech signal to obtain audio feature information of the original speech signal.

[0078] The audio feature information may be information related to the audio features in the original speech signal, such as the spectrum, intonation, and rhythm of the original speech signal. The specific audio analysis method may, for example, calculate certain audio parameters of the original speech signal to reflect the audio feature information of the original speech signal through the audio parameters, or extract features from the original speech signal to obtain an N-dimensional standard log-mel spectrum, and downsample the N-dimensional standard log-mel spectrum to obtain the audio feature information of the original speech signal. For example, the original speech signal may be converted into an N-dimensional log-mel filter bank feature, and an input feature matrix may be generated by stacking N-dimensional continuous frames and downsampling N times to use as the audio feature information of the original speech signal. Among them, the standard log-mel spectrum can be considered as a spectrum diagram obtained after converting the original speech signal into a Mel scale, which can simulate the perception characteristics of the human ear to sound, better characterize the characteristics of the audio signal, enhance the model's ability to capture audio features, and reduce the redundant information of the features, thereby improving the signal processing efficiency.

[0079] S204: Perform feature recognition on the audio feature information to obtain emotional features and semantic features of the original speech signal.

[0080] S205, using the timbre features of the target speech sample, performing timbre conversion processing on the emotional features and the semantic features to obtain a processed speech signal, wherein the timbre features are used to characterize the timbre of the target speech sample.

[0081] In this step, feature recognition can be performed on the obtained audio feature information. The content of feature recognition can include emotional features and semantic features, and can also include other related features. Different audio features can correspond to different feature recognition methods. For example, semantic feature recognition can be performed on the audio feature information to obtain the semantic features of the original speech signal. For example, the audio feature information can be converted into a bag-of-words matrix based on a bag-of-words model, and the theme or semantics of the audio feature information can be determined by counting the number of times each word appears. Alternatively, each word in the audio feature information can be mapped to a low-dimensional vector based on a word embedding model, and the relationship between each word can be determined by calculating the similarity, thereby analyzing the semantics of the audio feature information.

[0082] At the same time, this embodiment can use a self-attention network to perform emotional feature recognition on audio feature information to obtain the emotional characteristics of the original speech signal. On this basis, by adopting a non-autoregressive structure and a low-latency self-attention network, it can ensure real-time response of emotion recognition and meet the needs of real-time speech emotion recognition and conversion.

[0083] In a specific implementation, the timbre features of the target speech sample can be used to perform timbre conversion processing on the emotional features and semantic features. The timbre conversion process can map the emotional features and semantic features to the timbre features of the target speech sample to obtain a processed speech signal.

[0084] S206: Control the speaker of the headset to output the processed voice signal.

[0085] A speech processing method provided in this embodiment obtains the emotional features and semantic features of the original speech signal by performing feature recognition on the audio feature information of the original speech signal, which can provide an accurate feature basis for timbre conversion processing and enhance the realism of the processed speech signal.

[0086] In some embodiments, the user can set the timbre used in different voice scenarios through the application app. In addition to determining the target voice sample through scene recognition, this embodiment can also set a default target voice sample, such as setting a candidate voice sample among multiple candidate voice samples as the voice sample that is prioritized for subsequent timbre processing.

[0087] In some embodiments, the speech processing method provided in this embodiment can determine different target speech samples according to the number of characters in the original speech signal. For example, when there is only one person, the latest uploaded timbre can be used. When there are multiple people, different timbres can be automatically used to distinguish different people, that is, different candidate speech samples can be selected for subsequent timbre processing.

[0088] In some embodiments, the speech processing method provided in this embodiment can also perform multi-task embedded input to accurately realize feature recognition, such as pre-setting specific task embedding, including for emotion recognition tasks, so that the speech processing method of this embodiment will focus on emotion feature recognition; further, it can include for language recognition tasks, so that the speech processing method of this embodiment will focus on semantic feature recognition. This embodiment can choose whether to enable the corresponding specific task according to actual application needs.

[0089] In some embodiments, the headset involved in this embodiment can reduce the delay in the timbre conversion process by optimizing the hardware and software architecture, ensure the real-time and synchronization of the user receiving the converted sound, and improve the user experience. For example, a high-performance chip using digital signal processing technology (Digital Signal Processing, DSP) can be used to specifically handle voice signal analysis and noise reduction tasks, and a chip with an embedded neural network processor (Neural-network Process Units, NPU) can be used to run the timbre conversion processing algorithm on it, thereby achieving low-latency and real-time timbre conversion.

[0090] In some embodiments, the headphones involved in this embodiment may provide an open application programming interface (API) interface, allowing third-party applications or devices to access and expand the application scope of the tone conversion function, such as integration with smart home devices, vehicle systems, etc., to achieve wider applications.

[0091] In some embodiments, during the collection, storage and processing of voice signals, advanced encryption technology can be used to protect the user's voice samples and converted voice signals, and ensure personal privacy and data security. For example, in the processing stage of voice data, homomorphic encryption technology can be introduced so that the encrypted data can be directly calculated without decryption. Alternatively, the voice processing method of this embodiment can be completed by the headset and the server together. For example, the headset sends the collected original voice signal to the server for timbre conversion processing, and receives the processed voice signal returned by the server. Through homomorphic encryption technology, it can be ensured that the server does not need to decrypt data when processing voice recognition and sentiment analysis, thereby protecting the privacy of the user's voice data during the analysis process, greatly improving the level of privacy protection.

[0092] The following is an exemplary description of a speech processing method provided in this embodiment:

[0093] First, the voice acquisition module of the headset can record and store the user's voice samples, and then the sound input module of the headset can continuously monitor and receive the input sound signals.

[0094] Secondly, the timbre conversion module of the headset can use speech synthesis and timbre conversion algorithms to convert external input sounds into the timbre of user voice samples, including noise reduction and signal enhancement of the received sound signals to ensure sound quality, as well as analysis of the spectrum, intonation, rhythm and other characteristics of external sounds. Using the timbre conversion algorithm, the characteristics of external sounds are mapped to the timbre characteristics of the user's voice samples, thereby generating a sound signal that matches the user's timbre.

[0095] Finally, the synthesized sound can be transmitted to the user through the sound output module, so that all external sounds are presented in the user's own timbre.

[0096] Through the collaborative work of the above modules in the headset, by recording a user's voice, all external sounds can be converted into the timbre of the input voice, providing users with a more personalized and immersive sound experience. Exemplarily, this embodiment can be applied to telephone or voice chat, converting the other party's voice into the user's timbre to protect the privacy of communication; it can also be applied to multiplayer game scenarios, and the voices of all teammates can be played with the timbre configured by the user, enhancing the sense of immersion in the game; it can also be applied to language learning or music learning, such as converting foreign language pronunciation into the user's timbre and translating it to help users better understand and imitate, or converting all sounds into their own timbre, so that they can better understand their own voices for easier pronunciation, etc.

[0097] From the above description, it can be found that the voice processing method provided in this embodiment can adaptively match and convert the timbre of the user according to the spectral characteristics of the external sound by introducing an advanced timbre conversion algorithm, ensuring that the converted sound is consistent with the original sound in intonation, rhythm and emotion, thereby enhancing the immersion of virtual reality and enhancing the personalized experience. At the same time, in the timbre conversion process, the recognition of the emotion of the original sound is added, and the corresponding emotional characteristics are maintained and presented in the converted sound, thereby enhancing the sense of reality of the communication; it can also support multiple sound input sources, including Bluetooth, telephone and ambient sound, with strong compatibility; in addition, the use of an efficient timbre conversion algorithm ensures the real-time and smoothness of the sound conversion, and ensures the real-time processing of the voice signal.

[0098] Corresponding to the speech processing method of the above embodiment, Figure 3This is a structural block diagram of a speech processing device provided in one embodiment of the present application. For the sake of convenience of explanation, only the parts related to the embodiment of the present application are shown.

[0099] Reference Figure 3 , the device comprises:

[0100] An acquisition module 301 is used to acquire an original voice signal collected by a microphone of the headset;

[0101] A scene recognition module 302 is used to perform scene recognition on the original speech signal and determine a target speech sample corresponding to the original speech signal, where the target speech sample is a candidate speech sample corresponding to the scene of the original speech signal among multiple candidate speech samples;

[0102] The timbre conversion module 303 is used to perform timbre conversion processing on the original voice signal according to the target voice sample to obtain a processed voice signal;

[0103] The output module 304 is used to control the speaker of the headset to output the processed voice signal.

[0104] The present embodiment provides a speech processing device, which acquires the original speech signal collected by the microphone of the headset through an acquisition module; performs scene recognition on the original speech signal through a scene recognition module to determine the target speech sample corresponding to the original speech signal, and the target speech sample is a candidate speech sample corresponding to the scene of the original speech signal among multiple candidate speech samples; performs timbre conversion processing on the original speech signal according to the target speech sample through a timbre conversion module to obtain a processed speech signal; and controls the speaker of the headset to output the processed speech signal through an output module. By using this device, the original speech signal is subjected to timbre conversion processing according to the target speech sample corresponding to the original speech signal, which enriches the processing method of the speech signal and realizes the timbre conversion of the speech signal, thereby meeting the user's needs in speech privacy, security or personalized experience.

[0105] Optionally, the timbre conversion module includes:

[0106] An audio analysis unit, used to perform audio analysis on the original speech signal to obtain audio feature information of the original speech signal;

[0107] A feature recognition unit is used to perform feature recognition on audio feature information to obtain emotional features and semantic features of the original speech signal;

[0108] The timbre conversion unit is used to adopt the timbre features of the target speech sample, perform timbre conversion processing on the emotional features and semantic features, and obtain a processed speech signal. The timbre features are used to characterize the timbre of the target speech sample.

[0109] Optionally, the audio analysis unit is specifically used for:

[0110] Extract features from the original speech signal to obtain an N-dimensional standard logarithmic Mel spectrum;

[0111] The N-dimensional standard logarithmic Mel spectrum is downsampled to obtain the audio feature information of the original speech signal.

[0112] Optionally, the feature recognition unit is specifically used for:

[0113] Perform semantic feature recognition on the audio feature information to obtain the semantic features of the original speech signal;

[0114] The self-attention network is used to perform emotional feature recognition on audio feature information to obtain the emotional characteristics of the original speech signal.

[0115] Optionally, the scene recognition module includes:

[0116] A preprocessing unit, used for preprocessing the original speech signal to obtain a preprocessed speech signal;

[0117] A scene recognition unit, used to perform scene recognition on the preprocessed speech signal to obtain a speech scene of the preprocessed speech signal;

[0118] The selection unit is used to select a candidate speech sample corresponding to the speech scene from a plurality of candidate speech samples as a target speech sample corresponding to the original speech signal.

[0119] Optionally, the preprocessing unit is specifically used for:

[0120] Performing signal enhancement on the original speech signal to obtain an enhanced speech signal;

[0121] The enhanced speech signal is denoised using a long short-term memory network to obtain a preprocessed speech signal.

[0122] Optionally, the output module is specifically used to:

[0123] Translating the processed speech signal based on a specified language to obtain a translated speech signal;

[0124] The microphone of the headset is controlled to output the translated voice signal.

[0125] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0126] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0127] The embodiment of the present application also provides a headset, Figure 4 is a schematic diagram of the structure of an earphone provided by an embodiment of the present application, such as Figure 4 As shown, the headset includes: at least one processor 401, a memory 402, an input device 403, an output device 404, and a computer program stored in the memory 402 and executable on at least one processor 401. When the processor 401 executes the computer program, the steps in any of the above method embodiments are implemented.

[0128] The input device 403 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the headset. The output device 404 may include a display device such as a display screen.

[0129] The embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by the processor 401, the steps in the above-mentioned method embodiments can be implemented.

[0130] An embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0131] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 401, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the device / headphone, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electric carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0132] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0133] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0134] In the embodiments provided in the present application, it should be understood that the disclosed devices / earphones and methods can be implemented in other ways. For example, the device / earphone embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0135] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0136] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A speech processing method, characterized in that: include: Obtain the original voice signal collected by the microphone of the headset; Performing scene recognition on the original speech signal to determine a target speech sample corresponding to the original speech signal, wherein the target speech sample is a candidate speech sample corresponding to the scene of the original speech signal among a plurality of candidate speech samples; According to the target speech sample, performing timbre conversion processing on the original speech signal to obtain a processed speech signal; The speaker of the earphone is controlled to output the processed voice signal.

2. The speech processing method according to claim 1, characterized in that: The step of performing timbre conversion processing on the original voice signal according to the target voice sample to obtain a processed voice signal includes: Performing audio analysis on the original speech signal to obtain audio feature information of the original speech signal; Performing feature recognition on the audio feature information to obtain emotional features and semantic features of the original speech signal; The timbre features of the target speech sample are used to perform timbre conversion processing on the emotional features and the semantic features to obtain the processed speech signal, wherein the timbre features are used to characterize the timbre of the target speech sample.

3. The speech processing method according to claim 2, characterized in that: The performing audio analysis on the original speech signal to obtain audio feature information of the original speech signal includes: Extracting features of the original speech signal to obtain an N-dimensional standard logarithmic Mel spectrum; Downsampling is performed on the N-dimensional standard logarithmic Mel spectrum to obtain audio feature information of the original speech signal.

4. The speech processing method according to claim 2, characterized in that: The step of performing feature recognition on the audio feature information to obtain the emotional features and semantic features of the original speech signal includes: Performing semantic feature recognition on the audio feature information to obtain semantic features of the original speech signal; A self-attention network is used to perform emotional feature recognition on the audio feature information to obtain the emotional features of the original speech signal.

5. The speech processing method according to claim 1, characterized in that: The performing scene recognition on the original speech signal to determine the target speech sample corresponding to the original speech signal includes: Preprocessing the original speech signal to obtain a preprocessed speech signal; Performing scene recognition on the preprocessed speech signal to obtain a speech scene of the preprocessed speech signal; A candidate speech sample corresponding to the speech scene is selected from a plurality of candidate speech samples as a target speech sample corresponding to the original speech signal.

6. The speech processing method according to claim 5, characterized in that: The preprocessing of the original speech signal to obtain a preprocessed speech signal includes: Performing signal enhancement on the original speech signal to obtain an enhanced speech signal; The enhanced speech signal is subjected to noise reduction processing by using a long short-term memory network to obtain a preprocessed speech signal.

7. The speech processing method according to claim 1, characterized in that: The controlling the speaker of the headset to output the processed voice signal comprises: Translating the processed speech signal based on a specified language to obtain a translated speech signal; The microphone of the earphone is controlled to output the translated voice signal.

8. A speech processing device, characterized in that: include: An acquisition module, used to acquire the original voice signal collected by the microphone of the headset; A scene recognition module, used to perform scene recognition on the original speech signal and determine a target speech sample corresponding to the original speech signal, wherein the target speech sample is a candidate speech sample corresponding to the scene of the original speech signal among multiple candidate speech samples; A timbre conversion module, used to perform timbre conversion processing on the original voice signal according to the target voice sample to obtain a processed voice signal; The output module is used to control the speaker of the earphone to output the processed voice signal.

9. A headset comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the headset implements the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, enables the method according to any one of claims 1 to 7 to be performed.