Voice processing method and device and electronic equipment

By introducing an accent replacement module into the accent conversion model, and using accent encoder and multiple accent decoders for accent conversion, the problem of low accent conversion accuracy caused by tone leakage in the prior art is solved, and efficient and accurate multi-accent conversion is achieved.

CN120108409APending Publication Date: 2025-06-06FACE CUTE CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311659609.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-05
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The accent conversion model in the prior art has a problem of timbre leakage, resulting in poor accuracy of accent conversion.

Method used

The accent conversion model containing an accent replacement module is adopted. This module converts the accent information in the voice information except tone information through an accent encoder and multiple accent decoders to realize the conversion of multiple accents to multiple accents.

Benefits of technology

It effectively avoids tone leakage, improves the accuracy and effect of accent conversion, and supports "one-to-one", "many-to-one" and "many-to-many" conversions between multiple accents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108409A_ABST
    Figure CN120108409A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voice processing method and device and electronic equipment. The method comprises the following steps: acquiring a first voice of a first accent; acquiring an identifier of the second accent; based on an accent conversion model and the identifiers of the first voice and the second accent, second voice is determined, the voice content of the second voice is the same as the voice content of the first voice, and the accent of the second voice is the second accent; wherein the accent conversion model comprises an accent replacement module, the accent replacement module is used for converting accent information in other voice information except timbre information, and the accent replacement module comprises an accent encoder and a plurality of accent decoders connected with the accent encoder; and the plurality of accent decoders are in one-to-one correspondence with a plurality of accent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of speech processing technology, and in particular to a speech processing method, device and electronic device. Background Art

[0002] Accent conversion refers to converting the accent in a speech segment into another accent, but the speech content of the speech segment remains unchanged. For example, an electronic device can convert the Mandarin "hello" into the Sichuan dialect "hello".

[0003] Currently, electronic devices can process speech based on accent conversion models and then convert the accent of speech. For example, a trained accent conversion model can convert speech with one accent into speech with another accent. However, the accent conversion model has the problem of timbre leakage and the accuracy of accent conversion is poor. Summary of the invention

[0004] The present disclosure provides a speech processing method, device and electronic device, which are used to solve one or more technical problems in the prior art.

[0005] In a first aspect, the present disclosure provides a speech processing method, the speech processing method comprising:

[0006] Obtain a first speech with a first accent;

[0007] Get the identity of the second accent;

[0008] Determine a second voice based on the accent conversion model, the first voice, and the identifier of the second accent, wherein the voice content of the second voice is the same as the voice content of the first voice, and the accent of the second voice is the second accent;

[0009] Among them, the accent conversion model includes an accent replacement module, which is used to convert accent information in other speech information except timbre information. The accent replacement module includes an accent encoder and multiple accent decoders connected to the accent encoder, and the multiple accent decoders correspond one-to-one to multiple accents.

[0010] In a second aspect, the present disclosure provides a speech processing device, the speech processing device comprising a first acquisition module, a second acquisition module and a determination module, wherein:

[0011] The first acquisition module is used to acquire a first speech with a first accent;

[0012] The second acquisition module is used to acquire an identifier of a second accent;

[0013] The determination module is used to determine the second voice based on the accent conversion model, the first voice and the identification of the second accent, the voice content of the second voice is the same as the voice content of the first voice, and the accent of the second voice is the second accent; wherein the accent conversion model includes an accent replacement module, the accent replacement module is used to convert accent information in other voice information except timbre information, the accent replacement module includes an accent encoder and multiple accent decoders connected to the accent encoder, and the multiple accent decoders correspond one-to-one to multiple accents.

[0014] In a third aspect, an embodiment of the present disclosure provides an electronic device including: a processor and a memory;

[0015] The memory stores computer-executable instructions;

[0016] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the speech processing method as described in the first aspect and various possible aspects of the first aspect.

[0017] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer execution instructions are stored. When a processor executes the computer execution instructions, the speech processing method as described in the first aspect and various possible aspects of the first aspect are implemented.

[0018] The present disclosure provides a speech processing method, device and electronic device. The electronic device can obtain a first speech with a first accent, obtain an identifier of a second accent, and determine a second speech based on an accent conversion model, the first speech and the identifier of the second accent, wherein the speech content of the second speech is the same as the speech content of the first speech, and the accent of the second speech is the second accent. The accent conversion model includes an accent replacement module, which is used to convert accent information in other speech information except timbre information. The accent replacement module includes an accent encoder and multiple accent decoders connected to the accent encoder, and the multiple accent decoders correspond one-to-one to multiple accents. In the above method, since multiple accent decoders correspond one-to-one to multiple accents, the accent conversion model can train samples of multiple accents in parallel, and the accent conversion model is not limited to converting one accent to another accent, but can achieve "one-to-one" accent conversion, "many-to-one" accent conversion and "many-to-many" accent conversion, thereby reducing the cost of accent conversion. In addition, since the accent replacement module can convert accent information in other voice information except timbre information, the problem of timbre leakage can be avoided, the timbre of the original voice can be retained during the accent conversion process, the accuracy of accent conversion is improved, and the effect of accent conversion is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0020] Figure 1 A schematic diagram of an application scenario provided by an embodiment of the present disclosure;

[0021] Figure 2 A flowchart of a speech processing method provided by an embodiment of the present disclosure;

[0022] Figure 3 A schematic diagram of a process for obtaining a first voice provided in an embodiment of the present disclosure;

[0023] Figure 4 A schematic diagram of another process of obtaining a first voice provided by an embodiment of the present disclosure;

[0024] Figure 5 A schematic diagram of a process for acquiring a second accent provided by an embodiment of the present disclosure;

[0025] Figure 6 A schematic diagram of an accent replacement module provided in an embodiment of the present disclosure;

[0026] Figure 7 A structural schematic diagram of an accent conversion model provided by an embodiment of the present disclosure;

[0027] Figure 8 A schematic diagram of a method for determining a second voice provided in an embodiment of the present disclosure;

[0028] Fig. 9 A schematic diagram of a process for determining a first accent feature provided by an embodiment of the present disclosure;

[0029] Fig.10 A schematic diagram of a process for determining a second accent feature provided by an embodiment of the present disclosure;

[0030] Fig.11 A schematic diagram of a processing process of a diffusion module provided in an embodiment of the present disclosure;

[0031] Fig.12 A schematic diagram of a method for voice style conversion provided by an embodiment of the present disclosure;

[0032] Fig.13 A process diagram of a speech processing method provided by an embodiment of the present disclosure;

[0033] Fig.14 A schematic diagram of the structure of a speech processing device provided by an embodiment of the present disclosure; and

[0034] Fig.15 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0035] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0036] To facilitate understanding, the concepts involved in the embodiments of the present disclosure are explained below.

[0037] Electronic device: is a device with wireless transceiver function. Electronic devices can be deployed on land, including indoors or outdoors, handheld, wearable or vehicle-mounted. The electronic device can be a mobile phone, a tablet computer, a computer with wireless transceiver function, a virtual reality (VR) electronic device, an augmented reality (AR) electronic device, a wireless terminal in industrial control, a vehicle-mounted electronic device, a wireless terminal in self-driving, a wireless electronic device in remote medical care, a wireless electronic device in smart grid, a wireless electronic device in transportation safety, a wireless electronic device in smart city, a wireless electronic device in smart home, a wearable electronic device, etc. The electronic device involved in the embodiments of the present disclosure can also be called a terminal, a user equipment (UE), an access electronic device, a vehicle-mounted terminal, an industrial control terminal, a UE unit, a UE station, a mobile station, a mobile station, a remote station, a remote electronic device, a mobile device, a UE electronic device, a wireless communication device, a UE agent or a UE device, etc. The electronic device can also be fixed or mobile.

[0038] Next, combine Figure 1 , the application scenarios of the embodiments of the present disclosure are described.

[0039] Figure 1A schematic diagram of an application scenario provided by an embodiment of the present disclosure. Figure 1 , including an electronic device. Among them, an accent conversion application can be installed in the electronic device. The electronic device can obtain Mandarin speech based on the application, and based on the application, convert the Mandarin speech into Shaanxi dialect speech, wherein the speech content in the Shaanxi dialect speech is the same as the speech content in the Mandarin speech, and the speech style information such as stress, rhythm, etc. in the Shaanxi dialect speech is also the same as the speech style information such as stress, rhythm, etc. in the Mandarin speech. In this way, the electronic device can convert the accent in the speech to obtain speech with other accents, and the content and style of the speech after the accent conversion are the same as the content and style of the speech before the accent conversion.

[0040] It should be noted that Figure 1 The above is only an example of the application scenario of the embodiment of the present disclosure, and is not intended to limit the application scenario of the embodiment of the present disclosure.

[0041] In the related art, accent conversion refers to converting the accent in a speech segment into another accent, but the speech content of the speech segment remains unchanged. For example, an electronic device can convert "hello" in Mandarin into "hello" in Shaanxi dialect. At present, electronic devices can process speech based on an accent conversion model, and then convert the accent of the speech. For example, a trained accent conversion model can convert speech with one accent into speech with another accent, while keeping the speech content unchanged. However, the training samples of the accent conversion model are trained based on the mel-spectrogram of the speech, and the timbre information in the speech information will not be separated. In this way, the accent conversion model will have the problem of timbre leakage in the process of accent conversion of speech, the speech effect is poor, and the accuracy of accent conversion is poor.

[0042] In order to solve the technical problems in the related art, the embodiments of the present disclosure provide a speech processing method, in which an electronic device can obtain a first speech with a first accent, obtain an identification of a second accent, perform convolution processing on the first speech based on a timbre separation module in an accent conversion model to obtain a timbre feature of the first speech, perform convolution processing on the first speech based on an automatic speech recognition module in the accent conversion model to obtain a first accent feature of the first speech, perform convolution processing on the first accent feature based on an accent replacement module in the accent conversion model and the identification of the second accent to obtain a second accent feature, and determine a second speech based on the timbre feature, the second accent feature and the accent conversion model, wherein the speech content of the second speech is the same as the speech content of the first speech, and the accent of the second speech is the second accent. In this way, since the accent replacement module includes an accent encoder and multiple accent decoders connected to the accent encoder, and the multiple accent decoders correspond one-to-one to multiple accents, the accent conversion model can realize multi-accent to multi-accent accent conversion processing, and since the automatic speech recognition module can separate the timbre information in the first speech, the remaining accent information and voice style information, the accent replacement module can accurately convert the accent information, thereby avoiding the problem of timbre leakage and improving the accuracy of accent conversion.

[0043] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.

[0044] Figure 2 A flow chart of a speech processing method provided by an embodiment of the present disclosure. Figure 2 , the method may include:

[0045] S201: Acquire a first speech with a first accent.

[0046] The execution subject of the embodiments of the present disclosure may be an electronic device, or a voice processing device provided in the electronic device. The voice processing device may be implemented based on software, or the voice processing device may be implemented based on a combination of software and hardware, which is not limited in the embodiments of the present disclosure. The electronic device may be any device with terminal computing capability. For example, the electronic device may be a computer, a server, a mobile device, etc., which is not limited in the embodiments of the present disclosure.

[0047] The first voice may be a voice to be processed. For example, the first voice may be a voice to be processed with an accent conversion. For example, if the electronic device performs an accent conversion process on voice 1, the electronic device may determine that voice 1 is the first voice, and if the electronic device performs an accent conversion process on voice 2, the electronic device may determine that voice 2 is the first voice.

[0048] Optionally, the electronic device may receive the first voice sent by other devices, the electronic device may record the first voice in real time, or the electronic device may obtain the first voice from a database, which is not limited in the embodiments of the present disclosure.

[0049] The first accent may be the accent of the first voice. For example, if the accent of the first voice is Mandarin, the first accent may be Mandarin; if the accent of the first voice is Shaanxi dialect, the first accent may be Shaanxi dialect. It should be noted that the above accents are only examples and are not intended to limit the accents in the embodiments of the present disclosure.

[0050] It should be noted that after the electronic device acquires the first voice, it can determine the first accent of the first voice based on any feasible implementation method, and the embodiments of the present disclosure are not limited to this.

[0051] Optionally, the electronic device may obtain the first speech with the first accent based on the following two feasible implementations:

[0052] A possible implementation method:

[0053] A voice acquisition page is displayed, wherein the voice acquisition page includes a voice acquisition control. In response to a touch operation on the voice acquisition control, a voice page is displayed, wherein the voice page includes multiple voices. In response to a touch operation on the multiple voices, a voice associated with the touch operation is determined as a first voice.

[0054] Next, combine Figure 3 , in this feasible implementation method, a process of the electronic device acquiring the first voice is described.

[0055] Figure 3 A schematic diagram of a process for obtaining a first voice provided by an embodiment of the present disclosure. Figure 3 , including: an electronic device. The display page of the electronic device includes page 301 and page 302. Page 301 may be a voice acquisition page, and the voice acquisition page includes a voice acquisition control. When the user clicks the voice acquisition control, the display page of the electronic device jumps from page 301 to page 302.

[0056] See also Figure 3, page 302 may be a voice page. The voice page may include voice A, voice B, voice C, and a confirmation control (it should be noted that the accents and voice styles of voice A, voice B, and voice C may be different, and the embodiments of the present disclosure are not limited to this). When the user clicks voice A, voice A may be highlighted, and after the user clicks the confirmation control, the electronic device may determine voice A as the first voice. In this way, the user can flexibly select the first voice, thereby improving the flexibility of obtaining the first voice.

[0057] Another possible implementation:

[0058] A voice collection page is displayed, wherein the voice collection page includes a voice recording control, and in response to a touch operation on the voice recording control, voice recording is performed to obtain a first voice.

[0059] Next, combine Figure 4 , in this feasible implementation method, a process of the electronic device acquiring the first voice is described.

[0060] Figure 4 Another schematic diagram of a process for obtaining a first voice provided by an embodiment of the present disclosure. Figure 4 , including: an electronic device. Among them, the display page of the electronic device may include page 401 and page 402. Page 401 may be a voice collection page, and the voice collection page includes a voice recording control. After the user clicks the voice recording control, the display page of the electronic device jumps from page 401 to page 402. Page 402 may be a recording page, and the progress of the voice recording may be displayed in the recording page, and the recording page may include a confirmation control. After the voice recording is completed, the user may click the confirmation control, and the electronic device may determine the recorded voice as the first voice. In this way, the user can record the first voice in real time, improve the flexibility of obtaining the first voice, and improve the user experience.

[0061] S202: Obtain an identifier of a second accent.

[0062] The second accent may be a target accent for the accent conversion. For example, if the electronic device converts Mandarin speech into Shaanxi dialect speech, the target accent may be Shaanxi dialect, that is, the second accent may be Shaanxi dialect; if the electronic device converts Mandarin speech into Sichuan dialect speech, the target accent may be Sichuan dialect, that is, the second accent may be Sichuan dialect.

[0063] Optionally, the electronic device may obtain the identification of the second accent based on the following feasible implementation method: displaying a voice acquisition page, which may include a voice acquisition control and multiple accents, and in response to a user's touch operation on the multiple accents, determining the accent associated with the touch operation as the second accent.

[0064] It should be noted that the identifier of the second accent may be the name of the second accent, the unique mark of the second accent, etc., which is not limited in the embodiments of the present disclosure. Moreover, after the electronic device determines the second accent, it may obtain the identifier of the second accent based on any feasible implementation method, which is not limited in the embodiments of the present disclosure.

[0065] Next, combine Figure 5 , the process of acquiring the second accent by the electronic device is explained.

[0066] Figure 5 A schematic diagram of a process for obtaining a second accent provided by an embodiment of the present disclosure. Figure 5 , including: an electronic device. The display page of the electronic device may be a voice acquisition page. The voice acquisition page may include accent A, accent B, accent C, and a voice acquisition control. When the user clicks accent A, accent A may be highlighted, and the electronic device may determine accent A as the second accent.

[0067] S203: Determine the second voice based on the accent conversion model, the first voice, and the identifier of the second accent.

[0068] The speech content of the second speech is the same as the speech content of the first speech, and the accent of the second speech is the second accent. For example, the first speech may be "hello" in accent A, and the electronic device may obtain the second speech after performing accent conversion processing on the first speech, and the second speech may be "hello" in accent B. For example, the first speech may be "the weather is great today" in Shaanxi dialect, and the electronic device may obtain the second speech after performing accent conversion processing on the first speech, and the second speech may be "the weather is great today" in Mandarin.

[0069] The accent conversion model includes an accent replacement module. Optionally, the accent replacement module is used to convert the accent information in other voice information except the timbre information. For example, the voice information may include timbre information, accent information and voice style information, wherein the voice style information may include the prosody information, rhythm information, stress information, etc. of the voice, and the embodiment of the present disclosure is not limited to this. Other voice information except the timbre information may include accent information and voice style information, and the accent replacement module may convert the accent information.

[0070] Optionally, the accent replacement module includes an accent encoder and a plurality of accent decoders connected to the accent encoder, wherein the accent encoder can be used to encode speech information except timbre information, and the interpretation decoder can be used to decode information output by the accent encoder.

[0071] Among them, multiple accent decoders correspond one-to-one to multiple accents. For example, accent decoder A can correspond to accent 1, accent decoder B can correspond to accent 2, and accent decoder C can correspond to accent 3. For example, if accent decoder A corresponds to accent 1, then the accent conversion model can generate the speech with accent 1 based on the information output by accent decoder A; if accent decoder B corresponds to accent 2, then the accent conversion model can generate the speech with accent 2 based on the information output by accent decoder B. In this way, the accent conversion model can realize the scenario of converting multiple accents to multiple accents and reduce the cost of accent conversion. For example, since the multiple accent encoders in the accent conversion model correspond one-to-one to multiple accents, the accent conversion model can convert the speech of Shaanxi dialect into the speech of Sichuan dialect (one-to-one accent conversion), and the accent conversion model can also convert the speech of Shaanxi dialect into the speech of Sichuan dialect or Cantonese dialect (one-to-many accent conversion), and the accent conversion model can also convert the speech of Shaanxi dialect into the speech of Sichuan dialect or Cantonese dialect, and convert the speech of Northeastern dialect into the speech of Sichuan dialect or Cantonese dialect (many-to-many accent conversion).

[0072] Next, combine Figure 6 , the structure of the accent replacement module is explained.

[0073] Figure 6 A schematic diagram of an accent replacement module provided in an embodiment of the present disclosure. Figure 6 The accent replacement module may include an accent encoder, an accent decoder 1, an accent decoder 2, ..., and an accent decoder n. The accent encoder is connected to the accent decoder 1, the accent decoder 2, ..., and the accent decoder n respectively, and each accent decoder may correspond to an accent.

[0074] It should be noted that the structure of the accent conversion module can be a BN-to-BN (BN2BN) neural network structure or other neural network structures, which is not limited in the embodiments of the present disclosure.

[0075] Optionally, the accent conversion model may include a timbre separation module, an automatic speech recognition module and a diffusion module. The timbre separation module is used to obtain timbre information in the first speech, the automatic speech recognition module is used to obtain other speech information in the first speech except the timbre information, and the diffusion module is used to generate the second speech.

[0076] Next, combine Figure 7 , the structure of the accent transfer model is explained.

[0077] Figure 7 This is a schematic diagram of the structure of an accent conversion model provided by an embodiment of the present disclosure. Figure 7, including: an accent conversion model. The accent conversion model may include a timbre separation module, an automatic speech recognition module, a timbre replacement module and a diffusion module. The timbre separation module may be connected to the diffusion module, the automatic speech recognition module may be connected to the timbre replacement module, and the timbre replacement module may also be connected to the diffusion module.

[0078] It should be noted that the network structure of the timbre classification module in the embodiment of the present disclosure may be any neural network structure capable of timbre separation, and the embodiment of the present disclosure is not limited to this.

[0079] It should be noted that the network structure of the automatic speech recognition module in the embodiment of the present disclosure may be a network structure built based on automatic speech recognition (ASR) technology, or may be any other feasible neural network structure, and the embodiment of the present disclosure is not limited to this.

[0080] It should be noted that the network structure of the diffusion module in the embodiment of the present disclosure may be the network structure of a conditional diffusion model, or may be any other feasible neural network structure, and the embodiment of the present disclosure is not limited to this.

[0081] Among them, the electronic device can determine the second voice based on the following feasible implementation methods: based on the timbre separation module in the accent conversion model, convolution processing is performed on the first voice to obtain the timbre characteristics of the first voice, based on the automatic speech recognition module in the accent conversion model, convolution processing is performed on the first voice to obtain the first accent characteristics of the first voice, and the second voice is determined based on the timbre characteristics, the first accent characteristics, the identification of the second accent and the accent conversion model. In this way, the electronic device can separate the timbre information and the accent information, avoid timbre leakage, and improve the accuracy of accent conversion. In this way, after the accent conversion model performs accent conversion, the timbre of the second voice is the same as the timbre of the first voice, thereby improving the effect of accent conversion.

[0082] It should be noted that after the electronic device acquires the second voice, it may or may not perform voice alignment processing with the first voice, and this is not limited in the embodiments of the present disclosure.

[0083] The disclosed embodiment provides a speech processing method, in which an electronic device can obtain a first speech with a first accent, obtain an identifier of a second accent, perform convolution processing on the first speech based on a timbre separation module in an accent conversion model to obtain the timbre feature of the first speech, perform convolution processing on the first speech based on an automatic speech recognition module in the accent conversion model to obtain a first accent feature of the first speech, and determine the second speech based on the timbre feature, the first accent feature, the identifier of the second accent and the accent conversion model. In this way, since the accent replacement module includes multiple accent decoders, the multiple accent decoders can correspond to multiple accents one by one, therefore, the accent conversion model can implement multiple accents to multiple accents accent conversion processing, and since the automatic speech recognition module can separate the timbre information in the first speech, the remaining accent information and the voice style information, therefore, the accent replacement module can accurately convert the accent information, thereby avoiding the problem of timbre leakage and improving the accuracy of accent conversion.

[0084] exist Figure 2 Based on the embodiment shown below, combined with Figure 8 , a method for determining the second voice based on the accent conversion model, the first voice and the identification of the second accent in the above-mentioned voice processing method is described in detail.

[0085] Figure 8 A schematic diagram of a method for determining a second voice provided by an embodiment of the present disclosure. Figure 8 The method flow includes:

[0086] S801. Based on the timbre separation module in the accent conversion model, convolution processing is performed on the first speech to obtain the timbre characteristics of the first speech.

[0087] Optionally, the timbre separation module can obtain the speaker information in the first speech, and then obtain the timbre characteristics of the first speech. For example, the timbre separation module may include multiple convolutional layers, and the timbre classification module can obtain the speaker information of the first speech after convolution processing on the first speech based on multiple convolutional layers, and then obtain the timbre characteristics (embedding) of the first speech. For example, the timbre classification module can be an ECAPA-TDNN model, based on which the first speech can be convolutionally processed to obtain the timbre characteristics of the first speech, wherein the ECAPA-TDNN model can be a speech recognition model, which is based on the combination of TDNN (Time Delay Neural Network) and ECAPA (Extended Context-Aware Parallel Attention).

[0088] It should be noted that the timbre separation module may be a trained module, and the training process of the timbre separation module will not be described in detail in the embodiment of the present disclosure.

[0089] S802: Based on the automatic speech recognition module in the accent conversion model, perform convolution processing on the first speech to obtain a first accent feature of the first speech.

[0090] The automatic speech recognition module may be an ASR module. The electronic device performs convolution processing on the first speech based on the automatic speech recognition module in the accent conversion model to obtain the first accent feature of the first speech, which may be specifically: determining a target convolution layer among multiple convolution layers of the automatic speech recognition module, performing convolution processing on the first speech based on the automatic speech recognition module, and determining the convolution result output by the target convolution layer as the first accent feature.

[0091] The convolution result output by the target convolution layer may include accent information and voice style information (the first accent feature may include accent information and voice style information of the first voice). For example, the automatic speech recognition module may be a trained ASR model, which can convert any segment of speech into text. Since the goal of the model is to convert speech into text, the ASR model will lose the timbre information in the speech during the convolution process. The convolution result output by one of the convolution layers may not include the timbre information. The convolution layer may be the target convolution layer, which can accurately remove the timbre information in the speech information and improve the accuracy of accent conversion.

[0092] It should be noted that the automatic speech recognition module may be a pre-trained ASR model, and the training process of the automatic speech recognition module will not be described in detail in the embodiment of the present disclosure.

[0093] Next, combine Fig. 9 , the process of determining the first accent feature is explained.

[0094] Fig. 9 A schematic diagram of a process for determining a first accent feature provided by an embodiment of the present disclosure. Fig. 9 , including: a first voice, an automatic voice recognition module and a text of the first voice. The automatic voice recognition module may include convolutional layer 1, convolutional layer 2, ..., convolutional layer 9, and the 9 convolutional layers are connected in sequence. Electronic device ( Fig. 9 (not shown) a first speech can be input into the automatic speech recognition module. After the automatic speech recognition module obtains the first speech, it can perform convolution processing on the first speech based on 9 convolution layers. The convolution layer 9 can output the text of the first speech. The electronic device can determine the convolution result output by the convolution layer 5 as the first accent feature.

[0095] It should be noted that in the embodiment of the present disclosure, the electronic device can determine the convolution layer in the middle position of the automatic speech recognition module as the target convolution layer, or can determine the target convolution layer based on any other feasible implementation method, and the embodiment of the present disclosure is not limited to this.

[0096] S803: Determine a second voice based on the timbre feature, the first accent feature, the identifier of the second accent, and the accent conversion model.

[0097] Among them, the electronic device can determine the second voice based on the following feasible implementation method: based on the accent replacement module and the identification of the second accent, convolution processing is performed on the first accent feature to obtain the second accent feature, and the second voice is determined based on the timbre feature, the second accent feature and the accent conversion model.

[0098] The second accent feature may include information about the second accent and voice style information (the voice style information may be the same as the voice style information included in the first accent feature). For example, after the accent replacement module determines the identifier of the second accent, it may perform convolution processing on the first accent feature to convert the accent information in the first accent feature into information about the second accent.

[0099] Among them, the electronic device performs convolution processing on the first accent feature based on the accent replacement module and the identification of the second accent to obtain the second accent feature. Specifically, it can be: based on the identification of the second accent, determine the target accent decoder corresponding to the identification of the second accent in the accent replacement module, perform convolution processing on the first accent feature based on the accent encoder to obtain the convolution feature, and perform convolution processing on the convolution feature based on the target accent decoder to obtain the second accent feature.

[0100] The convolution feature may be a feature obtained by convolving the first accent feature by the accent encoder. The target accent decoder may be a decoder corresponding to the second accent. For example, if the second accent corresponds to accent decoder 1, the electronic device may determine accent decoder 1 as the target accent decoder, and if the second accent corresponds to accent decoder 2, the electronic device may determine accent decoder 2 as the target accent decoder.

[0101] Optionally, since multiple accent decoders correspond one-to-one to multiple accents, the electronic device can obtain the correspondence between the accent decoders and the accents, and determine the target accent decoder corresponding to the second accent based on the correspondence. For example, the correspondence between the accent decoders and the accents may include that accent 1 corresponds to accent decoder A, accent 2 corresponds to accent decoder B, and accent 3 corresponds to accent decoder C. If the second accent is accent 1, the electronic device can determine that the target accent decoder is accent decoder A, if the second accent is accent 2, the electronic device can determine that the target accent decoder is accent decoder B, and if the second accent is accent 3, the electronic device can determine that the target accent decoder is accent decoder C. In this way, the electronic device can accurately perform accent conversion processing on the first speech and improve the accuracy of accent conversion.

[0102] It should be noted that the above examples are merely examples of the correspondence between the accent decoder and the accent in the embodiments of the present disclosure, and are not limitations of the correspondence between the accent decoder and the accent.

[0103] Next, combine Fig.10 , the process of determining the second accent characteristics is explained.

[0104] Fig.10 A schematic diagram of a process for determining a second accent feature provided by an embodiment of the present disclosure. Fig.10 , including: a first accent feature, a second accent, and an accent replacement module. The accent replacement module may include an accent encoder, an accent decoder 1, an accent decoder 2, ..., an accent decoder n. Electronic device ( Fig.10 The first accent feature and the second accent (not shown) may be input to the accent replacement module, and the accent replacement module may determine that the second accent corresponds to the accent decoder 2. The accent encoder may perform convolution processing on the first accent feature to obtain a convolution feature, and send the convolution feature to the accent decoder 2, and the accent decoder 2 may perform convolution processing on the convolution feature to obtain a second accent feature.

[0105] The following is an explanation of the training process of the accent replacement module through specific examples.

[0106] The electronic device can process the Shaanxi dialect "Hello" based on the ASR model to obtain accent feature A (the output result of the middle convolution layer). The electronic device can process the Mandarin "Hello" based on the ASR model to obtain accent feature B (the output result of the middle convolution layer). In this way, the electronic device can obtain a group of samples, which may include accent feature A and accent feature B.

[0107] If the accent decoder corresponding to the Shaanxi dialect accent is accent decoder 1 (i.e., accent decoder 1 is trained as the Shaanxi dialect accent decoder), then the accent encoder in the accent replacement module can perform convolution processing on the accent feature B, and input the convolution result to the accent decoder 1, and the accent decoder 1 can output the accent feature C. The electronic device can determine the loss function based on the loss between the accent feature A and the accent feature C, and train the accent replacement module based on the loss function, and repeat the above steps based on multiple groups of samples until the training of the accent decoder 1 converges, that is, the training of the accent decoder 1 in the accent replacement module is completed.

[0108] Among them, the electronic device determines the second voice based on the timbre characteristics, the second accent characteristics and the accent conversion model, which can be specifically: obtaining a blurred image with Gaussian noise added, processing the timbre characteristics, the second accent characteristics and the blurred image based on the diffusion module in the accent conversion model to obtain a Mel-spectrogram image, and obtaining the second voice based on the Mel-spectrogram image.

[0109] The blurred image may be an image with Gaussian noise added. For example, the electronic device may add Gaussian noise to any image until the image is completely distorted, and the electronic device may obtain a blurred image. The electronic device may also obtain a blurred image based on any other feasible implementation method, and the embodiments of the present disclosure are not limited to this.

[0110] Optionally, the electronic device processes the timbre features, the second accent features and the blurred image based on the diffusion module in the accent conversion module to obtain a mel-spectrogram image. Specifically, for the first processing, based on the diffusion module, the timbre features and the second accent features, the Gaussian noise in the blurred image is denoised to obtain the first image to be processed; for the i-th processing, based on the diffusion module, the timbre features and the second accent features, the Gaussian noise in the i-1 images to be processed is denoised to obtain the i-th image to be processed; until the Gaussian noise in the i-th image to be processed is less than or equal to a preset threshold, a mel-spectrogram image is obtained.

[0111] Here, i is 2, 3, ..., N in sequence, and N is the number of times the diffusion module processes. For example, if the step length of the diffusion module is 1000 times, then N can be 1000.

[0112] For example, in actual application, the diffusion module can gradually restore (denoise) the blurred image based on the timbre features and the second accent features. After multiple processing steps, a clear image can be obtained. Since the image is restored based on the timbre features and the second accent features, the image can be a mel-spectrogram.

[0113] Next, combine Fig.11, the processing process of the diffusion module is explained.

[0114] Fig.11 A schematic diagram of the processing process of a diffusion module provided in an embodiment of the present disclosure. Fig.11 , including: timbre features, blurred images, second accent features, processing information (the first time, the information is: the first processing, which can be encoded based on the encoder and the result can be spliced ​​with the above information) and diffusion module. Among them, electronic equipment ( Fig.11 (not shown) the timbre feature, blurred image, first processed information and second accent feature can be input into the module of step 1 in the diffusion module (processing module of step 1), and the module of step 1 can output the denoised image 1 to be processed.

[0115] See also Fig.11 , the electronic device can input the timbre feature, the image 2 to be processed, the information processed for the second time, and the second accent feature to the module of step 2 in the diffusion module, and the module of step 2 can output the image 2 to be processed after the denoising of the image 1 to be processed. The electronic device can loop the above steps until the timbre feature, the image 999 to be processed, the information processed for the 1000th time, and the second accent feature are input to the module of step 1000 in the diffusion module, and the module of step 1000 can obtain the mel spectrum image after denoising the image 999 to be processed. In this way, based on multiple denoising steps, the accuracy of the mel spectrum image can be improved, and then the accuracy of the second speech can be improved.

[0116] It should be noted that after the electronic device acquires the mel-spectrogram image, it can obtain the second speech associated with the mel-spectrogram image based on any feasible implementation method, and the embodiments of the present disclosure are not limited to this.

[0117] The disclosed embodiment provides a method for determining a second voice, wherein the electronic device performs convolution processing on a first voice based on a timbre separation module in an accent conversion model to obtain timbre features of the first voice, determines a target convolution layer among multiple convolution layers of an automatic speech recognition module, performs convolution processing on the first voice based on the automatic speech recognition module, and determines the convolution result output by the target convolution layer as a first accent feature, performs convolution processing on the first accent feature based on an accent replacement module and an identifier of a second accent to obtain a second accent feature, and determines the second voice based on the timbre features, the second accent features, and the accent conversion model. In this way, since the second accent feature does not include timbre information, the accuracy of the second accent feature is relatively high, and therefore, the electronic device can accurately perform accent conversion processing on the first voice, improve the effect of the second voice, and improve the accuracy of the second voice.

[0118] Based on any of the above embodiments, before determining the second voice based on the timbre feature, the second accent feature and the accent conversion model, the above voice processing method may further include a voice style conversion method. Fig.12 , the method of speech style transfer is explained in detail.

[0119] Fig.12 A schematic diagram of a method for voice style conversion provided by an embodiment of the present disclosure. Fig.12 The method flow includes:

[0120] S1201: Obtain target voice style.

[0121] Optionally, the target voice style may be the voice style to be converted. For example, if the electronic device converts the voice style of the first voice into voice style 1, voice style 1 may be the target voice style, and if the electronic device converts the voice style of the first voice into voice style 2, voice style 2 may be the target voice style.

[0122] It should be noted that the voice style may indicate any voice information unrelated to timbre and accent, such as the rhythm, prosody, stress and intonation of the voice, and the embodiments of the present disclosure are not limited to this.

[0123] It should be noted that the electronic device can acquire the target voice style based on any feasible implementation method (for example, the method for the electronic device to acquire the target voice style is similar to the method for acquiring the second accent, which will not be elaborated in the embodiments of the present disclosure), and the embodiments of the present disclosure are not limited to this.

[0124] S1202: Determine a target voice style replacement module corresponding to a target voice style among a plurality of voice style replacement modules in the accent conversion model.

[0125] Optionally, the accent conversion model may include multiple voice style replacement modules, wherein each voice style replacement module may correspond to a voice style. For example, the voice style replacement module corresponds to voice style 1, and the electronic device may generate voice of voice style 1 based on the information output by the voice style replacement module; the voice style replacement module corresponds to voice style 2, and the electronic device may generate voice of voice style 2 based on the information output by the voice style replacement module.

[0126] The voice style replacement module is used to convert the voice style information in the accent feature. For example, the accent feature may include accent information and voice style information, and the accent replacement module may convert the accent information in the accent feature, and the voice style replacement module may convert the voice style information in the accent feature.

[0127] The target voice style replacement module may be a module corresponding to the target voice style. For example, voice style 1 corresponds to voice style replacement module A, and voice style 2 corresponds to voice style replacement module B. If the electronic device determines that the target voice style is voice style 1, the electronic device may determine that the target voice style replacement module is voice style replacement module A. If the electronic device determines that the target voice style is voice style 2, the electronic device may determine that the target voice style replacement module is voice style replacement module B.

[0128] It should be noted that the method for an electronic device to determine a target voice style replacement module is similar to the method for an electronic device to determine a target accent decoder, and the embodiments of the present disclosure are not limited thereto.

[0129] It should be noted that the model structure of the speech style replacement module may be the same as the model structure of the accent style replacement module, and the embodiments of the present disclosure are not limited to this.

[0130] Optionally, the training process of the voice style replacement module may be similar to the training process of the accent style replacement module. The training process of the voice style replacement module is described in detail below through specific examples.

[0131] The electronic device can process the Mandarin "Hello" with voice style 1 based on the ASR model to obtain voice style feature A (the output result of the middle convolution layer). The electronic device can process the Mandarin "Hello" with voice style 2 based on the ASR model to obtain voice style feature B (the output result of the middle convolution layer). In this way, the electronic device can obtain a group of samples, which may include voice style feature A and voice style feature B.

[0132] If the speech style replacement module is trained as a speech style replacement module of speech style 2, the electronic device can perform convolution processing on the speech style feature A based on the speech style replacement module to obtain the speech style feature C. The electronic device can construct a loss function based on the loss between the speech style feature B and the speech style feature C, and train the speech style replacement module based on the loss function. The above steps are repeated based on multiple groups of samples until the training of the speech style replacement module converges and the training of the speech style replacement module is completed.

[0133] S1203: Perform convolution processing on the second accent feature based on the target speech style replacement module.

[0134] Optionally, after the electronic device determines the target voice style replacement module, it can perform convolution processing on the second accent feature based on the target voice style replacement module, and then convert the voice style information in the second accent feature. The convolved second accent feature can participate in the generation process of the second voice. In this way, the electronic device can not only convert the accent of the first voice into the second accent, but also convert the voice style of the first voice into the target voice style.

[0135] The disclosed embodiment provides a method for voice style conversion, wherein an electronic device can obtain a target voice style, determine a target voice style replacement module corresponding to the target voice style among multiple voice style replacement modules in an accent conversion model, and perform convolution processing on a second accent feature based on the target voice style replacement module. In this way, the second accent feature can include not only information about the second accent, but also information about the target voice style, thereby improving the function of the accent conversion model, improving the effect of the second voice, and improving the accuracy of the second voice.

[0136] Based on any of the above embodiments, Fig.13 , the process of the speech processing method is explained.

[0137] Fig.13 A process diagram of a speech processing method provided by an embodiment of the present disclosure. Fig.13 , including speech of a first accent, a timbre separation module, an automatic speech recognition module, an accent replacement module, a voice style replacement module and a diffusion module. Among them, the timbre classification module, the automatic speech recognition module, the accent replacement module, the voice style replacement module and the diffusion module can be modules in the accent conversion model.

[0138] See also Fig.13 The electronic device may convert the speech of the first accent into a mel-spectrogram, and input the mel-spectrogram into the timbre separation module and the automatic speech recognition module. The timbre separation module may process the mel-spectrogram to obtain the timbre feature of the first speech. The automatic speech recognition module may process the mel-spectrogram, determine the output of the intermediate convolutional layer as the first accent feature, and input the first accent feature into the accent replacement module.

[0139] See also Fig.13 The accent replacement module can perform convolution processing on the first accent feature and input the convolution result to the voice style replacement module, which can obtain the second accent feature by performing convolution processing on the convolution result. The electronic device can obtain the Gaussian blurred image and convert the timbre feature, the second accent feature and the time step feature (which can be obtained based on the encoder, that is, respectively) into Fig.11 The information processed for the first time, the information processed for the second time, ..., the information processed for the 1000th time).

[0140] See also Fig.13 The diffusion module can perform multiple denoising processes on the Gaussian module image based on the second accent feature, timbre feature and time step feature to obtain a predicted mel spectrogram. The electronic device can obtain the speech of the second accent based on the predicted mel spectrogram, and the style of the speech of the second accent is the same as the speech style corresponding to the speech style replacement module.

[0141] In this way, the accent conversion model can not only realize "one-to-one", "one-to-many" and "many-to-many" accent conversion processing, but also replace the voice style of the original speech. Moreover, since the automatic speech recognition module can remove the timbre information in the speech information, the timbre leakage can be avoided in the process of accent conversion, thereby improving the accuracy of accent conversion. In this way, the accent conversion model can retain the timbre of the original speech and realize the conversion of accent or voice style (that is, without changing the speaker's timbre, changing the speaker's accent or speaking style), thereby improving the accuracy of accent conversion and the effect of accent conversion.

[0142] Fig.14 This is a schematic diagram of the structure of a speech processing device provided by an embodiment of the present disclosure. Fig.14 The speech processing device 140 includes a first acquisition module 141, a second acquisition module 142 and a determination module 143, wherein:

[0143] The first acquisition module 141 is used to acquire a first speech with a first accent;

[0144] The second acquisition module 142 is used to acquire an identifier of a second accent;

[0145] The determination module 143 is used to determine the second voice based on the accent conversion model, the identification of the first voice and the second accent, the voice content of the second voice is the same as the voice content of the first voice, and the accent of the second voice is the second accent; wherein the accent conversion model includes an accent replacement module, the accent replacement module is used to convert the accent information in other voice information except the timbre information, the accent replacement module includes an accent encoder and multiple accent decoders connected to the accent encoder, and the multiple accent decoders correspond one-to-one to multiple accents.

[0146] According to one or more embodiments of the present disclosure, the determining module 143 is specifically configured to:

[0147] Based on the timbre separation module in the accent conversion model, convolution processing is performed on the first speech to obtain the timbre characteristics of the first speech;

[0148] Based on the automatic speech recognition module in the accent conversion model, performing convolution processing on the first speech to obtain a first accent feature of the first speech;

[0149] The second speech is determined based on the timbre feature, the first accent feature, an identifier of the second accent, and the accent conversion model.

[0150] According to one or more embodiments of the present disclosure, the determining module 143 is specifically configured to:

[0151] Based on the accent replacement module and the identifier of the second accent, performing convolution processing on the first accent feature to obtain a second accent feature;

[0152] The second voice is determined based on the timbre feature, the second accent feature and the accent conversion model.

[0153] According to one or more embodiments of the present disclosure, the determining module 143 is specifically configured to:

[0154] Get the blurred image with Gaussian noise added;

[0155] Based on the diffusion module in the accent conversion model, the timbre feature, the second accent feature and the blurred image are processed to obtain a mel-spectrogram image;

[0156] The second speech is obtained based on the mel-spectrogram image.

[0157] According to one or more embodiments of the present disclosure, the determining module 143 is specifically configured to:

[0158] For the first treatment;

[0159] Based on the diffusion module, the timbre feature and the second accent feature, denoising the Gaussian noise in the blurred image to obtain a first image to be processed;

[0160] For the i-th processing;

[0161] Based on the diffusion module, the timbre feature and the second accent feature, denoising the Gaussian noise in the (i-1)th image to be processed to obtain the (i)th image to be processed, until the Gaussian noise in the (i)th image to be processed is less than or equal to a preset threshold, thereby obtaining the Mel-spectrogram image;

[0162] The i is 2, 3, ..., N in sequence, and N is the number of processing times of the diffusion module.

[0163] According to one or more embodiments of the present disclosure, the determining module 143 is specifically configured to:

[0164] Based on the identifier of the second accent, determining, in the accent replacement module, a target accent decoder corresponding to the identifier of the second accent;

[0165] Performing convolution processing on the first accent feature based on the accent encoder to obtain a convolution feature;

[0166] The convolution feature is convolved based on the target accent decoder to obtain the second accent feature.

[0167] According to one or more embodiments of the present disclosure, the determining module 143 is specifically configured to:

[0168] Determining a target convolutional layer among a plurality of convolutional layers of the automatic speech recognition module;

[0169] The first speech is subjected to convolution processing based on the automatic speech recognition module, and the convolution result output by the target convolution layer is determined as the first accent feature.

[0170] According to one or more embodiments of the present disclosure, the first acquisition module 141 is specifically used to:

[0171] Displaying a voice acquisition page, wherein the voice acquisition page includes a voice acquisition control;

[0172] In response to a touch operation on the voice acquisition control, displaying a voice page, wherein the voice page includes a plurality of voices;

[0173] In response to a touch operation on the multiple voices, the voice associated with the touch operation is determined as the first voice.

[0174] According to one or more embodiments of the present disclosure, the first acquisition module 141 is specifically used to:

[0175] Displaying a voice collection page, wherein the voice collection page includes a voice recording control;

[0176] In response to a touch operation on the voice recording control, voice recording is performed to obtain the first voice.

[0177] According to one or more embodiments of the present disclosure, the determining module 143 is further configured to:

[0178] Acquire the target voice style;

[0179] Determining a target speech style replacement module corresponding to the target speech style among a plurality of speech style replacement modules in the accent conversion model;

[0180] performing convolution processing on the second accent feature based on the target speech style replacement module;

[0181] The speech style replacement module is used to convert the speech style information in the accent feature.

[0182] The speech processing device provided in the embodiment of the present disclosure can be used to execute the technical solution of the above-mentioned method embodiment. Its implementation principle and technical effect are similar, and this embodiment will not be repeated here.

[0183] Fig.15 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Fig.15 , which shows a schematic diagram of the structure of an electronic device 1500 suitable for implementing the embodiment of the present disclosure. The electronic device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Fig.15 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0184] like Fig.15 As shown, the electronic device 1500 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 1501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1502 or a program loaded from a storage device 1508 to a random access memory (RAM) 1503. Various programs and data required for the operation of the electronic device 1500 are also stored in the RAM 1503. The processing device 1501, the ROM 1502, and the RAM 1503 are connected to each other via a bus 1504. An input / output (I / O) interface 1505 is also connected to the bus 1504.

[0185] Typically, the following devices may be connected to the I / O interface 1505: input devices 1506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 1507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 1508 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 1509. The communication devices 1509 may allow the electronic device 1500 to communicate with other devices wirelessly or by wire to exchange data. Although Fig.15 The electronic device 1500 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0186] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 1509, or installed from a storage device 1508, or installed from a ROM 1502. When the computer program is executed by the processing device 1501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0187] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0188] The computer-readable medium may be included in the electronic device, or may exist independently without being installed in the electronic device.

[0189] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.

[0190] An embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the speech processing method that may be involved in various embodiments above is implemented.

[0191] The embodiments of the present disclosure provide a computer program product, including a computer program. When the computer program is executed by a processor, the speech processing method that may be involved in the above embodiments is implemented.

[0192] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0193] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0194] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware. The name of a unit does not limit the unit itself in some cases. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses".

[0195] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0196] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0197] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0198] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0199] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0200] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can independently choose whether to provide personal information to software or hardware such as an electronic device, application, server or storage medium that performs the operation of the technical solution of the present disclosure based on the prompt message. As an optional but non-limiting implementation method, in response to receiving an active request from a user, the method of sending a prompt message to the user can be, for example, a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0201] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0202] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of the relevant laws and regulations. The data may include information, parameters and messages, such as flow switching indication information.

[0203] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.

[0204] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0205] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.

Claims

1. A speech processing method, It is characterized in that include: Obtaining a first voice with a first accent; Get the identity of the second accent; Determine a second voice based on the accent conversion model, the first voice, and the identifier of the second accent, wherein the voice content of the second voice is the same as the voice content of the first voice, and the accent of the second voice is the second accent; Among them, the accent conversion model includes an accent replacement module, which is used to convert accent information in other speech information except timbre information. The accent replacement module includes an accent encoder and multiple accent decoders connected to the accent encoder, and the multiple accent decoders correspond one-to-one to multiple accents.

2. The method according to claim 1, It is characterized in that The determining the second voice based on the accent conversion model, the first voice and the identifier of the second accent includes: Based on the timbre separation module in the accent conversion model, convolution processing is performed on the first speech to obtain the timbre characteristics of the first speech; Based on the automatic speech recognition module in the accent conversion model, performing convolution processing on the first speech to obtain a first accent feature of the first speech; The second speech is determined based on the timbre feature, the first accent feature, an identifier of the second accent, and the accent conversion model.

3. The method according to claim 2, It is characterized in that The determining the second voice based on the timbre feature, the first accent feature, the identifier of the second accent and the accent conversion model includes: Based on the accent replacement module and the identifier of the second accent, performing convolution processing on the first accent feature to obtain a second accent feature; The second voice is determined based on the timbre feature, the second accent feature and the accent conversion model.

4. The method according to claim 3, It is characterized in that Determining the second voice based on the timbre feature, the second accent feature, and the accent conversion model includes: Get the blurred image with Gaussian noise added; Based on the diffusion module in the accent conversion model, the timbre feature, the second accent feature and the blurred image are processed to obtain a mel-spectrogram image; The second speech is obtained based on the mel-spectrogram image.

5. The method according to claim 4, It is characterized in that The step of processing the timbre feature, the second accent feature and the blurred image based on the diffusion module in the accent conversion model to obtain a mel spectrum image includes: For the first treatment; Based on the diffusion module, the timbre feature and the second accent feature, denoising the Gaussian noise in the blurred image to obtain a first image to be processed; For the i-th processing; Based on the diffusion module, the timbre feature and the second accent feature, denoising the Gaussian noise in the (i-1)th image to be processed to obtain the (i)th image to be processed, until the Gaussian noise in the (i)th image to be processed is less than or equal to a preset threshold, thereby obtaining the Mel-spectrogram image; The i is 2, 3, ..., N in sequence, and N is the number of processing times of the diffusion module.

6. The method according to any one of claims 3 to 5, It is characterized in that The step of performing convolution processing on the first accent feature based on the accent replacement module and the identifier of the second accent to obtain the second accent feature includes: Based on the identifier of the second accent, determining, in the accent replacement module, a target accent decoder corresponding to the identifier of the second accent; Performing convolution processing on the first accent feature based on the accent encoder to obtain a convolution feature; The convolution feature is convolved based on the target accent decoder to obtain the second accent feature.

7. The method according to any one of claims 2 to 5, It is characterized in that The automatic speech recognition module based on the accent conversion model performs convolution processing on the first speech to obtain a first accent feature of the first speech, including: Determining a target convolutional layer among a plurality of convolutional layers of the automatic speech recognition module; The first speech is convolved based on the automatic speech recognition module, and the convolution result output by the target convolution layer is determined as the first accent feature.

8. The method according to any one of claims 1 to 5, It is characterized in that The step of acquiring the first speech with the first accent includes: Displaying a voice acquisition page, wherein the voice acquisition page includes a voice acquisition control; In response to a touch operation on the voice acquisition control, displaying a voice page, wherein the voice page includes a plurality of voices; In response to a touch operation on the multiple voices, the voice associated with the touch operation is determined as the first voice.

9. The method according to any one of claims 1 to 5, It is characterized in that The step of acquiring the first speech with the first accent includes: Displaying a voice collection page, wherein the voice collection page includes a voice recording control; In response to a touch operation on the voice recording control, voice recording is performed to obtain the first voice.

10. The method according to any one of claims 3 to 5, It is characterized in that Before determining the second voice based on the timbre feature, the second accent feature and the accent conversion model, the method further includes: Acquire the target voice style; Determining a target speech style replacement module corresponding to the target speech style among a plurality of speech style replacement modules in the accent conversion model; performing convolution processing on the second accent feature based on the target speech style replacement module; The speech style replacement module is used to convert the speech style information in the accent feature.

11. A speech processing device, It is characterized in that It includes a first acquisition module, a second acquisition module and a determination module, wherein: The first acquisition module is used to acquire a first speech with a first accent; The second acquisition module is used to acquire an identifier of a second accent; The determination module is used to determine the second voice based on the accent conversion model, the first voice and the identification of the second accent, the voice content of the second voice is the same as the voice content of the first voice, and the accent of the second voice is the second accent; wherein the accent conversion model includes an accent replacement module, the accent replacement module is used to convert accent information in other voice information except timbre information, the accent replacement module includes an accent encoder and multiple accent decoders connected to the accent encoder, and the multiple accent decoders correspond one-to-one to multiple accents.

12. An electronic device, It is characterized in that include: Processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the speech processing method according to any one of claims 1 to 10.

13. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer-executable instructions. When the processor executes the computer-executable instructions, the speech processing method according to any one of claims 1 to 10 is implemented.

Citation Information

Cited By

  • Voice evaluation method and device, medium and program product

    CN120319271A