Voice processing method and apparatus, and electronic device

By introducing a tone separation module and an automatic speech recognition module into the accent conversion model, combined with the accent replacement module, the problem of low accent conversion accuracy caused by tone leakage in the prior art is solved, and a high-accent conversion is achieved.

WO2025118848A1PCT designated stage expired Publication Date: 2025-06-12LEMON INC(GB) +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/126399
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-10-22
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The existing accent conversion model has timbre leakage problems, resulting in poor accuracy of accent conversion.

Method used

The accent conversion model including a tone separation module, an automatic speech recognition module and an accent replacement module is adopted to obtain the tone characteristics and accent characteristics through convolution processing, and the multi-accent to multi-accent conversion is performed based on these characteristics.

Benefits of technology

It effectively avoids tone leakage and improves the accuracy of accent conversion, so that the converted voice content and style are the same as the original voice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024126399_12062025_PF_FP_ABST
    Figure CN2024126399_12062025_PF_FP_ABST
Patent Text Reader

Abstract

A voice processing method and apparatus, and an electronic device. The method comprises: acquiring a first voice of a first accent (S201); acquiring an identifier of a second accent (S202); and determining a second voice on the basis of an accent conversion model, the first voice, and the identifier of the second accent (S203), wherein the voice content of the second voice is the same as the voice content of the first voice, and the accent of the second voice is the second accent.
Need to check novelty before this filing date? Find Prior Art

Description

Voice processing method, device and electronic equipment

[0001] This application claims priority to Chinese Patent Application No. 202311659609.5 filed on December 5, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field

[0002] The embodiments of the present disclosure relate to a speech processing method, apparatus, and electronic device. Background Art

[0003] Accent conversion is the process of converting one accent in a speech segment into another, without changing the speech content. For example, an electronic device can convert the Mandarin "hello" into the Sichuan dialect.

[0004] Currently, electronic devices can process speech based on accent conversion models to convert the accent of the speech. For example, a trained accent conversion model can convert speech with one accent into speech with another. However, accent conversion models suffer from timbre leakage, resulting in poor accent conversion accuracy.

[0005] Summary of the Invention

[0006] The present disclosure provides a speech processing method, apparatus, and electronic device.

[0007] In a first aspect, the present disclosure provides a speech processing method, the speech processing method comprising:

[0008] Obtaining a first speech with a first accent;

[0009] Get the identity of the second accent;

[0010] Determining a second voice based on the accent conversion model, the first voice, and the identifier of the second accent, wherein the voice content of the second voice is the same as the voice content of the first voice, and the accent of the second voice is the second accent;

[0011] Among them, the accent conversion model includes an accent replacement module, which is used to convert accent information in other speech information except timbre information. The accent replacement module includes an accent encoder and multiple accent decoders connected to the accent encoder, and the multiple accent decoders correspond one-to-one to multiple accents.

[0012] In a second aspect, the present disclosure provides a speech processing device, the speech processing device comprising a first acquisition module, a second acquisition module, and a determination module, wherein:

[0013] The first acquisition module is used to acquire a first speech with a first accent;

[0014] The second acquisition module is used to obtain an identifier of a second accent;

[0015] The determination module is used to determine the second voice based on the accent conversion model, the first voice and the identification of the second accent, the voice content of the second voice is the same as the voice content of the first voice, and the accent of the second voice is the second accent; wherein the accent conversion model includes an accent replacement module, the accent replacement module is used to convert accent information in other voice information except timbre information, the accent replacement module includes an accent encoder and multiple accent decoders connected to the accent encoder, and the multiple accent decoders correspond one-to-one to multiple accents.

[0016] In a third aspect, an embodiment of the present disclosure provides an electronic device comprising: a processor and a memory;

[0017] The memory stores computer-executable instructions;

[0018] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the speech processing method as described in the first aspect and various possible aspects of the first aspect.

[0019] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer execution instructions are stored. When a processor executes the computer execution instructions, the speech processing method as described in the first aspect and various possible aspects of the first aspect are implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, a brief introduction will be given below to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0021] FIG1 is a schematic diagram of an application scenario provided by an embodiment of the present disclosure;

[0022] FIG2 is a flow chart of a speech processing method provided by an embodiment of the present disclosure;

[0023] FIG3 is a schematic diagram of a process for obtaining a first voice according to an embodiment of the present disclosure;

[0024] FIG4 is a schematic diagram of another process of obtaining a first voice according to an embodiment of the present disclosure;

[0025] FIG5 is a schematic diagram of a process for acquiring a second accent provided by an embodiment of the present disclosure;

[0026] FIG6 is a schematic diagram of an accent replacement module provided by an embodiment of the present disclosure;

[0027] FIG7 is a schematic diagram of the structure of an accent conversion model provided by an embodiment of the present disclosure;

[0028] FIG8 is a schematic diagram of a method for determining a second voice provided by an embodiment of the present disclosure;

[0029] FIG9 is a schematic diagram of a process for determining a first accent feature according to an embodiment of the present disclosure;

[0030] FIG10 is a schematic diagram of a process for determining a second accent feature according to an embodiment of the present disclosure;

[0031] FIG11 is a schematic diagram of a processing process of a diffusion module provided by an embodiment of the present disclosure;

[0032] FIG12 is a schematic diagram of a method for voice style conversion provided by an embodiment of the present disclosure;

[0033] FIG13 is a process diagram of a speech processing method provided by an embodiment of the present disclosure;

[0034] FIG14 is a schematic structural diagram of a speech processing device provided by an embodiment of the present disclosure; and

[0035] FIG15 is a schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0037] To facilitate understanding, the concepts involved in the embodiments of the present disclosure are explained below.

[0038] Electronic device: is a device with wireless transceiver function. The electronic device can be deployed on land, including indoors or outdoors, handheld, wearable or vehicle-mounted. The electronic device can be a mobile phone, a tablet computer, a computer with wireless transceiver function, a virtual reality (VR) electronic device, an augmented reality (AR) electronic device, a wireless terminal in industrial control, an on-board electronic device, a wireless terminal in self-driving, a wireless electronic device in remote medical care, a wireless electronic device in smart grid, a wireless electronic device in transportation safety, a wireless electronic device in smart city, a wireless electronic device in smart home, a wearable electronic device, etc. The electronic device involved in the embodiments of the present disclosure can also be referred to as a terminal, user equipment (UE), an access electronic device, a vehicle-mounted terminal, an industrial control terminal, a UE unit, a UE station, a mobile station, a mobile station, a remote station, a remote electronic device, a mobile device, a UE electronic device, a wireless communication device, a UE agent or a UE device, etc. The electronic device can also be fixed or mobile.

[0039] The application scenario of the embodiment of the present disclosure is described below with reference to FIG1 .

[0040] FIG1 is a schematic diagram of an application scenario provided by an embodiment of the present disclosure. Please refer to FIG1 , which includes an electronic device. Among them, an accent conversion application can be installed in the electronic device. The electronic device can obtain Mandarin speech based on the application, and based on the application, convert the Mandarin speech into Shaanxi dialect speech, wherein the speech content in the Shaanxi dialect speech is the same as the speech content in the Mandarin speech, and the speech style information such as stress and rhythm in the Shaanxi dialect speech is also the same as the speech style information such as stress and rhythm in the Mandarin speech. In this way, the electronic device can convert the accent in the speech to obtain speech with other accents, and the content and style of the speech after the accent conversion are the same as the content and style of the speech before the accent conversion.

[0041] It should be noted that FIG1 is only an example of an application scenario of the embodiment of the present disclosure, and does not limit the application scenario of the embodiment of the present disclosure.

[0042] In the related art, accent conversion refers to converting the accent in a segment of speech into another accent, but the speech content of the segment remains unchanged. For example, an electronic device can convert the Mandarin "hello" into the Shaanxi dialect "hello". Currently, electronic devices can process speech based on an accent conversion model, and then convert the accent of the speech. For example, a trained accent conversion model can convert speech with one accent into speech with another accent, while keeping the speech content unchanged. However, the training samples of the accent conversion model are trained based on the mel-spectrogram of the speech, and the timbre information in the speech information will not be separated. In this way, the accent conversion model will have the problem of timbre leakage in the process of accent conversion of speech, the speech effect is poor, and the accuracy of accent conversion is poor.

[0043] In order to solve the technical problems in the related art, the embodiment of the present disclosure provides a speech processing method, in which the electronic device can obtain a first speech with a first accent, obtain an identification of a second accent, perform convolution processing on the first speech based on the timbre separation module in the accent conversion model to obtain the timbre characteristics of the first speech, perform convolution processing on the first speech based on the automatic speech recognition module in the accent conversion model to obtain the first accent characteristics of the first speech, perform convolution processing on the first accent characteristics based on the accent replacement module in the accent conversion model and the identification of the second accent to obtain the second accent characteristics, and determine the second speech based on the timbre characteristics, the second accent characteristics and the accent conversion model, wherein the speech content of the second speech is the same as the speech content of the first speech, and the accent of the second speech is the second accent. In this way, since the accent replacement module includes an accent encoder and multiple accent decoders connected to the accent encoder, and the multiple accent decoders correspond one-to-one to multiple accents, the accent conversion model can realize multi-accent to multi-accent accent conversion processing, and since the automatic speech recognition module can separate the timbre information in the first speech, the remaining accent information and voice style information, the accent replacement module can accurately convert the accent information, thereby avoiding the problem of timbre leakage and improving the accuracy of accent conversion.

[0044] The following detailed description of the technical solution of the present disclosure and how the technical solution of the present disclosure solves the above-mentioned technical problems is provided with specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The embodiments of the present disclosure will be described below in conjunction with the accompanying drawings.

[0045] FIG2 is a flow chart of a speech processing method provided by an embodiment of the present disclosure. Referring to FIG2 , the method may include:

[0046] S201: Obtain a first speech with a first accent.

[0047] The execution subject of the embodiments of the present disclosure may be an electronic device, or a voice processing device provided in the electronic device. The voice processing device may be implemented based on software, or based on a combination of software and hardware, which is not limited in the embodiments of the present disclosure. The electronic device may be any device with end-to-end computing capabilities. For example, the electronic device may be a computer, a server, a mobile device, etc., which is not limited in the embodiments of the present disclosure.

[0048] The first voice may be a voice to be processed. For example, the first voice may be a voice to be processed with an accent conversion. For example, if the electronic device performs an accent conversion on voice 1, the electronic device may determine that voice 1 is the first voice; if the electronic device performs an accent conversion on voice 2, the electronic device may determine that voice 2 is the first voice.

[0049] Optionally, the electronic device may receive the first voice sent by another device, the electronic device may record the first voice in real time, or the electronic device may obtain the first voice from a database, which is not limited in the embodiments of the present disclosure.

[0050] The first accent may be the accent of the first speech. For example, if the first speech has a Mandarin accent, the first accent may be Mandarin. If the first speech has a Shaanxi dialect accent, the first accent may be Shaanxi dialect. It should be noted that the above accents are merely examples and are not intended to limit the accents in the embodiments of the present disclosure.

[0051] It should be noted that after the electronic device acquires the first voice, it can determine the first accent of the first voice based on any feasible implementation method, and the embodiments of the present disclosure are not limited to this.

[0052] Optionally, the electronic device may obtain the first speech in the first accent based on the following two feasible implementation methods:

[0053] A possible implementation:

[0054] A voice acquisition page is displayed, wherein the voice acquisition page includes a voice acquisition control. In response to a touch operation on the voice acquisition control, a voice page is displayed, wherein the voice page includes multiple voices. In response to the touch operation on the multiple voices, the voice associated with the touch operation is determined as the first voice.

[0055] 3 , the process of the electronic device acquiring the first voice in this feasible implementation is described below.

[0056] Figure 3 is a schematic diagram of a process for acquiring a first voice according to an embodiment of the present disclosure. Referring to Figure 3 , an electronic device is included. The display page of the electronic device includes page 301 and page 302. Page 301 may be a voice acquisition page, which includes a voice acquisition control. When a user clicks the voice acquisition control, the display page of the electronic device jumps from page 301 to page 302.

[0057] Please refer to Figure 3, page 302 may be a voice page. The voice page may include voice A, voice B, voice C, and a confirmation control (it should be noted that the accents and voice styles of voice A, voice B, and voice C may be different, and the embodiments of the present disclosure are not limited to this). When the user clicks voice A, voice A may be highlighted, and after the user clicks the confirmation control, the electronic device may determine voice A as the first voice. In this way, the user can flexibly select the first voice, thereby improving the flexibility of obtaining the first voice.

[0058] Another possible implementation:

[0059] A voice collection page is displayed, wherein the voice collection page includes a voice recording control, and in response to a touch operation on the voice recording control, voice recording is performed to obtain a first voice.

[0060] 4 , the process of the electronic device acquiring the first voice in this feasible implementation is described below.

[0061] Figure 4 is another schematic diagram of the process of obtaining the first voice provided by an embodiment of the present disclosure. Please refer to Figure 4, which includes: an electronic device. Among them, the display page of the electronic device may include page 401 and page 402. Page 401 can be a voice acquisition page, which includes a voice recording control. After the user clicks the voice recording control, the display page of the electronic device jumps from page 401 to page 402. Page 402 can be a recording page, which can display the progress of the voice recording, and can include a confirmation control. After the voice recording is completed, the user can click the confirmation control, and the electronic device can determine the recorded voice as the first voice. In this way, the user can record the first voice in real time, improve the flexibility of obtaining the first voice, and improve the user experience.

[0062] S202: Obtain an identifier of a second accent.

[0063] The second accent may be a target accent for the accent conversion. For example, if the electronic device converts Mandarin speech into Shaanxi dialect speech, the target accent may be Shaanxi dialect, i.e., the second accent may be Shaanxi dialect; if the electronic device converts Mandarin speech into Sichuan dialect speech, the target accent may be Sichuan dialect, i.e., the second accent may be Sichuan dialect.

[0064] Optionally, the electronic device may obtain the identification of the second accent based on the following feasible implementation method: displaying a voice acquisition page, which may include a voice acquisition control and multiple accents, and in response to the user's touch operation on the multiple accents, determining the accent associated with the touch operation as the second accent.

[0065] It should be noted that the identifier of the second accent can be the name of the second accent, the unique mark of the second accent, etc., and the embodiments of the present disclosure are not limited to this. Moreover, after the electronic device determines the second accent, it can obtain the identifier of the second accent based on any feasible implementation method, and the embodiments of the present disclosure are not limited to this.

[0066] The following describes a process of the electronic device acquiring the second accent with reference to FIG5 .

[0067] Figure 5 is a schematic diagram of a process for acquiring a second accent according to an embodiment of the present disclosure. Referring to Figure 5 , the electronic device includes an electronic device. The display page of the electronic device may be a voice acquisition page. The voice acquisition page may include accent A, accent B, accent C, and a voice acquisition control. When a user clicks accent A, accent A may be highlighted, and the electronic device may determine accent A as the second accent.

[0068] S203: Determine the second voice based on the accent conversion model, the first voice, and the identifier of the second accent.

[0069] The second speech has the same content as the first speech, and the second speech has the second accent. For example, the first speech may be "hello" in accent A. After the electronic device performs accent conversion processing on the first speech, the second speech may be obtained. The second speech may be "hello" in accent B. For example, the first speech may be "The weather is great today" in Shaanxi dialect. After the electronic device performs accent conversion processing on the first speech, the second speech may be obtained. The second speech may be "The weather is great today" in Mandarin.

[0070] The accent conversion model includes an accent replacement module. Optionally, the accent replacement module is used to convert accent information in other speech information besides timbre information. For example, speech information may include timbre information, accent information, and voice style information, where voice style information may include prosody information, rhythm information, stress information, etc., although this is not limited in the present embodiment. Other speech information besides timbre information may include accent information and voice style information, and the accent replacement module may convert the accent information.

[0071] Optionally, the accent replacement module includes an accent encoder and multiple accent decoders connected to the accent encoder, wherein the accent encoder can be used to encode speech information except timbre information, and the interpretation decoder can be used to decode the information output by the accent encoder.

[0072] Multiple accent decoders correspond one-to-one to multiple accents. For example, accent decoder A can correspond to accent 1, accent decoder B can correspond to accent 2, and accent decoder C can correspond to accent 3. For example, if accent decoder A corresponds to accent 1, then the accent conversion model can generate speech with accent 1 based on the information output by accent decoder A. If accent decoder B corresponds to accent 2, then the accent conversion model can generate speech with accent 2 based on the information output by accent decoder B. In this way, the accent conversion model can realize multi-accent conversion scenarios, reducing the cost of accent conversion. For example, since the multiple accent encoders in the accent conversion model correspond one-to-one to multiple accents, the accent conversion model can convert the speech of Shaanxi dialect into the speech of Sichuan dialect (one-to-one accent conversion), the accent conversion model can also convert the speech of Shaanxi dialect into the speech of Sichuan dialect or Cantonese dialect (one-to-many accent conversion), the accent conversion model can also convert the speech of Shaanxi dialect into the speech of Sichuan dialect or Cantonese dialect, and convert the speech of Northeastern dialect into the speech of Sichuan dialect or Cantonese dialect (many-to-many accent conversion).

[0073] The structure of the accent replacement module will be described below with reference to FIG6 .

[0074] Figure 6 is a schematic diagram of an accent replacement module provided by an embodiment of the present disclosure. Referring to Figure 6 , the accent replacement module may include an accent encoder, accent decoder 1, accent decoder 2, ..., and accent decoder n. The accent encoder is connected to accent decoder 1, accent decoder 2, ..., and accent decoder n, respectively, and each accent decoder can correspond to a specific accent.

[0075] It should be noted that the structure of the accent conversion module can be a BN-to-BN (BN2BN) neural network structure or other neural network structures, which is not limited in the embodiments of the present disclosure.

[0076] Optionally, the accent conversion model may include a timbre separation module, an automatic speech recognition module, and a diffusion module. The timbre separation module is configured to obtain timbre information from the first speech, the automatic speech recognition module is configured to obtain other speech information from the first speech in addition to the timbre information, and the diffusion module is configured to generate the second speech.

[0077] The structure of the accent conversion model is described below with reference to FIG7 .

[0078] Figure 7 is a schematic diagram of the structure of an accent conversion model provided by an embodiment of the present disclosure. Referring to Figure 7 , the accent conversion model includes: an accent conversion model. The accent conversion model may include a timbre separation module, an automatic speech recognition module, a timbre replacement module, and a diffusion module. The timbre separation module may be connected to the diffusion module, the automatic speech recognition module may be connected to the timbre replacement module, and the timbre replacement module may also be connected to the diffusion module.

[0079] It should be noted that the network structure of the timbre classification module in the embodiment of the present disclosure can be any neural network structure with timbre separation, and the embodiment of the present disclosure is not limited to this.

[0080] It should be noted that the network structure of the automatic speech recognition module in the embodiment of the present disclosure can be a network structure built based on automatic speech recognition (ASR) technology, or it can be any other feasible neural network structure, and the embodiment of the present disclosure is not limited to this.

[0081] It should be noted that the network structure of the diffusion module in the embodiment of the present disclosure may be the network structure of a conditional diffusion model, or may be any other feasible neural network structure, and the embodiment of the present disclosure is not limited thereto.

[0082] The electronic device can determine the second voice based on the following feasible implementation: performing convolution processing on the first voice based on the timbre separation module in the accent conversion model to obtain the timbre characteristics of the first voice; performing convolution processing on the first voice based on the automatic speech recognition module in the accent conversion model to obtain the first accent characteristics of the first voice; and determining the second voice based on the timbre characteristics, the first accent characteristics, the identifier of the second accent, and the accent conversion model. In this way, the electronic device can separate the timbre information from the accent information, avoid timbre leakage, and improve the accuracy of accent conversion. Thus, after the accent conversion model performs accent conversion, the timbre of the second voice is the same as that of the first voice, thereby improving the effect of accent conversion.

[0083] It should be noted that after the electronic device acquires the second voice, it may or may not perform voice alignment processing with the first voice, and this embodiment of the present disclosure does not limit this.

[0084] The disclosed embodiments provide a speech processing method. An electronic device can obtain a first speech with a first accent, obtain an identifier of a second accent, perform convolution processing on the first speech based on a timbre separation module in an accent conversion model to obtain timbre features of the first speech, perform convolution processing on the first speech based on an automatic speech recognition module in the accent conversion model to obtain first accent features of the first speech, and determine the second speech based on the timbre features, the first accent features, the identifier of the second accent, and the accent conversion model. In this way, because the accent replacement module includes multiple accent decoders, each of which can correspond one-to-one to multiple accents, the accent conversion model can implement multi-accent to multi-accent accent conversion processing. Furthermore, because the automatic speech recognition module can separate the timbre information from the first speech, leaving only the accent information and voice style information, the accent replacement module can accurately convert the accent information, thereby avoiding timbre leakage and improving the accuracy of accent conversion.

[0085] Based on the embodiment shown in FIG2 , the method for determining the second speech based on the accent conversion model, the identification of the first speech and the second accent in the above speech processing method will be described in detail below in conjunction with FIG8 .

[0086] FIG8 is a schematic diagram of a method for determining a second voice according to an embodiment of the present disclosure. Referring to FIG8 , the method process includes:

[0087] S801: Based on the timbre separation module in the accent conversion model, perform convolution processing on the first speech to obtain the timbre characteristics of the first speech.

[0088] Optionally, the timbre separation module can obtain the speaker information in the first speech, and then obtain the timbre characteristics of the first speech. For example, the timbre separation module may include multiple convolution layers, and the timbre classification module can obtain the speaker information of the first speech after performing convolution processing on the first speech based on multiple convolution layers, and then obtain the timbre characteristics (embedding) of the first speech. For example, the timbre classification module can be an ECAPA-TDNN model, based on which the first speech can be convolutionally processed to obtain the timbre characteristics of the first speech, wherein the ECAPA-TDNN model can be a speech recognition model, which is obtained based on the combination of TDNN (Time Delay Neural Network) and ECAPA (Extended Context-Aware Parallel Attention).

[0089] It should be noted that the timbre separation module may be a trained module, and the training process of the timbre separation module will not be described in detail in the embodiment of the present disclosure.

[0090] S802: Based on the automatic speech recognition module in the accent conversion model, perform convolution processing on the first speech to obtain a first accent feature of the first speech.

[0091] The automatic speech recognition module may be an ASR module. The electronic device, based on the automatic speech recognition module in the accent conversion model, performs convolution processing on the first speech to obtain a first accent feature of the first speech. Specifically, the electronic device may perform convolution processing on the first speech based on the automatic speech recognition module in the accent conversion model to obtain a first accent feature of the first speech. Specifically, the electronic device may perform convolution processing on the first speech based on the automatic speech recognition module, determine a target convolution layer among multiple convolution layers of the automatic speech recognition module, perform convolution processing on the first speech based on the automatic speech recognition module, and determine the convolution result output by the target convolution layer as the first accent feature.

[0092] The convolution result output by the target convolution layer may include accent information and voice style information (the first accent feature may include accent information and voice style information of the first speech). For example, the automatic speech recognition module may be a trained ASR model that can convert any speech segment into text. Since the goal of the model is to convert speech into text, the ASR model will lose the timbre information in the speech during the convolution process. The convolution result output by one of the convolution layers may not include timbre information. This convolution layer may be the target convolution layer. This can accurately remove the timbre information from the speech information and improve the accuracy of accent conversion.

[0093] It should be noted that the automatic speech recognition module can be a pre-trained ASR model, and the training process of the automatic speech recognition module will not be described in detail in the embodiment of the present disclosure.

[0094] The process of determining the first accent feature is described below with reference to FIG9 .

[0095] FIG9 is a schematic diagram of a process for determining a first accent feature provided by an embodiment of the present disclosure. Please refer to FIG9 , which includes: a first voice, an automatic voice recognition module, and a text of the first voice. The automatic voice recognition module may include convolution layer 1, convolution layer 2, ..., convolution layer 9, and the nine convolution layers are connected in sequence. An electronic device (not shown in FIG9 ) can input the first voice into the automatic voice recognition module. After the automatic voice recognition module obtains the first voice, it can perform convolution processing on the first voice based on the nine convolution layers. The convolution layer 9 can output the text of the first voice. The electronic device can determine the convolution result output by the convolution layer 5 as the first accent feature.

[0096] It should be noted that in the embodiment of the present disclosure, the electronic device can determine the convolution layer in the middle position in the automatic speech recognition module as the target convolution layer, or can determine the target convolution layer based on any other feasible implementation method, and the embodiment of the present disclosure is not limited to this.

[0097] S803: Determine the second voice based on the timbre feature, the first accent feature, the identifier of the second accent, and the accent conversion model.

[0098] Among them, the electronic device can determine the second voice based on the following feasible implementation method: based on the accent replacement module and the identification of the second accent, convolution processing is performed on the first accent feature to obtain the second accent feature, and the second voice is determined based on the timbre feature, the second accent feature and the accent conversion model.

[0099] The second accent feature may include information about the second accent and voice style information (the voice style information may be the same as the voice style information included in the first accent feature). For example, after the accent replacement module determines the identifier of the second accent, it may perform convolution processing on the first accent feature to convert the accent information in the first accent feature into information about the second accent.

[0100] Among them, the electronic device performs convolution processing on the first accent feature based on the accent replacement module and the identification of the second accent to obtain the second accent feature. Specifically, it can be: based on the identification of the second accent, determine the target accent decoder corresponding to the identification of the second accent in the accent replacement module, perform convolution processing on the first accent feature based on the accent encoder to obtain the convolution feature, and perform convolution processing on the convolution feature based on the target accent decoder to obtain the second accent feature.

[0101] The convolution feature may be a feature obtained by convolving the first accent feature with the accent encoder. The target accent decoder may be a decoder corresponding to the second accent. For example, if the second accent corresponds to accent decoder 1, the electronic device may determine accent decoder 1 as the target accent decoder. If the second accent corresponds to accent decoder 2, the electronic device may determine accent decoder 2 as the target accent decoder.

[0102] Optionally, since multiple accent decoders correspond one-to-one to multiple accents, the electronic device can obtain the correspondence between the accent decoders and the accents and, based on the correspondence, determine the target accent decoder corresponding to the second accent. For example, the correspondence between the accent decoders and the accents may include accent 1 corresponding to accent decoder A, accent 2 corresponding to accent decoder B, and accent 3 corresponding to accent decoder C. If the second accent is accent 1, the electronic device can determine the target accent decoder to be accent decoder A; if the second accent is accent 2, the electronic device can determine the target accent decoder to be accent decoder B; and if the second accent is accent 3, the electronic device can determine the target accent decoder to be accent decoder C. In this way, the electronic device can accurately perform accent conversion processing on the first speech, thereby improving the accuracy of accent conversion.

[0103] It should be noted that the above examples are merely examples of the correspondence between the accent decoder and the accent in the embodiments of the present disclosure, and are not intended to limit the correspondence between the accent decoder and the accent.

[0104] The process of determining the second accent feature is described below with reference to FIG10 .

[0105] Figure 10 is a schematic diagram of a process for determining a second accent feature provided by an embodiment of the present disclosure. Please refer to Figure 10, which includes: a first accent feature, a second accent, and an accent replacement module. The accent replacement module may include an accent encoder, an accent decoder 1, an accent decoder 2, ..., and an accent decoder n. An electronic device (not shown in Figure 10) can input the first accent feature and the second accent into the accent replacement module, and the accent replacement module can determine that the second accent corresponds to the accent decoder 2. The accent encoder can perform convolution processing on the first accent feature to obtain a convolution feature, and send the convolution feature to the accent decoder 2. The accent decoder 2 can perform convolution processing on the convolution feature to obtain a second accent feature.

[0106] The following describes the training process of the accent replacement module through specific examples.

[0107] The electronic device can process the Shaanxi dialect "hello" based on the ASR model to obtain accent feature A (the output result of the middle convolution layer). The electronic device can process the Mandarin dialect "hello" based on the ASR model to obtain accent feature B (the output result of the middle convolution layer). In this way, the electronic device can obtain a set of samples, which may include accent feature A and accent feature B.

[0108] If the accent decoder corresponding to the Shaanxi dialect accent is accent decoder 1 (that is, accent decoder 1 is trained as the Shaanxi dialect accent decoder), then the accent encoder in the accent replacement module can perform convolution processing on the accent feature B and input the convolution result to the accent decoder 1. The accent decoder 1 can output the accent feature C. The electronic device can determine the loss function based on the loss between the accent feature A and the accent feature C, and train the accent replacement module based on the loss function. The above steps are repeated based on multiple groups of samples until the training of the accent decoder 1 converges, that is, the training of the accent decoder 1 in the accent replacement module is completed.

[0109] Among them, the electronic device determines the second voice based on the timbre features, the second accent features and the accent conversion model, which can be specifically: obtaining a blurred image with Gaussian noise added, processing the timbre features, the second accent features and the blurred image based on the diffusion module in the accent conversion model to obtain a mel-spectrogram image, and obtaining the second voice based on the mel-spectrogram image.

[0110] The blurred image may be an image to which Gaussian noise is added. For example, the electronic device may add Gaussian noise to any image until the image is completely distorted, thereby obtaining a blurred image. The electronic device may also obtain a blurred image based on any other feasible implementation method, which is not limited in the present embodiment.

[0111] Optionally, the electronic device processes the timbre features, the second accent features and the blurred image based on the diffusion module in the accent conversion module to obtain a mel-spectrogram image. Specifically, for the first processing, the Gaussian noise in the blurred image is denoised based on the diffusion module, the timbre features and the second accent features to obtain the first image to be processed. For the i-th processing, the Gaussian noise in the i-1 images to be processed is denoised based on the diffusion module, the timbre features and the second accent features to obtain the i-th image to be processed. This continues until the Gaussian noise in the i-th image to be processed is less than or equal to a preset threshold to obtain a mel-spectrogram image.

[0112] Here, i is 2, 3, ..., N, and N is the number of times the diffusion module processes. For example, if the step size of the diffusion module is 1000 times, then N can be 1000.

[0113] For example, in actual application, the diffusion module can gradually restore (denoise) the blurred image based on the timbre features and the second accent features. After multiple processing steps, a clear image can be obtained. Since the image is restored based on the timbre features and the second accent features, the image can be a mel-spectrogram.

[0114] Next, the processing process of the diffusion module will be described with reference to FIG11 .

[0115] FIG11 is a schematic diagram of the processing process of a diffusion module provided by an embodiment of the present disclosure. Referring to FIG11 , it includes: timbre features, blurred images, second accent features, processing information (the information for the first time is: the first processing, which can be encoded based on the encoder and the result can be spliced ​​with the above information) and a diffusion module. Among them, the electronic device (not shown in FIG11 ) can input the timbre features, blurred images, the first processing information and the second accent features to the module of step 1 in the diffusion module (the processing module of step 1), and the module of step 1 can output the denoised image 1 to be processed.

[0116] Referring to Figure 11 , the electronic device can input the timbre characteristics, image 2 to be processed, the information processed for the second time, and the second accent characteristics into the module in step 2 of the diffusion module. The module in step 2 can then output image 2 to be processed after denoising image 1. The electronic device can repeat the above steps until the module in step 1000 of the diffusion module inputs the timbre characteristics, image 999 to be processed, the information processed for the 1000th time, and the second accent characteristics. The module in step 1000 can then denoise image 999 to obtain a mel-spectrogram image. In this way, based on multiple denoising steps, the accuracy of the mel-spectrogram image can be improved, thereby improving the accuracy of the second speech.

[0117] It should be noted that after the electronic device obtains the mel-spectrum image, it can obtain the second speech associated with the mel-spectrum image based on any feasible implementation method, and the embodiments of the present disclosure are not limited to this.

[0118] The disclosed embodiments provide a method for determining a second speech. An electronic device, based on a timbre separation module in an accent conversion model, performs convolution processing on a first speech to obtain timbre features of the first speech. A target convolution layer is determined among multiple convolution layers of an automatic speech recognition module. The first speech is convoluted based on the automatic speech recognition module, and the convolution result output by the target convolution layer is determined as a first accent feature. Based on an accent replacement module and an identifier of a second accent, the first accent feature is convoluted to obtain a second accent feature. The second speech is determined based on the timbre features, the second accent features, and the accent conversion model. In this way, because the second accent features do not include timbre information, the second accent features have a higher accuracy. Therefore, the electronic device can accurately perform accent conversion processing on the first speech, improving the effect and accuracy of the second speech.

[0119] Based on any of the above embodiments, before determining the second voice based on the timbre features, the second accent features and the accent conversion model, the above voice processing method may further include a method for voice style conversion. Below, the method for voice style conversion is described in detail in conjunction with Figure 12.

[0120] FIG12 is a schematic diagram of a method for voice style conversion provided by an embodiment of the present disclosure. Referring to FIG12 , the method process includes:

[0121] S1201: Obtain a target voice style.

[0122] Optionally, the target voice style may be the voice style to be converted. For example, if the electronic device converts the voice style of the first voice into voice style 1, voice style 1 may be the target voice style; and if the electronic device converts the voice style of the first voice into voice style 2, voice style 2 may be the target voice style.

[0123] It should be noted that the voice style may indicate any voice information unrelated to timbre and accent, such as the rhythm, prosody, stress and intonation of the voice, and the embodiments of the present disclosure are not limited to this.

[0124] It should be noted that the electronic device can obtain the target voice style based on any feasible implementation method (for example, the method by which the electronic device obtains the target voice style is similar to the method by which the second accent is obtained, which will not be described in detail in the embodiments of the present disclosure), and the embodiments of the present disclosure do not limit this.

[0125] S1202: Determine a target speech style replacement module corresponding to a target speech style among multiple speech style replacement modules in the accent conversion model.

[0126] Optionally, the accent conversion model may include multiple voice style replacement modules, each of which may correspond to a voice style. For example, a voice style replacement module may correspond to voice style 1, and the electronic device may generate voice in voice style 1 based on the information output by the voice style replacement module; a voice style replacement module may correspond to voice style 2, and the electronic device may generate voice in voice style 2 based on the information output by the voice style replacement module.

[0127] The voice style replacement module is used to convert the voice style information in the accent feature. For example, the accent feature may include accent information and voice style information. The accent replacement module can convert the accent information in the accent feature, and the voice style replacement module can convert the voice style information in the accent feature.

[0128] The target voice style replacement module may be a module corresponding to the target voice style. For example, voice style 1 corresponds to voice style replacement module A, and voice style 2 corresponds to voice style replacement module B. If the electronic device determines that the target voice style is voice style 1, the electronic device may determine that the target voice style replacement module is voice style replacement module A. If the electronic device determines that the target voice style is voice style 2, the electronic device may determine that the target voice style replacement module is voice style replacement module B.

[0129] It should be noted that the method for an electronic device to determine a target voice style replacement module is similar to the method for an electronic device to determine a target accent decoder, and the embodiments of the present disclosure are not limited thereto.

[0130] It should be noted that the model structure of the voice style replacement module may be the same as the model structure of the accent style replacement module, and this embodiment of the present disclosure does not limit this.

[0131] Optionally, the training process of the voice style replacement module may be similar to the training process of the accent style replacement module. The following describes the training process of the voice style replacement module in detail through specific examples.

[0132] The electronic device can process the Mandarin "Hello" with voice style 1 based on the ASR model to obtain voice style feature A (the output result of the middle convolution layer). The electronic device can process the Mandarin "Hello" with voice style 2 based on the ASR model to obtain voice style feature B (the output result of the middle convolution layer). In this way, the electronic device can obtain a group of samples, which may include voice style feature A and voice style feature B.

[0133] If the voice style replacement module is trained as a voice style replacement module of voice style 2, the electronic device can perform convolution processing on the voice style feature A based on the voice style replacement module to obtain the voice style feature C. The electronic device can construct a loss function based on the loss between the voice style feature B and the voice style feature C, and train the voice style replacement module based on the loss function. The above steps are repeated based on multiple groups of samples until the training of the voice style replacement module converges and the training of the voice style replacement module is completed.

[0134] S1203: Perform convolution processing on the second accent feature based on the target speech style replacement module.

[0135] Optionally, after the electronic device determines the target voice style replacement module, it can perform convolution processing on the second accent feature based on the target voice style replacement module, and then convert the voice style information in the second accent feature. The convolved second accent feature can participate in the generation process of the second voice. In this way, the electronic device can not only convert the accent of the first voice into the second accent, but also convert the voice style of the first voice into the target voice style.

[0136] The disclosed embodiments provide a method for voice style conversion. An electronic device can obtain a target voice style, determine a target voice style replacement module corresponding to the target voice style from among multiple voice style replacement modules in an accent conversion model, and perform convolution processing on a second accent feature based on the target voice style replacement module. This allows the second accent feature to include not only information about the second accent but also information about the target voice style, thereby improving the functionality of the accent conversion model, the quality of the second voice, and the accuracy of the second voice.

[0137] Based on any of the above embodiments, the process of the speech processing method will be described below with reference to FIG13 .

[0138] FIG13 is a schematic diagram of a speech processing method provided by an embodiment of the present disclosure. Referring to FIG13 , the method includes speech with a first accent, a timbre separation module, an automatic speech recognition module, an accent replacement module, a voice style replacement module, and a diffusion module. The timbre classification module, the automatic speech recognition module, the accent replacement module, the voice style replacement module, and the diffusion module may be modules within an accent conversion model.

[0139] Referring to Figure 13 , the electronic device can convert speech with a first accent into a mel-spectrogram and input the mel-spectrogram into the timbre separation module and the automatic speech recognition module. The timbre separation module can process the mel-spectrogram to obtain timbre features of the first speech. The automatic speech recognition module can process the mel-spectrogram, determine the output of the intermediate convolutional layer as the first accent feature, and input the first accent feature into the accent replacement module.

[0140] Referring to FIG13 , the accent replacement module can perform convolution processing on the first accent feature and input the convolution result to the voice style replacement module. The voice style replacement module can then perform convolution processing on the convolution result to obtain a second accent feature. The electronic device can obtain a Gaussian blurred image and convert the timbre feature, the second accent feature, and the time step feature (which can be obtained based on the encoder, i.e., the information processed for the first time, the information processed for the second time, ..., and the information processed for the 1000th time in FIG11 ).

[0141] Referring to Figure 13 , the diffusion module can perform multiple denoising operations on the Gaussian module image based on the second accent features, timbre features, and time step features to generate a predicted mel-spectrogram. Based on the predicted mel-spectrogram, the electronic device can generate speech with the second accent, where the style of the speech with the second accent is the same as the style of the speech generated by the speech style replacement module.

[0142] In this way, the accent conversion model can not only realize "one-to-one", "one-to-many" and "many-to-many" accent conversion processing, but also replace the voice style of the original speech. Moreover, since the automatic speech recognition module can remove the timbre information in the speech information, the timbre leakage can be avoided during the accent conversion process, thereby improving the accuracy of the accent conversion. In this way, the accent conversion model can retain the timbre of the original speech and realize the conversion of accent or voice style (that is, without changing the speaker's timbre, but changing the speaker's accent or speaking style), thereby improving the accuracy of the accent conversion and the effect of the accent conversion.

[0143] FIG14 is a schematic diagram of the structure of a speech processing device provided by an embodiment of the present disclosure. Referring to FIG14 , the speech processing device 140 includes a first acquisition module 141, a second acquisition module 142, and a determination module 143, wherein:

[0144] The first acquisition module 141 is used to acquire a first speech with a first accent;

[0145] The second acquisition module 142 is used to obtain an identifier of a second accent;

[0146] The determination module 143 is used to determine the second voice based on the accent conversion model, the identification of the first voice and the second accent, wherein the voice content of the second voice is the same as the voice content of the first voice, and the accent of the second voice is the second accent; wherein the accent conversion model includes an accent replacement module, and the accent replacement module is used to convert the accent information in other voice information except the timbre information, and the accent replacement module includes an accent encoder and multiple accent decoders connected to the accent encoder, and the multiple accent decoders correspond one-to-one to multiple accents.

[0147] According to one or more embodiments of the present disclosure, the determining module 143 is specifically configured to:

[0148] performing convolution processing on the first speech based on the timbre separation module in the accent conversion model to obtain timbre features of the first speech;

[0149] performing convolution processing on the first speech based on the automatic speech recognition module in the accent conversion model to obtain a first accent feature of the first speech;

[0150] The second speech is determined based on the timbre feature, the first accent feature, the identifier of the second accent, and the accent conversion model.

[0151] According to one or more embodiments of the present disclosure, the determining module 143 is specifically configured to:

[0152] performing convolution processing on the first accent feature based on the accent replacement module and the identifier of the second accent to obtain a second accent feature;

[0153] The second voice is determined based on the timbre feature, the second accent feature, and the accent conversion model.

[0154] According to one or more embodiments of the present disclosure, the determining module 143 is specifically configured to:

[0155] Get the blurred image with Gaussian noise added;

[0156] processing the timbre feature, the second accent feature, and the blurred image based on the diffusion module in the accent conversion model to obtain a mel-spectrogram image;

[0157] The second speech is obtained based on the mel-spectrogram image.

[0158] According to one or more embodiments of the present disclosure, the determining module 143 is specifically configured to:

[0159] For the first treatment;

[0160] performing denoising on Gaussian noise in the blurred image based on the diffusion module, the timbre feature, and the second accent feature to obtain a first image to be processed;

[0161] For the i-th processing;

[0162] performing denoising on the Gaussian noise in the (i-1)th image to be processed based on the diffusion module, the timbre feature, and the second accent feature to obtain the i-th image to be processed, until the Gaussian noise in the i-th image to be processed is less than or equal to a preset threshold, thereby obtaining the mel-spectrogram image;

[0163] Here, the value of i is 2, 3, ..., N in sequence, and N is the number of times the diffusion module processes.

[0164] According to one or more embodiments of the present disclosure, the determining module 143 is specifically configured to:

[0165] Based on the identifier of the second accent, determining, in the accent replacement module, a target accent decoder corresponding to the identifier of the second accent;

[0166] performing convolution processing on the first accent feature based on the accent encoder to obtain a convolution feature;

[0167] The convolution feature is convolved based on the target accent decoder to obtain the second accent feature.

[0168] According to one or more embodiments of the present disclosure, the determining module 143 is specifically configured to:

[0169] Determining a target convolutional layer among a plurality of convolutional layers of the automatic speech recognition module;

[0170] Convolution processing is performed on the first speech based on the automatic speech recognition module, and the convolution result output by the target convolution layer is determined as the first accent feature.

[0171] According to one or more embodiments of the present disclosure, the first acquisition module 141 is specifically configured to:

[0172] Displaying a voice acquisition page, wherein the voice acquisition page includes a voice acquisition control;

[0173] In response to a touch operation on the voice acquisition control, displaying a voice page, wherein the voice page includes a plurality of voices;

[0174] In response to a touch operation on the multiple voices, the voice associated with the touch operation is determined as the first voice.

[0175] According to one or more embodiments of the present disclosure, the first acquisition module 141 is specifically configured to:

[0176] Displaying a voice collection page, wherein the voice collection page includes a voice recording control;

[0177] In response to a touch operation on the voice recording control, voice recording is performed to obtain the first voice.

[0178] According to one or more embodiments of the present disclosure, the determining module 143 is further configured to:

[0179] Acquire the target voice style;

[0180] Determining a target speech style replacement module corresponding to the target speech style among a plurality of speech style replacement modules in the accent conversion model;

[0181] performing convolution processing on the second accent feature based on the target speech style replacement module;

[0182] The voice style replacement module is used to convert the voice style information in the accent feature.

[0183] The speech processing device provided in the embodiment of the present disclosure can be used to execute the technical solution of the above-mentioned method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.

[0184] FIG15 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. Please refer to FIG15 , which shows a schematic diagram of the structure of an electronic device 1500 suitable for implementing an embodiment of the present disclosure. The electronic device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., as well as fixed terminals such as digital TVs, desktop computers, etc. The electronic device shown in FIG15 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.

[0185] As shown in FIG15 , the electronic device 1500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1502 or a program loaded from a storage device 1508 into a random access memory (RAM) 1503. Various programs and data required for the operation of the electronic device 1500 are also stored in the RAM 1503. The processing device 1501, the ROM 1502, and the RAM 1503 are connected to each other via a bus 1504. An input / output (I / O) interface 1505 is also connected to the bus 1504.

[0186] Typically, the following devices may be connected to the I / O interface 1505: an input device 1506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1509. The communication device 1509 may allow the electronic device 1500 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 15 shows an electronic device 1500 having various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0187] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 1509, or installed from the storage device 1508, or installed from the ROM 1502. When the computer program is executed by the processing device 1501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0188] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0189] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0190] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.

[0191] An embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the speech processing method that may be involved in various embodiments above is implemented.

[0192] An embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the various speech processing methods that may be involved in the above embodiments.

[0193] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).

[0194] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0195] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."

[0196] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0197] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0198] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0199] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0200] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0201] For example, in response to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thereby, the user can independently choose whether to provide personal information to software or hardware such as an electronic device, application, server or storage medium that performs the operation of the technical solution of the present disclosure based on the prompt message. As an optional but non-limiting implementation method, in response to receiving an active request from the user, the method of sending a prompt message to the user can be, for example, a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0202] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0203] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws and regulations. Data may include information, parameters and messages, such as flow switching indication information.

[0204] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0205] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0206] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A speech processing method, comprising: Obtain a first speech with a first accent; Get the identity of the second accent; Determine a second voice based on the accent conversion model, the first voice, and the identifier of the second accent, wherein the voice content of the second voice is the same as the voice content of the first voice, and the accent of the second voice is the second accent; Among them, the accent conversion model includes an accent replacement module, which is used to convert accent information in other speech information except timbre information. The accent replacement module includes an accent encoder and multiple accent decoders connected to the accent encoder, and the multiple accent decoders correspond one-to-one to multiple accents.

2. The method according to claim 1, wherein: The determining the second voice based on the accent conversion model, the first voice and the identifier of the second accent includes: Based on the timbre separation module in the accent conversion model, convolution processing is performed on the first speech to obtain the timbre characteristics of the first speech; Based on the automatic speech recognition module in the accent conversion model, performing convolution processing on the first speech to obtain a first accent feature of the first speech; The second speech is determined based on the timbre feature, the first accent feature, an identifier of the second accent, and the accent conversion model.

3. The method according to claim 2, wherein: The determining the second voice based on the timbre feature, the first accent feature, the identifier of the second accent and the accent conversion model includes: Based on the accent replacement module and the identifier of the second accent, performing convolution processing on the first accent feature to obtain a second accent feature; The second voice is determined based on the timbre feature, the second accent feature and the accent conversion model.

4. The method according to claim 3, wherein: Determining the second voice based on the timbre feature, the second accent feature, and the accent conversion model includes: Get the blurred image with Gaussian noise added; Based on the diffusion module in the accent conversion model, the timbre feature, the second accent feature and the blurred image are processed to obtain a mel-spectrogram image; The second speech is obtained based on the mel-spectrogram image.

5. The method according to claim 4, wherein: The step of processing the timbre feature, the second accent feature and the blurred image based on the diffusion module in the accent conversion model to obtain a mel spectrum image includes: For the first treatment; Based on the diffusion module, the timbre feature and the second accent feature, the Gaussian noise in the blurred image is Perform denoising to obtain the first image to be processed; For the i-th processing; Based on the diffusion module, the timbre feature and the second accent feature, denoising the Gaussian noise in the (i-1)th image to be processed to obtain the (i)th image to be processed, until the Gaussian noise in the (i)th image to be processed is less than or equal to a preset threshold, thereby obtaining the Mel-spectrogram image; The i is 2, 3, ..., N in sequence, and N is the number of processing times of the diffusion module.

6. The method according to any one of claims 3 to 5, wherein: The step of performing convolution processing on the first accent feature based on the accent replacement module and the identifier of the second accent to obtain the second accent feature includes: Based on the identifier of the second accent, determining, in the accent replacement module, a target accent decoder corresponding to the identifier of the second accent; Performing convolution processing on the first accent feature based on the accent encoder to obtain a convolution feature; The convolution feature is convolved based on the target accent decoder to obtain the second accent feature.

7. The method according to any one of claims 2 to 6, wherein: The automatic speech recognition module based on the accent conversion model performs convolution processing on the first speech to obtain a first accent feature of the first speech, including: Determining a target convolutional layer among a plurality of convolutional layers of the automatic speech recognition module; The first speech is subjected to convolution processing based on the automatic speech recognition module, and the convolution result output by the target convolution layer is determined as the first accent feature.

8. The method according to any one of claims 1 to 7, wherein: The step of acquiring the first speech with the first accent includes: Displaying a voice acquisition page, wherein the voice acquisition page includes a voice acquisition control; In response to a touch operation on the voice acquisition control, displaying a voice page, wherein the voice page includes a plurality of voices; In response to a touch operation on the multiple voices, the voice associated with the touch operation is determined as the first voice.

9. The method according to any one of claims 1 to 7, wherein: The step of acquiring the first speech with the first accent includes: Displaying a voice collection page, wherein the voice collection page includes a voice recording control; In response to a touch operation on the voice recording control, voice recording is performed to obtain the first voice.

10. The method according to any one of claims 3 to 9, wherein: Before determining the second voice based on the timbre feature, the second accent feature and the accent conversion model, the method further includes: Acquire the target voice style; Determining a target speech style replacement module corresponding to the target speech style among a plurality of speech style replacement modules in the accent conversion model; performing convolution processing on the second accent feature based on the target speech style replacement module; The speech style replacement module is used to convert the speech style information in the accent feature.

11. A speech processing device, comprising a first acquisition module, a second acquisition module and a determination module, wherein: The first acquisition module is configured to acquire a first speech with a first accent; The second acquisition module is configured to acquire an identifier of a second accent; The determination module is configured to determine the second voice based on the accent conversion model, the first voice and the identification of the second accent, the voice content of the second voice is the same as the voice content of the first voice, and the accent of the second voice is the second accent; wherein the accent conversion model includes an accent replacement module, the accent replacement module is used to convert accent information in other voice information except timbre information, the accent replacement module includes an accent encoder and multiple accent decoders connected to the accent encoder, and the multiple accent decoders correspond one-to-one to multiple accents.

12. An electronic device comprising a processor and a memory, wherein: The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the speech processing method according to any one of claims 1 to 10.

13. A computer-readable storage medium storing computer-executable instructions, wherein: When the processor executes the computer-executable instruction, the speech processing method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Method and device for generating audio, equipment and medium

    CN111899719A

  • Voice conversion method, system, electronic equipment and readable storage medium

    CN113571039A

  • Voice conversion method and related equipment

    CN114299908A

  • Image Enhancement via Iterative Refinement based on Machine Learning Models

    US20230153959A1

  • Generating data items using off-the-shelf guided generative diffusion processes

    WO2023144386A1