A personalized speech synthesis method and terminal based on artificial intelligence
By collecting and analyzing standard verification speech in a multi-person voice environment, calculating the short-time energy and spectrum conversion resonance peaks of the speech, and using artificial intelligence technology to screen and synthesize speech with similar timbre, the problem of difficulty in distinguishing speech with similar timbre in a multi-person voice environment is solved, and fast and effective personalized speech synthesis is achieved.
Patent Information
- Application Number
- CN202410681389.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-29
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-05-29
AI Technical Summary
Existing technologies have difficulty in quickly and effectively automatically recognizing and distinguishing voices with similar timbres in a multi-person voice environment, making personalized speech synthesis difficult.
By collecting standard verification voices of multiple voice output objects, calculating the short-time energy of the voice, screening similar verification voices, performing spectrum conversion and resonance peak comparison, screening voices with similar timbre, and using artificial intelligence technology for personalized voice synthesis processing.
It can quickly and effectively automatically identify and distinguish voices with similar timbres in a multi-person voice environment, achieving fast and effective processing of personalized speech synthesis.
Smart Images

Figure CN118471188B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech synthesis, and in particular relates to a personalized speech synthesis method and terminal based on artificial intelligence. Background Art
[0002] As a vital information carrier, speech has played a crucial role in the development and progress of human civilization. Speech synthesis, a crucial component of voice interaction, is now widely used in daily life and production, primarily in audiobooks, voice interaction, and entertainment.
[0003] In the existing technology, personalized speech synthesis can only perform corresponding personalized speech synthesis processing according to human needs. However, in some multi-person speech environments (for example, the game voice environment of a multi-person game team), if at least two people have similar timbres, it is difficult to distinguish the voices of different people. In this case, the existing speech synthesis technology cannot perform fast and effective automatic recognition, judgment and speech synthesis processing. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a personalized speech synthesis method and terminal based on artificial intelligence, aiming to solve the technical problems existing in the existing technology mentioned in the background technology.
[0005] The embodiment of the present invention is implemented as follows:
[0006] A personalized speech synthesis method based on artificial intelligence, the method specifically comprising the following steps:
[0007] Determining a plurality of voice output objects, presenting standard verification information to the plurality of voice output objects, and collecting standard verification voices of the plurality of voice output objects in response to the standard verification information;
[0008] Analyzing a plurality of the standard verification voices, calculating a plurality of corresponding voice short-time energies, comparing the plurality of the voice short-time energies, and screening a plurality of similar verification voices;
[0009] Performing spectrum conversion on the multiple similar verification voices, comparing multiple initial formant positions, multiple initial formant bandwidths, and multiple initial formant amplitudes, and screening multiple voices with similar timbre;
[0010] screening similar short-time energies corresponding to multiple voices with similar timbre from the multiple voice short-time energies, and selecting multiple target similar voices according to the multiple similar short-time energies;
[0011] From the multiple speech output objects, target output objects corresponding to multiple target similar speech are selected, and based on artificial intelligence technology, different speech synthesis processing is performed on the multiple target output objects.
[0012] As a further limitation of the technical solution of the embodiment of the present invention, the steps of determining multiple voice output objects, presenting standard verification information to the multiple voice output objects, and collecting standard verification voices of the multiple voice output objects in response to the standard verification information specifically include the following steps:
[0013] determining a plurality of speech output objects;
[0014] generating and presenting standard verification information to a plurality of said voice output objects;
[0015] Providing voice response prompts to the plurality of voice output objects;
[0016] Collect a plurality of standard verification voices of the voice output objects in response to the standard verification information.
[0017] As a further limitation of the technical solution of the embodiment of the present invention, the analyzing of the plurality of standard verification voices, calculating the plurality of corresponding voice short-time energies, comparing the plurality of voice short-time energies, and screening the plurality of similar verification voices specifically comprises the following steps:
[0018] intercepting a plurality of speech signal samples from a plurality of the standard verification voices;
[0019] Verify multiple speech signal samples corresponding to the speech according to multiple standards, and calculate multiple corresponding speech short-time energies;
[0020] Calculating the short-time energy difference between the plurality of speech short-time energies;
[0021] Comparing the multiple short-time energy differences with a preset standard energy difference to screen multiple similar energy differences;
[0022] Based on the multiple similar energy differences, multiple similar verification voices are screened from the multiple standard verification voices.
[0023] As a further limitation of the technical solution of the embodiment of the present invention, the calculation formula of the multiple speech short-time energies is:
[0024]
[0025] Among them, i represents i standard verification speech, E i is the short-time energy of the i-th standard verification speech, n represents n speech signal samples, x i (n) is the n speech signal samples of the i-th standard verification speech, and L is the window length.
[0026] As a further limitation of the technical solution of the embodiment of the present invention, performing spectrum conversion on the multiple similar verification voices, comparing multiple initial formant positions, multiple initial formant bandwidths, and multiple initial formant amplitudes, and screening multiple voices with similar timbre specifically includes the following steps:
[0027] Performing spectrum conversion on the plurality of similar verification voices to obtain a plurality of similar voice spectra;
[0028] Marking a plurality of initial formants of the spectrum in the plurality of similar speech spectra according to a preset number of formants;
[0029] comparing a plurality of initial resonance peak positions, a plurality of initial resonance peak bandwidths, and a plurality of initial resonance peak amplitudes of the plurality of initial resonance peaks of the spectrum, and recording a resonance peak comparison result;
[0030] According to the formant comparison result, a plurality of voices with similar timbre are screened from the plurality of similar verification voices.
[0031] As a further limitation of the technical solution of the embodiment of the present invention, the screening of similar short-time energies corresponding to multiple voices with similar timbre from the multiple voice short-time energies, and selecting multiple target similar voices according to the multiple similar short-time energies specifically includes the following steps:
[0032] Screening similar short-time energies corresponding to multiple voices with similar timbre from the multiple voice short-time energies;
[0033] Comparing the multiple similar short-time energies and recording the energy comparison results;
[0034] selecting a plurality of target short-time energies from the plurality of similar short-time energies according to the energy comparison result;
[0035] A plurality of target similar voices corresponding to a plurality of target short-time energies are selected from the plurality of voices with similar timbre.
[0036] As a further limitation of the technical solution of the embodiment of the present invention, the selecting, from the plurality of speech output objects, target output objects corresponding to the target similar speech, and performing different speech synthesis processing on the plurality of target output objects based on artificial intelligence technology specifically comprises the following steps:
[0037] Selecting a plurality of target output objects corresponding to a plurality of target similar voices from the plurality of voice output objects;
[0038] Select multiple target timbre models from the preset timbre model database;
[0039] Based on artificial intelligence technology, different speech synthesis processes are performed on the multiple target output objects according to the multiple target timbre models.
[0040] A personalized speech synthesis terminal based on artificial intelligence, comprising a verification speech acquisition module, a short-time energy comparison module, a spectrum conversion analysis module, a target speech selection module, and a speech synthesis processing module, wherein:
[0041] a verification voice collection module, configured to determine a plurality of voice output objects, present standard verification information to the plurality of voice output objects, and collect standard verification voices of the plurality of voice output objects in response to the standard verification information;
[0042] A short-time energy comparison module is used to analyze the plurality of standard verification voices, calculate the short-time energies of the plurality of corresponding voices, compare the short-time energies of the plurality of voices, and screen a plurality of similar verification voices;
[0043] a spectrum conversion analysis module, configured to perform spectrum conversion on the plurality of similar verification voices, compare a plurality of initial formant positions, a plurality of initial formant bandwidths, and a plurality of initial formant amplitudes, and screen a plurality of voices with similar timbre;
[0044] a target speech selection module, configured to screen similar short-time energies corresponding to multiple speech with similar timbre from the multiple speech short-time energies, and select multiple target similar speech according to the multiple similar short-time energies;
[0045] The speech synthesis processing module is used to select target output objects corresponding to multiple target similar speech from the multiple speech output objects, and perform different speech synthesis processing on the multiple target output objects based on artificial intelligence technology.
[0046] As a further limitation of the technical solution of the embodiment of the present invention, the short-time energy comparison module specifically includes:
[0047] A signal interception unit, configured to intercept a plurality of voice signal samples from each of the plurality of standard verification voices;
[0048] An energy calculation unit, configured to verify a plurality of speech signal samples corresponding to the speech according to a plurality of the standards, and calculate a plurality of corresponding speech short-time energies;
[0049] An energy difference calculation unit, configured to calculate a short-time energy difference between a plurality of the short-time energies of the speech;
[0050] an energy difference comparison unit, configured to compare the plurality of short-time energy differences with a preset standard energy difference, and screen a plurality of similar energy differences;
[0051] The first speech screening unit is used to screen a plurality of similar verification speech from a plurality of the standard verification speech according to a plurality of the similar energy differences.
[0052] As a further limitation of the technical solution of the embodiment of the present invention, the spectrum conversion analysis module specifically includes:
[0053] a spectrum conversion unit, configured to perform spectrum conversion on the plurality of similar verification voices to obtain a plurality of similar voice spectra;
[0054] A formant marking unit, configured to mark a plurality of initial formants of the spectrum in the plurality of similar speech spectra according to a preset number of formants;
[0055] a resonance peak comparison unit, configured to compare a plurality of initial resonance peak positions, a plurality of initial resonance peak bandwidths, and a plurality of initial resonance peak amplitudes of the plurality of initial resonance peaks of the spectrum, and record the resonance peak comparison result;
[0056] The second speech screening unit is used to screen multiple speech with similar timbre from the multiple similar verification speech according to the formant comparison result.
[0057] Compared with the prior art, the present invention has the following beneficial effects:
[0058] The embodiment of the present invention determines multiple speech output objects, collects multiple standard verification speech; calculates multiple speech short-time energies, screens multiple similar verification speech; performs spectrum conversion and comparison on multiple similar verification speech, screens multiple speech with similar timbre; selects multiple target similar speech; and based on artificial intelligence technology, performs different speech synthesis processing on multiple target output objects. It is capable of collecting multiple standard verification speech of speech output objects, calculating multiple speech short-time energies, performing spectrum conversion and initial resonance peak comparison, screening multiple speech with similar timbre, selecting multiple target output objects, and performing different speech synthesis processing. Thus, it is possible to quickly and effectively automatically perform recognition judgment and speech synthesis processing in a multi-person speech environment with similar timbre, making it easier to distinguish the speech of different people. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 A flowchart of a personalized speech synthesis method based on artificial intelligence provided by an embodiment of the present invention is shown;
[0060] Figure 2 A flowchart of collecting multiple standard verification voices in the method provided by an embodiment of the present invention is shown;
[0061] Figure 3 A flowchart of screening multiple similar verification voices in the method provided by an embodiment of the present invention is shown;
[0062] Figure 4 A flowchart of screening multiple voices with similar timbre in a method provided by an embodiment of the present invention is shown;
[0063] Figure 5 A flowchart of selecting multiple target similar voices in the method provided by an embodiment of the present invention is shown;
[0064] Figure 6 A flowchart of object speech synthesis processing in the method provided by an embodiment of the present invention is shown;
[0065] Figure 7 The following is a diagram showing the application architecture of an artificial intelligence-based personalized speech synthesis terminal provided by an embodiment of the present invention;
[0066] Figure 8 The following is a diagram showing the application architecture of a short-time energy comparison module in a terminal provided by an embodiment of the present invention;
[0067] Figure 9 The diagram shows an application architecture of a spectrum conversion analysis module in a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0068] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0069] It is understandable that in the existing technology, personalized speech synthesis can only perform corresponding personalized speech synthesis processing based on human needs. In some multi-person voice environments (for example, the game voice environment of a multi-person game team), if at least two people have similar timbres, it is difficult to distinguish the voices of different people. In this case, the existing speech synthesis technology cannot perform fast and effective automatic recognition, judgment and speech synthesis processing.
[0070] To solve the above problems, an embodiment of the present invention discloses an artificial intelligence-based personalized speech synthesis method and terminal, which determines multiple speech output objects, displays standard verification information to the multiple speech output objects, and collects standard verification speech of the multiple speech output objects in response to the standard verification information; analyzes the multiple standard verification speech, calculates multiple corresponding speech short-time energies, compares the multiple speech short-time energies, and screens multiple similar verification speech; performs spectrum conversion on the multiple similar verification speech, compares multiple initial resonance peak positions, multiple initial resonance peak bandwidths and multiple initial resonance peak amplitudes, and screens multiple speech with similar timbre; from the multiple speech short-time energies, screens similar short-time energies corresponding to multiple speech with similar timbre, and selects multiple target similar speech according to the multiple similar short-time energies; selects target output objects corresponding to multiple target similar speech from the multiple speech output objects, and performs different speech synthesis processing on the multiple target output objects based on artificial intelligence technology. It can collect standard verification voices of multiple voice output objects, calculate the short-time energy of multiple voices, compare the spectrum conversion with the initial resonance peak, screen multiple voices with similar timbres, select multiple target output objects, and perform different voice synthesis processing. In this way, it can quickly and effectively automatically perform recognition judgment and voice synthesis processing when there are similar timbres in a multi-person voice environment, making it easier to distinguish the voices of different people.
[0071] Specifically, Figure 1 The figure shows a flow chart of a personalized speech synthesis method based on artificial intelligence provided by an embodiment of the present invention.
[0072] In a preferred embodiment of the present invention, a personalized speech synthesis method based on artificial intelligence is provided, wherein the method specifically comprises the following steps:
[0073] Step S101: determine a plurality of voice output objects, display standard verification information to the plurality of voice output objects, and collect standard verification voices of the plurality of voice output objects in response to the standard verification information.
[0074] In an embodiment of the present invention, multiple voice output addresses are obtained in a multi-person voice environment, and multiple corresponding voice output objects are determined according to the multiple voice output addresses. By generating standard verification information, the standard verification information is sent to the multiple voice output objects, and voice response prompts are given to the multiple voice output objects. The multiple voice output objects can read the standard verification information according to the voice response prompts, and in the process of the multiple voice output objects reading the standard verification information, voice collection is performed on the multiple voice output objects to obtain multiple standard verification voices.
[0075] Specifically, Figure 2 The flowchart of collecting multiple standard verification voices in the method provided by the embodiment of the present invention is shown.
[0076] In a preferred embodiment of the present invention, determining multiple voice output objects, presenting standard verification information to the multiple voice output objects, and collecting standard verification voices of the multiple voice output objects in response to the standard verification information specifically include the following steps:
[0077] Step S1011: Determine multiple voice output objects.
[0078] Step S1012: Generate and display standard verification information to the plurality of voice output objects.
[0079] Step S1013: Providing voice response prompts to the multiple voice output objects.
[0080] Step S1014: Collect standard verification voices of the plurality of voice output objects in response to the standard verification information.
[0081] Furthermore, the personalized speech synthesis method based on artificial intelligence further includes the following steps:
[0082] Step S102: Analyze the plurality of standard verification voices, calculate the corresponding short-time energies of the voices, compare the short-time energies of the voices, and select a plurality of similar verification voices.
[0083] In an embodiment of the present invention, n speech signal samples are intercepted from multiple standard verification voices, and multiple corresponding speech short-time energies are calculated according to the n speech signal samples corresponding to the multiple standard verification voices. Then, the multiple speech short-time energies are compared with each other, and the short-time energy differences between the multiple speech short-time energies are calculated. Then, with a preset standard energy difference as a reference, the multiple short-time energy differences are compared with the standard energy difference, and the short-time energy differences that are smaller than the standard energy difference are marked as similar energy differences. In this way, multiple similar energy differences can be screened from the multiple short-time energy differences, and then, based on the multiple similar energy differences, multiple corresponding similar verification voices are screened from the multiple standard verification voices. Specifically, the calculation formula for calculating the multiple corresponding speech short-time energies is:
[0084]
[0085] Among them, i represents i standard verification speech, E i is the short-time energy of the i-th standard verification speech, n represents n speech signal samples, x i (n) is the n speech signal samples of the i-th standard verification speech, and L is the window length.
[0086] Specifically, Figure 3 A flowchart of screening multiple similar verification voices in the method provided by an embodiment of the present invention is shown.
[0087] Among them, in the preferred embodiment provided by the present invention, the analyzing of the plurality of standard verification voices, calculating the plurality of corresponding voice short-time energies, comparing the plurality of voice short-time energies, and screening the plurality of similar verification voices specifically includes the following steps:
[0088] Step S1021: intercept a plurality of voice signal samples from each of the plurality of standard verification voices.
[0089] Step S1022: Verify multiple speech signal samples corresponding to the speech according to multiple standards, and calculate multiple corresponding speech short-time energies.
[0090] Step S1023: Calculate the short-time energy difference between the multiple speech short-time energies.
[0091] Step S1024: compare the multiple short-time energy differences with a preset standard energy difference, and screen multiple similar energy differences.
[0092] Step S1025: Filter multiple similar verification voices from the multiple standard verification voices based on the multiple similar energy differences.
[0093] Furthermore, the personalized speech synthesis method based on artificial intelligence further includes the following steps:
[0094] Step S103 : performing spectrum conversion on the multiple similar verification voices, comparing multiple initial formant positions, multiple initial formant bandwidths, and multiple initial formant amplitudes, and screening multiple voices with similar timbre.
[0095] In an embodiment of the present invention, multiple similar speech spectra are obtained by performing spectrum conversion on multiple similar verification speech, and then multiple spectral initial resonance peaks are identified and marked in the multiple similar speech spectra according to a preset number of resonance peaks. The multiple spectral initial resonance peaks in the multiple similar speech spectra are identified and compared in terms of multiple initial resonance peak positions, multiple initial resonance peak bandwidths and multiple initial resonance peak amplitudes, and the resonance peak comparison results are recorded. Then, based on the resonance peak comparison results, multiple speech with similar timbre are screened from the multiple similar verification speech.
[0096] It can be understood that among multiple voices with similar timbre, the initial resonance peak positions, initial resonance peak bandwidths and initial resonance peak amplitudes of multiple corresponding spectral initial resonance peaks meet the preset similarity standards. Therefore, it can be determined that the timbre of multiple voice output objects corresponding to multiple voices with similar timbre is similar. In a multi-person voice environment, it is difficult to distinguish the voices of different people.
[0097] Specifically, Figure 4 The flowchart of screening multiple voices with similar timbre in the method provided by the embodiment of the present invention is shown.
[0098] In a preferred embodiment of the present invention, performing spectrum conversion on a plurality of similar verification voices, comparing a plurality of initial formant positions, a plurality of initial formant bandwidths, and a plurality of initial formant amplitudes, and screening a plurality of voices with similar timbre specifically comprises the following steps:
[0099] Step S1031: Perform spectrum conversion on the multiple similar verification voices to obtain multiple similar voice spectra.
[0100] Step S1032: Mark a plurality of initial formants in the plurality of similar speech spectra according to a preset number of formants.
[0101] Step S1033 : comparing the multiple initial resonance peak positions, the multiple initial resonance peak bandwidths, and the multiple initial resonance peak amplitudes of the multiple initial resonance peaks of the frequency spectrum, and recording the resonance peak comparison result.
[0102] Step S1034: Filter multiple voices with similar timbre from the multiple similar verification voices based on the formant comparison result.
[0103] Furthermore, the personalized speech synthesis method based on artificial intelligence further includes the following steps:
[0104] Step S104: Filtering similar short-time energies corresponding to multiple voices with similar timbre from the multiple voice short-time energies, and selecting multiple target similar voices according to the multiple similar short-time energies.
[0105] In an embodiment of the present invention, by marking the speech short-time energy corresponding to the timbre-similar speech as similar short-time energy, it is possible to screen multiple similar short-time energies from multiple speech short-time energies, and then compare the multiple similar short-time energies, record the energy comparison results, and then, according to the energy comparison results, eliminate the similar short-time energies with the largest energy values from the multiple similar short-time energies, and mark the remaining multiple similar short-time energies as multiple target short-time energies, and then mark the timbre-similar speech corresponding to the target short-time energy as target similar speech, so that multiple target similar speech can be selected from multiple timbre-similar speech.
[0106] Specifically, Figure 5 A flowchart of selecting multiple target similar voices in the method provided by an embodiment of the present invention is shown.
[0107] Among them, in the preferred embodiment provided by the present invention, the screening of similar short-time energies corresponding to multiple voices with similar timbre from the multiple voice short-time energies, and the selection of multiple target similar voices according to the multiple similar short-time energies specifically include the following steps:
[0108] Step S1041 : Filtering similar short-time energies corresponding to multiple voices with similar timbre from the multiple voice short-time energies.
[0109] Step S1042: Compare the multiple similar short-time energies and record the energy comparison results.
[0110] Step S1043: Select multiple target short-time energies from the multiple similar short-time energies according to the energy comparison result.
[0111] Step S1044: Select a plurality of target similar voices corresponding to a plurality of the target short-time energies from the plurality of voices with similar timbre.
[0112] Furthermore, the personalized speech synthesis method based on artificial intelligence further includes the following steps:
[0113] Step S105: Select multiple target output objects corresponding to target similar voices from the multiple voice output objects, and perform different voice synthesis processing on the multiple target output objects based on artificial intelligence technology.
[0114] In an embodiment of the present invention, the speech output object corresponding to the target similar speech is marked as the target output object, so that multiple target output objects can be selected from multiple speech output objects, and then multiple different target timbre models can be selected from a preset timbre model database. Based on artificial intelligence technology, according to the multiple target timbre models, when multiple target output objects have speech output, different speech synthesis processing is performed on the multiple target output objects.
[0115] Specifically, Figure 6 The flowchart of the object speech synthesis processing in the method provided by the embodiment of the present invention is shown.
[0116] In a preferred embodiment of the present invention, selecting a plurality of target output objects corresponding to target similar voices from the plurality of voice output objects, and performing different voice synthesis processing on the plurality of target output objects based on artificial intelligence technology specifically includes the following steps:
[0117] Step S1051: Select multiple target output objects corresponding to multiple target similar voices from the multiple voice output objects.
[0118] Step S1052: Select multiple target timbre models from a preset timbre model database.
[0119] Step S1053: Based on artificial intelligence technology, different speech synthesis processes are performed on the multiple target output objects according to the multiple target timbre models.
[0120] Further, Figure 7 The application architecture diagram of the personalized speech synthesis terminal based on artificial intelligence provided by an embodiment of the present invention is shown.
[0121] In another preferred embodiment of the present invention, a personalized speech synthesis terminal based on artificial intelligence includes:
[0122] The verification voice collection module 101 is used to determine multiple voice output objects, display standard verification information to the multiple voice output objects, and collect standard verification voices of the multiple voice output objects in response to the standard verification information.
[0123] In an embodiment of the present invention, the verification voice collection module 101 obtains multiple voice output addresses in a multi-person voice environment, determines multiple corresponding voice output objects according to the multiple voice output addresses, generates standard verification information, sends the standard verification information to the multiple voice output objects, and provides voice response prompts to the multiple voice output objects. The multiple voice output objects can read the standard verification information according to the voice response prompts, and in the process of the multiple voice output objects reading the standard verification information, voice collection is performed on the multiple voice output objects to obtain multiple standard verification voices.
[0124] The short-time energy comparison module 102 is used to analyze the plurality of standard verification voices, calculate the corresponding short-time energies of the voices, compare the short-time energies of the plurality of voices, and select a plurality of similar verification voices.
[0125] In an embodiment of the present invention, the short-time energy comparison module 102 intercepts n voice signal samples from multiple standard verification voices, calculates multiple corresponding voice short-time energies according to the n voice signal samples corresponding to the multiple standard verification voices, and then compares the multiple voice short-time energies with each other, calculates the short-time energy difference between the multiple voice short-time energies, and then compares the multiple short-time energy differences with the standard energy difference with a preset standard energy difference as a reference, and marks the short-time energy difference that is smaller than the standard energy difference as a similar energy difference, so that multiple similar energy differences can be screened from the multiple short-time energy differences. Then, based on the multiple similar energy differences, multiple corresponding similar verification voices are screened from the multiple standard verification voices. Specifically, the calculation formula for calculating the multiple corresponding voice short-time energies is:
[0126]
[0127] Among them, i represents i standard verification speech, E i is the short-time energy of the i-th standard verification speech, n represents n speech signal samples, x i (n) is the n speech signal samples of the i-th standard verification speech, and L is the window length.
[0128] Specifically, Figure 8 FIG. 1 shows an application architecture diagram of the short-time energy comparison module 102 in the terminal provided by an embodiment of the present invention.
[0129] In a preferred embodiment of the present invention, the short-time energy comparison module 102 specifically includes:
[0130] The signal interception unit 1021 is configured to intercept a plurality of voice signal samples from each of the plurality of standard verification voices.
[0131] The energy calculation unit 1022 is used to verify multiple speech signal samples corresponding to the speech according to multiple standards, and calculate multiple corresponding speech short-time energies.
[0132] The energy difference calculation unit 1023 is configured to calculate the short-time energy difference between the multiple short-time energies of the speech.
[0133] The energy difference comparison unit 1024 is configured to compare the multiple short-time energy differences with a preset standard energy difference, and screen multiple similar energy differences.
[0134] The first speech screening unit 1025 is configured to screen a plurality of similar verification speech sounds from the plurality of standard verification speech sounds according to the plurality of similar energy differences.
[0135] Furthermore, the personalized speech synthesis terminal based on artificial intelligence also includes:
[0136] The spectrum conversion analysis module 103 is used to perform spectrum conversion on the multiple similar verification voices, compare multiple initial formant positions, multiple initial formant bandwidths and multiple initial formant amplitudes, and screen multiple voices with similar timbre.
[0137] In an embodiment of the present invention, the spectrum conversion analysis module 103 performs spectrum conversion on multiple similar verification speech to obtain multiple similar speech spectra, and then identifies and marks multiple spectrum initial resonance peaks in the multiple similar speech spectra according to a preset number of resonance peaks. The multiple spectrum initial resonance peaks in the multiple similar speech spectra are identified and compared in terms of multiple initial resonance peak positions, multiple initial resonance peak bandwidths and multiple initial resonance peak amplitudes, and the resonance peak comparison results are recorded. Then, based on the resonance peak comparison results, multiple speech with similar timbre are screened from the multiple similar verification speech.
[0138] Specifically, Figure 9 FIG. 1 shows an application architecture diagram of the spectrum conversion analysis module 103 in the terminal provided by an embodiment of the present invention.
[0139] In a preferred embodiment of the present invention, the spectrum conversion analysis module 103 specifically includes:
[0140] The spectrum conversion unit 1031 is configured to perform spectrum conversion on the plurality of similar verification voices to obtain a plurality of similar voice spectrums.
[0141] The formant marking unit 1032 is configured to mark a plurality of initial formants in the plurality of similar speech spectra according to a preset number of formants.
[0142] The formant comparison unit 1033 is configured to compare the multiple initial formant positions, the multiple initial formant bandwidths, and the multiple initial formant amplitudes of the multiple initial formant peaks of the frequency spectrum, and record the formant comparison result.
[0143] The second speech screening unit 1034 is configured to screen a plurality of speech sounds with similar timbre from the plurality of similar verification speech sounds according to the formant comparison result.
[0144] Furthermore, the personalized speech synthesis terminal based on artificial intelligence also includes:
[0145] The target speech selection module 104 is configured to screen similar short-time energies corresponding to multiple speech with similar timbre from the multiple speech short-time energies, and select multiple target similar speech according to the multiple similar short-time energies.
[0146] In an embodiment of the present invention, the target speech selection module 104 marks the speech short-time energy corresponding to the timbre-similar speech as similar short-time energy, thereby being able to screen multiple similar short-time energies from multiple speech short-time energies, and then compare the multiple similar short-time energies, record the energy comparison results, and then, according to the energy comparison results, eliminate the similar short-time energies with the largest energy values from the multiple similar short-time energies, mark the remaining multiple similar short-time energies as multiple target short-time energies, and then mark the timbre-similar speech corresponding to the target short-time energy as target similar speech, so that multiple target similar speech can be selected from multiple timbre-similar speech.
[0147] The speech synthesis processing module 105 is used to select target output objects corresponding to multiple target similar speech from the multiple speech output objects, and perform different speech synthesis processing on the multiple target output objects based on artificial intelligence technology.
[0148] In an embodiment of the present invention, the speech synthesis processing module 105 marks the speech output object corresponding to the target similar speech as the target output object, so that it can select multiple target output objects from multiple speech output objects, and then select multiple different target timbre models from a preset timbre model database. Based on artificial intelligence technology, according to the multiple target timbre models, when multiple target output objects have speech output, different speech synthesis processing is performed on the multiple target output objects.
[0149] It should be understood that, although the various steps in the flow chart of each embodiment of the present invention are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence according to the order indicated by the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in order, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.
[0150] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0151] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A personalized speech synthesis method based on artificial intelligence, characterized in that: The method specifically comprises the following steps: Determining a plurality of voice output objects, presenting standard verification information to the plurality of voice output objects, and collecting standard verification voices of the plurality of voice output objects in response to the standard verification information; Analyzing a plurality of the standard verification voices, calculating a plurality of corresponding voice short-time energies, comparing the plurality of the voice short-time energies, and screening a plurality of similar verification voices; Performing spectrum conversion on the multiple similar verification voices, comparing multiple initial formant positions, multiple initial formant bandwidths, and multiple initial formant amplitudes, and screening multiple voices with similar timbre; screening similar short-time energies corresponding to multiple voices with similar timbre from the multiple voice short-time energies, and selecting multiple target similar voices according to the multiple similar short-time energies; Selecting target output objects corresponding to multiple target similar voices from the multiple voice output objects, and performing different voice synthesis processing on the multiple target output objects based on artificial intelligence technology; The performing spectrum conversion on the multiple similar verification voices, comparing the multiple initial formant positions, the multiple initial formant bandwidths and the multiple initial formant amplitudes, and screening the multiple voices with similar timbre specifically comprises the following steps: Performing spectrum conversion on the plurality of similar verification voices to obtain a plurality of similar voice spectra; Marking a plurality of initial formants of the spectrum in the plurality of similar speech spectra according to a preset number of formants; comparing a plurality of initial resonance peak positions, a plurality of initial resonance peak bandwidths, and a plurality of initial resonance peak amplitudes of the plurality of initial resonance peaks of the spectrum, and recording a resonance peak comparison result; Based on the formant comparison result, screening multiple voices with similar timbre from the multiple similar verification voices; The step of screening similar short-time energies corresponding to multiple voices with similar timbre from the multiple voice short-time energies, and selecting multiple target similar voices according to the multiple similar short-time energies specifically comprises the following steps: Screening similar short-time energies corresponding to multiple voices with similar timbre from the multiple voice short-time energies; Comparing the multiple similar short-time energies and recording the energy comparison results; selecting a plurality of target short-time energies from the plurality of similar short-time energies according to the energy comparison result; Selecting a plurality of target similar voices corresponding to a plurality of target short-time energies from the plurality of voices with similar timbre; The step of selecting target output objects corresponding to target similar voices from the plurality of voice output objects and performing different voice synthesis processing on the plurality of target output objects based on artificial intelligence technology specifically comprises the following steps: Selecting a plurality of target output objects corresponding to a plurality of target similar voices from the plurality of voice output objects; Select multiple target timbre models from the preset timbre model database; Based on artificial intelligence technology, different speech synthesis processes are performed on the multiple target output objects according to the multiple target timbre models.
2. The personalized speech synthesis method based on artificial intelligence according to claim 1, characterized in that: The steps of determining a plurality of voice output objects, presenting standard verification information to the plurality of voice output objects, and collecting standard verification voices of the plurality of voice output objects in response to the standard verification information specifically include the following steps: determining a plurality of speech output objects; generating and presenting standard verification information to a plurality of said voice output objects; Providing voice response prompts to the plurality of voice output objects; Collect a plurality of standard verification voices of the voice output objects in response to the standard verification information.
3. The personalized speech synthesis method based on artificial intelligence according to claim 1, characterized in that: The analyzing of the plurality of standard verification voices, calculating the corresponding short-time energies of the plurality of voices, comparing the short-time energies of the plurality of voices, and screening the plurality of similar verification voices specifically comprises the following steps: intercepting a plurality of speech signal samples from a plurality of the standard verification voices; Verify multiple speech signal samples corresponding to the speech according to multiple standards, and calculate multiple corresponding speech short-time energies; Calculating the short-time energy difference between the plurality of speech short-time energies; Comparing the multiple short-time energy differences with a preset standard energy difference to screen multiple similar energy differences; Based on the multiple similar energy differences, multiple similar verification voices are screened from the multiple standard verification voices.
4. The personalized speech synthesis method based on artificial intelligence according to claim 3, characterized in that: The calculation formula of multiple speech short-time energies is: ; in, represent A standard verification voice, For the The standard verifies the short-term energy of speech. represent Speech signal samples, For the Standard verification of speech Speech signal samples, To add window length.
5. A personalized speech synthesis terminal based on artificial intelligence, characterized in that: The terminal includes a verification voice acquisition module, a short-time energy comparison module, a spectrum conversion analysis module, a target voice selection module and a voice synthesis processing module, wherein: a verification voice collection module, configured to determine a plurality of voice output objects, present standard verification information to the plurality of voice output objects, and collect standard verification voices of the plurality of voice output objects in response to the standard verification information; A short-time energy comparison module is used to analyze the plurality of standard verification voices, calculate the short-time energies of the plurality of corresponding voices, compare the short-time energies of the plurality of voices, and screen a plurality of similar verification voices; a spectrum conversion analysis module, configured to perform spectrum conversion on the plurality of similar verification voices, compare a plurality of initial formant positions, a plurality of initial formant bandwidths, and a plurality of initial formant amplitudes, and screen a plurality of voices with similar timbre; a target speech selection module, configured to screen similar short-time energies corresponding to multiple speech with similar timbre from the multiple speech short-time energies, and select multiple target similar speech according to the multiple similar short-time energies; a speech synthesis processing module, configured to select target output objects corresponding to multiple target similar speech objects from the multiple speech output objects, and perform different speech synthesis processing on the multiple target output objects based on artificial intelligence technology; The spectrum conversion analysis module specifically includes: a spectrum conversion unit, configured to perform spectrum conversion on the plurality of similar verification voices to obtain a plurality of similar voice spectra; A formant marking unit, configured to mark a plurality of initial formants of the spectrum in the plurality of similar speech spectra according to a preset number of formants; a resonance peak comparison unit, configured to compare a plurality of initial resonance peak positions, a plurality of initial resonance peak bandwidths, and a plurality of initial resonance peak amplitudes of the plurality of initial resonance peaks of the spectrum, and record the resonance peak comparison result; A second speech screening unit is configured to screen a plurality of speech with similar timbre from the plurality of similar verification speech according to the formant comparison result; The method of screening similar short-time energies corresponding to multiple voices with similar timbre from the multiple voice short-time energies and selecting multiple target similar voices according to the multiple similar short-time energies is specifically as follows: Screening similar short-time energies corresponding to multiple voices with similar timbre from the multiple voice short-time energies; Comparing the multiple similar short-time energies and recording the energy comparison results; selecting a plurality of target short-time energies from the plurality of similar short-time energies according to the energy comparison result; Selecting a plurality of target similar voices corresponding to a plurality of target short-time energies from the plurality of voices with similar timbre; The step of selecting target output objects corresponding to target similar voices from the plurality of voice output objects and performing different voice synthesis processing on the plurality of target output objects based on artificial intelligence technology is specifically as follows: Selecting a plurality of target output objects corresponding to a plurality of target similar voices from the plurality of voice output objects; Select multiple target timbre models from the preset timbre model database; Based on artificial intelligence technology, different speech synthesis processes are performed on the multiple target output objects according to the multiple target timbre models.
6. The personalized speech synthesis terminal based on artificial intelligence according to claim 5, characterized in that: The short-time energy comparison module specifically includes: A signal interception unit, configured to intercept a plurality of voice signal samples from each of the plurality of standard verification voices; An energy calculation unit, configured to verify a plurality of speech signal samples corresponding to the speech according to a plurality of the standards, and calculate a plurality of corresponding speech short-time energies; An energy difference calculation unit, configured to calculate a short-time energy difference between a plurality of the short-time energies of the speech; an energy difference comparison unit, configured to compare the plurality of short-time energy differences with a preset standard energy difference, and screen a plurality of similar energy differences; The first speech screening unit is used to screen a plurality of similar verification speech from a plurality of the standard verification speech according to a plurality of the similar energy differences.
Citation Information
Patent Citations
Speech recognition and authentication method and system
CN110473552A
Speaker recognition device, speaker recognition method, and recording medium
CN111009248A