An audio generation method for an audio book based on speech synthesis

CN122761802APending Publication Date: 2026-09-15HEYI SIYUAN (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611084894.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-09-15

Smart Images

  • Figure CN122761802A_ABST
    Figure CN122761802A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of audio generation, and particularly relates to an audio generation method for audio novels based on speech synthesis, which comprises the following steps: obtaining novel text and analyzing the novel text to form text unit data, determining a speaking role through narrative relationship analysis, configuring a stable voice identity and a dynamic expression state, and generating a candidate audio in combination with front and rear boundary constraints; then performing text, role expression and boundary quality evaluation on the candidate audio, screening and locally correcting abnormalities, and finally completing audio connection according to semantic pauses and low-energy positions. The method can reduce the misrecognition of roles caused by the omission of the speaking subject, maintain the continuity of the tone and emotion of the same role, reduce the sound mutation, discontinuity and popping of adjacent audios, and reduce the processing amount caused by the overall regeneration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio generation, and more particularly to a method, electronic device, and storage medium for generating audiobooks based on speech synthesis. Background Technology

[0002] In scenarios where long, multi-character novels are mass-produced for audio on online literature platforms, audio content production platforms, or digital publishing platforms, the novel texts typically include narration, main characters, secondary characters, and continuous dialogue, and may contain character aliases, omitted subjects, cross-paragraph references, and scene transitions. The platform needs to convert the novel texts into audio novels with clearly distinguished characters, natural emotional expression, and seamless chapter transitions, and to make localized corrections to any errors in character portrayal, pronunciation, or emotional expression.

[0003] Currently, existing technologies typically segment novel text by paragraphs or punctuation marks, then identify narration and dialogue through rule matching or language models, assign corresponding voice timbres to the identified characters, and generate audio segment by segment using a speech synthesis model. Finally, the audio segments are concatenated according to the text order. When generation errors are found, the production team usually modifies the corresponding text, character tags, or synthesis parameters, regenerates the erroneous segment, and replaces the original audio.

[0004] It is evident that existing technologies primarily determine the voice actor based on the character name, quotation mark position, or local dialogue content in the current paragraph. This makes it difficult to accurately identify the character corresponding to omitted subjects or cross-paragraph references by combining contextual references, dialogue turns, and character relationships, which can easily lead to mismatched character timbre. Furthermore, each audio segment is usually generated independently, and when regenerating locally, there is a lack of constraints on the timbre, speech rate, emotion, and pauses of adjacent audio segments. This can easily cause sudden loudness changes, intonation breaks, or abnormal pauses at replacement positions, affecting the continuity of audio in long audio novels. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies, such as the difficulty in accurately identifying the voice actors in long multi-character novels and the difficulty in ensuring the continuous connection between locally regenerated audio and adjacent audio. The invention proposes a method, electronic device, and storage medium for generating audio for audiobooks based on speech synthesis.

[0006] To at least solve some of the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: Firstly, this invention provides a method for generating audiobooks based on speech synthesis, including: The novel text is retrieved and structured parsing is performed to obtain text unit data; Narrative relationship analysis is performed on the text unit data to obtain voice character data; Configure the character's voice based on the text unit data and the voice character data, and generate character voice control data; Perform front and back boundary constraint synthesis on the character voice control data to obtain candidate audio data; Perform a quality evaluation on the candidate audio data to obtain audio quality data; Based on the audio quality data, candidate audio data is filtered and anomalies are locally corrected to generate target audio data; The text unit data and the target audio data are subjected to audio splicing processing to generate audiobook audio.

[0007] Preferably, narrative relationship analysis is performed on the text unit data to obtain voice actor data, including: Role relationship extraction is performed on the text unit data to generate role relationship data; Perform candidate role evaluation on the role relationship data to generate corresponding role data; Based on the corresponding role data, role confirmation is performed, and voice-speaking role data is generated.

[0008] Preferably, based on the corresponding role data, role confirmation is performed to generate voice role data, including: Obtain manually annotated verification text, perform threshold calibration on the data corresponding to the role, and obtain the role confirmation threshold and the role differentiation threshold; The corresponding data for each role is compared with the role confirmation threshold and the role differentiation threshold to generate voice-generating role data.

[0009] Preferably, character voices are configured based on the text unit data and voice character data, and character voice control data is generated, including: Based on the voice actor data, configure a stable voice identity and generate character voice identity data; The expression state is updated on the text unit data to generate character expression state data; The text unit data is processed for pronunciation and pause adjustment to generate text control data; The character voice identity data, character expression status data, and text control data are collected to generate character voice control data.

[0010] Preferably, the character voice control data is subjected to front and rear boundary constraint synthesis to obtain candidate audio data, including: The current text unit to be synthesized is determined from the text unit data, and the preceding boundary is determined to generate the preceding boundary data. Extract the voice control content of subsequent text units from the character voice control data to generate subsequent boundary data; The preceding boundary data and subsequent boundary data are aggregated to generate boundary constraint data; Speech synthesis is performed based on the character voice control data and boundary constraint data to generate candidate audio data.

[0011] Preferably, the preceding boundary data and subsequent boundary data are aggregated to generate boundary constraint data, including: Perform a correspondence analysis on the voice roles of the current text unit to be synthesized and the adjacent text units to generate voice relationship results; Based on the vocal relationship results, boundary constraints are processed on the preceding boundary data, subsequent boundary data, character voice control data, and text control data. Among them, continuous constraints are processed for relationships of the same character, and switching constraints are processed for relationships of different characters, thus generating boundary constraint data.

[0012] Preferably, a quality evaluation is performed on the candidate audio data to obtain audio quality data, including: The candidate audio data is evaluated for text correspondence, and text correspondence data is generated. The candidate audio data is evaluated for roles and expressions to generate voice matching data; Perform boundary continuity evaluation on the candidate audio data to generate boundary evaluation data; The text correspondence data, sound correspondence data, and boundary evaluation data are aggregated to generate audio quality data.

[0013] Preferably, the process of filtering candidate audio data based on the audio quality data and locally correcting anomalies to generate target audio data includes: Based on the audio quality data, the candidate audio data is filtered to generate audio filtering results; In response to the audio filtering results meeting the quality requirements, the target audio data is determined; In response to the audio screening results not meeting the quality conditions, abnormal correction data is generated; Based on the anomaly correction data, perform local re-synthesis and quality review to generate target audio data.

[0014] Preferably, based on the anomaly correction data, local re-synthesis and quality review are performed to generate target audio data, including: Extract the anomaly type, anomaly location, and pause information from the anomaly correction data; The abnormal text unit is determined based on the aforementioned anomaly correction data; Based on the anomaly type, anomaly location, and pause information, the local regeneration range is determined; The character voice control data corresponding to the local regeneration range is updated based on the anomaly correction data, and the updated character voice control data is synthesized by front and back boundary constraints to generate local corrected audio. The quality of the locally corrected audio is checked to generate target audio data.

[0015] Preferably, audio splicing processing is performed on the text unit data and the target audio data to generate audiobook audio, including: Semantic pause recognition is performed on the text unit data to generate pause range data; Low-energy location extraction is performed on the target audio data to generate low-energy location data; Perform position mapping on the pause range data and low energy position data to generate connection position data; The target audio data is connected based on the connection location data to generate audiobook audio.

[0016] In a second aspect, the present invention provides an electronic device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method in the first aspect.

[0017] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method in the first aspect.

[0018] Compared with the prior art, the present invention has the following beneficial effects: I. This invention extracts vocal actions, dialogue turns, titles, scene presence, and preceding vocal relationships, weights each relationship value to form a role correspondence value, and combines a role confirmation threshold and a role differentiation threshold to confirm the vocal role. This helps reduce role misidentification based solely on nearby names and avoids directly confirming the wrong role when there is insufficient vocal evidence or when candidate roles are closely matched.

[0019] Second, by using the same voiceprint vector for the same character in each text unit, and adjusting the emotion, speech rate, pitch, energy and pauses based on the character's expression state data, this invention helps the character maintain the original timbre and basic vocal habits when expressing different emotions, while making emotional changes more natural and maintaining the continuity of expression across chapters.

[0020] Third, this invention extracts the actual end state of the preceding target audio and the target vocal state of the subsequent text unit, and performs continuous constraints and switching constraints according to the same role relationship and different role relationships respectively. Then, it combines the semantic pause range and low energy position to determine the connection position, which helps to reduce the sound abrupt change when the same role vocalizes continuously, and avoids the occurrence of too short pause, loudness abrupt change, discontinuity, popping sound and rhythm abrupt change when the role switches positions.

[0021] Fourth, this invention performs text correspondence evaluation, role and expression evaluation and boundary continuity evaluation on candidate audio, determines the local regeneration range according to the anomaly type and performs local resynthesis and quality review, which is beneficial to correct pronunciation errors, while preserving the original character timbre and emotional expression, and avoids regenerating the entire chapter, which increases the processing load and changes the already qualified audio. Attached Figure Description

[0022] Figure 1 This is a flowchart of the method of the present invention.

[0023] Figure 2 This is a comparison diagram of the effects of the present invention and the prior art. Detailed Implementation

[0024] To make the technical means, creative features, objectives, and effects of this invention easier to understand, the invention is further described below with reference to specific embodiments. However, the following embodiments are merely preferred embodiments of this invention and not all of them. Other embodiments obtained by those skilled in the art based on the embodiments described herein without creative effort are all within the protection scope of this invention.

[0025] A specific embodiment of the present invention discloses a method for generating audiobooks based on speech synthesis. Please refer to [link to relevant documentation]. Figure 1 ,include: S1. Obtain the novel text and perform structured parsing to obtain text unit data; S2. Perform narrative relationship analysis on the text unit data to obtain voice character data; S3. Configure the character voice based on the text unit data and the voice character data, and generate character voice control data; S4. Perform front and rear boundary constraint synthesis on the character voice control data to obtain candidate audio data; S5. Perform a quality evaluation on the candidate audio data to obtain audio quality data; S6. Based on the audio quality data, filter candidate audio data and locally correct anomalies to generate target audio data; S7. Perform audio splicing processing on the text unit data and the target audio data to generate audiobook audio.

[0026] It should be noted that this embodiment is applied to the batch audio conversion of long, multi-character suspense novels on an audio content production platform. The example text for this implementation is the following three sentences: "Did you go to the dock last night?" Lin Zheng asked, staring at Zhou Yuan. "No, I've been at home the whole time." "The person in the surveillance footage looks a lot like you."

[0027] The novel also includes the phrase "Zhou Yuan felt that there was danger all around him," and there is a chase scene that continues across chapters with the speaker panting. This can be used to explain the process of omitting the speaker, mispronouncing polyphonic characters, and maintaining a state of mind.

[0028] As an optional implementation method, step S1 includes the following specific content: The processing terminal retrieves the novel text from the platform's content library, first standardizing the text format and identifying chapter titles, then dividing the text according to chapter order. Next, it identifies scene changes such as time changes, location changes, and character entry / exit to determine the scene scope. Then, it divides the text by paragraphs and sentences, and uses quotation marks, colon positions, verbs, and context to correct dialogue boundaries, forming continuously arranged text units. This avoids the separation of dialogue and action descriptions caused by segmenting solely by punctuation, making the text segmentation result more consistent with the actual narrative relationships of the novel.

[0029] Subsequently, for each text unit, the following information is recorded: text unit number, chapter identifier, scene identifier, text content, text type, preceding text unit number, subsequent text unit number, character name, character alias, action description, title information, emotion description, and pause information, resulting in text unit data. The pause information includes pause position, punctuation type, and basic pause duration. Using clear and uniform recording fields ensures the complete preservation of the sequential relationships between text units and the character's narrative information.

[0030] It should be noted that the text type is determined using existing text classification models, character names and aliases are extracted using existing named entity recognition models, the action subject, the object of action, and the title relationship are determined using dependency parsing, and the emotion description is identified using an emotion classification model. The processing terminal determines the pause position based on punctuation type and sentence boundaries, and queries a preset punctuation pause correspondence table to determine the basic pause duration. For example, commas correspond to shorter pauses, while periods, question marks, and exclamation marks correspond to longer pauses. In this way, the novel text is converted into text unit data containing sequential relationships, narrative relationships, and pause requirements.

[0031] For example, in the example text, there are 3 consecutive text units. The first text unit records Lin Zheng as the subject of the action and Zhou Yuan as the person being questioned. The second text unit records the dialogue text between the characters. The sequential relationship between the dialogue and the adjacent dialogue is established by the preceding text unit number and the subsequent text unit number, so as to preserve the continuous dialogue and its action prompt relationship and avoid losing the clues for character judgment during the text segmentation stage.

[0032] As an alternative implementation, the second sentence, "No, I've been at home," does not directly contain the character's name; it only uses nearby names for identification. This could easily lead to the sentence being incorrectly assigned to Lin Zheng. Step S2 can solve this problem, specifically: S2. Perform narrative relationship analysis on the text unit data to obtain voice character data; S2.1. Extract role relationships from the text unit data to generate role relationship data; The processing terminal reads text type, character name, character alias, action description, title information, preceding text unit number, subsequent text unit number, and scene identifier from text unit data within the same scene. It uses an existing relation extraction model to identify the action subject, the object being acted upon, the speaking action, title reference, and pronoun reference relationships. Subsequently, it determines the dialogue turn relationship based on the preceding and subsequent text unit numbers, the scene presence relationship based on the scene identifier, and the preceding voice relationship based on the confirmed preceding voice actor.

[0033] The degree of association between the action subject, the object being acted upon, the speaking action and the candidate role is then converted into a vocal action relationship value, which forms a dialogue round relationship value, a title-pointing relationship value, a scene presence relationship value and a preceding vocal relationship value, generating role relationship data.

[0034] The relationship value identified by the model is the predicted probability output by the model. It is recorded as 1 when the rule is clearly true and 0 when it is not true. Using a unified relationship value can clearly represent the degree of association between different narrative clues and candidate characters, and reduce character misidentification based solely on nearby names.

[0035] S2.2, Perform candidate role evaluation on the role relationship data to generate corresponding role data; For the current dialogue text, roles within the same scene are selected as candidate roles from the role relationship data. For example, if the current scene is an interrogation room, and the text has already determined that Lin Zheng and Zhou Yuan are present, while Zhang Ming is located in another office, the processing terminal will identify Lin Zheng and Zhou Yuan as candidate roles for the dialogue "No, I've been at home the whole time," without specifying a speaker, and will not include Zhang Ming in the candidate pool.

[0036] Subsequently, the voice action relationship value, dialogue turn relationship value, title pointing relationship value, scene presence relationship value and previous voice relationship value corresponding to each candidate character are read in sequence. Each relationship value is multiplied by its corresponding weight and then summed to obtain the character correspondence value of the candidate character, which is used to represent the strength of the correspondence between a single candidate character and the current dialogue.

[0037] Then, each candidate character and its corresponding value are written into the same record to generate character-corresponding data, which is used to represent the evaluation results of all candidate characters. By comparing multiple candidate characters with unified values, it is possible to avoid unclear character judgments caused by the inability to directly compare different narrative threads.

[0038] This implementation uses pre-defined fixed weights. Before system deployment, the weighted combinations are traversed using verification text manually labeled with the actual voice actors. Among the combinations where the sum of all weights is 1, the group with the highest voice actor recognition accuracy is selected; if the recognition accuracy is the same, the group with the lower false confirmation rate is selected. For example, the weights of the selected voice action relationship, dialogue turn relationship, title orientation relationship, scene presence relationship, and preceding voice relationship are set to 0.25, 0.30, 0.20, 0.10, and 0.15, respectively.

[0039] S2.3. Based on the corresponding role data, perform role confirmation and generate voice role data; S2.3.1 Obtain manually annotated verification text, perform threshold calibration on the data corresponding to the role, and obtain the role confirmation threshold and the role differentiation threshold; The processing terminal obtains the verification text that has been marked with the actual voice actor, and calculates the role correspondence value of each candidate role in each verification text according to S2.1 and S2.2. Then, it determines the highest role correspondence value and the second highest role correspondence value, and subtracts the second highest role correspondence value from the highest role correspondence value to obtain the role distinction difference value. The highest role correspondence value represents the strength of the judgment basis of the most likely voice actor, and the role distinction difference value represents the degree of distinction of the role compared with other candidate roles.

[0040] Subsequently, the minimum and maximum values ​​of the highest-ranking character's corresponding value in the verification text are statistically analyzed, and candidate character confirmation thresholds are set at fixed intervals within this range. Then, the minimum and maximum values ​​of the character differentiation differences are statistically analyzed, and candidate character differentiation thresholds are set in the same manner. The fixed interval is determined based on the precision of the character's corresponding value; for example, if the character's corresponding value is retained to two decimal places, the interval can be set to 0.05. Each candidate character confirmation threshold is then paired sequentially with each candidate character differentiation threshold to form candidate threshold combinations. In each combination, the first value is used to determine the highest-ranking character's corresponding value, and the second value is used to determine the character differentiation difference, thus clearly comparing the confirmation effects of different threshold combinations.

[0041] For each candidate threshold combination, only verification texts that simultaneously reach two threshold values ​​are confirmed as speaking roles, and the confirmation results are compared with manually annotated results. Then, the correct confirmation rate and the incorrect confirmation rate are calculated separately. The upper limit of the incorrect confirmation rate is predetermined according to the proportion of role misidentification acceptable to the platform, for example, set to 5%, which can avoid introducing too many incorrect roles by increasing the number of automatic confirmations by relaxing the threshold.

[0042] Finally, candidate threshold combinations with false confirmation rates exceeding the upper limit are excluded, and the group with the highest correct confirmation rate is selected from the remaining combinations. The candidate role confirmation cutoff value in this group is determined as the role confirmation threshold, and the candidate role differentiation cutoff value in this group is determined as the role differentiation threshold. When correct confirmation rates are the same, the group with the larger number of automatically confirmed texts is selected, thereby reducing the amount of manual review while controlling role misidentification.

[0043] The character confirmation threshold sets the minimum value that the highest character's corresponding value should reach, while the character differentiation threshold sets the minimum value that the character differentiation difference should reach. The candidate character corresponding to the highest character's corresponding value is confirmed only when both the highest character's corresponding value and the character differentiation difference reach the character differentiation threshold. This combined assessment reduces false confirmations caused by weak vocal evidence or when candidate character corresponding values ​​are close.

[0044] S2.3.2. Compare the corresponding data of the roles with the role confirmation threshold and the role differentiation threshold respectively to generate voice role data; The processing terminal reads each candidate character and its corresponding value from the character correspondence data of the novel to be processed, determines the highest and second highest character correspondence values, and subtracts the second highest character correspondence value from the highest character correspondence value to obtain the character differentiation difference of the current dialogue text to be confirmed.

[0045] When the highest role correspondence value reaches the role confirmation threshold and the role differentiation difference reaches the role differentiation threshold, the role corresponding to the highest role correspondence value and the text unit number of the current dialogue text to be confirmed are written into the voice role data, thereby avoiding direct confirmation of the wrong role when the voice basis is insufficient or the candidate role correspondence is close.

[0046] For example, in the dialogue text "No, I've been at home" in the novel to be processed, Zhou Yuan's role correspondence value is 0.76, and Lin Zheng's role correspondence value is 0.42. Therefore, the highest role correspondence value is 0.76, the second highest is 0.42, and the role differentiation difference is 0.34. If the role confirmation threshold determined beforehand through manual annotation and text verification is 0.60, and the role differentiation threshold is 0.20, then both the highest role correspondence value and the role differentiation difference corresponding to Zhou Yuan meet the requirements. Therefore, the text unit number of Zhou Yuan and this dialogue text is written into the voice role data.

[0047] If any comparison is not satisfied, the preceding text units in the same scene are added sequentially along the preceding text unit number, and the action subject, title reference, pronoun reference, dialogue turn, scene presence and preceding voice relationship are extracted again to update the role relationship data; then the role correspondence value of the candidate role is recalculated according to S2.2 to obtain the updated role correspondence data; by expanding within the same scene, the judgment criteria directly related to the current dialogue can be increased, and the interference caused by roles in other scenes can be reduced.

[0048] Subsequently, two threshold comparisons are performed again on the updated role-corresponding data. If the conditions are met, the corresponding role is written into the voice-acting role data; if the conditions are still not met after expanding to the current scene start position, the processing terminal sends the current dialogue text to be confirmed and the candidate roles to the manual review interface, and receives the manual confirmation data returned by the manual review interface. The manual confirmation data includes the current text unit number and the voice-acting role confirmed by the human. The voice-acting role and text unit number in the manual confirmation data are written into the voice-acting role data, thereby ensuring that dialogues that cannot be automatically determined can still form clear voice-acting role results, avoiding input gaps in subsequent role voice configuration.

[0049] As an alternative implementation method, the same character will change their speech rate, pitch, and energy when under states of tension, anger, or fatigue. If fixed timbre parameters are directly changed, the same character may sound like different people. In view of this, the present invention proposes the following working method: S3. Configure the character voice based on the text unit data and the voice character data, and generate character voice control data; S3.1 Configure a stable voice identity based on the voice actor data, and generate character voice identity data; Before the novel is produced, the platform receives sample audio files uploaded by the production staff for each character and stores them in the character audio library according to the character identifier. When no sample audio file has been uploaded for a character, the processing terminal selects a target timbre from the preset timbres provided by the speech synthesis model, obtains the corresponding timbre encoding parameters, and uses the preset timbre to synthesize a preset standard text to generate a substitute character sample audio file.

[0050] The processing terminal retrieves the character sample audio uploaded by the production staff or the alternative character sample audio generated according to the preset timbre from the character audio library based on the character identifier in the voice character data. Then, the character sample audio is divided into frames, and existing speech activity detection methods are used to exclude silent frames and non-speech noise frames to obtain valid voice frames containing the character's voice. By excluding invalid audio content, the impact of silence and environmental noise on the character's voice characteristics can be reduced.

[0051] Then, the existing speaker feature extraction model is used to process the effective speech frames to obtain the character voiceprint vector, which is used to represent the relatively stable voice identity of the character. The same character uses the same voiceprint vector in each text unit, which can avoid the character being synthesized into different characters' timbre due to emotional changes.

[0052] For example, when a character speaks calmly and questions angrily, the processing terminal calls the same character's voiceprint vector to maintain their deep, resonant voice identity; the pitch, energy, and speed of the voice are increased only when the character is questioning angrily, so that although the character's tone changes, it still sounds like the same character and will not become a different timbre due to increased emotion.

[0053] Subsequently, the fundamental frequency of each effective vocal frame is calculated using the fundamental frequency detection method, and the basic pitch range is determined according to the fundamental frequency distribution to represent the pitch variation range when the character speaks normally. The basic speech rate is obtained by dividing the number of effective syllables in the character sample audio by the effective vocal duration. Then, the frame energy of each effective vocal frame is calculated, and the basic energy range is determined according to the frame energy distribution to represent the sound intensity range when the character speaks normally.

[0054] Finally, the character's voiceprint vector, basic pitch range, basic speech rate, basic energy range, and timbre encoding parameters corresponding to the preset timbre are collected to generate the character's voice identity data; when the preset timbre is not used, the timbre encoding parameters are empty.

[0055] The same character consistently uses the same voiceprint vector across all text units, with voice variations limited by a basic pitch range, basic speech rate, and basic energy range. Adjustments to emotions are controlled separately by the character's expressive state data, ensuring that the character maintains their original timbre and basic vocal habits when expressing different emotions.

[0056] S3.2, Update the expression state of the text unit data to generate character expression state data; The processing terminal inputs the text content, emotion description, action description, and punctuation type in the pause information of the current text unit into the existing emotion classification model to obtain the target emotion category and the corresponding target emotion intensity. The target emotion intensity is a value between 0 and 1, and the larger the value, the more obvious the corresponding emotion.

[0057] Subsequently, the most recently generated character expression state data for the same voice actor is read, and the recorded emotional intensity is used as the previous emotional intensity. If the current text unit is the first time the character has spoken, the initial intensity corresponding to the calm state is used as the previous emotional intensity. The retention ratio of the previous emotional intensity is then determined according to the state continuation coefficient, and the adoption ratio of the target emotional intensity is determined by subtracting the state continuation coefficient from 1. The two parts are then added together to obtain the current emotional intensity, so that the character's current emotion both continues the original state and reflects the emotional changes in the current text.

[0058] Among them, the state continuation coefficient represents the degree to which the previous emotion is retained in the current text unit.

[0059] For example, if the character's previous tension intensity was 0.30, the current target tension intensity is 0.80, and the state continuation coefficient is 0.40, then: The current tension level is 0.30 multiplied by 0.40, then 0.80 multiplied by 0.60, resulting in 0.60. This indicates that the character's tension level has increased from 0.30 to 0.60, instead of jumping directly to 0.80, which helps to make the emotional changes more natural.

[0060] The platform pre-acquires manually labeled emotional speech samples. These samples can be recorded by production staff or generated by a speech synthesis model and then manually verified. Each sample records the emotion category and intensity. The processing terminal calculates the speech rate, pitch, energy, and pauses of each sample and compares them with samples in a calm state to obtain the range of variation corresponding to different emotion categories and intensities, forming an emotion expression correspondence table.

[0061] Subsequently, the corresponding emotion expression table is queried according to the current emotion category and current emotion intensity to obtain the corresponding range of speech rate, pitch, energy and pause changes. This data is then written into the character expression status data along with the current emotion category and current emotion intensity, thereby converting the text emotion into specific speech expression requirements.

[0062] When switching chapters but the time and scene are continuous, the emotional intensity of the same character at the end of the previous chapter is used; when there is a clear time jump or calm description in the text, the emotional intensity is gradually reduced according to the above calculation, so as to maintain the continuity of expression across chapters and avoid changing the character's voice identity.

[0063] S3.3 Perform pronunciation and pause organization on the text unit data to generate text control data; The processing terminal uses a polyphonic character disambiguation model to analyze the contextual semantics of the current text content, and combines the pronunciations of personal names, place names, and novel-specific terms entered into the platform to generate pronunciation information.

[0064] Subsequently, the current text content and pause information are extracted from the text unit data, and the current text content, pronunciation information, and pause information are aggregated into text control data.

[0065] S3.4. Collect the character voice identity data, character expression status data, and text control data to generate character voice control data; The processing terminal uses the text unit number as an association identifier to write the voice actor data, voice identity data, expression status data, and text control data into the same control record, generating voice control data for the character. The voice identity data is used to define the stable timbre of the corresponding character, while the expression status data is used to define the emotion, speech rate, pitch, energy, and pauses of the current text unit.

[0066] For example, Lin Zheng uses the same voiceprint vector when calmly stating his position and when interrogating forcefully, while only adjusting his speech rate, pitch, and intensity when interrogating forcefully. This configuration allows the character to maintain a stable and distinguishable voice identity while expressing different emotions.

[0067] As an optional implementation, when each text unit is synthesized individually, the speech synthesis model only considers the character's voice and expression requirements for that sentence, without considering the actual end state of the previous audio segment or the target vocal state of the next audio segment. Therefore, although adjacent audio segments may satisfy the text and character requirements respectively, their connection positions may still be discontinuous. Step S4 can solve the above problem, including: S4. Perform front and rear boundary constraint synthesis on the character voice control data to obtain candidate audio data; S4.1 Determine the current text unit to be synthesized from the text unit data, and perform the preceding boundary determination to generate the preceding boundary data; The processing terminal processes each text unit sequentially according to its text unit number. After the preceding text unit completes steps S4 to S6 and generates the target audio data, the current text unit to be synthesized is processed. Based on the preceding text unit number, the processing terminal obtains the target audio corresponding to the preceding text unit from the target audio data. Subsequently, it extracts the actual speech rate and character voiceprint vector from the target audio, and extracts the fundamental frequency change trend, energy change trend, and end pause length from the audio frame at the end of the target audio to generate preceding boundary data.

[0068] The fundamental frequency and energy change trends were obtained by linearly fitting the fundamental frequency and energy of the last audio frame in chronological order, respectively. An increase in the fitting result indicates sound enhancement, while a decrease indicates sound weakening. Furthermore, the last valid speech frame was determined through general speech activity detection, and the duration from its end to the end of the audio frame was taken as the end pause length.

[0069] When the current text unit to be synthesized is the first text unit of a chapter, the processing terminal reads the character's voiceprint vector or timbre encoding parameters, basic pitch range, basic speech rate and basic energy range from the corresponding character's voice control data, and uses the platform's preset 1-second to 3-second chapter start pause range to generate the preceding boundary data corresponding to the chapter's starting position, thereby ensuring that the first sentence of the chapter also has clear generation conditions.

[0070] S4.2 Extract the voice control content of subsequent text units from the character voice control data to generate subsequent boundary data; The processing terminal reads the character voice control data corresponding to the subsequent text unit based on the number of the subsequent text unit to be synthesized, and extracts the voice character data, character voice identity data, character expression status data and text control data from it.

[0071] Subsequently, the basic pitch range and the pitch variation range are superimposed to obtain the subsequent target pitch range; the basic speech rate and the speech rate variation range are superimposed to obtain the subsequent target speech rate range; the basic energy range and the energy variation range are superimposed to obtain the subsequent target energy range; and the subsequent target pause range is determined based on the pause information and the pause variation range.

[0072] For example, if the basic pitch range is 100 to 120 Hz, and the pitch variation range is an increase of 10 to 20 Hz, then the subsequent target pitch range is 110 to 140 Hz; if the basic speech rate is 4 words per second, and the speech rate variation range is an increase of 0.5 to 1 word per second, then the subsequent target speech rate range is 4.5 to 5 words per second.

[0073] Finally, the subsequent voice actor, voiceprint vector, target pitch range, target speech rate range, target energy range, and target pause range are written into the same record to generate subsequent boundary data. This is beneficial for limiting the ending method of the current audio by utilizing its sound control requirements before the next audio segment is generated.

[0074] S4.3. Collect the preceding boundary data and subsequent boundary data to generate boundary constraint data; S4.3.1 Perform a correspondence analysis on the voice roles of the current text unit to be synthesized and the adjacent text units to generate voice relationship results; The processing terminal compares the vocal roles of the current text unit to be synthesized with those of the two adjacent text units. When the role identifiers are the same, they are recorded as the same role relationship; when the role identifiers are different, they are recorded as different role relationships. A vocal relationship result is generated, and the two vocal relationship results record the role relationships corresponding to the beginning and end of the current audio respectively, thereby avoiding the use of the same continuity rule when the roles of the preceding and following text units are different.

[0075] S4.3.2. Based on the vocal relationship results, perform boundary constraint processing on the preceding boundary data, subsequent boundary data, character voice control data, and text control data. Among them, continuous constraint processing is performed on the same character relationship, and switching constraint processing is performed on different character relationships to generate boundary constraint data. The processing terminal reads the character's voice identity data, character's expression status data, and text control data from the character's voice control data corresponding to the current text unit to be synthesized. Then, it superimposes the basic pitch range with the pitch variation range to obtain the current target pitch range; it superimposes the basic speech rate with the speech rate variation range to obtain the current target speech rate range; it superimposes the basic energy range with the energy variation range to obtain the current target energy range; and finally, based on the basic pause duration and pause variation range, it obtains the current target pause range.

[0076] When the current text unit to be synthesized and the preceding text unit are spoken by the same character, the processing terminal reads the actual speech rate, character voiceprint vector, fundamental frequency change trend, energy change trend and end pause length from the preceding boundary data, and takes the end state of the preceding target audio as the starting point, within the current target pitch range, current target speech rate range, current target energy range and current target pause range, determines the allowable variation range of the candidate audio beginning.

[0077] When the current text unit to be synthesized and the subsequent text units are spoken by the same character, the target pitch range, target speech rate range, target energy range and character voiceprint vector in the subsequent boundary data are read to determine the sound range that the end of the current candidate audio should reach.

[0078] Subsequently, the allowed range of variation at the beginning of the candidate audio and the range of sound that should be reached at the end are written into the boundary constraint data to ensure that the voiceprint of the same character remains consistent when the same character speaks continuously, and to limit the sudden changes in pitch, speech rate and energy.

[0079] When the current text unit to be synthesized and the preceding text unit are spoken by different characters, the processing terminal determines the character switching pause range and energy transition range of the candidate audio head based on the end pause length and energy change trend in the preceding boundary data, as well as the current target pause range and the current target energy range.

[0080] When the current text unit to be synthesized and the subsequent text units are spoken by different characters, the processing terminal determines the character switching pause range and energy transition range at the end of the candidate audio based on the current target pause range, the current target energy range, and the subsequent target pause range and subsequent target energy range in the subsequent boundary data.

[0081] Subsequently, the pause range for character switching and the energy transition range are written into the boundary constraint data, thereby preserving the timbre differences between different characters while avoiding excessively short pauses, sudden loudness changes, and unnatural dialogue rhythms during character switching.

[0082] S4.4. Perform speech synthesis based on the character voice control data and boundary constraint data to generate candidate audio data; The processing terminal inputs the character voice control data and boundary constraint data corresponding to the text unit to be synthesized into the speech synthesis model. When uploading character sample audio, the corresponding character voiceprint vector is fixed; when using preset timbre, the corresponding timbre encoding parameters are fixed, and the speech rate, pitch, energy, and pauses are adjusted within the allowed range to generate multiple audio versions.

[0083] Subsequently, a candidate audio identifier is assigned to each audio version, and the candidate audio identifier, the corresponding text unit number, and the audio version are written into the same record to generate candidate audio data, thereby preserving the basis for the generation of each candidate audio and facilitating accurate comparison of different versions.

[0084] As an optional implementation, the fact that the candidate audio can be played in its entirety does not necessarily mean that the text content, character voice, expression state, and continuity all meet the requirements. Therefore, this invention proposes the following working method: S5. Perform a quality evaluation on the candidate audio data to obtain audio quality data; S5.1 Perform text correspondence evaluation on the candidate audio data to generate text correspondence data; The processing terminal reads the current text content and pronunciation information in the corresponding text control data based on the text unit number in the candidate audio data; it processes the candidate audio using a general speech recognition model to obtain the recognized text and the start and end times of each recognized character in the candidate audio, and extracts the actual pronunciation of each character using a general pronunciation evaluation model.

[0085] Subsequently, the identified text is compared with the current text content to determine the omission, repetition, and misreading locations. The actual pronunciation is then compared with the target pronunciation in the pronunciation information to identify the abnormal pronunciation locations. After deduplication of each abnormal location, the corresponding number of erroneous characters is added together to obtain the number of erroneous characters. Then, the ratio of the number of erroneous characters to the number of characters in the current text is subtracted from 1 to obtain the text correspondence degree. If the calculation result is less than 0, it is counted as 0.

[0086] Finally, the degree of text correspondence, omissions, repetitions, misreadings, pronunciation abnormalities, and the corresponding start and end times of the word level are written into the same record to generate text correspondence data, which helps to clarify whether the candidate audio completely and accurately reads the current text.

[0087] S5.2. Perform role and expression evaluation on the candidate audio data to generate sound matching data; The processing terminal uses the speaker feature extraction model in S3.1 to process candidate audio, obtain candidate voiceprint vectors, and then calculates the cosine similarity between the candidate voiceprint vector and the corresponding role voiceprint vector to obtain the role consistency degree. The higher the role consistency degree, the closer the candidate audio is to the stable voice identity of the target role.

[0088] Subsequently, the processing terminal superimposes the basic pitch range, basic speech rate, and basic energy range from the character's voice identity data onto the pitch, speech rate, and energy change range from the character's expression state data to obtain the current target pitch range, current target speech rate range, and current target energy range; then, based on the basic pause duration and pause change range, it obtains the current target pause range.

[0089] Subsequently, the actual speech rate, average fundamental frequency, average short-time energy, and actual pause length of the candidate audio are calculated. When any actual value is within the corresponding target range, the corresponding deviation is recorded as 0. When the actual value exceeds the target range, the difference between the actual value and the nearest endpoint of the range is divided by the width of the corresponding target range to obtain the speech rate deviation, pitch deviation, energy deviation, or pause deviation.

[0090] Then, the average value of speech rate deviation, pitch deviation, energy deviation and pause deviation is subtracted from 1 to obtain the degree of expression conformity. If the calculation result is less than 0, it is counted as 0.

[0091] Finally, the consistency of the character, the degree of expression, the actual speech rate, the average fundamental frequency, the average short-term energy, the actual pause length, and various deviations are written into the same record to generate voice consistency data. This processing can identify the character's timbre deviation and emotional expression inconsistency.

[0092] S5.3 Perform boundary continuity evaluation on the candidate audio data to generate boundary evaluation data; The processing terminal reads the corresponding boundary constraint data based on the text unit number in the candidate audio data, and extracts the fundamental frequency, energy, speech rate, voiceprint and pause status of the beginning and end of the candidate audio respectively.

[0093] Under the same role relationship, the processing terminal compares the fundamental frequency, energy, and speech rate at the corresponding boundary of the candidate audio with the corresponding allowed range. When the actual value is within the allowed range, the corresponding normalized boundary difference is recorded as 0. When it exceeds the allowed range, the difference between the actual value and the nearest endpoint of the range is divided by the width of the corresponding allowed range to obtain the normalized boundary difference of the fundamental frequency, energy, and speech rate. Subsequently, the cosine similarity between the candidate voiceprint vector and the corresponding role voiceprint vector in the boundary constraint data is calculated, and the cosine similarity is subtracted from 1 to obtain the voiceprint normalized boundary difference.

[0094] Under different role relationships, the processing terminal compares the actual pause length and actual energy at the corresponding boundary of the candidate audio with the pause range and energy connection range of the role switching, and obtains the pause normalization boundary difference and energy connection normalization boundary difference in the above manner. Then, each normalization boundary difference is multiplied by the boundary weight and summed to obtain the boundary difference value; the boundary continuity is obtained by subtracting the boundary difference value limited to 0 to 1 from 1.

[0095] The boundary weights are pre-defined using manually labeled audio recordings of natural and anomalous transitions. For example, under the same role relationship, the weights for fundamental frequency, energy, speech rate, and voiceprint are set to 0.30, 0.30, 0.20, and 0.20, respectively; under different role relationships, the weights for pauses and energy are set to 0.60 and 0.40, respectively. This allows for a focus on evaluating the continuity of sound for the same role, while focusing on evaluating transitions, pauses, and loudness transitions for different roles.

[0096] Finally, the candidate audio identifier, beginning or end identifier, boundary continuity, and various normalized boundary differences are written into the same record to generate boundary evaluation data. This process can identify different types of transition anomalies according to the actual vocal relationships.

[0097] S5.4. Collect the text corresponding data, sound matching data and boundary evaluation data to generate audio quality data; The processing terminal uses candidate audio identifiers as the basis for association, reads the degree of text correspondence, the degree of role consistency, the degree of expression conformity, and the degree of boundary continuity, and then multiplies them by the corresponding quality weights and sums them to obtain the comprehensive quality value.

[0098] Quality weights are pre-calibrated using manually verified preferred candidate audio. Among the candidate combinations where the sum of all weights is 1, the group with the highest hit rate of preferred audio is selected; when verification audio is missing, equal weights are used, and recalibration is performed after accumulating manually verified audio.

[0099] Subsequently, the four degree values, various deviations, and the overall quality value are written into the same record to generate audio quality data, thereby completely preserving the various quality results of the candidate audio and avoiding judging the overall quality based on only a single indicator.

[0100] As an alternative implementation, the phrase "fraught with danger" in the novel might be synthesized into a stressed syllable indicating weight. Regenerating the entire chapter would increase processing complexity and alter already valid audio. Therefore, this invention proposes the following working method: S6. Based on the audio quality data, filter candidate audio data and locally correct anomalies to generate target audio data; S6.1. Based on the audio quality data, perform filtering on the candidate audio data to generate audio filtering results; The platform obtains manually labeled qualified audio and abnormal audio in advance, and separately calculates the text correspondence degree, role consistency degree, expression compliance degree and boundary continuity degree.

[0101] Subsequently, candidate threshold values are set at an interval of 0.01 and verified one by one. For example, it is finally determined that the text correspondence degree is not less than 0.95, the role consistency degree is not less than 0.85, the expression compliance degree is not less than 0.80, and the boundary continuity degree is not less than 0.82. The above minimum requirements together constitute the quality condition.

[0102] The processing terminal compares the audio quality data corresponding to each candidate audio with the quality condition item by item. When all four degrees meet the corresponding minimum requirements, the audio is recorded as qualified; when any degree fails to meet the corresponding minimum requirement, the audio is recorded as abnormal, and the evaluation items that fail to meet the requirements are recorded.

[0103] Finally, the candidate audio identifier, qualification status, unqualified evaluation items and comprehensive quality value are written into the same record to generate an audio screening result. Through item-by-item screening, audio with high comprehensive quality but still obvious misreading, abnormal timbre and abnormal connection is prevented from being directly adopted.

[0104] S6.2. In response to said audio screening result satisfying the quality condition, determine target audio data; When there is a qualified candidate audio in the audio screening result, the processing terminal selects the candidate audio with the highest comprehensive quality value, writes the candidate audio, the corresponding text unit number and the actual audio duration into the same record, and generates target audio data, so as to select the audio with better overall performance after eliminating individual quality defects, and enable subsequent text units to accurately obtain the corresponding preceding target audio.

[0105] S6.3. In response to said audio screening result not satisfying the quality condition, generate abnormality correction data; The processing terminal determines the abnormal position and corresponding deviation in the text correspondence data, sound compliance data or boundary evaluation data respectively based on the evaluation items that fail to meet the requirements, with the candidate audio identifier as the association basis, and determines the abnormal type and the content to be corrected.

[0106] When the text correspondence degree fails to meet the requirement, the missing reading position, repeated reading position, misreading position or abnormal pronunciation position are determined according to the text correspondence data. In case of missing reading, repeated reading or misreading, the corresponding original text segment and abnormal position are written into the forced regeneration content; in case of abnormal pronunciation, the pronunciation information in the text control data is updated. For example, when the pronunciation of "wei xian chong chong" is wrong, the manually confirmed correct pronunciation is written into the corresponding text unit.

[0107] When the role consistency degree fails to meet the requirement, the role voiceprint vector in the role voice identity data is read and used as the fixed voice identity for local re-synthesis.

[0108] When the expression does not meet the requirements, adjust the corresponding range of change according to the direction of deviation in speech rate, pitch, energy, or pauses; if the actual change is higher than the upper limit of the range, lower the range; if it is lower than the lower limit of the range, raise the range.

[0109] When the continuity of the boundary does not meet the requirements, adjust the corresponding boundary constraints according to the location of the boundary difference; when there is a sudden change in the head, tighten the allowable range of change in the head; when the tail does not reach the state of the subsequent sound, adjust the range of sound that the tail should reach.

[0110] Subsequently, the abnormal text unit number, abnormal type, abnormal location, forced regeneration content, and updated content are written into the same record to generate abnormal correction data.

[0111] S6.4. Based on the anomaly correction data, perform local re-synthesis and quality verification to generate target audio data; S6.4.1 Determine the abnormal text unit based on the aforementioned abnormal correction data; The processing terminal reads the abnormal text unit number, abnormal type, abnormal location, forced regeneration content and updated content from the abnormal correction data, and obtains the corresponding text content, the preceding text unit number and the following text unit number from the text unit data to determine the abnormal text unit.

[0112] S6.4.2. Based on the anomaly type, anomaly location, and pause information, determine the local regeneration range; When the anomaly type is a missed read, a duplicate, a misread, or a pronunciation error, the processing terminal reads the pause position in the text control data corresponding to the abnormal text unit and determines the text between the two nearest pause positions before and after the abnormal character position as the local regeneration range; if no pause position is identified before or after, the local regeneration range is determined by expanding the preset number of characters forward and backward from the abnormal character position as the center.

[0113] When the exception type is a role consistency exception or expression exception, the entire text of the exception text unit is determined as the local regeneration range; when the exception type is a boundary exception, the text between the beginning or end of the exception and the nearest pause position in the corresponding direction is determined as the local regeneration range.

[0114] The processing terminal maps the local regeneration range to the local audio interval in the original candidate audio based on the start and end times of the character level in the corresponding text data, and retains the original candidate audio 300 milliseconds before and after the interval as boundary overlapping audio.

[0115] S6.4.3. Update the character voice control data corresponding to the local regeneration range based on the anomaly correction data, and perform front and rear boundary constraint synthesis on the updated character voice control data to generate local correction audio. The processing terminal updates the pronunciation information, character expression status data, or boundary constraint data corresponding to the local regeneration range according to the updated content in the anomaly correction data.

[0116] If the character consistency does not meet the requirements, the corresponding character's voiceprint vector is reread and used as the fixed voice identity.

[0117] Subsequently, the text content corresponding to the local regeneration range, the updated character voice control data, and the boundary constraint data are input into the speech synthesis model to generate locally corrected audio; the boundary overlapping audio is not input into the speech synthesis model, but is only used to replace and connect the locally corrected audio with the original candidate audio.

[0118] For example, if the only error in the local audio is the pronunciation of "dangerous", only the pronunciation information of that word is updated, while the character's voiceprint vector, speech rate, pitch, energy, and boundary constraints remain unchanged. Then, the local area containing that word is resynthesized, which can correct the pronunciation error while preserving the original character's timbre and emotional expression.

[0119] S6.4.4 Perform a quality check on the locally corrected audio to generate target audio data; The processing terminal replaces the corresponding abnormal segments in the original candidate audio with the locally corrected audio, generates corrected candidate audio, and performs text correspondence evaluation, role and expression evaluation, boundary continuity evaluation and comprehensive quality calculation on the corrected candidate audio according to S5.1 to S5.4.

[0120] When the corrected candidate audio meets all quality conditions, the corrected candidate audio, its corresponding text unit number, and the actual audio duration are written into the target audio data. If the corrected candidate audio still fails to meet the quality conditions, the abnormal correction data is updated based on the evaluation items that did not meet the requirements and reprocessed. If the quality conditions are still not met after reprocessing a preset number of times, the corrected candidate audio and abnormal items are sent to the manual review interface, and the manually confirmed audio is written into the target audio data.

[0121] This process can prevent the introduction of new mispronunciations, timbre issues, and transition abnormalities after local resynthesis.

[0122] As an alternative implementation method, if the locally corrected target audio is directly connected along the text boundaries, it is easy to produce breaks, pops, and rhythmic jumps at obvious vocal positions. In view of this, the present invention proposes the following working method: S7. Perform audio splicing processing on the text unit data and target audio data to generate audiobook audio; S7.1 Perform semantic pause recognition on the text unit data to generate pause range data; The processing terminal arranges each text unit according to its text unit number and reads the corresponding pause information, voice actor data, scene identifier, and preceding and following text unit numbers. When the pause information of the current text unit indicates the end of a statement, or when the voice actors of adjacent text units are different, or when the scene identifiers are different, the connection position of the adjacent text units is determined as the semantic pause position.

[0123] Subsequently, the processing terminal determines the preset duration audio before the end of the previous target audio as the preceding pause search range, and determines the same preset duration audio after the start of the next target audio as the following pause search range. It then writes the preceding text unit number, the following text unit number, the semantic pause boundary, the preceding pause search range, and the following pause search range into the same record to generate pause range data.

[0124] For example, text unit 12 is Lin Zheng's voice saying "Did you go to the dock last night?", and text unit 13 is Zhou Yuan's voice saying "No, I've been at home the whole time". The processing terminal determines the first 100 milliseconds of the target audio corresponding to text unit 12 as the search range for the preceding pause, and determines the first 100 milliseconds of the target audio corresponding to text unit 13 as the search range for the following pause.

[0125] S7.2. Perform low-energy location extraction on the target audio data to generate low-energy location data; The processing terminal divides the target audio into frames, and then divides the sum of squares of the sample values ​​of each frame by the number of sample points to obtain the short-time energy of each audio frame.

[0126] The platform pre-acquires audio frames with natural pauses that have been manually confirmed, and uses the upper quantile value of their short-time energy distribution as the low-energy threshold, for example, the 90th percentile value. Then, the text unit number, the location of the audio frame with short-time energy not higher than the low-energy threshold, and its short-time energy are written into the same record to generate low-energy location data. Low-energy location indicates the location where the actual sound is weak, which can reduce the popping sound caused by direct connection in high-loudness sound areas.

[0127] S7.3 Perform position mapping on the pause range data and low energy position data to generate connection position data; The processing terminal filters low-energy positions within the preceding pause search range from the low-energy position data based on the previous text unit number to obtain the preceding candidate positions; then, based on the following text unit number, it filters low-energy positions within the following pause search range to obtain the following candidate positions.

[0128] The processing terminal combines each front candidate position with each rear candidate position to form a candidate connection position pair, and calculates the boundary difference value of each candidate connection position pair according to the normalization calculation method in S5.3.

[0129] Subsequently, the candidate connection position pair with the smallest boundary difference value is selected, and the previous text unit number, the next text unit number, the previous target audio connection position, and the next target audio connection position are written into the same record to generate connection position data.

[0130] This processing takes into account both the meaning of the language and the actual sound state, and can select the position with the smallest change in auditory perception from multiple natural pause positions.

[0131] For example, text unit 12 detects low-energy positions A1 and A2 within the search range of the preceding pauses in the target audio, and text unit 13 detects low-energy positions B1 and B2 within the search range of the following pauses in the target audio. The boundary difference value of the candidate connection position pair (A1, B1) is 0.18, and the boundary difference value of the candidate connection position pair (A2, B2) is 0.07. Therefore, the positions corresponding to A2 and B2 are written into the connection position data.

[0132] S7.4 Connect the target audio data according to the connection location data to generate audiobook audio; The processing terminal retains the audio before the previous target audio connection position and the audio after the next target audio connection position based on the connection position data; it extracts 100 milliseconds of audio before the previous target audio connection position as the front overlapping audio and extracts 100 milliseconds of audio after the next target audio connection position as the rear overlapping audio, aligns the two audio segments according to the sampling time, and forms a 100-millisecond short-time overlapping area.

[0133] At the beginning of the short-time overlap region, the weights of the preceding and following target audio are 1 and 0, respectively; at the quarter position, they are 0.75 and 0.25, respectively; at the middle position, they are 0.5 and 0.5, respectively; at the three-quarter position, they are 0.25 and 0.75, respectively; and at the end position, they are 0 and 1, respectively. The weights at the remaining positions change proportionally.

[0134] For example, at the one-quarter mark of the short-time overlap zone, the sample value of the preceding overlapping audio is 0.6, and the sample value of the following overlapping audio is 0.4. Therefore, the connection sample value at this position is 0.6 × 0.75 + 0.4 × 0.25 = 0.55. The processing terminal uses 0.55 as the output sample value of the audiobook audio at this sampling position and calculates the output sample values ​​at other sampling positions within the short-time overlap zone in the same way.

[0135] Subsequently, the target audios are connected sequentially according to the text unit number, and the chapter audios are combined according to the chapter identifier to generate audiobook audio, thereby making the waveforms and loudness of adjacent audios transition smoothly and reducing dropouts, pops and abrupt changes in rhythm.

[0136] like Figure 2As shown, compared with existing technologies, this invention improves the accuracy of voice character recognition, the consistency of voice characters, the naturalness of boundary transitions, and the efficiency of anomaly correction. Specifically, the joint judgment of multiple narrative relationships reduces character misidentification when the voice actor is omitted; the separation of stable voice identity from dynamic expression state helps maintain consistency in the voice of the same character throughout; boundary constraints and low-energy connection position filtering improve the transition between adjacent audio segments; and local resynthesis of abnormal positions reduces the processing load of overall regeneration and improves anomaly correction efficiency.

[0137] The embodiments of the present invention described above are subject to modification and change of method by those skilled in the art without departing from the embodiments and broader aspects of the present invention. The appended claims are intended to include all such modifications and changes of method that do not depart from the present invention.

[0138] Those skilled in the art should understand that the embodiments of the present invention can be implemented using a pure hardware architecture, a pure software architecture, or an integrated hardware and software architecture. The present invention can be prepared as a computer program product, which can be stored in various non-volatile computer-readable storage media, including but not limited to solid-state drives, flash memory chips, mobile storage devices, optical discs, cloud storage servers, and other standardized storage media, and is not limited to traditional storage media.

[0139] Based on the foregoing description in conjunction with the accompanying drawings, those skilled in the art will understand that the embodiments of this application can also be implemented by software programs. Therefore, this application also provides a computer-readable storage medium. This computer-readable storage medium stores computer-readable instructions thereon, which, when executed by one or more processors, implement the method described above in conjunction with the accompanying drawings.

[0140] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0141] It should be noted that although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart can be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0142] It should be understood that when the terms "first," "second," "third," and "fourth," etc., are used in the claims, specification, and drawings of this application, they are used only to distinguish different objects and not to describe a specific order. The terms "comprising" and "including" as used in the specification and claims of this application indicate the presence of the described features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof.

[0143] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application. As used in this specification and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this specification and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.

[0144] Although the embodiments of this application are described above, the content is merely an example adopted for the purpose of facilitating understanding of this application and is not intended to limit the scope and application scenarios of this application. Any person skilled in the art described in this application may make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed in this application, but the scope of patent protection of this application shall still be determined by the scope defined in the appended claims.

Claims

1. A method for generating audiobook audio based on speech synthesis, characterized in that, include: The novel text is retrieved and structured parsing is performed to obtain text unit data; Narrative relationship analysis is performed on the text unit data to obtain voice character data; Configure the character's voice based on the text unit data and the voice character data, and generate character voice control data; Perform front and back boundary constraint synthesis on the character voice control data to obtain candidate audio data; Perform a quality evaluation on the candidate audio data to obtain audio quality data; Based on the audio quality data, candidate audio data is filtered and anomalies are locally corrected to generate target audio data; The text unit data and the target audio data are subjected to audio splicing processing to generate audiobook audio.

2. The method for generating audiobooks based on speech synthesis according to claim 1, characterized in that, Narrative relationship analysis is performed on the text unit data to obtain voice character data, including: Role relationship extraction is performed on the text unit data to generate role relationship data; Perform candidate role evaluation on the role relationship data to generate corresponding role data; Based on the corresponding role data, role confirmation is performed, and voice-speaking role data is generated.

3. The method for generating audiobooks based on speech synthesis according to claim 2, characterized in that, Based on the corresponding role data, role confirmation is performed, and voice-speaking role data is generated, including: Obtain manually annotated verification text, perform threshold calibration on the data corresponding to the role, and obtain the role confirmation threshold and the role differentiation threshold; The corresponding data for each role is compared with the role confirmation threshold and the role differentiation threshold to generate voice-generating role data.

4. The method for generating audiobooks based on speech synthesis according to claim 1, characterized in that, Configure character voices based on the text unit data and voice character data, and generate character voice control data, including: Based on the voice actor data, configure a stable voice identity and generate character voice identity data; The expression state of the text unit data is updated to generate character expression state data; The text unit data is processed for pronunciation and pause adjustment to generate text control data; The character voice identity data, character expression status data, and text control data are collected to generate character voice control data.

5. The method for generating audiobooks based on speech synthesis according to claim 4, characterized in that, Perform front and rear boundary constraint synthesis on the character voice control data to obtain candidate audio data, including: The current text unit to be synthesized is determined from the text unit data, and the preceding boundary is determined to generate the preceding boundary data. Extract the voice control content of subsequent text units from the character voice control data to generate subsequent boundary data; The preceding boundary data and subsequent boundary data are aggregated to generate boundary constraint data; Speech synthesis is performed based on the character voice control data and boundary constraint data to generate candidate audio data.

6. The method for generating audiobooks based on speech synthesis according to claim 5, characterized in that, The preceding boundary data and subsequent boundary data are aggregated to generate boundary constraint data, including: Perform a correspondence analysis on the voice roles of the current text unit to be synthesized and the adjacent text units to generate voice relationship results; Based on the vocal relationship results, boundary constraints are processed on the preceding boundary data, subsequent boundary data, character voice control data, and text control data. Among them, continuous constraints are processed for relationships of the same character, and switching constraints are processed for relationships of different characters, thus generating boundary constraint data.

7. The method for generating audiobooks based on speech synthesis according to claim 1, characterized in that, Perform quality evaluation on the candidate audio data to obtain audio quality data, including: The candidate audio data is evaluated for text correspondence, and text correspondence data is generated. The candidate audio data is evaluated for roles and expressions to generate voice matching data; Perform boundary continuity evaluation on the candidate audio data to generate boundary evaluation data; The text correspondence data, sound correspondence data, and boundary evaluation data are aggregated to generate audio quality data.

8. The method for generating audiobooks based on speech synthesis according to claim 1, characterized in that, Based on the audio quality data, candidate audio data is filtered and anomalies are locally corrected to generate target audio data, including: Based on the audio quality data, the candidate audio data is filtered to generate audio filtering results; In response to the audio filtering results meeting the quality requirements, the target audio data is determined; In response to the audio screening results not meeting the quality conditions, abnormal correction data is generated; Based on the anomaly correction data, perform local re-synthesis and quality review to generate target audio data.

9. The method for generating audiobooks based on speech synthesis according to claim 8, characterized in that, Based on the aforementioned anomaly correction data, local re-synthesis and quality review are performed to generate target audio data, including: Extract the anomaly type, anomaly location, and pause information from the anomaly correction data; The abnormal text unit is determined based on the aforementioned anomaly correction data; Based on the anomaly type, anomaly location, and pause information, the local regeneration range is determined; The character voice control data corresponding to the local regeneration range is updated based on the anomaly correction data, and the updated character voice control data is synthesized by front and back boundary constraints to generate local corrected audio. The quality of the locally corrected audio is checked to generate target audio data.

10. The method for generating audiobooks based on speech synthesis according to claim 1, characterized in that, Perform audio concatenation processing on the text unit data and target audio data to generate audiobook audio, including: Semantic pause recognition is performed on the text unit data to generate pause range data; Low-energy location extraction is performed on the target audio data to generate low-energy location data; Perform position mapping on the pause range data and low energy position data to generate connection position data; The target audio data is connected based on the connection location data to generate audiobook audio.