Speech synthesis method based on multi-format file, terminal equipment and storage medium

By extracting, preprocessing, segmenting and splicing text data in multi-format files, the problems of information loss and speech coherence during text extraction in traditional technologies are solved, and an efficient speech synthesis method is realized.

CN120071889AInactive Publication Date: 2025-05-30SHENZHEN MAIFENG TECH CO LTD

Patent Information

Application Number
CN202510528411.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional techniques tend to lose or misprocess structured content when performing text extraction, resulting in a lack of coherence in the generated audio information.

Method used

A speech synthesis method based on multi-format files is provided, by extracting original text data in multi-format files, pre-processing to remove interfering content, segmenting text data to preserve semantic structures, and generating spliced ​​audio clips to achieve coherent audio output.

Benefits of technology

It effectively solves the problem of information loss in complex document format processing, maintains text semantic coherence, and realizes seamless connection of ultra-long text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071889A_ABST
    Figure CN120071889A_ABST
Patent Text Reader

Abstract

The invention is suitable for the field of data processing, and discloses a speech synthesis method based on a multi-format file, terminal equipment and a storage medium. The speech synthesis method based on the multi-format file comprises the following steps: extracting original text data in the multi-format file, wherein the multi-format file comprises an electronic document and an image; preprocessing the original text data to obtain target text data, the preprocessing including interference content filtering operation; segmenting the target text data to obtain a plurality of text segments, the segmentation operation being performed based on text semantic structure hierarchy and character number limitation; generating a plurality of audio clips corresponding to the plurality of text clips; and splicing the plurality of audio clips to obtain a target audio file. According to the method, the audio file synthesized based on the multi-format file can completely present the original text structural characteristics, and seamless connection of the super-long text can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing, and particularly relates to a voice synthesis method, a terminal device, and a storage medium based on multi-format files. Background Art

[0002] With the rapid development of artificial intelligence technology, traditional work has begun to be automated. For example, in the fields of voice synthesis and text-to-speech, the application of artificial intelligence has made remarkable progress.

[0003] Since the formats and structures of processing targets may be very complex, such as including document files in different formats, and there are contents such as titles, paragraphs, headers, and footers in the documents, traditional technologies are prone to losing or misprocessing these structured contents during text extraction, resulting in the lack of coherence of the generated audio information. A new technical means is needed to solve the above technical problems. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a voice synthesis method, a terminal device, and a storage medium based on multi-format files, which can solve the problem that structured contents are easily lost or misprocessed during text extraction in related technologies, resulting in the lack of coherence of the generated audio information.

[0005] The first aspect of the present invention provides a voice synthesis method based on multi-format files, including: Extracting the original text data from the multi-format files, where the multi-format files include electronic documents and images; Performing preprocessing on the original text data to obtain target text data, where the preprocessing includes an interference content filtering operation; Segmenting the target text data to obtain multiple text segments, where the segmentation operation is based on the text semantic structure level and the character number limit; Generating multiple audio segments corresponding to the multiple text segments; Stitching the multiple audio segments to obtain a target audio file.

[0006] Optionally, in the first implementation manner of the first aspect of the present invention, the step of performing preprocessing on the original text data to obtain target text data includes: Identifying and deleting non-body content elements in the text to obtain the target text data, where the non-body content elements include at least one of headers, footers, and page numbers, and the identification operation is implemented based on a layout analysis algorithm.

[0007] Optionally, in the second implementation manner of the first aspect of the present invention, the step of segmenting the target text data to obtain multiple text segments includes: Perform multi-level segmentation on the target text data to obtain multiple text segments. The first-level segmentation in the multi-level segmentation includes dividing text blocks based on the chapter and section directory structure, and the second-level segmentation in the multi-level segmentation includes secondary segmentation of text blocks exceeding a preset character threshold based on paragraphs.

[0008] Optionally, in the third implementation manner of the first aspect of the present invention, the step of generating multiple audio segments corresponding to the multiple text segments includes: Dynamically select a target platform from multiple speech synthesis platforms; Call the speech synthesis interface of the target platform, and input the multiple text segments in parallel into the corresponding speech synthesis model to generate audio segments with consistent timbre characteristics. The speech synthesis parameters include at least one of a timbre identifier, a speech rate parameter, and an emotion label.

[0009] Optionally, in the fourth implementation manner of the first aspect of the present invention, the step of dynamically selecting a target platform from multiple speech synthesis platforms includes: Receive a timbre selection instruction initiated by the user, and the selection instruction triggers the invocation of a cross-platform timbre library; Dynamically select the target platform from the multiple speech synthesis platforms according to the speech synthesis parameters in the selection instruction, where the speech synthesis parameters include at least one of a timbre identifier, a speech rate parameter, and an emotion label.

[0010] Optionally, in the fifth implementation manner of the first aspect of the present invention, the step of extracting the original text data from the multi-format file includes: Call a format conversion interface to perform text extraction on the electronic document format file in the multi-format file, and perform optical character recognition processing on the image format file in the multi-format file to obtain the original text data.

[0011] Optionally, in the sixth implementation manner of the first aspect of the present invention, the step of splicing the multiple audio segments to obtain a target audio file includes: Analyze the encoding format, bit rate, sampling rate, and channel parameters of the multiple audio segments, and uniformly convert the multiple audio segments into a target encoding format, a target sampling rate, and a target channel to obtain each target audio segment; Splice each of the target audio segments to obtain an audio file; Perform dynamic bit rate equalization processing on the audio file to obtain the target audio file.

[0012] Optionally, in the seventh implementation manner of the first aspect of the present invention, the step of extracting the original text data from the multi-format file includes: If the electronic document format file is a Word document, extract the final version text from the revision records of the Word document; If the electronic document format file is a PDF document, identify the mixed content of the text layer and the scanned layer of the PDF document, and preferentially extract the text layer text.

[0013] In a second aspect, an embodiment of the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned voice synthesis method based on multi-format files are implemented.

[0014] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned voice synthesis method based on multi-format files are implemented.

[0015] In a fourth aspect, an embodiment of the present invention provides a computer program product. When the computer program product runs on a terminal device, the terminal device is enabled to execute the above-mentioned voice synthesis method based on multi-format files.

[0016] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows: By supporting the multi-modal input of files and combining structured text extraction and interference filtering, the problem of information loss in the processing of complex document formats is effectively solved; Based on the dual segmentation of semantic levels and character limits, compared with traditional methods, the problem of character number limits can be solved, and the semantic coherence of the text can be maintained through the segmentation logic of chapter / paragraph priority; Cooperating with multi-platform audio generation and splicing technology, the finally synthesized audio file can not only completely present the original text structure features, but also achieve seamless connection of ultra-long texts. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a schematic diagram of an embodiment of the voice synthesis method based on multi-format files in an embodiment of the present invention; Figure 2 It is a schematic diagram of a specific embodiment of step S101 of the voice synthesis method based on multi-format files in an embodiment of the present invention; Figure 3It is a schematic diagram of a specific embodiment of step S1041 of the speech synthesis method based on multi-format files in the embodiment of the present invention; Figure 4 It is a schematic diagram of a specific embodiment of step S105 of the speech synthesis method based on multi-format files in the embodiment of the present invention; Figure 5 It is a schematic diagram of an embodiment of a terminal device in the embodiment of the present invention. Detailed implementation manners

[0019] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0020] It should be noted that the terms "including", "comprising" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusion. For example, a process, method, terminal, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices. In the terms of the present invention in the claims, specification and specification drawings, relational terms such as "first" and "second" are only used to distinguish one entity / operation / object from another entity / operation / object, and do not necessarily require or imply any such actual relationship or order between these entities / operations / objects.

[0021] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0022] With the rapid development of artificial intelligence technology, traditional work has begun to be automatically processed. For example, in the fields of speech synthesis and text-to-speech, significant progress has been made in the application of artificial intelligence.

[0023] Since the formats and structures of processing targets may be complex, for example, including document files in different formats, and there are contents such as titles, paragraphs, headers, and footers in the documents, traditional technologies are prone to losing or misprocessing these structured contents during text extraction, resulting in the lack of coherence in the generated audio information. A new technical means is needed to solve the above technical problems.

[0024] In view of this, the embodiments of the present invention provide a voice synthesis method, a terminal device, and a storage medium based on multi-format files. By supporting multi-modal input of files and combining structured text extraction and interference filtering, the problem of information loss in complex document format processing is effectively solved; based on the dual segmentation of semantic levels and character limits, compared with traditional methods, the problem of character number limits can be solved, and the text semantic coherence can be maintained through the segmentation logic of chapter / paragraph priority; combined with multi-platform audio generation and splicing technologies, the finally synthesized audio file can not only completely present the original text structure features but also achieve seamless connection of ultra-long texts.

[0025] In order to illustrate the technical solution of the present invention, the following will be described through specific embodiments.

[0026] Figure 1 The figure shows a schematic flowchart of a voice synthesis method based on multi-format files provided by the embodiments of the present invention. This method can be applied to a terminal device. The terminal device can be a mobile phone, a tablet computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, etc.

[0027] Specifically, the above voice synthesis method based on multi-format files may include the following steps S101 to step S105.

[0028] Step S101, extract the original text data in the multi-format file. The multi-format file includes electronic documents and images.

[0029] In the embodiment of the present invention, the terminal device (such as a voice synthesis device based on multi-format files) receives the electronic documents (including formats such as Word and PDF) and image files uploaded by the user and starts the multi-modal processing process. For electronic documents, call the format conversion interface to parse the file structure (such as extracting the revised text of the Word document and preferentially extracting the text layer text of the PDF document); perform OCR recognition on the image file, extract the text content, and generate the original text data. This stage completes the unified text processing of different format files.

[0030] Step S102, perform preprocessing on the original text data to obtain target text data. The preprocessing includes an interference content filtering operation.

[0031] In an embodiment of the present invention, for the extracted original text data, interference content filtering is performed. For example, non-text elements such as headers, footers, and page numbers are identified through a layout analysis algorithm, and redundant information (such as document watermarks, format symbols) is deleted based on semantic rules. The output is the target text data that only retains core content such as titles and paragraphs, ensuring the purity and semantic integrity of subsequent processing.

[0032] Step S103: Split the target text data to obtain multiple text segments. The splitting operation is performed based on the text semantic structure level and the character number limit.

[0033] In an embodiment of the present invention, based on the chapter structure of the target text (such as the table of contents, title level), the first-level splitting is performed to generate logical text blocks.

[0034] If the number of characters in a single block exceeds the preset threshold (such as 20,000 words), the second-level splitting is triggered, and it is further split according to paragraph or full stop boundaries to avoid semantic breaks. The split text segments are stored in the queue order, providing input for segmented speech synthesis.

[0035] Step S104: Generate multiple audio segments corresponding to the multiple text segments.

[0036] In an embodiment of the present invention, according to the voice parameters selected by the user (such as voice identifier, speech rate, emotion label), the interfaces of multiple speech synthesis platforms are dynamically matched; the text segment queue is sent to the selected platform in parallel, and its speech synthesis model is called to generate audio segments. Through parameter synchronization technology, the consistency of the voice and speech rate of each segment is ensured, and the sound quality difference in cross-platform synthesis is eliminated.

[0037] Step S105: Concatenate the multiple audio segments to obtain the target audio file.

[0038] In an embodiment of the present invention, the encoding format, bit rate, and sampling rate parameters of each audio segment are parsed and uniformly converted into the target format (such as MP3, 44.1 kHz sampling rate); through the timestamp alignment and dynamic bit rate equalization algorithm, the silent gaps and volume fluctuations between segments are eliminated to achieve seamless concatenation. Finally, the complete target audio file is output, and the target audio file can retain the original text semantic structure and have no obvious concatenation traces.

[0039] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows: By supporting the multi-modal input of files and combining structured text extraction and interference filtering, the problem of information loss in the processing of complex document formats is effectively solved; based on the dual segmentation of semantic levels and character limits, compared with traditional methods, the problem of character number limits can be solved, and the semantic coherence of the text can be maintained through the segmentation logic of chapter / paragraph priority; combined with the multi-platform audio generation and splicing technology, the finally synthesized audio file can not only completely present the original text structure features, but also achieve seamless connection of ultra-long texts.

[0040] When traditional text-to-speech technology processes complex-format documents, it often lacks the ability of structured analysis, resulting in the inability to distinguish the main text area from the non-main text area, and mistakenly synthesizing interference content such as headers and page numbers into the audio, causing information redundancy. Based on this, the present invention proposes an alternative embodiment.

[0041] Step S102 further includes the following specific embodiments.

[0042] Step S1021, identify and delete the non-main text content elements in the text to obtain target text data. The non-main text content elements include at least one of headers, footers, and page numbers, and the identification operation is implemented based on a layout analysis algorithm.

[0043] In the embodiment of the present invention, when the original text extraction of the multi-format file is completed, the preprocessing process is started. At this time, the text data carries the original layout information (such as headers, footers, page numbers, etc.) and enters the interference content identification stage. Call the pre-trained layout analysis algorithm to parse the layout structure of the text. The algorithm distinguishes the main text area from the non-main text area based on feature recognition technology and locates the coordinate ranges of headers, footers, and page numbers.

[0044] According to the layout analysis result, mark and delete the identified non-main text content elements. For example, delete the company name in the header, the copyright statement in the footer, and the page number numbers scattered on each page of the document, while retaining the core main text content such as titles and paragraphs.

[0045] Output the target text data containing only the main text to ensure that the subsequent segmentation and speech synthesis stages are not interfered by redundant information. The text data generated in this stage has removed format symbols (such as page break symbols) and irrelevant content, forming a semantically coherent input source.

[0046] In the embodiments of the present invention, through the layout analysis algorithm, accurate recognition and deletion of non-text content are achieved, effectively solving the problem of audio information loss caused by misprocessing of structured content in traditional technologies. This embodiment can improve the accuracy of extracting the document text, avoid interference elements such as headers / footers from being mixed into the speech synthesis process, enable the generated audio file to fully present the core content of the original text, and at the same time reduce speech pauses or logical breaks caused by redundant information.

[0047] Traditional text segmentation technologies have significant defects. For example, using single-character count segmentation results in the forced breakup of chapter content, leading to logical jumps during audio playback. Based on this, the present invention proposes an alternative embodiment.

[0048] Step S103 also includes the following specific implementation manners.

[0049] Step S1031, perform multi-level segmentation on the target text data to obtain multiple text segments. The first-level segmentation in the multi-level segmentation includes dividing text blocks based on the chapter and section directory structure, and the second-level segmentation in the multi-level segmentation includes secondary segmentation of text blocks exceeding a preset character threshold based on paragraphs.

[0050] In the embodiments of the present invention, when the target text data is preprocessed, the multi-level segmentation engine is activated. First, read the global structural features of the text, determine whether structured segmentation is required, and load the preset character count threshold parameter (such as 20,000 characters). Identify the chapter and section directory structure of the document based on the semantic analysis algorithm, including chapter titles (such as marks like "Chapter 1", "2.1", etc.) and directory levels (through typesetting features such as font size, bold style, etc.). Cut the text into multiple logical blocks according to the recognition results, with each block corresponding to an independent chapter or sub-chapter, so that the segmented text blocks are self-contained units semantically.

[0051] Count the number of characters in each chapter text block. If the number of characters in a certain block exceeds the preset threshold (such as 20,000 words), trigger the second-level segmentation mechanism. Dynamically adjust the threshold range according to actual needs (for example, allow a ±10% floating range) to avoid unreasonable segmentation granularity caused by absolute thresholds.

[0052] Perform paragraph-level segmentation on the over-limit chapter blocks, and first identify paragraph delimiters (such as line breaks, first-line indentation) or natural full-stop boundaries, so that each sub-segment maintains paragraph integrity; if there is no clear paragraph mark, select the best segmentation point according to the semantic coherence algorithm (such as based on the continuity of context keywords) to prevent semantic breaks. Store the text segments after multi-level segmentation into the processing queue in the original chapter order, and at the same time attach structure tags (such as "Chapter3_Paragraph2") to maintain the text logical relationship.

[0053] In the embodiments of the present invention, through a multi-level segmentation strategy guided by chapter headings and contents, the traditional fixed-length segmentation method is improved to semantic-driven dynamic cutting. The primary segmentation based on the chapter structure can preserve the logical integrity of the text, making the content after audio synthesis exactly correspond to the original text structure; the secondary paragraph segmentation triggered by the character threshold can break through the capacity limit of single-file processing and avoid semantic interruption problems caused by mechanical segmentation.

[0054] Traditional speech synthesis solutions are limited by the capacity of the voice library on a single platform and cannot meet diverse requirements. Moreover, the serial synthesis mode causes a significant increase in the time-consuming for processing large texts. When synthesizing in batches, due to platform switching or parameter deviation, there are obvious differences in tone / speech rate between audio segments. Based on this, the present invention proposes an alternative embodiment.

[0055] Refer to Figure 2 , Figure 2 which is a schematic diagram of a specific embodiment of step S104 of the speech synthesis method based on multi-format files in the embodiments of the present invention. Step S104 further includes the following specific implementations.

[0056] Step S1041, dynamically select a target platform from multiple speech synthesis platforms.

[0057] In the embodiments of the present invention, when the text segmentation is completed and a fragment queue is generated, the multi-platform speech synthesis scheduling module is activated. At this time, the speech synthesis parameter configuration interface is initialized, and the list of available voice libraries (including self-developed and third-party platform voice resources) is loaded to prepare to receive user instructions.

[0058] Receive the voice selection instruction submitted by the user through the interaction interface, including configuration information such as voice identifier (such as "female voice - broadcast accent"), speech rate parameter (such as 1.2 times speed), and emotion label (such as "serious", "pleased"). Convert the user's requirements into a standardized set of speech synthesis parameters and trigger the cross-platform voice matching process.

[0059] According to the voice parameters configured by the user, dynamically select a target platform from the registered multiple speech synthesis platforms. The matching rules include but are not limited to calculating the similarity of voice characteristics, detecting the availability of platform services, and evaluating the synthesis delay. Preferably, select the platform with a parameter matching degree ≥ 90% and the fastest response speed.

[0060] Step S1042, call the speech synthesis interface of the target platform, and input multiple text fragments in parallel into the corresponding speech synthesis model to generate audio segments with consistent voice characteristics. The speech synthesis parameters include at least one of voice identifier, speech rate parameter, and emotion label.

[0061] In an embodiment of the present invention, the text segment queue is distributed to the speech synthesis interface of the target platform according to priorities, and parallel processing is achieved through multi-threading technology. Each thread independently calls the platform API, binds the text segment with user parameters (such as timbre identifier, speech rate, etc.), and inputs them into the synthesis model to generate an audio segment file with unified timbre characteristics.

[0062] Perform parameter synchronization detection on the generated audio segments. The timbre consistency can be verified through voiceprint analysis; the speech rate fluctuation can be calibrated based on timestamps. Trigger the resynthesis mechanism for unqualified segments to make all audio segments seamlessly connect in terms of timbre and rhythm.

[0063] In the embodiments of the present invention, through multi-platform dynamic scheduling and parallel synthesis technology, the limitations of timbre selection and low processing efficiency of traditional single speech engines are solved. Based on the optional timbre library, users can match the best timbre across platforms; at the same time, the parallel processing technology improves the audio generation speed, and the parameter synchronization mechanism can ensure the consistency of timbre and speech rate of each segment, avoiding the problem of sound quality jump caused by traditional batch-by-batch synthesis.

[0064] In traditional speech synthesis systems, timbre selection is limited by the storage capacity of a single platform. Users often compromise and use non-optimal solutions because they cannot find a matching timbre; moreover, parameter adjustment requires manual switching of platforms or repeated configuration, which is cumbersome and there is no real-time preview. Based on this, the present invention proposes an optional embodiment.

[0065] Refer to Figure 3 , Figure 3 is a schematic diagram of a specific embodiment of step S1041 of the speech synthesis method based on multi-format files in the embodiments of the present invention. Step S1041 further includes the following specific implementation manners.

[0066] Step S10411, receive the timbre selection instruction initiated by the user, and the selection instruction triggers the invocation of the cross-platform timbre library.

[0067] In an embodiment of the present invention, the user can initiate a timbre selection instruction through the interaction interface, such as browsing the cross-platform timbre library; selecting the target timbre identifier; customizing the speech parameters. The terminal device captures the user operations in real time and generates a structured parameter instruction set.

[0068] Step S10412, dynamically select the target platform from multiple speech synthesis platforms according to the speech synthesis parameters in the selection instruction, where the speech synthesis parameters include at least one of a timbre identifier, a speech rate parameter, and an emotion label.

[0069] In an embodiment of the present invention, the timbre identifier selected by the user is matched with a multi-platform timbre feature database, for example, through a quantized comparison of voiceprint features (such as MFCC coefficient matching), a list of candidate platforms is screened out, and the service status of each platform is verified (such as API availability, response delay).

[0070] Multi-dimensional weight calculation is performed based on user parameters (speech speed, emotion label) and platform performance indicators (synthesis quality score, real-time load). For example, platforms that support emotion labels and whose speech speed adjustment range covers user needs are given priority. If multiple platforms meet the conditions, the target platform with the lowest latency (e.g. <200ms) or the highest synthesis accuracy (e.g. 98% similarity) is selected.

[0071] Encapsulate the speech synthesis parameters (timbre ID, speech rate, emotion label) configured by the user into platform-compatible API request parameters, and call the speech synthesis interface of the target platform. The parameter format differences between different platforms can be automatically processed.

[0072] The timbre synthesis results of the target platform are fed back to the user-side preview module in real time, supporting segment audition and parameter fine-tuning. If the user modifies the parameters (such as switching the emotional label to "calm"), the platform matching and interface call process will be re-triggered to form a closed-loop interaction.

[0073] In the embodiment of the present invention, a user-driven cross-platform dynamic timbre selection mechanism is used to break through the timbre limitations of a single engine and improve user matching. At the same time, based on a platform optimization strategy with multi-dimensional weights, synthesis quality and response speed are balanced.

[0074] Traditional text extraction solutions cannot distinguish text layers from scanned images when processing PDFs, resulting in duplicate OCR or missing text; and Word document extraction ignores revision records, which may output unfinalized erroneous content. Based on this, the present invention proposes an optional embodiment.

[0075] Step S101 also includes the following specific implementation methods.

[0076] Step S1011 , calling a format conversion interface to perform text extraction on an electronic document format file in a multi-format file, and performing optical character recognition processing on an image format file in a multi-format file, so as to obtain original text data.

[0077] In an embodiment of the present invention, a mixed format file (such as Word, PDF, JPG / PNG image) uploaded by a user is received, and a format recognition module is started. File types are distinguished according to file extensions or binary feature codes, for example, electronic documents (Word, PDF) are classified into a document processing flow, and image files (JPG, PNG, etc.) are classified into an image processing flow, thereby realizing automatic classification of multimodal input.

[0078] Call the Office format conversion interface (such as the Apache POI library) for Word documents, parse the.docx file structure, extract the body text (including table text), the final version content in the revision record, and retain the paragraph marks; for PDF documents, use tools such as PDFBox to parse the text layer content, give priority to extracting editable text, and skip the image area if there is a scanned layer (to avoid duplicate processing with the subsequent OCR process). Start the optical character recognition (OCR) engine (such as Tesseract) for image files, perform preprocessing operations, such as performing image binarization and noise reduction; perform layout analysis to identify text areas; switch multi-language models (automatically select the training library according to the text language); output text data with coordinate information, retaining paragraph line breaks and the original text layout order of the original image.

[0079] Unify the encoding (such as converting to UTF-8) of the text stream extracted from the electronic document and the OCR recognition result, and merge them into the original text data set. For the case of mixed electronic documents and scanned pages in the same file (such as a PDF containing scanned picture pages), perform automatic separation processing, for example, directly extract the text layer pages and transfer the scanned pages to the OCR process.

[0080] Perform preliminary cleaning on the extracted original text, such as deleting invisible characters and fixing encoding errors (such as replacing garbled symbols), and store it in the temporary database in the original file order to provide structured input for the subsequent preprocessing stage.

[0081] In the embodiments of the present invention, through the collaborative processing mechanism of electronic document format parsing and image OCR, high-precision text extraction of multi-format files can be achieved. Retaining the revision version information for Word documents avoids text omission caused by traditional solutions ignoring revision records; the priority extraction strategy for the PDF text layer can improve the text recognition accuracy of mixed documents. The intelligent preprocessing of the OCR process can reduce the character error rate of image-to-text conversion.

[0082] In traditional audio splicing technology, differences in audio formats / parameters synthesized by different platforms are likely to cause splicing failures or sudden changes in sound quality; simple chronological splicing will produce perceptible silent intervals or volume jumps; manual adjustment relies on professional software, which is time-consuming and difficult to ensure consistency. Based on this, an alternative embodiment of the present invention is proposed.

[0083] Refer to Figure 4 , Figure 4 is a schematic diagram of a specific embodiment of step S105 of the voice synthesis method based on multi-format files in the embodiments of the present invention, and step S105 further includes the following specific implementation manners.

[0084] Step S1051: Analyze the encoding format, bit rate, sampling rate, and channel parameters of multiple audio segments, and uniformly convert the multiple audio segments into a target encoding format, a target sampling rate, and a target channel to obtain respective target audio segments.

[0085] In an embodiment of the present invention, read all generated audio segment files, and call an audio codec library to analyze the metadata of each segment, including the encoding format (such as MP3, WAV), bit rate (such as 128 kbps), sampling rate (such as 44.1 kHz), and number of channels (such as stereo / mono). Generate a unified conversion instruction set according to preset target parameters (such as MP3 format, 48 kHz sampling rate, dual channels).

[0086] Perform format conversion on audio segments with inconsistent encoding formats, use an audio resampling algorithm to adjust the sampling rate, match the target number of channels through channel merging / splitting techniques, and unify different bit rates into a target bit rate (such as a constant 192 kbps) through a dynamic bit rate compression algorithm. Preserve the audio timestamp information during the conversion process to ensure accurate subsequent splicing timing.

[0087] Step S1052: Splice the respective target audio segments to obtain an audio file.

[0088] In an embodiment of the present invention, load the standardized audio file in the original order of the text segments and perform head-to-tail connection based on timestamps. Eliminate the silent gaps between segments according to the crossfade algorithm. For example, take 200 ms of audio data at the end of the previous segment and the beginning of the next segment, and achieve a smooth transition through linear superposition to avoid the abruptness caused by mechanical splicing.

[0089] Step S1053: Perform dynamic bit rate equalization processing on the audio file to obtain a target audio file.

[0090] In an embodiment of the present invention, perform global equalization optimization on the spliced complete audio file. For example, analyze the volume fluctuation curve of the entire file, and use the dynamic range compression (DRC) algorithm to control the volume difference within ±3 dB; detect and eliminate residual background noise pulses (such as clicks, current sounds); adaptively adjust the bit rate distribution according to the target file size requirement (such as low bit rate for dialogue parts and high bit rate for music parts) to achieve an optimal balance between audio quality and file size.

[0091] After generating the final target audio file, perform automatic verification. For example, compare the deviation between the audio duration and the estimated text duration (error ≤ 1%); detect the waveform continuity at the splicing points (transition smoothness score ≥ 95 points); verify the metadata integrity (such as tag information, copyright statement). If the verification fails, trigger the re-synthesis or splicing process of the specified segments.

[0092] In the embodiments of the present invention, through audio format standardization and intelligent splicing technology, the industry problems of poor compatibility of multi-source audio segments and obvious splicing traces are solved.

[0093] When extracting text from traditional electronic documents, revision records are often ignored, resulting in the output text containing unconfirmed modification traces or missing the content of the final version; when parsing PDF, it is impossible to distinguish between the text layer and the scanned image, either only part of the editable text is extracted, or the whole document is forced to perform OCR. Based on this, an alternative embodiment of the present invention is proposed.

[0094] Step S101 also includes the following specific implementation manners.

[0095] Step S1012, if the electronic document format file is a Word document, extract the final version text in the revision record of the Word document.

[0096] In the embodiments of the present invention, when it is detected that the file uploaded by the user is a Word document, call the Office document parsing interface (such as Apache POI) to traverse the document revision record. Extract all the content of the final version after accepting the revisions, and at the same time filter out the rejected modification traces and annotation information, so that the output text only contains the final valid content confirmed by the author.

[0097] Perform layer parsing on the PDF document, and identify the text layer (selectable and editable text objects) and the scanned layer (uneditable content embedded in the form of a picture) in the document through a PDF parsing library (such as PDFBox). Give priority to extracting the text data in the text layer, and at the same time mark the page numbers where the scanned layer is located to avoid repeated OCR processing of the image area.

[0098] For PDF pages containing a mixed layer of text and scanned content, start a hierarchical processing mechanism. The content of the text layer is directly extracted and the original layout order is retained; the scanned layer area triggers image preprocessing (such as sharpening and contrast adjustment), and then call the OCR engine for text recognition; merge the OCR results of the text layer and the scanned layer in page order to make the text stream consistent with the original layout.

[0099] Perform integrity verification on the extraction results, for example, check the revision record coverage rate of the Word document; verify the text layer extraction rate of the PDF document; output an error log to prompt the user to process the abnormal page.

[0100] Step S1013, if the electronic document format file is a PDF document, identify the mixed content of the text layer and the scanned layer in the PDF document, and give priority to extracting the text in the text layer.

[0101] In an embodiment of the present invention, the standardized audio files are loaded in the original order of the text segments and joined head to tail based on timestamps. According to the waveform crossfade algorithm, the silent gaps between segments are eliminated. For example, 200 ms of audio data is taken from the end of the previous segment and the beginning of the next segment respectively, and smooth transition is achieved through linear superposition, avoiding the abrupt feeling caused by mechanical splicing.

[0102] In an embodiment of the present invention, through the differential processing strategy for Word revision versions and PDF layers, the accuracy and integrity of text extraction from electronic documents can be significantly improved.

[0103] As Figure 5 shown, it is a schematic diagram of a terminal device provided by an embodiment of the present invention. The terminal device 500 may include: a processor 501, a memory 502, and a computer program 503 stored in the memory 502 and executable on the processor 501, such as a speech synthesis program based on multi-format files. When the processor 501 executes the computer program 503, the steps in the above-mentioned various speech synthesis embodiments based on multi-format files are implemented.

[0104] The computer program may be divided into one or more modules / units. One or more modules / units are stored in the memory 502 and executed by the processor 501 to complete the present invention. One or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the terminal device.

[0105] The terminal device may include, but is not limited to, a processor 501 and a memory 502. Those skilled in the art can understand that Figure 5 this is only an example of the terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the terminal device may also include input / output devices, network access devices, buses, etc.

[0106] The so-called processor 501 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0107] The memory 502 may be an internal storage unit of the terminal device, such as the hard disk or memory of the terminal device. The memory 502 may also be an external storage device of the terminal device, such as a plug-in hard disk equipped on the terminal device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 502 may also include both an internal storage unit and an external storage device of the terminal device. The memory 502 is used to store computer programs and other programs and data required by the terminal device. The memory 502 may also be used to temporarily store data that has been output or will be output.

[0108] It should be noted that for the convenience and brevity of description, the structure of the above terminal device may also refer to the specific description of the structure in the method embodiment, which will not be elaborated here.

[0109] The embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above voice synthesis method based on multi-format files can be implemented.

[0110] The embodiment of the present invention provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can be made to execute the steps in the above voice synthesis method based on multi-format files when executed.

[0111] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0112] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0113] In the embodiments provided by the present invention, it should be understood that the disclosed terminal device and method can be implemented in other ways. For example, the above-described terminal device embodiments are merely illustrative. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0114] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0115] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, may exist physically alone for each unit, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0116] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0117] The above-mentioned embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A speech synthesis method based on multi-format files, characterized in that: include: Extracting original text data from a multi-format file, the multi-format file including an electronic document and an image; Performing preprocessing on the original text data to obtain target text data, wherein the preprocessing includes an interference content filtering operation; Segmenting the target text data to obtain a plurality of text segments, wherein the segmentation operation is performed based on the text semantic structure level and the character quantity limit; Generating a plurality of audio segments corresponding to the plurality of text segments; The multiple audio clips are concatenated to obtain a target audio file.

2. The speech synthesis method based on multi-format files as claimed in claim 1, characterized in that: The step of performing preprocessing on the original text data to obtain target text data comprises: Identify and delete non-text content elements in the text to obtain the target text data, wherein the non-text content elements include at least one of a header, a footer, and a page number, and the identification operation is implemented based on a layout analysis algorithm.

3. The speech synthesis method based on multi-format files as claimed in claim 1, characterized in that: The step of segmenting the target text data to obtain a plurality of text segments comprises: Multi-level segmentation is performed on the target text data to obtain multiple text fragments, wherein the first level segmentation in the multi-level segmentation includes dividing the text blocks based on the chapter directory structure, and the second level segmentation in the multi-level segmentation includes secondary segmentation of the text blocks exceeding a preset character threshold based on paragraphs.

4. The speech synthesis method based on multi-format files as claimed in claim 1, characterized in that: The step of generating a plurality of audio segments corresponding to the plurality of text segments comprises: Dynamically select a target platform from multiple speech synthesis platforms; The speech synthesis interface of the target platform is called, and the multiple text segments are input into the corresponding speech synthesis model in parallel to generate audio segments with consistent timbre characteristics, wherein the speech synthesis parameters include at least one of a timbre identifier, a speech rate parameter, and an emotion label.

5. The speech synthesis method based on multi-format files as claimed in claim 4, characterized in that: The step of dynamically selecting a target platform from a plurality of speech synthesis platforms comprises: Receiving a timbre selection instruction initiated by a user, wherein the selection instruction triggers a cross-platform timbre library call; The target platform is dynamically selected from the multiple speech synthesis platforms according to the speech synthesis parameters in the selection instruction, wherein the speech synthesis parameters include at least one of a timbre identifier, a speech speed parameter, and an emotion tag.

6. The method for speech synthesis based on multi-format files as claimed in claim 1, characterized in that: The step of extracting original text data from the multi-format file comprises: The format conversion interface is called to perform text extraction on the electronic document format file in the multi-format file, and optical character recognition processing is performed on the image format file in the multi-format file to obtain the original text data.

7. The method for speech synthesis based on multi-format files as claimed in claim 1, characterized in that: The step of splicing the multiple audio clips to obtain a target audio file comprises: Parsing the encoding formats, bit rates, sampling rates, and channel parameters of the multiple audio clips, and converting the multiple audio clips into a target encoding format, a target sampling rate, and a target channel to obtain each target audio clip; Splicing the target audio clips to obtain an audio file; Dynamic bit rate equalization processing is performed on the audio file to obtain the target audio file.

8. The method for speech synthesis based on multi-format files as claimed in claim 1, characterized in that: The step of extracting original text data from the multi-format file comprises: If the electronic document format file is a Word document, extracting the final version text in the revision record of the Word document; If the electronic document format file is a PDF document, the mixed content of the text layer and the scan layer of the PDF document is identified, and the text of the text layer is preferentially extracted.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the multi-format file-based speech synthesis method as claimed in any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the speech synthesis method based on multi-format files as claimed in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Entertainment audio file for text-only application

    CN103200309A

  • Method and apparatus for merging multimedia files

    CN108184079A

  • Scene-based aloud-reading audio production method and system based on handheld intelligent terminal

    CN108536655A

  • Speech synthesis method and device, computer equipment and storage medium

    CN112634865A

  • Speech synthesis proxy method and device, electronic equipment and readable storage medium

    CN113327571A

Cited By

  • Streaming audio synthesis method and device, storage medium and electronic device

    CN121600905A