Audio processing method and device, electronic equipment and storage medium
By segmenting, classifying, and splicing text or voice input, and selecting timbres and audio files according to different scenario types, the problem of poor user experience caused by the single timbre in existing technologies is solved, thus improving the user experience of speech synthesis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN LUKA DR TECHNOLOGY CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies suffer from a lack of variety in timbre and pitch when processing speech synthesis in scenarios involving mixed Chinese and English, poetry, nursery rhymes, and stories, resulting in a poor user experience.
By dividing, classifying, and splicing text or voice input, and selecting appropriate timbres and audio files according to different scene types, the audio data required for synthesis can be generated.
It enhances the emotional impact and usage scenarios of the output audio, optimizes the single timbre and tone, and improves the user experience.
Smart Images

Figure CN122116866A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and more particularly to an audio processing method, apparatus, electronic device, and storage medium. Background Technology
[0002] Currently, when using text-to-speech or speech-to-speech services, especially in scenarios involving mixed Chinese and English text, poems, nursery rhymes, or stories, only one timbre and tone can be used. This makes the entire speech playback process very monotonous, rigid, and awkward, resulting in a poor user experience. Summary of the Invention
[0003] This invention provides an audio processing method that aims to offer an efficient and intelligent approach to improve user experience. By differentiating text or voice input for different business scenarios, selecting different timbres or audio files, and synthesizing the required audio file, the method enhances the emotional impact and relevance of the output audio, optimizes the limitations of single-timbre and tone, and further improves the user experience.
[0004] In a first aspect, embodiments of the present invention provide an audio processing method, the method comprising the following steps:
[0005] Acquire data to be processed, which may be text data or audio data;
[0006] The data to be processed is divided into sentences to obtain multiple sentences corresponding to the data to be processed;
[0007] The sentences to be processed are classified to obtain the scene types of the sentences to be processed, and different scene types correspond to different timbres;
[0008] Based on the scene type, obtain the sentence audio data corresponding to the sentence to be processed, wherein the timbre of the sentence audio data corresponds to the timbre of the scene type;
[0009] The audio data of the sentences to be processed are concatenated according to the order of the sentences in the data to be processed to obtain the target audio data corresponding to the data to be processed.
[0010] Optionally, dividing the data to be processed into sentences to obtain multiple sentences corresponding to the data to be processed includes:
[0011] The data to be processed is divided into sentences to obtain multiple initial sentences;
[0012] Determine the language corresponding to each of the initial sentences;
[0013] Among the multiple initial sentences, the initial sentence in Chinese is identified as the sentence to be processed, thus obtaining multiple sentences to be processed corresponding to the data to be processed.
[0014] Optionally, classifying the sentence to be processed to obtain the scenario type of the sentence to be processed includes:
[0015] Construct prompt word templates for classifying sentences by scenario;
[0016] The sentence to be processed is filled into the prompt word template to obtain the category prompt words for the sentence to be processed, and each sentence to be processed corresponds to one category prompt word;
[0017] The classification prompt words are input into a large language model, and the large language model outputs the scene type of the sentence to be processed.
[0018] Optionally, obtaining the sentence audio data corresponding to the sentence to be processed based on the scenario type includes:
[0019] In a preset audio database, matching is performed based on the scene type and the sentence to be processed. The audio file library includes audio data corresponding to different sentences under different scene types.
[0020] If the match is successful, the matched audio file will be identified as the sentence audio data corresponding to the sentence to be processed.
[0021] If the matching fails, then based on the scene type and the sentence to be processed, the corresponding sentence audio data is generated.
[0022] Optionally, generating sentence audio data corresponding to the sentence to be processed based on the scene type and the sentence to be processed includes:
[0023] Based on the scenario type, determine the target audio parameters corresponding to the sentence to be processed;
[0024] Based on the target audio parameters, the sentence to be processed is generated to obtain the sentence audio data corresponding to the sentence to be processed.
[0025] Optionally, the step of concatenating the audio data of the sentences to be processed according to their order in the data to be processed to obtain the target audio data corresponding to the data to be processed includes:
[0026] The order of each sentence to be processed in the data to be processed is determined as the concatenation order of each sentence to be processed;
[0027] The audio data of each sentence to be processed is post-processed to obtain audio data to be spliced, and each audio data to be spliced corresponds to one sentence to be processed.
[0028] The audio data to be spliced is spliced in the splicing order to obtain the target audio data corresponding to the data to be processed.
[0029] Optionally, the step of post-processing the audio data corresponding to each of the sentences to be processed to obtain the audio data to be concatenated includes:
[0030] The audio data corresponding to each of the sentences to be processed are subjected to format unification and parameter unification processing to obtain the audio data of the sentences with unified format and parameters as the audio to be spliced.
[0031] Secondly, embodiments of the present invention also provide an audio processing apparatus, the audio processing apparatus comprising:
[0032] The first acquisition module is used to acquire data to be processed, which is text data or audio data.
[0033] The segmentation module is used to segment the data to be processed into sentences, thereby obtaining multiple sentences to be processed corresponding to the data to be processed.
[0034] The classification module is used to classify the sentence to be processed to obtain the scene type of the sentence to be processed, and different scene types correspond to different timbres;
[0035] The second acquisition module is used to acquire the sentence audio data corresponding to the sentence to be processed based on the scene type, wherein the timbre of the sentence audio data corresponds to the timbre of the scene type;
[0036] The processing module is used to concatenate the audio data of the sentences to be processed in the order of the sentences in the data to be processed, so as to obtain the target audio data corresponding to the data to be processed.
[0037] Thirdly, embodiments of the present invention provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps in the audio processing method provided in embodiments of the present invention.
[0038] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the audio processing method provided in the embodiments of the present invention.
[0039] In this embodiment of the invention, data to be processed is acquired, which may be text data or audio data; the data to be processed is divided into sentences to obtain multiple sentences to be processed; the sentences to be processed are classified to obtain scene types, with different scene types corresponding to different timbres; based on the scene types, sentence audio data corresponding to the sentences to be processed is acquired, with the timbre of the sentence audio data corresponding to the timbre of the scene type; the sentence audio data is concatenated according to the order of the sentences to be processed in the data to obtain target audio data corresponding to the data to be processed. By dividing the data to be processed into multiple sentences to be processed, classifying the sentences to be processed to obtain scene types, acquiring sentence audio data corresponding to the sentences to be processed according to scene types, and concatenating the sentence audio data according to the order of the sentences to be processed in the data to obtain target audio data corresponding to the data to be processed, the emotional content and usage scenarios of the output audio are enhanced, the problem of single timbre tone is optimized, and the user experience is further improved. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a flowchart of an audio processing method provided in an embodiment of the present invention;
[0042] Figure 2 This is a structural diagram of the unified audio data format processing provided in the embodiments of the present invention;
[0043] Figure 3 This is a structural diagram of an audio processing method provided in an embodiment of the present invention;
[0044] Figure 4 This is a structural diagram of another audio processing method provided in an embodiment of the present invention;
[0045] Figure 5 This is a schematic diagram of the structure of an audio processing device provided in an embodiment of the present invention;
[0046] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0047] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] like Figure 1 As shown, Figure 1 This is a flowchart of an audio processing method provided by an embodiment of the present invention, which includes the following steps:
[0049] 101. Obtain the data to be processed.
[0050] In this embodiment of the invention, the audio processing method described above can be applied to a speech governance platform, and the audio processing can be constructed using a server or a distributed server. The speech governance platform includes interfaces for calling databases, data processing applications, knowledge databases, large oracle models, or large language models. Data to be processed can be obtained through data interfaces or through user uploads.
[0051] The data to be processed is either text data or audio data.
[0052] The text data mentioned above can be documents, articles, etc. that require dubbing or narration, such as poems, nursery rhymes, stories, articles, etc.; the audio data mentioned above can be recordings or voice files.
[0053] It should be noted that the above text data can be Chinese text data, English text data, or a mixture of Chinese and English text data; the above audio data can be Chinese audio files, English audio files, or a mixture of Chinese and English audio files.
[0054] In one possible embodiment, the data to be processed could be: "The weather is nice today. Shall we go to the park? Do you like playing basketball?"; or "Farmers toil under the scorching sun, their sweat drips onto the soil. Who knows that every grain in the bowl is the result of hard work?"; or "Xiao Ming went to the supermarket and bought some fruit."; or "Xiao Ming went to the supermarket and bought some fruit."
[0055] 102. Divide the data to be processed into sentences to obtain multiple sentences corresponding to the data to be processed.
[0056] In this embodiment of the invention, the data to be processed can be segmented into sentences to obtain multiple sentences corresponding to the data to be processed.
[0057] The above division can be understood as the process of splitting a piece of text or data into sentences, so that each segment is a complete sentence.
[0058] In one possible implementation, for example, the data to be processed is: "The weather is nice today. Shall we go to the park? Do you like playing basketball?" can be divided into three sentences:
[0059] "The weather is great today."
[0060] "Shall we go to the park?"
[0061] Do you like playing basketball?
[0062] For example, the data to be processed is: "Hoeing the fields at noon, sweat drips onto the soil. Who knows that every grain in the bowl is the result of hard work?" This can be divided into:
[0063] "Hoeing the fields at noon, sweat drips onto the soil beneath the crops."
[0064] "Who knows that every grain of rice on the plate is the result of hard work?"
[0065] For example, the data to be processed is: "Xiaoming went to the supermarket and bought some fruits." This can be divided into:
[0066] "Xiaoming went to the supermarket and bought some fruit."
[0067] "Some fruits."
[0068] It should be noted that when dividing the data to be processed into sentences, sentences in Chinese are identified as the multiple processing sentences corresponding to the data to be processed. Specifically, English-to-speech requires playback of the English audio, while Chinese-to-speech requires playback of the audio with poetic phrasing and emotional expression.
[0069] 103. Classify the sentences to be processed to obtain the scenario type of the sentences to be processed.
[0070] In this embodiment of the invention, different scene types correspond to different timbres. These scene types include poetry, nursery rhymes, idioms, stories, and reading. For example, poetry has rhythm and cadence, requiring a softer and more melodious timbre; nursery rhymes are designed for children, requiring a bright, cheerful, and easily understood timbre; idioms require a clear and powerful timbre; stories require a warm and friendly timbre; and reading is more formal or academic, requiring a serious and professional timbre.
[0071] Furthermore, keywords can be extracted from the sentence to be processed. These keywords are then input into the Large Language Model (LLM), which outputs the scene type corresponding to the sentence. The aforementioned Large Language Model (LLM) is a deep learning-based model primarily used for natural language tasks such as text generation, classification, summarization, and translation.
[0072] 104. Based on the scene type, obtain the sentence audio data corresponding to the sentence to be processed.
[0073] In this embodiment of the invention, the timbre of the sentence audio data corresponds to the timbre of the scene type.
[0074] Based on the scene type of the sentence to be processed, the corresponding scene type timbre can be obtained.
[0075] Specifically, when the scene type is poetry scene type, the sound of poetry scene type is obtained; when the scene type is children's song scene type, the sound of children's song scene type is obtained; when the scene type is idiom scene type, the sound of idiom scene type is obtained; when the scene type is story scene type, the sound of story scene type is obtained; when the scene type is reading scene type, the sound of reading scene type is obtained, and so on.
[0076] 105. Concatenate the audio data of the sentences in the order they appear in the data to be processed to obtain the target audio data corresponding to the data to be processed.
[0077] In this embodiment of the invention, multiple sentence audios can be concatenated in the order they appear in the data to be processed, thereby generating the target audio corresponding to the data to be processed.
[0078] In one possible implementation, for example, the sentences to be processed are "Hoeing the fields at noon, sweat drips onto the soil" and "Who knows that every grain in the bowl is the result of hard work?". The audio of these two sentences is obtained separately, and then they are combined into corresponding audio data in the order of "Hoeing the fields at noon, sweat drips onto the soil. Who knows that every grain in the bowl is the result of hard work?".
[0079] In this embodiment of the invention, the invention distinguishes different business scenarios for text or voice input, selects different timbres or audio files, synthesizes the required audio file, enhances the emotion and usage scenarios of the output audio, optimizes the problem of single timbre tone, and further improves the user experience.
[0080] In this embodiment of the invention, data to be processed is acquired, which may be text data or audio data. The data to be processed is divided into sentences to obtain multiple sentences to be processed. The sentences to be processed are classified to obtain scene types, with different scene types corresponding to different timbres. Based on the scene type, the audio data of the sentences to be processed is acquired, with the timbres of the sentence audio data corresponding to the scene type. The sentence audio data is then concatenated according to the order of the sentences in the data to be processed to obtain the target audio data corresponding to the data to be processed. By dividing the data to be processed into multiple sentences to be processed, classifying the sentences to be processed to obtain scene types, and acquiring the corresponding sentence audio data according to the scene types, and concatenating the sentence audio data according to the order of the sentences in the data to obtain the target audio data corresponding to the data to be processed, the emotional content and usage scenarios of the output audio are enhanced, the problem of single timbre tone is optimized, and the user experience is further improved.
[0081] It is understood that in the specific implementation of this application, data such as text data, voice data, and task data are involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required. Furthermore, the collection, use, and processing of related data, as well as the training, deployment, and invocation of large language models, must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0082] Optionally, in the step of dividing the data to be processed into sentences to obtain multiple sentences to be processed, the data to be processed can be divided into sentences to obtain multiple initial sentences; the language corresponding to each initial sentence can be determined; among the multiple initial sentences, the initial sentence with the language of Chinese can be determined as the sentence to be processed, thus obtaining multiple sentences to be processed corresponding to the data to be processed.
[0083] In this embodiment of the invention, the above-mentioned segmentation by sentence can be understood as dividing a piece of text or data according to sentence boundaries, so that each segmented part is a complete sentence. Sentence boundaries include periods, question marks, exclamation marks, etc.
[0084] The initial sentence mentioned above can be understood as a sentence obtained after division, which is a complete sentence in the original text.
[0085] The above-mentioned languages can be understood as the types of languages, including Chinese, English, etc.
[0086] Furthermore, sentences in Chinese can be selected from multiple initial sentences as the sentences to be processed.
[0087] It should be noted that Chinese poetry and prose are rich and diverse, and the vocal timbre needed to express emotions is also rich and diverse. For example, Chinese poetry requires vocal timbre to express the emotions conveyed through phrasing.
[0088] Optionally, in the step of classifying the sentence to be processed to obtain the scene type of the sentence, a prompt word template for classifying the sentence by scene can be constructed; the sentence to be processed is filled into the prompt word template to obtain the classification prompt words of the sentence to be processed; the classification prompt words are input into the large language model, and the scene type of the sentence to be processed is output by the large language model.
[0089] In this embodiment of the invention, each sentence to be processed corresponds to a category prompt word. The aforementioned scene types include poetry, nursery rhymes, idioms, stories, reading, and other scene types.
[0090] The aforementioned prompt word template includes keywords or phrases corresponding to scene types. This template guides the large language model in understanding and analyzing the input sentence and determining its scene type. The prompt word can be something like, "This sentence XXXX (the input sentence to be processed) is in which of the five categories: [poetry, children's song, idiom, story, reading]? Only the category word is needed to indicate the category." For example, if the sentence XXXX (the input sentence to be processed) is a poem, the answer is "poetry"; if the sentence XXXX (the input sentence to be processed) is a children's song, the answer is "children's song"; if the sentence XXXX (the input sentence to be processed) is an idiom, the answer is "idiom," and so on.
[0091] The aforementioned large language model is a deep learning-based model primarily used for natural language tasks such as text generation, classification, summarization, and translation.
[0092] In one possible implementation, for example, after filling the prompt word template with "Hoeing the fields at noon, sweat drips onto the soil," the prompt words for the poem are obtained. The prompt words for the poem are then input into the large language model, and the large language model outputs the poem scene type.
[0093] In another possible implementation, for example, after filling the prompt word template with "Two tigers, two tigers, running fast, running fast", the prompt words of the nursery rhyme are obtained, and the prompt words of the nursery rhyme are input into the large language model, and the large language model outputs the nursery rhyme scene type.
[0094] Optionally, in the step of obtaining the sentence audio data corresponding to the sentence to be processed based on the scene type, a matching can be performed in a preset audio database based on the scene type and the sentence to be processed; if the matching is successful, the matched audio file is determined as the sentence audio data corresponding to the sentence to be processed; if the matching fails, the sentence audio data corresponding to the sentence to be processed is generated based on the scene type and the sentence to be processed.
[0095] In this embodiment of the invention, the audio file library includes audio data corresponding to different sentences under different scene types.
[0096] The aforementioned preset audio database is used for storing, managing, and retrieving audio files. This audio database includes various types of audio files, such as audio data for different poems in a poetry context, audio data for different idioms in a idiom context, and audio data for different children's songs in a children's song context. The audio database provides search and retrieval functions.
[0097] In one possible implementation, for example, the scene type is poetry and the phrase "Hoeing the fields at noon, sweat drips onto the soil" is entered into a preset audio database for searching. If the corresponding audio file is found, it is determined to be the audio data of the sentence "Hoeing the fields at noon, sweat drips onto the soil".
[0098] In another possible implementation, for example, the scene type is poetry and the phrase "Hoeing the fields at noon, sweat drips onto the soil" is entered into a preset audio database for searching. If no corresponding audio file is found, it can be matched one by one in the priority order of [Poetry] > [Children's Songs] > [Idioms] > [Stories] > [Reading], and the reading type is used as a fallback.
[0099] It should be noted that if no scenario type of the sentence to be processed is matched, it can be matched one by one in the priority order of [Poetry] > [Children's Songs] > [Idioms] > [Stories] > [Reading], and the reading type can be used as a fallback.
[0100] Optionally, in the step of generating the sentence audio data corresponding to the sentence to be processed based on the scene type and the sentence to be processed, the target audio parameters corresponding to the sentence to be processed can be determined based on the scene type; based on the target audio parameters, the sentence to be processed is generated to obtain the sentence audio data corresponding to the sentence to be processed.
[0101] In this embodiment of the invention, the aforementioned scene types include poetry, nursery rhymes, idioms, stories, and reading. Different scene types correspond to different timbres. For example, poetry has rhythm and cadence, requiring a softer and more melodious timbre; nursery rhymes are designed for children and require a bright, cheerful, and easily understood timbre; idioms require a clear and powerful timbre; stories require a warm and friendly timbre; and reading is more formal or academic, requiring a serious and professional timbre.
[0102] The audio parameters mentioned above include the audio sampling rate, bit rate, number of channels, etc.
[0103] The above generation process can be understood as the process of converting text into audio data.
[0104] Optionally, in the step of concatenating the sentence audio data according to the order of the sentences to be processed in the data to be processed to obtain the target audio data corresponding to the data to be processed, the order of each sentence to be processed in the data to be processed can be determined as the concatenation order of each sentence to be processed; the sentence audio data corresponding to each sentence to be processed is post-processed to obtain the audio data to be concatenated; the audio data to be concatenated is concatenated according to the concatenation order to obtain the target audio data corresponding to the data to be processed.
[0105] In this embodiment of the invention, each audio data to be spliced corresponds to a sentence to be processed.
[0106] The order in which the sentences to be processed are concatenated can be determined based on the sentence order in the data to be processed.
[0107] The above post-processing can be understood as noise reduction, equalization, and other processing operations to improve the quality and consistency of audio data.
[0108] The above splicing can be understood as connecting the audio data of multiple sentences together in sequence to form a complete audio data processing process.
[0109] For example, the order of "Hoeing the fields at noon, sweat drips onto the soil" precedes the order of "Who knows that every grain in the bowl is the result of hard work"; therefore, the order of the audio data for "Hoeing the fields at noon, sweat drips onto the soil" also precedes the audio data for "Who knows that every grain in the bowl is the result of hard work".
[0110] Furthermore, the audio data of "Hoeing the fields at noon, sweat drips onto the soil" and "Who knows that every grain in the bowl is the result of hard work" are spliced together in the order of "Hoeing the fields at noon, sweat drips onto the soil" before "Who knows that every grain in the bowl is the result of hard work". This results in the audio data of "Hoeing the fields at noon, sweat drips onto the soil" being in the order of "Who knows that every grain in the bowl is the result of hard work".
[0111] Optionally, in the step of post-processing the audio data corresponding to each sentence to be processed to obtain the audio data to be spliced, the audio data corresponding to each sentence to be processed can be subjected to format unification processing and parameter unification processing to obtain the audio data with unified format and parameters as the audio to be spliced.
[0112] In this embodiment of the invention, the above-mentioned format unification processing can be understood as a process of converting audio data from different sources or in different formats into the same format. For example, converting all audio data into WAV format can ensure the compatibility and consistency of audio data.
[0113] The above parameter unification process can be understood as adjusting key parameters such as the sampling rate and bit rate of the audio data to a consistent level to ensure that the audio data of each sentence can be successfully combined during splicing.
[0114] It should be noted that standardizing the format and parameters of audio data can improve its smoothness, compression efficiency, and transmission efficiency. Commonly used audio formats include mp3, wav, pcm, and acc.
[0115] like Figure 2 As shown, Figure 2 This is a structural diagram of the unified audio data format processing provided in this embodiment of the invention. Specifically, step ①: convert different audio data into WAV audio data; step ②: merge; step ③: MP3 audio stream.
[0116] In this embodiment, the different audio data mentioned above include mp3, wav, pcm, acc, etc. Audio data with different formats and parameters can be uniformly converted into WAV audio data with unified parameters and a unified format using ffmpeg. Specifically, the parameters of the WAV audio data can be: a) encoder: libmp3lame, c) bitrate: 128000, d) number of channels: 1, e) sample rate: 16kbps, f) volume: different for different scenarios. ffmpeg is an open-source multimedia framework used to process audio, video, and other media files. ffmpeg provides a complete set of tools and libraries, supporting encoding, decoding, transcoding, streaming media transmission, recording, and playback of various audio and video formats. The Java service calls the capabilities of this FFmpeg system component through the ws.schild component.
[0117] Step ② above can be understood as merging the obtained audio data in a unified format (wav). This merging operation can be performed using the AudioSystem utility class provided by Java.
[0118] Step ③ above can be understood as using ffmpeg to convert the merged audio file from WAV to MP3 format. This MP3 format offers good compression with minimal loss of sound quality. Finally, the converted MP3 audio stream is sent to the front end.
[0119] In this embodiment, by using the MP3 audio format, the compression rate of audio data can be improved, thus increasing transmission efficiency.
[0120] like Figure 3 As shown, Figure 3This is a structural diagram of an audio processing method provided by an embodiment of the present invention. Specifically, it includes step ①: segmenting the input data to be processed by sentence; step ②: organizing scenes; step ③: classifying; and step ④: English.
[0121] In this embodiment, the data to be processed is text data or audio data.
[0122] Step ① above can be understood as dividing the input data to be processed into sentences, separating Chinese and English to obtain the Chinese sentences to be processed, and obtaining the English sentences in step ④. It should be noted that Chinese poetry and prose are rich and varied, requiring a wide range of emotional tones to express emotions; for example, Chinese poetry requires the emotional tone of poetic phrasing. English, on the other hand, only needs the tone of the English flow.
[0123] Step ② above can be understood as organizing model prompts for the large language model. These prompts can be phrased as, "Which of the five categories is this sentence XXXX (the input sentence to be processed) [poetry, children's song, idiom, story, reading]? Only the category word is needed to indicate the category." For example, if the sentence XXXX (the input sentence to be processed) is an idiom, then the answer is "idiom."
[0124] Step ③ above can be understood as executing a prompt request to obtain the scene type corresponding to the sentence to be processed. It should be noted that if no scene type is matched, it can be matched one by one in the priority order of [Poetry] > [Children's Songs] > [Idioms] > [Stories] > [Reading], and the reading type can be used as a fallback.
[0125] Step ④ above can be understood as obtaining English data, and English-to-speech conversion requires the timbre of the English process.
[0126] In this embodiment, the present invention divides the data to be processed into sentences, distinguishes different business scenarios, selects different timbres or audio files, and synthesizes the required audio files, thereby optimizing the problem of single timbre tone and solving the problem of poor user experience to a certain extent.
[0127] like Figure 4 As shown, Figure 4 This is a structural diagram of another audio processing method provided in an embodiment of the present invention. Specifically, it includes step ①: audio content governance platform; step ②: obtaining audio files; step ③: generating different audio streams for different scene types; step ④: generating audio streams; step ⑤: audio file collection.
[0128] In this embodiment, the aforementioned audio content governance platform can be understood as a management system for manually maintaining audio files corresponding to text and type.
[0129] Step ② above can be understood as querying audio data in a manually maintained management system that corresponds to the text and type of audio files, based on the input data to be processed and the corresponding scene type. Once the corresponding audio data is found, the corresponding audio data is obtained. The scene types mentioned above include poetry, children's songs, idioms, stories, reading, and other scene types.
[0130] Step ③ above can be understood as follows: when a match is not found in the manually maintained management system for text and corresponding audio files, the corresponding TTS can be invoked according to different scenario types to dynamically generate the corresponding audio stream file. TTS stands for "Text-to-Speech," which translates to "text-to-speech" or "speech synthesis." The goal of TTS is to convert text content into natural and fluent speech output, enabling computers to "speak."
[0131] Step ④ above can be understood as directly calling the English-language TTS to generate an audio stream when the data to be processed is in English. Specifically, English data requires English-language audio.
[0132] Step ⑤ above can be understood as collecting all audio streams in the order of the data to be processed to obtain the audio data corresponding to the data to be processed.
[0133] In this embodiment, the present invention differentiates the data to be processed into different business scenarios and languages, selects different timbres or audio files, and synthesizes the required audio file, thereby enhancing the emotion and usage scenarios of the output audio and optimizing the problem of single timbre tone.
[0134] like Figure 5 As shown, an embodiment of the present invention provides an audio processing device, which includes:
[0135] The first acquisition module 501 is used to acquire data to be processed, wherein the data to be processed is text data or audio data.
[0136] The segmentation module 502 is used to segment the data to be processed into sentences to obtain multiple sentences to be processed corresponding to the data to be processed.
[0137] The classification module 503 is used to classify the sentence to be processed to obtain the scene type of the sentence to be processed, and different scene types correspond to different timbres;
[0138] The second acquisition module 504 is used to acquire the sentence audio data corresponding to the sentence to be processed based on the scene type, wherein the timbre of the sentence audio data corresponds to the timbre of the scene type;
[0139] The processing module 505 is used to concatenate the audio data of the sentences to be processed in the order of the sentences in the data to be processed, so as to obtain the target audio data corresponding to the data to be processed.
[0140] Optionally, the segmentation module 502 is further configured to segment the data to be processed into sentences to obtain multiple initial sentences; determine the language corresponding to each initial sentence; and among the multiple initial sentences, determine the initial sentence whose language is Chinese as the sentence to be processed, thereby obtaining multiple sentences to be processed corresponding to the data to be processed.
[0141] Optionally, the classification module 503 is further configured to construct a prompt word template for classifying sentences by scene; fill the sentence to be processed into the prompt word template to obtain the classification prompt words for the sentence to be processed, with each sentence to be processed corresponding to one classification prompt word; input the classification prompt words into a large language model, and output the scene type of the sentence to be processed through the large language model.
[0142] Optionally, the second acquisition module 504 is further configured to perform matching in a preset audio database based on the scene type and the sentence to be processed, wherein the audio file database includes audio data corresponding to different sentences under different scene types; if the matching is successful, the matched audio file is determined as the sentence audio data corresponding to the sentence to be processed; if the matching fails, the sentence audio data corresponding to the sentence to be processed is generated based on the scene type and the sentence to be processed.
[0143] Optionally, the second acquisition module 504 is further configured to determine the target audio parameters corresponding to the sentence to be processed based on the scene type; and to perform generation processing on the sentence to be processed based on the target audio parameters to obtain the sentence audio data corresponding to the sentence to be processed.
[0144] Optionally, the processing module 505 is further configured to determine the order of each sentence to be processed in the data to be processed as the splicing order of each sentence to be processed; to perform post-processing on the audio data of each sentence to be processed to obtain audio data to be spliced, each audio data to be spliced corresponding to one sentence to be processed; and to splice the audio data to be spliced according to the splicing order to obtain the target audio data corresponding to the data to be processed.
[0145] Optionally, the processing module 505 is further configured to perform format unification processing and parameter unification processing on the sentence audio data corresponding to each sentence to be processed, so as to obtain the sentence audio data with unified format and parameters as the audio to be spliced.
[0146] like Figure 6As shown, this embodiment of the invention also provides an electronic device, including a processor, which can execute any of the above-described audio processing methods.
[0147] Specifically, it includes a processor 601 and a memory 602, as well as a computer program stored in the memory 602 and capable of running on the processor 601 to perform audio processing methods, wherein:
[0148] The processor 601 executes the calculator program containing the audio processing method stored in the memory 602, and performs the following steps:
[0149] Acquire data to be processed, which may be text data or audio data;
[0150] The data to be processed is divided into sentences to obtain multiple sentences corresponding to the data to be processed;
[0151] The sentences to be processed are classified to obtain the scene types of the sentences to be processed, and different scene types correspond to different timbres;
[0152] Based on the scene type, obtain the sentence audio data corresponding to the sentence to be processed, wherein the timbre of the sentence audio data corresponds to the timbre of the scene type;
[0153] The audio data of the sentences to be processed are concatenated according to the order of the sentences in the data to be processed to obtain the target audio data corresponding to the data to be processed.
[0154] Optionally, the step of processor 601 dividing the data to be processed into sentences to obtain multiple sentences to be processed corresponding to the data to be processed includes:
[0155] The data to be processed is divided into sentences to obtain multiple initial sentences;
[0156] Determine the language corresponding to each of the initial sentences;
[0157] Among the multiple initial sentences, the initial sentence in Chinese is identified as the sentence to be processed, thus obtaining multiple sentences to be processed corresponding to the data to be processed.
[0158] Optionally, the processor 601 performs the classification of the sentence to be processed to obtain the scene type of the sentence to be processed, including:
[0159] Construct prompt word templates for classifying sentences by scenario;
[0160] The sentence to be processed is filled into the prompt word template to obtain the category prompt words for the sentence to be processed, and each sentence to be processed corresponds to one category prompt word;
[0161] The classification prompt words are input into a large language model, and the large language model outputs the scene type of the sentence to be processed.
[0162] Optionally, the step of processor 601 executing the step of obtaining the sentence audio data corresponding to the sentence to be processed based on the scene type includes:
[0163] In a preset audio database, matching is performed based on the scene type and the sentence to be processed. The audio file library includes audio data corresponding to different sentences under different scene types.
[0164] If the match is successful, the matched audio file will be identified as the sentence audio data corresponding to the sentence to be processed.
[0165] If the matching fails, then based on the scene type and the sentence to be processed, the corresponding sentence audio data is generated.
[0166] Optionally, the step of generating sentence audio data corresponding to the sentence to be processed based on the scene type and the sentence to be processed, executed by the processor 601, includes:
[0167] Based on the scenario type, determine the target audio parameters corresponding to the sentence to be processed;
[0168] Based on the target audio parameters, the sentence to be processed is generated to obtain the sentence audio data corresponding to the sentence to be processed.
[0169] Optionally, the process executed by processor 601 to concatenate the audio data of the sentences to be processed in the order they appear in the data to be processed, to obtain the target audio data corresponding to the data to be processed, includes:
[0170] The order of each sentence to be processed in the data to be processed is determined as the concatenation order of each sentence to be processed;
[0171] The audio data of each sentence to be processed is post-processed to obtain audio data to be spliced, and each audio data to be spliced corresponds to one sentence to be processed.
[0172] The audio data to be spliced is spliced in the splicing order to obtain the target audio data corresponding to the data to be processed.
[0173] Optionally, the post-processing of the sentence audio data corresponding to each of the sentences to be processed, performed by the processor 601 to obtain the audio data to be concatenated, includes:
[0174] The audio data corresponding to each of the sentences to be processed are subjected to format unification and parameter unification processing to obtain the audio data of the sentences with unified format and parameters as the audio to be spliced.
[0175] This invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the audio processing method or application-side audio processing method provided in this invention, and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0176] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0177] The above description discloses only preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.
Claims
1. An audio processing method, characterized in that, The method includes the following steps: Acquire data to be processed, which may be text data or audio data; The data to be processed is divided into sentences to obtain multiple sentences corresponding to the data to be processed; The sentences to be processed are classified to obtain the scene types of the sentences to be processed, and different scene types correspond to different timbres; Based on the scene type, obtain the sentence audio data corresponding to the sentence to be processed, wherein the timbre of the sentence audio data corresponds to the timbre of the scene type; The audio data of the sentences to be processed are concatenated according to the order of the sentences in the data to be processed to obtain the target audio data corresponding to the data to be processed.
2. The audio processing method as described in claim 1, characterized in that, The step of dividing the data to be processed into sentences to obtain multiple sentences corresponding to the data to be processed includes: The data to be processed is divided into sentences to obtain multiple initial sentences; Determine the language corresponding to each of the initial sentences; Among the multiple initial sentences, the initial sentence in Chinese is identified as the sentence to be processed, thus obtaining multiple sentences to be processed corresponding to the data to be processed.
3. The audio processing method as described in claim 1, characterized in that, The process of classifying the sentences to be processed to obtain the scenario types of the sentences to be processed includes: Construct prompt word templates for classifying sentences by scenario; The sentence to be processed is filled into the prompt word template to obtain the category prompt words for the sentence to be processed, and each sentence to be processed corresponds to one category prompt word; The classification prompt words are input into a large language model, and the large language model outputs the scene type of the sentence to be processed.
4. The audio processing method as described in claim 3, characterized in that, The step of obtaining the sentence audio data corresponding to the sentence to be processed based on the scenario type includes: In a preset audio database, matching is performed based on the scene type and the sentence to be processed. The audio file library includes audio data corresponding to different sentences under different scene types. If the match is successful, the matched audio file will be identified as the sentence audio data corresponding to the sentence to be processed. If the matching fails, then based on the scene type and the sentence to be processed, the corresponding sentence audio data is generated.
5. The audio processing method as described in claim 4, characterized in that, The step of generating sentence audio data corresponding to the sentence to be processed based on the scene type and the sentence to be processed includes: Based on the scenario type, determine the target audio parameters corresponding to the sentence to be processed; Based on the target audio parameters, the sentence to be processed is generated to obtain the sentence audio data corresponding to the sentence to be processed.
6. The audio processing method according to any one of claims 1 to 5, characterized in that, The step of concatenating the audio data of the sentences to be processed according to their order in the data to be processed to obtain the target audio data corresponding to the data to be processed includes: The order of each sentence to be processed in the data to be processed is determined as the concatenation order of each sentence to be processed; The audio data of each sentence to be processed is post-processed to obtain audio data to be spliced, and each audio data to be spliced corresponds to one sentence to be processed. The audio data to be spliced is spliced in the splicing order to obtain the target audio data corresponding to the data to be processed.
7. The audio processing method as described in claim 6, characterized in that, The step of post-processing the audio data corresponding to each of the sentences to be processed to obtain the audio data to be concatenated includes: The audio data corresponding to each of the sentences to be processed are subjected to format unification and parameter unification processing to obtain the audio data of the sentences with unified format and parameters as the audio to be spliced.
8. An audio processing apparatus, characterized in that, The audio processing device includes: The first acquisition module is used to acquire data to be processed, which is text data or audio data. The segmentation module is used to segment the data to be processed into sentences, thereby obtaining multiple sentences to be processed corresponding to the data to be processed. The classification module is used to classify the sentence to be processed to obtain the scene type of the sentence to be processed, and different scene types correspond to different timbres; The second acquisition module is used to acquire the sentence audio data corresponding to the sentence to be processed based on the scene type, wherein the timbre of the sentence audio data corresponds to the timbre of the scene type; The processing module is used to concatenate the audio data of the sentences to be processed in the order of the sentences in the data to be processed, so as to obtain the target audio data corresponding to the data to be processed.
9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the audio processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the audio processing method as described in any one of claims 1 to 7.