Method and system for generating expressive audio based on textual data
By splitting the text and synthesizing sound effect tags, the problem of insufficient auditory experience in text-to-audio technology is solved, and more emotional and immersive audio generation is achieved.
Patent Information
- Application Number
- CN202511575821.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-10-31
Smart Images

Figure CN121034284B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing technology, and specifically to a method and system for generating expressive audio based on text data. Background Technology
[0002] With the rapid development of artificial intelligence and speech synthesis technology, text-to-speech applications have been widely used in various fields, especially in radio broadcasters or radio programs, where the automatic generation of speech based on text by artificial intelligence can greatly improve audio production efficiency.
[0003] Current text-to-audio technologies primarily focus on the accuracy of the resulting audio, neglecting the emotional and immersive listening experience, leading to a harsh listening experience for the audience. Summary of the Invention
[0004] The main objective of this invention is to provide a method and system for generating expressive audio based on text data, aiming to solve the problem that current text-to-audio technologies neglect the emotional and immersive experience of the auditory experience, resulting in a stiff listening experience for the audience.
[0005] The technical solution proposed in this invention is as follows:
[0006] A method for generating expressive audio based on text data, applied to a system for generating expressive audio based on text data; the system includes a management terminal and a server that are communicatively connected to each other; the method includes:
[0007] The management terminal acquires the text to be converted and sends it to the server;
[0008] The server splits the text to be converted into multiple corresponding text segments and determines the sound effect tags corresponding to each text segment.
[0009] The server obtains the audio of the sound effect corresponding to each sound effect tag;
[0010] The server performs speech conversion on the text to be converted to generate the first target audio.
[0011] The server determines the insertion time of each sound effect audio in the first target audio, and synthesizes each sound effect audio and the first target audio based on the insertion time of the sound effect audio in the first target audio to obtain the second target audio.
[0012] Preferably, the server runs a TTS model; the server stores a sound effect material library; the server splits the text to be converted into multiple corresponding text segments, and determines the sound effect tags corresponding to each text segment, including:
[0013] The server uses predetermined punctuation marks in the text to be converted as delimiters to divide the text into multiple text segments, wherein the predetermined punctuation marks include commas, periods, and semicolons.
[0014] The server performs speech-to-speech conversion on the text to be converted to generate the first target audio, including:
[0015] The server uses the TTS model to perform speech conversion on the text to be converted to generate the first target audio.
[0016] The server obtains the audio of the sound effect corresponding to each sound effect tag, including:
[0017] The server obtains the audio of each audio effect tag based on the audio effect material library.
[0018] Preferably, the server further includes splitting the text to be converted into multiple corresponding text segments and determining the sound effect tags corresponding to each text segment:
[0019] The server constructs a large language model based on a neural network and obtains a training sample set, which includes multiple training texts and sound effect tags corresponding to each training text. The sound effect tags are sound effects that correspond to the semantics of the training text. When the training text is descriptive text, the corresponding sound effect tag is none.
[0020] The server uses the training text in the training sample set as input parameters and the corresponding sound effect tags of the training text as output parameters to train the large language model based on neural networks.
[0021] The server inputs each text segment into the trained neural network-based large language model to obtain the output sound effect tags, and uses the output sound effect tags as the corresponding sound effect tags for each text segment.
[0022] Preferably, the server stores a sound effect matching word library, which includes multiple sound effect tags and matching words corresponding to each sound effect tag; the server determines the insertion time of each sound effect audio in the first target audio, and synthesizes each sound effect audio and the first target audio based on the insertion time of the sound effect audio in the first target audio to obtain the second target audio, including:
[0023] The server obtains the sound effect tags corresponding to the text segment and marks them as target tags;
[0024] The server determines whether the text segment contains matching words corresponding to the target tag based on the sound effect matching word library;
[0025] If so, the server marks the matching words in the text segment as target words, and marks the time period in which the target words in the text segment appear in the first target audio as the first target time period;
[0026] The server marks the end time of the first target time period as the first start time, and marks the end time of the time period in which the last character in the text segment appears in the first target audio as the first end time.
[0027] The server determines the time period between the first start time and the first end time as the insertion time period of the sound effect audio corresponding to the target label in the first target audio.
[0028] Preferably, the server determines whether the text segment contains matching words corresponding to the target tag based on the sound effect matching word library, and then further includes:
[0029] If not, the server will mark the text segment as a custom text segment and send the custom text segment to the management terminal;
[0030] The management terminal obtains manually input labeled words corresponding to the custom text segment and sends the labeled words to the server;
[0031] The server marks the time period in which the tagged words of the custom text segment appear in the first target audio as the second target time period;
[0032] The server marks the end time of the second target time period as the second start time, and marks the end time of the time period in which the last character in the custom text segment appears in the first target audio as the second end time.
[0033] The server determines the time period between the second start time and the second end time as the insertion period of the sound effect audio corresponding to the target label in the first target audio.
[0034] Preferably, the server marks the end time of the first target time period as the first start time, and marks the end time of the time period in which the last character in the text segment appears in the first target audio as the first end time, and then further includes:
[0035] The server determines whether the duration between the first starting time and the first ending time is less than a first preset duration.
[0036] If so, the server marks the time elapsed after the first starting time for a first preset duration as the adjusted ending time;
[0037] The server determines the time period between the first starting time and the adjusted ending time as the insertion time of the sound effect audio corresponding to the target label in the first target audio.
[0038] Preferably, the system further includes a user terminal communicatively connected to the server; the user terminal includes an in-ear player; the in-ear player includes a speaker, a first microphone, and a first vibration sensor; the method further includes:
[0039] The server sends the second target audio to the user terminal;
[0040] The user terminal controls the speaker to play the second target audio;
[0041] During the playback of the second target audio by the speaker, the user terminal acquires the first ambient sound signal collected in real time by the first microphone, and the first vibration signal of the user's head collected in real time by the first vibration sensor.
[0042] The user terminal determines whether the first condition is met, wherein the first condition is: the average amplitude of the first vibration signal within the past second preset time period is greater than the first preset amplitude, and the average intensity of the first sound signal within the past second preset time period is greater than the first preset intensity.
[0043] If the first condition is met, the user terminal marks the current volume of the second target audio being played by the speaker as the original volume and generates a volume reduction command;
[0044] When a volume reduction command is present, the user terminal controls the speaker to reduce the volume of the second target audio being played based on the volume reduction command;
[0045] When a volume reduction command is present, the user terminal determines whether a second condition is met, wherein the second condition is: the average amplitude of the first vibration signal within the past second preset time period is less than the second preset amplitude, and the average intensity of the first sound signal within the past second preset time period is less than the second preset intensity, the second preset amplitude is less than the first preset amplitude, and the second preset intensity is less than the first preset intensity.
[0046] If the second condition is met, the user terminal generates a recovery command and deletes the volume reduction command;
[0047] When a recovery command is available, the user terminal controls the speaker to restore the volume of the second target audio being played to the original volume based on the recovery command.
[0048] Preferably, the user terminal body is further provided with a second vibration sensor and a second microphone; the method further includes:
[0049] The user terminal acquires the second vibration signal collected in real time by the second vibration sensor;
[0050] The user terminal acquires the second sound signal collected in real time by the second microphone;
[0051] The user terminal determines whether the first condition is met, and then the process further includes:
[0052] If the first condition is met, the user terminal marks the generation time of the volume reduction instruction as the target time, and determines a third target time period based on the target time, wherein the middle time of the third target time period is the target time, and the duration of the third target time period is twice the second preset duration.
[0053] The user terminal marks the second sound signal within the third target time period as the terminal sound signal, and the first sound signal within the third target time period as the headphone sound signal;
[0054] The user terminal obtains the ratio of the average intensity value of the terminal sound signal to the average intensity value of the headphone sound signal, and marks it as the intensity ratio.
[0055] The user terminal marks the second vibration signal within the third target time period as the terminal vibration signal, and the first vibration signal within the third target time period as the earphone vibration signal;
[0056] The user terminal acquires the amplitude values of the terminal vibration signal and the headphone vibration signal at each sampling time, and calculates the amplitude ratio of the terminal vibration signal and the headphone vibration signal at different sampling times:
[0057] The user terminal calculates the variance of the amplitude ratio of the terminal vibration signal and the headphone vibration signal at different sampling times;
[0058] If the intensity ratio is less than a first preset ratio and greater than a second preset ratio, and the variance is less than a preset variance, the user terminal generates a normal command and deletes the tone reduction command.
[0059] When a normal instruction is present, the user terminal will no longer generate a volume reduction instruction and will control the speaker to restore the volume of the second target audio being played to the original volume.
[0060] This invention also proposes a system for generating expressive audio based on text data, and applies a method for generating expressive audio based on text data; the system includes a management terminal and a server that are communicatively connected to each other.
[0061] The above technical solution can achieve the following beneficial effects:
[0062] The proposed method for generating expressive audio based on text data enhances the emotional and immersive listening experience, avoiding the problem of a harsh listening experience. First, the audio to be converted is acquired. Then, the text to be converted is split into multiple text segments, and the corresponding sound effect tags for each text segment are determined. Next, the sound effect audio corresponding to each sound effect tag is acquired, and the text to be converted is converted into speech to generate a first target audio. Then, the insertion time of each sound effect audio in the first target audio is determined, and based on the insertion time of the sound effect audio in the first target audio, the sound effect audio and the first target audio are synthesized to obtain a second target audio. This method combines text-to-speech technology and sound effect synthesis technology, resulting in a more immersive and richer second target audio, thus enhancing the user's listening experience. Attached Figure Description
[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0064] Figure 1 This is a flowchart illustrating the first embodiment of a method for generating expressive audio based on text data proposed in this invention. Detailed Implementation
[0065] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0066] This invention proposes a method and system for generating expressive audio based on text data.
[0067] As attached Figure 1 As shown, in the first embodiment of the method for generating expressive audio based on text data proposed in this invention, this method is applied to a system for generating expressive audio based on text data; the system includes a management terminal and a server that are communicatively connected to each other; this embodiment includes the following steps:
[0068] Step S110: The management terminal obtains the text to be converted and sends it to the server.
[0069] Specifically, the audio generation method provided in this disclosure is applied to scenarios where audio data is generated based on text to be converted. For example, it can intelligently generate audio programs based on text information and broadcast them; such as automatic audiobook reading, story reading, novel reading, news broadcasting, and other scenarios.
[0070] The management terminal here is a smart terminal operated by administrators, such as a personal computer; the server here can be a cloud server or a local server. The management terminal establishes a connection with the server through the Internet. Administrators can directly input the text to be converted through the input device of the management terminal, or receive the text to be converted from other terminals on the Internet; the management terminal has computing processing capabilities; the server is responsible for processing audio operations, including text-to-speech, text analysis, sound effect generation, and audio synthesis.
[0071] Step S120: The server splits the text to be converted into multiple corresponding text segments and determines the sound effect tags corresponding to each text segment.
[0072] Specifically, since the text to be converted generally contains multiple sentences, it needs to be segmented to obtain multiple text segments, so as to determine the corresponding sound effect tags for each text segment. The sound effect category tags corresponding to the text segments are related to the semantic information of the text segments. For example, when the text segment contains onomatopoeic words (such as boom, ding-ding) or specific images that produce sounds (such as thunder, rain, car driving), it will correspond to a specific type of sound effect tag. When the text segment is a general descriptive statement, such as "the price of the top is twice the price of the pants", then the corresponding sound effect tag for the text segment is none.
[0073] Specifically, when the audio effect tag corresponding to a text segment is "news interview", the original audio corresponding to that text segment is obtained (the original audio here is the actual voice of the person involved collected by the reporter during the news interview).
[0074] Step S130: The server obtains the audio of the sound effect corresponding to each sound effect tag.
[0075] Specifically, when the sound effect tag corresponding to a text segment is "rain", the corresponding sound effect audio is the sound of rain.
[0076] Step S140: The server performs speech conversion on the text to be converted to generate the first target audio.
[0077] Step S150: The server determines the insertion time of each sound effect audio in the first target audio, and synthesizes each sound effect audio and the first target audio based on the insertion time of the sound effect audio in the first target audio to obtain the second target audio.
[0078] For example, when the sound effect tag corresponding to a text segment is "news interview", the original audio corresponding to that text segment is marked as the third target audio, and the third target audio and the first target audio corresponding to the text segment with the sound effect tag "news interview" are synthesized to obtain the second target audio, which can enhance the audience's sense of presence.
[0079] Specifically, the sound effects audio should be inserted at an appropriate position within the primary target audio to enhance the listener's listening experience.
[0080] The proposed method for generating expressive audio based on text data enhances the emotional and immersive listening experience, avoiding the problem of a harsh listening experience. First, the audio to be converted is acquired. Then, the text to be converted is split into multiple text segments, and the corresponding sound effect tags for each text segment are determined. Next, the sound effect audio corresponding to each sound effect tag is acquired, and the text to be converted is converted into speech to generate a first target audio. Then, the insertion time of each sound effect audio in the first target audio is determined, and based on the insertion time of the sound effect audio in the first target audio, the sound effect audio and the first target audio are synthesized to obtain a second target audio. This method combines text-to-speech technology and sound effect synthesis technology, resulting in a more immersive and richer second target audio, thus enhancing the user's listening experience.
[0081] In a second embodiment of the method for generating expressive audio based on text data proposed in this invention, based on the first embodiment, the server runs a TTS model; the server stores a sound effects material library; step S120 includes the following steps:
[0082] Step S210: The server uses predetermined punctuation marks in the text to be converted as delimiters to divide the text to be converted into multiple text segments, wherein the predetermined punctuation marks include commas, periods, and semicolons.
[0083] Specifically, the text to be converted is segmented using punctuation marks to obtain multiple text segments.
[0084] Step S140 includes the following steps:
[0085] Step S220: The server uses the TTS model to perform speech conversion on the text to be converted to generate the first target audio.
[0086] Specifically, the TTS (Text-to-Speech) model is a large model that converts text into natural speech and is widely used in fields such as intelligent assistants, audiobooks, and navigation broadcasts.
[0087] Step S130 includes the following steps:
[0088] Step S230: The server obtains the audio of each audio tag based on the audio material library.
[0089] Specifically, the sound effects library includes a variety of sound effects audio.
[0090] In a third embodiment of the method for generating expressive audio based on text data proposed in this invention, based on the first embodiment, step S120 further includes the following steps:
[0091] Step S310: The server constructs a large language model based on a neural network and obtains a training sample set, wherein the training sample set includes multiple training texts and sound effect tags corresponding to each training text. The sound effect tags are sound effects that correspond to the semantics of the training text. When the training text is descriptive text, the corresponding sound effect tag is none.
[0092] Specifically, a Large Language Model (LLM) is a deep learning model trained on a large amount of text data that can generate natural language text or understand the meaning of language text. Large Language Models can handle various natural language tasks, such as text classification, question answering, and dialogue, and are an important pathway to artificial intelligence.
[0093] Step S320: The server uses the training text in the training sample set as input parameters and the corresponding sound effect tags of the training text as output parameters to train the large language model based on the neural network.
[0094] Step S330: The server inputs each text segment into the trained neural network-based large language model to obtain the output sound effect tags, and uses the output sound effect tags as the corresponding sound effect tags for each text segment.
[0095] Specifically, this embodiment uses a large language model based on neural networks to obtain the sound effect tags corresponding to each text segment; for example, when the text segment is "the wind blows across the grass", the corresponding sound effect tag is "the sound of wind".
[0096] In the fourth embodiment of the method for generating expressive audio based on text data proposed in this invention, based on the first embodiment, the server stores a sound effect matching word library, which includes multiple sound effect tags and matching words corresponding to each sound effect tag; step S150 includes the following steps:
[0097] Step S410: The server obtains the sound effect tags corresponding to the text segment and marks them as target tags.
[0098] Step S420: The server determines whether the text segment contains matching words corresponding to the target tag based on the sound effect matching word library.
[0099] Specifically, the matching words here are the words for which the corresponding sound effect audio needs to be played. When a matching word appears in the text segment, the corresponding sound effect audio is played. The matching word is generally a verb, which is the verb that appears when the corresponding sound effect begins. For example, if the text segment is "A gentle breeze suddenly blew across the grass", and the corresponding sound effect tag is "wind sound", then the word "blow" is the matching word for the sound effect tag. When the first target audio reads the word "blow" in the text segment, the sound effect audio "wind sound" is inserted as background sound into the first target audio.
[0100] In addition, each sound effect tag has at least one matching word, and usually multiple matching words; for example, when the sound effect tag is "wind sound", the corresponding matching words include, but are not limited to: blow, sound, scrape, rise, brush, sweep, roll, float.
[0101] If so, proceed to step S430: The server marks the matching words in the text segment as target words, and marks the time period in which the target words in the text segment appear in the first target audio as the first target time period.
[0102] Specifically, each word needs to be read aloud for a certain amount of time, so the time period in which the target word appears in the first target audio is marked as the first target time period.
[0103] Step S440: The server marks the end time of the first target time period as the first start time, and marks the end time of the time period in which the last character in the text segment appears in the first target audio as the first end time.
[0104] Specifically, the insertion period for the sound effect audio corresponding to the target tag is from the position of the target word in the text segment until the end of the text segment.
[0105] Step S450: The server determines the time period between the first start time and the first end time as the insertion time period of the sound effect audio corresponding to the target label in the first target audio.
[0106] In the fifth embodiment of the method for generating expressive audio based on text data proposed in this invention, based on the fourth embodiment, after step S420, the following steps are further included:
[0107] If not, proceed to step S510: The server marks the text segment as a custom text segment and sends the custom text segment to the management terminal.
[0108] Specifically, if not, it means that there are no matching words corresponding to the sound effect tags in the text segment, but the audio text segment does have corresponding sound effect tags. Therefore, it is necessary to manually determine the insertion time of the sound effect audio in the text segment. Thus, the text segment is marked as a custom text segment and sent to the management terminal.
[0109] Step S520: The management terminal obtains the manually input labeled words corresponding to the custom text segment and sends the labeled words to the server.
[0110] Specifically, the manually marked words here indicate the starting position where the sound effect audio should be inserted.
[0111] Step S530: The server marks the time period in which the tagged words of the custom text segment appear in the first target audio as the second target time period.
[0112] Step S540: The server marks the end time of the second target time period as the second start time, and marks the end time of the time period in which the last character in the custom text segment appears in the first target audio as the second end time.
[0113] Specifically, the insertion period for the audio effect corresponding to the target tag is from the position of the marked word in the text segment until the end of the text segment.
[0114] Step S550: The server determines the time period between the second start time and the second end time as the insertion time period of the sound effect audio corresponding to the target label in the first target audio.
[0115] Specifically, this embodiment provides a technical solution for manually determining the insertion time of sound effects audio in text segments where no matching words appear.
[0116] In the sixth embodiment of the method for generating expressive audio based on text data proposed in this invention, based on the fourth embodiment, after step S440, the following steps are further included:
[0117] Step S610: The server determines whether the duration between the first starting time and the first ending time is less than the first preset duration (e.g., 1 second).
[0118] Specifically, the first preset duration here is the appropriate minimum duration of the sound effect audio, which is generally set to 1 second. That is, the duration of the sound effect audio should not be too short to avoid affecting the listener's listening experience.
[0119] If so, proceed to step S620: The server marks the time elapsed after the first starting time as the adjusted ending time.
[0120] Specifically, if the duration between the first starting point and the first ending point is less than the first preset duration, it means that the duration of the sound effect audio is too short. In order to ensure the listening experience of the audience, the duration of the sound effect audio needs to be extended, that is, the duration is adjusted to the first preset duration.
[0121] Step S630: The server determines the time period between the first starting time and the adjusted ending time as the insertion time period of the sound effect audio corresponding to the target label in the first target audio.
[0122] Specifically, because the duration of the sound effect audio has been adjusted, it is necessary to redetermine the insertion time of the sound effect audio corresponding to the target tag in the first target audio.
[0123] In a seventh embodiment of a method for generating expressive audio based on text data proposed in this invention, based on the first embodiment, the system further includes a user terminal (e.g., a smartphone terminal) communicatively connected to the server; the user terminal includes an in-ear player (e.g., earbuds communicatively connected to the user terminal via Bluetooth); the in-ear player includes a speaker, a first microphone, and a first vibration sensor; specifically, the first microphone and the first vibration sensor are disposed at the end of the in-ear player furthest from the speaker to reduce sound interference from the speaker to the first microphone and the first vibration sensor; this embodiment also includes the following steps:
[0124] Step S701: The server sends the second target audio to the user terminal.
[0125] Step S702: The user terminal controls the speaker to play the second target audio.
[0126] Step S703: During the playback of the second target audio by the speaker, the user terminal acquires the first ambient sound signal collected in real time by the first microphone, and the first vibration signal of the user's head collected in real time by the first vibration sensor.
[0127] Step S704: The user terminal determines whether the first condition is met, wherein the first condition is: the average amplitude of the first vibration signal within the past second preset time period (e.g., 1 second) is greater than the first preset amplitude, and the average intensity of the first sound signal within the past second preset time period is greater than the first preset intensity value.
[0128] Specifically, when a user speaks, their vocal cords vibrate, so the first vibration sensor can detect a significant vibration signal, and the first sound sensor can detect a significant sound signal; therefore, the first preset amplitude is set to the standard amplitude of the vibration signal that the first vibration sensor can detect when a human is speaking normally; the first preset intensity value is set to the standard intensity value of the sound signal that the first sound sensor can detect when a human is speaking normally; when the first condition is met, it indicates that the user is currently speaking.
[0129] Step S705: If the first condition is met, the user terminal marks the current volume of the second target audio being played by the speaker as the original volume and generates a volume reduction command.
[0130] Step S706: When a volume reduction command is present, the user terminal controls the speaker to reduce the volume of the second target audio being played based on the volume reduction command.
[0131] Specifically, when a user is currently speaking, it is assumed that the user is likely talking to other people around them. Therefore, the user terminal sends a volume reduction command and controls the speaker to reduce the volume of the second target audio being played, thereby reducing the volume of the second target audio and avoiding interference with the user's conversation.
[0132] Step S707: When a volume reduction command is present, the user terminal determines whether the second condition is met, wherein the second condition is: the average amplitude of the first vibration signal within the past second preset time period is less than the second preset amplitude, and the average intensity value of the first sound signal within the past second preset time period is less than the second preset intensity value, the second preset amplitude is less than the first preset amplitude, and the second preset intensity value is less than the first preset intensity value.
[0133] Specifically, the second preset amplitude is less than the first preset amplitude, and the second preset intensity value is less than the first preset intensity value. The second preset intensity value is set as the standard intensity value of the sound signal that the first sound sensor can detect when a human is talking normally and other people around him are talking, and only other people are talking. Therefore, when the second condition is met, it means that the user has stopped talking and there are no other people talking around him. Therefore, it is presumed that the user has stopped talking, and the playback volume of the second target audio can be restored.
[0134] Step S708: If the second condition is met, the user terminal generates a recovery command and deletes the tone reduction command.
[0135] Step S709: When a recovery command is present, the user terminal controls the speaker to restore the volume of the second target audio being played to the original volume based on the recovery command.
[0136] In the eighth embodiment of the method for generating expressive audio based on text data proposed in this invention, based on the seventh embodiment, the user terminal body is further provided with a second vibration sensor and a second microphone; this embodiment also includes the following steps:
[0137] Step S801: The user terminal acquires the second vibration signal collected in real time by the second vibration sensor.
[0138] Step S802: The user terminal acquires the second sound signal collected in real time by the second microphone.
[0139] Step S704, followed by the following steps:
[0140] Step S803: If the first condition is met, the user terminal marks the generation time of the volume reduction instruction as the target time, and determines a third target time period based on the target time, wherein the middle time of the third target time period is the target time, and the duration of the third target time period is twice the second preset duration.
[0141] Specifically, it can be seen that the duration of the third target time period determined in the above steps is twice the second preset duration. The longer duration makes it easier to more accurately determine whether the user is currently in an overall vibration environment (such as continuous external environmental vibration caused by taking a train, bus or other means of transportation) rather than the user's own speech, thus avoiding misjudgment and reducing the playback volume of the second target audio.
[0142] Step S804: The user terminal marks the second sound signal within the third target time period as the terminal sound signal, and marks the first sound signal within the third target time period as the headphone sound signal.
[0143] Specifically, when a user uses a user terminal (smartphone) to listen to a second target audio, the user terminal is farther away from the user's head than an in-ear player (e.g., held in the hand or placed in another position); therefore, the terminal's sound signal focuses more on reflecting the sound of the entire surrounding environment, while the headphone's sound signal focuses more on reflecting the sound of the surrounding environment near the user's head.
[0144] Step S805: The user terminal obtains the ratio of the average intensity value of the terminal sound signal to the average intensity value of the headphone sound signal, and marks it as the intensity ratio.
[0145] Specifically, when the user is speaking normally, the average intensity value of the terminal's sound signal should be less than the average intensity value of the headphone's sound signal (e.g., intensity ratio less than 0.8); when the user is not speaking but is in a relatively noisy external environment, the average intensity value of the terminal's sound signal should be close to the average intensity value of the headphone's sound signal (e.g., intensity ratio less than 1.1 and greater than 0.9).
[0146] Step S806: The user terminal marks the second vibration signal within the third target time period as the terminal vibration signal, and marks the first vibration signal within the third target time period as the earphone vibration signal.
[0147] Specifically, under normal circumstances, when a user uses a user terminal (smartphone) to listen to a second target audio, the user terminal is farther away from the user's head compared to an in-ear player; therefore, the terminal vibration signal focuses more on reflecting the vibration of the entire surrounding environment, while the headphone sound signal focuses more on reflecting the vibration of the surrounding environment near the user's head.
[0148] Step S807: The user terminal acquires the amplitude value of the terminal vibration signal at each sampling time (for example, in this embodiment, the sampling period of the vibration signal is 0.01 seconds, so the interval between adjacent sampling times is 0.01 seconds), and the amplitude value of the headphone vibration signal at each sampling time, and calculates the amplitude ratio of the terminal vibration signal and the headphone vibration signal at different sampling times:
[0149] ,
[0150] In the formula, Let be the amplitude ratio of the terminal vibration signal and the headphone vibration signal at the i-th sampling moment; Let be the amplitude value of the terminal vibration signal at the i-th sampling time, 1≤i≤N, where N is the total number of sampling times of the terminal vibration signal, and the total number of sampling times of the headphone vibration signal is the same as the total number of sampling times of the terminal vibration signal. Let be the amplitude value of the terminal vibration signal at the i-th sampling time.
[0151] Step S808: The user terminal calculates the variance of the amplitude ratio of the terminal vibration signal and the headphone vibration signal at different sampling times.
[0152] Specifically, when a user is speaking normally, the user terminal focuses more on monitoring the vibration of the overall external environment, and therefore hardly detects the vibration signal of the user's head caused by speaking. As a result, the difference between the amplitude ratio of the terminal vibration signal and the headphone vibration signal at different sampling times will be large (i.e., there is no obvious similarity between the terminal vibration signal and the headphone vibration signal), resulting in a large variance value. Conversely, when the variance value is small, it indicates that there is an obvious similarity between the terminal vibration signal and the headphone vibration signal. In other words, the vibration signal detected by the user terminal is similar to the vibration signal detected by the in-ear player, but the user terminal cannot actually detect the vibration signal of the user's head. Therefore, it is only possible that both detected the overall vibration of the external environment (such as the overall vibration felt by the user when riding a train). In this case, even if the first condition is met, the situation where the user is speaking should be excluded (i.e., the playback volume of the second target audio should not be reduced).
[0153] Step S809: If the intensity ratio is less than a first preset ratio (e.g., 0.9) and greater than a second preset ratio (e.g., 1.1), and the variance is less than a preset variance (e.g., 0.5), the user terminal generates a normal instruction and deletes the tone reduction instruction.
[0154] Specifically, in summary, when the intensity ratio is less than the first preset ratio and greater than the second preset ratio, and the variance is less than the preset variance, it indicates that the user is in a relatively noisy and vibrating external environment (such as riding in a car or boat). Although the first condition is met, the user is not actually in a state of conversation or speaking, so correction is needed (i.e., the playback volume of the second target audio is no longer reduced). Therefore, a normal instruction is generated and the volume reduction instruction is deleted.
[0155] Step S810: When a normal instruction exists, the user terminal no longer generates a volume reduction instruction and controls the speaker to restore the volume of the second target audio being played to the original volume.
[0156] Specifically, when a normal instruction is present, the user terminal will no longer generate a volume reduction instruction until the user leaves the overall noisy and vibrating external environment.
[0157] This invention also proposes a system for generating expressive audio based on text data, and applies a method for generating expressive audio based on text data; the system includes a management terminal and a server that are communicatively connected to each other.
[0158] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0159] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A method for generating expressive audio based on text data, characterized in that, An application to a system for generating expressive audio based on text data; the system includes a management terminal and a server that are communicatively connected to each other, and a user terminal that is communicatively connected to the server; The user terminal includes an in-ear player; the in-ear player includes a speaker, a first microphone, and a first vibration sensor; The method includes: The management terminal acquires the text to be converted and sends it to the server; The server splits the text to be converted into multiple corresponding text segments and determines the sound effect tags corresponding to each text segment. The server obtains the audio of the sound effect corresponding to each sound effect tag; The server performs speech conversion on the text to be converted to generate the first target audio. The server determines the insertion time of each sound effect audio in the first target audio, and synthesizes each sound effect audio and the first target audio based on the insertion time of the sound effect audio in the first target audio to obtain the second target audio; The server sends the second target audio to the user terminal; The user terminal controls the speaker to play the second target audio. During the playback of the second target audio, the user terminal acquires the first ambient sound signal collected in real time by the first microphone, and the first vibration signal of the user's head collected in real time by the first vibration sensor. The user terminal determines whether the first condition is met, wherein the first condition is: the average amplitude of the first vibration signal within the past second preset time period is greater than the first preset amplitude, and the average intensity of the first sound signal within the past second preset time period is greater than the first preset intensity. If the first condition is met, the user terminal marks the current volume of the second target audio being played by the speaker as the original volume and generates a volume reduction command; When a volume reduction command is present, the user terminal controls the speaker to reduce the volume of the second target audio being played based on the volume reduction command; The server stores a sound effect matching word library, which includes multiple sound effect tags and matching words corresponding to each sound effect tag. The server determines the insertion time of each sound effect audio in the first target audio, and synthesizes each sound effect audio and the first target audio based on the insertion time of the sound effect audio in the first target audio to obtain the second target audio, including: The server obtains the sound effect tags corresponding to the text segment and marks them as target tags; The server determines whether the text segment contains matching words corresponding to the target tag based on the sound effect matching word library; If so, the server marks the matching words in the text segment as target words, and marks the time period in which the target words in the text segment appear in the first target audio as the first target time period; The server marks the end time of the first target time period as the first start time, and marks the end time of the time period in which the last character in the text segment appears in the first target audio as the first end time. The server determines the time period between the first start time and the first end time as the insertion time period of the sound effect audio corresponding to the target label in the first target audio.
2. The method for generating expressive audio based on text data according to claim 1, characterized in that, The server runs a TTS model; the server stores a sound effect material library; the server splits the text to be converted into multiple corresponding text segments, and determines the sound effect tags corresponding to each text segment, including: The server uses predetermined punctuation marks in the text to be converted as delimiters to divide the text into multiple text segments, wherein the predetermined punctuation marks include commas, periods, and semicolons. The server performs speech-to-speech conversion on the text to be converted to generate the first target audio, including: The server uses the TTS model to perform speech conversion on the text to be converted to generate the first target audio. The server obtains the audio of the sound effect corresponding to each sound effect tag, including: The server obtains the audio of each audio effect tag based on the audio effect material library.
3. The method for generating expressive audio based on text data according to claim 1, characterized in that, The server splits the text to be converted into multiple corresponding text segments and determines the corresponding sound effect tags for each text segment, and also includes: The server constructs a large language model based on a neural network and obtains a training sample set, which includes multiple training texts and sound effect tags corresponding to each training text. The sound effect tags are sound effects that correspond to the semantics of the training text. When the training text is descriptive text, the corresponding sound effect tag is none. The server uses the training text in the training sample set as input parameters and the corresponding sound effect tags of the training text as output parameters to train the large language model based on neural networks. The server inputs each text segment into the trained neural network-based large language model to obtain the output sound effect tags, and uses the output sound effect tags as the corresponding sound effect tags for each text segment.
4. The method for generating expressive audio based on text data according to claim 1, characterized in that, The server determines whether the text segment contains matching words corresponding to the target tag based on the sound effect matching word library, and then further includes: If not, the server will mark the text segment as a custom text segment and send the custom text segment to the management terminal; The management terminal obtains manually input labeled words corresponding to the custom text segment and sends the labeled words to the server; The server marks the time period in which the tagged words of the custom text segment appear in the first target audio as the second target time period; The server marks the end time of the second target time period as the second start time, and marks the end time of the time period in which the last character in the custom text segment appears in the first target audio as the second end time. The server determines the time period between the second start time and the second end time as the insertion period of the sound effect audio corresponding to the target label in the first target audio.
5. The method for generating expressive audio based on text data according to claim 1, characterized in that, The server marks the end time of the first target time period as the first start time, and marks the end time of the time period in which the last character in the text segment appears in the first target audio as the first end time, and then includes: The server determines whether the duration between the first starting time and the first ending time is less than a first preset duration. If so, the server marks the time elapsed after the first starting time for a first preset duration as the adjusted ending time; The server determines the time period between the first starting time and the adjusted ending time as the insertion time of the sound effect audio corresponding to the target label in the first target audio.
6. The method for generating expressive audio based on text data according to claim 1, characterized in that, If the first condition is met, the user terminal marks the current volume of the second target audio being played by the speaker as the original volume and generates a volume reduction command, and then further includes: When a volume reduction command is present, the user terminal determines whether a second condition is met, wherein the second condition is: the average amplitude of the first vibration signal within the past second preset time period is less than the second preset amplitude, and the average intensity of the first sound signal within the past second preset time period is less than the second preset intensity, the second preset amplitude is less than the first preset amplitude, and the second preset intensity is less than the first preset intensity. If the second condition is met, the user terminal generates a recovery command and deletes the volume reduction command; When a recovery command is available, the user terminal controls the speaker to restore the volume of the second target audio being played to the original volume based on the recovery command.
7. A method for generating expressive audio based on text data according to claim 6, characterized in that, The user terminal body is also equipped with a second vibration sensor and a second microphone; the method further includes: The user terminal acquires the second vibration signal collected in real time by the second vibration sensor; The user terminal acquires the second sound signal collected in real time by the second microphone; The user terminal determines whether the first condition is met, and then the process further includes: If the first condition is met, the user terminal marks the generation time of the volume reduction instruction as the target time, and determines a third target time period based on the target time, wherein the middle time of the third target time period is the target time, and the duration of the third target time period is twice the second preset duration. The user terminal marks the second sound signal within the third target time period as the terminal sound signal, and the first sound signal within the third target time period as the headphone sound signal; The user terminal obtains the ratio of the average intensity value of the terminal sound signal to the average intensity value of the headphone sound signal, and marks it as the intensity ratio. The user terminal marks the second vibration signal within the third target time period as the terminal vibration signal, and the first vibration signal within the third target time period as the earphone vibration signal; The user terminal acquires the amplitude values of the terminal vibration signal and the headphone vibration signal at each sampling time, and calculates the amplitude ratio of the terminal vibration signal and the headphone vibration signal at different sampling times: The user terminal calculates the variance of the amplitude ratio of the terminal vibration signal and the headphone vibration signal at different sampling times; If the intensity ratio is less than a first preset ratio and greater than a second preset ratio, and the variance is less than a preset variance, the user terminal generates a normal command and deletes the tone reduction command. When a normal instruction is present, the user terminal will no longer generate a volume reduction instruction and will control the speaker to restore the volume of the second target audio being played to the original volume.
8. A system for generating expressive audio based on text data, characterized in that, The system employs a method for generating expressive audio based on text data as described in any one of claims 1-7; the system includes a management terminal and a server that are communicatively connected to each other.
Citation Information
Patent Citations
Adding background sound to speech-containing audio data
CN107464555A
Audio book manufacturing method and device and storage medium
CN116403561A