Data conversion methods, apparatus and electronic equipment
By encoding and mapping the input data to a vector space of preset dimensions, and using a speech generator to convert the decoded information into audio signals, the problems of high cost and insufficient quality in podcast audio generation are solved, and efficient and automated podcast audio generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-03-06
AI Technical Summary
Existing podcast audio production is costly, relies on manual recording and post-processing, and is difficult to automate on a large scale. Furthermore, existing speech synthesis technology is inadequate in terms of emotional expression and sound quality, making it difficult to generate high-quality podcast content.
By acquiring input data, encoding it, mapping it to a vector space of preset dimensions, and using a speech generator to convert the decoded information into continuous audio signals, the system achieves end-to-end generation of audio signals from various data types, simplifying the operation process and improving production efficiency.
It enables end-to-end generation of audio signals from various data types, simplifies the operation process, improves the efficiency of audio signal production, and generates natural, fluent, and emotionally resonant podcast content while ensuring audio conversion quality.
Smart Images

Figure CN118918877B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a data conversion method, apparatus and electronic device. Background Technology
[0002] Podcasts are a form of audio content distribution and dissemination, shared and subscribed to through internet platforms. Podcast audio production requires the participation of hosts, guests, and technical personnel, especially experts in the relevant fields, which is costly. Furthermore, high-quality recording equipment and professional recording studio environments increase the hardware costs of audio content generation, making large-scale automated production of podcast audio impossible. Summary of the Invention
[0003] The purpose of this disclosure is to provide a data conversion method, apparatus, and electronic device to convert any video data, audio data, and text data into podcast audio in a single manner, while ensuring the quality of audio conversion and enabling the mass production of podcast audio.
[0004] In a first aspect, this disclosure provides a data conversion method, which includes: acquiring input data; encoding the input data to obtain an encoded vector; wherein the input data includes at least one of the following: text data, audio data, and video data; mapping the encoded vector to a vector space of a preset dimension to obtain a first vector corresponding to the encoded vector; decoding the first vector to obtain decoded information; wherein the decoded information is used to indicate semantic information corresponding to the input data; and converting the decoded information into a continuous audio signal through a speech generator; wherein the audio signal matches the semantic information corresponding to the input data.
[0005] Secondly, this disclosure provides a data conversion apparatus, comprising: a data acquisition module for acquiring input data and encoding the input data to obtain an encoded vector; wherein the input data includes at least one of the following: text data, audio data, and video data; a vector mapping module for mapping the encoded vector to a vector space of a preset dimension to obtain a first vector corresponding to the encoded vector; a vector decoding module for decoding the first vector to obtain decoded information; wherein the decoded information is used to indicate semantic information corresponding to the input data; and a speech generation module for converting the decoded information into a continuous audio signal through a speech generator; wherein the audio signal matches the semantic information corresponding to the input data.
[0006] Thirdly, this disclosure provides an electronic device including a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor, the processor executing the machine-executable instructions to implement the above-described data conversion method.
[0007] Fourthly, this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when invoked and executed by a processor, cause the processor to implement the aforementioned data conversion method.
[0008] The embodiments disclosed herein bring the following beneficial effects:
[0009] This disclosure provides a data conversion method, apparatus, and electronic device. First, input data is acquired and encoded to obtain an encoded vector. The input data includes at least one of the following: text data, audio data, and video data. The encoded vector is then mapped to a vector space of a preset dimension to obtain a first vector corresponding to the encoded vector. The first vector is then decoded to obtain decoded information, which indicates the semantic information corresponding to the input data. Finally, a speech generator converts the decoded information into a continuous audio signal, where the audio signal matches the semantic information corresponding to the input data. This method achieves end-to-end generation of audio signals from input data of various data types, eliminating intermediate steps and additional processes, simplifying the speech generation process. Simultaneously, this method improves the production efficiency of audio signals while ensuring audio conversion quality.
[0010] Other features and advantages of this disclosure will be set forth in the following description, or some features and advantages may be inferred from the description or determined without doubt, or may be learned by practicing the techniques described above.
[0011] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the specific embodiments of this disclosure or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0013] Figure 1 A flowchart of a data conversion method provided in this embodiment of the disclosure;
[0014] Figure 2 This is a schematic diagram of the structure of a data conversion device provided in an embodiment of the present disclosure;
[0015] Figure 3This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0017] Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely to illustrate selected embodiments of the disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0018] A podcast is a form of audio content distribution and delivery, shared and subscribed to through internet platforms. Users can listen to podcasts anytime, anywhere via mobile devices or computers, accessing information and entertainment. Podcast content is typically presented in the form of conversations, interviews, storytelling, and feature reports, covering a wide range of topics, including education, technology, news, entertainment, culture, and health. The characteristics of the podcast format include the following:
[0019] 1. The characteristics of knowledge dissemination: users can listen to podcasts anytime, anywhere, making them suitable for listening during fragmented time such as driving, exercising, or doing housework; podcast programs usually explore specific topics in depth, providing more detailed and in-depth content than other media; covering a wide range of topics to meet the needs and interests of different users.
[0020] 2. Engaging: Many podcasts use an interview format, with the host and guests interacting and discussing topics in a lively and interesting way; podcast hosts can create programs based on their personal style and expression, forming a unique auditory experience; they can convey emotions through sound and establish an emotional connection with the audience.
[0021] However, podcast audio requires manual recording, meaning it necessitates the participation of hosts, guests, and technical personnel, especially experts in the relevant fields. This is costly, and high-quality recording equipment and professional recording studio environments further increase hardware costs. Additionally, podcast audio recording has a long cycle, requiring pre- and post-production processing. Pre-production, including topic selection, scriptwriting, and guest invitations, is time-consuming. The recording process requires multiple trials and adjustments to ensure content quality and effect. Post-production, including editing, mixing, and audio quality optimization, requires professional personnel. Moreover, podcast audio editing is challenging. First, the recorded audio needs to be edited to remove redundant content and flaws, ensuring program coherence and smoothness. Then, noise reduction, equalization, and dynamic range compression are required to ensure clarity and sound quality.
[0022] Therefore, most podcast content generation relies on manual editing, which is inefficient and cannot achieve large-scale automated production. Furthermore, existing speech synthesis technologies have shortcomings in emotional expression, sound quality, and personalization. For example, text-to-speech systems often lack natural emotional expression and high-fidelity sound quality in the generated speech, making it difficult to produce natural, fluent, and emotionally resonant podcast content. Moreover, the generated content often lacks coherence and logical structure, making it difficult to maintain a high-quality user experience.
[0023] To address the aforementioned issues, embodiments of the present invention provide a data conversion method, apparatus, and electronic device. This technology can be applied to scenarios where video data, audio data, and text data are converted into audio signals of a unified format.
[0024] To facilitate understanding of the embodiments of the present invention, a data conversion method disclosed in the embodiments of the present invention will first be described in detail, such as... Figure 1 As shown, the method includes the following specific steps:
[0025] Step S102: Obtain input data and encode the input data to obtain an encoding vector; wherein the input data includes at least one of the following types of data: text data, audio data, and video data.
[0026] In practical implementation, input data can be text, audio, or video data; it can also be text and audio, audio and video, text and video, or a combination of both. To uniformly convert input data of various data types into podcast audio content, preprocessing and encoding of the input data of different data types is first necessary. This step is crucial because it transforms various forms of input into a unified representation, allowing for subsequent processing and generation within the same framework.
[0027] Because different data types correspond to different characteristics of input data, the preprocessing rules and encoding rules for different data types are also different. However, the format of the encoding vectors for different data types is the same. The preprocessing rules and encoding rules for each data type are pre-set according to R&D needs and are not specifically limited here. Subsequent embodiments will describe the preprocessing and encoding methods for input data in detail.
[0028] Specifically, the input data can include casually recorded videos, trending news, chat logs from friends, phone recordings, TV programs, etc., which can be combined in any way to generate audio signals. This flexibility allows users to make full use of data from various sources, enhancing the richness of the audio signal content. Compared to traditional podcast recording, which mainly relies on a single modality (such as text or audio) and requires manual integration of different data types, this method can achieve unified processing and generation of multiple data types through automation.
[0029] Step S104: Map the encoded vector to a vector space of a preset dimension to obtain the first vector corresponding to the encoded vector.
[0030] In practical implementation, to achieve a unified representation of the encoded vectors, they need to be mapped to a vector space of a preset dimension to obtain the encoded vector mapped to this preset dimension, which is also known as the first vector. The preset dimension can be determined based on R&D requirements. This method maps encoded vectors corresponding to input data of different data types to a vector control of a unified dimension, facilitating subsequent decoding and speech conversion processing.
[0031] Step S106: Decode the first vector to obtain decoded information; wherein the decoded information is used to indicate the semantic information corresponding to the input data.
[0032] In a specific implementation, the first vector is a vector of a preset dimension. The first vector is input into a pre-trained decoder. Based on the analysis of the contextual relationship in the first vector by the decoder, the semantic information corresponding to the input data can be obtained, so that the decoder outputs decoded information containing the speech information corresponding to the input data.
[0033] In step S108, the decoded information is converted into a continuous audio signal by a speech generator; wherein the audio signal is matched with the semantic information corresponding to the input data.
[0034] Specifically, the decoded information output by the decoder first needs to be passed to the speech generator. The speech generator uses speech synthesis technology to convert the decoded information into a natural and fluent speech signal. During the speech signal generation process, the intonation, rhythm, and other aspects of the audio output can be dynamically adjusted based on the input characteristics to ensure that the generated audio is natural and fluent. For example, if the decoder outputs "Welcome to our podcast," the speech generator will convert this information into a natural and fluent speech signal.
[0035] The aforementioned data conversion method enables end-to-end generation of audio signals from input data of various data types, eliminating the need for intermediate steps and additional processes. This simplifies the speech generation process and, while ensuring audio conversion quality, improves the efficiency of audio signal production.
[0036] The following examples describe a method for preprocessing and encoding input data.
[0037] Specifically, the process of encoding the input data to obtain the encoded vector described above can be achieved through the following steps 10-12:
[0038] Step 10: Standardize the input data according to preset rules to obtain standardized input data; wherein, the preset rules include standardization rules corresponding to multiple preset data types, including text type, audio type and video type.
[0039] In practical implementation, the above-mentioned preset rules include standardized processing rules corresponding to each of the multiple preset data types. Different preset data types have different standardized processing rules, and the standardized processing rules corresponding to each preset data type can be determined according to R&D needs.
[0040] For example, standardization rules for text types include: removing redundant spaces and special characters from text data, and standardizing case. These rules allow for the standardization of text input data, specifically removing or replacing special characters, including but not limited to HTML (HyperText Markup Language) tags and URLs (Uniform Resource Locators). Uppercase letters in the input data are then converted to lowercase for easier processing, and regular expressions are used to remove redundant spaces and punctuation. Standardization rules for audio types include noise reduction and volume normalization. Since audio data is a continuous signal, standardization rules are necessary to remove background noise and adjust volume to ensure consistent loudness across different audio streams. Standardization rules for video data include frame rate adjustment and resolution normalization. Video data can be standardized by using standardized processing rules corresponding to the video data. This means adjusting the frame rate of the video data so that different video segments have a consistent frame rate (e.g., 30 frames per second), and then adjusting the video resolution to make it have a consistent size (e.g., 224x224 pixels).
[0041] Step 11: Divide the standardized input data into multiple sub-data.
[0042] In practical implementation, since the standardized input data varies in length, it is necessary to divide the standardized input data into multiple sub-data of equal length so that the sub-data can be encoded uniformly. The division method differs for different preset data types of input data; for example, text data requires text segmentation, video data requires video frame segmentation, and audio data requires audio frame segmentation.
[0043] Text segmentation is the process of dividing a continuous piece of text data into meaningful words or sub-words. For example, the Byte-Pair Encoding (BPE) algorithm can be used. It breaks down text data into smaller units by progressively merging the most frequent character or sub-word pairs, preserving good semantic representation. For instance, "hello world" can be broken down into "he", "ll", "o", "wor", and "ld". Sub-words are text units that fall between words and characters. Sub-words divide a continuous piece of text into smaller, meaningful units. These units can be entire words or parts of words. Sub-word segmentation methods strike a balance between words and characters, handling word-level information while effectively processing new or rare words by segmenting them into sub-word units.
[0044] Audio framing involves dividing continuous audio data into multiple smaller segments to facilitate further processing. For example, audio data can be divided into frames of 20-50 milliseconds in length, with some overlap between frames (e.g., 10 milliseconds). A 10-second audio dataset could be divided into 500 20-millisecond frames. Specifically, audio data can be divided into multiple short-segment frames using tools including, but not limited to, audio processing tools. For instance, functions from the Librosa library can be used to frame audio data.
[0045] Video framing involves dividing video data into multiple independent video frames, each of which is an image. Dividing video data into independent video frames requires ensuring that each second of video contains the same number of frames. For example, a 10-second video at a frame rate of 30 frames per second would be divided into 300 frames.
[0046] Step 12: For multiple sub-data, encode the sub-data to obtain the encoding vector corresponding to each sub-data.
[0047] In practice, different encoding rules are used for sub-data of different preset data types, and the encoding rules for each preset data type are pre-set. According to the preset encoding rules, each sub-data can be encoded into a string (equivalent to the encoding vector mentioned above) so that the string can be uniformly decoded later.
[0048] The encoding rule for text data is to convert the segmented text (i.e., the aforementioned sub-data) into vector representations so that they can be processed by machine learning models. Specifically, embedding layers of various pre-trained language models (such as BERT, GPT, Transformer, etc.) can be used to convert each text segment into a fixed-dimensional vector representation. For example, "hello" might be converted into a vector of length 768: [0.1, 0.2, 0.3, ..., 0.768].
[0049] In the encoding and processing of audio data, the first step is to generate a spectrogram corresponding to each sub-data unit, and then convert the spectrogram into a fixed-dimensional vector representation. Specifically, the spectrogram is the frequency domain representation of the audio data, used to describe the frequency changes of the audio signal over time. A short-time Fourier transform (SFT) can be used to convert the time-domain signal of each frame in the sub-data unit into a frequency-domain representation, resulting in the spectrogram. For example, the audio data is first divided into several short time segments (frames), each typically 20-50 milliseconds. A window function (such as a Hamming window) is applied to each frame to reduce edge effects. Then, a fast Fourier transform is performed on each frame to convert the time-domain signal into a frequency-domain signal. The amplitude of each frequency-domain signal is then squared to obtain the power spectrum. The power spectra of each frame are arranged together to form a two-dimensional image, with time on the horizontal axis and frequency on the vertical axis. Color or brightness represents power. After obtaining the spectrogram, it needs to be converted into a vector representation so that it can be processed by machine learning models. Specifically, a pre-trained audio model (such as the Audio Spectrogram Transformer, AST) can be used to encode the spectrogram, generating a fixed-dimensional vector representation. For example, a spectrum can be encoded as a vector of length 512: [0.1,0.2,0.3,...,0.512].
[0050] In the encoding and processing of video data, the first step is to extract feature vectors from the sub-data. Then, these feature vectors are converted into a unified vector representation, aligning them with the vector representations of text and audio data. Specifically, for each video frame, a pre-trained visual encoder (ViT) is used to extract feature vectors from the corresponding image. ViT is a model that uses a Transformer architecture to process image data and extracts global visual features. Its working principle is as follows: First, the image is divided into fixed-size blocks, such as 16x16 pixels. Then, each block is flattened into a one-dimensional vector and mapped to a high-dimensional space through a linear transformation. Positional encoding is added to each block to preserve spatial information. The block embeddings and positional encodings are then input into the Transformer, which uses a self-attention mechanism to calculate the relationships between each block, generating a global representation of the image. After obtaining the feature vectors of the video frames, a linear transformation is performed to ensure they reside in the same vector space as the vector representations of text and audio. For example, the feature vectors of the video frames are mapped to a vector space of length 768, aligning with the text representation.
[0051] The above approach, through preprocessing and encoding of text, audio, and video data, transforms input data from different modalities into a unified vector representation. These vector representations are then passed to a multimodal model for further processing and generation, ultimately achieving end-to-end podcast audio content output. These steps ensure that data from various input modalities can be processed uniformly within the same framework, laying a solid foundation for subsequent multimodal generation.
[0052] The following examples are used to describe a unified representation of vectors.
[0053] Specifically, the process of mapping the encoded vector to a vector space of a preset dimension to obtain the first vector corresponding to the encoded vector may include: mapping the encoded vector to a preset vector space to obtain the mapped encoded vector; and multiplying the mapped encoded vector with a target matrix of a preset dimension to map the mapped encoded vector to a vector space of a preset dimension to obtain the first vector.
[0054] In practical implementation, to uniformly process different preset data types, it is necessary to map the encoding vectors corresponding to different preset data types to the same representation space, which is the aforementioned vector space of preset dimensions. Specifically, cross-modal attention mechanisms can be used to map data from different modalities (here, modality is equivalent to preset data type) to the same representation space. For example, cross-modal attention networks (MCAN) can be used to align text, audio, and video features.
[0055] In a specific implementation, the process of mapping the encoding vector to a preset vector space to obtain the mapped encoding vector may include: generating position codes and modal tags corresponding to the encoding vector; wherein, the position codes are used to indicate the order relationship of each element in the encoding vector, and the modal tags are used to indicate the data type of the input data corresponding to the encoding vector; embedding the position codes and modal tags into the encoding vector to obtain the mapped encoding vector.
[0056] Specifically, the positional encoding described above is used to add positional information to each encoded vector to preserve the order of elements within the vector. This positional information is needed in subsequent decoding because it inherently lacks order information. In practical applications, the process of generating positional codes can include: generating corresponding positional information for each input element (such as a word in text, a frame in audio, or a frame in video). Positional codes are typically generated using sine and cosine functions to ensure the uniqueness of the encoding at different positions. Then, the generated positional information is added to each encoded vector to form a feature representation with positional information. For example, for text input, the generated positional code is added to the text's feature vector; that is, for each position, the generated sine and cosine values are added to the corresponding dimension of the feature vector to obtain the encoded vector with positional encoding.
[0057] The modality tagging described above adds label information to the encoding vector of each modality to distinguish data of different modalities. Here, modality is also a data type, including text, audio, and video. The process of generating modality tags involves: generating a fixed embedding vector for each modality as its tag; for example, the tag vector for text can be a specific fixed vector, while the tag vectors for audio and video can be other fixed vectors. Then, the modality tag is embedded into the corresponding modality's encoding vector, enabling the model to recognize and distinguish data of different modalities. For example, the text encoding vector is added to EtextE_{text}Etext (equivalent to the modality tag), and the audio encoding vector is added to EaudioE_{audio}Eaudio (equivalent to the modality tag).
[0058] After obtaining the mapped encoded vectors, it is necessary to map them to a unified vector space of a predetermined dimension using a target matrix of a predetermined dimension, resulting in the first vector. This allows for unified processing of encoded vectors from different modalities. Here, a linear transformation is used to map feature vectors from different modalities to a vector space of the same dimension. During the linear transformation, a linear transformation layer (fully connected layer) needs to be defined for each modality to map the input feature vector to a predetermined dimension. For example, the encoded vectors for audio, video, and text are all mapped to a 768-dimensional vector space. The target matrix corresponding to the linear transformation layer is learned during training.
[0059] The above approach utilizes cross-modal attention mechanisms and linear transformation techniques to achieve the alignment and fusion of data from different modalities, ensuring the coherence and consistency of the generated content. This technique enables multimodal data to be processed within the same representation space, which is rare in existing technologies.
[0060] The following examples are used to describe the feature fusion method.
[0061] Specifically, the input data includes data of at least two types, that is, the input data is multimodal data. Based on this, after mapping the encoding vector to a vector space of a preset dimension to obtain the first vector corresponding to the encoding vector, it is also necessary to perform feature fusion processing on the first vector corresponding to the input data of at least two types to obtain the first vector after feature fusion, so that the first vector after feature fusion can be decoded to obtain the decoded information.
[0062] For example, the input data includes text data, audio data, and video data. This requires feature fusion processing of the first vector corresponding to the text, the first vector corresponding to the video, and the first vector corresponding to the audio to generate a unified encoding vector in order to obtain a unified vector representation.
[0063] In a specific real-time scenario, the process of performing feature fusion processing on the first vector corresponding to input data of at least two data types to obtain the first vector after feature fusion may include: determining the fusion weight corresponding to the first vector of each of the at least two data types; multiplying the first vector of each data type with the corresponding fusion weight to obtain the multiplication result corresponding to each data type; and adding the multiplication results corresponding to each data type to obtain the first vector after feature fusion.
[0064] Specifically, an attention layer can be designed to calculate the fusion weights between the first vectors corresponding to different data types. Then, the attention mechanism is used to fuse the first vectors corresponding to different data types into a unified feature vector representation. For example, the fusion weights between the first vectors corresponding to text and audio can be calculated, merging their corresponding first vectors into a unified 768-dimensional vector. The attention mechanism is a method for calculating relationships between vectors, particularly for handling relationships between different modalities (such as text, audio, and video) of data. By calculating fusion weights, it highlights important information and reduces the influence of irrelevant information on the model. Through fusion weight calculation, vectors of different data types can be aligned and fused in the same representation space, ensuring that the information between vectors corresponding to multiple data types can complement each other, thus improving the model's performance.
[0065] The purpose of calculating fusion weights is to highlight important information in vectors from different data models, align vectors corresponding to different data types, and ensure they are represented and processed in the same vector space. Through an attention mechanism, the model can dynamically focus on the most important parts of the input data for the generation task, improving the generation effect and quality. In a specific embodiment, the steps for calculating fusion weights include: the input data includes text data, audio data, and video data, and the first vector includes text feature vectors, audio feature vectors, and video feature vectors; for the feature vectors corresponding to each data type, weight matrices Q and K are obtained by performing a linear transformation on the feature vectors; the dot product of weight matrices Q and K is calculated and then divided by a scaling factor to obtain the fusion weight matrix; then the fusion matrix is normalized to obtain the normalized fusion weight matrix.
[0066] It should be noted that during the encoding stage, the encoding vectors corresponding to each data type have been mapped to a unified vector space of a preset dimension. This simply converts data of different data types into vector representations of the same dimension for subsequent processing. However, even though the vectors are unified in dimension, the information and semantics they contain may not be completely aligned. Therefore, feature fusion is needed to ensure that information from different data types is consistently represented and combined in the semantic space.
[0067] In one specific embodiment, feature fusion may include: using a cross-modal attention mechanism to enable information between vectors of different data types to reference and complement each other, and to align the vectors. The alignment process ensures the semantic consistency of text, audio, and video features. The aligned vectors are then weighted and summed to generate a unified feature representation, which is the fused vector. This vector contains information from all data types, improving the performance of subsequent generation tasks.
[0068] Among the feature fusion methods described above, the fused vector can better represent the overall picture of the input data, improving the model's ability to understand the input data. Simultaneously, through alignment and fusion, the model can more accurately generate the target output (such as high-quality audio), improving the performance of the generation task.
[0069] The following examples are used to describe the decoding method and the method of generating audio signals.
[0070] Specifically, the process of decoding the first vector to obtain the decoded information may include: inputting the first vector into a pre-trained decoder, obtaining the correlation between the elements in the first vector through the decoder, and encoding the corresponding position of the first vector based on the correlation and the position of the first vector to obtain the decoded information corresponding to the first vector.
[0071] In practical implementation, self-attention can be used to calculate the correlation between each element in the first vector and other elements. This involves performing multi-head self-attention computation on the first vector to obtain the context representation at each time step, thereby determining the relationships between elements in the first vector and capturing global dependencies. Since the decoder lacks order information, positional encoding is used to provide positional information between elements in the first vector. Combining the relationships between elements in the first vector with the positional information yields the decoded information corresponding to the first vector, which includes the semantic information of the input data corresponding to the first vector.
[0072] In one specific embodiment, the process of obtaining the decoded information corresponding to the first vector based on the association relationship and the positional encoding corresponding to the first vector may include: dividing the first vector into multiple sub-vectors according to the positional encoding corresponding to the first vector; for each sub-vector, generating decoded information corresponding to the current sub-vector based on the association relationship between the elements in the current sub-vector using a decoder, and generating decoded information corresponding to the next sub-vector based on the association relationship between the decoded information corresponding to the current sub-vector and the elements in the next sub-vector of the current sub-vector; combining the decoded information corresponding to multiple sub-vectors to obtain the decoded information corresponding to the first vector. In this method, the decoded information generated by each sub-vector is used as the input to its corresponding next sub-vector, thereby gradually obtaining the entire decoded information.
[0073] In an optional embodiment, the decoder, based on an autoregressive generation mechanism, predicts the audio signal at each time step in the first vector after feature fusion, thereby progressively generating semantic and audio signals. For example, assuming the first vector represents a podcast segment introducing a certain topic, the decoder will progressively generate the audio signal of that podcast segment. The autoregressive mechanism generates the output progressively, with the output of each step serving as the input for the next step. This generation method ensures the coherence and consistency of the generated content.
[0074] Furthermore, in order to generate high-quality audio signals, the decoded information needs to be denoised and refined to obtain high-fidelity decoded information, so that the speech generator can perform speech generation processing on the high-fidelity decoded information.
[0075] In practical implementation, a latent diffusion model can be used for denoising and refinement. This involves diffusing the decoded information within the latent space to progressively denoise and generate high-quality target data, which is the aforementioned high-fidelity decoded information. Specifically, in audio generation, the latent diffusion model first randomly samples the decoded information in the latent space to generate an initial low-resolution audio signal. Then, through a diffusion process, it progressively improves the audio quality and detail, generating a high-resolution audio signal. In other words, through a series of diffusion steps, each step builds upon the output of the previous step, progressively improving the audio quality and detail.
[0076] The diffusion process can also be achieved through the following steps: the latent diffusion model learns the distribution of data by progressively adding and removing noise. That is, noise is initially added to generate a low-resolution audio signal, and then noise is progressively removed to generate a high-resolution audio signal. For example, assuming the initially generated low-resolution audio signal represents "Welcome to our podcast," the diffusion process progressively improves the audio quality, eventually generating a high-fidelity audio signal. Then, inputting the high-fidelity audio signal into a speech generator yields a naturally flowing audio signal.
[0077] The above method, based on an autoregressive generation mechanism and a potential diffusion model, enables the direct generation of high-quality audio content from multimodal input, avoiding intermediate steps and extra processes, and ensuring that the generated audio is natural, smooth, and coherent.
[0078] The following examples are used to provide a detailed description of the training of the model used in this invention.
[0079] We will begin with an introduction to dataset preparation and preprocessing.
[0080] The dataset includes a large amount of text, audio, and video data. The text data includes press releases, lecture transcripts, podcast scripts, etc., covering a wide range of topics; the audio data includes podcast recordings, speech recordings, and conversation recordings; and the video data includes lecture videos, interview videos, and news videos. Specifically, web crawlers can be used to collect data from publicly available podcast platforms, news websites, video sharing platforms, etc., and high-quality raw content can also be collected using audio and video recording equipment.
[0081] After acquiring the data in the dataset, it is necessary to label the data: text data needs to be labeled with themes, key points, and sentiment; audio data needs to be labeled with speech, music, and noise; and video data needs to be labeled with important frames, characters, scenes, etc. After labeling, the data also needs to be cleaned and preprocessed. Cleaning is to remove noise and irrelevant information to ensure data quality, and also to ensure the time synchronization of text, audio, and video data by aligning the data of each data type using timestamps.
[0082] To ensure that multi-data types can be effectively processed and trained by multimodal language models, data preprocessing is crucial. This preprocessing includes data alignment and data normalization. Data alignment is achieved as follows:
[0083] 1. Text-audio alignment
[0084] Add timestamps to both text and audio data to mark the start and end times of each sentence. For example, in a lecture recording, the text is labeled as: "Hello everyone, welcome to the lecture. [00:00:05-00:00:10]". Based on the timestamps, the audio data is segmented into corresponding sentence segments and aligned with the text sentences. For example, a segment of audio in the lecture recording is synchronized with a corresponding text sentence. In practice, timestamp annotation can be performed using tools including but not limited to audio processing tools (such as Praat). Taking Praat as an example, timestamps can be added manually or automatically, and then the audio can be segmented based on the annotation results to obtain audio segments synchronized with the text, thereby achieving the data alignment goal.
[0085] 2. Text-Video Alignment
[0086] First, keyframes are extracted from the video data. These frames typically correspond to important events or scenes in the video, such as the start and end frames of each speech in a news video. Then, a timestamp is added to each keyframe, marking the specific time the frame appeared; for example, a timestamp is added to each keyframe of a speech. Finally, based on the timestamps, the video frames are aligned with the text content. For example, a text description in a news video is synchronized with its corresponding video frame. In practice, keyframe extraction and time stamping can be performed using tools including but not limited to video processing tools (such as FFmpeg). Taking FFmpeg as an example, keyframes can be extracted from the video, each frame can be timestamped, and then these frames can be aligned with the text content to achieve our data alignment goal.
[0087] When standardizing data, the mean and standard deviation of the feature vectors for each data type are first calculated. For example, the mean and standard deviation of each frequency component in an audio spectrogram are calculated. Then, the feature vectors are standardized so that the mean is 0 and the standard deviation is 1. For example, the mean of each frequency component in the audio spectrogram is subtracted, and then the result is divided by the standard deviation. In practice, feature standardization can be performed using tools including but not limited to data processing tools (such as NumPy). Taking NumPy as an example, the mean and standard deviation of the feature vectors can be calculated, and then standardized to achieve our feature standardization goal.
[0088] After preparing the dataset, the next step is to use it to train a multimodal language model. The following are the detailed technical steps, describing the process of each step from a technical perspective:
[0089] 1. Data loading and batch processing:
[0090] Data loading refers to loading the preprocessed dataset, including feature vectors from text, audio, and video data. Specifically, the preprocessed dataset is read from the storage system, and the data is then organized into a process-friendly format, such as tensors or arrays, ensuring that the feature vectors and labels remain consistent throughout the loading process. To facilitate model training, a data loader is used to load the data from the dataset into memory in batches.
[0091] To reduce memory usage and improve computational efficiency, the data in the loaded dataset is divided into several batches. First, the data in the dataset is divided into fixed-size batches (e.g., 32 data samples per batch). Then, it is ensured that the data in each batch is consistent in data type and timing. Specifically, batch processing techniques are used to divide the data into smaller chunks suitable for parallel computing, thereby improving training efficiency and speed.
[0092] 2. Model Initialization
[0093] Multimodal language models consist of an encoder, decoder, fusion layer, and generation layer. Based on a predefined model architecture, the encoder (e.g., text encoder, audio encoder, video encoder), fusion layer (e.g., cross-modal attention mechanism for feature fusion of vectors), decoder, and generation layer (including autoregressive generators, latent diffusion models, and speech generators) are initialized. Specifically, deep learning frameworks (e.g., TensorFlow, PyTorch) can be used to build and initialize the network structures of each part of the model.
[0094] Next, the model's parameters (such as weights and biases) are initialized to prepare for the training process. Initialization is performed using parameters from a pre-trained model (e.g., loading pre-trained parameters from models like BERT or ViT), while random initialization methods (such as Xavier initialization) are used for newly added layers or untrained parts. In practice, a combination of pre-training and random initialization ensures that the model has good performance and convergence in the early stages of training.
[0095] 3. Forward propagation steps:
[0096] First, the batch-processed data is input into the initialized model to begin forward propagation computation. Specifically, the feature vectors of text, audio, and video data are input into their respective encoders to obtain the high-dimensional feature vectors output by each encoder (equivalent to the aforementioned encoding vectors). Through forward propagation, the input data is processed through each layer of the network to obtain the output of each layer. Then, a cross-modal attention mechanism is used to fuse the feature vectors of different data types, generating a unified multimodal feature vector. The cross-modal attention layer then calculates the fusion weights between the text, audio, and video feature vectors, and these feature vectors are weighted and fused to obtain a unified feature representation. The cross-modal attention mechanism ensures that information from different modalities is comprehensively considered after fusion, improving the effectiveness of the feature representation.
[0097] Then, the fused feature vectors are transformed into the target output (such as an audio signal) through a decoder and a generation layer. Specifically, an autoregressive generation mechanism is used to generate the target audio signal step by step, and then a latent diffusion model is used to further optimize the generated audio signal and improve its quality.
[0098] 4. Calculate the loss function
[0099] First, a suitable loss function needs to be selected to evaluate the difference between the model output and the true labels. For audio generation tasks, mean squared error (MSE) or spectral reconstruction error can be used as loss functions. These loss functions quantify the error between the model output and the true target, providing guidance for optimizing model parameters. Then, the audio output generated by the model is compared with the true audio labels, and the loss value is calculated.
[0100] 5. Backpropagation and parameter update
[0101] First, the backpropagation algorithm is used to calculate the gradient of the loss function with respect to the model parameters. Specifically, starting from the output of the loss function, the gradient is calculated layer by layer, and the gradient is passed to the parameters of each layer. The chain rule is used to calculate the impact of each layer's parameters on the loss function, obtaining the gradient information. Then, optimization algorithms (such as Adam and SGD) are used to update the model parameters based on the calculated gradients. Specifically, according to the set learning rate, the model parameters are adjusted to optimize them in the direction of minimizing the loss, thereby improving the model's performance and accuracy.
[0102] 6. Model Evaluation and Validation
[0103] During training, the model is periodically evaluated using a validation set to check its generalization performance. Specifically, validation set data is input into the model, and its loss value and evaluation metrics (such as accuracy, precision, and recall) are calculated. Evaluating the model's performance on unseen data using the validation set ensures its good generalization ability. Then, based on the evaluation results, the model architecture, parameters, or training strategy are adjusted to further optimize model performance; that is, the learning rate is adjusted, the number of model layers is increased or decreased, and the loss function is modified based on the validation set evaluation results. Through iterative optimization, the model is continuously adjusted and improved to enhance its final generation results.
[0104] 7. Model Saving and Deployment
[0105] Save the trained model parameters and architecture to the file system for easy use and deployment later. Specifically, you can use the save function provided by the deep learning framework to export the model parameters and architecture as a file. Saving the trained model ensures that it can be effectively loaded and used in future inference and generation processes. Then, deploy the trained model to the production environment to provide real-time audio generation services. Alternatively, you can deploy the model to a server or cloud platform to build an inference service interface for users. Through deployment, the model can be invoked in practical applications to achieve end-to-end generation from multimodal input to audio output.
[0106] By following the steps above, the training process of a multimodal language model can be completed, from data loading, model initialization, forward propagation, loss calculation, backpropagation, parameter update to final model evaluation, saving and deployment, thus achieving the automated generation of high-quality podcast audio content.
[0107] The following examples illustrate how to add personalized audio styles.
[0108] Specifically, based on preset input features, the audio signal is adjusted to obtain the final audio signal; wherein the input features include at least one of the following: emotional characteristics, style features, opening features, ending features, and sound quality features. The input features need to be processed together with the input data to obtain an audio signal that matches the input features.
[0109] In practical implementation, fine-tuning a pre-trained multimodal large language model to achieve high-quality audio generation requires targeted processing during the preparation and preprocessing of training data. The optimization focuses on five key aspects: naturalness and emotional expression of speech, clarity and sound quality of speech, speech style and personalization, content coherence and logical structure, and distinctive openings and closings.
[0110] 1. Naturalness of speech and emotional expression
[0111] Collect audio materials containing different emotional expressions, such as joy, sadness, and anger, ensuring these are common emotional expressions in podcast content. Then, perform sentiment annotation on the training data, labeling the sentiment type and intensity of each audio or text segment. This sentiment annotation information is then used as additional input features to train the model, using sentiment embedding techniques including, but not limited to, sentiment recognition.
[0112] During training, an emotion transfer mechanism is incorporated, enabling the model to generate audio content with corresponding emotions based on input features. For example, if the input text is labeled as "pleasant," the audio generated by the model should exhibit a pleasant tone and emotion.
[0113] 2. Speech clarity and sound quality
[0114] Collect high-quality audio materials, ensuring high standards of recording equipment and environment. For example, clear audio recorded in a studio environment. Perform noise reduction on the audio materials to remove background noise and interference, ensuring clear speech. Use spectral enhancement techniques to improve the spectral detail of the audio signal. For example, spectral smoothing and enhancement algorithms can be used to make the audio signal clearer. Then, apply dynamic range compression techniques to balance the loudness of the audio signal, ensuring speech clarity and consistency.
[0115] 3. Voice style and personalization
[0116] Audio materials from various announcers, hosts, and other individuals are collected, covering diverse voice styles and personalities. The audio data is then style-labeled, annotating the voice style and personality characteristics of each audio segment. This style-labeled information is used as input features and fed into the model for training. Style transfer techniques are employed to enable the model to generate audio content with the corresponding style based on the input features.
[0117] During the model generation process, a personalized optimization mechanism is incorporated to give the generated audio content a specific voice style and personality. For example, by adjusting the model parameters, the generated audio can simulate the voice style of a particular presenter.
[0118] 4. Content coherence and logical structure
[0119] Collect coherent and logically structured text data to ensure logical coherence and consistency of the content. Specifically, collect conversational audio materials to ensure smooth and coherent dialogue. Then, segment and logically annotate the text data to ensure logical coherence of the input content. Use semantic alignment technology to semantically align the input text and generated audio to ensure logical coherence and consistency of the generated content.
[0120] 5. Podcast-style openings and endings
[0121] We collect and organize common podcast opening and closing templates to ensure the generated content retains the podcast's unique characteristics. We standardize the opening and closing sequences, creating template data with a fixed format. In implementation, we embed these templates into the model's input features, ensuring the generated podcast content has standardized openings and closings. During generation, the model automatically generates openings and closings based on the input features, resulting in more professional and complete podcast content.
[0122] For example, each audio or text segment can be labeled with its sentiment type and intensity, such as happy (high), sad (medium), and angry (low), generating sentiment-labeled training data. This sentiment-labeled information is then embedded into the model as additional input features, enabling the model to generate audio that expresses the corresponding emotions. Similarly, each audio segment can be labeled with its speech style, such as formal, informal, or humorous, generating style-labeled training data. This style-labeled information is then embedded into the model as input features, giving the model audio a specific speech style and personality.
[0123] The above method, through emotion and style embedding, enables the generated audio to be personalized and emotionally expressive, meeting the needs of different users. Traditional methods struggle to achieve personalization and emotionalization of audio content; this solution, through model optimization and data annotation, achieves personalized and emotionally rich audio generation.
[0124] This solution incorporates various optimization measures in the data preprocessing and annotation stages, including sentiment annotation, style annotation, and semantic alignment, ensuring high-quality and consistent input data and laying a solid foundation for subsequent generation. This solution not only automates podcast generation but also uses personalized tuning techniques to imbue the generated audio content with individuality and emotional expression. This innovative approach, combining automation and personalization, meets users' demands for high-quality, personalized content.
[0125] By leveraging pre-trained multimodal large models, this solution achieves automated podcast generation, significantly reducing manual intervention and time costs. Users can generate high-quality podcast content without complex manual editing. Traditional methods rely on manual recording and editing, which is inefficient and costly. This solution improves production efficiency through automation, adapting to the needs of large-scale content generation.
[0126] Corresponding to the above method embodiments, this invention also provides a data conversion device, such as... Figure 2 As shown, the device includes:
[0127] The data acquisition module 30 is used to acquire input data and encode the input data to obtain an encoding vector; wherein the input data includes at least one of the following types of data: text data, audio data, and video data.
[0128] The vector mapping module 31 is used to map the encoded vector to a vector space of a preset dimension to obtain the first vector corresponding to the encoded vector.
[0129] The vector decoding module 32 is used to decode the first vector to obtain decoded information; wherein the decoded information is used to indicate the semantic information corresponding to the input data.
[0130] The speech generation module 33 is used to convert the decoded information into a continuous audio signal through a speech generator; wherein the audio signal is matched with the semantic information corresponding to the input data.
[0131] The aforementioned data conversion device enables end-to-end generation of audio signals from input data of various data types, eliminating the need for intermediate steps and additional processes. This simplifies the voice generation process and, while ensuring audio conversion quality, improves the production efficiency of audio signals.
[0132] Specifically, the data acquisition module 30 is used to: standardize the input data according to preset rules to obtain standardized input data; wherein the preset rules include standardization rules corresponding to multiple preset data types, and the multiple preset data types include text type, audio type and video type; divide the standardized input data to obtain multiple sub-data; and encode the sub-data to obtain the encoding vector corresponding to each sub-data.
[0133] In an optional embodiment, the vector mapping module 31 is used to: map the encoded vector to a preset vector space to obtain the mapped encoded vector; and multiply the mapped encoded vector with a target matrix of a preset dimension to map the mapped encoded vector to a vector space of a preset dimension to obtain a first vector.
[0134] Furthermore, the vector mapping module 31 described above is also used to: generate position codes and modal markers corresponding to the encoding vector; wherein, the position codes are used to indicate the order relationship of each element in the encoding vector, and the modal markers are used to indicate the data type of the input data corresponding to the encoding vector; and embed the position codes and modal markers into the encoding vector to obtain the mapped encoding vector.
[0135] In a specific implementation, the input data includes data of at least two data types; based on this, the device further includes a feature fusion module, used to: after mapping the encoded vector to a vector space of a preset dimension to obtain a first vector corresponding to the encoded vector, perform feature fusion processing on the first vector corresponding to the input data of at least two data types to obtain a first vector after feature fusion, so that the first vector after feature fusion can be decoded to obtain decoded information.
[0136] In an optional embodiment, the feature fusion module is further configured to: determine the fusion weight corresponding to the first vector of each of at least two data types; multiply the first vector of each data type with the corresponding fusion weight to obtain the multiplication result corresponding to each data type; and add the multiplication results corresponding to each data type to obtain the first vector after feature fusion.
[0137] In an optional embodiment, the vector decoding module 32 is configured to: input the first vector into a pre-trained decoder, obtain the association relationship between the elements in the first vector through the decoder, and obtain the decoding information corresponding to the first vector based on the association relationship and the position encoding corresponding to the first vector.
[0138] Furthermore, the aforementioned vector decoding module 32 is also used to: divide the first vector into multiple sub-vectors according to the position encoding corresponding to the first vector; for the multiple sub-vectors, generate decoding information corresponding to the current sub-vector based on the relationship between the elements in the current sub-vector using a decoder, and generate decoding information corresponding to the next sub-vector based on the relationship between the decoding information corresponding to the current sub-vector and the elements in the next sub-vector of the current sub-vector; and combine the decoding information corresponding to the multiple sub-vectors to obtain the decoding information corresponding to the first vector.
[0139] Furthermore, the aforementioned device also includes an information optimization module, used to: denoise and refine the decoded information before it is converted into a continuous audio signal by the speech generator, to obtain high-fidelity decoded information, so that the speech generator can perform speech generation processing on the high-fidelity decoded information.
[0140] In a specific implementation, the above device also includes an audio adjustment module, used to: adjust the audio signal based on preset input features to obtain the final audio signal; wherein the input features include at least one of the following: emotional characteristics, style characteristics, opening characteristics, ending characteristics, and sound quality characteristics.
[0141] The data conversion device provided in this disclosure has the same implementation principle and technical effect as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment.
[0142] This invention also provides an electronic device, such as... Figure 3 As shown, the electronic device includes a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, which executes the machine-executable instructions to implement the data conversion method described above.
[0143] Specifically, the above data conversion method includes: acquiring input data; encoding the input data to obtain an encoded vector; wherein the input data includes at least one of the following: text data, audio data, and video data; mapping the encoded vector to a vector space of a preset dimension to obtain a first vector corresponding to the encoded vector; decoding the first vector to obtain decoded information; wherein the decoded information is used to indicate the semantic information corresponding to the input data; and converting the decoded information into a continuous audio signal through a speech generator; wherein the audio signal matches the semantic information corresponding to the input data.
[0144] The aforementioned data conversion method enables end-to-end generation of audio signals from input data of various data types, eliminating the need for intermediate steps and additional processes. This simplifies the speech generation process and, while ensuring audio conversion quality, improves the efficiency of audio signal production.
[0145] In an optional embodiment, the step of encoding the input data to obtain an encoding vector includes: standardizing the input data according to preset rules to obtain standardized input data; wherein the preset rules include standardization rules corresponding to multiple preset data types, and the multiple preset data types include text type, audio type and video type; dividing the standardized input data to obtain multiple sub-data; and encoding the sub-data to obtain an encoding vector corresponding to each sub-data.
[0146] In an optional embodiment, the step of mapping the encoded vector to a vector space of a preset dimension to obtain the first vector corresponding to the encoded vector includes: mapping the encoded vector to a preset vector space to obtain the mapped encoded vector; and multiplying the mapped encoded vector with a target matrix of a preset dimension to map the mapped encoded vector to a vector space of a preset dimension to obtain the first vector.
[0147] In an optional embodiment, the step of mapping the encoding vector to a preset vector space to obtain the mapped encoding vector includes: generating position codes and modal tags corresponding to the encoding vector; wherein, the position codes are used to indicate the order relationship of each element in the encoding vector, and the modal tags are used to indicate the data type of the input data corresponding to the encoding vector; embedding the position codes and modal tags into the encoding vector to obtain the mapped encoding vector.
[0148] In an optional embodiment, the input data includes data of at least two data types; after the step of mapping the encoded vector to a vector space of a preset dimension to obtain the first vector corresponding to the encoded vector, the method further includes: performing feature fusion processing on the first vector corresponding to the input data of at least two data types to obtain the first vector after feature fusion, so that the first vector after feature fusion can be decoded to obtain decoded information.
[0149] In an optional embodiment, the step of performing feature fusion processing on the first vector corresponding to input data of at least two data types to obtain the first vector after feature fusion includes: determining the fusion weight corresponding to the first vector of each of the at least two data types; multiplying the first vector of each data type with the corresponding fusion weight to obtain the multiplication result corresponding to each data type; and adding the multiplication results corresponding to each data type to obtain the first vector after feature fusion.
[0150] In an optional embodiment, the above-described step of decoding the first vector to obtain decoded information includes: inputting the first vector into a pre-trained decoder, obtaining the correlation between elements in the first vector through the decoder, and encoding the corresponding position of the first vector based on the correlation and the position of the first vector to obtain the decoded information corresponding to the first vector.
[0151] In an optional embodiment, the step of obtaining the decoding information corresponding to the first vector based on the association relationship and the position encoding corresponding to the first vector includes: dividing the first vector into multiple sub-vectors according to the position encoding corresponding to the first vector; for the multiple sub-vectors, generating the decoding information corresponding to the current sub-vector based on the association relationship between the elements in the current sub-vector using a decoder, and generating the decoding information corresponding to the next sub-vector of the current sub-vector based on the decoding information corresponding to the current sub-vector and the association relationship between the elements in the next sub-vector of the current sub-vector; and combining the decoding information corresponding to the multiple sub-vectors to obtain the decoding information corresponding to the first vector.
[0152] In an optional embodiment, before the step of converting the decoded information into a continuous audio signal by the speech generator, the method further includes: denoising and refining the decoded information to obtain high-fidelity decoded information, so that the speech generator can perform speech generation processing on the high-fidelity decoded information.
[0153] In an optional embodiment, the above method further includes: adjusting the audio signal based on preset input features to obtain the final audio signal; wherein the input features include at least one of the following: emotional characteristics, style characteristics, opening characteristics, ending characteristics, and sound quality characteristics.
[0154] Furthermore, Figure 3 The electronic device shown also includes a bus 102 and a communication interface 103, with the processor 101, the communication interface 103 and the memory 100 connected via the bus 102.
[0155] The memory 100 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 103 (which can be wired or wireless), such as the Internet, wide area network, local area network, or metropolitan area network. The bus 102 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0156] Processor 101 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 101 or by instructions in software form. Processor 101 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a readily available storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 100, and processor 101 reads information from memory 100 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.
[0157] This invention also provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are called and executed by a processor, they cause the processor to implement the aforementioned data conversion method. For specific implementation details, please refer to the method embodiments, which will not be repeated here.
[0158] Specifically, the above data conversion method includes: acquiring input data; encoding the input data to obtain an encoded vector; wherein the input data includes at least one of the following: text data, audio data, and video data; mapping the encoded vector to a vector space of a preset dimension to obtain a first vector corresponding to the encoded vector; decoding the first vector to obtain decoded information; wherein the decoded information is used to indicate the semantic information corresponding to the input data; and converting the decoded information into a continuous audio signal through a speech generator; wherein the audio signal matches the semantic information corresponding to the input data.
[0159] The aforementioned data conversion method enables end-to-end generation of audio signals from input data of various data types, eliminating the need for intermediate steps and additional processes. This simplifies the speech generation process and, while ensuring audio conversion quality, improves the efficiency of audio signal production.
[0160] In an optional embodiment, the step of encoding the input data to obtain an encoding vector includes: standardizing the input data according to preset rules to obtain standardized input data; wherein the preset rules include standardization rules corresponding to multiple preset data types, and the multiple preset data types include text type, audio type and video type; dividing the standardized input data to obtain multiple sub-data; and encoding the sub-data to obtain an encoding vector corresponding to each sub-data.
[0161] In an optional embodiment, the step of mapping the encoded vector to a vector space of a preset dimension to obtain the first vector corresponding to the encoded vector includes: mapping the encoded vector to a preset vector space to obtain the mapped encoded vector; and multiplying the mapped encoded vector with a target matrix of a preset dimension to map the mapped encoded vector to a vector space of a preset dimension to obtain the first vector.
[0162] In an optional embodiment, the step of mapping the encoding vector to a preset vector space to obtain the mapped encoding vector includes: generating position codes and modal tags corresponding to the encoding vector; wherein, the position codes are used to indicate the order relationship of each element in the encoding vector, and the modal tags are used to indicate the data type of the input data corresponding to the encoding vector; embedding the position codes and modal tags into the encoding vector to obtain the mapped encoding vector.
[0163] In an optional embodiment, the input data includes data of at least two data types; after the step of mapping the encoded vector to a vector space of a preset dimension to obtain the first vector corresponding to the encoded vector, the method further includes: performing feature fusion processing on the first vector corresponding to the input data of at least two data types to obtain the first vector after feature fusion, so that the first vector after feature fusion can be decoded to obtain decoded information.
[0164] In an optional embodiment, the step of performing feature fusion processing on the first vector corresponding to input data of at least two data types to obtain the first vector after feature fusion includes: determining the fusion weight corresponding to the first vector of each of the at least two data types; multiplying the first vector of each data type with the corresponding fusion weight to obtain the multiplication result corresponding to each data type; and adding the multiplication results corresponding to each data type to obtain the first vector after feature fusion.
[0165] In an optional embodiment, the above-described step of decoding the first vector to obtain decoded information includes: inputting the first vector into a pre-trained decoder, obtaining the correlation between elements in the first vector through the decoder, and encoding the corresponding position of the first vector based on the correlation and the position of the first vector to obtain the decoded information corresponding to the first vector.
[0166] In an optional embodiment, the step of obtaining the decoding information corresponding to the first vector based on the association relationship and the position encoding corresponding to the first vector includes: dividing the first vector into multiple sub-vectors according to the position encoding corresponding to the first vector; for the multiple sub-vectors, generating the decoding information corresponding to the current sub-vector based on the association relationship between the elements in the current sub-vector using a decoder, and generating the decoding information corresponding to the next sub-vector of the current sub-vector based on the decoding information corresponding to the current sub-vector and the association relationship between the elements in the next sub-vector of the current sub-vector; and combining the decoding information corresponding to the multiple sub-vectors to obtain the decoding information corresponding to the first vector.
[0167] In an optional embodiment, before the step of converting the decoded information into a continuous audio signal by the speech generator, the method further includes: denoising and refining the decoded information to obtain high-fidelity decoded information, so that the speech generator can perform speech generation processing on the high-fidelity decoded information.
[0168] In an optional embodiment, the above method further includes: adjusting the audio signal based on preset input features to obtain the final audio signal; wherein the input features include at least one of the following: emotional characteristics, style characteristics, opening characteristics, ending characteristics, and sound quality characteristics.
[0169] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal device, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0170] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0171] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A data conversion method characterized by, The method comprises: obtaining input data, and performing encoding processing on the input data to obtain an encoded vector; wherein the input data comprises at least one of the following data: text data, audio data, and video data; mapping the encoded vector to a vector space of a preset dimension to obtain a first vector corresponding to the encoded vector; performing decoding processing on the first vector to obtain decoding information; wherein the decoding information is used to indicate semantic information corresponding to the input data; converting the decoding information into continuous audio signals through a voice generator; wherein the audio signals match the semantic information corresponding to the input data; the step of mapping the encoded vector to a vector space of a preset dimension to obtain a first vector corresponding to the encoded vector comprises: generating position encoding and modal labels corresponding to the encoded vector; wherein the position encoding is used to indicate the order relationship of each element in the encoded vector, and the modal labels are used to indicate the data type of the input data corresponding to the encoded vector; embedding the position encoding and the modal labels into the encoded vector to obtain the mapped encoded vector; multiplying the mapped encoded vector by a target matrix of a preset dimension to map the mapped encoded vector to the vector space of the preset dimension to obtain the first vector.
2. The method of claim 1, wherein, The step of performing encoding processing on the input data to obtain an encoded vector comprises: performing standardization processing on the input data according to a preset rule to obtain standardized input data; wherein the preset rule comprises standardization processing rules corresponding to a plurality of preset data types, and the plurality of preset data types comprise text type, audio type, and video type; performing division processing on the standardized input data to obtain a plurality of divided sub-data; performing encoding processing on the sub-data to obtain an encoded vector corresponding to each sub-data.
3. The method of claim 1, wherein, The input data comprises data of at least two data types; after the step of mapping the encoded vector to a vector space of a preset dimension to obtain a first vector corresponding to the encoded vector, the method further comprises: performing feature fusion processing on the first vectors corresponding to the input data of the at least two data types to obtain a feature-fused first vector, so that subsequent decoding processing is performed on the feature-fused first vector to obtain decoding information.
4. The method of claim 3, wherein, The step of performing feature fusion processing on the first vectors corresponding to the input data of the at least two data types to obtain a feature-fused first vector comprises: determining a fusion weight corresponding to the first vector of each data type in the at least two data types; multiplying the first vector of each data type by the corresponding fusion weight to obtain a multiplication result corresponding to each data type, and adding the multiplication results corresponding to each data type to obtain the feature-fused first vector.
5. The method of claim 1, wherein, The step of performing decoding processing on the first vector to obtain decoding information comprises: The first vector is input into a pre-trained decoder, the decoder is used to obtain the correlation between the elements in the first vector, and the decoding information corresponding to the first vector is obtained based on the correlation and the position encoding corresponding to the first vector.
6. The method of claim 5, wherein, The step of obtaining the decoding information corresponding to the first vector based on the correlation and the position encoding corresponding to the first vector comprises: According to the position encoding corresponding to the first vector, the first vector is divided into a plurality of sub-vectors; For the plurality of sub-vectors, the decoding information corresponding to the current sub-vector is generated based on the correlation between the elements in the current sub-vector by the decoder, and the decoding information corresponding to the next sub-vector of the current sub-vector is generated based on the correlation between the elements in the next sub-vector of the current sub-vector and the decoding information corresponding to the current sub-vector. The decoding information corresponding to the plurality of sub-vectors is combined to obtain the decoding information corresponding to the first vector.
7. The method of claim 1, wherein, Before the step of converting the decoding information into continuous audio signals by the voice generator, the method further comprises: The decoding information is denoised and refined to obtain high-fidelity decoding information, so that the voice generator performs voice generation processing on the high-fidelity decoding information.
8. The method of claim 1, wherein, The method further comprises: Based on the preset input features, the audio signal is adjusted to obtain a final audio signal; wherein the input features include at least one of the following: emotional characteristics, style features, opening features, ending features and sound quality features.
9. A data conversion device, characterized by comprising: The device comprises: A data acquisition module is configured to acquire input data, encode the input data, and obtain an encoded vector; wherein the input data includes at least one of the following data: text data, audio data and video data; A vector mapping module is configured to map the encoded vector to a vector space of a preset dimension to obtain a first vector corresponding to the encoded vector; A vector decoding module is configured to decode the first vector to obtain decoding information; wherein the decoding information is used to indicate semantic information corresponding to the input data; A voice generation module is configured to convert the decoding information into continuous audio signals by a voice generator; wherein the audio signals match the semantic information corresponding to the input data; The vector mapping module is further configured to generate position encoding and modal markers corresponding to the encoded vector; wherein the position encoding is used to indicate the order of each element in the encoded vector, and the modal markers are used to indicate the data type of the input data corresponding to the encoded vector; the position encoding and the modal markers are embedded into the encoded vector to obtain a mapped encoded vector; the mapped encoded vector is mapped to a vector space of a preset dimension by multiplying the mapped encoded vector with a target matrix of a preset dimension to obtain the first vector.
10. An electronic device, comprising: The electronic device comprises a processor and a memory, the memory stores machine executable instructions capable of being executed by the processor, and the processor executes the machine executable instructions to implement the data conversion method in any one of claims 1 to 8.
11. A computer readable storage medium characterized by, The computer readable storage medium stores computer executable instructions, and when the computer executable instructions are called and executed by a processor, the computer executable instructions cause the processor to implement the data conversion method in any one of claims 1 to 8.
Citation Information
Patent Citations
Voice conversion method and device, equipment and storage medium
CN113470664A
Voice recognition method and system based on audio and video dual modes
CN114974215A