Data processing method and apparatus, electronic device, computer readable storage medium and computer program product
By fusing and mapping acoustic and semantic features obtained from interactive information, and using latent representation and diffusion models to generate more accurate speech signals, the problem of inaccurate acoustic representation in existing technologies is solved, and self-supervised learning and efficient speech generation are achieved.
Patent Information
- Application Number
- PCT/CN2025/085706
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-29
- Filing Date
- 2025-03-28
- Publication Date
- 2025-11-06
AI Technical Summary
In interactive generation, existing technologies have low accuracy in acoustic representation of speech signals, especially in interactive dialogues. Reliance on explicit text prompts leads to inaccurate acoustic representation. Furthermore, traditional methods are highly dependent on data annotation and are inefficient, making it difficult to train on a wide range of datasets in a self-supervised manner.
By acquiring historical interaction information and predicting the acoustic and semantic features of interactive text, a fusion mapping process is performed. Then, a latent expression model and a diffusion model are used for denoising to generate more accurate speech signals, achieving self-supervised learning without manual annotation.
It improves the acoustic dimension accuracy of speech generation in interactive dialogues, and the generated speech signals can characterize the semantic features and paralinguistic information of the predicted interactive text, thereby enhancing the user experience and interaction warmth.
Smart Images

Figure CN2025085706_06112025_PF_FP_ABST
Abstract
Description
Data processing method and device, electronic equipment, computer readable storage medium and computer program product
[0001] Cross-reference to related applications
[0002] The present application is based on and claims priority to Chinese Patent Application No. 2024105332702, filed on April 29, 2024, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD
[0003] The present application relates to the field of artificial intelligence, and in particular, to a data processing method and device, electronic equipment, computer readable storage medium and computer program product. BACKGROUND
[0004] Artificial intelligence (AI) is a comprehensive technology of computer science, which studies the design principles and implementation methods of various intelligent machines to enable machines to have perception, reasoning and decision-making functions. Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, such as natural language processing technology and machine learning / deep learning. With the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0005] Interactive generation is increasingly valued, for example, related technologies provide a multi-functional, user-friendly artificial intelligence chat robot platform that integrates large language models to improve the creativity of dialogue automation and support the creation of highly personalized and engaging chat experiences, but the solutions in related technologies pay less attention to the accuracy of acoustic expression of predicted voice signals in interactive dialogue processes. Even if the acoustic expression of the predicted voice signal is concerned, it is based on explicit text cues (e.g., happy, angry, etc.) to control the acoustic expression, thereby making the accuracy of the acoustic expression low. SUMMARY
[0006] The embodiments of the present application provide a data processing method, device, electronic equipment, computer readable storage medium and computer program product, which can predict the acoustic quality of voice signals corresponding to interactive text.
[0007] The technical solutions of the embodiments of the present application are implemented as follows:
[0008] The embodiments of the present application provide a data processing method, which is executed by an electronic device, comprising:
[0009] Obtaining historical interaction information and predicted interactive text corresponding to the historical interaction information;
[0010] extract a first acoustic feature of the historical interaction information and a first semantic feature of the historical interaction information, and extract a second semantic feature of the predicted interaction text;
[0011] perform fusion mapping processing on the first acoustic feature, the first semantic feature, and the second semantic feature to obtain a first paralinguistic feature;
[0012] perform denoising processing on initial noise based on the second semantic feature and the first paralinguistic feature to obtain a second acoustic feature of the predicted interaction text;
[0013] generate a voice signal corresponding to the predicted interaction text based on the second acoustic feature.
[0014] Embodiments of the present application provide a data processing apparatus, comprising:
[0015] an acquisition module configured to acquire historical interaction information and predicted interaction text corresponding to the historical interaction information;
[0016] an extraction module configured to extract a first acoustic feature of the historical interaction information and a first semantic feature of the historical interaction information, and extract a second semantic feature of the predicted interaction text;
[0017] a paralinguistic module configured to perform fusion mapping processing on the first acoustic feature, the first semantic feature, and the second semantic feature to obtain a first paralinguistic feature;
[0018] a denoising module configured to perform denoising processing on initial noise based on the second semantic feature and the first paralinguistic feature to obtain a second acoustic feature of the predicted interaction text;
[0019] a prediction module configured to generate a voice signal corresponding to the predicted interaction text based on the second acoustic feature.
[0020] Embodiments of the present application provide a data processing method, which is executed by an electronic device and comprises:
[0021] acquire a plurality of historical interaction information samples, and extract interaction context information samples and to-be-predicted information samples from the plurality of historical interaction information samples;
[0022] extract a third acoustic feature of the interaction context information samples and a fourth semantic feature of the interaction context information samples, and extract a fifth semantic feature of the to-be-predicted information samples;
[0023] perform the following processing by a latent expression model: perform fusion mapping processing based on the third acoustic feature, the fourth semantic feature, and the fifth semantic feature to obtain a second paralinguistic feature;
[0024] The following processing is performed by a diffusion model: based on the fifth semantic feature and the second paralinguistic feature, denoising processing is performed on initial noise to obtain a fourth acoustic feature of the to-be-predicted information sample;
[0025] Based on the fourth acoustic feature and a labeled acoustic feature of the to-be-predicted information sample, an acoustic loss is determined;
[0026] Based on the acoustic loss, parameters of the latent expression model and parameters of the diffusion model are updated to obtain an updated latent expression model and an updated diffusion model.
[0027] Embodiments of the present application provide a data processing apparatus, which comprises:
[0028] The second acquisition module is configured to acquire a plurality of historical interaction information samples, and extract an interaction context information sample and a to-be-predicted information sample from the plurality of historical interaction information samples;
[0029] The second extraction module is configured to extract a third acoustic feature of the interaction context information sample and a fourth semantic feature of the interaction context information sample, and extract a fifth semantic feature of the to-be-predicted information sample;
[0030] The second paralanguage module is configured to perform the following processing by a latent expression model: based on the third acoustic feature, the fourth semantic feature and the fifth semantic feature, fusion mapping processing is performed to obtain a second paralanguage feature;
[0031] The second denoising module is configured to perform the following processing by a diffusion model: based on the fifth semantic feature and the second paralanguage feature, denoising processing is performed on initial noise to obtain a fourth acoustic feature of the to-be-predicted information sample;
[0032] The loss module is configured to determine an acoustic loss based on the fourth acoustic feature and a labeled acoustic feature of the to-be-predicted information sample;
[0033] The training module is configured to update parameters of the latent expression model and parameters of the diffusion model based on the acoustic loss to obtain an updated latent expression model and an updated diffusion model.
[0034] Embodiments of the present application provide an electronic device, which comprises:
[0035] The memory is configured to store computer executable instructions;
[0036] The processor is configured to execute the computer executable instructions stored in the memory to implement the data processing method provided by the embodiments of the present application.
[0037] An embodiment of the present application provides a computer readable storage medium, which stores computer executable instructions, and when the computer executable instructions are executed by a processor, a data processing method provided by the embodiment of the present application is implemented.
[0038] An embodiment of the present application provides a computer program product, which comprises computer executable instructions, and when the computer executable instructions are executed by a processor, a data processing method provided by the embodiment of the present application is implemented.
[0039] The embodiment of the present application has the following beneficial effects:
[0040] By the embodiment of the present application, the historical interaction information and the predicted interaction text corresponding to the historical interaction information are obtained, the first acoustic feature of the historical interaction information and the first semantic feature of the historical interaction information are extracted, and the second semantic feature of the predicted interaction text is extracted. Here, the semantic and acoustic features of the historical interaction information can be obtained, and the semantic feature of the predicted interaction text can also be obtained. The first paralanguage feature is obtained by fusion mapping processing based on the first acoustic feature, the first semantic feature and the second semantic feature. The first paralanguage feature is different from the explicit feature, and can more flexibly and variously represent paralanguage information. The second acoustic feature of the predicted interaction text is obtained by denoising processing of the initial noise based on the second semantic feature and the first paralanguage feature, so that the second acoustic feature can carry more accurate paralanguage information. The voice signal corresponding to the predicted interaction text is generated based on the second acoustic feature. The finally generated voice signal can represent the semantic feature of the predicted interaction text, and can also represent the paralanguage information that the predicted interaction text should have, thereby improving the accuracy of voice generation in the acoustic dimension in the interactive dialogue. BRIEF DESCRIPTION OF DRAWINGS
[0041] FIGS. 1A-1D are structural schematic diagrams of a voice generation system provided by the related art;
[0042] FIG. 2 is a structural schematic diagram of a voice generation system provided by an embodiment of the present application;
[0043] FIG. 3 is an application schematic diagram of a data processing method provided by an embodiment of the present application;
[0044] FIG. 4 is a structural schematic diagram of a data processing apparatus provided by an embodiment of the present application;
[0045] FIGS. 5A-5D are flow schematic diagrams of a data processing method provided by an embodiment of the present application;
[0046] FIG. 6 is a principle schematic diagram of a data processing method provided by an embodiment of the present application;
[0047] FIG. 7 is a first architecture schematic diagram of a data processing method provided by an embodiment of the present application;
[0048] Fig. 8 is a second architecture schematic diagram of the data processing method provided by the embodiments of the present application;
[0049] Fig. 9 is a third architecture schematic diagram of the data processing method provided by the embodiments of the present application;
[0050] Fig. 10 is a structure schematic diagram of a potential expression model of the data processing method provided by the embodiments of the present application;
[0051] Fig. 11 is a structure schematic diagram of a first diffusion model of the data processing method provided by the embodiments of the present application;
[0052] Fig. 12 is a structure schematic diagram of a second diffusion model of the data processing method provided by the embodiments of the present application. DETAILED DESCRIPTION
[0053] In order to make the objects, technical solutions and advantages of the present application clearer, the following will further describe the present application in conjunction with the accompanying drawings, and the described embodiments should not be regarded as limiting the present application, and all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0054] In the following description, "some embodiments" are related to a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subset of all possible embodiments, and can be combined with each other without conflict.
[0055] In the following description, the terms "first\second\third" are only to distinguish similar objects, and do not represent a specific order of the objects, and it can be understood that "first\second\third" can be interchanged with a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application.
[0057] Before the embodiments of the present application are further described in detail, the terms and phrases involved in the embodiments of the present application are explained, and the terms and phrases involved in the embodiments of the present application are applicable to the following explanations.
[0058] 1) Large Language Model (LLM): It is a deep learning-based natural language processing model with strong text generation and understanding capabilities. LLM can learn rich language knowledge and reasoning ability through training on large-scale corpus, so it can generate fluent and natural text and try to answer various language processing tasks. In recent years, with the continuous development of technology, LLM has made significant progress in natural language processing and has shown broad application prospects in many fields. Therefore, LLM refers to a large language model with strong text generation and understanding capabilities.
[0059] 2) Embedding: Through Embedding technology, high-dimensional vectors can be converted to relatively low-dimensional space, making machine learning easier and more efficient. It is mainly used in natural language processing and machine learning fields. It refers to converting a high-dimensional sparse vector (such as one-hot encoding in bag-of-words model) into a low-dimensional dense real number vector. This process is also called word embedding or vector embedding, which can encode semantic information into the low-dimensional vector. This conversion is usually completed through neural network training. In practical applications, Embedding can help machines better understand text data and improve the efficiency and accuracy of natural language processing.
[0060] 3) In-context learning (ICL): It is a special prompt engineering method. It gives the model a demonstration of the task within the natural language prompt template. Through in-context learning, a ready-made large language model can be directly used to solve new tasks without fine-tuning the large language model for new tasks.
[0061] 4) Paralinguistic latent expression: Paralinguistic latent expression is used to represent the emotional intent of speech, etc. Unlike explicit expression of emotions such as joy, anger, sadness, happiness, and neutrality, paralinguistic latent expression is a feature of emotion in the latent space that can express more actual emotional intent that cannot be directly expressed with emotional labels.
[0062] The related art provides a large text-to-speech model NaturalSpeech2, which uses an audio neural network codec with a residual vector quantizer to obtain quantized latent vectors of speech, and then uses a diffusion model to generate the latent vectors conditioned on text input. A speech cue mechanism is designed in NaturalSpeech2 to enable context learning for the diffusion model and duration and pitch predictors, supporting diverse and zero-shot speech generation. The architecture diagram of NaturalSpeech2 can be seen in FIG. 1A, which includes an encoder, a decoder, and a latent diffusion model, and the prior condition is calculated by a phoneme encoder and a duration predictor. The speech cue mechanism in NaturalSpeech2 is a method for context learning of the duration / pitch predictor and the latent diffusion model. As shown in FIG. 1B, in the training process, a random segment in the target speech is used as a speech cue, and the diffusion model is only used to predict the remaining part. In the inference process, a reference speech of a certain speaker (more accurately, a latent representation obtained by processing the original speech signal by a speech encoder) is given as a cue.
[0063] The related art also provides a large multilingual translation generation model SeamlessExpressive, as shown in FIG. 1C, which is based on a large multilingual translation generation model SeamlessM4Tv2, and internally combines an expressivity encoder to analyze the translation source speech to obtain appropriate rhythm, speed, pause, etc. to guide the target speech generation. SeamlessExpressive uses a module to generate mel filter bank output for a vocoder. As shown in FIG. 1D, the module decouples the semantic component and the expressive component of the source speech through an unsupervised speech reconstruction pre-training process. Then, SeamlessExpressive conditions on the translated source speech audio signal to transfer its expressive components (intonation, emotional expression, and voice style) to the translated target speech.
[0064] Naturalspeech2 and SeamlessExpressive both rely on pre-trained models and pre-processing algorithms to calculate several predefined expressive-related features as learning objectives. Specifically, the prior conditions of Naturalspeech2 are calculated by a phoneme encoder and a duration / pitch predictor, and the true duration and pitch information are calculated in advance and used as learning objectives to train the duration / pitch predictor. In the training process of the Naturalspeech2 generation model, the true duration and pitch information are used as prior conditions, while in the inference process, the duration and pitch information predicted by the duration / pitch predictor are used as prior conditions. This leads to the following disadvantages: 1) the prior condition information is mismatched between the training and inference processes; 2) several prior features related to speech expression need to be predefined and limited; 3) based on these predefined prior features, the model relies on special pre-processing algorithms to extract prior features related to speech expression, such as pitch extraction, speech classification, intensity estimation, etc., so the performance of the model is limited by the performance of these pre-processing algorithms; 4) some pre-processing algorithms require high-quality speech, such as pitch estimation algorithms, which to some extent makes the system lack robustness to speech data with background sound and noise, thus greatly limiting the expandability of the training data.
[0065] SeamlessExpressive also has the above-mentioned disadvantages, and is limited to translation scenarios, that is, it relies on the source speech audio signal of translation as a condition to migrate its intonation, emotional expression and voice style to the translation target speech. In this process, not only the sentence-level rhythm and intonation of the source speech need to be preserved, but also the semantic-level rhythm (such as pauses) need to be preserved, so training the rhythm identification module requires the speech training data to be aligned in terms of rhythm, and the fundamental frequency, voiced / unvoiced, segmentation / energy, etc. information needs to be annotated by humans or pre-processing algorithms, which greatly limits the acquisition of available training data.
[0066] The expressiveness of the speech generation model in the related art is very dependent on data annotation. As the model size becomes larger and larger, the demand for large-scale data also becomes larger and larger, and the problems of traditional annotation methods being expensive, time-consuming and inefficient are increasingly prominent. Unlike the current artificial intelligence paradigm, an excellent actor or director, a professional voice actor, while reading the script, can infer a lively speech expression that matches it. People can not only understand what the dialogue says while reading a lively dialogue description in a novel, but also imagine how the characters express the dialogue in their minds.
[0067] In view of the above-mentioned shortcomings and limitations, the Naturalspeech2 and Seamless Expressive schemes are still difficult to break through the limitations of the existing artificial intelligence paradigm and are difficult to perform self-supervised training on a large number of publicly available broadcast, film and television programs, short video and short drama, Internet audio data sets and the like, which are the problems that the technical solutions provided in the embodiments of the present application are committed to solve.
[0068] Inspired by this, the embodiments of the present application aim to overcome the limitations of the above-mentioned traditional artificial intelligence paradigm, and can perform self-supervised training on a large number of publicly available broadcast, film and television programs, short video and short drama, Internet audio data sets and the like, and do not require any artificial annotation of paralinguistic annotations or speech expression tags. The technical solutions provided in the embodiments of the present application not only solve the problems of related technical solutions, such as high cost, time-consuming, low efficiency and dependence on traditional annotation, but more importantly, the technical solutions provided in the embodiments of the present application no longer rely on a few data annotation personnel, and no longer rely on the rigid, inaccurate, incomplete, single-dimensional (for example, limited to the emotional dimension), oversimplified (for example, limited to 5 to 11 categories such as joy, anger, sadness, happiness, and neutral) and biased artificial annotation based on a predefined annotation vocabulary, on the contrary, the technical solutions provided in the embodiments of the present application learn potential expressions in a completely self-supervised manner.
[0069] The embodiments of the present application provide a data processing method, device, electronic equipment, computer readable storage medium and computer program product, which can improve the acoustic quality of the predicted interaction text corresponding to the speech signal.
[0070] The data processing method provided in the embodiments of the present application can be implemented by the terminal / server alone; or can be implemented by the terminal and the server cooperatively, for example, the server alone bears the data processing method below, or the terminal sends the plurality of historical interaction information to the server, and the server performs the data processing method according to the received plurality of historical interaction information.
[0071] The data processing method provided in the embodiments of the present application can be implemented by the terminal / server alone; or can be implemented by the terminal and the server cooperatively, for example, the terminal alone bears the data processing method below, or the terminal sends the data processing request (carrying the plurality of historical interaction information) to the server, and the server performs the data processing method according to the received data processing request.
[0072] Referring to FIG. 2, FIG. 2 is an architecture schematic diagram of a data processing system provided by the embodiments of the present application, the terminal 400 is connected to the server 200 through the network 300, and the network 300 can be a wide area network or a local area network, or a combination of the two.
[0073] The terminal 400 can be used to obtain a data processing request, and the data processing request specifically refers to a voice signal requesting to generate a predicted interaction text based on historical interaction information in an interactive dialogue. The terminal 400 sends the data processing request to the server 200, and the plurality of historical interaction information can be obtained directly by the terminal 400 and carried to the data processing request. The server 200 obtains the predicted interaction text corresponding to the historical interaction information; extracts the first acoustic feature of the historical interaction information and the first semantic feature of the historical interaction information, and extracts the second semantic feature of the predicted interaction text; performs fusion mapping processing based on the first acoustic feature, the first semantic feature and the second semantic feature to obtain the first paralanguage feature; based on the second semantic feature and the first paralanguage feature, the initial noise is denoised to obtain the second acoustic feature of the predicted interaction text; and based on the second acoustic feature, the voice signal of the corresponding predicted interaction text is generated. The server 200 sends the voice signal to the terminal 400, and outputs the voice signal through the terminal 400.
[0074] The technical scheme provided by the embodiment of the application can be applied to various projects and product applications including an artificial intelligence chat robot platform, an audio and video conference system, a virtual human voice assistant, a vehicle-mounted voice interaction system and the like. The embodiment of the application provides a method for learning paralanguage hidden representation and a multi-round interactive generation device. The method can utilize the capability of a large-scale language model, and utilize unsupervised learning on massive unlabeled voice interaction data to provide voice representation with strong generalization, fine granularity, robustness and universality, thereby improving the performance of product applications and enhancing user experience.
[0075] The following is an example of applying the technical scheme provided by the embodiment of the application to multiple scenarios:
[0076] In the application scenario of intelligent customer service, the user feeds back a question (such as "Why is my order delayed?") through the intelligent customer service terminal 400. The terminal automatically obtains the history of the last three rounds of dialogue. The server 200 extracts the speech speed fluctuation (first acoustic feature) and complaint keywords (first semantic feature) when the user asks a question, generates a predicted reply "being processed urgently, and expected to be delivered within 24 hours" (second semantic feature), performs fusion mapping processing on the above features to obtain a paralanguage feature (such as a hidden space feature representing a neutral emotion) that can represent the attitude of a customer service personnel, and finally generates a synthesized voice with a soothing emotion but maintaining professionalism. Compared with a mechanical reply, the user complaint rate can be effectively reduced by 40%.
[0077] In the application scenario of online education intelligent accompanying practice, the student practices English conversation through the education tablet terminal 400, the terminal records the pronunciation error history, the server extracts the vowel bias (the first acoustic feature) and the grammar error (the first semantic feature) of the student's pronunciation, generates the correct formal reply "pay attention to the pronunciation of / th / needs the tongue tip to resist the teeth" (the second semantic feature), and the fusion mapping processing of the above features obtains the encouraging paralanguage feature (such as the hidden space feature representing the encouraging emotion) which can represent the teacher style, and generates the guiding voice which not only corrects the error but also maintains the encouragement. The pronunciation accuracy of the student is improved by 25%, and the practice enthusiasm is improved.
[0078] In the application scenario of medical inquiry voice assistant, the patient describes the symptoms through the hospital terminal 400, the system records the anxious tone feature, extracts the rapid breathing sound (the first acoustic feature) and the keyword "chest pain" (the first semantic feature), generates the suggestion "suggest to perform electrocardiogram examination immediately" (the second semantic feature), and the fusion mapping processing of the above features obtains the authoritative paralanguage feature (such as the hidden space feature representing the serious emotion) which can represent the chief physician, and generates the medical order voice which is not only professional and serious but also has a calming effect. The patient's medical order compliance is improved by 30%, and the communication efficiency is improved by 50%.
[0079] The above application scheme can realize the leap from mechanical response to personification interaction, significantly improve the interaction temperature while maintaining the semantic accuracy.
[0080] Fig. 3 shows the application architecture of the technical solutions provided by the embodiments of the present application as functional modules in an artificial intelligence chat robot platform and product. Next, the most typical artificial intelligence chat robot platform is introduced. The artificial intelligence chat robot platform is an indispensable tool for modern enterprises, which combines automation, efficiency and personalized customer experience. Leading artificial intelligence platforms and products provide a wealth of choices to meet the needs of every user and business entering the digital age. Botpress, as a versatile and user-friendly artificial intelligence chat robot platform, integrates large language models to improve the creativity of dialogue automation, and supports the creation of highly personalized and engaging chat experiences; Dialogflow is also an artificial intelligence chat robot platform that provides convenience for 24 / 7 customer self-service through virtual agents and interactive voice response systems, and can handle routine tasks and query interactions; Amazon Lex uses voice and text to build conversational interfaces in applications, using advanced deep learning features such as automatic speech recognition and natural language understanding to create engaging user experiences with realistic interactions; Yellow.ai uses advanced technologies such as natural language processing and generative artificial intelligence to provide empathetic and personalized interactions. The platform provides rapid deployment through pre-built templates and integrations to meet the needs of different industries; Gupshup is a comprehensive conversation engagement platform designed to enhance interactions between businesses and customers through various information channels, and it integrates artificial intelligence to create human-like conversations aimed at improving customer satisfaction and driving business growth. Gupshup's powerful natural language understanding models and sentiment analysis tools also enable the creation of responsive, intuitive and compassionate conversational agents, making Gupshup a versatile tool for businesses seeking to enhance customer engagement and personalize communication across multiple touchpoints.
[0081] The electronic device for performing the data processing method provided by the embodiments of the present application can be various types of terminal devices or servers. In some embodiments, the server 200 can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, etc. Basic cloud computing services. The terminal 400 can be a smartphone, tablet computer, notebook computer, desktop computer, smart speaker, smart watch, smart voice interaction device, smart home appliance, vehicle-mounted terminal, flight, etc., but is not limited thereto. Terminals and servers can be connected directly or indirectly through wired or wireless communication methods, and the embodiments of the present application do not make any limitations.
[0082] In some embodiments, a terminal or a server can implement the data processing method provided in the embodiments of the present application by running a computer program. For example, the computer program can be a native program or a software module in an operating system; can be a native application program (APP), that is, a program that needs to be installed in an operating system to run, such as an instant messaging APP; can also be a mini-program, that is, a program that only needs to be downloaded into a browser environment to run; and can also be a mini-program that can be embedded into any APP. In summary, the above computer program can be any form of application program, module or plug-in.
[0083] The embodiments of the present application can be applied to various scenarios, including but not limited to artificial intelligence scenarios. Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include, for example, sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-training model technology, operation / interaction system, mechatronics, etc. Among them, the pre-training model is also called large model or basic model, which can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.
[0084] With the research and progress of artificial intelligence technology, artificial intelligence technology has been researched and applied in many fields, such as common smart home, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned vehicles, autonomous vehicles, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, intelligent medical care, intelligent customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0085] Referring to FIG. 4, FIG. 4 is a structural schematic diagram of a terminal 400 provided by an embodiment of the present application. The terminal 400 shown in FIG. 4 includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in the terminal 400 are coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between the components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, all the buses are marked as the bus system 440 in FIG. 4.
[0086] The processor 410 can be an integrated circuit chip that has a processing capability of signals, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc., wherein the general-purpose processor can be a microprocessor or any conventional processor.
[0087] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432 that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0088] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 optionally includes one or more storage devices physically located in proximity to the processor 410.
[0089] The memory 450 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. Non-volatile memory can be read only memory (ROM), and volatile memory can be random access memory (RAM). The memory 450 described in embodiments of the present application is intended to include any suitable type of memory.
[0090] In some embodiments, the memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or a subset or superset thereof, which are exemplarily illustrated below.
[0091] The operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0092] The network communication module 452 is used to reach other computing devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including Bluetooth, wireless compatibility certification (WiFi), and universal serial bus (USB), etc.
[0093] a presentation module 453 for enabling presentation of information (e.g., a user interface for operating a peripheral device and displaying content and information) via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430;
[0094] an input processing module 454 for detecting and translating one or more user inputs or interactions from one or more input devices 432.
[0095] In some embodiments, the apparatus provided by the embodiments of the present application can be implemented in software. FIG. 4 shows a data processing apparatus 455-1 stored in the memory 450, which can be software in the form of programs and plug-ins, etc., including the following software modules: a first acquisition module 4551, a first extraction module 4552, a first auxiliary language module 4553, a first denoising module 4554, and a prediction module 4555. FIG. 4 shows a data processing apparatus 455-2 stored in the memory 450, which can be software in the form of programs and plug-ins, etc., including the following software modules: a second acquisition module 4556, a second extraction module 4557, a second auxiliary language module 4558, a second denoising module 4559, a loss module 45510, and a training module 45511. These modules are logical, and thus can be combined or further split according to the implemented functions. The functions of the various modules will be described below.
[0096] As mentioned above, the method of data processing provided by the embodiments of the present application can be implemented by various types of electronic devices. Referring to FIG. 5A, FIG. 5A is a flowchart of the method of data processing provided by the embodiments of the present application, which is explained in combination with steps 101 to 105 shown in FIG. 5A.
[0097] In step 101, historical interaction information and predicted interaction text corresponding to the historical interaction information are acquired.
[0098] As an example, referring to FIG. 7, the historical interaction information here is the conversation that has occurred, for example, when the value of l in FIG. 7 is 4, the interaction information of the first round to the fourth round is the historical interaction information, and the fifth round interaction information needs to be inferred based on the previous four rounds of historical interaction information. The text content of the fifth round interaction information is the predicted interaction text. Here, a large language model can be used to predict the text content of the fifth round of conversation based on the conversation that has occurred (historical interaction information). The historical interaction information involved here can be in text form or in voice form.
[0099] As an example, the following is a complete technical solution for predicting the next round of conversation text based on historical interaction information by a large language model (LLM). First, data preprocessing and context modeling are performed to serialize the conversation history into a formatted Prompt, multi-dimensional feature extraction is performed, and sequence-based text prediction is performed on the multi-dimensional features, thereby outputting the next round of conversation text. In step 102, the first acoustic feature of the historical interaction information and the first semantic feature of the historical interaction information are extracted, and the second semantic feature of the predicted interaction text is extracted.
[0100] As an example, for historical interaction information, if the historical interaction information itself is speech, the acoustic feature x 1,...,l here can be extracted using a mel filter bank. If the historical interaction information here is text, the acoustic feature here can be predicted based on the historical interaction information corresponding to the text using the method provided by the embodiments of the present application (steps 101 to 104 are performed), for example, the acoustic feature of the third round of historical interaction information (text form) is predicted based on the first and second rounds of historical interaction information. For the first round of historical interaction information, if the first round of historical interaction information is in text form, the acoustic feature of the first round of historical interaction information can be predicted by a large language model on the text content of the first round of historical interaction information, or randomly initialized.
[0101] As an example, the semantic feature is converted by a tokenizer (Tokenizer Model) for each historical interaction information, which can convert the original audio segment or text segment into discrete unit token representation u. More specifically, as shown in FIG. 7, the tokenizer (Tokenizer Model) provided by the embodiments of the present application adopts the Multitask-UnitY2 module (hereinafter referred to as UnitY) in the large language translation model Seamless, which converts speech or text into discrete unit token representation, thereby reducing the dimension and achieving higher quality speech generation. That is, the original audio or text sequence is converted into discrete unit token representation If multiple languages need to be supported, language discrete representation can be combined to identify the language used in the conversation, as shown in FIG. 7, where eng represents English.
[0102] As an example, the first semantic feature u 1,...,l is obtained based on the audio segment or the text segment, and the second semantic feature u l+1 is obtained based on the text segment.
[0103] As an example, the specific implementation process of the Tokenizer Model can refer to the UNITY2 module architecture in the large multi-language translation generation model Seamless to implement the application and training process, and use a pre-trained multi-language self-supervised learning large model to obtain the UnitY2 module (the embodiments of the present application also completely support using any module that converts from the voice or text modal to discrete representation as an alternative solution).
[0104] As an example, the semantic feature is a discrete sequence feature for representing semantic information, and the acoustic feature is a discrete sequence feature for representing acoustic information.
[0105] In step 103, the first acoustic feature, the first semantic feature, and the second semantic feature are fused and mapped to obtain a first paralanguage feature.
[0106] As an example, the fusion and mapping process refers to splicing the first acoustic feature, the first semantic feature, and the second semantic feature, and then performing mapping processing on the spliced feature to map the spliced feature from the original high-dimensional sequence feature space to a low-dimensional latent space or hidden space to obtain a vector representation type feature. The first paralanguage feature is a vector representation of the latent space including paralanguage information (paralanguage latent expression), and the paralanguage latent expression is used to represent the emotional intent of the voice and other information. Unlike explicit expressions such as joy, anger, sadness, happiness, and neutrality as emotional labels, the paralanguage latent expression is an emotional feature in the latent space, which can express more actual emotional intent that cannot be directly represented by emotional labels.
[0107] As an example, step 103 is implemented by calling a latent expression model. The input of the latent expression model (LEM) includes the acoustic feature (the first acoustic feature) of the interactive dialogue context (historical interaction information) and the unit token representation u 1,...,l (the first semantic feature) from the Tokenizer Model, and the unit token representation u l+1 (the second semantic feature) of the predicted interactive text. The discrete unit token representation u mainly contains language semantic information, but lacks paralanguage information related to expression. Therefore, the embodiments of the present application propose a latent expression model LEM for predicting the paralanguage latent expression (the first paralanguage feature) between the interactive segments, D is the vector space dimension of the paralanguage latent expression, L is the number of interaction information, and T is the frame sequence length of the discrete unit token representation.
[0108] In step 104, based on the second semantic feature and the first paralinguistic feature, the initial noise is denoised to obtain the second acoustic feature of the predicted interaction text.
[0109] As an example, step 104 can be implemented by calling a paralinguistic token acoustic model (Para-Token Acoustic Model), which predicts the acoustic feature of the next interaction segment based on the paralinguistic latent representation e (the first paralinguistic feature) and the unit token representation u (the second semantic feature) of the next interaction segment as conditions.
[0110] In step 105, a speech signal corresponding to the predicted interaction text is generated based on the second acoustic feature.
[0111] As an example, a high-fidelity generative adversarial vocoder (HiFi-GAN vocoder) synthesizes a speech signal based on the acoustic feature of the next interaction segment, and the embodiments of the present application also do not exclude using any vocoder module from acoustic representation to synthesized speech as an alternative.
[0112] In some embodiments, the predicted interaction text of the historical interaction information is attitude predicted by a language model to obtain a response attitude text of the predicted interaction text; a third semantic feature of the response attitude text is extracted; and correspondingly, in step 104, the initial noise is denoised based on the second semantic feature and the first paralinguistic feature to obtain the second acoustic feature of the predicted interaction text, which can be implemented by the following technical solution: the initial noise is denoised based on the second semantic feature, the third semantic feature, and the first paralinguistic feature to obtain the second acoustic feature of the predicted interaction text.
[0113] As an example, based on the first semantic feature and the second semantic feature, the predicted interaction text of the historical interaction information can also be attitude predicted by a language model to obtain a response attitude text of the predicted interaction text, and then a tokenizer (Tokenizer Model) extracts a third acoustic feature (meta unit token representation) of the response attitude text. A large language model infers the expression information of the current predicted interaction text based on the historical interaction information, where the current predicted interaction text refers to "what to say", and the expression information refers to "how to say". The unit token representation u and the meta unit token representation of the predicted interaction text are extracted by the UnitY2 module described in the foregoing, and are input into the subsequent LEM model. l+1The representation says what, and the meta-language unit word represents how it says what it says. The large language model, as an agent, not only infers the text content of the currently predicted interactive text, but also infers the expressive information of the currently predicted interactive text, and combines it with LEM to predict the latent expressions of the para-language (first para-language features).
[0114] The large language model is used as an example to infer expressive information. For instance, the contextual prompt could be designed as "Please analyze this person's response and, in conjunction with the context of the interactive dialogue, describe how they expressed themselves, including but not limited to their expression style, emotions, tone of voice, speed, and emphasis." In the example given in Figure 9, the interactive text of the current dialogue character A is "I'm sorry," and the corresponding expressive information is "I'm sorry and feel guilty." Next, the expressive information obtained through LLM inference, such as "I'm sorry and feel guilty," is extracted as text by the Unity2 module to obtain meta-unit tokens.
[0115] By performing secondary attitude prediction on the predicted text through a language model, a semantic-attitude dual-dimensional understanding is formed, enabling the generated response text to not only conform to the dialogue logic but also dynamically adapt to the user's emotional state. Multi-level feature fusion enhancement and the synergistic effect of the third and second semantic features achieve content accuracy and emotional adaptability. Combined with paralinguistic features, the emotional expressiveness of speech synthesis is improved.
[0116] This application's embodiments obtain historical interaction information and corresponding predicted interaction text; extract the first acoustic features and first semantic features of the historical interaction information, and extract the second semantic features of the predicted interaction text. This is equivalent to obtaining the semantic and acoustic features of the historical interaction information, as well as the semantic features of the predicted interaction text. Based on the first acoustic features, first semantic features, and second semantic features, a fusion mapping process is performed to obtain the first secondary language feature. This first secondary language feature differs from explicit features and can more flexibly and diversely represent secondary language information. Based on the second semantic features and the first secondary language feature, noise reduction processing is performed on the initial noise to obtain the second acoustic feature of the predicted interaction text. This allows the second acoustic feature to carry more accurate secondary language information. Based on the second acoustic feature, a speech signal corresponding to the predicted interaction text is generated. The final generated speech signal can simultaneously represent the semantic features of the predicted interaction text and also represent the secondary language information that the predicted interaction text should have, improving the accuracy of speech generation in the acoustic dimension during interactive dialogue.
[0117] The following section introduces a technical solution for denoising the initial noise based on the second semantic feature, the third semantic feature, and the first sub-language feature to obtain the second acoustic feature of the predicted interactive text.
[0118] In some embodiments, the denoising processing is implemented by a first diffusion model, the first diffusion model comprising P cascaded multi-semantics denoising networks, P being no less than 2; referring to FIG. 5B, the denoising processing on the initial noise based on the second semantic feature, the third semantic feature and the first paralinguistic feature to obtain the second acoustic feature of the predicted interactive text can be implemented by steps 1041 to 1042 shown in FIG. 5B.
[0119] As an example, the first diffusion model here is a paralinguistic word representation acoustic model, which supports predicting the second acoustic feature y l+1 (second semantic feature) and meta-linguistic unit word representation (third semantic feature) as conditions.
[0120] As an example, the second semantic feature here is used to represent the content semantics of the predicted interactive text (for example, “I am very sorry”), and the third semantic feature is used to represent the attitude semantics of the predicted interactive text (for example, guilt).
[0121] In step 1041, the input of the pth multi-semantics denoising network in the P cascaded multi-semantics denoising networks is subjected to multi-semantics denoising processing, and the pth multi-semantics denoising result output by the pth multi-semantics denoising network is transmitted to the (p+1)th multi-semantics denoising network for continuous multi-semantics denoising processing, to obtain the (P+1)th multi-semantics denoising result corresponding to the (P+1)th multi-semantics denoising network.
[0122] As an example, referring to FIG. 12, the first diffusion model comprises P cascaded multi-semantics denoising networks, which is equivalent to performing P times of multi-semantics denoising processing (since each time of denoising involves the second semantic feature and the third semantic feature, it is called multi-semantics denoising processing), each time of multi-semantics denoising processing is performed according to the multi-semantics denoising result obtained from the previous multi-semantics denoising processing, and then input into the next multi-semantics denoising network for multi-semantics denoising processing.
[0123] As an example, p is an integer variable starting from 1 and increasing, and the value range of P is 1≤p<P, when p is 1, the input of the pth multi-semantics denoising network is the initial noise, the second semantic feature, the third semantic feature and the first paralinguistic feature, and when p is 2≤p<P, the input of the pth multi-semantics denoising network is the (p-1)th multi-semantics denoising result output by the (p-1)th multi-semantics denoising network, the second semantic feature, the third semantic feature and the first paralinguistic feature.
[0124] As an example, taking N as 3 as an example, through the first multi-semantics denoising network, multi-semantics denoising processing is performed on the initial noise (for example, Gaussian noise) based on the second semantic feature, the third semantic feature and the first paralinguistic feature, to obtain a first multi-semantics denoising result, through the second multi-semantics denoising network, multi-semantics denoising processing is performed on the first multi-semantics denoising result based on the second semantic feature, the third semantic feature and the first paralinguistic feature, to obtain a second multi-semantics denoising result, through the third multi-semantics denoising network, multi-semantics denoising processing is performed on the second multi-semantics denoising result based on the second semantic feature, the third semantic feature and the first paralinguistic feature, to obtain a third multi-semantics denoising result. Each multi-semantics denoising result obtained through the above manner can be a speech noise or a latent space encoding of the speech noise, and the multi-semantics denoising processing performed by each multi-semantics denoising network is equivalent to multi-semantics denoising processing of one time step.
[0125] In some embodiments, the multi-semantics denoising processing performed by the p th multi-semantics denoising network in the P cascaded multi-semantics denoising networks on the input of the p th multi-semantics denoising network can be implemented through the following technical solution: fusing the input of the p th multi-semantics denoising network and the cascade identifier corresponding to the p th multi-semantics denoising network to obtain a first fusion result; performing convolution processing on the first fusion result to obtain a first convolution result; fusing the first convolution result, the second semantic feature and the first paralinguistic feature to obtain a second fusion result; fusing the second fusion result and the third semantic feature to obtain a third fusion result; performing activation processing based on a gating mechanism on the third fusion result to obtain a first activation result; fusing the first activation result and the input of the p th multi-semantics denoising network to obtain the output of the p th multi-semantics denoising network. Through the embodiments of the present application, the third semantic feature, the second semantic feature and the first paralinguistic feature can be used as conditions for step-by-step denoising processing, so that the result obtained by each denoising is more capable of realizing ideal acoustic characteristics.
[0126] As an example, when the value of p satisfies the interval condition, the second fusion result and the third semantic feature are fused to obtain a third fusion result, otherwise the second fusion result is directly activated based on a gating mechanism to obtain a first activation result.
[0127] As an example, in the specific implementation of the paralinguistic word representation acoustic model shown in Figure 12, taking a value of p=10 as an example, the input of the 10th multi-semantic denoising network is fused with the concatenated identifier t (where h represents the timestamp 10) corresponding to the 10th multi-semantic denoising network to obtain a first fusion result. The first fusion result is then convolved to obtain a first convolution result (this convolution is dilated convolution). Finally, the first convolution result and the unit word representation u of the interactive text predicting the interactive information are combined. l+1 The second semantic feature and the paralinguistic latent expression e (the first paralinguistic feature) are added together to obtain the second fusion result h. Every two multi-semantic denoising networks (with interval conditions), a feature-wise linear modulation (FiLM) layer is added to the third multi-semantic denoising network to fuse the metalinguistic unit word representations predicted by the LLM agent. The embedding feature c corresponding to the third semantic feature is used to perform activation processing based on a gating mechanism on the third fusion result to obtain the first activation result. Here, activation processing is performed using the tanh activation function and the ReLU activation function respectively. Then, the two activation results are processed based on a gating mechanism to obtain the first activation result. The first activation result is fused with the input of the 10th semantic denoising network to obtain the output of the 10th semantic denoising network. Finally, the input of the 10th semantic denoising network is added to the first activation result to obtain the output of the 10th semantic denoising network.
[0128] In some embodiments, the above-described fusion processing of the second fusion result and the third semantic feature to obtain the third fusion result can be implemented through the following technical solution: embedding the third semantic feature to obtain the embedded feature of the third semantic feature; performing a first linear mapping process on the embedded feature of the third semantic feature to obtain a first linear mapping result, and performing a second linear mapping process on the embedded feature of the third semantic feature to obtain a second linear mapping result; using the first linear mapping result as a scaling parameter and the second linear mapping result as a bias parameter, mapping the second fusion result to obtain the third fusion result. Through the embodiments of this application, the third semantic feature can be mapped to the third fusion result through feature linear modulation, thereby effectively combining it with the second fusion result and improving feature representation capability.
[0129] As an example, the third semantic feature is first embedded to obtain the embedded feature c of the third semantic feature. Then, referring to formulas (1) and (2), the embedded feature c of the third semantic feature is linearly mapped: γ = f1(c)·θ γ(1); β = f2(c)·θ β , (2);
[0130] wherein f1, f2 represent linear mapping functions respectively; γ and β represent scaling parameters and bias parameters respectively; θ γ and θ β represent their corresponding learnable scalar parameters respectively.
[0131] As an example, referring to formula (3), the first linear mapping result γ is taken as a scaling parameter, and the second linear mapping result β is taken as a bias parameter, the second fusion result h is mapped to obtain the third fusion result FiLM(h,c) = (γ+1)·h+β (3).
[0132] wherein FiLM(h,c) is the third fusion result, γ and β represent scaling parameters and bias parameters respectively, and h is the second fusion result.
[0133] In step 1042, the second acoustic feature of the predicted interaction text is determined based on the Pth multi-semantics denoising result output by the Pth multi-semantics denoising network.
[0134] As an example, referring to FIG. 12, the Pth multi-semantics denoising result output by the Pth multi-semantics denoising network is subjected to activation processing based on a relu activation function, and then the activation processing result is subjected to convolution processing to obtain the second acoustic feature of the predicted interaction text
[0135] Through the embodiments of the present application, the parameters in the generation process, such as the second semantic feature, the third semantic feature, and the first paralanguage feature, can be flexibly adjusted, so as to control the diversity and style of the generated second acoustic feature, meet the needs of different users, and inject the second semantic feature, the third semantic feature, and the first paralanguage feature in each denoising generation process, so as to ensure that the generated second acoustic feature can conform to the corresponding semantics and emotion at all times.
[0136] In some embodiments, the output of the first diffusion model can also be input to a post-processing network (PostNet) for post-processing, and the processing result obtained by the post-processing is taken as the second acoustic feature of the predicted interaction text Thus, it can help to compensate for the residual signal that is difficult to capture in the diffusion stage.
[0137] In some embodiments, the fusion mapping processing is implemented by a latent expression model, the latent expression model comprising N cascaded attention networks, N being no less than 2; referring to FIG. 5C, the fusion mapping processing on the first acoustic feature, the first semantic feature and the second semantic feature to obtain the first paralinguistic feature in step 103 can be implemented by steps 1031 to 1033 shown in FIG. 5C.
[0138] As an example, the latent expression model (LEM) here supports predicting the first paralinguistic feature 1,...,l based on the first acoustic feature (Mel feature x 1,...,l , the first semantic feature u l+1 and the second semantic feature u D is the vector space dimension of the paralinguistic latent expression, T is the sequence length of the first acoustic feature, the first semantic feature and the second semantic feature, and L is the total number of interaction information (i.e. the number of historical interaction information plus 1). Referring to FIG. 10, to balance the model capability and the computational complexity, the embodiment of the present application adopts a spectrum-interaction-transformer (SI-transformer) modified based on a spatial-temporal-transformer (ST-transformer) in the specific architecture implementation of the LEM. As shown in FIG. 10, the SI-transformer comprises N spectrum-interaction blocks.
[0139] In step 1031, the first acoustic feature, the first semantic feature and the second semantic feature are spliced to obtain an acoustic spliced feature.
[0140] In step 1032, the input of the nth attention network in the N cascaded attention networks is processed by attention, and the nth attention result output by the nth attention network is transmitted to the (n+1)th attention network for further attention processing to obtain the (n+1)th attention result corresponding to the (n+1)th attention network.
[0141] As an example, referring to FIG. 10, the latent expression model comprises N cascaded attention networks, which is equivalent to performing N times of attention processing, each time being attention processing according to the attention result obtained from the previous attention and then inputting to the next attention network for attention processing.
[0142] As an example, n is an integer variable starting from 1 and increasing, the value range of n is 1≤n<N, when n is 1, the input of the nth attention network is the acoustic concatenation feature, when n is 2≤n<N, the input of the nth attention network is the (n-1)th attention result output by the (n-1)th attention network.
[0143] As an example, taking N as 3 as an example, the acoustic concatenation feature is processed by the first attention network to obtain the first attention result, the first attention result is processed by the second attention network to obtain the second attention result, and the second attention result is processed by the third attention network to obtain the third attention result.
[0144] In some embodiments, the above-mentioned attention processing of the input of the nth attention network in the N cascaded attention networks can be realized by the following technical solutions: performing self-attention processing on the acoustic concatenation feature corresponding to each interaction information to obtain the first self-attention result of the acoustic concatenation feature, wherein the interaction information is the historical interaction information or the predicted interaction text; performing self-attention processing on the first self-attention result of the acoustic concatenation feature corresponding to each feature element to obtain the second self-attention result of the acoustic concatenation feature; and performing full connection processing on the second self-attention result of the acoustic concatenation feature to obtain the nth attention result output by the nth attention network.
[0145] As an example, each spectrum interaction module (attention network) is composed of a spectrum attention (Spectrum attention) layer and an interaction attention (Interaction attention) layer, and then connected with a feed-forward (FFW) layer. The spectrum attention layer performs self-attention calculation on each interaction information based on the spectrum dimension, and the interaction attention layer performs self-attention calculation on the vector sequence between multiple interaction information.
[0146] As an example, the Spectrum attention layer in the SI-transformer only performs self-attention calculation on the input of each interaction information (performs self-attention processing on the acoustic concatenation feature corresponding to each interaction information), that is, performs self-attention calculation on the acoustic concatenation feature of 1×T×(80+D u ) dimension, and D uis the dimension of semantic features, 80 is the dimension of acoustic features, T is the sequence length of semantic features and acoustic features, and since there are L Mel concatenation features, L times of calculation are required; next, the Interaction attention layer only performs self-attention calculation on the input corresponding to each spectrum tile (self-attention processing of the first self-attention result of the acoustic concatenation features corresponding to each feature element), that is, self-attention calculation is performed on an Lx1x1-dimensional sequence, and since there are Tx(80+D u ) sequences, Tx(80+D u ) times of calculation are required. Here, the feature element refers to each feature value in the first self-attention result.
[0147] Through the embodiments of the present application, the calculation complexity mainly lies in the Spectrum attention layer, and is linearly related to the number of interactions L, while the calculation complexity of the classic Transformer is related to the square of L, so the calculation complexity of the attention processing involved in the embodiments of the present application is relatively low, and the content generation task of multiple rounds of interaction can be efficiently supported.
[0148] In step 1033, the first auxiliary language feature is determined based on the Nth attention result output by the Nth attention network.
[0149] As an example, the Nth attention result f o output by the Nth attention network is finally mapped to the first auxiliary language feature , where D is the vector space dimension of the auxiliary language latent expression.
[0150] Through the embodiments of the present application, multi-modal deep fusion can be realized, for example, cross-domain feature jointing: through the concatenation of acoustic features (speech physical properties) and bilingual semantic features (historical content + predicted content), acoustic-semantic cross-modal alignment is realized, the problem of speech and text fragmentation in traditional methods is solved, complementary gain between features can be realized, through the multi-layer attention mechanism, progressive feature extraction can be performed, the first layer of attention focuses on the primary association of acoustic-semantic (such as stress position and keyword matching), the middle layer can mine deep temporal dependencies (such as the relationship between dialogue rhythm and emotion progression), the Nth layer can generate a globally consistent expression (to ensure that the auxiliary language feature is self-consistent within the sentence), each layer of attention automatically adjusts the contribution weight of the three types of features, and the auxiliary language representation is enhanced, through cascaded attention to capture long-distance dependencies, the position accuracy of the emphasized stress is improved, and the cascaded structure can optimize the calculation efficiency, the parameter amount is only 72% of the parallel structure, and the memory occupation is reduced by 40% (through layer-by-layer feature dimension reduction).
[0151] In some embodiments, the denoising processing is implemented through a second diffusion model, the second diffusion model comprising M cascaded single semantic denoising networks, M being no less than 2; and the denoising processing of the initial noise based on the second semantic feature and the first paralinguistic feature in step 104 can be implemented through the following technical solution: the input of the mth single semantic denoising network in the M cascaded single semantic denoising networks is subjected to single semantic denoising processing, and the mth single semantic denoising result output by the mth single semantic denoising network is transmitted to the (m+1)th single semantic denoising network for continuous single semantic denoising processing, to obtain the (m+1)th single semantic denoising result corresponding to the (m+1)th single semantic denoising network; and the second acoustic feature of the predicted interactive text is determined based on the Mth single semantic denoising result output by the Mth single semantic denoising network. Through the embodiments of the present application, the initial noise can be subjected to multiple iterative denoising processing, so as to improve the accuracy of the second acoustic feature.
[0152] As an example, the second diffusion model herein is a paralinguistic word representation acoustic model, which supports predicting a second acoustic feature y l+1 as a condition
[0153] As an example, referring to FIG. 11, the second diffusion model comprises M cascaded single semantic denoising networks, so as to perform M times of single semantic denoising processing (since each time of denoising involves the second semantic feature, it is called single semantic denoising processing), each time of single semantic denoising processing is performed based on the single semantic denoising result obtained from the previous single semantic denoising processing, and then the single semantic denoising result is input into the next single semantic denoising network for single semantic denoising processing.
[0154] As an example, m is an integer variable starting from 1 and increasing, and the value range of M is 1≤m<M, when m is 1, the input of the mth single semantic denoising network is the initial noise, the second semantic feature and the first paralinguistic feature, and when m is 2≤m<M, the input of the mth single semantic denoising network is the (m-1)th single semantic denoising result output by the (m-1)th single semantic denoising network, the second semantic feature and the first paralinguistic feature.
[0155] As an example, taking M as 3 as an example, the initial noise (for example, Gaussian noise) is processed by the first single semantic denoising network based on the second semantic feature and the first paralinguistic feature to obtain a first single semantic denoising result, the first single semantic denoising result is processed by the second single semantic denoising network based on the second semantic feature and the first paralinguistic feature to obtain a second single semantic denoising result, and the second single semantic denoising result is processed by the third single semantic denoising network based on the second semantic feature and the first paralinguistic feature to obtain a third single semantic denoising result. Each single semantic denoising result obtained by the above manner can be a speech noise or a latent space encoding of the speech noise, and the single semantic denoising processing performed by each single semantic denoising network is equivalent to a single semantic denoising processing of one time step.
[0156] In some embodiments, the single semantic denoising processing of the input of the mth single semantic denoising network in the M cascaded single semantic denoising networks can be implemented by the following technical solution: fusing the input of the mth single semantic denoising network and the cascade identifier corresponding to the mth single semantic denoising network to obtain a fourth fusion result; performing convolution processing on the fourth fusion result to obtain a second convolution result; fusing the second convolution result, the second semantic feature, and the first paralinguistic feature to obtain a fifth fusion result; performing activation processing based on a gating mechanism on the fifth fusion result to obtain a second activation result; and fusing the second activation result and the input of the mth single semantic denoising network to obtain the output of the mth single semantic denoising network. Through the embodiments of the present application, the second semantic feature and the first paralinguistic feature can be used as conditions to guide each denoising process, so that each denoising result gradually approaches the ideal second acoustic feature.
[0157] As an example, in the specific implementation of the paralinguistic word representation acoustic model shown in FIG. 11, taking m as 10 as an example, the input of the 10th single semantic denoising network is fused with the cascade identifier t corresponding to the 10th single semantic denoising network (here, h represents that the time stamp is 10) to obtain a fourth fusion result, and the fourth fusion result is processed by convolution to obtain a second convolution result. Here, the convolution processing is a dilated convolution, and then the second convolution result, the unit word representation u l+1The second semantic feature and the paralinguistic latent expression e (the first paralinguistic feature) are added to obtain a fifth fusion result h, and the fifth fusion result is activated based on a gating mechanism to obtain a second activation result, where the activation is performed by a tanh activation function and a relu activation function, respectively, and the two activation results are processed based on a gating mechanism to obtain the second activation result; the second activation result and the input of the 10th single semantic denoising network are fused to obtain the output of the pth single semantic denoising network, and finally the input of the 10th single semantic denoising network and the second activation result are added to obtain the output of the pth single semantic denoising network.
[0158] As mentioned above, the data processing method provided by the embodiments of the present application can be implemented by various types of electronic devices. Referring to FIG. 5D, FIG. 5D is a flowchart of a data processing method provided by an embodiment of the present application, which is described in combination with steps 201 to 206 shown in FIG. 5D.
[0159] In step 201, a plurality of historical interaction information samples are obtained, and interaction context information samples and to-be-predicted information samples are extracted from the plurality of historical interaction information samples.
[0160] As an example, there are 10 interaction segments in the training sample, and if the 5th interaction segment is taken as the to-be-predicted information sample, the previous 4 interaction segments are taken as the historical interaction information sample.
[0161] In step 202, a third acoustic feature of the interaction context information sample and a fourth semantic feature of the interaction context information sample are extracted, and a fifth semantic feature of the to-be-predicted information sample is extracted.
[0162] In step 203, the following processing is performed by a latent expression model: fusion mapping processing based on the third acoustic feature, the fourth semantic feature, and the fifth semantic feature to obtain a second paralinguistic feature.
[0163] In step 204, the following processing is performed by a diffusion model: denoising processing of initial noise based on the fifth semantic feature and the second paralinguistic feature to obtain a fourth acoustic feature of the to-be-predicted information sample.
[0164] As an example, the implementation of steps 202 to 204 can refer to the implementation of steps 102 to 104.
[0165] In step 205, an acoustic loss is determined based on the fourth acoustic feature and a labeled acoustic feature of the to-be-predicted information sample.
[0166] In some embodiments, determining the acoustic loss based on the fourth acoustic feature and the labeled acoustic feature of the to-be-predicted information sample in step 205 can be implemented by the following technical solutions: determining a first mean absolute error and a first mean square error between the fourth acoustic feature and the labeled acoustic feature of the to-be-predicted information sample; and performing fusion processing on the first mean absolute error and the first mean square error to obtain the acoustic loss.
[0167] As an example, refer to formula (4):
[0168] wherein, is a calculation manner of the first mean absolute error, is a calculation manner of the first mean square error, is the fourth acoustic feature, x l+1 is the labeled acoustic feature of the to-be-predicted information sample. The L1 loss function is also called the mean abs error, that is, the mean absolute error, which is the absolute value of the difference between the predicted value and the true value. The L2 loss function is also called the mean square error, that is, the mean square error, which is the square of the difference between the predicted value and the true value.
[0169] Through the embodiments of the present application, the acoustic loss is constrained from two scales of L1 and L2, so that the diffusion model and the latent expression model trained based on the above acoustic loss have better feature expression capability.
[0170] In step 206, the parameters of the latent expression model and the parameters of the diffusion model are updated based on the acoustic loss to obtain an updated latent expression model and an updated diffusion model.
[0171] By the embodiment of the present application, a plurality of historical interaction information samples are obtained, and an interaction context information sample and a to-be-predicted information sample are extracted from the plurality of historical interaction information samples; a third acoustic feature of the interaction context information sample and a fourth semantic feature of the interaction context information sample are extracted, and a fifth semantic feature of the to-be-predicted information sample is extracted; the following processing is performed by a latent expression model: fusion mapping processing is performed based on the third acoustic feature, the fourth semantic feature, and the fifth semantic feature to obtain a second paralanguage feature; the second paralanguage feature herein is different from an explicit feature, and can more flexibly and variously represent paralanguage information; the following processing is performed by a diffusion model: based on the fifth semantic feature and the second paralanguage feature, initial noise is denoised to obtain a fourth acoustic feature of the to-be-predicted information sample; the final fourth acoustic feature can represent semantic features of a predicted interaction text, and can also represent paralanguage information that the predicted interaction text should have, thereby improving the accuracy of voice generation in the acoustic dimension in an interactive dialogue; based on the fourth acoustic feature and a labeled acoustic feature of the to-be-predicted information sample, an acoustic loss is determined; parameters of the latent expression model and parameters of the diffusion model are updated based on the acoustic loss to obtain an updated latent expression model and an updated diffusion model; after forward propagation, the loss function is used for backward updating, so that the latent expression model and the diffusion model can better complete the second paralanguage feature generation task and the fourth acoustic feature generation task.
[0172] In some embodiments, in step 206, based on the acoustic loss, the parameters of the latent expression model and the parameters of the diffusion model are updated to obtain an updated latent expression model and an updated diffusion model, which can be implemented by the following technical solution: obtaining a paralanguage loss and obtaining a linear modulation loss; at least one of the paralanguage loss and the linear modulation loss is fused with the acoustic loss to obtain a fusion loss; based on the fusion loss, the parameters of the latent expression model and the parameters of the diffusion model are updated to obtain an updated latent expression model and an updated diffusion model. Through the embodiment of the present application, the training effect can be improved from multiple dimensions, so that the updated latent expression model and the updated diffusion model have better model effect.
[0173] As an example, the fusion loss function is The LEM and the paralanguage word representation acoustic model (implemented as a diffusion model, for example, a first diffusion model or a second diffusion model) are jointly trained, as shown in formula (5):
[0174] wherein, and respectively represent the acoustic loss, the paralanguage loss, and the linear modulation loss; λl and λ f respectively represent and corresponding weights, for example, λ l = 1.0, λ f = 0.001.
[0175] In some embodiments, the above-mentioned obtaining the paralinguistic loss can be implemented by the following technical solutions: extracting a fifth acoustic feature of the to-be-predicted information sample; performing the following processing through an auxiliary latent expression model: performing fusion mapping processing based on the third acoustic feature, the fifth acoustic feature, the fourth semantic feature, and the fifth semantic feature to obtain a third paralinguistic feature; and determining the paralinguistic loss based on the third paralinguistic feature and the second paralinguistic feature. Through the embodiments of the present application, more stable and efficient training can be ensured.
[0176] As an example, to ensure more stable and efficient training, the embodiments of the present application provide a teacher-student network-based training framework, as shown in FIG. 8. In the MIGS training process, the model parameters of the teacher-latent expression model (LEM teacher) are obtained by calculating the exponential moving average (EMA) of the model parameters of the student-latent expression model (LEM student), and the LEM teacher does not participate in gradient backpropagation, that is, it does not calculate the gradient. In addition, the input information of the LEM teacher is complete, that is, the complete context information without causal masking processing. Therefore, the paralinguistic latent expression e predicted by the LEM teacher can be used as another target for training the LEM student, that is, the LEM student model is trained or jointly trained through the knowledge distillation of the teacher-student network.
[0177] As an example, a third acoustic feature x 1,...,l of the interaction context information sample is extracted 1,...,l and a fourth semantic feature u l+1 of the interaction context information sample is extracted l+1The auxiliary latent expression model is a LEM teacher, and the fusion mapping processing based on the third acoustic feature, the fifth acoustic feature, the fourth semantic feature, and the fifth semantic feature to obtain the third paralanguage feature can refer to step 103. The difference is that the acoustic splicing feature is obtained by splicing the third acoustic feature, the fifth acoustic feature, the fourth semantic feature, and the fifth semantic feature, and the acoustic splicing feature is processed by the auxiliary latent expression model to obtain the third paralanguage feature.
[0178] In some embodiments, the fusion mapping processing based on the third acoustic feature, the fifth acoustic feature, the fourth semantic feature, and the fifth semantic feature to obtain the third paralanguage feature can be realized by the following technical solution: performing splicing processing on the third acoustic feature, the fifth acoustic feature, the fourth semantic feature, and the fifth semantic feature to obtain a feature splicing result; performing attention processing on the input of an h-th attention network in H cascaded attention networks through the h-th attention network, and transmitting an h-th attention result output by the h-th attention network to an (h+1)-th attention network to continue attention processing to obtain an (h+1)-th attention result corresponding to the (h+1)-th attention network; determining the third paralanguage feature based on an H-th attention result output by an H-th attention network; wherein h is an integer variable with a value increasing from 1, and the value range of h is 1≤h
[0179] As an example, the auxiliary latent expression model includes H cascaded attention networks, so that H times of attention processing are performed, and each time the attention processing is performed according to the attention result obtained by the previous attention processing, and then the attention processing is performed in the next attention network.
[0180] As an example, h is an integer variable with a value increasing from 1, and the value range of h is 1≤h As an example, h is an integer variable with a value increasing from 1, and the value range of h is 1≤h
[0181] Through the embodiments of the present application, multi-modal deep fusion can be realized, for example, cross-domain feature jointing: through splicing of acoustic features (speech physical properties) and bilingual semantic features (historical content + predicted content), acoustic-semantic cross-modal alignment is realized, the problem of speech and text being split in traditional methods is solved, complementary gain between features can be realized, through multi-layer attention mechanism, progressive feature extraction can be performed, the first layer of attention focuses on the primary association of acoustic-semantic (such as accent position and keyword matching), the middle layer can mine deep temporal dependence (such as the relationship between dialogue rhythm and emotion progression), the Nth layer can generate a globally consistent expression (to ensure that paralinguistic features are self-consistent within a sentence), each layer of attention automatically adjusts the contribution weight of the three types of features, and the paralinguistic representation is enhanced, through cascaded attention to capture long-distance dependence, the position accuracy of the emphasized accent is improved, and the cascaded structure can optimize the calculation efficiency, the parameter quantity is only 72% of the parallel structure, and the memory occupation is reduced by 40% (through layer-by-layer feature dimension reduction).
[0182] In some embodiments, the determining of the paralinguistic loss based on the third paralinguistic feature and the second paralinguistic feature can be realized by the following technical solution: determining a second mean square error between the third paralinguistic feature and the second paralinguistic feature; and obtaining a paralinguistic loss positively correlated with the second mean square error.
[0183] As an example, refer to formula (6):
[0184] wherein, is the third paralinguistic feature output by the teacher-Lem, e is the second paralinguistic feature output by the student-Lem, and the student-Lem used in step 203 is the latent expression model, is a calculation method of the second mean square error.
[0185] Through the embodiments of the present application, the difference between the paralinguistic features is quantified by the second mean square error (MSE), which forces the model to maintain feature consistency, which is equivalent to realizing knowledge transfer between models, thereby improving the training effect.
[0186] In some embodiments, the obtaining of the linear modulation loss can be realized by the following technical solution: obtaining a linear parameter used for performing linear mapping in the diffusion model; and performing fusion processing on the linear parameter to obtain the linear modulation loss.
[0187] As an example, refer to formula (7):
[0188] where θ γ and θ β respectively represent the scalar parameters that are learnable by the FiLM layer, which can be seen in combination with equations (1) to (3), for the first diffusion model, the feature-wise linear modulation processing is used in the first diffusion model, and θ γ and θ β are used in the process of feature-wise linear modulation processing.
[0189] The feature-wise linear modulation (FiLM) layer is a neural network module that can be used to implement linear adjustment of features. The main function of the FiLM layer is to scale and shift the input features, and this scaling and shifting is learnable. Through the embodiments of the present application, the parameters used in the feature-wise linear modulation processing can be minimized, so as to control the first diffusion model to perform only subtle feature-wise linear modulation processing, thereby improving the overall training effect.
[0190] In the following, an exemplary application of the embodiments of the present application in an actual application scenario will be described.
[0191] In the interactive dialogue automatic generation scene, for example, in the human-computer dialogue scene, the data processing request can be generated in response to the user input dialogue content, and the terminal can be used to obtain the data processing request, where the data processing request specifically refers to a voice signal requesting to generate predicted interactive text based on historical interactive information in the interactive dialogue. The terminal sends the data processing request to the server, and the plurality of historical interactive information can be obtained directly by the terminal and carried to the data processing request, the server obtains the predicted interactive text corresponding to the historical interactive information; the first acoustic feature of the historical interactive information and the first semantic feature of the historical interactive information are extracted, and the second semantic feature of the predicted interactive text is extracted; based on the first acoustic feature, the first semantic feature and the second semantic feature, a fusion mapping processing is performed to obtain a first paralinguistic feature; based on the second semantic feature and the first paralinguistic feature, noise is removed from the initial noise to obtain the second acoustic feature of the predicted interactive text; based on the second acoustic feature, a voice signal corresponding to the predicted interactive text is generated, the server sends the voice signal to the terminal, and the voice signal is output through the terminal.
[0192] Referring to FIG. 6, a large amount of interactive content data can be obtained by mining a large number of publicly disclosed Internet audio data sets, radio film and television programs, short video short dramas, etc., but the corresponding paralanguage label data of the interactive content is very little, and the cost of obtaining paralanguage label annotation is high. The technical scheme provided by the embodiment of the present application aims to overturn the traditional method, without any manual annotation of speech, and learns the paralanguage latent expression information in a self-supervised manner. The words related to the description of "latent expression" given in FIG. 6 are only for illustrative purposes, i.e., "sarcastic tone", "fake happy" and the like are only for illustrative purposes, and do not represent that the embodiment of the present application outputs explicit expression description in the processing process. The technical scheme of the embodiment of the present application focuses on inferring the implicit paralanguage expression.
[0193] Referring to FIG. 7, the multi-turn interactive speech model (MIGS) shown in FIG. 7 generates speech in combination with the above hints of the multi-turn interactive environment. The input of the latent expression model (LEM) includes the acoustic features of the interactive dialogue context and the unit token representation from the tokenizer (Tokenizer Model); the paralanguage latent expression predicted by the LEM is input into the paralanguage token acoustic model (Para-Token Acoustic Model) in combination with the unit token representation of the next interactive segment to generate the acoustic features of the next interactive segment, and is returned to the input side as new interactive content for continuing the next round of iteration processing, while being synthesized into a speech signal by a high-fidelity generative adversarial vocoder (HiFi-GAN vocoder).
[0194] As shown in FIG. 7, the MIGS is composed of four main modules: a latent expression model (LEM) whose input includes the acoustic features of the interactive dialogue context and the unit token representation from the tokenizer (Tokenizer Model), inferring the paralanguage latent expression e of each interactive segment; a tokenizer (Tokenizer Model) for converting original audio segments or text segments into discrete unit token representations u; a paralanguage token acoustic model (Para-Token Acoustic Model) predicting the acoustic features of the next interactive segment based on the paralanguage latent expression e and the unit token representation u of the next interactive segment as a condition; and a high-fidelity generative adversarial vocoder (HiFi-GAN vocoder) synthesizing a speech signal based on the acoustic features of the next interactive segment.
[0195] More specifically, as shown in FIG. 7, the tokenizer provided by the embodiments of the present application adopts the Multitask-UnitY2 module (hereinafter referred to as UnitY) in the large language translation model Seamless to convert the voice or text into discrete unit token representations, thereby reducing the dimension and achieving higher quality voice generation. That is, the original audio or text sequence is taken as input to generate discrete unit token representations If the system needs to support multiple languages, the language discrete representation can be combined to identify the language used in the dialogue, as shown in FIG. 7, where eng represents English.
[0196] Regarding the acoustic feature x, the 80-dimensional Mel filter bank feature (hereinafter referred to as Mel feature) of the original audio can be used Where L, T s ,T τ ,T,T x are the number of interactive segment sequences, the length of the original audio segment, the length of the text sequence, the length of the frame sequence of the discrete unit token representation, and the length of the Mel feature sequence, respectively. In the UnitY module, the lengths of the discrete unit token representation and the Mel feature sequence are aligned, so in order to facilitate description, the embodiments of the present application can assume that T x = T.
[0197] Mel feature is a common acoustic representation used in speech synthesis reconstruction, as it contains both linguistic semantic information and paralinguistic information. The technical solutions provided by the embodiments of the present application use Mel feature for illustration, but this does not mean that the embodiments of the present application are limited to using Mel feature as acoustic representation. It should be noted that the technical solutions provided by the embodiments of the present application can use any acoustic representation that contains both linguistic semantic information and paralinguistic information, for example, the quantized hidden vector calculated by the residual vector quantizer in the neural network audio encoder, which is another acoustic representation used by the embodiments of the present application.
[0198] The discrete unit token representation u mainly contains linguistic semantic information, but lacks paralinguistic information related to expression. Therefore, the embodiments of the present application propose a latent expression model LEM to predict the paralinguistic latent expression between interactive segments D is the vector space dimension of the paralinguistic latent expression.
[0199] In the training phase, the input of LEM includes the Mel feature x 1,...,l of the interactive context 1,...,l and the discrete unit token representation u l+1, the LEM outputs a paralinguistic latent expression e (containing paralinguistic information). The LEM output is combined with the paralinguistic information e related to the expression and the discrete unit token representation u l+1 In combination, referred to as a paralinguistic token representation, the combination of paralinguistic information and unit token representation representing semantics is input to a paralinguistic token representation acoustic model as an input condition to predict Mel features of the next interaction segment The entire training process is based on predicting Mel features of the next interaction segment, without any manually annotated speech, and learns paralinguistic latent expression e from massive Internet broadcast and movie program interaction content data in a self-supervised manner.
[0200] To ensure more stable and efficient training, an embodiment of the present application provides a teacher-student network-based training framework, as shown in FIG. 8. In the MIGS training process, the model parameters of the teacher-latent expression model (LEM teacher) are obtained by calculating the exponential moving average (Exponential Moving Average, EMA) of the model parameters of the student-latent expression model (LEM student), and the LEM teacher does not participate in gradient backpropagation, that is, it does not calculate the gradient. In addition, the input information of the LEM teacher is complete, that is, the complete context information without causal masking processing. The so-called causal masking processing is to mask the Mel features of the current predicted interaction segment and the subsequent interaction segment during training, and to mask the unit token representation of the current predicted interaction segment and the subsequent interaction segment, so as to ensure that the input information of the LEM student is consistent in the training stage and the inference stage. Therefore, the paralinguistic latent expression e predicted by the LEM teacher can be used as another target for training the LEM student, that is, the LEM student model is trained or jointly trained through knowledge distillation of the teacher-student network.
[0201] The MIGS architecture combined with LLM is introduced below, as shown in FIG. 9. In the MIGS training process, the LLM combined with the context prompt engineering infers expression information based on historical interaction text. Here, the current predicted interaction text refers to "what to say", and the expression information refers to "how to say". After the unit token representation and meta unit token are extracted by the UnitY2 module described in the foregoing, they are input to the subsequent LEM model. In the MIGS interactive inference process, the LLM as an intelligent agent not only infers the text content of the current predicted interaction text, but also infers the expression information of the current predicted interaction text. The paralinguistic latent expression is predicted by combining the LEM, and the speech signal is generated again and returned to the input side as new interaction content for the next round of processing.
[0202] As shown in FIG. 9, in the MIGS training process, the prior training data is used as the current interactive text, and the context prompt engineering is combined to obtain the expression information by LLM inference. For example, the context prompt can be designed as "please analyze the response of the person, and describe how he expresses in combination with the context of the interactive dialogue, including but not limited to expression style, emotion, tone, speed, lightness, etc.", in the example given in FIG. 9, the interactive text of the current dialogue role A is "I am very sorry", and the corresponding expression information is "sorry and guilty". Next, the expression information obtained by LLM inference, such as "sorry and guilty", is input into the subsequent paralinguistic token acoustic model in the form of text through the UnitY2 module to extract meta unit token.
[0203] In the interactive inference process of MIGS, LLM as an agent not only infers the text content of the current predicted interactive text, but also infers the expression information of the interactive text. These text information is extracted by the UnitY2 module to obtain unit tokens and meta unit tokens, combined with the paralinguistic latent expression predicted by LLM, and the Mel feature is generated by the paralinguistic token acoustic model, (the Mel feature is synthesized into a speech signal by a vocoder), and is returned to the input side as new interactive content for the next round of processing. Similarly, based on the MIGS basic architecture given in FIG. 8, a MIGS architecture combining LLM as an agent can also be extended.
[0204] The specific implementation of each module in MIGS is introduced below.
[0205] The specific implementation process of the tokenizer (Tokenizer Model) can refer to the UNITY2 module architecture in the large multi-language translation generation model Seamless to realize the application and training process, and the pre-trained multi-language self-supervised learning large model is used to obtain the UnitY2 module (the embodiments of the present application also completely support using any module that converts from speech or text modal to discrete representation as an alternative solution). The vocoder (Vocoder) can adopt the HiFi-GAN vocoder architecture in Seamless to realize the application and training process, which will not be described again in the embodiments of the present application (the embodiments of the present application also completely do not exclude using any vocoder module that converts from acoustic representation to synthesized speech as an alternative solution).
[0206] The specific implementation of LEM is introduced below. Referring to FIG. 10, in order to balance the model capability and the calculation complexity, the Spectrum-Interaction-transformer (SI-transformer) modified based on the Spatial-Temporal-transformer (ST-transformer) is adopted in the specific architecture implementation of LEM. As shown in FIG. 10, the SI-transformer includes N Spectrum-Interaction blocks, each of which is composed of a Spectrum attention layer and an Interaction attention layer, and then connected with a Feed-Forward (FFW) layer. The Spectrum attention layer performs self-attention calculation on each interaction segment based on the spectrum dimension, and the Interaction attention layer performs self-attention calculation on the vector sequence between multiple interaction segments.
[0207] The SI-transformer is adopted to implement LEM, and the causal masking processing is adopted, that is, the Mel features of the current predicted interaction segment and the subsequent interaction segments are masked during training, and the unit word representation corresponding to the D u embedding features of the subsequent interaction segments are masked during training. Then, the Mel features x 1,...,L and the unit word representation u 1,...,L of the total L interaction segments are spliced to obtain the Mel spliced features n After the causal masking processing, the Mel spliced features are used as the input in the training stage.
[0208] The Spectrum attention layer in the SI-transformer is different from the attention module of the classic transformer. The Spectrum attention layer in the SI-transformer only performs self-attention calculation on the input of each interaction segment, that is, self-attention calculation is performed on the Mel spliced features with the dimension of 1×T×(80+D u ). Since there are L Mel spliced features, L times of calculation are required. Then, the Interaction attention layer only performs self-attention calculation on the input corresponding to each spectrum tile, that is, self-attention calculation is performed on the sequence with the dimension of L×1×1. Since there are T×(80+D u ) sequences, T×(80+D u ) times of calculation are required.
[0209] The SI-transformer outputs a feature f o mapped to the paralinguistic latent representation where D is the vector space dimension of the paralinguistic latent representation.
[0210] The computational complexity of the SI-transformer is mainly in the Spectrum attention layer, which is linearly related to the number of interactions L, while the computational complexity of the classic Transformer is related to the square of L, so the SI-transformer can efficiently support content generation tasks with multiple rounds of interaction.
[0211] The specific implementation of the paralinguistic word representation acoustic model is introduced below.
[0212] Corresponding to the MIGS architecture shown in FIGS. 7 and 8, the paralinguistic word representation acoustic model supports generating a paralinguistic latent representation e and a unit word representation u l+1 As a condition, the Mel feature of the target segment is predicted
[0213] More specifically, referring to FIG. 11, the paralinguistic word representation acoustic model adopts a WaveNet architecture to build a diffusion model, which includes 40 WaveNet layers in total, each layer inputting a 1-dimensional dilated convolution layer (Dilated Conv), each layer having a kernel size of 3, a filter size of 1024, and an expansion size of 2. The input of the Dilated Conv includes the input of the WaveNet layer and the embedding feature corresponding to the timestamp t of the current diffusion stage. The output of the Dilated Conv is added to the paralinguistic latent representation e and the unit word representation u l+1 The corresponding embedding feature is obtained to obtain the hidden layer representation h.
[0214] The output of the WaveNet can be input to a post-processing network (PostNet) for post-processing, which can help to compensate for residual signals that are difficult to capture in the diffusion stage.
[0215] Corresponding to the MIGS architecture combined with the LLM in FIG. 9, another specific implementation of the paralinguistic word representation acoustic model is given in the embodiments of the present application, which can support generating a paralinguistic latent representation e and a unit word representation u l+1 As a fine-grained (segment-level) condition, in particular, it also supports fusing the meta-language unit word representation predicted from the LLM As a coarse-grained (interaction-level) condition.
[0216] Here, the diffusion model is also built using the WaveNet architecture, containing a total of 40 WaveNet layers. The difference is that in the specific implementation of the paralinguistic word representation acoustic model shown in Figure 12, every two WaveNet layers, a Feature-wise Linear Modulation (FiLM) layer is added to the hidden representation h of the third WaveNet layer. This layer is used to fuse the metalinguistic unit word representations predicted by the LLM agent. The corresponding embedding feature c. The implementation of the FiLM layer can be found in formulas (8)-(10): FiLM(h,c)=(γ+1)·h+β (8); γ=f1(c)·θ γ (9); β=f2(c)·θ β , (10);
[0217] Where f1 and f2 represent linear mapping functions, respectively; γ and β represent the scaling and bias parameters of the FiLM layer, respectively; θ γ and θ β These represent their corresponding learnable scalar parameters.
[0218] The embodiments of this application can input the output of WaveNet into a PostNet for post-processing, thereby helping to compensate for residual signals that are difficult to capture during the diffusion stage.
[0219] The embodiments of this application may also support the use of other types of conditional generation models as alternatives.
[0220] The loss function involved in training MIGS as provided in the embodiments of this application is described below.
[0221] By comprehensive loss function Jointly train the LEM and the paralinguistic word representation acoustic model, see formula (11):
[0222] in, and Let λ represent the Mel feature loss, the knowledge distillation loss of the teacher-student network, and the L2 regularization loss of the feature linear modulation, respectively; l and λ f They represent and The corresponding weights, for example, λ l =1.0, λ f =0.001. The formulas for each loss are defined as follows:
[0223] wherein, and correspond to L1 loss function and L2 loss function, respectively; and represent predicted Mel features and actual Mel features, respectively, γ and represent scalar parameters learned by FiLM layers, respectively; β e correspond to paralinguistic latent representations predicted by LEM student and LEM teacher, respectively.
[0224] The output of WaveNet in the embodiments of the present application can be input to PostNet for post-processing output. Then See formula (15):
[0225] wherein, and correspond to L1 loss function and L2 loss function, respectively; and represent predicted Mel features and actual Mel features, respectively, represent Mel features output by PostNet for post-processing.
[0226] The beneficial effects of the technical solutions provided by the embodiments of the present application mainly include the following aspects:
[0227] 1. Scalable to large-scale data: without any manual annotation of speech, a large amount of public Internet audio, radio film and television programs, short videos, short dramas and other massive interactive content data are mined to solve the problems of few available paralinguistic label data, high paralinguistic label annotation cost and difficult to guarantee quality, and implicit expression can be learned in a self-supervised manner;
[0228] 2. Achieve diversity and zero-shot learning: combined with the large-scale training data of the previous item, the diversity and zero-shot learning performance of the system is relatively improved compared with the prior art;
[0229] 3. Robustness to data and algorithms: unlike existing speech generation large models, it does not rely on pre-defined speech expression-related explicit expression features, nor does it rely on any special pre-processing algorithm to calculate these pre-defined prior features, for example, it does not need to extract pitch / fundamental frequency, classify voiced / unvoiced, estimate intensity, etc. Avoid relying on algorithms with high requirements for speech quality, and take advantage of the scalability of large-scale data to increase the robustness of speech data with background sound and noise;
[0230] 4. Fine-grained instruction automatic generation and control: The technical solution provided in the embodiments of the present application introduces a reasoning process that automatically predicts paralinguistic potential expression e, especially fine-grained potential expression, without the need for people to explicitly input text instructions for control. At the same time, although not within the scope of the description of the technical solution provided in the embodiments of the present application, it is easy to see that the technical solution provided in the embodiments of the present application has the flexibility to expand to explicit control. If it is necessary to increase user controllable functions, assuming that the instruction set and small-scale labeled data are fine-tuned, an additional module that maps explicit instructions to paralinguistic potential expression can be added during training, that is, the function of explicit control can be achieved.
[0231] 5. Open generation: The target application scenario of the embodiments of the present application is more inclined to open generation tasks, that is, there is no standard answer as a reference. For example, unlike in translation tasks, it is necessary to rely on the reference source voice to measure the goodness of the generated target voice. At the same time, the embodiments of the present application have higher requirements for new measurement standards such as the matching degree of the generated expression and the context semantics, liveliness, improvisational openness, and richness.
[0232] It can be understood that in the embodiments of the present application, data related to user information is involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of related data need to comply with relevant laws, regulations, and standards in relevant countries and regions.
[0233] The following continues to illustrate an exemplary structure of the implementation of the data processing apparatus 455-1 provided in the embodiments of the present application as a software module. In some embodiments, as shown in FIG. 4, the software module stored in the data processing apparatus 455-1 in the memory 450 can include: a first acquisition module configured to acquire historical interaction information and predicted interaction text corresponding to the historical interaction information; a first extraction module configured to extract first acoustic features of the historical interaction information and first semantic features of the historical interaction information, and extract second semantic features of the predicted interaction text; a first paralanguage module configured to perform fusion mapping processing on the first acoustic features, the first semantic features, and the second semantic features to obtain first paralanguage features; a first denoising module configured to perform denoising processing on initial noise based on the second semantic features and the first paralanguage features to obtain second acoustic features of the predicted interaction text; and a prediction module configured to generate a voice signal corresponding to the predicted interaction text based on the second acoustic features.
[0234] In some embodiments, the first extraction module is further configured to: perform attitude prediction processing on the predicted interaction text of the historical interaction information by a language model to obtain response attitude text of the predicted interaction text; and extract third semantic features of the response attitude text. The first denoising module is further configured to: perform denoising processing on the initial noise based on the second semantic features, the third semantic features, and the first paralinguistic features to obtain second acoustic features of the predicted interaction text.
[0235] In some embodiments, the denoising processing is implemented by a first diffusion model, and the first diffusion model includes P cascaded multi-semantic denoising networks, where P is not less than 2. The first denoising module is further configured to: perform multi-semantic denoising processing on an input of a pth multi-semantic denoising network in the P cascaded multi-semantic denoising networks by the pth multi-semantic denoising network, and transmit a pth multi-semantic denoising result output by the pth multi-semantic denoising network to a (p+1)th multi-semantic denoising network to continue multi-semantic denoising processing to obtain a Pth multi-semantic denoising result corresponding to the (P+1)th multi-semantic denoising network; and determine the second acoustic features of the predicted interaction text based on a Pth multi-semantic denoising result output by the Pth multi-semantic denoising network, where p is an integer variable that starts from 1 and increases by 1, and the value range of P is 1≤p<P, when p takes the value of 1, the input of the pth multi-semantic denoising network is the initial noise, the second semantic features, the third semantic features, and the first paralinguistic features, and when p takes the value of 2≤p<P, the input of the pth multi-semantic denoising network is a (p-1)th multi-semantic denoising result output by the (p-1)th multi-semantic denoising network, the second semantic features, the third semantic features, and the first paralinguistic features.
[0236] In some embodiments, the first denoising module is further configured to: perform fusion processing on the input of the pth single semantic denoising network and a cascade identifier corresponding to the pth single semantic denoising network to obtain a first fusion result; perform convolution processing on the first fusion result to obtain a first convolution result; perform fusion processing on the first convolution result, the second semantic features, and the first paralinguistic features to obtain a second fusion result; perform fusion processing on the second fusion result and the third semantic features to obtain a third fusion result; perform activation processing based on a gating mechanism on the third fusion result to obtain a first activation result; and perform fusion processing on the first activation result and the input of the pth single semantic denoising network to obtain an output of the pth single semantic denoising network.
[0237] In some embodiments, the first denoising module is further configured to: perform embedding processing on the third semantic feature to obtain an embedded feature of the third semantic feature; perform first linear mapping processing on the embedded feature of the third semantic feature to obtain a first linear mapping result, and perform second linear mapping processing on the embedded feature of the third semantic feature to obtain a second linear mapping result; perform mapping processing on the second fusion result by taking the first linear mapping result as a scaling parameter and the second linear mapping result as a bias parameter to obtain the third fusion result.
[0238] In some embodiments, the fusion mapping processing is implemented through a latent expression model, and the latent expression model includes N cascaded attention networks, where N is not less than 2; the first auxiliary language module is configured to: perform concatenation processing on the first acoustic feature, the first semantic feature, and the second semantic feature to obtain an acoustic concatenated feature; perform attention processing on an input of an nth attention network in the N cascaded attention networks through the nth attention network, and transmit an nth attention result output by the nth attention network to an (n+1)th attention network to continue attention processing to obtain an (n+1)th attention result corresponding to the (n+1)th attention network; and determine the first auxiliary language feature based on an Nth attention result output by an Nth attention network; where n is an integer variable that starts from 1 and increases, and the value range of n is 1≤n<N, when n is 1, the input of the nth attention network is the acoustic concatenated feature, and when n is 2≤n<N, the input of the nth attention network is an (n-1)th attention result output by an (n-1)th attention network.
[0239] In some embodiments, the first auxiliary language module is configured to: perform self-attention processing on the acoustic concatenated feature corresponding to each interaction information to obtain a first self-attention result of the acoustic concatenated feature, where the interaction information is the historical interaction information or the predicted interaction text; perform self-attention processing on the first self-attention result of the acoustic concatenated feature corresponding to each feature element to obtain a second self-attention result of the acoustic concatenated feature; and perform full connection processing on the second self-attention result of the acoustic concatenated feature to obtain the nth attention result output by the nth attention network.
[0240] In some embodiments, the denoising processing is implemented through a second diffusion model, the second diffusion model comprising M cascaded single semantic denoising networks, M being no less than 2; the first denoising module is further configured to: through an mth single semantic denoising network in the M cascaded single semantic denoising networks, perform single semantic denoising processing on an input of the mth single semantic denoising network, and transmit an mth single semantic denoising result output by the mth single semantic denoising network to an (m+1)th single semantic denoising network to continue the single semantic denoising processing, to obtain an (m+1)th single semantic denoising result corresponding to the (m+1)th single semantic denoising network; and determine the second acoustic feature of the predicted interaction text based on an Mth single semantic denoising result output by an Mth single semantic denoising network; wherein m is an integer variable starting from 1 and increasing, and the value range of M is 1≤m<M, when m is 1, the input of the mth single semantic denoising network is the initial noise, the second semantic feature and the first paralinguistic feature, and when m is 2≤m<M, the input of the mth single semantic denoising network is an (m-1)th single semantic denoising result output by an (m-1)th single semantic denoising network, the second semantic feature and the first paralinguistic feature.
[0241] In some embodiments, the first denoising module is further configured to: fuse the input of the mth single semantic denoising network and a cascade identifier corresponding to the mth single semantic denoising network to obtain a fourth fusion result; perform convolution processing on the fourth fusion result to obtain a second convolution result; fuse the second convolution result, the second semantic feature and the first paralinguistic feature to obtain a fifth fusion result; perform activation processing based on a gating mechanism on the fifth fusion result to obtain a second activation result; and fuse the second activation result and the input of the mth single semantic denoising network to obtain the output of the mth single semantic denoising network.
[0242] The following continues to illustrate an example structure of the data processing apparatus 455-2 provided by the embodiments of the present application implemented as a software module. In some embodiments, as shown in FIG. 4, the software module stored in the data processing apparatus 455-2 of the memory 450 can include: a second acquisition module configured to acquire a plurality of historical interaction information samples, and extract an interaction context information sample and a to-be-predicted information sample from the plurality of historical interaction information samples; a second extraction module configured to extract a third acoustic feature of the interaction context information sample and a fourth semantic feature of the interaction context information sample, and extract a fifth semantic feature of the to-be-predicted information sample; a second paralanguage module configured to perform the following processing by using a latent expression model: performing fusion mapping processing based on the third acoustic feature, the fourth semantic feature, and the fifth semantic feature to obtain a second paralanguage feature; a second denoising module configured to perform the following processing by using a diffusion model: performing denoising processing on an initial noise based on the fifth semantic feature and the second paralanguage feature to obtain a fourth acoustic feature of the to-be-predicted information sample; a loss module configured to determine an acoustic loss based on the fourth acoustic feature and a labeled acoustic feature of the to-be-predicted information sample; and a training module configured to update parameters of the latent expression model and parameters of the diffusion model based on the acoustic loss to obtain an updated latent expression model and an updated diffusion model.
[0243] In some embodiments, the loss module is configured to: determine a first mean absolute error and a first mean square error between the fourth acoustic feature and the labeled acoustic feature of the to-be-predicted information sample; and perform fusion processing on the first mean absolute error and the first mean square error to obtain the acoustic loss.
[0244] In some embodiments, the training module is configured to: acquire a paralanguage loss and acquire a linear modulation loss; perform fusion processing on at least one of the paralanguage loss and the linear modulation loss and the acoustic loss to obtain a fusion loss; and update the parameters of the latent expression model and the parameters of the diffusion model based on the fusion loss to obtain an updated latent expression model and an updated diffusion model.
[0245] In some embodiments, the training module is configured to: extract a fifth acoustic feature of the to-be-predicted information sample; perform the following processing by using an auxiliary latent expression model: performing fusion mapping processing based on the third acoustic feature, the fifth acoustic feature, the fourth semantic feature, and the fifth semantic feature to obtain a third paralanguage feature; and determining the paralanguage loss based on the third paralanguage feature and the second paralanguage feature.
[0246] In some embodiments, the training module is configured to: determine a second average square error between the third sub-language feature and the second sub-language feature; and obtain a sub-language loss positively correlated with the second average square error.
[0247] In some embodiments, the training module is configured to: obtain a linear parameter in the diffusion model for performing linear mapping; and perform fusion processing on the linear parameter to obtain the linear modulation loss.
[0248] Embodiments of the present application provide a computer program product, which includes computer executable instructions stored in a computer readable storage medium. A processor of an electronic device reads the computer executable instructions from the computer readable storage medium, and the processor executes the computer executable instructions to cause the electronic device to perform the data processing method provided by the embodiments of the present application.
[0249] Embodiments of the present application provide a computer readable storage medium storing computer executable instructions, wherein the computer executable instructions, when executed by a processor, cause the processor to perform the data processing method provided by the embodiments of the present application, for example, the data processing method shown in FIG. 5A.
[0250] In some embodiments, the computer readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM, etc.; or various devices including one or any combination of the above memories.
[0251] In some embodiments, the computer executable instructions can be in the form of programs, software, software modules, scripts or codes, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and can be deployed in any form, including being deployed as independent programs or being deployed as modules, components, subroutines or other units suitable for use in a computing environment.
[0252] As an example, the computer executable instructions can but not necessarily correspond to files in a file system, can be stored in part of a file storing other programs or data, for example, stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program in question, or stored in multiple cooperating files (for example, files storing one or more modules, subroutines or code portions).
[0253] By way of example, computer-executable instructions can be deployed to be executed on one electronic device or on multiple electronic devices at one location or distributed across multiple locations and executed on multiple electronic devices networked together.
[0254] To sum up, the historical interaction information and the predicted interaction text corresponding to the historical interaction information are acquired by the embodiments of the present application; the first acoustic feature of the historical interaction information and the first semantic feature of the historical interaction information are extracted, and the second semantic feature of the predicted interaction text is extracted. Here, the semantic and acoustic features of the historical interaction information can be acquired, and the semantic feature of the predicted interaction text can also be acquired. The first paralanguage feature is obtained by fusion mapping processing based on the first acoustic feature, the first semantic feature and the second semantic feature. The first paralanguage feature is different from the explicit feature, and can more flexibly and diversely represent the paralanguage information. The second acoustic feature of the predicted interaction text is obtained by denoising processing of the initial noise based on the second semantic feature and the first paralanguage feature. Thus, the second acoustic feature can carry more accurate paralanguage information. The voice signal corresponding to the predicted interaction text is generated based on the second acoustic feature. The finally generated voice signal can represent the semantic feature of the predicted interaction text, and can also represent the paralanguage information that the predicted interaction text should have. The accuracy of the voice generation in the acoustic dimension in the interactive dialogue is improved.
[0255] The above merely describes the embodiments of the present application, but is not used to limit the protection scope of the present application. Any modification, equivalent replacement and improvement within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A data processing method, the method being performed by an electronic device, the method comprising: obtaining historical interaction information and predicted interaction text corresponding to the historical interaction information; extracting first acoustic features of the historical interaction information and first semantic features of the historical interaction information, and extracting second semantic features of the predicted interaction text; performing fusion mapping processing on the first acoustic features, the first semantic features, and the second semantic features to obtain first paralinguistic features; performing denoising processing on initial noise based on the second semantic features and the first paralinguistic features to obtain second acoustic features of the predicted interaction text; generating a speech signal corresponding to the predicted interaction text based on the second acoustic features.
2. The method of claim 1, wherein, The method further comprises: performing attitude prediction processing on the predicted interaction text of the historical interaction information by a language model to obtain response attitude text of the predicted interaction text; extracting third semantic features of the response attitude text; The denoising processing on the initial noise based on the second semantic features and the first paralinguistic features to obtain the second acoustic features of the predicted interaction text comprises: performing denoising processing on the initial noise based on the second semantic features, the third semantic features, and the first paralinguistic features to obtain the second acoustic features of the predicted interaction text.
3. The method of claim 2, wherein, The denoising processing is realized by a first diffusion model, and the first diffusion model comprises P cascaded multi-semantic denoising networks, P being not less than 2; The denoising processing on the initial noise based on the second semantic features, the third semantic features, and the first paralinguistic features to obtain the second acoustic features of the predicted interaction text comprises: performing multi-semantic denoising processing on an input of a pth multi-semantic denoising network in the P cascaded multi-semantic denoising networks, and transmitting a pth multi-semantic denoising result output by the pth multi-semantic denoising network to an (p+1)th multi-semantic denoising network to continue multi-semantic denoising processing to obtain an (P+1)th multi-semantic denoising result corresponding to the (p+1)th multi-semantic denoising network; determining the second acoustic features of the predicted interaction text based on an Pth multi-semantic denoising result output by an Pth multi-semantic denoising network; wherein p is an integer variable starting from 1 and increasing, and the value range of P is 1≤p 4. The method of claim 3, wherein, The multi-semantic denoising processing on the input of the pth multi-semantic denoising network in the P cascaded multi-semantic denoising networks comprises: The input of the pth multi-semantics denoising network is fused with the cascade identifier corresponding to the pth multi-semantics denoising network to obtain a first fusion result; The first fusion result is subjected to convolution processing to obtain a first convolution result; The first convolution result, the second semantic feature, and the first auxiliary language feature are fused to obtain a second fusion result; The second fusion result and the third semantic feature are fused to obtain a third fusion result; The third fusion result is subjected to activation processing based on a gating mechanism to obtain a first activation result; The first activation result and the input of the pth multi-semantics denoising network are fused to obtain an output of the pth multi-semantics denoising network.
5. The method of claim 4, wherein, The second fusion result and the third semantic feature are fused to obtain a third fusion result, including: The third semantic feature is subjected to embedding processing to obtain an embedded feature of the third semantic feature; The embedded feature of the third semantic feature is subjected to first linear mapping processing to obtain a first linear mapping result, and is subjected to second linear mapping processing to obtain a second linear mapping result; The second fusion result is subjected to mapping processing with the first linear mapping result as a scaling parameter and the second linear mapping result as a bias parameter to obtain the third fusion result.
6. The method according to any one of claims 1 to 5, wherein, The fusion mapping processing is implemented through a latent expression model, and the latent expression model includes N cascaded attention networks, and N is not less than 2; The first acoustic feature, the first semantic feature, and the second semantic feature are fused to obtain a first auxiliary language feature, including: The first acoustic feature, the first semantic feature, and the second semantic feature are spliced to obtain acoustic spliced features; An nth attention network in the N cascaded attention networks performs attention processing on an input of the nth attention network, and an nth attention result output by the nth attention network is transmitted to an (n+1)th attention network to continue attention processing to obtain an (n+1)th attention result corresponding to the (n+1)th attention network; The first auxiliary language feature is determined based on an Nth attention result output by an Nth attention network; wherein n is an integer variable that increases from 1, and the value range of n is 1≤n<N, when n takes the value 1, the input of the nth attention network is the acoustic spliced feature, and when n takes the value 2≤n<N, the input of the nth attention network is an (n-1)th attention result output by an (n-1)th attention network.
7. The method of claim 6, wherein, The input of the nth attention network in the N cascaded attention networks is subjected to self-attention processing corresponding to each interaction information to obtain a first self-attention result of the nth attention network, wherein the interaction information is the historical interaction information or the predicted interaction text; performing self-attention processing on the first self-attention result of the nth attention network corresponding to each feature element to obtain a second self-attention result of the nth attention network; performing full connection processing on the second self-attention result of the nth attention network to obtain an nth attention result output by the nth attention network.
8. The method according to any one of claims 1 to 7, wherein, The denoising processing is implemented through a second diffusion model, and the second diffusion model includes M cascaded single semantic denoising networks, where M is not less than 2. The denoising processing on the initial noise based on the second semantic feature and the first paralinguistic feature to obtain the second acoustic feature of the predicted interactive text includes: performing single semantic denoising processing on the input of the mth single semantic denoising network in the M cascaded single semantic denoising networks, and transmitting an mth single semantic denoising result output by the mth single semantic denoising network to an (m+1)th single semantic denoising network to continue the single semantic denoising processing, to obtain an (m+1)th single semantic denoising result corresponding to the (m+1)th single semantic denoising network; determining the second acoustic feature of the predicted interactive text based on an Mth single semantic denoising result output by an Mth single semantic denoising network. wherein m is an integer variable starting from 1 and increasing, and the value range of M is 1≤m<M, when m is 1, the input of the mth single semantic denoising network is the initial noise, the second semantic feature and the first paralinguistic feature, and when m is 2≤m<M, the input of the mth single semantic denoising network is an (m-1)th single semantic denoising result output by an (m-1)th single semantic denoising network, the second semantic feature and the first paralinguistic feature.
9. The method of claim 8, wherein, The single semantic denoising processing on the input of the mth single semantic denoising network in the M cascaded single semantic denoising networks includes: performing fusion processing on the input of the mth single semantic denoising network and a cascade identifier corresponding to the mth single semantic denoising network to obtain a fourth fusion result; performing convolution processing on the fourth fusion result to obtain a second convolution result; performing fusion processing on the second convolution result, the second semantic feature and the first paralinguistic feature to obtain a fifth fusion result; performing activation processing based on a gating mechanism on the fifth fusion result to obtain a second activation result; performing fusion processing on the second activation result and the input of the mth single semantic denoising network to obtain the output of the mth single semantic denoising network.
10. A data processing method, the method being performed by an electronic device, and the method comprising: obtaining a plurality of historical interaction information samples, and extracting an interaction context information sample and a to-be-predicted information sample from the plurality of historical interaction information samples; extracting a third acoustic feature of the interaction context information sample and a fourth semantic feature of the interaction context information sample, and extracting a fifth semantic feature of the to-be-predicted information sample; performing fusion mapping processing based on the third acoustic feature, the fourth semantic feature and the fifth semantic feature through a latent expression model to obtain a second paralinguistic feature; The following processing is performed by the diffusion model: based on the fifth semantic feature and the second paralinguistic feature, denoising processing is performed on the initial noise to obtain a fourth acoustic feature of the to-be-predicted information sample; based on the fourth acoustic feature and the labeled acoustic feature of the to-be-predicted information sample, an acoustic loss is determined; based on the acoustic loss, the parameters of the latent representation model and the parameters of the diffusion model are updated to obtain an updated latent representation model and an updated diffusion model.
11. The method of claim 10, wherein, The acoustic loss is determined based on the fourth acoustic feature and the labeled acoustic feature of the to-be-predicted information sample, comprising: determining a first mean absolute error and a first mean square error between the fourth acoustic feature and the labeled acoustic feature of the to-be-predicted information sample; the first mean absolute error and the first mean square error are fused to obtain the acoustic loss.
12. The method of claim 10 or 11, wherein, The parameters of the latent representation model and the parameters of the diffusion model are updated based on the acoustic loss to obtain an updated latent representation model and an updated diffusion model, comprising: obtain a paralinguistic loss and obtain a linear modulation loss; fuse at least one of the paralinguistic loss and the linear modulation loss with the acoustic loss to obtain a fusion loss; based on the fusion loss, the parameters of the latent representation model and the parameters of the diffusion model are updated to obtain an updated latent representation model and an updated diffusion model.
13. The method of claim 12, wherein, The paralinguistic loss is obtained, comprising: extracting a fifth acoustic feature of the to-be-predicted information sample; The following processing is performed by the auxiliary latent representation model: based on the third acoustic feature, the fifth acoustic feature, the fourth semantic feature and the fifth semantic feature, fusion mapping processing is performed to obtain a third paralinguistic feature; based on the third paralinguistic feature and the second paralinguistic feature, the paralinguistic loss is determined.
14. The method of claim 13, wherein, The paralinguistic loss is determined based on the third paralinguistic feature and the second paralinguistic feature, comprising: determining a second mean square error between the third paralinguistic feature and the second paralinguistic feature; obtaining a paralinguistic loss positively correlated with the second mean square error.
15. The method according to any one of claims 12 to 14, wherein, The linear modulation loss is obtained, comprising: obtaining a linear parameter in the diffusion model for performing linear mapping; fuse the linear parameter to obtain the linear modulation loss.
16. The method of claim 13, wherein, The third paralinguistic feature is obtained by performing fusion mapping processing on the third acoustic feature, the fifth acoustic feature, the fourth semantic feature and the fifth semantic feature, comprising: splicing the third acoustic feature, the fifth acoustic feature, the fourth semantic feature and the fifth semantic feature to obtain a feature splicing result; by the hth attention network in the H cascaded attention networks, the input of the hth attention network is processed by attention, and the hth attention result output by the hth attention network is transmitted to the h+1th attention network for further attention processing to obtain the h+1th attention result corresponding to the h+1th attention network; determine the third paralinguistic feature based on an Hth attention result output by an Hth attention network; wherein h is an integer variable starting from 1 and increasing by 1, and the value of h ranges from 1 to H, when h is 1, the input of the hth attention network is the feature concatenation result, and when h is 2 to H, the input of the hth attention network is an (h-1)th attention result output by an (h-1)th attention network.
17. A data processing apparatus, comprising: a first acquisition module configured to acquire historical interaction information and predicted interaction text corresponding to the historical interaction information; a first extraction module configured to extract first acoustic features of the historical interaction information and first semantic features of the historical interaction information, and extract second semantic features of the predicted interaction text; a first paralinguistic module configured to perform fusion mapping processing on the first acoustic features, the first semantic features, and the second semantic features to obtain first paralinguistic features; a first denoising module configured to perform denoising processing on initial noise based on the second semantic features and the first paralinguistic features to obtain second acoustic features of the predicted interaction text; a prediction module configured to generate a speech signal corresponding to the predicted interaction text based on the second acoustic features.
18. A data processing apparatus, comprising: a second acquisition module configured to acquire a plurality of historical interaction information samples, and extract an interaction context information sample and a to-be-predicted information sample from the plurality of historical interaction information samples; a second extraction module configured to extract third acoustic features of the interaction context information sample and fourth semantic features of the interaction context information sample, and extract fifth semantic features of the to-be-predicted information sample; a second paralinguistic module configured to perform the following processing by a latent expression model: performing fusion mapping processing based on the third acoustic features, the fourth semantic features, and the fifth semantic features to obtain second paralinguistic features; a second denoising module configured to perform the following processing by a diffusion model: performing denoising processing on initial noise based on the fifth semantic features and the second paralinguistic features to obtain fourth acoustic features of the to-be-predicted information sample; a loss module configured to determine an acoustic loss based on the fourth acoustic features and labeled acoustic features of the to-be-predicted information sample; a training module configured to update parameters of the latent expression model and parameters of the diffusion model based on the acoustic loss to obtain an updated latent expression model and an updated diffusion model.
19. An electronic device, comprising: a memory configured to store computer executable instructions; a processor configured to execute the computer executable instructions stored in the memory to implement the data processing method in any one of claims 1 to 9 or 10 to 16.
20. A computer readable storage medium storing computer executable instructions, the computer executable instructions being executed by a processor to implement the data processing method in any one of claims 1 to 9 or 10 to 16.
21. A computer program product comprising computer executable instructions to implement the data processing method of any one of claims 1 to 9 or 10 to 16 when executed by a processor.
Citation Information
Patent Citations
Data processing method and device for man-machine interaction conversation, medium and electronic equipment
CN113761156A
Customer service verbal skill guiding method and device, equipment and storage medium
CN114220461A
Human-computer interaction method, device and system
CN116088788A
Voice interaction system and method, electronic equipment and storage medium
CN116417003A
Exploiting acoustic and lexical properties of phonemes to recognize valence from speech
US20200298873A1
Cited By
Risk dynamic processing method and device in real-time communication scene, equipment and medium
CN121151505A
Self-adaptive denoising and intelligent optimization processing system and method for geophysical exploration data
CN121388400A
Language disfluency detection method based on multi-mode and window attention mechanism
CN121393426A
Semantic understanding method and engine based on multi-model collaborative reasoning and self-supervised learning
CN121659960A
Underwater scene multi-modal generation method based on diffusion model
CN122176113A